Summary
A full reindex of a project whose notes have not changed re-embeds every observation. On a note with more than 256 observations, some observations are never embedded at all, because the deferred shards are recorded but nothing runs them.
Both come from the same fact: observation rows get new ids every time a note is republished, and the vector chunk key is built from the row id.
All references are to main at 22d31e96.
What happens
ObservationRepository replaces observations by deleting and re-inserting them (repository/observation_repository.py:145, delete_by_fields then insert). Each republish gives every observation a new row id.
- The semantic chunk key includes that row id:
chunk_key = f"{row.type}:{row.id}:{chunk_index}" (repository/semantic_chunking.py:96).
- So after any republish, every observation chunk has a key the vector table has never seen. The
source_hash would match, but the key lookup misses first, and the chunk is embedded again.
project add always runs a full index (cli/commands/command_utils.py:142, run_project_index(project, force_full=True, ...)). A container that runs project add at startup therefore republishes every note, and re-embeds every observation, on every boot, even when nothing changed.
The oversized-entity shard never resumes
OVERSIZED_ENTITY_VECTOR_SHARD_SIZE = 256 (repository/semantic_vector_sync.py:28) splits a large entity into bounded work units. That is deliberate and a good idea (#723, which describes it as sharding oversized entity vector sync "into bounded resumable work units"). entity.vector_sync_deferred_at records that the entity has unfinished shards (search_repository_base.py:427).
As far as I can find, only readiness reads that column (services/project_readiness.py:308). Nothing schedules the next shard. The remaining chunks are embedded only if something else triggers another sync of that entity, and when that happens the observation ids have been re-minted, so the sync starts again from shard one.
Reproduction
Local, default embedding settings:
- A project with 60 observations across a few notes, no edits between runs. Every
bm reindex (full) re-embeds exactly 60 observation chunks.
- One note with 300 observations. Three full passes: each reports
entities_deferred=1, and after all three, 44 of the 300 observations have no vector.
On a real deployment that runs project add at container start, the number of chunks re-embedded on each restart matched the number of observations in each project (6 of 6, 43 of 43, and so on), and a project with about 800 observations was busy re-embedding for roughly a quarter of an hour after each restart.
Suggested directions
These are independent; any one of them helps.
(A) Key observation chunks on something stable
The observation's identity for embedding purposes is its entity plus what it says, not its row id. Duplicate text inside one note is legal and meaningful (the comment just after observation_repository.py:145 explains why), so add an occurrence index to keep twins distinct.
# semantic_chunking.py, when building chunks for an observation row
def observation_identity(row, occurrence: int) -> str:
digest = sha256(f"{row.category}\x00{row.content}".encode()).hexdigest()[:16]
return f"{row.entity_id}:{digest}:{occurrence}"
# occurrence = how many earlier observations in this entity have the same
# (category, content); computed once per entity while iterating rows in order.
if row.type == "observation":
chunk_key = f"observation:{observation_identity(row, occurrence)}:{chunk_index}"
else:
chunk_key = f"{row.type}:{row.id}:{chunk_index}"
With that, an unchanged observation keeps its key across a republish, the existing source_hash check sees the same text, and the embedding is skipped. The row-id replacement semantics in ObservationRepository can stay as they are; the temporal projection depends on them (indexing/relation_persistence.py:254).
(B) Schedule the continuation for deferred entities
# after a vector sync pass finishes, or on a periodic tick
deferred = await session.execute(
select(Entity.id)
.where(Entity.project_id == project_id)
.where(Entity.vector_sync_deferred_at.is_not(None))
.order_by(Entity.vector_sync_deferred_at)
.limit(batch_size)
)
for entity_id in deferred.scalars():
await enqueue_vector_sync(entity_id) # resumes at the next unscheduled shard
Without (A) this still re-embeds from the start after any republish, but it does finish, which fixes the "44 of 300 never embedded" case.
(C) Don't force a full index when project add finds a known project
# cli/commands/command_utils.py
already_indexed = project_row is not None and project_row.last_indexed_at is not None
await run_project_index(project, force_full=not already_indexed, run_in_background=False)
A new project still gets its full first index. An already-registered, already-indexed project gets the normal incremental pass, which leaves unchanged notes alone.
Environment
main at 22d31e96
- Postgres backend
Happy to test a patch against the reproduction above.
Summary
A full reindex of a project whose notes have not changed re-embeds every observation. On a note with more than 256 observations, some observations are never embedded at all, because the deferred shards are recorded but nothing runs them.
Both come from the same fact: observation rows get new ids every time a note is republished, and the vector chunk key is built from the row id.
All references are to
mainat22d31e96.What happens
ObservationRepositoryreplaces observations by deleting and re-inserting them (repository/observation_repository.py:145,delete_by_fieldsthen insert). Each republish gives every observation a new row id.chunk_key = f"{row.type}:{row.id}:{chunk_index}"(repository/semantic_chunking.py:96).source_hashwould match, but the key lookup misses first, and the chunk is embedded again.project addalways runs a full index (cli/commands/command_utils.py:142,run_project_index(project, force_full=True, ...)). A container that runsproject addat startup therefore republishes every note, and re-embeds every observation, on every boot, even when nothing changed.The oversized-entity shard never resumes
OVERSIZED_ENTITY_VECTOR_SHARD_SIZE = 256(repository/semantic_vector_sync.py:28) splits a large entity into bounded work units. That is deliberate and a good idea (#723, which describes it as sharding oversized entity vector sync "into bounded resumable work units").entity.vector_sync_deferred_atrecords that the entity has unfinished shards (search_repository_base.py:427).As far as I can find, only readiness reads that column (
services/project_readiness.py:308). Nothing schedules the next shard. The remaining chunks are embedded only if something else triggers another sync of that entity, and when that happens the observation ids have been re-minted, so the sync starts again from shard one.Reproduction
Local, default embedding settings:
bm reindex(full) re-embeds exactly 60 observation chunks.entities_deferred=1, and after all three, 44 of the 300 observations have no vector.On a real deployment that runs
project addat container start, the number of chunks re-embedded on each restart matched the number of observations in each project (6 of 6, 43 of 43, and so on), and a project with about 800 observations was busy re-embedding for roughly a quarter of an hour after each restart.Suggested directions
These are independent; any one of them helps.
(A) Key observation chunks on something stable
The observation's identity for embedding purposes is its entity plus what it says, not its row id. Duplicate text inside one note is legal and meaningful (the comment just after
observation_repository.py:145explains why), so add an occurrence index to keep twins distinct.With that, an unchanged observation keeps its key across a republish, the existing
source_hashcheck sees the same text, and the embedding is skipped. The row-id replacement semantics inObservationRepositorycan stay as they are; the temporal projection depends on them (indexing/relation_persistence.py:254).(B) Schedule the continuation for deferred entities
Without (A) this still re-embeds from the start after any republish, but it does finish, which fixes the "44 of 300 never embedded" case.
(C) Don't force a full index when
project addfinds a known projectA new project still gets its full first index. An already-registered, already-indexed project gets the normal incremental pass, which leaves unchanged notes alone.
Environment
mainat22d31e96Happy to test a patch against the reproduction above.