Skip to content

Unchanged observations are re-embedded on every full reindex, and deferred oversized-entity shards never resume #1605

Description

@sammywachtel

Summary

A full reindex of a project whose notes have not changed re-embeds every observation. On a note with more than 256 observations, some observations are never embedded at all, because the deferred shards are recorded but nothing runs them.

Both come from the same fact: observation rows get new ids every time a note is republished, and the vector chunk key is built from the row id.

All references are to main at 22d31e96.

What happens

  1. ObservationRepository replaces observations by deleting and re-inserting them (repository/observation_repository.py:145, delete_by_fields then insert). Each republish gives every observation a new row id.
  2. The semantic chunk key includes that row id: chunk_key = f"{row.type}:{row.id}:{chunk_index}" (repository/semantic_chunking.py:96).
  3. So after any republish, every observation chunk has a key the vector table has never seen. The source_hash would match, but the key lookup misses first, and the chunk is embedded again.
  4. project add always runs a full index (cli/commands/command_utils.py:142, run_project_index(project, force_full=True, ...)). A container that runs project add at startup therefore republishes every note, and re-embeds every observation, on every boot, even when nothing changed.

The oversized-entity shard never resumes

OVERSIZED_ENTITY_VECTOR_SHARD_SIZE = 256 (repository/semantic_vector_sync.py:28) splits a large entity into bounded work units. That is deliberate and a good idea (#723, which describes it as sharding oversized entity vector sync "into bounded resumable work units"). entity.vector_sync_deferred_at records that the entity has unfinished shards (search_repository_base.py:427).

As far as I can find, only readiness reads that column (services/project_readiness.py:308). Nothing schedules the next shard. The remaining chunks are embedded only if something else triggers another sync of that entity, and when that happens the observation ids have been re-minted, so the sync starts again from shard one.

Reproduction

Local, default embedding settings:

  • A project with 60 observations across a few notes, no edits between runs. Every bm reindex (full) re-embeds exactly 60 observation chunks.
  • One note with 300 observations. Three full passes: each reports entities_deferred=1, and after all three, 44 of the 300 observations have no vector.

On a real deployment that runs project add at container start, the number of chunks re-embedded on each restart matched the number of observations in each project (6 of 6, 43 of 43, and so on), and a project with about 800 observations was busy re-embedding for roughly a quarter of an hour after each restart.

Suggested directions

These are independent; any one of them helps.

(A) Key observation chunks on something stable

The observation's identity for embedding purposes is its entity plus what it says, not its row id. Duplicate text inside one note is legal and meaningful (the comment just after observation_repository.py:145 explains why), so add an occurrence index to keep twins distinct.

# semantic_chunking.py, when building chunks for an observation row
def observation_identity(row, occurrence: int) -> str:
    digest = sha256(f"{row.category}\x00{row.content}".encode()).hexdigest()[:16]
    return f"{row.entity_id}:{digest}:{occurrence}"

# occurrence = how many earlier observations in this entity have the same
# (category, content); computed once per entity while iterating rows in order.
if row.type == "observation":
    chunk_key = f"observation:{observation_identity(row, occurrence)}:{chunk_index}"
else:
    chunk_key = f"{row.type}:{row.id}:{chunk_index}"

With that, an unchanged observation keeps its key across a republish, the existing source_hash check sees the same text, and the embedding is skipped. The row-id replacement semantics in ObservationRepository can stay as they are; the temporal projection depends on them (indexing/relation_persistence.py:254).

(B) Schedule the continuation for deferred entities

# after a vector sync pass finishes, or on a periodic tick
deferred = await session.execute(
    select(Entity.id)
    .where(Entity.project_id == project_id)
    .where(Entity.vector_sync_deferred_at.is_not(None))
    .order_by(Entity.vector_sync_deferred_at)
    .limit(batch_size)
)
for entity_id in deferred.scalars():
    await enqueue_vector_sync(entity_id)   # resumes at the next unscheduled shard

Without (A) this still re-embeds from the start after any republish, but it does finish, which fixes the "44 of 300 never embedded" case.

(C) Don't force a full index when project add finds a known project

# cli/commands/command_utils.py
already_indexed = project_row is not None and project_row.last_indexed_at is not None
await run_project_index(project, force_full=not already_indexed, run_in_background=False)

A new project still gets its full first index. An already-registered, already-indexed project gets the normal incremental pass, which leaves unchanged notes alone.

Environment

  • main at 22d31e96
  • Postgres backend

Happy to test a patch against the reproduction above.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions