Skip to content

fix(kb): re-ingest a synced source that changed on a managed knowledge base - #1365

Merged
philmerrell merged 1 commit into
developfrom
fix/kb-sync-managed-reingest
Sep 27, 2026
Merged

philmerrell merged 1 commit into
developfrom
fix/kb-sync-managed-reingest

Conversation

@philmerrell

Copy link
Copy Markdown
Contributor

Problem

When a KB sync finds that a Drive file (or a crawled page) has changed, it writes the new bytes over the document's existing S3 object. It then relies on the ObjectCreated event to re-ingest them. The legacy pipeline handles that. On a managed knowledge base, though, the event goes to the ingestion consumer. By then the DOC# row is already complete with byteCapSettled, which looks exactly like a redelivery, so the consumer returned already-settled. Even without that early exit, Bedrock reports the document INDEXED (the old version), and the consumer would have declined to re-submit.

As a result:

  • the knowledge base kept serving the version the file had at import
  • the size change never reached storedBytes/committedBytes, so a delete refunded the old size

The new test test_the_new_bytes_reach_the_knowledge_base reproduced this on develop: the consumer returned already-settled and Bedrock was never called.

Fix

The sync marks the version it staged. Both sync paths (Drive and web re-crawl) write stagedContentHash before the S3 overwrite. Nothing else writes it, so a document that was never synced behaves exactly as before.

The consumer tells a change from a redelivery. It re-ingests a complete or failed row whenever stagedContentHash differs from the ingestedContentHash recorded by the last re-ingest. Otherwise the settled early exit applies as it does now. The re-ingest (_reingest_changed_document) runs in four steps:

  1. Reserve the growth, then claim. It measures the object with an S3 HEAD. Only growth beyond committedBytes is reserved against the cap, and that happens before anything is submitted. If the growth would exceed the cap, the change is refused, Bedrock never receives it, and the previous version keeps serving. The sync gates (sourceEtag, contentHash) are cleared so the next sync run stages it again and retries. The claim (byte_cap.claim_reingest) is conditional per version hash. A redelivery therefore reserves nothing, and a newer change replaces an unfinished claim and returns its reservation.
  2. Submit once. reingestSubmittedAt means a redelivery waits on the ingestion already running instead of restarting it.
  3. Wait for the new version. Bedrock already holds a document under this id, so a status last updated before the submission belongs to the old version. wait_until_indexed(not_before=…) waits it out.
  4. Complete. One conditional DOC# write sets status complete, re-stamps committedBytes, records ingestedContentHash and removes the claim. The ledger then moves by the difference (byte_cap.settle_reingest): growth is committed from the reservation and shrinkage is refunded.

The status stays complete throughout, so the document keeps answering from its previous version until the new one replaces it. A failed document now recovers when its source changes. A delete during a re-ingest returns the growth reservation (settle_bytes_on_delete → release_reingest_once). The consumer's completion write is refused on a deleting row, so exactly one side settles.

Legacy (S3 Vectors) is unaffected. The consumer sends legacy documents to their own pipeline before it reads the row. The only change on that path is the extra stagedContentHash attribute, which the legacy pipeline ignores. previousChunkCount shrinkage cleanup is unchanged, and a test pins this.

Interaction with #1361

This branch is based on develop and does not depend on byte_cap.reserved_at_request. The re-ingest path never reads _declared_bytes, because a sync reserves nothing at request time and the growth is reserved at ingestion. Its merge with feature/kb-byte-counter-settle-unexplained is conflict-free. The new tests, plus the consumer, lifecycle and sync-worker suites, pass on the merged tree (138 passed).

Known edges (not changed here)

  • Reconciler over-counts a refused change. When a change is refused on the cap, S3 already holds the larger bytes. The reconciler sizes counted documents from S3, so it over-counts by the growth until the next change. That errs in the safe direction (the documented one).
  • Stage timing. The sync writes contentHash before _stage_to_s3, so a failed S3 put is never re-staged by gate 2. This is pre-existing and affects both engines.

Tests

  • New file tests/lambdas/test_kb_sync_managed_reingest.py (16 tests). It runs the real sync worker (Drive and the vault stubbed at their seams), which stages to a real moto S3 object, and the real consumer handles the resulting event against a fake Bedrock that models per-document status. It covers growth and shrinkage deltas, a delete refunding the new size, redelivery after completion and while indexing (no re-submit, no re-reserve), stale pre-submit status, a newer change superseding an unfinished one, cap refusal plus retry on the next sync, a delete mid-re-ingest, Bedrock failing the new version, a failed document recovering, the unsynced-redelivery early exit, and legacy routing.
  • Mutation-checked: removing the re-ingest branch, the not_before guard, the delete-path release, or the per-version claim check each fails a named test.
  • test_kb_sync_worker.py now asserts the marker on changed Drive files and pages, and its absence on same-bytes runs.
  • Every test file referencing the byte cap, consumer, kb-sync or document service passed: 709 passed.

🤖 Generated with Claude Code

…e base

A KB sync that finds its Drive file (or crawled page) changed overwrites the
document's S3 object in place and relies on the ObjectCreated event to
re-ingest. On a managed knowledge base that event reaches the ingestion
consumer, which found a complete row with byteCapSettled and returned
"already-settled" — the early exit that makes redelivery idempotent. The new
bytes were never ingested, and the size change never reached the byte ledger.

The sync now writes stagedContentHash before staging. The consumer re-ingests
while it differs from the ingestedContentHash the last re-ingest recorded:
it reserves the size growth against the cap before submitting (a file that
grew past the cap is refused and its previous version keeps serving; the next
sync retries), submits once per version, waits out Bedrock statuses that
predate the submission, and on completion re-stamps committedBytes and
commits or refunds the difference. A per-version claim on the DOC# row keeps
all of it idempotent under EventBridge redelivery; a delete mid-flight
returns the growth reservation. Legacy knowledge bases are routed away before
any of this, as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@philmerrell
philmerrell merged commit a4e7e19 into develop Sep 27, 2026
7 checks passed
@philmerrell
philmerrell deleted the fix/kb-sync-managed-reingest branch September 27, 2026 17:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant