fix(kb): re-ingest a synced source that changed on a managed knowledge base - #1365
Merged
Merged
Conversation
…e base A KB sync that finds its Drive file (or crawled page) changed overwrites the document's S3 object in place and relies on the ObjectCreated event to re-ingest. On a managed knowledge base that event reaches the ingestion consumer, which found a complete row with byteCapSettled and returned "already-settled" — the early exit that makes redelivery idempotent. The new bytes were never ingested, and the size change never reached the byte ledger. The sync now writes stagedContentHash before staging. The consumer re-ingests while it differs from the ingestedContentHash the last re-ingest recorded: it reserves the size growth against the cap before submitting (a file that grew past the cap is refused and its previous version keeps serving; the next sync retries), submits once per version, waits out Bedrock statuses that predate the submission, and on completion re-stamps committedBytes and commits or refunds the difference. A per-version claim on the DOC# row keeps all of it idempotent under EventBridge redelivery; a delete mid-flight returns the growth reservation. Legacy knowledge bases are routed away before any of this, as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
When a KB sync finds that a Drive file (or a crawled page) has changed, it writes the new bytes over the document's existing S3 object. It then relies on the
ObjectCreatedevent to re-ingest them. The legacy pipeline handles that. On a managed knowledge base, though, the event goes to the ingestion consumer. By then theDOC#row is alreadycompletewithbyteCapSettled, which looks exactly like a redelivery, so the consumer returnedalready-settled. Even without that early exit, Bedrock reports the documentINDEXED(the old version), and the consumer would have declined to re-submit.As a result:
storedBytes/committedBytes, so a delete refunded the old sizeThe new test
test_the_new_bytes_reach_the_knowledge_basereproduced this ondevelop: the consumer returnedalready-settledand Bedrock was never called.Fix
The sync marks the version it staged. Both sync paths (Drive and web re-crawl) write
stagedContentHashbefore the S3 overwrite. Nothing else writes it, so a document that was never synced behaves exactly as before.The consumer tells a change from a redelivery. It re-ingests a
completeorfailedrow wheneverstagedContentHashdiffers from theingestedContentHashrecorded by the last re-ingest. Otherwise the settled early exit applies as it does now. The re-ingest (_reingest_changed_document) runs in four steps:committedBytesis reserved against the cap, and that happens before anything is submitted. If the growth would exceed the cap, the change is refused, Bedrock never receives it, and the previous version keeps serving. The sync gates (sourceEtag,contentHash) are cleared so the next sync run stages it again and retries. The claim (byte_cap.claim_reingest) is conditional per version hash. A redelivery therefore reserves nothing, and a newer change replaces an unfinished claim and returns its reservation.reingestSubmittedAtmeans a redelivery waits on the ingestion already running instead of restarting it.wait_until_indexed(not_before=…)waits it out.DOC#write sets statuscomplete, re-stampscommittedBytes, recordsingestedContentHashand removes the claim. The ledger then moves by the difference (byte_cap.settle_reingest): growth is committed from the reservation and shrinkage is refunded.The status stays
completethroughout, so the document keeps answering from its previous version until the new one replaces it. Afaileddocument now recovers when its source changes. A delete during a re-ingest returns the growth reservation (settle_bytes_on_delete→release_reingest_once). The consumer's completion write is refused on adeletingrow, so exactly one side settles.Legacy (S3 Vectors) is unaffected. The consumer sends legacy documents to their own pipeline before it reads the row. The only change on that path is the extra
stagedContentHashattribute, which the legacy pipeline ignores.previousChunkCountshrinkage cleanup is unchanged, and a test pins this.Interaction with #1361
This branch is based on
developand does not depend onbyte_cap.reserved_at_request. The re-ingest path never reads_declared_bytes, because a sync reserves nothing at request time and the growth is reserved at ingestion. Its merge withfeature/kb-byte-counter-settle-unexplainedis conflict-free. The new tests, plus the consumer, lifecycle and sync-worker suites, pass on the merged tree (138 passed).Known edges (not changed here)
contentHashbefore_stage_to_s3, so a failed S3 put is never re-staged by gate 2. This is pre-existing and affects both engines.Tests
tests/lambdas/test_kb_sync_managed_reingest.py(16 tests). It runs the real sync worker (Drive and the vault stubbed at their seams), which stages to a real moto S3 object, and the real consumer handles the resulting event against a fake Bedrock that models per-document status. It covers growth and shrinkage deltas, a delete refunding the new size, redelivery after completion and while indexing (no re-submit, no re-reserve), stale pre-submit status, a newer change superseding an unfinished one, cap refusal plus retry on the next sync, a delete mid-re-ingest, Bedrock failing the new version, a failed document recovering, the unsynced-redelivery early exit, and legacy routing.not_beforeguard, the delete-path release, or the per-version claim check each fails a named test.test_kb_sync_worker.pynow asserts the marker on changed Drive files and pages, and its absence on same-bytes runs.🤖 Generated with Claude Code