Skip to content

Managed KB: reconciler over-counts a synced change that was refused on the byte cap #1372

Description

@philmerrell

Summary

When a KB sync stages a changed source whose growth would exceed the owner's byte cap, the ingestion consumer correctly refuses it. The previous version stays indexed and served, and the ledger is left alone. But the new, larger bytes are already in S3, because the sync overwrote the object before the consumer ran. The reconciler's storedBytes anchor sizes counted documents from S3, so its next pass counts the refused version's size instead of the size actually committed. The owner is charged for growth that was never indexed.

This was introduced by the refuse-before-submit behaviour in #1365 and is listed there under "Known edges".

How it happens

  1. A managed document is complete with committedBytes = C (its S3 object is C bytes).
  2. A KB sync finds the source changed and overwrites the S3 object with C + g bytes (kb_sync/worker.py → _stage_to_s3).
  3. The consumer (kb_migration/ingestion_consumer.py → _claim_reingest) tries to reserve g, gets ByteCapExceeded, and calls _record_reingest_over_cap. Nothing is submitted, committedBytes stays C, and the sync gates are cleared so the next sync retries.
  4. reconciler.stored_bytes_from_s3 counts the document (it has committedBytes and no byteCapRefunded) but sums its S3 size, C + g. So storedBytes/totalBytes go up by g on the next pass.

Impact

  • Direction: it over-counts, which the byte-cap design treats as the safe direction. The owner is under-permitted, never over-permitted.
  • Owner-visible effect: the phantom g eats into the remaining allowance. The retry on the next sync then needs room for g twice (once phantom, once real), so a user who frees exactly g bytes still can't get the change through.
  • Duration: until the document's next successful re-ingest, or its deletion. On delete, refund_once returns committedBytes = C, and the reconciler stops counting the object once it is gone.
  • Scope: only documents with a sync policy, on managed knowledge bases, whose source grew past the cap. Rare, but it gets worse the closer an owner is to their cap.

Possible fixes

  • Anchor on the ledger, not S3, for documents with a pending refused version. For example, the reconciler uses committedBytes when stagedContentHash != ingestedContentHash (the version in S3 was never ingested). This is the smallest change and keeps S3 as the source of truth otherwise.
  • Anchor on committedBytes everywhere. This is simpler, but drops the S3 measurement the reconciler exists to provide, so it probably isn't the right call.
  • Stage elsewhere. The sync could stage changed bytes to a side key and only promote to the document's key after the cap check. This is a larger change: the ingestion trigger and the legacy pipeline both key on the document's S3 path.

Tests

The moto fixtures in backend/tests/lambdas/test_kb_sync_managed_reingest.py already reproduce the refused-change state (TestTheByteCapGatesTheNewVersion). A new test could run the reconciler's refresh_stored_bytes / stored_bytes_from_s3 after the refusal and assert storedBytes equals the committed size.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions