perf: reuse one Bedrock client across side-channel calls - #1375
Merged
Merged
Conversation
The tool-batch summarizer, compaction summary, document abstract and
embeddings each built boto3.client("bedrock-runtime") per call: ~250ms
the first time in a process (botocore loads the service model), run on
the event loop before asyncio.to_thread, and a fresh connection pool —
a new TCP+TLS handshake — on every call.
The summarizer, compaction and digest now use the process-cached
apis.shared.aws_clients.get_client (locked, keyed by region), fetched
inside the worker thread so a first build stays off the loop.
bedrock_embeddings keeps its own lazily built, locked module client:
the kb-sync and rag-ingestion Lambda images copy that package without
aws_clients.py. Embeddings sit on the knowledge-base search path, ahead
of the model's first token.
Tests reset the cached client per test (and per fake install, where a
test swaps fakes mid-test), and pin one client build across calls.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Four side-channel Bedrock callers built
boto3.client("bedrock-runtime")on every call:apis/shared/tool_summaries/summarizer.py: once per tool batch, concurrent with the agent streamagents/main_agent/session/compaction_summary.py: compression and extraction callsapis/shared/files/document_digest.py: document abstractapis/shared/embeddings/bedrock_embeddings.py: every knowledge-base search and ingestion batchEach build ran synchronously on the event loop, before
asyncio.to_thread/run_in_executor, and each per-call client got a fresh connection pool, so every call paid a new TCP+TLS handshake.The cost of a build depends on the process. The first
bedrock-runtimeclient in a process takes about 250ms locally (botocore parses the service model); later builds reuse the parsed model and cost a few milliseconds. The inference-api Runtime already pays that first parse at startup (apis/inference_api/warmup.py, logged aswarmup step=boto:bedrock-runtime ms=6), so in the Runtime this PR saves about 6ms of event-loop work per call, not 250ms. The 250ms applies to app-api, which has no warmup: its first document digest per task used to build the client on the event loop, holding up every other request on that task.This applies the fix used for the conversation-title path in #1374: build one client lazily under a lock, and fetch it inside the worker thread so the first build stays off the event loop.
apis.shared.aws_clients.get_client, which is process-cached, locked, and keyed by(service, region). Keying by region matters for compaction, whose functions take aregionargument.apis/shared/embeddings/but notaws_clients.py, so importing it there would fail at runtime in those Lambdas, and tests would not catch it. Embeddings sit on the knowledge-base search path, ahead of the model's first token.The Converse/InvokeModel payloads are unchanged, so there is no prompt or cache-prefix impact. The effect on time to first token is small and only in the good direction: the knowledge-base search that runs before the first token no longer builds a client on the event loop.
Measured impact in dev
Validated on dev after merge; every path exercised worked with no new errors. The speed-up is below what dev can measure:
The change is kept for event-loop hygiene (app-api's first digest; the pre-first-token search path) and to put these callers on the repo's shared
aws_clientspattern, not for a measured latency win.Known gap
The first builds are serialised by separate locks (
aws_clients, embeddings, and the title client on #1374). Two different first builds can still reach boto3's default session at the same moment. That window is once per process, the same exposure the title fix already accepts.Test plan
test_summarizer.py,test_compaction_summary.py,test_document_digest.pyandtest_side_channel_inference_config.py. The helpers that install a fakeboto3reset it too, since one digest test swaps fakes mid-test.tests/shared/test_bedrock_embeddings_client.py). All four fail without the source change.tests/shared tests/agents tests/apis tests/lambdas tests/architecturerun serially (to catch cached-client leaks between tests): 7,882 passedruff checkclean🤖 Generated with Claude Code