Skip to content

perf: reuse one Bedrock client across side-channel calls - #1375

Merged
philmerrell merged 1 commit into
developfrom
feature/side-channel-bedrock-client-reuse
Sep 28, 2026
Merged

philmerrell merged 1 commit into
developfrom
feature/side-channel-bedrock-client-reuse

Conversation

@philmerrell

@philmerrell philmerrell commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Four side-channel Bedrock callers built boto3.client("bedrock-runtime") on every call:

  • apis/shared/tool_summaries/summarizer.py: once per tool batch, concurrent with the agent stream
  • agents/main_agent/session/compaction_summary.py: compression and extraction calls
  • apis/shared/files/document_digest.py: document abstract
  • apis/shared/embeddings/bedrock_embeddings.py: every knowledge-base search and ingestion batch

Each build ran synchronously on the event loop, before asyncio.to_thread / run_in_executor, and each per-call client got a fresh connection pool, so every call paid a new TCP+TLS handshake.

The cost of a build depends on the process. The first bedrock-runtime client in a process takes about 250ms locally (botocore parses the service model); later builds reuse the parsed model and cost a few milliseconds. The inference-api Runtime already pays that first parse at startup (apis/inference_api/warmup.py, logged as warmup step=boto:bedrock-runtime ms=6), so in the Runtime this PR saves about 6ms of event-loop work per call, not 250ms. The 250ms applies to app-api, which has no warmup: its first document digest per task used to build the client on the event loop, holding up every other request on that task.

This applies the fix used for the conversation-title path in #1374: build one client lazily under a lock, and fetch it inside the worker thread so the first build stays off the event loop.

  • Summarizer, compaction, digest use the existing apis.shared.aws_clients.get_client, which is process-cached, locked, and keyed by (service, region). Keying by region matters for compaction, whose functions take a region argument.
  • Embeddings has its own lazily built, locked module-level client. The kb-sync and rag-ingestion Lambda images copy apis/shared/embeddings/ but not aws_clients.py, so importing it there would fail at runtime in those Lambdas, and tests would not catch it. Embeddings sit on the knowledge-base search path, ahead of the model's first token.

The Converse/InvokeModel payloads are unchanged, so there is no prompt or cache-prefix impact. The effect on time to first token is small and only in the good direction: the knowledge-base search that runs before the first token no longer builds a client on the event loop.

Measured impact in dev

Validated on dev after merge; every path exercised worked with no new errors. The speed-up is below what dev can measure:

  • KB search embeddings: 133, 125 and 125ms after deploy, against a median of about 137ms over 35 searches in the prior two weeks. The old timings exclude the client build (it ran before the log line that starts the window), so the old cost was slightly higher than 137ms, but three samples are within noise. Titan's own latency dominates.
  • Document digest: 460–600ms per abstract, dominated by Nova. Only one earlier sample exists, so there's no baseline.
  • Not exercised in dev: compaction (threshold 100,000 tokens) and the rag-ingestion Lambda's embeddings step (new knowledge bases default to the managed engine, which skips it).

The change is kept for event-loop hygiene (app-api's first digest; the pre-first-token search path) and to put these callers on the repo's shared aws_clients pattern, not for a measured latency win.

Known gap

The first builds are serialised by separate locks (aws_clients, embeddings, and the title client on #1374). Two different first builds can still reach boto3's default session at the same moment. That window is once per process, the same exposure the title fix already accepts.

Test plan

  • Autouse fixtures reset the cached client per test in test_summarizer.py, test_compaction_summary.py, test_document_digest.py and test_side_channel_inference_config.py. The helpers that install a fake boto3 reset it too, since one digest test swaps fakes mid-test.
  • New tests check that repeated calls build one client: summarizer, digest, compaction (including per-region keying) and embeddings (tests/shared/test_bedrock_embeddings_client.py). All four fail without the source change.
  • Full backend suite under xdist: 10,942 passed, 3 skipped
  • tests/shared tests/agents tests/apis tests/lambdas tests/architecture run serially (to catch cached-client leaks between tests): 7,882 passed
  • ruff check clean

🤖 Generated with Claude Code

The tool-batch summarizer, compaction summary, document abstract and
embeddings each built boto3.client("bedrock-runtime") per call: ~250ms
the first time in a process (botocore loads the service model), run on
the event loop before asyncio.to_thread, and a fresh connection pool —
a new TCP+TLS handshake — on every call.

The summarizer, compaction and digest now use the process-cached
apis.shared.aws_clients.get_client (locked, keyed by region), fetched
inside the worker thread so a first build stays off the loop.
bedrock_embeddings keeps its own lazily built, locked module client:
the kb-sync and rag-ingestion Lambda images copy that package without
aws_clients.py. Embeddings sit on the knowledge-base search path, ahead
of the model's first token.

Tests reset the cached client per test (and per fake install, where a
test swaps fakes mid-test), and pin one client build across calls.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@philmerrell
philmerrell merged commit bb70ec7 into develop Sep 28, 2026
7 checks passed
@philmerrell
philmerrell deleted the feature/side-channel-bedrock-client-reuse branch September 28, 2026 03:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant