Skip to content

docs+refactor: map the turn path to first token and extract the route's phases - #1388

Merged
philmerrell merged 3 commits into
developfrom
claude/gracious-hawking-9969be
Sep 30, 2026
Merged

philmerrell merged 3 commits into
developfrom
claude/gracious-hawking-9969be

Conversation

@philmerrell

Copy link
Copy Markdown
Contributor

Why

A turn is the most complicated process in the app and nobody had one map of it. Three specs each covered one slice (the preamble, the deferred build, the V2 migration), and the handler in inference_api/chat/routes.py was ~2,200 lines in one function with two generators closing over ~40 locals. PR #1377 proposes the next TTFT change on that path, and deciding it needs the whole picture.

What changed

Commit 1 — docs/specs/turn-path-ttft.md. Every stage from the SPA's send to the first token, in four tables (before the handler, inside it, inside the stream, inside the agent build): what runs, what is network, what blocks the event loop, which turn_prelude mark times it, and the measured cost where one exists. Then:

  • Findings. A fresh boto3.Session() re-parses service models every time (~150ms per client, laptop), and the AgentCore Memory SDK, MemoryClient, strategy discovery and Strands' BedrockModel all build fresh sessions — so a cold build parses the same models four to seven times and warmup.py parses into a session none of them use. StreamCoordinator._get_initial_message_count runs one ListEvents over the whole history on the head of every turn just for len(), in the prepared→thinking gap no clock covers. session_mgr and tools inside the build are independent and serial. Agent turns carry two extra META reads and three writes on the critical path. agent_build.tools (2039ms cold) is still undecomposed.
  • PR perf: A/B the first-turn agent build (shared boto3 session built at warm-up), default off #1377 assessment. Merge the instrumentation and the shared-clients arm, but build the shared session, clients and strategy ids at warm-up (created lazily by the first build, the first turn still pays two parses and the strategy-id round trip). The off-loop arm does not reduce TTFT by itself; its real value is overlapping the KB search with the build, which the A/B does not measure.
  • Plan, measure-first and V2-leaning: P0 run the A/B → P1 split tools and extend the clock to the first token → P2 one shared boto session warmed at start → P3 overlap session_mgr ‖ tools in the constructor and start the KB search ahead of the build → P4 bookkeeping off the path → P5 the V2 track ending in prewarm-on-intent.
  • Regression plan for hand-off (§6): ten invariants with the tests that pin them, a per-change checklist with acceptance thresholds, the dev recipes that settle a timing claim, and an end-to-end matrix.

Pointers added in CLAUDE.md and at the top of the preamble spec.

Commit 2 — the route refactor. Behaviour-preserving. The handler's contiguous phases become named module-level functions with typed results, in the order the map lists them: _resolve_turn_attachments, _prepare_session_state, _check_turn_quota, _resolve_turn_model, _resolve_effective_tools, _build_turn_tools, _agent_for_app_dispatch (the two MCP App dispatches shared 60 identical lines), _validate_resume_interrupts, _build_citations, and _refuse_turn / _forbidden_turn for the six conversational early exits. The handler drops from ~2,240 to ~1,640 lines and its docstring is now the phase index keyed to the turn_prelude marks. Four comments that still described the pre-PR-2 preamble (including the disproved "~53ms per GSI query" claim) are corrected.

What a reviewer should know

  • Nothing here changes behaviour, the prompt, toolConfig, restored history or any mark. Every module-level seam the tests patch (is_quota_enforcement_enabled, get_quota_checker, ensure_session_metadata_exists, _load_user_settings, get_agent, …) is kept in routes.py for that reason.
  • One test pin moved: test_get_agent_call_sites.py expects the main get_agent call's memory_binding to read turn_tools.memory_binding_key. The property it protects (tools and key element derive from the same project_memory) still holds inside _build_turn_tools.
  • The extraction was done by line range with an anchor assertion at every boundary, then reviewed by hand; the generators (stream_with_quota_warning, _guarded_stream) are untouched apart from reading attachment fields off the new TurnAttachments, to keep the merge with perf: A/B the first-turn agent build (shared boto3 session built at warm-up), default off #1377 (which edits _guarded_stream) small.
  • The spec's findings are labelled Measured / Read / Hypothesis deliberately; none of the numbers in the plan's "expected" columns are claimed until the recipe in §6C produces them.

Tests

Full backend suite from the worktree with the main venv: 10950 passed, 3 skipped, 0 failed. Route-area subset (inference_api, route-level invocations, MCP App dispatch, base-agent registration): 488 passed.

🤖 Generated with Claude Code

philmerrell and others added 2 commits September 29, 2026 10:20
Adds docs/specs/turn-path-ttft.md: every stage of a turn (before the
handler, inside it, inside the stream, inside the agent build) with what
runs where, what is network, what blocks the event loop and which
turn_prelude mark times it; the findings that map exposes (fresh boto3
sessions re-parse service models so warm-up never reaches the build, a
full-history ListEvents on the head of every turn just to count it, the
serial session_mgr/tools build, agent-turn writes on the critical path,
the still-undecomposed agent_build.tools); an assessment of PR #1377; an
ordered, measure-first plan that leans toward Runtime V2; and a
regression plan a cheaper model can execute. Pointers from CLAUDE.md and
the preamble spec.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Behaviour-preserving. The invocations handler's contiguous phases become
named module-level functions with typed results, in the order the
turn-path map lists them: _resolve_turn_attachments, _prepare_session_state,
_check_turn_quota, _resolve_turn_model, _resolve_effective_tools,
_build_turn_tools, _agent_for_app_dispatch (the two MCP App dispatches
shared 60 identical lines), _validate_resume_interrupts, _build_citations,
and _refuse_turn / _forbidden_turn for the six conversational early exits.
The handler drops from ~2,240 to ~1,640 lines and its docstring is now the
phase index keyed to the turn_prelude marks. Every module-level seam the
tests patch is kept; the one AST pin on the main get_agent call now reads
the binding off turn_tools. Four comments that still described the
pre-PR-2 preamble (including the disproved ~53ms-per-GSI-query claim) are
corrected.

Full backend suite: 10950 passed, 3 skipped.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@philmerrell
philmerrell merged commit 127467f into develop Sep 30, 2026
7 checks passed
@philmerrell
philmerrell deleted the claude/gracious-hawking-9969be branch September 30, 2026 17:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant