Skip to content

chore(deps): strands-agents 1.57.1 + bedrock-agentcore 1.23.1 + boto 1.43.103 (paired bump) - #1367

Open
philmerrell wants to merge 2 commits into
developfrom
chore/strands-1-57-agentcore-1-23-1
Open

philmerrell wants to merge 2 commits into
developfrom
chore/strands-1-57-agentcore-1-23-1

Conversation

@philmerrell

Copy link
Copy Markdown
Contributor

Summary

One paired dependency bump, per kaizen review 2026-09-25 Proposal 4 ("Ship"). It also fixes the places where the bump broke our own code. The biggest of those is voice mode, which the research pass had called non-breaking.

Package Before After
strands-agents / strands-agents[bidi] 1.55.0 1.57.1
bedrock-agentcore 1.21.0 1.23.1
boto3 / botocore (locked) 1.43.68 1.43.103 (agentcore 1.23.1 needs 1.43.72 or later)
mcp (transitive, locked) 1.28.1 1.30.0 (still below 2, held by the declared [tool.uv] constraint)
awscrt (transitive, [bidi]) 0.32.0 0.32.2 (strands 1.57.1 pins it exactly)
strands-agents-tools 0.8.8 0.8.8 (unchanged: 0.8.9 needs mcp 2.x)

The Lambda images that pip-install from their own requirements files move with the lock: scheduled-runs worker and dispatcher, kb-sync, and kb-migration. I checked kb-migration's load-bearing boto pin against 1.43.103. The MANAGED KB type, managedKnowledgeBaseConfiguration, and all four *KnowledgeBaseDocuments operations are still there. The RAG ingestion image has its own hash-locked requirements.lock (boto3 1.42.73). It carries neither strands nor agentcore, so I left it alone.

Guard rails, now tests (backend/tests/supply_chain/test_strands_agentcore_pairing.py)

  • Pairing. strands 1.56 or later if and only if bedrock-agentcore 1.23.1 or later. The test checks the declared pins, uv.lock, and the Lambda requirements files (agentcore/boto3/botocore must match the lock). It also requires strands-agents and strands-agents[bidi] to be pinned to the same version.
  • mcp<2 comes from a declared constraint, not by accident. [tool.uv].constraint-dependencies must exclude every 2.x release. uv.lock's [manifest] must record that constraint, and the locked mcp must be below 2.
  • I mutation-tested all five failure modes locally: agentcore half-bump, bidi/base drift, constraint removed, constraint loosened to <3, and a Lambda file drifting. Each one fails the suite.

Breakage the bump caused, and the fixes

1. Voice mode. Without this PR, every voice turn would break after deploy. Strands 1.56 and 1.57 rewrote the experimental Bidi API:

  • Provider constructor. BedrockNovaSonicModel(audio=...) now takes {"input": {"sample_rate"}, "output": {"sample_rate"}}, and voice is its own kwarg. Our five-key dict still constructs, but it only warns, so the voice silently fell back to matthew.
  • Input. BidiAgent.send() accepts only {"audio_delta": {...bytes}}, {"text": ...} or a str. Our {"type": "bidi_audio_input", ...} dicts now raise ValueError.
  • Output. Every event was renamed or reshaped: bidi_audio_delta, bidi_transcript_start/delta/stop, bidi_barge_in, bidi_response_stop (with no stop_reason) and bidi_connection_stop. bidi_response_start now fires at the user's first content, before any user speech has been transcribed. The FINAL assistant transcript pass is gone.

Fix: a VoiceWireAdapter in voice_agent.py translates 1.57 events back to the existing WebSocket contract, so the SPA and voice_routes.py are unchanged. The legacy bidi_response_start is emitted when the assistant starts, which is where 1.55 emitted it. Without that, the SPA would file each user utterance one turn late. Assistant transcript deltas go out with is_final=True, because the speculative pass is now the only one. Turn and usage accounting runs on the translated stream.

One side effect: response_start_count is now one per turn. Under 1.55 it was about four per turn, because it counted every Nova contentStart. _finalize_voice_session takes max(completed, started), so voice sessions' message_count stops being inflated.

2. TurnBasedSessionManager now sees the voice BidiAgent. A BidiAgent now drives the ordinary session hooks (AgentInitializedEvent → initialize, MessageAddedEvent → retrieve_customer_context) instead of the dedicated Bidi callbacks it used before. Two guards:

  • initialize restores a BidiAgent through the SDK and returns. Without this, the text agent's session-level compaction checkpoint would be applied to the voice agent's own message list, which slices the wrong history.
  • retrieve_customer_context skips a BidiAgent. This matches the guard agentcore 1.23.1 added to the method we override. Without it, every voice transcript would trigger an LTM retrieval spliced into the live Bidi history.

3. The context-window fallback now resolves the hosted OpenAI ids. 1.56's nested-prefix strip (#4221) resolves us.openai.gpt-6-astra and us.openai.gpt-5.6-* to 1,050,000. The curated rows' deliberate maxInputTokens: 272_000 pricing cap still wins, so each pair now logs one context_window_disagreement warning per process, as expected. I updated the tests and the docstring.

⚠️ For the reviewer: a managed OpenAI-family row with maxInputTokens absent would now fall through to 1.05M instead of None. That means compaction would cut past the 272K price tier. I did not add a guard, because that is a policy call. Worth checking that no dev or prod row lacks the field.

harness-sdk#4618 assessment: no Bedrock usage change in 1.57.1

  • strands/models/bedrock.py changed only in formatting between 1.55.0 and 1.57.1 (two reasoning-delta dicts were re-wrapped).
  • strands/telemetry/metrics.py did not change. _total_prompt_tokens, the tell the kaizen gate names, is still present.
  • Bedrock Converse usage still arrives disjoint, so usage_normalization.py's "leave Bedrock untouched" contract holds.
  • A new test in test_usage_normalization.py fails the day _total_prompt_tokens disappears.
  • #4193 (cache_write_tokens mapping) merged after 1.57.1 was cut, so it is not in this pin.

#4361 did land, for Chat Completions only. OpenAIModel now reports cacheWriteInputTokens from prompt_tokens_details.cache_write_tokens, next to an inputTokens that is still inclusive. Our usage_normalized wrapper subtracts it exactly once. With 10,000 prompt tokens, 6,000 cached and 3,000 written, the result is inputTokens=1000, read 6000, write 3000, which is disjoint. A test pins this against the real SDK class. On the Chat Completions path, cache writes are now priced at the write rate instead of as plain input. Responses still needs our own mapping, which is unchanged.

Upstream diff: what touches our surface (1.55.0 to 1.57.1, read in site-packages)

  • Prompt caching. No change to CacheConfig behaviour on the Bedrock path. cache_key now auto-derives strands-<session_id> for OpenAI/LiteLLM/Mistral when unset (1.56 #4083), but we only set cache_config on Bedrock, and bedrock_responses sets its own prompt_cache_key first. probe_bedrock_cache_point_support.py --offline-only gives an identical table on 1.55.0 and 1.57.1: Haiku gets a tools point, and our system point survives on all five probe models.
  • Context manager (standing watch). We don't use ContextManager. Upstream changed its default strategies (truncate at 1,500 with a 750 preview; summarize at 0.85 with preserve_recent=4). The Stash now always namespaces per (session_id, agent_id), so trap 2 still stands and #4367 is still the blocker. clear_session() is new. The retrieval tool now returns media instead of an error. The ContextOffloader plugin we subclass, S3Storage's write/read path and SlidingWindowConversationManager are unchanged. #4254's session integration carries Strands' own stash into its session snapshots; nothing in it addresses our DynamoDB checkpoint or truncation anchor.
  • Interrupts and hooks. The event loop now raises interrupts registered during BeforeModelCallEvent hooks. BeforeModelCallEvent is still not _Interruptible, and none of our hooks interrupt. InterruptException is no longer wrapped in EventLoopException. We have no except EventLoopException pause detection. interventions now propagates handler interrupts regardless of on_error (#4371). MessageUpdatedEvent is new. AfterToolsEvent / MessageAddedEvent semantics are unchanged for steering. SessionManager became Generic[_SessionAgentT] and its Bidi callbacks moved (see fix 2).
  • MCP Apps seam. strands/tools/mcp/mcp_client.py is byte-identical, so the ClientSession substitution is unaffected.
  • Tool specs. #4426 builds a normalized copy instead of mutating the caller's spec. The wire bytes are unchanged.
  • Telemetry. Agent spans now carry gen_ai.system_instructions through _redact, so our gen_ai_unredacted_attributes= allowlist redacts them like the model spans.
  • Import time (the 30 s Runtime init budget). A warm import of strands + agentcore session manager + our session/voice modules takes 2.4–3.0 s on 1.57.1, against 2.3–3.2 s on 1.55.0.

Capabilities worth adopting later (not in this PR):

  • strands.vended_tools.handoff_to_user has the same interrupt shape as our ask_user_question. Read it before building a third interrupt tool.
  • The Bedrock requestTimeout in the 1.57.0 notes is the TypeScript SDK (#4408). The Python BedrockModel has no such option; we already bound reads through boto_client_config.
  • Other new pieces: mcp_router and a2a_client vended tools, Agent.shutdown(), background tasks, and bedrock_mantle_config.endpoint="bedrock-runtime", which accepts us.openai.* CRIS ids.

Tests

  • Full backend suite, in the worktree's own env (uv sync --extra agentcore --extra dev, then again with --extra bidi): 10891 passed, 3 skipped, 0 failed in both envs (CI-shaped without bidi, and with bidi so the voice path is exercised against the real 1.57.1 Bidi classes).
  • Before the fixes, the bump alone left 5 failures, all caused by the bump: 2 voice contract tests and 3 context-window tests. Those led to fixes 1 and 3.
  • Root tests/supply_chain/: 48 passed.

Dev validation after merge (develop auto-deploys dev)

A missed pairing does not fail at import in a unit test. On the Runtime it shows up as "Runtime initialization time exceeded (30s)" → 502. So validate real turns, not an import:

  1. Text turn with a tool call. Run a multi-turn chat on the default Haiku, where at least one turn calls a tool, and a follow-up turn that should read the cache.
    • Expect: no 502 or init timeout, and the stream completes.
    • Then open GET /admin/costs/sessions/{id}/calls. The C# rows should have the same shape as before: cacheStatus present, turn 2 a hit, cacheReadInputTokens non-zero, and toolConfigHash / systemPromptHash stable between the two turns.
    • inputTokens should be small relative to cacheReadInputTokens. If they are comparable, Bedrock usage has gone inclusive.
  2. Voice. Open voice mode, speak two turns, interrupt the assistant once, then close.
    • Expect: audio plays at the right speed in the configured voice (tiffany, not matthew); each user utterance shows before its reply; barge-in stops playback.
    • Afterwards the conversation shows both turns and the voice session's cost and metadata rows are written.
  3. A GPT-5.6 Chat Completions turn, if one is enabled in dev: cacheWriteInputTokens may now be non-zero, and the three input buckets should still sum to the call's total input.
  4. Runtime logs: one context_window_disagreement warning per OpenAI model id per process is expected. Anything more is not.

Kaizen docs are deliberately untouched; the queue entry moves in a separate PR.

🤖 Generated with Claude Code

philmerrell and others added 2 commits September 27, 2026 11:13
…boto 1.43.103

One paired bump (kaizen review 2026-09-25, Proposal 4):
- strands-agents / strands-agents[bidi] 1.55.0 -> 1.57.1
- bedrock-agentcore 1.21.0 -> 1.23.1 (strands >=1.56 removed the Bidi hook
  events agentcore <=1.23.0 imports at module load)
- boto3/botocore 1.43.68 -> 1.43.103 (agentcore 1.23.1 floor is 1.43.72)
- mcp 1.28.1 -> 1.30.0, still held below 2 by the declared uv constraint
- Lambda requirements files follow the lock

The two guard rails from the superseded 1.56 entry are now assertions in
tests/supply_chain/test_strands_agentcore_pairing.py: the strands/agentcore
pairing (declared, locked, and in the Lambda images), and mcp<2 held by a
declared [tool.uv] constraint rather than incidentally by the idna pin.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s 1.57

Breakage caused by the strands-agents bump:
- Voice: 1.56/1.57 rewrote the experimental Bidi API (provider audio config
  and voice kwarg, BidiAgent.send() input shapes, renamed and reshaped
  output events). VoiceWireAdapter translates the new events back to the
  existing WebSocket contract, so the SPA and voice routes are unchanged.
- TurnBasedSessionManager: a BidiAgent now drives the ordinary session hooks.
  initialize() restores it without the text-only compaction/repair path, and
  retrieve_customer_context() skips it, mirroring agentcore 1.23.1's guard.
- Context window: the SDK table now resolves the hosted OpenAI ids via its
  nested prefix strip; tests and docstring updated (the curated 272K cap
  still wins).
- Usage: pin #4361's Chat Completions cache-write mapping (normalized once)
  and add a tripwire for harness-sdk#4618 (_total_prompt_tokens).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant