Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CLAUDE.MD
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,7 @@ npx cdk deploy {prefix}-PlatformStack
- **Signal-based state** throughout frontend (`signal()`, `computed()`)
- **Prompt-cache stability is a contract.** Bedrock prompt caching is exact-prefix-match: any list that reaches the system prompt or `toolConfig` (skills, tools, models, MCP tool listings) must be deterministically ordered at its source, and restored conversation history must be byte-stable between compaction-state changes (see the truncation anchor in `TurnBasedSessionManager`). An order flip or history mutation between turns silently re-writes a 30k–150k-token prefix at the cache-write premium, which is **1.25× the model's own base input rate** (2× at 1h TTL; a cache read is 0.1×) — there is no flat per-MTok figure, so price it against the model actually in play: $1.375/MTok on our default Haiku 4.5, $4.125 on Sonnet 4.6. Our `us.*` ids are **Regional (CRIS)** inference profiles and price ~10% above `global.*` for the same model; rates live in `curated-models.ts`. Per-model AWS **model cards** stay the primary source — they carry context windows, caching support, parameter tables and cutoffs, which no pricing API does — but the **Price List API now corroborates almost all of them**, and a rate worth trusting should agree in both. **Which offer file matters more than which vendor:** `--service-code AmazonBedrockFoundationModels` (372 us-west-2 SKUs) carries Claude, Cohere, Palmyra, TwelveLabs, Luma, Stability; `--service-code AmazonBedrock` (1052) carries Nova, xAI, Google, DeepSeek, Qwen and Moonshot. Querying the wrong one returns nothing and reads exactly like an unpublished model — which is how an earlier revision of this note concluded the API held "10 Claude SKUs, none newer than Claude 3." **That is no longer true** (re-verified 2026-09-21): every Claude model we curate is present and matches its card to the cent — Haiku 4.5 $1.10/$5.50 Regional and $1.00/$5.00 Global, Sonnet 4.6 $3.30/$16.50, Opus 4.7 $5.50/$27.50, Sonnet 5 $2.00/$10.00 Global, Fable 5.1 $10/$50 Global with a cache read of $0.25 that independently confirms its 0.025× multiplier. The 1.100× Regional premium reproduces throughout. **Two traps remain.** The `model` attribute is gone from the product schema — match on `servicename` (FoundationModels) or `usagetype` (AmazonBedrock), or every filter silently returns zero. And the hosted OpenAI family (`us.openai.gpt-5.6-*`, `us.openai.gpt-6-astra`, `openai.gpt-5.4`) is genuinely absent from **both** offer files, so its cards remain the only source; the old explanation that Marketplace billing is uncovered is wrong — Claude's rows are `MP:` usagetypes and are covered. `PROMPT_CACHE_OBSERVABILITY_ENABLED=false` disables the observability layer (fingerprint hook, cacheStatus derivation, EMF metrics) — the caching itself stays on
- **Token cost effectiveness is a design tenet — engineer against waste, not against context.** Before merging a change that touches the model call path, answer: what does this add to the prompt, on every turn, for the life of every session? (1) Anything in the cacheable prefix (system prompt, `toolConfig`, restored history) must be deterministic and append-only — see the prompt-cache contract above. (2) Per-turn payloads (tool results, MCP responses, retrieved documents) should be bounded or offloaded, never unbounded pass-through. (3) Don't guess — verify: `cacheStatus` + fingerprint hashes on the session's `C#` rows, `GET /admin/costs/sessions/{id}/calls`, and the `AgentCoreStack/PromptCache` EMF metrics exist to prove a change's cost impact. The balance: "cost effective" means eliminating waste (avoidable cache re-writes, duplicated context, oversized payloads) — it never means stripping context the model needs for a quality answer. When cost and answer quality genuinely conflict, quality wins; look for the cheaper path to the *same* quality, not a cheaper answer
- **The turn path is mapped, stage by stage, in `docs/specs/turn-path-ttft.md`** — every stage from the SPA's send to the first token, with what runs where, what is network, what blocks the event loop, and which `turn_prelude` mark times it. Read it before touching `inference_api/chat/routes.py`, `chat/service.py`, `base_agent.py`, `session_factory.py` or the stream coordinator; the route's phases (`_resolve_turn_attachments`, `_prepare_session_state`, `_check_turn_quota`, `_resolve_turn_model`, `_resolve_effective_tools`, `_build_turn_tools`) are named for that map. It also carries the regression plan for this path (§6): the invariants, the tests per change, and the dev recipes that settle a timing claim
- **Time to first token (TTFT) is a hard budget — add no latency before the first token.** That path is everything between a request arriving and the first streamed token: request handling in `inference_api/chat/routes.py`, agent build/cache lookup, MCP pre-flight, and every `BeforeInvocationEvent` / `BeforeModelCallEvent` hook — the latter fires before *every* model call, so its cost repeats on each tool round. Default pattern: a hook captures raw facts only, and anything derived from them is computed after the model answers, memoized (e.g. `get_context_breakdown(agent, itemized=True)` itemizes on read, not in `ContextAttributionHook`). If a change genuinely cannot avoid adding latency there, measure it at realistic scale (dozens of tools, a production-sized system prompt — not a toy fixture), state the number in the PR, and get the developer driving the change to accept it explicitly before merging. A millisecond or two sounds free; the preamble was cut from 455ms to 22–37ms, so each such addition is several percent of that budget, and they compound
- **One session can be served by more than one agent — never cache session state on an agent instance.** The agent cache keys on *configuration* (system prompt, tools, model, skills), so an `@`-mention turn builds a second `Agent`, and each `Agent` builds its own `TurnBasedSessionManager`. Both write the same DynamoDB session row and neither knows the other exists, while `initialize()` never re-runs on a cache hit — so anything a manager loads once and holds goes stale silently. This has bitten twice: conversation history (#741, fixed by aliasing the message list in `_adopt_session_conversation`) and compaction state (#751, fixed by re-reading it every turn in `_adopt_persisted_compaction_state`). Per-session state must be aliased across instances or re-read per turn, and must never move backwards — a clobbered checkpoint or truncation anchor is a prompt-cache **cost** bug before it is a correctness one
- **All dependencies use exact version pins** — no `^`, `~`, or `>=`
Expand Down
Loading