fix(gcp-to-aws): align Sonnet 5 pricing, guidance, and report estimates - #269
Conversation
…w the standard price Anthropic's pricing page states the $2/$10 introductory rate for Claude Sonnet 5 "is now the standard price" and that "the previously scheduled increase to $3/$15 on September 1, 2026 will not occur." The Bedrock pricing page (Anthropic tab, read 2026-09-03) lists Sonnet 5 at $2.00/$10.00 with no promotional qualifier. Our documents had been pricing Sonnet 5 at the cancelled $3/$15 "steady-state" rate in every comparison, overstating it by a third. Every Sonnet 5 comparison is re-derived at $2/$10 (blended 2:1 in:out, per the house rule), and several conclusions flip: - GPT-5.6 Terra vs Sonnet 5: "Terra 19% cheaper" -> Sonnet 20% cheaper - GPT-5.4 vs Sonnet 5: "near parity" -> Sonnet 36% cheaper - GPT-5.5 vs Sonnet 5: 52% -> 68% cheaper - GPT-4o vs Sonnet 5: "Sonnet 29% more expensive" -> Sonnet 7% cheaper - GPT-5.2 vs Sonnet 5: "source 17% cheaper" -> Sonnet 20% cheaper - GPT-5.1/5 and GPT-4.1: source advantage shrinks to 11% / 14% - GPT-4 Turbo / GPT-4: 58% / 82% -> 72% / 88% cheaper - Gemini 3.1 Pro vs Sonnet 5: "+24%" -> -13%; Sonnet 5 is now the cheaper side Files: pricing-cache (narrative + quick-reference status), clarify-ai and clarify-ai-only baseline tables, ai-openai/ai-gemini/ai-anthropic design refs, and the test comment in test_bedrock_pricing.py (the STATIC_FALLBACK already carried 0.002/0.010). The "do not default to Fable 5" sentences in the openai and gemini refs now also name Fable 5.1 (see awslabs#267). Both plugin copies updated.
…ine, not one version (review on awslabs#267)
leon1418
left a comment
There was a problem hiding this comment.
[🤖 AI review 🤖]
Scope: Full review of all 14 changed files (7 advisor + 7 migrate twins). Every human-written changed line inspected. Design, functionality, arithmetic, documentation, tests, and cross-plugin parity all checked.
Change summary: Claude Sonnet 5's $2/$10 launch rate became the permanent standard price on Sep 1, 2026 (the scheduled step-up to $3/$15 was cancelled). This PR re-derives every Sonnet 5 comparison across all pricing reference docs at the now-standard $2/$10 rate. Several conclusions flip direction (e.g. GPT-5.6 Terra was "19% cheaper" → Sonnet is now 20% cheaper; GPT-5.2 was "source 17% cheaper" → Bedrock is now 20% cheaper). The Fable exclusion list is also extended to cover Fable 5.1.
Design: Sound. The change is well-motivated — documents that priced Sonnet 5 at the cancelled $3/$15 rate were actively misleading. Updating all references at once, with both plugin copies, is the right approach. The conservative Tier-0 guidance (same-model default, Sonnet 5 as cost alternative) is preserved, which is correct — the price change strengthens the cross-family cost case without changing the risk calculus.
Arithmetic verification (2:1 blended in:out, independently recomputed):
- GPT-5.6 Terra vs Sonnet 5: 20% ✅
- GPT-5.5 vs Sonnet 5: 68% ✅
- GPT-5.4 vs Sonnet 5: 36% ✅
- GPT-4o vs Sonnet 5: 7% ✅
- GPT-5.2 vs Sonnet 5: Bedrock 20% ✅ (direction flip correct)
- GPT-5.1/5 vs Sonnet 5: Source ~11% ✅
- GPT-4.1 vs Sonnet 5: Source ~14% ✅
- GPT-4 Turbo: Bedrock 72% ✅
- GPT-4: Bedrock 88% ✅
- Gemini 3.1 Pro monthly: -13% ✅ (12.5% rounds to 13%)
- Gemini 3.1 Pro blended: ~12% ✅
All percentage claims are arithmetically correct.
Cross-plugin parity: ✅ PASS — all 7 file pairs (advisor/aws-startup-advisor ↔ migrate/migration-to-aws) are byte-identical at the current head. No drift allowlist changes.
Tests: The test comment update in test_bedrock_pricing.py is documentation-only (the test itself already asserted 0.002/0.010 rates, which are correct). No test logic changed.
Findings: 0 mandatory blockers, 0 Nits.
The change improves overall code health: it corrects pricing data that was actively wrong, the arithmetic is verified, and both plugin copies are in sync. Clean.
CI: 5 checks QUEUED (build, gitleaks, bandit, semgrep, checkov) — not yet completed.
Approvals: 0 (REVIEW_REQUIRED).
Merge state: BLOCKED (pending CI + approval).
Head: 6379a2181b7b4c13403d43fa0d3788c7e31232f9 (2 commits).
herosjourney
left a comment
There was a problem hiding this comment.
Reviewed the full 14-file diff (7 advisor + 7 migrate twins) against Anthropic's live pricing page and the Bedrock Sonnet 5 model card, and recomputed every flipped percentage at 2:1 blended in:out.
Upstream claim
Confirmed on Anthropic pricing (read today):
The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur.
Table row is $2 / $10 with cache write $2.50 / $4 and cache read $0.20. The Bedrock Sonnet 5 model card defers dollar rates to the Bedrock pricing page (Anthropic tab is JS-rendered and did not come through in a static fetch). Treating Bedrock on-demand as matching the Anthropic list rate is consistent with how this cache is already sourced.
Arithmetic (2:1 in:out, independently recomputed)
| Comparison | Claimed | Recomputed | Verdict |
|---|---|---|---|
| Terra 2.20/13.20 vs Sonnet 2/10 | Sonnet 20% | 20.45% | ✅ |
| GPT-5.5 5.50/33 vs Sonnet | Sonnet 68% | 68.18% | ✅ |
| GPT-5.4 2.75/16.50 vs Sonnet | Sonnet 36% | 36.36% | ✅ |
| GPT-4o 2.50/10 vs Sonnet | Bedrock 7% | 6.67% | ✅ |
| GPT-5.2 1.75/14 vs Sonnet | Bedrock 20% | 20.00% | ✅ |
| GPT-5.1/5 1.25/10 vs Sonnet | Source 11% | 10.71% | ✅ |
| GPT-4.1 2/8 vs Sonnet | Source 14% | 14.29% | ✅ |
| GPT-4 Turbo 10/30 vs Sonnet | Bedrock 72% | 72.00% | ✅ |
| GPT-4 30/60 vs Sonnet | Bedrock 88% | 88.33% | ✅ |
| Gemini 3.1 Pro 150M+75M | $1,200 → $1,050 (−13%) | −12.5% | ✅ |
| Gemini blended vs Sonnet | ~12% | 12.5% | ✅ |
Unchanged rows (Sol 12%, Opus 20%, Luna/Nova) left alone, correctly.
Design
Right call to keep Tier-0 same-model GPT as the default and present Sonnet 5 as the cost alternative now that the margin is real. The Fable 5.1 / Mythos exclusion wording is consistent with #267 and does not change the default.
Finding — leftover "source is cheaper" list is now false
ai-openai-to-bedrock.md (both copies), immediately under the Option B table that this PR rewrites:
For the rows where the source is cheaper (GPT-5.2, GPT-5.1/5, GPT-4.1, GPT-4o), note that these models are on a
GPT-5.2 and GPT-4o just flipped to Bedrock cheaper. Leaving them in that list undoes the honesty fix two paragraphs up. Suggested:
For the rows where the source is still cheaper (GPT-5.1/5, GPT-4.1), note that these models are on a
Same class of stale-claim problem this PR exists to close. I'd want this fixed before merge.
Nits (non-blocking)
ai-gemini-to-bedrock.md"Migration case by tier" still says Gemini 3.1 Pro → Bedrock is "NOT cost", while the same file's cost bullet and monthly table now show Sonnet ~12–13% cheaper. Soften to matchclarify-ai.md("modest edge; case is reliability/ecosystem more than cost").Recommend defaults (Jul 2026)is now a stale month stamp on a Sep 3 price change. Drop the month or move it to Sep 2026.pricing-cache.mdLast updated jumps 2026-08-24 → 2026-09-03 for a Sonnet 5 status edit (the $2/$10 cells were already correct). That resets the 30-day staleness clock for the whole cache, including infra rows that were not re-pulled. Fine if intentional; worth a one-liner in the commit/PR if it is.
Parity / scope
Advisor ↔ migrate twins are byte-identical in the diff. Test change is comment-only; STATIC_FALLBACK already asserted 0.002/0.010. Sonnet 4.6 $3/$15 and the Sol promo note are correctly left alone.
Verdict: approve the rate change; one leftover sentence needs a two-word list edit before merge.
herosjourney
left a comment
There was a problem hiding this comment.
Inline notes for the review posted above. One leftover list is wrong after the table flip; the rest are nits.
…5 re-derivation; address review nits - ai-openai-to-bedrock: GPT-5.2 and GPT-4o flipped to Bedrock-cheaper in the Option B table, so the follow-on sentence now lists only GPT-5.1/5 and GPT-4.1. - ai-gemini-to-bedrock: the "Migration case by tier" bullet for Gemini 3.1 Pro said "NOT cost" while the same file's numbers show Sonnet 5 ~12% cheaper; now "modest cost edge, not cost-led", matching clarify-ai. - "Recommend defaults (Jul 2026)" -> (Sep 2026) in the openai and gemini refs. - pricing-cache "Last updated" restored to 2026-08-24: this PR changes no numeric cell, so the whole-cache staleness clock should not reset here.
|
All five inline notes addressed in c0e2cc0 (both plugin copies, drift:check OK): the leftover "source is cheaper" list trimmed to GPT-5.1/5 and GPT-4.1; the Gemini 3.1 Pro tier bullet softened to match the file's own numbers; |
herosjourney
left a comment
There was a problem hiding this comment.
Re-reviewed at c0e2cc0. All five items from the earlier pass are in the tip:
- Blocking leftover list is now
GPT-5.1/5, GPT-4.1(both twins); GPT-5.2 / GPT-4o stay Bedrock-cheaper in the table. - Gemini 3.1 Pro bullet matches the ~12% edge and is no longer "NOT cost".
- Recommend-defaults stamp is Sep 2026.
- Cache Last updated restored to 2026-08-24 so the staleness clock does not reset on a status-only edit.
Arithmetic from the first pass still holds. No remaining $3/$15 comparison rows. Approve.
|
Both plugin copies. |
…ables, Clarify Q19 shortcuts, base-of-comparison wording Audit of every Sonnet 5 mention in both plugins (141 lines, 30 files) after the second review on awslabs#269. The previous sweep matched `3.00 / 15.00` and `$3/$15` but not `$3.00 / $15.00`, which is how the Gemini mapping tables are written, and it could not see conclusions stated without a price at all (Clarify's Q19 shortcut table). Corrected at $2/$10, 2:1 blended: - ai-gemini-to-bedrock: Gemini 3.1 Pro vs Sonnet 5 "Gemini 24% cheaper" -> Bedrock 13% cheaper (now agrees with the monthly table); 2.5 Pro 40% -> 11%; 3.5 Flash 33% -> 14%; price cells at the Flash Thinking and 1.5 Pro rows; "+31% cost" -> -13%; volume table $1,575 (+425%) -> $1,050 (+250%). - clarify.md Q19 shortcuts: GPT-5.5 "53% savings" -> 68%; GPT-5.4 "near price parity" -> ~36% cheaper (matches clarify-ai); GPT-4 Turbo "70% cheaper on input" -> 80%. - clarify-ai GPT-5/5.1/5.2 row: GPT-5.2 is ~25% more expensive measured against Sonnet 5 (the earlier 20% used GPT-5.2 as the base); "roughly neutral" replaced with the split verdict. - 12.5% rounded consistently to 13% (was ~12% in prose, -13% in the table). Acceptance: no Sonnet 5 line carries a $3/$15-shaped price except the sentences that explain the cancelled increase; no old derived figure remains.
|
Fixed in 4df8200 (both copies): Q19 shortcuts now read GPT-5.5 → "Sonnet 5 for 68% savings", GPT-5.4 → "~36% cheaper blended" (matching clarify-ai:329), GPT-4 Turbo → "80% cheaper on input". You're right that this is the first table a user sees. Two things in the same rows I did not touch, to keep this PR to the Sonnet 5 change: "Claude Opus 4.6 — Bedrock 17% cheaper on output" (the repo's default is Opus 4.8 and 25 vs 33 is 24%, not 17%) and the "Opus 4.6 for hardest" phrasing at :576/:581 — those are a separate Opus 4.6→4.8 consistency fix. |
ayn-builds
left a comment
There was a problem hiding this comment.
Reviewed at 4df8200. The re-derivation holds up. I checked the comparisons the notes below rest on, found no stale $3/$15 left under skills/, and confirmed all seven changed doc pairs are still byte-identical across the two plugins. The bedrock_pricing.py change is comment-only and the fallback already had 0.002/0.010. I couldn't run test_bedrock_pricing.py: my Python is 3.9 and line 99 uses dict | None.
Three notes below are the same underlying thing, prices got updated but the conclusions drawn from them didn't. Since the plugin copies are identical, each fix lands twice.
One more, outside the diff so treat it as a heads-up rather than a request. fixtures/migration-report-reference.html:777 still prices Sonnet 5 at $3/$15, and the token table gives it away: 60M input billed at $180 and 40M output at $600 only work at the old rate. At $2/$10 it's $120 and $400, a $520 total and about -58% instead of -38%.
I'd normally leave a fixture alone, except report-decision-core.md:40 points agents at this file as the example of "real artifact data" with actual numeric costs, so it's what gets copied into customer-facing cost tables. Haiku and Nova in the same table are already correct, which makes Sonnet 5 the one row that got missed, and nothing tests these values. Fine as a follow-up if you'd rather keep this PR to the eight files.
|
@ayn-builds Addressed the report-fixture follow-up from your latest review in 40bf8a8, in both plugin copies.
All four new inline comments also have replies against the same commit. Validation: drift/shared/fixture checks, Markdown lint and dprint passed; Node suites passed 62 + 62 + 12 tests; the llm-to-bedrock suite passed all 249 tests via uv on Python 3.11.13, including the pricing tests. I independently checked the fixture arithmetic and verified the updated values in the rendered report. The PR description now reflects the complete change and validation. |
|
Rechecked every review comment and merged current main (
Validation: 136 Node tests, 249 llm-to-bedrock tests, and 134 report-validator tests passed; drift/shared/fixture checks, Markdown lint, and formatting passed. Independently reconciled HTML and JSON token costs, monthly/annual totals, row deltas, unique sections, and TOC targets. |
Summary
Claude Sonnet 5 references still used the canceled $3/$15 rate, and several recommendation sentences and report totals did not follow the corrected price cells. This PR aligns both plugins' comparison tables, Clarify guidance, and report fixture with the standard $2/$10 input/output price per 1M tokens.
Pricing and recommendations
Comparisons use the existing 2:1 input/output token ratio unless a table states its own volume. The resulting Sonnet 5 savings are:
GPT-5/5.1 and GPT-4.1 remain cheaper at source, by approximately 11% and 14% respectively when measured against Sonnet 5. Tier-0 same-model GPT moves remain the default on risk grounds; Sonnet 5 is an optional cost alternative.
The Gemini Thinking guidance no longer promises comparable or lower costs from using Sonnet 5. At the table's $0.30/$3.50 upper-rate example, Gemini 2.5 Flash is approximately $1.37/M blended versus $4.67/M for Sonnet 5. The recommendation now treats Sonnet as a capability alternative and requires profiling each model's billed thinking tokens. The Gemini 3.5 Flash Thinking sentence and the full-thinking cost override follow the same rule.
The generic Gemini Pro shortcut retains the neutral "Comparable tier" description because the cost direction differs between 2.5 Pro and 3.1 Pro. Fable 5.1 names and the undefined "Covered Models" term are deferred to #267; the existing frontier-model exclusion remains.
Report fixture
Both copies of
migration-report-reference.htmlnow use $120 for 60M input tokens and $400 for 40M output tokens: $520 for Sonnet 5, or $540 including the existing $20 image estimate. The companionestimation-ai-reference.jsonfiles also use $540 and a 57% reduction. The headline, CFO summary, monthly and annual tables, and model-downsizing estimates agree with these values:The infrastructure delta is corrected to $53 ($165 minus $112) so row deltas reconcile with the combined total. The cache-savings estimate is marked workload-dependent because the fixture specifies no cache-hit or write assumptions.
Scope and provenance
The report preserves the current main branch's standalone Cost Optimization and Savings Plans/Reserved Instances sections, their TOC links, and warnings against stacking commitment savings onto the Optimized tier. Corrected Sonnet downsizing and workload-dependent caching guidance appear in the moved opportunity table.
Both plugin copies stay byte-identical for the changed files. The pricing-cache
Last updatedremains 2026-08-24: its numeric price cells were already $2/$10, so a status-only correction does not reset the entire cache's freshness clock. Sonnet 4.6 pricing and the separate Sol promotional rate are unchanged.Sources recorded for the rate change on 2026-09-03:
Validation