feat(azure-to-aws): add Azure→AWS migration skill (infra + AI/agentic tracks) - #304
Conversation
…/ai/
A third migration skill (azure-to-aws) needs the same Bedrock-side AI content
gcp-to-aws carries, and a copy would be a drift surface — `shared:check` only
protects a file that has a canonical home. So eight files move to the
plugin-neutral `skills/shared/ai/` and are vendored back into gcp-to-aws at
`references/vendored/ai/`, byte-identical, the same lockfile pattern the DSL
and estimate schemas already use.
Promoted (source-cloud agnostic by inspection — the target is Bedrock or
AgentCore, and the mapping keys off the SDK the code uses, not its host):
ai-model-lifecycle.md ai-migration-guardrails.md bedrock-quotas.md
ai-openai-to-bedrock.md ai-anthropic-to-bedrock.md
design-ref-harness.md design-ref-agentic-to-agentcore.md
sdk-capability-map.json
Deliberately NOT promoted: `ai-gemini-to-bedrock.md` and `schema-discover-ai.md`,
both written against one provider's SDK surface, and `ai.md`, which is gcp's
category rubric.
`sync-vendored-shared.ts` needed no change: it walks each skill's vendored side
and requires a canonical match, so partial vendoring works and skills need not
vendor identical subsets.
Two prose lines in ai-openai-to-bedrock.md said "GCP-hosted applications"; they
now name the source cloud generically and state that Azure OpenAI routes here
too, since the Bedrock target does not depend on which endpoint served the
calls. ai-model-lifecycle.md's "Mapping Guides" heading no longer names the
one guide that stayed behind in gcp.
Consumers updated in both plugin trees: gcp's SKILL.md conditional-load table
and file tree, design-refs/{index,ai}.md, phases/{design,discover,clarify,
estimate}, shared/pricing-{cache,fallback}.md, and agent-advisor's
migration-plan refs. Bare same-directory filenames became explicit paths —
those files are no longer siblings.
Two hardcoded paths outside the skills also moved: agent-advisor's
`test_scoring.py` lifecycle drift-guard, and `model-id-lint.py`'s catalog
allowlist. The lint now collapses `skills/<skill>/references/vendored/<rel>`
to `skills/shared/<rel>` before matching, so one canonical allowlist entry
covers every vendored mirror — otherwise the next skill to vendor a catalog
file would fail the lint until someone remembered to add its path, which is
the per-copy manifest vendoring exists to avoid.
A documented gap: the promoted files still reference
`references/shared/pricing-{cache,fallback}.md`, which stay per-skill because
pricing-cache.md doubles as gcp's AWS-infrastructure rate card. The new
`skills/shared/ai/README.md` states that as a consumer contract so it is not a
latent dangling reference.
`phase-status.schema.json` sets `additionalProperties: false`, so a skill whose artifact-generating phase is opt-in has nowhere to record that the user opted in. gcp-to-aws already works this way and azure-to-aws will gate Generate on it. Added as an OPTIONAL enum (`decide` | `decide_and_execute`) rather than a defaulted one, because the absent state is load-bearing: absent means the decision gate has not been presented yet, and must never read as consent. Skills with an unconditional Generate keep omitting the key, which is why this is not `default: "decide"`. The description also pins the write ordering — `decide_and_execute` is written BEFORE the generate phase loads — so a session that dies mid-Generate resumes as an Execute run instead of re-asking a question the user already answered. Vendored copies re-synced.
Step 1 of the build sequence: prove the wiring before any content lands. Seven
phases — discover → clarify → design → estimate → generate on the backbone, plus
the `workshop` and `feedback` sidebars — each with real fragments, exactly one
assembler, gates, re-entry guards, and the artifact graph the whole skill hangs
off. `lint:frontmatter` reports 7 phase files checked and passes in both trees.
The non-zero count is the point. `parse.ts` returns null for a file that does not
start with `---`, so a prose-shaped skill passes `lint:frontmatter` green while
being completely unvalidated — that is exactly what gcp-to-aws does today (0 phase
files checked). Building this as a gcp-style clone would have shipped a skill the
team believed was DSL-checked and was not.
Structural decisions the skeleton locks in, because retrofitting them changes
every downstream contract:
- **The artifact graph.** `azure-resource-inventory.json` + `azure-resource-
clusters.json` → `preferences.json` → `aws-design.json` → `estimation-infra.json`
→ the generate set, with `scenarios/index.json` off the workshop sidebar. Every
phase's `_input` resolves to an upstream `_produces`; the validator enforces it.
- **ARM as the canonical vocabulary.** `azure_id` is the full ARM resource ID and
`azure_type` a `Microsoft.*` string. Four of the five discovery sources speak
ARM natively; only Terraform needs translating. The ARM ID also embeds
subscription and resource group, so one field supplies the cluster seed key, the
environment scope, and uniqueness with no derivation.
- **`_exec: { _agent: rw }` + `_interactive: false` on discover and generate.** Both
are bulky, file-only and self-contained. This is also what forces live `az`
capture to become main-window pre-work later rather than a fragment: a dispatched
worker cannot prompt for consent.
- **Generate is gated on `run_mode: decide_and_execute`** plus
`_check_phase_completed: workshop` (the sidebar declares `_gates: generate`).
The `run_mode` check is an `_assert`, which CI binds but never evaluates — that
is stated in `generate.md` so a future reader does not assume it is enforced.
- **Postconditions encode the FINISHED contract, not the skeleton's behaviour.** A
category or table landing in a later step cannot land silently: its assert fails
until the unit exists. Every unit file carries a `## Status` block naming what it
does today and which step fills it in.
Three shared refs ship with real shapes rather than stubs, because downstream
phases are written against them: `schema-discover-azure.md` (inventory, the typed
edge table, the drift record, clusters), `schema-preferences.md` (the assumption
sheet's four dispositions, and why a considered `N/A` is written explicitly rather
than omitted), and `schema-workshop-scenarios.md`.
Repo wiring per the plan's §1 checklist: `mise.toml` `lint:frontmatter` (both
trees), `cross-plugin-drift.ts` `SKILLS`, the three plugin.json flavours in both
trees, three marketplace manifests, `advisor/AGENTS.md`, both READMEs, and the
gcp/heroku SKILL.md descriptions now name azure-to-aws instead of excluding Azure
with no forwarding address. `architect-for-startups` keeps
`references/migration-azure-to-aws.md` for the pre-decision advisory conversation
and hands off to the skill on real migration intent.
Telemetry is wired now, including the two `emit.mjs` defects a third registered
skill triggers — both silent, both fixed here because registering AZURE_TO_AWS is
what activates them:
- The run-ownership guard resolved a single "other" skill via `.find()`, which was
only correct while exactly two were registered. With three, a GCP-registered hook
looking at an Azure run could resolve `other` to HEROKU_TO_AWS, whose inventory
is absent, so the guard did not fire and the Azure run was reported as
`sourceProvider: GCP`. It now checks every other registered skill.
- `costContainer` read `current_costs ?? gcp_baseline`. The baseline key is now
looked up from the registering skill (`SKILL_INVENTORY[skill].baseline`), because
a hardcoded key reproduces the $685M dropped-spend bug the comment above it
documents, once per skill that is not the one it was written for.
`hooks/telemetry/cursor/hooks.json` registered GCP_TO_AWS only; it now registers
all three, which also fixes HEROKU_TO_AWS having been missing there since it
shipped.
The migrate copy of `SKILL.md` carries no telemetry block, matching gcp and heroku:
that tree ships no `hooks/` directory, so the hook command would resolve to a
nonexistent file, and the emitter-path search pattern names the advisor plugin,
which is not one of the drift tool's normalized token classes. It is allowlisted
with that reasoning recorded, alongside the same three vendored files heroku
already allowlists.
Deliberately NOT in this commit, per the sequence: Bicep/ARM/RDfA/live/billing/
app-code fragments (step 2), the mapping tables (step 3), real clustering and the
pattern catalog (step 4), rubric and knowledge content (step 5), the dual estimate
output and the generators (step 6), fixtures (step 7). The `ai/` vendored subtree
is also deferred to the AI-route work, because the promoted refs depend on a
per-skill `references/shared/pricing-cache.md` this skill does not ship yet.
… external oracle Pulls the canonicalization table forward from step 3 to before step 2's discovery breadth, because until it exists there is nothing to TEST. The reason is the DSL's central limitation: `_assert` postconditions verify shape, never correctness, and the model both produces the artifact and evaluates the assertion against it. There is no independent oracle. So `discover-iac.md` with no extraction rules at all still passes: a capable model reads `azurerm_linux_web_app`, emits `Microsoft.Web/sites` from pretraining, inspects its own well-formed output against "every entry has a canonical Microsoft.* type", and emits HANDOFF_OK. That is a FALSE GREEN — structure satisfied, contract satisfied, zero skill content consulted, output stamped `confidence: deterministic` when it is a model prior, and unreproducible across runs. You would conclude discovery works. Landing: - `references/shared/arm-type-canonicalization.md` — the `azurerm_* → Microsoft.*` table across compute/data/networking/identity/messaging/observability, plus the six traps where a plausible guess is wrong: `Microsoft.Web/functionApps` does not exist (a function app is `sites` + `kind`), the plan is `serverfarms` lowercase while `serverFarmId` is the camelCase property, `Microsoft.Cache/Redis` carries a capital R, Cosmos DB's provider is still `DocumentDB`, Azure OpenAI is a Cognitive Services account, and a resource group's own ID has no `/providers/` segment. Also the `azure_id` reconstruction rules, including the `<subscription-unknown>` placeholder — inventing a GUID makes the ID unstable across runs and breaks drift comparison against a live capture, which is the one thing the ID exists to support. - `references/shared/extract-terraform.md` — extraction, per-type attributes (only what a downstream table actually reads), the edge table, and the secret boundary. `count`/`for_each` become ONE entry with the expression recorded rather than being fanned out, because the count is usually a variable and fanning out invents resources. `tfstate` is never read. Registry modules absent from the workspace get a warning, that being the commonest reason a Terraform-only inventory is confidently incomplete. - `discover-iac.md` rewritten with real dialect detection and canonicalization, and a **halt guard**: a dialect present whose `extract-*.md` ref is missing now emits GATE_FAIL instead of exiting cleanly. Previously "no dialect present" and "the dialect is present but this skill was never taught to read it" were indistinguishable, and the second is exactly the hole improvisation walks through. So a `.bicep` or ARM workspace stops today rather than half-discovering — a partial inventory presented as complete makes the estimate confidently wrong about the size of the estate. - `fixtures/azure-iac-terraform/` — a 28-resource synthetic estate, `expected-*.json`, and a stdlib asserter, both trees. Synthetic rather than a real repo because a real repo will not contain the traps: this one deliberately puts five web apps on one S1 plan, an idle plan with zero apps, an app in `rg-app` with its database in `rg-data`, a second resource group holding two unrelated workloads, a function app, an SMB file share, a Kafka-enabled Event Hubs namespace, a Mongo-API Cosmos account, a private endpoint, a cost-bearing type absent from the table, and an unresolvable module. The asserter pins ONLY facts where improvisation and correctness diverge; facts a model gets right by accident are deliberately not asserted, since they cost review attention and prove nothing. Verified non-vacuous against a hand-written inventory carrying `Microsoft.Web/functionApps` and `Microsoft.Cache/redis` — it fails and names both by the rule they violate. The secret sentinel is a greppable token rather than a realistic credential because a realistic one trips gitleaks. Smoke-only in `run-asserters.py`: the corpus is committed, the run tree is not, since producing one requires an agent run. CI runs it against an empty dir and requires a clean non-zero exit as a bitrot guard. Gates: lint:frontmatter 7 phase files in both trees, shared:check, model-id lint, fixtures:check (93 json, 9 asserters), fixtures:assert (4 golden, 5 smoke) all green. drift:check clean for azure-to-aws; the two pre-existing gcp/heroku SKILL.md failures from the telemetry work are untouched and still outstanding.
…ive defects it found
The corpus already existed; this actually runs it. Extraction was performed by hand
against `extract-terraform.md` + `arm-type-canonicalization.md`, the resulting
27-resource inventory is committed as a golden tree, and the asserter is promoted from
smoke-only to a GOLDEN CI gate.
Read this as a consistency test, not a capability test. Whoever authored the rules also
performed the extraction, so it proves the rules, the corpus and the oracle agree — not
that the skill teaches a model that has never seen them. It still found five real
defects, three of which would have made the oracle pass while testing nothing.
**1. Identity was not unique.** The asserter indexed resources by `tf_resource_name`.
Terraform namespaces local names PER TYPE, and this corpus collides on five of them:
`core`, `storefront`, `reporting`, `data`, `store`. So five of twenty-one type
assertions were silently pointed at whichever resource happened to be written last,
passing or failing by accident of iteration order. Identity is now the full Terraform
address (`azurerm_subnet.data`), recorded as `config.tf_address`, with the asserter
failing loudly on a missing or duplicate one. Six previously-unassertable resources
became assertable as a side effect.
**2. The horizontal-resource-group case was not actually expressed.** The corpus claimed
to exercise the merge of an app in `rg-app` with its database in `rg-data`, but the app
never referenced the database — `DATABASE_URL` was a literal. The only edges crossing a
group boundary were subnet ones, which are a far weaker signal and are not what the case
is for. The app now interpolates the server's `fqdn`, the edge taxonomy gained the
`data_ref` type it had no home for, and the assertion is a specific directed pair
(`storefront` → `store`) rather than "some edge crosses a boundary", which a broken
extractor could satisfy with the subnet edges alone.
**3. Child resources were being nulled out.** The rules said an absent
`resource_group_name` means `resource_group: null` plus a warning. Most child types
carry no `resource_group_name` — storage shares and containers, SQL databases, Event
Hubs, Service Bus queues, Cosmos databases — and reference their parent instead. Since
a resource with no group cannot be clustered, that rule collapsed the cluster seed for
most of a real estate. Children now inherit the parent's group via the parent reference,
and the fixture asserts it on the storage share.
**4. Redaction-in-place escaped the secret check.** Asserting on the sentinel's text
missed a truncated-sentinel-plus-`***` mutation, which the rules forbid but a substring
check cannot see. The oracle now asserts the VALUES CONTAINER does not exist at all —
no `app_settings`, `connection_string`, `value`, `admin_password`, or access-key field
on any resource's config.
**5. Nothing asserted the child resource group.** Covered by 3.
Verified by mutation: 18 realistic failures were injected into the passing inventory.
Sixteen were caught first time; the two that escaped are defects 4 and 5, and both are
now caught. Every mutation is named by the rule it violates rather than by a diff, so a
failure tells you which decision was made wrong: the function app typed as
`Microsoft.Web/functionApps`, `kind` dropped, Redis lower-cased, the plan camelCased, a
resource-group id given a `/providers/` segment, an invented subscription GUID, one and
then all five `hosted_on` edges dropped, the app-to-database edge dropped, `private_link`
dropped, a secret value recorded and then redacted in place, an unknown type guessed
rather than reported, a module silently omitted, `enabled_protocol` and `kafka_enabled`
dropped, a null child group, and `tf_address` absent.
One finding NOT fixed, recorded for review: nearly every `azure_id` here contains a
`tf:<local>` name segment, because every resource in the corpus names itself with a
`${var.prefix}` expression. That is honest and stable across runs, but it means
`azure_id` is not a real ARM ID for most Terraform-sourced resources and cannot be
matched against a live capture. `metadata.subscription_id_source` records the
subscription half of the problem; the name half has no equivalent flag yet. A
per-resource `azure_id_synthetic` marker is the obvious fix and belongs with the live
`az` path, where the drift comparison it protects actually happens.
Gates: lint:frontmatter 7 phase files both trees, shared:check, model-id lint,
fixtures:check (94 json), fixtures:assert (5 golden, 4 smoke) green. drift:check clean
for azure-to-aws; the two pre-existing gcp/heroku SKILL.md failures are untouched.
Untracked across two sessions; §11 (resolved open questions) and §12 (sequencing change: canonicalization before discovery breadth) only exist here, so losing the file loses the reasoning behind the build order.
Build step 3 (plan §7a). Pass 1 of Design is real; pass 2 (the category rubrics) lands in step 5, and until it does the phase HALTS rather than mapping compute and databases from model priors. The disposition table is knowledge/*.json rather than design-refs prose, per plan §0 and following heroku's fast-path-addons.json. That choice buys something specific: the asserter can validate the TABLE, so a design row claiming `confidence: deterministic` is checked against a row that actually exists — the one direction that catches an improvised label. Three deltas from the plan, all flagged rather than quiet: * The 10-row Direct Mappings table becomes 17. Four rows are additions of necessity: canonicalization emits subnets and the storage-service children as their own child-typed resources, so without a row each one reaches the unknown-type policy, matches "sits in a network/data provider namespace", and STOPS the design on every real estate. §7a.3 fixed the row set before the child-type vocabulary existed. * Three rows carry a mechanical condition on one config field and are still architecture-invariant, because the condition reads a property of the resource itself: SMB/NFS share, kafka_enabled, and the Cosmos API. Owner decisions 11.4 and 11.5 say protocol is the WHOLE rubric for the first two, which means there is no rubric left — only a lookup. Same shape as gcp's google_sql_database_instance (SQL Server) row. * Decision 11.5 disturbs no row. Canonicalization already makes a file share its own resources[] entry, so a share is never inside the account's mapping unit and cannot pull the account's target anywhere. The account keeps Always -> S3 for its blob surface; the only condition it needs is account_kind == FileStorage, where there is no blob surface at all. An untranslated type is now cost-bearing BY DEFAULT and STOPs. §7a.4's three-part test cannot clear it — no SKU, no consumption row, no provider namespace — and the skill cannot demonstrate that a resource it could not name is free. Design therefore reads iac_metadata.untranslated_types as a separate input, because Discover drops those resources from resources[] entirely and iterating resources[] cannot find them. A STOP writes aws-design.json carrying everything determined plus a `halt` object, then fails its gate. Discarding the work would make the user re-run the phase to learn one missing table row, and would hide which resources were already fine. Oracle: after-design-halted/ + check_expected_design.py, registered golden. 30 injected mutations, 30 caught first pass, each naming the rule violated — including the three famous wrong answers (DynamoDB for Mongo Cosmos, EFS for an SMB share, Kinesis for a Kafka namespace) and the 5x App Service Plan fan-out. Identity is the Terraform address resolved through the run's own inventory, so the expectations survive the azure_id reconstruction change that the live `az` path will bring. Deliberately untested and said so in the fixture README: the five specialist gate rows (the corpus fires none), and clusters[]/pattern_id (step 4). Gates: lint:frontmatter 7 files both trees (and verified non-vacuous — a broken _knowledge path fails it), shared:check, fixtures:check, fixtures:assert 6 golden / 4 smoke both trees, lint:model-ids all green. drift:check still RED on the two pre-existing gcp/heroku SKILL.md files; not this branch's, not allowlisted. lint:md and fmt:check UNVERIFIED — dprint and markdownlint are not installed on this box.
A fresh-context agent ran Discover over the committed corpus and PASSED the Terraform oracle — every casing trap, the fan-in edges, the cross-RG data_ref, child resource-group inheritance, the secret boundary, all correct from the refs alone. It then diverged from the golden tree on six things no ref stated, and no existing assertion noticed any of them. The run was correct by its own instructions; the instructions were incomplete. 1. `warnings[]` was mandated by three files and DEFINED BY NONE. Absent from schema-discover-azure.md entirely: no location, no entry shape, no code vocabulary. The run invented all three and picked `module_not_discovered` where the golden says `module_not_resolved`. Now defined once, with a closed seven-code vocabulary and a required subject (azure_id or identifier) — a warning nobody can attribute to anything is noise. 2. `subscription_id_source` was specified in BOTH `metadata` (canonicalization ref) and `iac_metadata` (discover-iac, the golden, the asserter). Resolved to `iac_metadata`: only Terraform needs the ID reconstructed, so only the IaC section has anything to say about where the subscription half came from. 3. `data_ref` — described as the app-to-data edge and the thing that makes the horizontal-RG merge work — was in extract-terraform.md and NOT in schema-discover-azure.md's typed-edge list. Step 4's clustering reads the canonical list, so the edge that matters most was missing from the file that consumes it. The list is now canonical and closed. Containment stays deliberately edge-less, and now says why: an ARM azure_id CONTAINS its parent's as a literal prefix, so a share's account is derivable by truncation for every source. Observability links (a component's workspace_id) live in config, not edges — an edge implies a dependency the architecture preserves, and that one does not survive the migration. 4. The `/fileServices/default/` implicit-singleton segment was stated nowhere. The run guessed it, and guessed RIGHT, which is worse than guessing wrong: the rule looked taught when it was not, for all four storage sub-services. Now in § Reconstructing azure_id, with the note that the segment belongs to the ID and never to the type string. 5. extract-terraform.md had no per-type row for Key Vault, Log Analytics, private endpoints, or App Insights, so a correct-by-the-ref extraction dropped sku_name, sku + retention_in_days, and subresource_names — all cost- or routing-bearing. Five rows added, plus two rules the table never stated: omit rather than null, and a type with no row still gets an entry. 6. discover.md's postcondition demanded "the edges[] set that justified the grouping" while discover-assemble.md instructs an EMPTY edges[]. A live contradiction that passed only because `_assert` has no teeth — the exact false-green class this project is organised around. Fixed by naming the real reason: clusters now carry `justification`, which is `seed:resource_group` today. Without it, "grouped by the seed" and "grouped for no recorded reason" produce identical output. The oracle now enforces all of it, which is the point — this class of defect is invisible to shape assertions by construction. New `contract_vocabulary` and `required_config_fields` blocks check the warning codes, the edge types, the no-null rule, and the 15 cost/routing fields per canonical type. Run against the capability test's own output they produce 23 named failures; run against the golden, PASS. Seven further mutations injected (invented code, missing detail, unattributable warning, absent array, invented edge type, null config, dropped field) — 7/7 caught, each naming its rule. Golden inventory updated where the ref and the golden disagreed: App Insights gains application_type + workspace_id, and warnings[] is rewritten to the defined shape. The golden had been over-delivering relative to its own ref, which is the mirror image of defect 5 and equally invisible. Design gets its own closed warning vocabulary in design-infra.md, kept separate from Discover's: Discover's codes say what could not be READ, Design's say what was DECIDED, and a cost-bearing or untranslated type produces a `halt` rather than a warning because a warning means the design continued. Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check, fixtures:assert 6 golden / 4 smoke both trees, lint:model-ids green. drift:check still RED on the two pre-existing gcp/heroku SKILL.md files. lint:md and fmt:check still UNVERIFIED — dprint and markdownlint absent.
…rinting the answers A second capability run — fresh context, answer keys blocked — passed the Discover oracle including all the new contract checks, so the previous commit's hygiene held. It then failed the Design oracle in seven places, and every one was my under-specification from bd398c7 rather than the model improvising. FIRST, THE GOOD NEWS: the halt guard works under pressure. Told to map resources whose rubric file is absent, the run produced NO mapping for six of them — three App Service Plans, the function app, the Windows VM, the Postgres server — and recorded each with its missing ref_file. Unprompted, it reported being strongly tempted not to: index.md's "Typical AWS target" column printed the answers and fast-path.md's Preferred-Target table printed the tie-breaker, so a complete and plausible compute and database design was available without opening compute.md at all, and every shape assertion would have passed it. That is the whole false-green thesis, live: the guard is the only thing standing between "no rubric needed" and "the rubric was never written", and it held. But it held against a hazard I built, so: * index.md's right-hand column is no longer an answer key. Rubric rows now carry an unordered CANDIDATE SET in braces — no defaults, no preference ordering — and the column note says plainly that being able to answer from the column alone means the row is wrong. Fast-path rows keep their target, because there the answer legitimately lives in the table. SECOND: aws-design.json had no schema, and five of the seven failures were that. Every other artifact has one. The run reverse-engineered the shape from twelve postconditions and scattered prose and invented reasonable-but-different key names, while the committed golden used a third set — so the oracle was asserting hosted_app_azure_ids and sizing_source that NO skill file required. That is exactly the golden-over-delivers-vs-ref defect I diagnosed for extract-terraform.md one commit earlier, reintroduced in design-infra.md in the same session. A hand-authored golden and a prose-only contract drift by construction. references/shared/schema-design-aws.md is now the thing both have to agree with, and it makes three things REQUIRED that were merely implied: fast_path_row on every deterministic entry (so the label is auditable rather than an unverifiable claim), and hosted_app_azure_ids + sizing_source on every compute unit — without which "five apps correctly fanned in" and "four fanned in, one silently dropped" produce identical artifacts. THIRD: two contradictions the run walked straight into. * index.md gave Microsoft.Web/sites two rows: one saying "none of its own — fans IN to its plan", the next giving kind=functionapp its own Lambda/Fargate target. It followed the second and emitted a standalone function-app entry, contradicting design.md's own postcondition. A Y1 consumption plan is exactly where the two readings collide. Resolved in favour of the plan: a site NEVER gets an entry, and a hosted site's kind narrows the plan's candidate set instead. * design.md's cluster postcondition demanded target_architecture while patterns.md does not exist, so it was unsatisfiable. The run wrote "UNDETERMINED" and flagged that a green-seeking agent writes a plausible architecture string there instead, which nothing downstream could catch. Fixed the same way cluster `justification` was: pattern_status now distinguishes `recognized` from `unclassified` (catalog consulted, nothing matched) from `catalog_absent` (never attempted), and target_architecture MUST be null unless recognized. The honest state is now expressible AND asserted. Also: postcondition 7 predated pending_rubric[] and so a correctly-halted design could not pass it — pending_rubric is permanent contract, not scaffolding, since adding an index.md row without its file can recur. And halt.blocking now requires one missing_rubric_file entry per distinct missing ref, because pending_rubric[] with no halt reads as complete to anything that only inspects services[]. Smaller: name_expression_unresolved is now one entry per run listing the affected addresses. Per-resource it produced 21 of 25 Discover warnings and buried the four actionable ones — the ref already granted that courtesy to subscription_id_unresolved and not to this. Oracle: 3 new assertion groups (cluster pattern_status, compute-unit required fields, halt-covers-pending). 7 mutations injected — 7/7 caught, including the two that matter most: clusters given a plausible recognized architecture, and a halt naming compute.md but not database.md. Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check, fixtures:assert 6 golden / 4 smoke both trees, lint:model-ids green. drift:check still RED on the two pre-existing gcp/heroku SKILL.md files. lint:md and fmt:check still UNVERIFIED.
The plan existed in two places and forked: §11 and §12 landed in the repo copy on 2026-09-04 while the Pippin artifact stayed at its 2026-09-03 content, so the mirror still described a build order §12 inverts and an Azure SQL disposition §11.3 overrules. Nothing enforces the sync, so the header says which way it flows. Pippin is now published from this file (v4).
Keeps the handoff in the repo alongside the plan, for the same reason: the Pippin copy forked once and nothing enforces the sync. Published to Pippin from this file. Corrects the previous handoff's claim that bedrock-quotas.md has zero inbound references — it has two real load references (gcp design-ai.md:192 and estimate-ai.md:108), and acting on the old claim would have deleted a live file.
…d inventory
Two things: the pass-2 rubrics the corpus tests, and the inventory/cluster work
that has to be right before a rubric reading them means anything.
INVENTORY — three gaps that only show up on a real repo
* Canonicalization coverage 77 -> 137 azurerm types. The old set was sized for
the synthetic corpus; a real repo routinely carries azurerm_firewall,
azurerm_route_table, azurerm_virtual_network_peering, azurerm_bastion_host,
azurerm_app_configuration, azurerm_key_vault_key and more, each of which was
silently dropped as untranslated. Also added § Coverage is not completeness,
because 137 is still not the provider surface and the honest framing matters:
the STOP is the mechanism, one table row is the fix.
* `.terraform/modules/` is now READ, and this is the biggest practical fix in
the commit. The blanket `.terraform/` ban was written to keep state files out
and was correct for state — but `modules/` holds nothing except downloaded
module SOURCE, exactly as safe as the local module path we already recurse
into. Modern Azure Terraform leans hard on Azure Verified Modules, so under
the old ban a repo could declare almost its whole estate through registry
modules and get back a nearly empty inventory plus a few warnings, with every
downstream phase then reasoning confidently about a fraction of the estate.
Still banned: any *.tfstate, anywhere.
* Association-only resources now have a rule. `azurerm_subnet_*_association`
and friends exist only in Terraform and have no ARM type at all. They were
being reported as untranslated — i.e. as gaps in this skill — which is wrong
and, worse, buries the real gaps: on an IaC-heavy repo the associations
outnumber the genuinely-missing types. They emit an EDGE and no entry.
CLUSTERING — real, not a seed
references/clustering/{clustering-algorithm,typed-edges-strategy,classification-rules,tiering}.md.
Seed per resource group, split a candidate whose members have no edges between
them, merge candidates joined by a crossing edge, then tier + primary + roles.
The load-bearing judgement is that two edge types are AMBIENT and must not
merge: `network` (everything in a VNet shares subnets) and `secret_ref` (one
Key Vault serves the estate). Merging on either collapses the whole estate into
one cluster and destroys the partition. They are still recorded and still count
for internal connectivity in the split step — sharing a subnet is weak evidence
that two resources are related and good evidence that two already-related ones
belong together, and that asymmetry is the point.
Rejected a weight-and-threshold scheme: the threshold would have no defensible
source, would need re-tuning per estate shape, and its failures would be silent.
Binary merges/does-not per edge type is explainable in one sentence per row.
RUBRICS
compute.md — eliminators (App Runner is never a candidate; Lambda's 15-minute
ceiling, Durable Functions, Windows containers on EB, GPU on Fargate), the six
criteria in order, the fan-in, x86_64, and right-sizing as post-selection.
database.md — the availability override gate first, because it is the thing that
decides RDS vs Aurora and it is NOT inferable from configuration. When the
answer is absent the default is RDS single-AZ, explicitly not Aurora: a source
`high_availability { ZoneRedundant }` says what they bought, not what they need,
and inferring Aurora from silence inflates the estimate with no visible cause.
The source HA posture surfaces as a FINDING instead
(`availability_downgrade_from_source`), which is a genuinely useful thing to
tell someone.
FIXTURE — the untranslated-type case was fragile and is now durable
Broadening the canonicalization table added azurerm_dev_test_lab, which was the
corpus's untranslated-type fixture — so a coverage pass silently disarmed the
STOP test. Any fixture that depends on a type being ABSENT has that failure
mode. Swapped to azurerm_iothub: out of this skill's scope by design, so a
coverage pass will not absorb it, and it carries a real sku block so the corpus
comment ("it carries a sku") is now literally true rather than aspirational.
The golden design is regenerated with both passes applied: 14 services[], zero
pending_rubric[], one halt blocker instead of three. The pending_rubric
assertion INVERTED — at step 3 a resource mapped past a missing rubric was
improvisation; now a resource still parked as pending is a run that did not load
rubric files that exist.
Oracle: 4 new assertion groups (rubric outcomes with their tempting wrong
answers, the x86_64 default, the source-HA finding, empty pending_rubric).
12 mutations injected, 12 caught — including Postgres -> Aurora inferred from
the source's own HA setting, which is the single most valuable assertion here.
One oracle bug found and fixed by its own golden: the architecture check
substring-matched the whole aws_config, so a `graviton_reason` field explaining
why Graviton was NOT chosen tripped it. Narrowed to architecture-bearing keys —
a blunt match would have pushed authors to stop explaining themselves, which is
backwards, since that explanation is what stops the x86 default reading as a bug.
Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check,
fixtures:assert 6 golden / 4 smoke both trees, lint:model-ids green.
drift:check still RED on the two pre-existing gcp/heroku SKILL.md files.
lint:md and fmt:check still UNVERIFIED.
A third capability run — fresh context, answer keys blocked — confirmed the three inventory fixes work end to end and then found that ONE of them had broken Design. Both regressions were mine, from the previous commit. WHAT WORKED (verified from a clean context, quoting the instruction it followed) * Registry modules: read `.terraform/modules/modules.json`, resolved the key to a directory, discovered both module resources, tagged them `tf_module`. The `module.cdn` case whose source is NOT on disk still warned. The distinction held. * Association-only resources: no entry, no untranslated warning, one `network` edge. * Every halt honoured. On the route table and bastion host it reported being "strongly tempted" — warning-and-continuing would have made the gate 1-of-12 instead of 2-of-12 and looked much better — and stopped anyway. * The database went to RDS single-AZ with no availability answer, and surfaced the source's ZoneRedundant HA as a finding rather than reading it as the answer. That is the exact behaviour database.md §1 was written to produce. REGRESSION 1 — the coverage pass left 53 canonical types with no disposition Taking the canonicalization table 77 -> 137 types made the INVENTORY better and DESIGN worse: each newly-translatable type is now discovered, matches no row, hits the cost-bearing provider-namespace clause, and STOPS. On this run a route table and a bastion host — both free-or-trivial in reality — halted the whole design. Every oracle stayed green throughout, because the corpus contains a dozen types out of 137. So the fix is not just the 53 rows (30 direct / 52 skip / 10 gate now) but the invariant: check_expected_design.py reads both files and asserts ZERO orphans. That check would have failed the moment the coverage pass landed. Two dispositions worth flagging as judgement calls: Azure Bastion -> Session Manager (no host at all, and free — an EC2 bastion would add a permanent cost line the target does not have), and Traffic Manager -> a Route 53 routing policy rather than a new service. Logic Apps, Batch, ML workspaces and Stream Analytics went to specialist gates rather than being mapped, because in each case the work is a rewrite that the resource does not describe. REGRESSION 2 — clustering over-fragmented, and my own worked example proved it The run produced 16 clusters where the file's example predicted 3. Cause: I had deliberately made containment NOT an edge (an ARM id contains its parent's, so it is derivable) and then wrote a split step that only looked at edges — so a VNet split from its subnets, a storage account from its share, and every edgeless observability resource became its own "workload". The example was written by hand instead of by tracing the algorithm, so the two disagreed and the algorithm shipped. Two fixes: containment counts as connectivity in the split step (derived from the id prefix, still not stored as an edge), and a split only happens when two or more components each contain a PRIMARY-ELIGIBLE resource — a component with nothing that could be its primary is a fragment, not a workload, and attaches to the largest component. Re-traced the example against the algorithm: 4 clusters. Also `split:*` no longer requires a non-empty `edges[]`. It cannot have one — its justification is an ABSENCE — and the old rule made a correct split unrepresentable, so the run had to mislabel nine clusters `seed:resource_group` to pass the gate, destroying the information the field exists to carry. THE CORPUS NOW GATES THE FIXES PERMANENTLY The module and association cases were tested once by a throwaway workspace. Both are now in the committed corpus: `.terraform/modules/naming/` with two resources, a `module "naming"` block, and an `azurerm_subnet_route_table_association`. 29 golden resources from 31 declared. Verified `.terraform/` is not gitignored here, or the fixture would have been silently empty. Oracle: 5 new assertion groups. 10 mutations, 10 caught — including the blanket `.terraform` ban being restored, module.naming warned as unresolved, the association given an entry, the association's edge dropped, and the bastion mapped to an EC2 instance. Known and NOT fixed: the coverage check verifies a disposition EXISTS, not that it is correct. A wrong row for a type absent from the corpus escapes, and one injected mutation demonstrates that. Closing it needs corpus breadth, not a cleverer asserter. Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check, fixtures:assert 6 golden / 4 smoke both trees, lint:model-ids green. drift:check still RED on the two pre-existing gcp/heroku SKILL.md files. lint:md and fmt:check still UNVERIFIED.
Build step 5, part 1 — the phase that was blocking everything. Clarify ran and then
failed four of its own postconditions because only clarify-global.md existed.
FOUR FRAGMENTS, following gcp's category model
clarify-compute.md compute target, plan isolation, CPU architecture, traffic,
long-lived connections, VM cutover (ESSENTIAL)
clarify-database.md availability (the override gate's input), DB cutover
(ESSENTIAL), traffic, storage I/O, Cosmos read/write split,
Redis modules
clarify-licensing.md CONDITIONAL — Windows model (ESSENTIAL), SQL model, AHUB
detection, the Azure Edition hard blocker
clarify-identity.md always fires, one shallow row
ONE STRUCTURAL CORRECTION toward gcp's flow
clarify.md said each fragment "presents its sheet section", which would give the user
FIVE sheets and five interleaved rounds of essentials. gcp runs ONE sheet as a single
mandatory gate, then batches the essentials — and this phase's own postcondition says
"every assumption-sheet row the user was shown", singular.
So fragments now compute rows and ask nothing; the assembler owns the whole
conversation in three gates: the consolidated sheet (batched five at a time, with each
fragment's consequence line), then the ESSENTIAL questions with their context, then the
recap. One place knows the full row set.
THE TWO RULES THAT CARRY THE WEIGHT
`ESSENTIAL` + `value: null` IS the completion gate. An essential row has no default on
purpose and the phase must not complete while one is unanswered. It is the only way the
contract can say "shown and not answered".
A value taken from its default STAYS `PROPOSED`. `DETECTED` means read from the estate.
Design's rationale prints "you chose Elastic Beanstalk" differently from "we assumed
Elastic Beanstalk" — but only if this file recorded which happened.
Three rows have no default at all, because they select different RUNBOOKS rather than
different numbers: VM cutover (MGN vs rebuild), DB cutover (DMS vs dump/restore), and
the Cosmos read/write split (which moves the DynamoDB conversion by multiples).
`data.availability` becomes ESSENTIAL when the source is zone-redundant, rather than
PROPOSED. Silently downgrading resilience someone pays for today is the expensive
mistake, and reading the source's own HA setting as the answer is the other one — it
says what they bought, not what they need.
CLUSTERING WAS REAL AND COMPLETELY UNGATED
Found while wiring preferences.clusters[], which keys on cluster_id: there was no
golden clusters artifact at all, so the split/merge rules were reviewed prose and
nothing more. The design golden also still carried the old seed-only 3 clusters, not
the 4 the re-traced algorithm produces — so the two disagreed and nothing noticed.
after-discover/azure-resource-clusters.json now exists: 4 clusters, 29 members,
justifications merge:cross_group_edges and split:no_internal_edges. The asserter pins
the merge (rg-app + rg-data joined only by app-to-data edges — seed-only clustering
would cost the app without its database), the split inside one resource group, the
legal single-member idle-plan cluster, that a split MAY have empty edges[] while a
merge may not, and that a Microsoft.Web/sites is never a cluster primary.
THE CLARIFY TESTING GAP, ADDRESSED RATHER THAN DEFERRED
Clarify is _interactive: true, so the capability-test method cannot reach it — a
dispatched agent has no user. clarify-answers.json stands in for one, which makes the
sheet's BRANCHING testable. Stated plainly in both the asserter and the fixture: this
tests nothing about wording, batching or tone. Those need a human. It tests which rows
fire, which are ESSENTIAL vs PROPOSED, and which are N/A — which is where the defects
are, because a firing rule is a judgement and prose cannot check itself.
The golden is a BLOCKED clarify on purpose: the scripted user declines to state Azure
spend, so an ESSENTIAL row is unanswered and the phase must GATE_FAIL. Pinning the gate
matters more than pinning the happy path.
15 mutations, 15 caught, including: availability PROPOSED instead of ESSENTIAL;
availability DETECTED from the source's own HA; the Cosmos RU question asked for a
Mongo-API account; isolation asked for the 1-app plan; licensing N/A despite a Windows
VM; identity N/A because no managed identities were found; spend invented to get past
the gate; and clarify_status COMPLETED with an essential row still null.
The asserter also caught an omission in my own golden — `ahub_in_use` was DETECTED with
no stated source, which is the promoted-default signature. It was genuinely read from
the estate; I just had not recorded what was read. AHUB matters because it does NOT
travel: if in use, the Azure baseline must be compared at the UNREDUCED rate or AWS
looks worse than it is.
Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check (101 json, 11
asserters), fixtures:assert 7 golden / 4 smoke both trees, lint:model-ids green.
drift:check still RED on the two pre-existing gcp/heroku SKILL.md files. lint:md and
fmt:check still UNVERIFIED.
Prepares a session handoff, and brings the authoritative plan up to date — it had not been touched since §11/§12 on 2026-09-04, so eight commits of decisions lived only in commit messages. Plan additions: §13 — the 21 decisions taken 2026-09-05/06, grouped by what they touch, with a reason per row. Includes the three corrections to earlier claims: 11.5 disturbs no Direct Mappings row, §7a.3's 10-row table is superseded by 30, and bedrock-quotas.md is NOT dead (two real load references — the earlier claim came from a grep that ate its own evidence). §14 — the sequencing change. Terraform end to end BEFORE source breadth, superseding both the original build order and §12's amendment. The previous order optimised for discovery breadth and produced a skill that discovers five ways and cannot produce an answer. Records the counter-argument too, because it is real: most startups have no azurerm_* Terraform, and SKILL.md commits to live-first — so this trades reach for completeness on purpose. §15 — how the skill is actually tested, which is the most transferable part. The extensional/intensional distinction and why only building the first left every contract disagreement invisible; the capability-test method and what each of the three runs found; why two of the three goldens are FAILING states. Handoff at .agents/scratchpad/azure-to-aws-handoff-2026-09-06.md, kept in the repo for the same reason as the plan: the previous handoff existed only in Pippin and drifted unnoticed until the plan had forked by 8k characters.
Verified against the design golden rather than asserted from memory: the corpus routes ZERO resources to a missing rubric file. Design already emits 16 services[] with pending_rubric[] empty, and halts on exactly one thing — the untranslated azurerm_iothub. So networking.md + messaging.md are needed before any REAL customer repo and block nothing on the fixture; they were wrongly placed next. Estimate and Generate are on the critical path for both definitions of the goal, networking.md for only one. And those two rubrics need corpus EXTENSION to be testable at all, since nothing in the fixture reaches them. This also promotes plan §13.7 #1 from an open question to THE gate on the end-to-end goal: estimate.md carries _check_phase_completed: design and Design GATE_FAILs on the corpus, so neither Estimate nor Generate can run there until the untranslated-type halt is resolved. Both paths written up, with the user override recommended and the STOP kept as the default.
…ring The summary box and §1 still named networking.md as next after §8 was corrected. Also distinguishes the two halt situations, which is the thing that was conflated: on the CORPUS the only blocker is the untranslated type (pending_rubric is empty, 16 services mapped); on an arbitrary REAL repo the 7 missing rubric files would also halt.
Rule 2 of arm-type-canonicalization.md claimed ARM type strings are compared case-sensitively, and named two casing "traps": a capital R in Microsoft.Cache/Redis and all-lowercase Microsoft.Web/serverfarms. Both were unsourced, and the second is unsourceable — Azure/bicep-types-az ships BOTH Microsoft.Web/serverFarms and Microsoft.Web/serverfarms in one generated index. For the cache type, that index and magodo/aztft (the mapping library behind Microsoft's supported Azure/aztfexport) both render Microsoft.Cache/redis. The capital-R form is what appears in azurerm RESOURCE IDs — a Terraform-provider artifact, not an ARM type. Nine of this file's 124 externally checkable rows differ from aztft by casing alone, so the discipline it claimed to enforce was wrong about 7% of its own content. Worse, the rule made a mis-cased type "silently fall through to the unknown-type policy", so a run emitting the CORRECT Microsoft.Cache/redis would have had its Redis treated as untranslated and STOPped the design. And expected-iac-terraform.json listed both correct forms under forbidden_types, failing them as "the signature of a guessed translation". Separate the two things the file conflated: - MATCHING folds case, everywhere (fast-path, Skip Mappings, index.md routing, rubric selection). A case-only difference must never reach the unknown-type policy. - EMISSION follows this file's spelling, as a stated CONVENTION justified by azure_id strings being joined by exact match — cluster membership, cluster keys, and the future drift comparison against a live capture all need one resource to yield one string. Explicitly not a claim about ARM. Deletes the Redis trap, reframes the serverfarms trap to keep its real content (serverFarmId is the property pointing AT the plan, not the plan's type name), drops four inline casing assertions that are all among the nine unsourceable rows, and records the evidence in a new section so it does not get "fixed" back. No golden churn and no strictness lost: the ~74 existing occurrences of the two display strings are correct under the convention and are untouched, and all three azure oracles still pass. Mutation-tested — sites->functionApps (8 fails), Redis->redis (2 fails, now reported as a convention violation rather than a guessed translation), DocumentDB->CosmosDB (2 fails). Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check, lint:model-ids, fixtures:assert (11 asserters) all green both trees. drift:check still red on only the two pre-existing gcp/heroku SKILL.md files. lint:md and fmt:check remain UNVERIFIED — dprint and markdownlint are absent.
…r branch
This branch is cut from main, where `hooks/telemetry/` does not exist — it is
created by feat/telemetry-hooks. So azure's SKILL.md declared PostToolUse and
Stop hooks invoking `${CLAUDE_PLUGIN_ROOT}/hooks/telemetry/emit.mjs`, a path
that cannot resolve, and carried a 60-line consent step driving a CLI that is
not present. A hook command pointing at a missing file is exactly the defect
that makes the two pre-existing gcp/heroku drift failures unfixable on the
telemetry branch; shipping azure with the same shape would add a third.
Removed here, not redesigned:
- the `hooks:` frontmatter block and its SessionEnd note (advisor tree; the
migrate copy never had one, since migrate/plugins/migration-to-aws ships no
hooks/ directory at all)
- the Telemetry consent section from SKILL.md § Execution, including the
$EMIT resolution ladder
- the `AZURE_TO_AWS` SKILL_INVENTORY entry and the Cursor hook registration,
which were hunks of 317c3ac against files that do not exist on main
- `.cursor-plugin/marketplace.json`, which was created by the telemetry branch
(2d09d36) rather than by main, so it is a telemetry artifact and not azure's
Plan §3a is annotated as DEFERRED rather than deleted: the design stands, and it
lands as a follow-up on feat/telemetry-hooks (or after it merges), together with
the two emit.mjs defects §3a documents and this branch no longer fixes.
Consequence to be explicit about: azure-to-aws currently emits NO telemetry.
Plan §3a calls that a shipping requirement, so this is a known gap on this
branch, deliberately taken to keep azure independently mergeable from main.
Rebasing 940aadd onto current main left `skills/shared/ai/` stale. That commit created the canonical copy as a pure ADD while MOVING the gcp copy to `references/vendored/ai/`, so git could follow the rename and carry main's subsequent edits into the vendored path, but had nothing to follow for the canonical one. The canonical file therefore kept the pre-rebase content. Only `ai-migration-guardrails.md` was actually affected, in both trees: main's 8c65c92 replaced the "shared 10,000 RPM account limit" risk table with the GPT-5.6 per-model TPM-only quotas (no RPM quota, cached input tokens excluded), and the canonical copy still carried the old table. `shared:check` caught it, which is the whole point of that gate — a stale canonical is worse than a stale vendored copy, because `shared:sync` would have propagated it outward and silently reverted main's facts in every consuming skill. The remaining seven promoted files were already identical. shared:check and drift:check are both green: 334 identical, 28 allowlisted.
…and defects Six things the plan did not carry, and the handoff either omitted or got wrong. §16.1 The Rule 2 casing correction, with the evidence, so it is not reverted: bicep-types-az ships BOTH serverFarms and serverfarms in one generated index, and both it and aztft render Microsoft.Cache/redis. Nine of 124 checkable rows differed by case alone. Records the two live consequences — a correct answer STOPped Design, and the oracle failed correct answers as "the signature of a guessed translation". §16.2 magodo/aztft as external prior art. It validates our mapping semantics on 123 of 124 rows, and settles the "1 to N mapping" question empirically: 1048 of 1089 azurerm types resolve to exactly ONE ARM type and ZERO resolve to more than one, so Terraform→ARM is a function and there is nothing for a model to choose. Also records that aztft does NOT cover 13.2d — 22 association entries, none flagged property-not-resource — and that it was deliberately NOT adopted as an intensional check. §16.3 The `confirm` phase, replacing §13.7 #1's open question and both paths the handoff proposed. A phase rather than post-Discover afterwork because resolving an untranslated type ADDS a resource and Clarify's fragment triggers read the inventory. Option (b), amend in place. Design's "untranslated_types is empty" postcondition passes UNCHANGED, so the STOP stays absolute. Carries the obligation that would otherwise have been silent: amending the inventory obliges re-deriving the clusters, because 64 azure_id occurrences live there in four roles. §16.4 Corrects the handoff's "ONE owner decision": TWO gates block the corpus, since after-clarify is BLOCKED_ON_ESSENTIAL independently of the Design halt. §16.5 Why TF→ARM→TF is not a lossy double translation — config.tf_address retains the azurerm type, Generate never reads Azure HCL, and the real loss is the config attribute allowlist, which is independent of ARM. §16.6 Four defects found and not fixed: azapi_resource unhandled, config.tf_file empty on 24 of 29 golden entries against a ref that mandates it, the missing clusters→inventory pairing check, and the canonical-staleness shape where a promoted file created as an ADD misses upstream edits its vendored twin receives. §16.7 The branch restructure, and the fact that drift:check is now GREEN — the blocker handoff §7 called four sessions old was caused by telemetry landing advisor-only, and basing on main removes it. Also states plainly that azure-to-aws currently emits no telemetry.
… AzAPI discovery Step 5b's three rubric files, and the one Discover gap that made real repos undiscoverable. Both trees. ## The three rubrics Each follows compute.md/database.md's anatomy: eliminators, the six criteria in order with first-match-wins, feature parity, cluster context, post-selection sizing, output contract, status. Confidence is `inferred`, never `deterministic`. App Runner appears in networking.md's eliminator table as ALWAYS eliminated. The content is not a mapping table with prose around it. Each file carries the trap that makes its category non-obvious: - **networking.md** — the DOUBLE-BALANCER trap. An Elastic Beanstalk load-balanced environment provisions its own ALB, so mapping an App Gateway whose backend pool is a web app to a *second* ALB double-counts the balancer. This is the networking analogue of the App Service Plan fan-in and it fails the same way: quietly, and only in the estimate. Also: Azure NAT Gateway is REGIONAL and AWS NAT Gateway is ZONAL, so one Azure resource becomes N AWS resources where N is the AZ count from preferences.json — an estimate that prices one NAT for a multi-AZ design is wrong in the flattering direction. And APIM's policy XML has no single AWS counterpart; the service maps, the policy layer is a rewrite. - **messaging.md** — Service Bus is one broker; AWS splits its capabilities across SQS, SNS, EventBridge and Amazon MQ, so a NAMESPACE maps to a SET derived from its children and carries no cost line of its own unless the answer is Amazon MQ. Eight eliminators are all readable from IaC (256 KB SQS ceiling vs Premium's 100 MB, the 5-minute FIFO dedup window vs Service Bus's 7 days, 15-minute DelaySeconds vs scheduled messages, AMQP/JMS clients). Event Hubs consumer groups are free on Azure and priced on Kinesis enhanced fan-out — named, because nothing else would surface it. - **analytics.md** — for AI Search, the index maps and the PIPELINE THAT FILLS IT does not: indexers and skillsets have no OpenSearch equivalent and become explicit ingestion infrastructure the source did not have. For Databricks the like-for-like default is Databricks on AWS, with EMR only when the workspace is confirmed plain Spark — EMR is Spark, not Databricks with a different bill. Says explicitly that Synapse/ADF/Stream Analytics/ML workspaces are specialist gates and must NOT get rubric rows here, since a gate may never be overridden. Corrected while writing: messaging.md initially claimed Event Grid types are Skip Mappings. They are DIRECT MAPPINGS to EventBridge. The status section now says so and explains why a rubric row for them would be unreachable. ## AzAPI `azapi_resource` had zero mentions anywhere in the skill (plan §16.6), yet it is common in modern Azure Terraform and its `type` argument states the canonical ARM type verbatim — `Microsoft.Consumption/budgets@2023-05-01`. So the easiest extraction in the skill was the one that halted. `extract-terraform.md` § Step 2a now covers it: split on `@`, use the left half AS GIVEN without consulting the canonicalization table, keep the API version in config, derive azure_id from parent_id, read `body` for routing attributes only and never for values, and stamp `config.declared_via: "azapi"`. `azapi_update_resource` and `azapi_resource_action` emit no entry — they mutate a resource declared elsewhere, the AzAPI analogue of the association-only class. And a cost-bearing AzAPI type with no DISPOSITION still STOPs: AzAPI clears the type-vocabulary problem, not the unknown-type policy. discover-iac.md now states that azapi is part of the terraform dialect, not a fourth one, so the halt guard does not misreport an azapi-only file as unreadable. ## New intensional check `check_index_reference_files_exist` — every .md named in an index.md ROUTING ROW must be on disk or declared in `known_pending` with a reason. This was on the handoff's list of unguarded pairs: index.md carries a HALT, so a deliberately pending rubric and a typo'd filename were indistinguishable at runtime, and a typo would have been diagnosed as "write that rubric" when the rubric already existed. Checked in BOTH directions — a known_pending entry that now EXISTS also fails, so the list cannot rot the way a fixture depending on an absent type does. Only `ai.md` is pending, and it belongs to deferred step 6. Mutation-tested: a typo'd routing row fails naming the file; a stale known_pending entry fails naming it. ## Effect Of five realistic example estates, Design now completes on three (was one). `legacy-lift` and `data-platform` newly clear. The two that still halt do so ONLY on untranslated types — azurerm_maps_account and azurerm_notification_hub_namespace — which are deliberately left absent as the `confirm`-phase test cases. Adding rows for them would silently disarm the only examples that exercise the STOP. Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check, lint:model-ids green; drift:check OK (337 identical, 28 allowlisted); fixtures:assert PASS both trees. lint:md and fmt:check remain UNVERIFIED — dprint and markdownlint absent.
… were wrong
Answering "for services we cannot map, are we not relying on the LLM's internal
knowledge?" — yes, and the largest instance of it was sizing.
Five of the six sizing tables the rubrics NAMED did not exist; only
fast-path-services.json was on disk. So every instance size in the committed
golden came from pretraining: S1 -> t3.medium, Standard_D4s_v5 -> m6i.xlarge,
GP_Standard_D2s_v3 -> db.t4g.medium. No skill file contained any of those
mappings.
Three compounding problems, all now closed:
1. The instruction was UNSATISFIABLE. compute.md, database.md and analytics.md
each said sizing "states the dev-tier default and says the table is absent
rather than inventing a number". With no table, the dev-tier default WAS the
invented number. The rule forbade exactly what it required.
2. `sizing_source` made an unsourced number look sourced. It records the Azure
INPUT ({"sku_name": "S1", "worker_count": 2}) and says nothing about where
t3.medium came from, so a reviewer sees a provenance field and reasonably
assumes provenance.
3. The oracle pinned NO sizes, so two runs could disagree on every number and
both stay green — section 12's false green in the one dimension Estimate is
built on.
## Eight tables, auditable by construction
Every row records vCPU and memory for BOTH sides, so a reviewer can check the
arithmetic without trusting the author. A row whose AWS side is smaller than its
Azure side on either axis is a bug by inspection.
appservice-eb-sizing.json 27 plan SKUs, incl. Y1/FC1/EP* routing to Lambda
vm-ec2-sizing.json 15 family classes + 24 explicit sizes
flexible-server-rds-sizing.json 14 SKUs, Graviton default with x86 per row
cosmos-dynamodb-conversion.json RU -> ops/s -> WCU/RCU, with the multipliers
azure-region-map.json 44 regions, same_country flagged per row
disk-ebs-sizing.json tier map + per-volume IOPS breakpoints
aks-eks-sizing.json only what DIFFERS between AKS and EKS
rightsizing-thresholds.json P95 bands + the aggressiveness slider
Wired into design.md's _knowledge with _when guards so a run loads only the
tables its inventory reaches. lint:frontmatter validates those paths, so a typo
fails the gate.
## Two of the three golden sizes were wrong
Writing the tables immediately contradicted the golden:
S1 (1 vCPU / 1.75 GiB) t3.medium (2/4) -> t3.small (2/2)
GP_Standard_D2s_v3 (2 vCPU / 8 GiB) db.t4g.medium (2/4) -> db.m6g.large (2/8)
Standard_D4s_v5 (4 vCPU / 16 GiB) m6i.xlarge (4/16) unchanged, correct
The RDS one shipped at HALF the source memory. The EB one at double. The third
was right — which is the point: an unasserted number is not the same as a
correct one, and nothing distinguished them.
## sizing_provenance
New REQUIRED field on any services[] entry whose aws_config carries a size:
`table` | `measured` | `user_stated` | `model_prior`.
`model_prior` is a legal value and must be used honestly — writing `table` for a
row that was never consulted is the failure the field exists to prevent, and the
only one a reviewer cannot detect from the artifact alone. A model_prior entry
SHOULD also carry a warnings[] entry naming the type with no sizing row, so the
gap is visible in the report and fixable in one place.
analytics.md still has no sizing table, deliberately, so it now says outright
that its node counts are `model_prior` rather than implying otherwise.
Also corrected in passing: plan section 0 records the EBS breakpoint as "gp3 to
80K IOPS". A single gp3 volume caps at 16,000 IOPS; 80,000 is an instance-level
aggregate. disk-ebs-sizing.json carries the per-volume limits, which is what a
disk maps to.
## Checks
check_sizing_provenance pins the three corrected sizes and requires the
provenance enum on every sized entry. Mutation-tested four ways: reverting either
wrong size fails naming the table row and the source SKU; dropping the field
fails; an invented enum value fails.
Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check,
lint:model-ids green; drift:check OK (345 identical); fixtures:assert PASS both
trees; all three azure oracles PASS both trees. lint:md and fmt:check remain
UNVERIFIED — dprint and markdownlint absent.
… a STOP "Should we not use the LLM to resolve what the json does not carry? We don't have such limitations in gcp-to-aws." Half right, and the half that was wrong is the more useful half. gcp-to-aws does NOT have broader knowledge. It lists 28 google_* types total (22 in index.md, 19 in fast-path.md) against azure's 65 index.md rows and 92 fast-path rows. What it has is a FALLBACK before the stop, at gcp design-infra.md:62-65: substring-match the type NAME to a category, run that category's rubric, and only STOP if no pattern matches. So azure was STRICTER than gcp on more than double the per-type coverage: a type we could name, in a namespace we understood, still halted Design. That is the real defect. But gcp's fallback is not the one to copy. Matching `"log" -> monitoring` against a type name also matches google_diaLOGflow_agent, which is a silent miscategorisation into the wrong rubric. Azure has a better signal for free: the provider namespace is a structural, authoritative segment of an ARM type string, not a substring guess — and the cost-bearing test already reads it. ## Two derived rules, ahead of the STOP Both are data in fast-path-services.json; design-infra.md § 2 places them. - **namespace_routing** — 54 provider namespaces -> 30 rubric routes, 16 skips, 8 gates. Rules, not rows: 54 rules cover a provider surface past a thousand types. - **child_type_rule** — a canonical type with 2+ segments after the provider is a child of the type formed by dropping its last segment; if the parent has a disposition, the child is a config_source of it. Derived from evidence, not invented: 27 of the 33 config_source rows in skip_mappings are child types and 0 of the 19 noise rows are, so the disposition was always derivable from the type path. Only the description of WHAT it contributes needed authoring. Three constraints on both: an explicit row always wins; the route is RECORDED; and a derived route can never produce confidence: deterministic. ## routing_provenance New REQUIRED field on every services[] entry: `table` | `index_md` | `child_type_rule` | `namespace_rule`. Asserted, with the deterministic guard asserted separately — that tier needs an authored direct_mappings row and a fast_path_row naming it, so a derived route claiming it is a defect. The field exists because the risk these rules introduce is not wrongness, it is INVISIBILITY: without it, a namespace-routed mapping is indistinguishable from curated judgement in the artifact. Same discipline as sizing_provenance. ## Two thin rubrics, because the rules made them load-bearing namespace_routing sends Microsoft.Storage/* and Microsoft.NetApp/* to storage.md and Microsoft.KeyVault/* and Microsoft.ManagedIdentity/* to identity.md, so leaving them unwritten would have converted the unknown-type STOP into a missing-rubric HALT — moving the failure, not fixing it. Written thin, per plan §14: - storage.md — the SMB/NFS protocol rule's reasoning, access tiers, and the replication finding that matters: GRS/GZRS becomes S3 Cross-Region Replication, a NEW cost line that Azure bundled into one SKU. - identity.md — the secret/key/certificate split, and an explicit REFUSAL to translate Azure RBAC into IAM policy. A translated policy is plausible and wrong, and the failure is either privilege escalation or an outage, both silent until exercised. Read the assignment for its identity_grant edge, record the intent, let a human author the policy. ai.md stays the one declared-pending route (build step 6). ## Also index.md's closing paragraph sent readers straight from "no row" to the STOP. It now names both derived rules, and explains why a speculative row is worse than no row: an authored row SHADOWS the namespace rule that would have routed the type correctly and recorded that the decision was derived. _coverage_contract revised: a type may now resolve by an authored row, an index.md row, child_type_rule, or namespace_routing. Requiring a per-type row made coverage growth O(n) in authored judgement against a provider surface past a thousand. ## Checks check_routing_provenance (enum + the deterministic guard) and check_namespace_routing_targets (every namespace route resolves to a file on disk or a declared-pending one). Mutation-tested three ways: a derived route claiming deterministic fails; a missing field fails; a namespace rule pointing at an unwritten rubric fails. Corpus expectation recorded honestly: this corpus routes NOTHING through a derived rule, so the derived counts are zero and the checks guard the contract rather than exercise it. Extending the corpus with a type only a namespace rule can place is the follow-up. Effect on the five example estates: legacy-lift and data-platform already cleared with the 5b rubrics; the two remaining halts are untranslated types only, which is confirm's job. Nothing now halts for want of a disposition. Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check, lint:model-ids green; drift:check OK (347 identical); fixtures:assert PASS both trees. lint:md and fmt:check remain UNVERIFIED — dprint and markdownlint absent.
…exception list
"Building this for all types is impossible and not a good design. Deterministic
mappings for what we are sure about, LLM for the rest." Measured, and the data
backs it completely.
Of the 132 rows in arm-type-canonicalization.md:
105 (79%) have a resource segment a PATTERN derives
104 of those also have a namespace already declared in namespace_routing
27 (20%) carry information no pattern can produce
So four fifths of the file was restating a rule, at 13% coverage of the azurerm
surface, decaying every provider release. The 27 that matter are the traps it
already documents: linux_web_app -> sites, service_plan -> serverfarms,
public_ip -> publicIPAddresses, lb -> loadBalancers, api_management -> service,
application_insights -> components, stream_analytics_job -> streamingjobs.
## Derivation is now the default path
1. Table first - because that is what makes derivation SAFE: the cases where a
guess goes wrong are enumerated there.
2. Not listed -> derive. Resource segment by camelCase-pluralise of the Terraform
suffix (79% provable). Namespace from knowledge, then CROSS-CHECKED against
fast-path-services.json namespace_routing, which declares 55 namespaces
independently of this file.
3. Namespace recognised -> keep the resource with its full config,
azure_type_provenance: "derived", type recorded in iac_metadata.derived_types.
4. Namespace NOT recognised -> STOP, recorded in untranslated_types. That field
now means "no recognised namespace", a far stronger signal than "no row".
The cross-check is the guard, not a hope: a derived namespace that an independent
artefact also carries is corroborated; one nobody recognises is exactly where the
model is inventing, and that is the case that stops.
## An unlisted type is no longer DROPPED
The old behaviour discarded the resource entirely - type string only, no config.
That is what made 87% of the provider surface a hard stop AND what destroyed the
sku/tier evidence Design needs for the cost-bearing test, which is why 13.1d had
to assume cost-bearing for anything unnamed. A derived resource keeps its place in
resources[] with config intact, so namespace_routing can place it and the SKU test
can actually run.
## The growth cap, so this is mechanical rather than aspirational
check_canonicalization_governance:
- max_derivable_rows 105 - a CEILING, not a target. Adding a derivable row means
the file is chasing completeness again and fails the gate. It should FALL as
redundant rows are pruned, never rise.
- min_divergent_rows 27 - a floor, so pruning cannot delete the informative rows.
- every namespace in the table must exist in namespace_routing, or the two
artefacts contradict each other and the cross-check would reject a type the
table asserts is real.
The derivable rows present today are a CLOSED CORE: verified, free at runtime, so
they stay. What changed is the admission test - a new row is admissible only where
derivation would be WRONG.
That third check found a real contradiction on its first run:
Microsoft.DevTestLab was in the table with no namespace rule. Added as a skip - a
DevTest Lab is a management wrapper around VMs that are inventoried in their own
right.
## One DELIBERATE test failure, and it is the point
check_halt_not_stale fails on after-design-halted, correctly.
Derivation resolves azurerm_iothub to Microsoft.Devices/iotHubs, and
Microsoft.Devices IS a namespace_routing gate - so the corpus's halt describes
behaviour the skill no longer produces, while every shape assertion still passes.
That is the hand-authored-golden-drifts-from-a-prose-ref trap in its dangerous
direction: green and wrong. A golden cannot notice that a prose rule changed
underneath it, so the check reads the rule and the artifact TOGETHER.
Left red on purpose. A knowingly-stale green is worse than a red with a precise
reason, and the message names the fix: regenerate the golden with a capability run.
Consequence worth stating plainly - Design should now COMPLETE on the corpus, which
makes Estimate reachable for the first time.
Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check,
lint:model-ids green; drift:check OK (347 identical). fixtures:assert has the one
intentional azure design failure above; the other 10 asserters pass in both trees.
lint:md and fmt:check remain UNVERIFIED - dprint and markdownlint absent.
…day's work Ran the capability test on the corpus: fresh isolated agent, answer keys hard prohibited, pointed only at SKILL.md. All three phases completed with HANDOFF_OK and NO halt, so the derivation design works end to end and Design completes on the corpus for the first time. It also found 10 file-level contradictions and 20 under-specifications. The worst were mine, from the same day. ## The one that mattered most (NOTES 2.8) extract-terraform.md said a type with no per-type-attributes row carries no sizing attributes. design-infra.md decides cost-bearing-ness by asking whether config has a SKU/tier/capacity property. A DERIVED type has no row by definition, so its SKU was dropped and it was then GUARANTEED to look benign — reintroducing, from the other end, the exact silent under-report that decision 13.1d exists to prevent. The run carried the sku anyway and recorded that no rule authorised it. It only escaped mattering here because the namespace gate fired first; on an estate whose derived type lands in a rubric namespace it would have understated a cost-bearing resource. Fixed by generalising the rule Step 2a already states for AzAPI: always extract sku / sku_name / tier / capacity / size, for EVERY type, listed or not. ## Contradictions closed (all introduced yesterday) - NOTES 1.1 — discover-iac.md still said an unlisted type is SKIPPED, against three files saying it is derived and retained. It is the FIRST file the fragment loads, so two readings of one estate gave a completed design versus a halt. Replaced with the derive-cross-check-or-stop rule in full. - NOTES 1.2 — extract-terraform.md's own closing checklist required every azure_type to appear in the canonicalization table, which Step 2b exists to admit exceptions to. A correct derived entry failed the checklist. Now states both provenances, and gains a checklist item for the SKU rule above. - NOTES 1.10 — two status blocks asserted files were missing that had landed hours earlier: clarify-global.md on azure-region-map.json (the run correctly used the table and suppressed the "absent" flag it was told to emit), and design.md listing five rubrics as "still to come". Both corrected, and design.md now says outright that the filesystem wins over the paragraph — a prose status list goes stale the moment a file lands, and the halt decision reads from disk. ## Skill files no longer cite their own oracle Four places named check_expected_design.py as the authority on rules they did not state. Three were mine from yesterday. The run reported being STRONGLY TEMPTED to open it, and put the reason precisely: "a skill that cites its own oracle as a normative reference is teaching by pointing at the answer key." That is worse than untidy — reading an asserter invalidates a capability run, which is how this skill is tested, so the citation actively degrades the test method. All four replaced by the rule stated where it belongs. design-infra.md now says: if a rule is only discoverable by reading test code, the rule is missing and the skill file is the bug. ## service_id was unreproducible, and the oracle keyed on it The schema specified it as "<stable slug, unique within the artifact>". Capability run 4 produced 8 of 14 ids differently from the golden — every one a valid stable slug — and three size assertions failed on a NAMING difference rather than a wrong size. An identifier nobody can reproduce cannot be referenced from a report, cross-checked between artifacts, or asserted by a fixture, and the instability is invisible because each run is internally consistent. Two fixes, both needed: - schema-design-aws.md gains a derivation rule: <aws-service-slug>-<local-name>, no abbreviating, because an abbreviation is a choice and choices are what made this unreproducible. - the asserter now pins by azure_id, which is built by a stated rule and is the same string in every artifact that mentions the resource. Nothing may key on service_id across artifacts. ## Still deliberately red, unchanged check_halt_not_stale still fails on after-design-halted, and the capability run still fails the no-halt and 29-resource expectations. Same single root cause: those expectations describe pre-derivation behaviour. Regenerating the goldens from this run — which also fixes the Cosmos clustering error the run exposed in the golden — is the next step and was scoped out of this change. Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check, lint:model-ids green; drift:check OK (347 identical). lint:md and fmt:check UNVERIFIED.
…ing halts for
want of knowledge
gcp-to-aws serves customers today, and checking WHY it handles every case settled
this. It is not mapping breadth: gcp lists 28 google_* types against azure's 65
index rows and 92 fast-path rows. It is three properties, and only one is knowledge:
1. complete CATEGORY coverage — 9 rubrics, all on disk
2. a permissive router — unknown type -> category -> rubric answers
3. NO veto — nothing can reject a type once a category is found
and no missing-rubric halt rule at all.
Azure had double the per-type coverage and stopped more often, which is the wrong
trade. It had acquired two azure-only obstacles:
- a 55-entry namespace list that could VETO a derived type. That is the same failure
as a 1089-row table, one level up: 14 of 15 sampled plausible namespaces were
absent, so Microsoft.Maps/accounts halted a run despite being unambiguous.
- an explicit prohibition, written by me yesterday: "Do NOT guess a category from
the type NAME."
## What changed
**The namespace cross-check is now a SIGNAL, not a veto.** Recognised ->
azure_type_provenance "derived". Unrecognised -> "derived_uncorroborated" plus a
type_derived_uncorroborated warning. Either way the resource is KEPT with its full
config. It records whether a second artefact agreed; it does not decide whether the
resource exists.
**An unrecognised namespace routes to a model-chosen category.** Pick the best-fit
rubric from those on disk, apply its six criteria like any other pass-2 resource,
record routing_provenance "model_category", warn with the namespace and the category
chosen, take confidence inferred. Not a stop.
**iac_metadata.untranslated_types now means one thing only: the skill cannot NAME
the service.** Not "no table row", not "namespace unrecognised" — both of those are
derived and retained. Expect it empty on almost every repo.
**A STOP survives only where the model genuinely cannot say what a service does.**
That is a real answer, and it is rare.
## Deviating from gcp in exactly one way, deliberately
gcp records NO provenance for a pattern-routed category — it uses the category and
sets confidence inferred, and nothing in the artifact says the category was guessed.
Azure keeps the provenance field, because that is what makes a derived decision
reviewable instead of indistinguishable from a curated one. Mimic the behaviour, not
the blind spot.
Azure's router is also better than gcp's on the same axis: gcp matches a SUBSTRING of
the type name ("log" -> monitoring), which also matches google_diaLOGflow_agent. The
provider namespace is structural and does not have that failure mode.
## The line this settles
Fall back to the model for FACTS about Azure — what a service is, what a SKU's vCPU
count is, which category a namespace belongs to. Never for this project's OPINIONS:
Elastic Beanstalk over Fargate for PaaS posture, x86_64 over Graviton, single-AZ plus
a visible finding rather than inferring Aurora from a zone-redundant source. A
model-chosen category is fact-shaped; the rubric it lands in still supplies the
opinion. That is why the missing-rubric halt stays and the namespace veto goes.
Recorded in design-infra.md so the reasoning survives, not just the rule.
## Effect
Nothing in the five example estates blocks Design any more — three-tier-saas's
azurerm_maps_account was the last one, and Microsoft.Maps not being in the list is now
irrelevant to whether it resolves.
Enums widened: azure_type_provenance gains derived_uncorroborated,
routing_provenance gains model_category, and the deterministic guard now names all
three derived routes explicitly.
Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check, lint:model-ids
green; drift:check OK (347 identical). The azure design asserter still fails on the
two pre-derivation expectations and the stale halt — same root cause, still scoped to
the golden regeneration. lint:md and fmt:check UNVERIFIED.
The whole suite is green again, and the red it clears was real: two goldens were describing behaviour the skill no longer produces, while every shape assertion still passed against them. ## after-discover — replaced, and it fixes an error the golden had Adopted run 4's output verbatim: 30 resources (was 29) because a derived type is now RETAINED, provenance per entry, and config.tf_file populated on 30 of 30 where the hand-authored golden had it empty on 24 of 29 against a ref that mandates it. Clusters 4 -> 5, and this one was a defect in the golden rather than a consequence of my change. The golden put the Cosmos account inside the 21-member cluster justified "merge:cross_group_edges" -- but that account has ZERO edges in either direction and none of the cluster's three justifying edges touches it. Run 4 split it out, which is what the algorithm produces: an edgeless primary-eligible resource in its own resource group is its own workload (13.4b), and a split is justified by an ABSENCE (13.3d) so its edges[] is legitimately empty. The old number came from intuition -- "the app must use the database" -- and the corpus declares no such reference. That is 13.4e firing a THIRD time: a golden or worked example must be TRACED against the algorithm, never authored by hand. ## after-design-halted -> after-design The halt became unreachable: azurerm_iothub derives to Microsoft.Devices/iothubs, Microsoft.Devices is a namespace_routing gate, and the resource defers cleanly. A golden for a state the skill cannot produce is worse than none, so the directory is renamed and the assertion INVERTED -- the oracle now requires that halt NOT occur, and requires the derived type in deferred[]. Mapped or deferred, never quietly skipped. The benign-skip guard survives unchanged and is still the sharpest assertion here. ## Two more instances of the oracle asserting what no skill file taught Run 4's clarify output failed three checks, and two were the oracle's fault: - clarify-compute.md showed a SINGLE-element app_service_plans example and never said one row per plan. Run 4 emitted only the plan it questioned, correctly by the ref. Now stated: one row per plan ALWAYS, including plans you do not ask about, with disposition N/A and a reason for 0- and 1-app plans -- because design-infra.md's fan-in rule looks a plan up here and a plan with no row is indistinguishable from a plan the user declined to split. - NOTHING required licensing._fired or _firing_reason. Now required, with the reason: "the category did not fire" and "it fired and the fragment forgot to write it down" otherwise produce the same artifact, and the second silently drops a licensing decision worth real money. clarify_status was in the same category -- asserted by the oracle, defined nowhere. Now defined in clarify-assemble.md as COMPLETE | BLOCKED_ON_ESSENTIAL, including the rule run 4 had to invent: a user answer that conflicts with a hard_blocker does NOT block. Record it as given with a conflict key, put the blocker in licensing.blockers[] with severity blocker, and keep status COMPLETE -- the customer answered, and the blocker is a prerequisite rather than a competing preference. ## What I deliberately did NOT do after-clarify-complete/ is NOT shipped. Run 4 produced one, but it predates the two rules above, and hand-editing it to comply would be authoring a golden by intuition -- the exact failure that has now bitten three times. clarify-answers-complete.json is committed with the completing answer set and a _no_golden_yet note explaining that the next capability run against the corrected files produces it properly. after-clarify/ (BLOCKED) is kept: the scripted user there declines to state spend, and that branch is what proves ESSENTIAL + value:null gates the phase. Its clusters[] list gained the new cluster id -- mechanically forced, since preferences.clusters[] must key exactly to the clusters artifact. ## Also fixed The canonical-type coverage check read EVERY Microsoft.* string in arm-type-canonicalization.md, so a type named in prose looked like a table row and reported a false orphan (Microsoft.Maps/accounts, which I had used as an example). It now reads table rows only. Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check, lint:model-ids green; drift:check OK (347 identical); fixtures:assert PASS in BOTH trees -- 11 asserters, 7 golden, 4 smoke, no deliberate failures left. lint:md and fmt:check UNVERIFIED.
…missing its root volume
Two findings from capability run 4, both of which understate or misdirect a real
estate while leaving every test green.
## NOTES 1.6 — database.md read the availability answer from the wrong path
database.md's override-gate table keyed on `design_constraints.availability`. The
answer lives at **`data.availability`** — clarify-database.md writes it there,
schema-preferences.md documents it there, and the committed golden has it there.
database.md was the ONLY reader, and it was looking in the place gcp-to-aws keeps it.
The bug was invisible rather than loud, which is what makes it worth a commit
message: looking in the wrong place finds nothing, and the gate's own "if the answer
is absent, default to RDS single-AZ" rule then produces the SAME answer the corpus
expects. So every oracle stayed green. On an estate where the customer answers
`multi-az-ha` it silently produces RDS where Aurora was chosen — and records the
customer's answer faithfully, right next to a target that ignores it.
Fixed, with the reason it stayed hidden written into the file so the next person does
not re-port the gcp path.
## NOTES 2.10 — the VM's root volume was absent from the design entirely
A VM's OS disk is normally an INLINE BLOCK, not an azurerm_managed_disk:
os_disk { caching = "ReadWrite", storage_account_type = "Premium_LRS" }
So no Microsoft.Compute/disks resource exists, a resource-block walk misses it, and
no EBS volume reaches the design. disk-ebs-sizing.json already carried the RULE (OS
disk becomes the instance's root volume in aws_config, never its own entry) — what
was missing was the EXTRACTION. Run 4 recorded a prose note because no rule let it do
better, and flagged that this systematically understates every VM estimate.
extract-terraform.md gains a section for inline blocks that are not resources, and
os_disk joins the VM and VMSS attribute rows. Design maps it to aws_config.root_volume
using the tier map: Premium_LRS -> gp3, a lookup rather than a judgement.
**When disk_size_gb is absent the size is the image default, and the rule now forbids
supplying one.** Carry size_gib null with size_source "image_default_unstated". A
remembered 127 GiB is indistinguishable from a measured one and Estimate would price
it as fact; a stated unknown is worth more than a plausible number.
## Goldens
Both edits to the goldens are mechanical reads of committed input, not judgements: the
inventory's VM gains config.os_disk verbatim from compute.tf (with disk_size_gb null
because the source omits it), and the design's EC2 entry trades its prose note for a
structured root_volume whose volume_type comes from the tier map.
check_inline_os_disk asserts it four ways, all mutation-tested: reverting to a prose
note fails; a wrong volume_type fails; INVENTING the image-default size fails; and
emitting the OS disk as its own services[] entry fails, because counting it twice is
the commonest way a VM estate's storage cost gets inflated.
Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check, lint:model-ids
green; drift:check OK; fixtures:assert PASS both trees. lint:md and fmt:check
UNVERIFIED.
… clarify branch Run 5: fresh isolated agent, answer keys hard-prohibited, pointed only at SKILL.md. All three phases HANDOFF_OK, no halt. Discover and Design PASSED their oracles unmodified — the first time a capability run has cleared them on first contact. ## What run 5 corroborated It independently produced exactly what I had hand-added yesterday: config.os_disk with disk_size_gb null, and a root_volume of gp3 / size_gib null / size_source "image_default_unstated" — plus the gp3 baseline_iops I had not bothered with. Same 5 clusters, same 16 services. So the goldens are now the RUN's output rather than my edits. Per 13.4e a golden must be TRACED, not authored, and an earned artifact that happens to match a hand edit is still worth more than the hand edit. ## The completing clarify branch finally exists after-clarify-complete/ + expected-clarify-complete.json, from run 5. Both branches of the phase are now asserted, and they are opposite states on purpose: - after-clarify/ — the scripted user DECLINES to state spend, so an ESSENTIAL row is null and the phase must GATE. Proves 13.5b. - after-clarify-complete/ — spend answered, so the phase must COMPLETE. This is what makes Design reachable, and it additionally pins the conflicting-answer rule. check_conflicting_answer is new: vm_cutover mgn against the Azure-Edition hard blocker must be recorded AS GIVEN with a conflict key, the blocker must appear in licensing.blockers[], and clarify_status must stay COMPLETE. Run 5 got all three right and cited clarify-assemble.md:106-111 as specifying it "completely". check_expected_clarify.py takes an optional spec argument so one asserter serves both; check_expected_clarify_complete.py is a thin wrapper only because run-asserters.py keys its registry on a bare path and cannot pass arguments. ## The finding worth the most: nine names for one field Run 5's clarify failed one assertion — 9 rows DETECTED with value == default and no stated source. The content was entirely correct; the KEY NAMES were `note` and `mapped_from` where the oracle accepted `source`, `forced_by` or `reason`. Nothing in the skill named the key. Across the golden and one run there were NINE different justification key names in play: source, reason, note, mapped_from, forced_by, context, conflict, source_ha_context, residency_warning. This is the third instance of the same defect class — service_id, then cluster_id, now this. A field a fixture keys on has to be named by a rule. schema-preferences.md now fixes the key BY DISPOSITION: DETECTED needs `source` (what in the estate was read), N/A needs `reason`, ESSENTIAL needs `context`, and `forced_by` takes precedence for a value a blocker removed the choice on. Extra keys are allowed alongside; they never substitute. The oracle narrowed to source|forced_by, and row_dispositions' requires_keys were aligned to follow the disposition — several rows were demanding `reason` on a DETECTED row, which is the N/A key, so the assertion was asking for the wrong field. cpu_architecture requires forced_by, because x86_64 there is not a free reading of the estate: a Windows workload removes Graviton as an option. Also fixed a bug in my own new check — it used get(), which returns None for anything that is not a dict, so a list-valued path read as empty and reported a blocker run 5 had recorded correctly. ## Adopted with two mechanical edits, declared The run's DETECTED rows had `note`/`mapped_from` renamed to `source` (a key rename; content untouched, and the name was unspecified when the run started), and 5 cluster pattern_id rows gained a `source` whose content is fixed by 13.3c plus the fact that patterns.md is absent — one possible reason, not a judgement. Both recorded in the artifact. Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check, lint:model-ids green; drift:check OK (347 identical); fixtures:assert PASS both trees — 12 asserters, 8 golden, 4 smoke. lint:md and fmt:check UNVERIFIED.
Verification evidence — round-2 report finding fixed (+ AI-only companion gap)Ran the skill end-to-end (headless 1. The 35-failure check now returns REPORT_OKRepo: The exact validator invocation from your finding: (Was: a 41-line, 4-section stub → 35 failures.) 2. Savings Plans / Reserved Instances section — the monetary omission — is present with product-specific rows
Not a single generic "commitment discounts" row, and it carries the contract caveats (incremental to Balanced on-demand, not additive on Optimized, RDS SP/RI mutual exclusivity). 3. Stable Appendix A–J lettering
4. Companion gap found + fixed: AI-only runs were producing no report at allTesting surfaced a different instance of the same silent-degradation class: on an app-code-only AI repo ( Fixed by making Mode is discriminating, not a rubber stamp — the same report fails full mode (missing infra sections). 67 existing validator unit tests still pass. Caveat (so the numbers aren't over-read)These runs priced via a search fallback (awspricing MCP unreachable in this environment) and ran Commits: |
herosjourney
left a comment
There was a problem hiding this comment.
User-facing surfaces: Azure is announced everywhere except the three plugin descriptions
Checked whether the PR notifies users of the new Azure→AWS capability across the README / marketplace / plugin-manifest surfaces (advisor tree only — confirmed the PR correctly touches zero migrate-tree files, which matches the plan to deprecate that tree).
Correctly updated — these announce Azure well:
.claude-plugin/marketplace.json— both plugin descriptions now lead with “Azure, GCP, or Heroku” and name the concrete mappings (App Service → EB, AKS → EKS, Azure OpenAI → Bedrock). ✅- top-level
README.md— azure-to-aws listed in both plugin rows. ✅ advisor/README.md— full azure-to-aws skill bullet with mappings, the intent-routing table, and the MCP-server line. ✅advisor/plugins/aws-startup-advisor/setup.md— install list + skill description. ✅advisor/AGENTS.md. ✅- All three plugin manifests’
keywordsarrays —azure,azure-to-aws,azure-sql,azure-openai. ✅
Gap — the human-readable description prose in all three plugin manifests still omits Azure, even though the keywords and the marketplace entry include it. This is the sentence a user actually reads in the Claude Code / Cursor / Codex plugin UI when deciding whether to install, so it now contradicts the marketplace entry:
| File | Field | Current prose | Fix |
|---|---|---|---|
.claude-plugin/plugin.json (line 5) |
description |
“Migrate to AWS from GCP or Heroku—…” | “Migrate to AWS from Azure, GCP, or Heroku—…” |
.cursor-plugin/plugin.json (line 4) |
description |
“Migrate to AWS from GCP or Heroku—…” | “Migrate to AWS from Azure, GCP, or Heroku—…” |
.codex-plugin/plugin.json (line 94) |
longDescription |
“Migrate from GCP, Heroku, OpenAI, or Gemini…” | “Migrate from Azure, GCP, Heroku, OpenAI, or Gemini…” (inline comment below) |
Suggest also naming an Azure primitive or two in the claude/cursor description for parity with the marketplace copy (e.g. “…including Azure App Service → Elastic Beanstalk, AKS → EKS, and Azure OpenAI → Bedrock…”), so the in-agent install card matches what the marketplace promises.
None of this blocks the skill — it works and is discoverable via keywords — but a user reading the plugin’s own description in-agent won’t see Azure listed as a supported source, which undercuts the launch you’re trying to announce.
Verification note: read against the PR head (916a9762) diffed vs main @ b30cb39b; the claude/cursor description lines fall outside the PR’s diff hunks (only their keyword arrays changed), so those two are flagged here in the body rather than inline.
| "displayName": "AWS Startup Advisor", | ||
| "shortDescription": "Build and migrate on AWS with startup-focused architecture, cost, and security guidance plus GCP/Heroku-to-AWS and AI-stack (Bedrock, agents) migration workflows.", | ||
| "longDescription": "Personalized architecture, cost, security, and migration guidance for startups. From day-one account setup and security baselines to production-ready infrastructure, cost optimization, and beyond. Includes AWS Activate Credits eligibility, 60+ exclusive startup offers, and multi-account multi-region support. Built on expertise from AWS Startup Solutions Architects and patterns from 350,000+ startups.\n\nFeatures\nVetted prompts for startups at every stage:\n\n\nDay-one account setup, security baseline, least-privilege roles\nScaffold a production architecture for your stack\nSet up GuardDuty, Security Hub, and a vulnerability scanner\nMigrate from GCP, Heroku, OpenAI, or Gemini — infrastructure, AI SDKs to Amazon Bedrock, and agentic systems\n\n\nAWS Activate Benefits\nCheck eligibility and explore offers available to startups on AWS:\n\n\nAWS Activate Credits: Check eligibility and learn how to apply for credits to offset costs across 200+ services, including infrastructure, data services, and AI/ML models on Bedrock.\nExclusive Startup Offers: Access discounts and extended trials on dozens of tools for payments, analytics, communications, and developer productivity.", | ||
| "longDescription": "Personalized architecture, cost, security, and migration guidance for startups. From day-one account setup and security baselines to production-ready infrastructure, cost optimization, and beyond. Includes AWS Activate Credits eligibility, 60+ exclusive startup offers, and multi-account multi-region support. Built on expertise from AWS Startup Solutions Architects and patterns from 350,000+ startups.\n\nFeatures\nVetted prompts for startups at every stage:\n\n\nDay-one account setup, security baseline, least-privilege roles\nScaffold a production architecture for your stack\nSet up GuardDuty, Security Hub, and a vulnerability scanner\nMigrate from GCP, Heroku, OpenAI, or Gemini \u2014 infrastructure, AI SDKs to Amazon Bedrock, and agentic systems\n\n\nAWS Activate Benefits\nCheck eligibility and explore offers available to startups on AWS:\n\n\nAWS Activate Credits: Check eligibility and learn how to apply for credits to offset costs across 200+ services, including infrastructure, data services, and AI/ML models on Bedrock.\nExclusive Startup Offers: Access discounts and extended trials on dozens of tools for payments, analytics, communications, and developer productivity.", |
There was a problem hiding this comment.
This longDescription still reads “Migrate from GCP, Heroku, OpenAI, or Gemini” — Azure isn’t listed, even though this PR adds azure/azure-to-aws/azure-sql/azure-openai to the keywords just below and the top-level marketplace.json now leads with “Azure, GCP, or Heroku.” This is the description a user reads in the Codex plugin UI, so it now contradicts the marketplace entry. Add Azure to the sentence, e.g. “Migrate from Azure, GCP, Heroku, OpenAI, or Gemini — infrastructure, AI SDKs to Amazon Bedrock, and agentic systems.” The same fix is needed in the description field of .claude-plugin/plugin.json and .cursor-plugin/plugin.json (both still say “GCP or Heroku”; flagged in the review body since those lines are outside this PR’s diff).
herosjourney
left a comment
There was a problem hiding this comment.
Three correctness findings, all verified against this branch (head 916a9762) vs main @ b30cb39b. None is a style nit; each is a way the skill either ships a false statement or advertises a capability that halts.
1. This PR promotes a lifecycle file that reasserts the “6-month Legacy” error #307 is fixing — and they clobber each other silently
Confirmed: #304 creates the canonical skills/shared/ai/ai-model-lifecycle.md (plus vendored copies under each skill’s references/vendored/ai/), and line 11 reads “Legacy (minimum 6 months before EOL)” with no mention of the 45-day policy, the 2026-09-07 split, model-card governance, or the EOL’d rows. That is exactly the false statement open PR #307 exists to correct.
The important nuance: the two PRs touch different paths, so git will not raise a merge conflict.
- #307 edits the old path
skills/gcp-to-aws/references/**shared**/ai-model-lifecycle.md. - #304 deletes that path and relocates to
skills/**shared**/ai/+references/**vendored**/ai/.
So whichever lands second silently wins, with no conflict marker to force a human to reconcile: if #304 lands after #307, the stale content silently reappears in the new canonical copy; if #307 lands after #304, it patches a path that no longer exists. #307 is currently OPEN and main still says “minimum 6 months,” so the risk is live.
Fix: a plain rebase is not enough (the paths differ, so #307’s diff won’t apply to #304’s new files). Port #307’s corrected content into #304’s new canonical skills/shared/ai/ai-model-lifecycle.md, then re-sync the vendored copies from it (they are byte-identical to the canonical today — keep them so). Verify the canonical copy carries the 45-day/2026-09-07 policy before merge, since it is the source every vendored copy derives from and no conflict will protect it.
2. The skill advertises Bicep / ARM / live az in its triggers, but those paths halt or don’t exist
Confirmed. SKILL.md’s description triggers include “migrate Bicep to Terraform,” “migrate ARM templates to Terraform,” and a first-class “read-only consent-gated live az CLI capture.” But:
discover-iac.md:208states plainly: “Until those two refs exist, a workspace containing.bicepor ARM templates halts.” Andreferences/shared/extract-bicep.md/extract-arm.mddo not exist on this branch.- There is no live-
azcapture fragment inreferences/phases/discover/(onlydiscover-iac.md,discover-app-code.md,discover-assemble.md) and no Azure live-capture fixture.
So a user who says “migrate my Bicep to Terraform” or expects the advertised live az capture gets matched, then hits a halt. Discover itself is honest — the problem is the marketing in the trigger list and the 7-phase description overpromising relative to what’s wired.
Fix: make the description Terraform-first and name Bicep / ARM / live-az as not-yet-supported (or remove them from the triggers) until extract-bicep.md, extract-arm.md, and the az fragment land. Inline comment on the trigger line below.
3. llm-to-bedrock still hardcodes a “gcp-to-aws” handoff with no Azure branch
Confirmed. skills/llm-to-bedrock/SKILL.md:127 reads: “I'm now invoking the gcp-to-aws Assess skill to discover your AI workloads and design…” — there is no Azure signal, azure-to-aws is not mentioned anywhere in that skill, and #304 does not touch llm-to-bedrock at all. Once the Azure skill ships, an Azure user routed into the AI path is told the tool is invoking “gcp-to-aws,” which is a user-visible inconsistency.
The PR body already identified this as the minimal fix and deferred it. That’s a defensible sequencing call (it pairs with #290, which is already merged) — but it should be an explicit decision, not a silent gap. Either add the one-line “Azure estate → azure-to-aws” branch/note to llm-to-bedrock/SKILL.md, or state in the PR that it is intentionally out of scope and tracked as a follow-up. Leaving the hardcoded “gcp-to-aws” sentence while shipping an Azure skill is the remaining user-visible footgun.
Verification note: findings 1 and 3 are flagged in this body rather than inline — finding 1’s file was relocated (the anchor path is ambiguous across the move) and finding 3’s file is untouched by this PR, so neither line sits in the diff. Finding 2 is anchored inline on the SKILL.md trigger line.
| @@ -0,0 +1,221 @@ | |||
| --- | |||
| name: azure-to-aws | |||
| description: "Migrate workloads from Microsoft Azure to AWS. Triggers on: migrate from Azure, Azure to AWS, move off Azure, migrate Azure app to AWS, migrate AKS to EKS, migrate App Service to AWS, migrate App Service to Elastic Beanstalk, migrate Azure VMs to EC2, migrate Azure SQL to RDS, migrate Azure Database for PostgreSQL to RDS, migrate Azure Database for MySQL to RDS, migrate Cosmos DB to DynamoDB, migrate Azure Cache for Redis to ElastiCache, migrate Blob Storage to S3, migrate Azure Functions to Lambda, migrate Service Bus to SQS, migrate Event Hubs to Kinesis or MSK, migrate Azure OpenAI to Bedrock, migrate my Azure OpenAI app to AWS, move Azure AI workloads to AWS, migrate Azure-hosted agentic workloads to AWS, migrate Azure-hosted LangChain to Bedrock, migrate Azure-hosted LangGraph to AWS, migrate Bicep to Terraform, migrate ARM templates to Terraform, leave Azure, estimate AWS costs for my Azure infrastructure, what-if workshop, reprice Azure migration, compare migration scenarios, workshop mode. Runs a 7-phase process: discover Azure resources from Terraform/Bicep/ARM templates, a read-only consent-gated live `az` CLI capture, an optional Resource Discovery for Azure report, application code, and optional billing exports; clarify migration requirements via an assumption sheet; design AWS architecture; estimate costs (both a 1:1 lift and a right-sized target); optionally reprice what-if scenarios in a workshop sidebar; generate migration artifacts when the user opts in; and collect optional feedback. Clarify must finish before Design, Estimate, or Generate, and Generate is opt-in — it runs only after the user chooses Execute at the post-Estimate decision gate. Uses resource-group-seeded clustering refined by typed edges from ARM resource IDs, canonical `Microsoft.*` ARM resource types as the mapping key, a deterministic fast-path table for architecture-invariant primitives, and a pattern catalog so recommendations describe workloads rather than isolated resources. Do not use for: GCP migrations (see gcp-to-aws), Heroku migrations (see heroku-to-aws), general AWS architecture advice without migration intent (see architect-for-startups), AWS-to-Azure reverse migration, on-premises-into-Azure discovery (that is Azure Migrate's job, not this skill's), or Azure-to-Azure refactoring." | |||
There was a problem hiding this comment.
These triggers advertise “migrate Bicep to Terraform,” “migrate ARM templates to Terraform,” and a “read-only consent-gated live az CLI capture,” but none of those paths is wired yet: discover-iac.md:208 says a workspace with .bicep or ARM halts until extract-bicep.md / extract-arm.md exist (they don’t on this branch), and there is no live-az capture fragment or fixture in references/phases/discover/. A user whose ask matches one of these triggers gets routed in and then stops. Make the description Terraform-first and name Bicep / ARM / live-az as not-yet-supported (or drop them from the trigger list) until those extractors and the az fragment land. The 7-phase summary sentence (“discover … from Terraform/Bicep/ARM … a read-only consent-gated live az CLI capture”) should be softened the same way.
Blocking checklist item: canonical AI lifecycle file must carry #307\u2019s 45-day policy before mergeSharpening finding #1 from my earlier review into a concrete, blocking action now that the plan is settled: #307 is being folded into this PR and closed as superseded. #307\u2019s fix is correct and needs to land within ~10-14 days; this PR merges (~7 days) inside that window, so it should carry the fix rather than have #307 merge separately on a path this PR deletes.\n\nThe risk this guards against: this PR\u2019s canonical |
Ready-to-apply: #307's lifecycle fix, merged into this PR's canonical filePer the plan in the earlier comment — #307 is being folded into this PR and closed as superseded. I've done the merge and verified it; below is the corrected canonical file plus one command to regenerate the vendored copies. This is a two-way merge, not a copy — I kept #307's facts and this PR's framing:
How to apply (edit the canonical, then sync)
mise run shared:sync # canonical -> the 4 vendored copies (both trees)
mise run shared:check # must pass
Verified before postingAgainst this branch ( One judgment call: I kept #307's Sep-21 dated table. Since this PR is ~a week out, those Do not close #307 until this content is committed here — closing it early while this file stays stale loses the fix silently (no merge conflict fires, and Merged canonical
|
Proposed:
|
…l ai-model-lifecycle.md (PR awslabs#304 merge blocker) herosjourney's blocking checklist (PR awslabs#304): awslabs#307 is being folded into awslabs#304 and closed as superseded. Because this PR relocates ai-model-lifecycle.md into the canonical skills/shared/ai/ tree (deleting gcp-to-aws/references/shared/...), no merge conflict fires — so if awslabs#304 merged with the stale canonical, shared:check would keep the WRONG content ("Legacy minimum 6 months before EOL", no 45-day policy) byte-identical across all 6 copies. This carries awslabs#307's fix in. Two-way merge (herosjourney's ready-to-apply content, comment 5770573792): kept awslabs#307's FACTS + this PR's canonical FRAMING. - From awslabs#307: the 45-day / 2026-09-07 model-card policy split and the "don't apply the 6-month assumption post-2026-09-07" warning; both reference links; EOL'd rows (Claude 3 Haiku, Nova Premier v1, Nova Sonic v1) moved from the live table into Removed; restricted / Covered-Model handling; the GetFoundationModel unverified-fallback guidance; dates refreshed to Sep 21. - From this PR: the canonical blockquote header (source-agnostic + edit-here- then-shared:sync) and the shared-location wording. Applied to BOTH tree canonicals (advisor + migrate), then shared:sync propagated to all vendored copies. All 6 ai-model-lifecycle.md files are byte-identical. Verified: shared:check OK (both trees), drift:check OK (397 identical), lint:model-ids OK, fmt:check clean, lint:md 0 errors. Only the 6 lifecycle files touched. Stale "minimum 6 months before EOL" gone; 45-day policy present in every copy; Haiku/Premier/Sonic in Removed. awslabs#307 to be closed as superseded only after this lands.
Blocking lifecycle item — done ✅ (
|
…ng-cache + phase-status conflicts) main advanced 39 commits (incl. awslabs#307 Bedrock lifecycle refresh merged as 39ca3d4, awslabs#291 heroku decision gate, awslabs#310/awslabs#314 report + plan-writer). Conflict resolution: - ai-model-lifecycle.md (gcp vendored, both trees): took OURS — same awslabs#307 facts, but at this PR's correct vendored/canonical path with the canonical header (main edited the old gcp-to-aws/references/shared/ path this PR deletes). - pricing-cache.md (gcp shared, both trees): hybrid — kept main's FRESHER lifecycle facts (Nova Canvas/Reel excluded; Nova Sonic past-EOL) but rewrote the path references from shared/ai-model-lifecycle.md -> vendored/ai/ai-model-lifecycle.md (this PR relocated the file; shared/ path would dangle). - phase-status.schema.json (shared + heroku vendored, both trees): took THEIRS (main's newer run_mode description prose); run_mode key preserved. - model-id-lint.py (both trees): updated the Haiku/Premier/Sonic allowlist paths from the deleted skills/gcp-to-aws/references/shared/ai-model-lifecycle.md to the canonical skills/shared/ai/ai-model-lifecycle.md (vendored copies inherit via canonicalize()) — main's linter didn't know this PR moved the file. Post-merge shared:sync re-propagated the shared phase-status schema into the azure vendored copy (canonical-staleness trap). Verified green both trees: shared:check OK, drift:check OK (398 identical), frontmatter 7 files, asserters 15 PASS, 92 validator unit tests pass, mise run build all 16 tasks clean (lint:model-ids OK, fmt clean, lint:md 0 errors, all security scanners clean). All six ai-model-lifecycle.md copies byte-identical.
… skill-list conflict) main advanced another 13 commits (awslabs#313, awslabs#316, cask-1555 cost-routes, pricing/plan.json fixes). Single conflict: advisor/README.md — main reworded the MCP-servers sentence ('live AWS data' / 'AWS documentation lookups') while this PR added azure-to-aws to the skill list. Resolved as a union: kept azure-to-aws in the list AND main's improved wording. Post-merge shared:sync clean (0 updates). Verified: shared:check OK, drift OK (398 identical), asserters 15 PASS, lint:model-ids OK, mise run build all 16 tasks clean.
…onsolidation conflicts) main awslabs#318 consolidated pricing onto a cache-only model (removed the live pricing MCP), touching pricing-mode.md, pricing-fallback.md, aws-infra-pricing.json, design-ai.md, and the README across both trees — 13 conflicts. Resolution: - pricing-mode.md / pricing-fallback.md (canonical + heroku vendored): took THEIRS (awslabs#318's cache-only rewrite; our PR does not touch these). - aws-infra-pricing.json (_instances_note): UNION — kept our azure-specific note (x86 families as the azure default for Windows/.NET fleets, flexible-server-rds-sizing.json ref) and adopted awslabs#318's cache-only correction ('verify against the public AWS pricing pages' instead of 'via the AWS Pricing MCP'). - design-ai.md: took awslabs#318's 'If MCP call fails' phrasing (no retry count — cache-only) but kept our vendored path references/vendored/ai/ai-migration-guardrails.md. - advisor/README.md: union — azure-to-aws in the skill list + awslabs#318's wording. Post-merge shared:sync reconciled vendored copies (6/tree). Verified: shared:check OK, drift OK (399 identical), frontmatter 7, asserters 15 PASS, 92 validator tests pass, mise run build all 16 tasks clean.
herosjourney
left a comment
There was a problem hiding this comment.
Review: add azure-to-aws migration skill
Reviewed at the design/behavioral level (393 files, +56k — not read line-by-line); findings below are verified against head 1df5cb095ed0. This is a strong foundation. Nothing here is architectural — the blockers are truthfulness of the SKILL.md triggers and the per-file Status blocks, plus one dangling reference.
Genuinely good (verified)
- Cost/estimate honesty is excellent. Estimate asserts every design service appears in the breakdown exactly once — priced, or carrying an
exclusion_reasonwithexcluded_from_total(nothing silently dropped); totals are explicit floors; traditional-AI workloads (Document Intelligence→Textract, etc.) land inservices_not_estimated[]rather than faked into token cost. The estimate golden fixture is deliberately "not clean" (5/16 unpriced) — it tests the states an implementation is most likely to fake tidy. - Pricing is cache-first and honest — split infra/AI caches with declared staleness windows and accuracy bands (±5-10% infra, ±15-25% AI), a
pricing:coveragegate, and an explicit us-east-1 region warning. - AI track reuses the shared canonical
skills/shared/ai/files (vendored byte-identical) instead of duplicating, and itsai-model-lifecycle.mdcarries the corrected 45-day policy (the #307 clash is not reintroduced). - Drift, pricing-coverage, and 5 golden asserters (~3k lines of independent oracles) pass live. The clarify golden is a BLOCKED run — it pins the completion gate, not just the happy path.
Findings
[P1 — honesty] SKILL.md triggers over-promise the discovery surface. The description advertises "migrate Bicep to Terraform," "migrate ARM templates to Terraform," "discover from Terraform/Bicep/ARM templates," and "live az CLI capture" — but on this head only Terraform is built: extract-bicep.md/extract-arm.md are absent and a .bicep/ARM workspace halts (discover-iac.md §Status), and the live-az path is "lands in step 2" (discover.md §Status). The halt-don't-fake behavior is the right mitigation, but the trigger list matches user intents the skill can't fulfill — a Bicep user is matched, then stopped. Soften the triggers to Terraform-first, naming Bicep/ARM/live-az as not-yet. (A separate PR already adds live-az and softens exactly these triggers; once it lands this narrows to Bicep/ARM only.)
[P2 — consistency] Two stale Status blocks misrepresent completeness in opposite directions.
references/phases/generate/generate.md:123calls itself "skeleton (build step 1)… emitters land in step 6," but the emitter fragmentgenerate-artifacts-infra.md:236says "implemented (build step: Generate)." The emitters are real — the orchestrator's Status under-reports.references/phases/design/design.md:146listsai.mdas "still to come," butreferences/design-refs/ai.mdexists as a real ~8KB rubric — over-reports the gap.
Refresh both; a reviewer trusting Status blocks misjudges completeness both ways.
[P2 — dangling reference] graviton.md is referenced but absent from the azure tree. SKILL.md:21 (and ~6 other spots: compute.md, database.md, clarify-compute.md, clarify-global.md, two knowledge JSONs) reference references/shared/graviton.md's escape path as firing "routinely," but that file is absent from both azure trees — it exists only in the gcp-to-aws sibling. The vendored README declares the skill must be self-contained (runnable lifted-out). These are prose mentions, not hard Load directives, so likely not a runtime halt — but it's a real broken path against the self-contained guarantee. Either vendor graviton.md into the azure tree or inline the escape criteria.
[P3 — disclosed, non-blocking] The run_mode: decide_and_execute Generate consent gate is an _assert that CI binds but never evaluates — honestly disclosed, but the opt-in-to-execute gate has no mechanical teeth and depends on the interpreter honoring it. Worth tracking, since Generate emits Terraform.
[P3 — deferred, confirmed] Cross-skill Azure routing isn't wired: llm-to-bedrock hardcodes gcp-to-aws for Assess and agent-advisor inline-executes gcp phase files via $GCP_BASE, neither with an Azure path. Named as deferred cross-skill work — fine to leave; flagging that a pure-Azure-AI user still touches gcp paths until it lands.
Net
The architecture is sound and the cost-honesty engineering is above the bar. Fix the SKILL.md trigger truthfulness (P1), the two Status blocks (P2), and the dangling graviton.md (P2), and the Terraform path is merge-ready — Bicep/ARM/live-az/AI-route can honestly land as sequenced follow-ups.
… + dangling graviton.md (PR awslabs#304 review) herosjourney's merge-readiness review (verified vs head 1df5cb0): P1 (honesty) — SKILL.md triggers over-promised the discovery surface. Only Terraform is built; Bicep/ARM extractors and the live-az capture fragment are not on this branch (a Bicep/ARM/live-only workspace halts). Fixes: - dropped 'migrate Bicep to Terraform' / 'migrate ARM templates to Terraform' triggers; - 7-phase summary now says discovery is from Terraform, naming Bicep/ARM/live-az as advertised triggers whose extractors are not yet on this branch (sequenced follow-ups); - softened the 'Live-first discovery' philosophy bullet to '(planned)'. P2 (consistency) — two stale Status blocks: - generate.md said 'skeleton (build step 1)… emitters land in step 6' but the emitters are implemented — refreshed to 'implemented (build step: Generate)', describing the infra/AI-only/both tracks and the validated report path. - design.md listed ai.md as 'still to come' but references/design-refs/ai.md exists (~8KB) — corrected to 'on disk'; only licensing.md/gpu-hpc.md/patterns.md remain. P2 (dangling reference) — graviton.md (referenced from SKILL.md + ~6 spots as firing 'routinely') was absent from the azure tree, breaking the self-contained guarantee. Vendored gcp's references/shared/graviton.md + its schema-graviton.md dependency into azure's references/shared/ (both trees, byte-identical) — the azure-private placement matching pricing-cache.md / report-decision-core.md. P3 items (run_mode _assert-no-teeth; cross-skill llm-to-bedrock/agent-advisor gcp- hardcoding) left as disclosed/deferred per the review. Verified both trees: drift OK (401 identical), shared:check OK, frontmatter 7, asserters 15 PASS, lint:model-ids OK, mise run build all 16 tasks clean.
Review findings addressed —
|
leon1418
left a comment
There was a problem hiding this comment.
[🤖 AI review 🤖]
Reviewed 9d57d393154aaa7c32a9987750c00299b3cfa5f0 against 6b48f8ad84fdd3928de41bf116e1e0f3a181cc6c with OCR delegation, an isolated general review, a separate repository review, and candidate verification. Requesting changes for the behavioral contracts below.
The two perspectives accounted for all 413 additions/modifications/deletions (397 GitHub file entries with rename compression). The current head contains changes in both plugin trees. The correction list is consolidated before publication; it includes the earlier fixes' remaining integration gaps.
Validation: 323 Agent Advisor tests and 152 report/policy tests pass per plugin; 15 fixture asserters pass per plugin; frontmatter, vendored/shared, drift, model-ID and pricing-vocabulary checks pass. All 25 source-contract/Heroku tests pass with Terraform 1.13 (the first run selected ambient 1.15.5). These validate source contracts and stored fixtures; I did not run a live Azure migration or deploy resources.
Already present: the IaC Cognitive AI producer, current lifecycle-policy content, refreshed Design/Generate Status blocks, and the missing Graviton references. The deferred Bicep/ARM/live discovery and cross-skill routing do not need to be implemented for this review.
Correction paths: plugin-relative paths below apply to both advisor/plugins/aws-startup-advisor/ and migrate/plugins/migration-to-aws/. Root documentation paths are stated separately. Change the deficient contract sites; preserve producers/consumers already correct.
F1 [P1] Make the declared AI-only route pass the backbone gates
An app-code-only Azure AI project completes Discover and the standalone AI-only Clarify flow. Discover intentionally leaves infrastructure artifacts absent, but Design requires inventory and clusters, Estimate requires infrastructure design and inventory, and Clarify still checks infrastructure-only fields. The documented AI-only chain cannot satisfy its own gates. Confirmed reentry also resets statuses while negative producers leave old optional files in place; artifact-presence routing can retain the previous track or fail the absence gate.
Correction: Carry track conditions through Clarify, Design, Estimate, their fragments/assemblers, the decision gate and Generate validation directory. Gate or adapt the workshop offer for AI-only runs. Preserve mandatory Clarify and execute consent; test fresh and resumed AI-only and mixed runs. On confirmed Discover reruns, retire obsolete phase-owned optional outputs and derived descendants based on current contributions, preserving a currently valid IaC AI producer.
Paths: skills/azure-to-aws/references/phases/clarify/clarify-ai-only.md, skills/azure-to-aws/references/phases/clarify/clarify-assemble.md, skills/azure-to-aws/references/phases/clarify/clarify.md, skills/azure-to-aws/references/phases/design/design-assemble.md, skills/azure-to-aws/references/phases/design/design-infra.md, skills/azure-to-aws/references/phases/design/design.md, skills/azure-to-aws/references/phases/discover/discover-app-code.md, skills/azure-to-aws/references/phases/discover/discover-assemble.md, skills/azure-to-aws/references/phases/discover/discover.md, skills/azure-to-aws/references/phases/estimate/estimate-assemble.md, skills/azure-to-aws/references/phases/estimate/estimate-infra.md, skills/azure-to-aws/references/phases/estimate/estimate.md, skills/azure-to-aws/references/phases/generate/generate.md, skills/azure-to-aws/references/phases/workshop/workshop.md, skills/azure-to-aws/references/shared/schema-discover-ai.md
Acceptance: A fresh AI-only project reaches a decision and, with execute consent, its report without fabricated infra artifacts. Mixed and infra runs retain their gates; decide-complete resume and workshop offer do not require absent files. AI-only validation runs against ai-migration rather than nonexistent terraform/. Confirmed mixed-to-infra-only and mixed-to-AI-only reruns leave the current route's artifact set; a valid current IaC AI contribution is retained.
F2 [P1] Do not prescribe unsupported Redis cluster encryption arguments
The new policy checks a standalone Redis aws_elasticache_cluster and suggests adding both encryption flags. at_rest_encryption_enabled is not in that Terraform resource schema. The new 'good' fixture passes POLICY_OK while carrying an unsupported argument; following the prescribed fix produces invalid Terraform.
Correction: Require an encrypted aws_elasticache_replication_group for Redis rather than inventing cluster arguments. Update checker, fix_hint, authoring posture, skill text, fixtures and tests in both trees; validate the accepted fixture against the supported provider schema.
Paths: skills/tf-best-practices/SKILL.md, skills/tf-best-practices/fixtures/terraform-policy/bad-elasticache-cluster-redis-unencrypted/cache.tf, skills/tf-best-practices/fixtures/terraform-policy/good-elasticache-cluster-redis-encrypted/cache.tf, skills/tf-best-practices/references/security-posture-rules.md, skills/tf-best-practices/scripts/test_validate_terraform_policy.py, skills/tf-best-practices/scripts/validate-terraform-policy.py
Acceptance: The accepted encrypted Redis fixture is valid for the supported AWS provider. The policy's repair guidance cannot recommend an unsupported cluster argument. Encrypted replication groups pass and unencrypted Redis is still rejected in both copies.
F3 [P2] Use one preference envelope across the Azure AI consumers
A mixed infrastructure/AI run confirms a European region and high AI token volume. Mixed Clarify writes global.target_region and top-level AI rows, while AI Design reads design_constraints.target_region with a us-east-1 fallback and Estimate reads ai_constraints.ai_token_volume.value. Recorded region/usage choices can be lost. The AI-only funding field is also nested where Generate reads a top-level field.
Correction: Normalize mixed and AI-only preferences to one documented shape or add explicit adapters. Trace region, all AI rows, agentic fields and startup program status through Design, Estimate, Generate and report; verify user choices survive.
Paths: skills/azure-to-aws/references/phases/clarify/clarify-ai-only.md, skills/azure-to-aws/references/phases/clarify/clarify-ai.md, skills/azure-to-aws/references/phases/clarify/clarify-assemble.md, skills/azure-to-aws/references/phases/design/design-ai.md, skills/azure-to-aws/references/phases/estimate/estimate-ai.md, skills/azure-to-aws/references/phases/generate/generate-artifacts-ai.md, skills/azure-to-aws/references/shared/schema-preferences.md
Acceptance: A mixed run's eu-west-1 region and high token volume are exactly what AI Design and Estimate consume. AI-only and mixed program-status and agentic preferences survive through the generated report/artifacts. Defaults apply only when the canonical field is absent, never because its producer wrote another shape.
F4 [P1] Run Generate dependencies and validation in their owning context
Generate is dispatched to generic-phase-worker-rw through its declared _exec tier. The worker has Read/Grep/Glob/Write/Edit only, but the infra fragment requires a skill invocation and the report fragment requires Python execution. The main-window finish step runs only Terraform validation. The report additionally consumes AI artifacts and an accounting ledger produced later in the declared order.
Correction: Load authoring posture before dispatch, complete emitters/accounting before report rendering, and run mandatory report and Terraform validators in the main window before postconditions. Preserve fail-closed behavior and both execution tracks.
Paths: skills/azure-to-aws/references/phases/generate/generate-artifacts-infra.md, skills/azure-to-aws/references/phases/generate/generate-artifacts-report.md, skills/azure-to-aws/references/phases/generate/generate-assemble.md, skills/azure-to-aws/references/phases/generate/generate.md
Acceptance: A strict rw worker finishes its assigned work without shell or skill calls. Authoring posture is supplied before emission; accounting/AI outputs exist before dependent report content. The main window runs both required validators and refuses completion on REPORT_FAIL or POLICY_FAIL.
F5 [P2] Admit the Terraform forms the extractor actually supports
The workspace contains only azapi_resource resources, JSON Terraform, or downloaded module resources referenced by root module blocks. Entry detection requires an azurerm resource and excludes the downloaded module directory before the extractor is loaded. A valid supported source can be treated as no IaC and stop without reaching the AzAPI/module extraction rules.
Correction: Align source admission and dialect detection with AzAPI, Terraform JSON and resolvable modules, while preserving state/secret exclusions and explicit halts for deferred Bicep/ARM sources. Add cold-start cases for each supported form.
Paths: skills/azure-to-aws/references/phases/discover/discover.md, skills/azure-to-aws/references/phases/discover/discover-iac.md, skills/azure-to-aws/references/shared/extract-terraform.md
Acceptance: AzAPI-only, JSON Terraform and downloaded-module-only supported inputs reach extraction. State files remain unread and Bicep/ARM remain explicitly deferred.
F6 [P2] Accept the provenance values emitted by the fallback routes
A named Azure resource uses an unlisted namespace, such as the documented Microsoft.Maps/accounts example. The producer emits derived_uncorroborated, but the Discover gate rejects it. Even after passing that gate, the model_category Design route is omitted from the Design schema checklist. The advertised non-halting fallback cannot complete honestly.
Correction: Synchronize producer, schemas, validation checklists and fixture vocabulary for derived_uncorroborated and model_category; retain warning and inferred-confidence requirements.
Paths: skills/azure-to-aws/references/phases/design/design-infra.md, skills/azure-to-aws/references/phases/discover/discover.md, skills/azure-to-aws/references/shared/extract-terraform.md, skills/azure-to-aws/references/shared/schema-design-aws.md, skills/azure-to-aws/references/shared/schema-discover-azure.md
Acceptance: The documented Maps/unlisted-namespace path can retain derived_uncorroborated and model_category without failing its own gates. Derived routes remain inferred and keep their provenance warnings.
F7 [P2] Preserve the routing fields consumed by the new rubrics
Terraform declares a Service Bus queue with requires_session=true, or uses separately declared AKS node pools or App Service container/runtime settings. The extraction allowlist has no Service Bus queue or child node-pool row and does not carry several explicit web runtime fields. Only generic SKU/tier/capacity survives for unlisted types. Downstream routing and sizing therefore lose session requirements, pool capacity and runtime evidence despite a structurally valid inventory.
Correction: Audit each currently supported rubric and sizing consumer against the extractor allowlist. Preserve their concrete routing/sizing fields and canonical shapes, including Service Bus entity settings, child node pools, web runtime settings and load-balancer configuration; test representative downstream decisions.
Paths: fixtures/azure-iac-terraform/check_expected_iac_terraform.py, fixtures/azure-iac-terraform/expected-iac-terraform.json, skills/azure-to-aws/references/design-refs/messaging.md, skills/azure-to-aws/references/shared/extract-terraform.md, skills/azure-to-aws/references/shared/schema-discover-azure.md
Acceptance: requires_session=true reaches the messaging rubric and produces the intended FIFO decision. A separately declared node pool retains VM size/count/spot/taints in its parent EKS design. All current rubric and sizing inputs have an extraction source or an explicit unavailable-data treatment.
F8 [P2] Honor an explicit App Service isolation split through Generate
The user selects separate environments per app in Clarify Q-C2. Design's phase manifest permits the split, but its schema checklist unconditionally prohibits site entries and Generate again requires exactly one compute target per plan. The requested isolation either fails validation or is folded away.
Correction: Make the no-split default conditional throughout schema, design, generation/accounting and estimation. With isolation_split=true, emit and price the intended separate environments; retain one plan target when false.
Paths: skills/azure-to-aws/references/phases/design/design-infra.md, skills/azure-to-aws/references/design-refs/index.md, skills/azure-to-aws/references/design-refs/compute.md, skills/azure-to-aws/references/shared/schema-design-aws.md, skills/azure-to-aws/references/phases/estimate/estimate-infra.md, skills/azure-to-aws/references/phases/generate/generate-artifacts-infra.md, skills/azure-to-aws/references/phases/generate/generate-assemble.md
Acceptance: An explicit two-app isolation split produces two corresponding environments and matching cost/accounting. The default unsplit plan still produces one compute target and one cost unit.
F9 [P2] Generate AI artifacts from each workload's source and target
An Azure-hosted application uses direct OpenAI or Anthropic, both explicitly supported discovery/design sources. Generate always emits an AzureOpenAI source branch, defaults AI_PROVIDER to azure_openai and runs comparison against Azure OpenAI. Such users lack that source endpoint/credential, so the default adapter and rollback do not call their existing provider. A traditional-AI target such as Textract is also ignored: the generator never reads target_aws_service and defaults to Bedrock scaffolding instead.
Correction: Derive source adapter, comparison client, credentials and rollback default from per-workload source provenance. Cover azure_openai, direct openai, anthropic and mixed sources without changing the working pre-cutover provider. Dispatch on each design block's target class; implement the selected non-Bedrock service or explicitly defer it in accounting and the report. Do not count unrelated Bedrock scaffolding as completion.
Paths: skills/azure-to-aws/references/phases/generate/generate-artifacts-ai.md, skills/azure-to-aws/references/phases/generate/generate-artifacts-report.md, skills/azure-to-aws/references/phases/generate/generate-assemble.md, skills/azure-to-aws/references/phases/generate/generate.md
Acceptance: Pre-cutover and rollback calls use the actually detected Azure OpenAI, OpenAI or Anthropic source. Mixed-provider workloads do not share an incorrect forced Azure client or credential. A Textract-only or mixed text/document design produces the selected implementation or an explicit visible deferral, with complete per-workload accounting.
F10 [P2] Ship the report validator with the supported skill-only install
The user installs the documented skill directories with npx skills add, including the sibling tf-best-practices skill. The mandatory report validator lives at plugin root and is absent from that install. The standalone skill therefore cannot satisfy the Generate report gate even though its own references are vendored.
Correction: Bundle/resolve the mandatory validator within the distributable skill dependency layout, or explicitly require and validate a complete plugin installation before execution. Test the documented installation layout through the report gate.
Paths: skills/azure-to-aws/references/phases/generate/generate-artifacts-report.md, skills/azure-to-aws/references/phases/generate/generate.md, skills/azure-to-aws/references/vendored/README.md, scripts/validate-migration-report.py
Acceptance: The documented install layout can resolve and execute the mandatory report validator. Missing dependencies produce an explicit preflight error; no skipped validator is reported as a pass.
F11 [P2] Finish the discovery-scope correction on consumer entry surfaces
A reader follows the README's Bicep, ARM, live az or billing-based Azure discovery promise. The README and AGENTS entry still promise those sources, and migrate/README.md describes a live discovery procedure as available. The implemented Discover phase still contains only Terraform and app-code producers and intentionally halts other paths.
Correction: Make all active Azure entry documentation Terraform/app-code first, and consistently label deferred sources. Update advisor/README.md, advisor/AGENTS.md, migrate/README.md and both Azure SKILL.md Defaults; do not implement the deferred capabilities merely to satisfy documentation.
Paths: skills/azure-to-aws/SKILL.md, advisor/README.md, advisor/AGENTS.md, migrate/README.md
Acceptance: Every active Azure source-discovery entrypoint agrees that Terraform/app code are implemented and the other sources are deferred.
F12 [P2] Use the current SQS message-size limit in the eliminator
A Service Bus workload needs messages larger than 256 KiB but no larger than 1 MiB. The hard eliminator incorrectly excludes SQS and forces a broker or S3 claim-check architecture for a payload SQS accepts directly.
Correction: Use the supported SQS limit and byte units in both rubric copies, preserving claim-check/broker handling above that limit; check the boundary without adding an unnecessary component.
Paths: skills/azure-to-aws/references/design-refs/messaging.md
Acceptance: A payload above 256 KiB and at or below 1 MiB does not eliminate SQS solely for size; larger payload handling remains explicit.
F13 [P2] Do not default New Zealand estates abroad because of a nonexistent-region claim
An estate in Azure newzealandnorth accepts the mapped default AWS region. The map sends it to Sydney and claims AWS has no New Zealand region, contrary to both the current region list and this table's same-country preference. This needlessly changes the country of the proposed deployment.
Correction: Use the available New Zealand region as the geographic default, subject to normal service availability and account opt-in checks; retain an explicit justified fallback where a needed service is unavailable.
Paths: skills/azure-to-aws/knowledge/design/azure-region-map.json
Acceptance: The New Zealand row no longer claims an absent AWS region and follows its same-country geographic default, with explicit availability/opt-in fallback.
F14 [P2] Construct complete and collision-safe ARM resource identities
Two invocations of the same local/downloaded module each declare azurerm_service_plan.this with name=var.name, different module inputs, and the same resource group. A top-level AzAPI resource also uses a resource-group parent. The prescribed unresolved name is tf:this in both instances. tf_address is also specified without the module prefix; only an auxiliary tf_module differs. The assembler deduplicates by the resulting identical azure_id, collapsing two real plans into one and understating the estate and cost. The AzAPI recipe separately omits /providers/ for that scope, producing an invalid ARM ID.
Correction: Use a stable fully qualified Terraform/module identity in synthetic unresolved ARM-name segments or as a separate collision-safe identity key. Include module instance addresses in tf_address, and resolve references through the same identity mapping without pretending unresolved names are real Azure names. Use the full ARM type when constructing top-level/extension IDs; only append remaining child segments under an already matching provider-qualified parent.
Paths: skills/azure-to-aws/references/phases/discover/discover-assemble.md, skills/azure-to-aws/references/shared/arm-type-canonicalization.md, skills/azure-to-aws/references/shared/extract-terraform.md, skills/azure-to-aws/references/shared/schema-discover-azure.md
Acceptance: Two instances of the same parameterized module retain two inventory and design entries. References into each module resolve to its own resource and stay stable across reruns. The documented budget example yields a provider-qualified ARM ID. Nested AzAPI children keep exactly one appropriate provider segment and join their parent.
F15 [P2] Apply the endpoint contract per GPT family rather than banning runtime
A GPT-5.6 workload needs the supported runtime/CRIS endpoint, for example Converse with Guardrails and permitted cross-region inference. The frozen fact reference explicitly supports GPT-5.6 runtime/CRIS and Converse Guardrails, but the new Azure design, schema and Generate gates classify every proprietary openai.gpt-* ID as Mantle-only. They require a model change or reject a supported same-model configuration, defeating the stated same-model selection policy.
Correction: Propagate the model/endpoint/feature matrix into Azure selection, schemas, pricing and generation. Permit GPT-5.6 runtime with the proper inference-profile IDs and IAM when its feature and residency constraints match; keep the actual Mantle-only restrictions for GPT-5.5/5.4.
Paths: skills/azure-to-aws/references/phases/design/design-ai.md, skills/azure-to-aws/references/phases/design/design.md, skills/azure-to-aws/references/phases/estimate/estimate-ai.md, skills/azure-to-aws/references/phases/generate/generate-artifacts-ai.md, skills/azure-to-aws/references/phases/generate/generate.md, skills/azure-to-aws/references/shared/pricing-cache.md, skills/azure-to-aws/references/shared/schema-design-aws-ai.md
Acceptance: A permitted GPT-5.6 Converse/Guardrails case keeps the same model and valid CRIS ID. GPT-5.5/5.4 runtime requests remain rejected, and Mantle-specific features still select Mantle. Generated credentials, IAM and endpoint agree with the chosen per-model path.
Provider/API facts were checked against HashiCorp's aws_elasticache_cluster schema (including AWS provider v5.80.0), AWS's current SQS message quotas and Region list, and the GPT-5.6 Terra/Sol model cards. The latter explicitly support runtime/CRIS and Converse; the blanket Mantle-only rule is not the current contract.
Approval remains for the human reviewer; this review does not approve or clear another review.
| _on_failure: _halt_and_inform | ||
| - _check_single_active_phase: true | ||
| _on_failure: _halt_and_inform | ||
| - _check_file_exists: [azure-resource-inventory.json, azure-resource-clusters.json, preferences.json] |
There was a problem hiding this comment.
[🤖 AI review 🤖] F1 [P1] Make the declared AI-only route pass the backbone gates
An app-code-only Azure AI project completes Discover and the standalone AI-only Clarify flow. Discover intentionally leaves infrastructure artifacts absent, but Design requires inventory and clusters, Estimate requires infrastructure design and inventory, and Clarify still checks infrastructure-only fields. The documented AI-only chain cannot satisfy its own gates. Confirmed reentry also resets statuses while negative producers leave old optional files in place; artifact-presence routing can retain the previous track or fail the absence gate.
Carry track conditions through Clarify, Design, Estimate, their fragments/assemblers, the decision gate and Generate validation directory. Gate or adapt the workshop offer for AI-only runs. Preserve mandatory Clarify and execute consent; test fresh and resumed AI-only and mixed runs. On confirmed Discover reruns, retire obsolete phase-owned optional outputs and derived descendants based on current contributions, preserving a currently valid IaC AI producer.
The review summary lists the full affected paths and acceptance checks for both plugin copies.
| continue # memcached / variable-driven / absent engine → exempt | ||
| if _has_own_attr(body, "replication_group_id"): | ||
| continue # node of a replication group — encryption enforced there | ||
| for attr in ("at_rest_encryption_enabled", "transit_encryption_enabled"): |
There was a problem hiding this comment.
[🤖 AI review 🤖] F2 [P1] Do not prescribe unsupported Redis cluster encryption arguments
The new policy checks a standalone Redis aws_elasticache_cluster and suggests adding both encryption flags. at_rest_encryption_enabled is not in that Terraform resource schema. The new 'good' fixture passes POLICY_OK while carrying an unsupported argument; following the prescribed fix produces invalid Terraform.
Require an encrypted aws_elasticache_replication_group for Redis rather than inventing cluster arguments. Update checker, fix_hint, authoring posture, skill text, fixtures and tests in both trees; validate the accepted fixture against the supported provider schema.
The review summary lists the full affected paths and acceptance checks for both plugin copies.
| "ai_monthly_spend": "$500-$2K", // DETECTED | PROPOSED | ||
| "ai_priority": "balanced", // PROPOSED | ||
| "ai_critical_feature": null, // PROPOSED | ||
| "ai_token_volume": "low", // PROPOSED |
There was a problem hiding this comment.
[🤖 AI review 🤖] F3 [P2] Use one preference envelope across the Azure AI consumers
A mixed infrastructure/AI run confirms a European region and high AI token volume. Mixed Clarify writes global.target_region and top-level AI rows, while AI Design reads design_constraints.target_region with a us-east-1 fallback and Estimate reads ai_constraints.ai_token_volume.value. Recorded region/usage choices can be lost. The AI-only funding field is also nested where Generate reads a top-level field.
Normalize mixed and AI-only preferences to one documented shape or add explicit adapters. Trace region, all AI rows, agentic fields and startup program status through Design, Estimate, Generate and report; verify user choices survive.
The review summary lists the full affected paths and acceptance checks for both plugin copies.
|
|
||
| > **Plan-share links are GATED OFF** (landing page not live). Do NOT emit a share link. | ||
|
|
||
| ## Step 5: Validate the rendered report (REQUIRED — mandatory gate) |
There was a problem hiding this comment.
[🤖 AI review 🤖] F4 [P1] Run Generate dependencies and validation in their owning context
Generate is dispatched to generic-phase-worker-rw through its declared _exec tier. The worker has Read/Grep/Glob/Write/Edit only, but the infra fragment requires a skill invocation and the report fragment requires Python execution. The main-window finish step runs only Terraform validation. The report additionally consumes AI artifacts and an accounting ledger produced later in the declared order.
Load authoring posture before dispatch, complete emitters/accounting before report rendering, and run mandatory report and Terraform validators in the main window before postconditions. Preserve fail-closed behavior and both execution tracks.
The review summary lists the full affected paths and acceptance checks for both plugin copies.
|
|
||
| | Dialect | Detection | | ||
| | ----------- | -------------------------------------------------------------------------- | | ||
| | `terraform` | a `**/*.tf` or `**/*.tf.json` file containing a `resource "azurerm_` block | |
There was a problem hiding this comment.
[🤖 AI review 🤖] F5 [P2] Admit the Terraform forms the extractor actually supports
The workspace contains only azapi_resource resources, JSON Terraform, or downloaded module resources referenced by root module blocks. Entry detection requires an azurerm resource and excludes the downloaded module directory before the extractor is loaded. A valid supported source can be treated as no IaC and stop without reaching the AzAPI/module extraction rules.
Align source admission and dialect detection with AzAPI, Terraform JSON and resolvable modules, while preserving state/secret exclusions and explicit halts for deferred Bicep/ARM sources. Add cold-start cases for each supported form.
The review summary lists the full affected paths and acceptance checks for both plugin copies.
| **Migrate to AWS** | ||
|
|
||
| - **`gcp-to-aws`** — Plan a migration from Google Cloud Platform — and OpenAI/Gemini AI workloads — to AWS, directly in your IDE. Runs a guided, multi-phase workflow: discover resources from Terraform/IaC, app code, and GCP billing exports; design an AWS architecture; estimate costs; and generate migration artifacts. AI-provider migration maps OpenAI/Gemini usage to closest-fit Amazon Bedrock model families. | ||
| - **`azure-to-aws`** — Plan a migration from Microsoft Azure — and Azure OpenAI / agentic AI workloads — to AWS. Runs a 7-phase workflow: discover resources from Terraform/Bicep/ARM, a consent-gated read-only `az` CLI capture, app code, and Azure billing exports; design an AWS architecture; estimate costs (1:1 lift and right-sized); optionally reprice what-if scenarios; and generate migration artifacts when the user opts in. App Service → Elastic Beanstalk, AKS → EKS, Cosmos DB → DynamoDB, Azure OpenAI → Bedrock. |
There was a problem hiding this comment.
[🤖 AI review 🤖] F11 [P2] Finish the discovery-scope correction on consumer entry surfaces
A reader follows the README's Bicep, ARM, live az or billing-based Azure discovery promise. The README and AGENTS entry still promise those sources, and migrate/README.md describes a live discovery procedure as available. The implemented Discover phase still contains only Terraform and app-code producers and intentionally halts other paths.
Make all active Azure entry documentation Terraform/app-code first, and consistently label deferred sources. Update advisor/README.md, advisor/AGENTS.md, migrate/README.md and both Azure SKILL.md Defaults; do not implement the deferred capabilities merely to satisfy documentation.
The review summary lists the full affected paths and acceptance checks for both plugin copies.
|
|
||
| | Candidate | Eliminated when | | ||
| | -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | ||
| | **SQS** (standard or FIFO) | the entity's `max_message_size_in_kilobytes` exceeds **256 KB**. SQS caps at 256 KB; Service Bus Premium allows 100 MB. Route to Amazon MQ, or to SQS with an S3 claim-check — and say which | |
There was a problem hiding this comment.
[🤖 AI review 🤖] F12 [P2] Use the current SQS message-size limit in the eliminator
A Service Bus workload needs messages larger than 256 KiB but no larger than 1 MiB. The hard eliminator incorrectly excludes SQS and forces a broker or S3 claim-check architecture for a payload SQS accepts directly.
Use the supported SQS limit and byte units in both rubric copies, preserving claim-check/broker handling above that limit; check the boundary without adding an unnecessary component.
The review summary lists the full affected paths and acceptance checks for both plugin copies.
| "australiasoutheast": { "aws": "ap-southeast-2", "same_country": true, "geo": "Melbourne -> Sydney, Australia" }, | ||
| "australiacentral": { "aws": "ap-southeast-2", "same_country": true, "geo": "Canberra -> Sydney, Australia" }, | ||
| "indonesiacentral": { "aws": "ap-southeast-3", "same_country": true, "geo": "Jakarta, Indonesia" }, | ||
| "newzealandnorth": { |
There was a problem hiding this comment.
[🤖 AI review 🤖] F13 [P2] Do not default New Zealand estates abroad because of a nonexistent-region claim
An estate in Azure newzealandnorth accepts the mapped default AWS region. The map sends it to Sydney and claims AWS has no New Zealand region, contrary to both the current region list and this table's same-country preference. This needlessly changes the country of the proposed deployment.
Use the available New Zealand region as the geographic default, subject to normal service availability and account opt-in checks; retain an explicit justified fallback where a needed service is unavailable.
The review summary lists the full affected paths and acceptance checks for both plugin copies.
| needs to decide whether the resource costs money. Derivation is the design, not a | ||
| fallback — but the table is consulted FIRST, because the cases where a guess goes wrong | ||
| are enumerated there. | ||
| 2. **Resolve `name`** from the block's `name` attribute. When it is an expression |
There was a problem hiding this comment.
[🤖 AI review 🤖] F14 [P2] Construct complete and collision-safe ARM resource identities
Two invocations of the same local/downloaded module each declare azurerm_service_plan.this with name=var.name, different module inputs, and the same resource group. A top-level AzAPI resource also uses a resource-group parent. The prescribed unresolved name is tf:this in both instances. tf_address is also specified without the module prefix; only an auxiliary tf_module differs. The assembler deduplicates by the resulting identical azure_id, collapsing two real plans into one and understating the estate and cost. The AzAPI recipe separately omits /providers/ for that scope, producing an invalid ARM ID.
Use a stable fully qualified Terraform/module identity in synthetic unresolved ARM-name segments or as a separate collision-safe identity key. Include module instance addresses in tf_address, and resolve references through the same identity mapping without pretending unresolved names are real Azure names. Use the full ARM type when constructing top-level/extension IDs; only append remaining child segments under an already matching provider-qualified parent.
The review summary lists the full affected paths and acceptance checks for both plugin copies.
| keeps the OpenAI SDK; only the base URL (`.../openai/v1/responses`), credential (a Bedrock API | ||
| key/token provider), model ID, and IAM (`bedrock-mantle:*`) change. Read | ||
| `references/shared/openai-on-bedrock.md` for exact values. If the source uses Chat Completions, | ||
| plan a reshape to `responses.create`. No Converse fallback exists for proprietary GPT models |
There was a problem hiding this comment.
[🤖 AI review 🤖] F15 [P2] Apply the endpoint contract per GPT family rather than banning runtime
A GPT-5.6 workload needs the supported runtime/CRIS endpoint, for example Converse with Guardrails and permitted cross-region inference. The frozen fact reference explicitly supports GPT-5.6 runtime/CRIS and Converse Guardrails, but the new Azure design, schema and Generate gates classify every proprietary openai.gpt-* ID as Mantle-only. They require a model change or reject a supported same-model configuration, defeating the stated same-model selection policy.
Propagate the model/endpoint/feature matrix into Azure selection, schemas, pricing and generation. Permit GPT-5.6 runtime with the proper inference-profile IDs and IAM when its feature and residency constraints match; keep the actual Mantle-only restrictions for GPT-5.5/5.4.
The review summary lists the full affected paths and acceptance checks for both plugin copies.
…xed; false positives noted) Triaged the 15-finding automated review (leon1418, verified vs head 9d57d39). Fixed the genuine bugs; documented push-backs on false positives / deferred items. REAL bugs fixed (both trees): F2 [P1] Redis cluster encryption arg — verified against the AWS provider schema (terraform providers schema -json): at_rest_encryption_enabled is NOT a valid argument on aws_elasticache_cluster (only transit_encryption_enabled is; at-rest requires aws_elasticache_replication_group). Our checker required it and the 'good' fixture set it, so following our guidance produced un-appliable Terraform. Fixed the checker to require transit_encryption_enabled on the standalone Redis cluster and direct at-rest to a replication group; corrected the fixture, fix_hint, header docstring, security-posture-rules.md, and the test comment. 60 policy tests pass. F13 [P2] New Zealand region — the region map claimed "No AWS region in New Zealand" and defaulted newzealandnorth to Sydney. AWS Asia Pacific (New Zealand), ap-southeast-6, has been GA since 2025-09-30 (confirmed via the internal region registry). Fixed to same-country ap-southeast-6 with a service-availability/opt-in fallback note. F6 [P2] derived_uncorroborated provenance — producers emit azure_type_provenance: "derived_uncorroborated" for an unrecognized namespace (the advertised non-halting fallback), but the discover.md gate _assert only allowed {table, derived, user_confirmed} — so that resource failed the gate. Added derived_uncorroborated to the enum. (The model_category half of the finding was already correct in the design schema — no change.) F5 [P2] Terraform-form admission — extract-terraform.md handles azapi_resource and downloaded modules, but discover-iac.md dialect detection only fired 'terraform' on a resource "azurerm_ block, so an AzAPI-only or module-only repo was seen as "no IaC" and exited before the extractor. Broadened the detection trigger to azurerm_/azapi_/module blocks. F7 [P2] Routing fields dropped by the extractor allowlist — messaging.md routes ServiceBus queues with requires_session=true to SQS FIFO, but the allowlist had no ServiceBus rows at all, and no separate agentPools (child AKS node pool) row. Added Microsoft.ServiceBus/namespaces{,/queues,/topics} (requires_session, size/TTL/dedup) and managedClusters/agentPools rows. F14 [P2] Collision-safe ARM identity — tf_address was <type>.<local_name> without the module prefix, and the tf: unresolved-name segment was the bare local name, so two invocations of the same module (each azurerm_service_plan.this, same RG) reconstructed the SAME azure_id and the assembler collapsed two real resources into one. Made tf_address the full module-qualified address and module-qualified the tf: segment. False positives (no change; will note on the PR): - F12 [P2] SQS message size — our rubric correctly caps SQS at 256 KB and routes larger payloads to claim-check/Amazon MQ. The finding's "SQS accepts up to 1 MiB directly" is incorrect (that is Service Bus's limit). Deferred / disclosed (out of this PR's scope per the human reviewer's sequencing): F1, F4, F8, F9, F15, F3 — AI-only backbone-gate carry-through, Generate worker-context ordering, App Service isolation split, per-workload AI source branching, GPT-5.6 CRIS/Mantle classification, and preference-envelope normalization. These are the deferred AI-route / cross-skill follow-ups; F11's doc-scope tail (advisor/AGENTS.md + migrate/README.md Terraform-first) is also included in this commit. Verified both trees: drift OK (401 identical), frontmatter 7, asserters 15 PASS, lint:model-ids OK, 60 policy tests + 92 report tests pass, mise run build all 16 tasks clean, dprint clean.
Response to the automated (open-code-review) findings —
|
…, not a false positive) Correcting my earlier triage: F12 from the open-code-review is REAL. Amazon SQS raised its max message payload from 256 KiB to 1 MiB on 2025-08-04 (standard AND FIFO, all commercial Regions + GovCloud; Lambda event source mapping updated to match) — https://aws.amazon.com/about-aws/whats-new/2025/08/amazon-sqs-max-payload-size-1mib/ messaging.md's hard eliminator still capped SQS at 256 KB, so a Service Bus workload with payloads between 256 KiB and 1 MiB was wrongly excluded from SQS and pushed to an unnecessary Amazon MQ broker or S3 claim-check. Fixed the boundary to 1 MiB (1024 KiB) in both rubric copies: <=1 MiB maps to SQS directly (no claim-check); only >1 MiB routes to Amazon MQ / S3 claim-check. Verified both trees: drift OK (401 identical), asserters 15 PASS, fmt clean, mise run build all 16 tasks clean.
Correction on F12 — it's a real bug, now fixed (
|
…ws#352) * feat(aws-startup-advisor): add live az CLI discovery to azure-to-aws azure-to-aws discovered only from Terraform/app code; a workspace with no azurerm_* IaC (the common startup case) halted. This adds the live `az` capture path the SKILL.md advertised and discover.md reserved a slot for — a second producer of the same inventory contract, keyed on canonical Microsoft.* ARM types (no translation table; `az` emits them natively). Two-part design, honoring discover.md's execution model (_interactive: false, _exec: { _agent: rw } — the dispatched worker is file-only and cannot prompt): - Part A (main-window pre-work, invoked from discover.md _preconditions): preflight `az`, consent gate, read-only `az resource list` fast-path with per-service enrichment fallthrough, writing $MIGRATION_DIR/live-capture/ + manifest.json. Read-only allowlist; never captures app-setting/connection- string/Key-Vault values or mints a token; explicit --subscription scope. - Part B (dispatched `live` fragment, this file): file-only parse of live-capture/ into azure-resource-inventory.json / clusters, edge inference from resolved ids, simplified clustering, and IaC-merge-with-drift when Terraform also ran. Entry-guards on the manifest — absent ⇒ no-op. Wiring: registered the `live` fragment (triggers on live-capture/manifest.json), added the _live_capture_prework precondition step, and relaxed the _unrecoverable assertion + two _postconditions so a live-only run is a valid producer (a no-IaC/no-code workspace with a reachable subscription now discovers instead of halting). Updated the discover.md Status/live-note and three SKILL.md claims that said live `az` was "not yet implemented." Corrections vs the source draft (awslabs/startups#304 comment): schema references point to the real schema-discover-azure.md (schema-discover-iac.md does not exist here); azure_type_provenance uses the closed set {table,derived,derived_uncorroborated,user_confirmed} and ai_source the closed set {azure_openai,openai,anthropic,both,other}; warnings use the closed schema-discover-azure.md vocabulary — so it satisfies discover.md's postconditions. Provenance: capture commands live-validated against a real subscription (az-cli 2.90.0) for the preflight, `az resource list` fast-path, Container Apps, PostgreSQL Flexible Server, Storage, ACR, Log Analytics, and Cognitive Services rows. Rows for services absent from the test subscription (webapp, aks, vm, mysql, cosmos, redis, service bus, event hubs, vnet, key vault, ML, dns) carry ported command shapes and should get one live pass before being relied on. Verified: mise run build exit 0 (markdownlint, manifests, v1.0.0 spec, skill-parity, gitleaks all green; description 1012 chars, within the 1024 limit). * fix(aws-startup-advisor): declare the live fragment in the discover assembler _reads Self-review of aws#352: discover.md registers `live` as a producing fragment, but discover-assemble.md's _reads block listed only iac + app-code. The assembler BODY already handles live `az` (it sits highest in the source-precedence rule and in the drift logic), so this was a declaration gap, not a logic gap — add `live` to _reads so the assembler's inputs match the phase's registered fragments. * fix(aws-startup-advisor): align discover-live.md to the real azure-to-aws schema Addresses the review of aws#352 — I had ported the source draft's vocabulary without validating it against this skill's schema/postconditions. Each of these would have failed the exact no-Terraform run the PR unlocks: Blocking: - Edge types: the draft's `network_membership` / `data_dependency` are not in schema-discover-azure.md § Typed edges, and the phase postcondition rejects any edge type not in that table. Use `network` and `data_ref` (kept `hosted_on`). - azure_id: the fast-path `az resource list --query` (and per-service rows) omitted the ARM id, so a live-only inventory could not fill the required `azure_id`. Project `id:id`; Step 3 copies it into `azure_id`. - Live-only vs Terraform-shaped postconditions: made the `iac_metadata` assertion conditional on an IaC dialect contributing, and split the producer-agreement assertion so a live-sourced Cognitive Services/ML resource satisfies it via detection_signals[].method live_az, infrastructure[] keyed by azure_id, sources_analyzed.terraform false, inferred_from_iac false, and no tf_address requirement. (models[].detected_via has no live value in the schema — the live signal rides on detection_signals[], not detected_via.) Should-fix: - Warnings: `reader_role_missing` is not in the closed vocabulary; a capture failure is recorded in the manifest note + live_metadata.capture_warnings, not warnings[]. Drift: written as resources[].drift `{field, values:[{source, value}], won:live}` per the schema, not a `live_metadata.drift.config_conflicts` shape with iac_value/live_value. - SQL row: `az sql db list` requires `--resource-group` as well as `--server`. Log Analytics row 11b marked mode E. Added `sql` to the not-yet-live-tested list. Nit: - confidence is the enum `deterministic|measured|inferred|billing_inferred`; live `az` without metrics is `inferred`, not `0.99`. Also (self-review, prior commit): declared the `live` fragment in the assembler _reads. Verified: markdownlint 0 issues, mise run build exit 0. * fix(aws-startup-advisor): id:id in every live capture command + add live_metadata to the schema Follow-up review of aws#352, two should-fixes: 1. The prior commit added id:id only to the fast path and the SQL row, with a prose note saying "each per-service --query must project id:id, stated once." But the cells ARE the commands, and rows that run only on fallthrough (service bus, event hubs, dns) never inherit the fast-path id — so a per_service run captured no ARM id and the azure_id postcondition would fail. Put id:id inside every --query in the capture table, including both steps of the SQL and Cognitive Services (account + deployment) rows. 2. The fragment wrote a top-level live_metadata object plus metadata.clustering_mode and resources[].unmanaged_by_iac / not_found_live, none of which the inventory schema (schema-discover-azure.md) defined — and the assembler owns the contract from that schema. Added them to the schema: a `## live_metadata` section (the live counterpart of iac_metadata, documenting capture_warnings/derived_types/ drift-rollup), clustering_mode under metadata, and the two optional resource flags on resources[]. So the live fragment's outputs are now schema-sanctioned rather than off-contract. Verified: markdownlint 0 issues, mise run build exit 0. * fix(aws-startup-advisor): project resourceGroup on every live capture command Follow-up review of aws#352 (same class as the id:id fix). Step 3 fills the inventory's required resource_group from the captured resourceGroup field, and the phase postcondition requires resource_group on every entry — but only the fast path and the SQL server list projected it. On a per_service fallthrough the capture had an ARM id but no resourceGroup, so Step 3 could not fill resource_group and the postcondition would fail the run. Add resourceGroup:resourceGroup to every top-level --query in the capture table (including az sql db list and the Cognitive Services deployment list). Nested sub-projections (AKS agentPoolProfiles[], vnet subnets[]) are left as name-only — they are not top-level resources and carry no id/resourceGroup of their own. Dropped the now-stale "stated once rather than in every cell" note since every row now carries both required fields explicitly. Verified: markdownlint 0 issues, mise run build exit 0. * fix(aws-startup-advisor): address 6 P2 review findings on live az discovery All six confirmed valid against the actual DSL contract and schema, not just plausible-sounding review prose: 1. discover.md's _preconditions named a _live_capture_prework check kind that does not exist in INTERPRETER.md's closed vocabulary (_check_phase_completed/_check_single_active_phase/_check_file_exists/ _validate_json/_assert) and had no _on_failure. Replaced with a real _assert predicate; moved the actual main-window action into prose in the phase body (the DSL has no check kind for 'perform an action,' only for verifying one). 2. A confirmed re-entry reuses (per INTERPRETER.md's single-run-directory rule), so declining live capture on a re-entry left the prior attempt's live-capture/manifest.json in place, and the live fragment's existence-based trigger fired on it anyway. Added an explicit re-entry check that invalidates/deletes the stale manifest on decline. 3. Several capture rows project field names/units that diverge from the canonical config keys extract-terraform.md defines for the same ARM type (sku/capacity -> sku_name/worker_count for App Service Plans; storageGb -> storage_mb, GB->MB, for Postgres/MySQL; etc.), so sizing and drift comparisons against an IaC-sourced entry would silently read those fields as absent. Added a rename/reshape table. Also fixed row 16's deployment capture: the --query projection already flattens properties.model.name into a plain string under the key 'model', but the prose read 'model.name' as if model were still an object. 4. The fast-path/enrichment merge joined on name+type, which is not unique within a subscription, even though every row already projects the full ARM id needed for an exact join. Changed the join key to azure_id. 5-6. schema-discover-ai.md's infrastructure[] only defined a Terraform-shaped entry, and discover.md's producer-agreement postcondition had two clauses that required sources_analyzed.terraform/inferred_from_iac to be both true (IaC resource present) and false (live resource present) in the same profile - unsatisfiable by any run with both. Added a live-sourced infrastructure[] entry shape (azure_id-keyed, no address/file), documented profile_source/sources_analyzed as OR'd across every qualifying resource rather than assigned exclusively per producer, and rewrote the postcondition plus discover-assemble.md's/ discover-app-code.md's merge rules to match. * fix(aws-startup-advisor): check ml extension presence non-interactively before row 17 Live-tested row 17 (az ml workspace list) against a real subscription without the ml extension installed: it hangs on an interactive Y/n dynamic-install prompt with no clean error, inside a 120s timeout with no way to answer - the exact same failure mode the file already warns about for az graph query. The doc said 'only if the ml extension is present' but never said how to check that without running the command first. Added an 'az extension list' pre-check (safe, non-interactive, names no resource) before row 17, matching the graph-query guardrail already in this file. * docs(aws-startup-advisor): README catalog row still said no live az capture SKILL.md's frontmatter description, top build-status blockquote, and Philosophy bullet were already updated (in the earlier merge-conflict resolution against aws#351) to say the live-az capture path is implemented, but the plugin README's azure-to-aws catalog row was never touched and still said discovery does not read 'a live Azure CLI capture.' Swept the rest of the plugin (SKILL.md itself, plugin.json keywords, other skills' cross-references, the top-level repo README) for the same stale claim - README.md's catalog row was the only remaining gap. * fix(aws-startup-advisor): address 4 P2 review findings on live az discovery - Reconcile IaC-declared and live-observed AI infrastructure entries that describe the same deployed resource: IaC-sourced infrastructure[] entries now carry the resource's reconstructed azure_id when resolvable, and the assembler merges an IaC entry with a live entry sharing that azure_id into one entry (live wins on config disagreement) instead of always treating them as two distinct resources. - Add a closed-vocabulary warning code (enrichment_id_unmatched) for the case where a live-capture enrichment row's id has no matching az resource list fast-path entry, so discover-live.md's diagnostic instruction can actually satisfy the phase's closed-warnings-vocabulary postcondition. - Mark Event Hubs (row 13) as an enrichment (E) row so kafka_enabled is captured on the successful az resource list fast path, not only in the per-service fallthrough — otherwise a fast-path Event Hubs namespace always routed to Kinesis regardless of its actual Kafka protocol setting. - Complete the live-capture field normalization table: rename sku->sku_name and haMode->high_availability for Postgres/MySQL flexible servers (rows 7/8), and kind->account_kind for storage accounts (row 11) — without the latter, fast-path-services.json's FileStorage exception could never match a live-captured storage account. * fix(aws-startup-advisor): resolve _init/_preconditions ordering cycle, finish live capture field normalization - discover.md: the pre-dispatch main-window action (live az consent + capture) runs BEFORE _preconditions, per this phase's own design -- but this phase carries _init: true, and INTERPRETER.md's generic _exec contract runs _init state setup AFTER _preconditions. On a cold start this is a dependency cycle: capture needs $MIGRATION_DIR, which only _init creates, but _init normally doesn't run until after the gate that requires capture to have already happened. Fixed by performing _init state setup as the pre-dispatch action's new Step 0, before the re-entry check -- _init is idempotent (a resumed run just reads the existing .phase-status.json), so this doesn't double-initialize or re-prompt resume-vs-fresh; the later "Step: Run the phase" _init call is now a no-op re-confirmation for this phase, documented as such. - discover-live.md: completed the field-normalization table with the remaining gaps a prior partial fix left uncovered -- rows 1/3 (plan/runtime/ app-setting/https aliases), 4 (Kubernetes version/node-pool fields, preserving every captured pool), 4a (replica bounds), 6 (SQL database sku/size), 7 (haMode must populate the nested high_availability.mode shape, not replace the object with a scalar), 11 (sku needs SPLITTING into account_tier + account_replication_type, not a straight rename -- and 11's captured accessTier 'tier' field is a different Azure concept from the canonical account_tier, despite the similar name), and 11b (retention_in_days). Also updated Step 4's edge-inference table to read the correct field names post-normalization (e.g. config.service_plan_id, not the raw sku/plan capture key) since edges are inferred after normalization runs. * fix(aws-startup-advisor): Step 4 edge table must read post-normalization field names for Redis and VM rows Redis row 10's edge-inference entry named the raw --query source path (subnetId) instead of the flat capture key that path is projected into (config.subnet) -- Step 3 only ever writes config.subnet, so a literal subnetId lookup in Step 4 can never find the value and the redis->vnet network edge could never be emitted. Fixed that row and proactively fixed the same latent mismatch on the VM row (row 5), which named networkProfile.networkInterfaces[0].id instead of its flat capture key config.subnet (note: despite the key's name it holds a NIC id, not a subnet id -- the existing resolution note on that row is unchanged). --------- Co-authored-by: Logan Kleier <lkleier@amazon.com>
What
Adds the
azure-to-awsmigration skill to both plugin trees (aws-startup-advisorandmigration-to-aws), built as "gcp-to-aws's product surface on heroku-to-aws's DSL machinery." It takes an Azure estate (Terraform IaC and/or application code) through Discover → Clarify → Design → Estimate → (opt-in) Generate, producing an AWS design, a cost estimate, and — when the user opts intodecide_and_execute— Terraform + a migration report.Includes a self-contained AI/agentic workload track (Azure OpenAI / Cognitive Services → Amazon Bedrock, with traditional-AI capabilities routed to Textract/Rekognition/etc.), mirroring gcp's self-contained AI subsystem.
Why
Azure is the remaining major source cloud without a first-class migration skill. This lands it on the shared, already-converged AWS-side product shape (decide-vs-execute gate,
run_mode, Assess handoff viaDECISION.md, tf-policy gate), so it plugs into the existing gate/report contracts rather than inventing new ones. Design posture: mirror gcp; diverge only for a documented Azure fact (ARM-ID edges, x86-default for Windows/.NET, App Service Plan fan-in), each recorded as a decision.Validation
claude -p, real plugin) across 14 Azure sample repos (8 infra, 2 AI iac+appcode, 4 AI app-code-only). 13/14 completed all phases withrun_mode: decide_and_executeand a full artifact set; the 1 miss was shared-git-tree corruption from concurrent runs, not a skill defect. Confirmed behaviors: App Service Plan fan-in preserved (1 EB env for N apps), Document Intelligence → Textract lands inservices_not_estimated[], Azure OpenAI gpt-4o → Bedrock Claude with honest cross-family caveats.mise run build:lint:md,lint:types,lint:frontmatter,shared:check,drift:check,fixtures:assert,fixtures:check,pricing:*,testall pass.fmt:checkresidual is limited to 4 pre-existing non-skill files (repo-root READMEs + gcpdesign-refs/index.md) that are unformatted onmaintoday — out of this skill's scope.Also in this branch
mainup to date (incl.run_id/owning_skilltelemetry seeding); resolved thephase-status.schema.jsoncollisions as a union (keptrun_mode, added the new telemetry keys).skills/shared/clarify/(region/compliance/availability/cost-appetite/multi-cloud), the gcp-to-aws Clarify fragment restructure, two closed gaps on this skill (compliance was never asked; no user-geography question), and the what-if workshop sidebar buildout.faa0ce88): addressed @herosjourney's review — fixed the IaC-only AI-producer bug (theiac_cognitiveproducer + a non-vacuous producer-agreement gate; an infra-only Azure-OpenAI estate no longer drops its AI layer / understates the estimate), narrowed the AI trigger phrases so they don't collide withgcp-to-aws, and addedazure-to-awstoadvisor/README.md/setup.md/ the three router skills' frontmatter. bandit/checkov bot findings were already resolved in2d6a1f2(nosec + checkov skip-path, mirroring gcp).Open items (why this is still a DRAFT)
llm-to-bedrockhardcodesgcp-to-awsfor Assess, andagent-advisor'smigration-plan.mdinline-executesgcp-to-aws's phase files via$GCP_BASE— both with no Azure path. Per @herosjourney's review we are not building a second delegated Assess path; azure's self-contained AI track already covers Azure-sourced AI workloads. The minimal fix (source-cloud detection inllm-to-bedrockPhase A routing to azure's AI track, or an explicit "assumes GCP/generic; Azure users → azure-to-aws" note) is a cross-skill change for whoever picks up the feat(llm-to-bedrock): accept Decide-complete as sufficient Assess handoff #290 / A0-router work. Named here so theagent-advisorinstance isn't a hidden follow-up.796ca5f(embeddings pricing) — @herosjourney read it: scoped to the AI pricing tables, no schema/cross-skill surface; only a pricing-accuracy check against current Bedrock/Azure OpenAI pages is worth doing before ready. Low risk.generatefinish-fragment extraction — with three skills (gcp/heroku/azure) now authoring the same "Finish Generate in the main window" step, @herosjourney recommends extracting toshared/as the next shared-infrastructure work. Not blocking this PR.Remaining before marking ready (from the review): update the consumer-facing
advisor/README.md+setup.md— done (faa0ce88); the optionalvalidate-producer-agreement.pyauthoring-time preflight is deferred (the non-vacuous_assertalready closes the specific hole).