Skip to content

feat(azure-to-aws): add Azure→AWS migration skill (infra + AI/agentic tracks) - #304

Merged
icarthick merged 84 commits into
awslabs:mainfrom
icarthick:feat/azure-to-aws
Sep 23, 2026
Merged

icarthick merged 84 commits into
awslabs:mainfrom
icarthick:feat/azure-to-aws

Conversation

@icarthick

@icarthick icarthick commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

What

Adds the azure-to-aws migration skill to both plugin trees (aws-startup-advisor and migration-to-aws), built as "gcp-to-aws's product surface on heroku-to-aws's DSL machinery." It takes an Azure estate (Terraform IaC and/or application code) through Discover → Clarify → Design → Estimate → (opt-in) Generate, producing an AWS design, a cost estimate, and — when the user opts into decide_and_execute — Terraform + a migration report.

Includes a self-contained AI/agentic workload track (Azure OpenAI / Cognitive Services → Amazon Bedrock, with traditional-AI capabilities routed to Textract/Rekognition/etc.), mirroring gcp's self-contained AI subsystem.

Why

Azure is the remaining major source cloud without a first-class migration skill. This lands it on the shared, already-converged AWS-side product shape (decide-vs-execute gate, run_mode, Assess handoff via DECISION.md, tf-policy gate), so it plugs into the existing gate/report contracts rather than inventing new ones. Design posture: mirror gcp; diverge only for a documented Azure fact (ARM-ID edges, x86-default for Windows/.NET, App Service Plan fan-in), each recorded as a decision.

Validation

  • First real end-to-end capability run: exercised the skill headless (claude -p, real plugin) across 14 Azure sample repos (8 infra, 2 AI iac+appcode, 4 AI app-code-only). 13/14 completed all phases with run_mode: decide_and_execute and a full artifact set; the 1 miss was shared-git-tree corruption from concurrent runs, not a skill defect. Confirmed behaviors: App Service Plan fan-in preserved (1 EB env for N apps), Document Intelligence → Textract lands in services_not_estimated[], Azure OpenAI gpt-4o → Bedrock Claude with honest cross-family caveats.
  • Gates green (both trees, by exit code): frontmatter, cross-plugin drift (374 identical / 28 allowlisted), shared-vendored check, fixture asserters (13: 9 golden + 4 smoke), pricing-coverage.
  • mise run build: lint:md, lint:types, lint:frontmatter, shared:check, drift:check, fixtures:assert, fixtures:check, pricing:*, test all pass. fmt:check residual is limited to 4 pre-existing non-skill files (repo-root READMEs + gcp design-refs/index.md) that are unformatted on main today — out of this skill's scope.

Also in this branch

  • Merged main up to date (incl. run_id/owning_skill telemetry seeding); resolved the phase-status.schema.json collisions as a union (kept run_mode, added the new telemetry keys).
  • Markdownlint + dprint cleanup across the azure surface; a small fix to make the azure estimate asserter's vocabulary regex tolerant of dprint column padding.
  • Merged feat(migrate): Sync migrate/ from GitFarm #6 (herosjourney): shared Clarify questions in skills/shared/clarify/ (region/compliance/availability/cost-appetite/multi-cloud), the gcp-to-aws Clarify fragment restructure, two closed gaps on this skill (compliance was never asked; no user-geography question), and the what-if workshop sidebar buildout.
  • Review round (faa0ce88): addressed @herosjourney's review — fixed the IaC-only AI-producer bug (the iac_cognitive producer + a non-vacuous producer-agreement gate; an infra-only Azure-OpenAI estate no longer drops its AI layer / understates the estimate), narrowed the AI trigger phrases so they don't collide with gcp-to-aws, and added azure-to-aws to advisor/README.md / setup.md / the three router skills' frontmatter. bandit/checkov bot findings were already resolved in 2d6a1f2 (nosec + checkov skip-path, mirroring gcp).

Open items (why this is still a DRAFT)

  1. Cross-skill source-cloud routing (was Shared wiring to converge before an azure-to-aws skill lands #292 item 3(b)). llm-to-bedrock hardcodes gcp-to-aws for Assess, and agent-advisor's migration-plan.md inline-executes gcp-to-aws's phase files via $GCP_BASE — both with no Azure path. Per @herosjourney's review we are not building a second delegated Assess path; azure's self-contained AI track already covers Azure-sourced AI workloads. The minimal fix (source-cloud detection in llm-to-bedrock Phase A routing to azure's AI track, or an explicit "assumes GCP/generic; Azure users → azure-to-aws" note) is a cross-skill change for whoever picks up the feat(llm-to-bedrock): accept Decide-complete as sufficient Assess handoff #290 / A0-router work. Named here so the agent-advisor instance isn't a hidden follow-up.
  2. Commit 796ca5f (embeddings pricing) — @herosjourney read it: scoped to the AI pricing tables, no schema/cross-skill surface; only a pricing-accuracy check against current Bedrock/Azure OpenAI pages is worth doing before ready. Low risk.
  3. Shared generate finish-fragment extraction — with three skills (gcp/heroku/azure) now authoring the same "Finish Generate in the main window" step, @herosjourney recommends extracting to shared/ as the next shared-infrastructure work. Not blocking this PR.

Remaining before marking ready (from the review): update the consumer-facing advisor/README.md + setup.md — done (faa0ce88); the optional validate-producer-agreement.py authoring-time preflight is deferred (the non-vacuous _assert already closes the specific hole).

…/ai/

A third migration skill (azure-to-aws) needs the same Bedrock-side AI content
gcp-to-aws carries, and a copy would be a drift surface — `shared:check` only
protects a file that has a canonical home. So eight files move to the
plugin-neutral `skills/shared/ai/` and are vendored back into gcp-to-aws at
`references/vendored/ai/`, byte-identical, the same lockfile pattern the DSL
and estimate schemas already use.

Promoted (source-cloud agnostic by inspection — the target is Bedrock or
AgentCore, and the mapping keys off the SDK the code uses, not its host):

  ai-model-lifecycle.md  ai-migration-guardrails.md  bedrock-quotas.md
  ai-openai-to-bedrock.md  ai-anthropic-to-bedrock.md
  design-ref-harness.md  design-ref-agentic-to-agentcore.md
  sdk-capability-map.json

Deliberately NOT promoted: `ai-gemini-to-bedrock.md` and `schema-discover-ai.md`,
both written against one provider's SDK surface, and `ai.md`, which is gcp's
category rubric.

`sync-vendored-shared.ts` needed no change: it walks each skill's vendored side
and requires a canonical match, so partial vendoring works and skills need not
vendor identical subsets.

Two prose lines in ai-openai-to-bedrock.md said "GCP-hosted applications"; they
now name the source cloud generically and state that Azure OpenAI routes here
too, since the Bedrock target does not depend on which endpoint served the
calls. ai-model-lifecycle.md's "Mapping Guides" heading no longer names the
one guide that stayed behind in gcp.

Consumers updated in both plugin trees: gcp's SKILL.md conditional-load table
and file tree, design-refs/{index,ai}.md, phases/{design,discover,clarify,
estimate}, shared/pricing-{cache,fallback}.md, and agent-advisor's
migration-plan refs. Bare same-directory filenames became explicit paths —
those files are no longer siblings.

Two hardcoded paths outside the skills also moved: agent-advisor's
`test_scoring.py` lifecycle drift-guard, and `model-id-lint.py`'s catalog
allowlist. The lint now collapses `skills/<skill>/references/vendored/<rel>`
to `skills/shared/<rel>` before matching, so one canonical allowlist entry
covers every vendored mirror — otherwise the next skill to vendor a catalog
file would fail the lint until someone remembered to add its path, which is
the per-copy manifest vendoring exists to avoid.

A documented gap: the promoted files still reference
`references/shared/pricing-{cache,fallback}.md`, which stay per-skill because
pricing-cache.md doubles as gcp's AWS-infrastructure rate card. The new
`skills/shared/ai/README.md` states that as a consumer contract so it is not a
latent dangling reference.
`phase-status.schema.json` sets `additionalProperties: false`, so a skill whose
artifact-generating phase is opt-in has nowhere to record that the user opted
in. gcp-to-aws already works this way and azure-to-aws will gate Generate on it.

Added as an OPTIONAL enum (`decide` | `decide_and_execute`) rather than a
defaulted one, because the absent state is load-bearing: absent means the
decision gate has not been presented yet, and must never read as consent.
Skills with an unconditional Generate keep omitting the key, which is why this
is not `default: "decide"`.

The description also pins the write ordering — `decide_and_execute` is written
BEFORE the generate phase loads — so a session that dies mid-Generate resumes
as an Execute run instead of re-asking a question the user already answered.

Vendored copies re-synced.
Step 1 of the build sequence: prove the wiring before any content lands. Seven
phases — discover → clarify → design → estimate → generate on the backbone, plus
the `workshop` and `feedback` sidebars — each with real fragments, exactly one
assembler, gates, re-entry guards, and the artifact graph the whole skill hangs
off. `lint:frontmatter` reports 7 phase files checked and passes in both trees.

The non-zero count is the point. `parse.ts` returns null for a file that does not
start with `---`, so a prose-shaped skill passes `lint:frontmatter` green while
being completely unvalidated — that is exactly what gcp-to-aws does today (0 phase
files checked). Building this as a gcp-style clone would have shipped a skill the
team believed was DSL-checked and was not.

Structural decisions the skeleton locks in, because retrofitting them changes
every downstream contract:

- **The artifact graph.** `azure-resource-inventory.json` + `azure-resource-
  clusters.json` → `preferences.json` → `aws-design.json` → `estimation-infra.json`
  → the generate set, with `scenarios/index.json` off the workshop sidebar. Every
  phase's `_input` resolves to an upstream `_produces`; the validator enforces it.
- **ARM as the canonical vocabulary.** `azure_id` is the full ARM resource ID and
  `azure_type` a `Microsoft.*` string. Four of the five discovery sources speak
  ARM natively; only Terraform needs translating. The ARM ID also embeds
  subscription and resource group, so one field supplies the cluster seed key, the
  environment scope, and uniqueness with no derivation.
- **`_exec: { _agent: rw }` + `_interactive: false` on discover and generate.** Both
  are bulky, file-only and self-contained. This is also what forces live `az`
  capture to become main-window pre-work later rather than a fragment: a dispatched
  worker cannot prompt for consent.
- **Generate is gated on `run_mode: decide_and_execute`** plus
  `_check_phase_completed: workshop` (the sidebar declares `_gates: generate`).
  The `run_mode` check is an `_assert`, which CI binds but never evaluates — that
  is stated in `generate.md` so a future reader does not assume it is enforced.
- **Postconditions encode the FINISHED contract, not the skeleton's behaviour.** A
  category or table landing in a later step cannot land silently: its assert fails
  until the unit exists. Every unit file carries a `## Status` block naming what it
  does today and which step fills it in.

Three shared refs ship with real shapes rather than stubs, because downstream
phases are written against them: `schema-discover-azure.md` (inventory, the typed
edge table, the drift record, clusters), `schema-preferences.md` (the assumption
sheet's four dispositions, and why a considered `N/A` is written explicitly rather
than omitted), and `schema-workshop-scenarios.md`.

Repo wiring per the plan's §1 checklist: `mise.toml` `lint:frontmatter` (both
trees), `cross-plugin-drift.ts` `SKILLS`, the three plugin.json flavours in both
trees, three marketplace manifests, `advisor/AGENTS.md`, both READMEs, and the
gcp/heroku SKILL.md descriptions now name azure-to-aws instead of excluding Azure
with no forwarding address. `architect-for-startups` keeps
`references/migration-azure-to-aws.md` for the pre-decision advisory conversation
and hands off to the skill on real migration intent.

Telemetry is wired now, including the two `emit.mjs` defects a third registered
skill triggers — both silent, both fixed here because registering AZURE_TO_AWS is
what activates them:

- The run-ownership guard resolved a single "other" skill via `.find()`, which was
  only correct while exactly two were registered. With three, a GCP-registered hook
  looking at an Azure run could resolve `other` to HEROKU_TO_AWS, whose inventory
  is absent, so the guard did not fire and the Azure run was reported as
  `sourceProvider: GCP`. It now checks every other registered skill.
- `costContainer` read `current_costs ?? gcp_baseline`. The baseline key is now
  looked up from the registering skill (`SKILL_INVENTORY[skill].baseline`), because
  a hardcoded key reproduces the $685M dropped-spend bug the comment above it
  documents, once per skill that is not the one it was written for.

`hooks/telemetry/cursor/hooks.json` registered GCP_TO_AWS only; it now registers
all three, which also fixes HEROKU_TO_AWS having been missing there since it
shipped.

The migrate copy of `SKILL.md` carries no telemetry block, matching gcp and heroku:
that tree ships no `hooks/` directory, so the hook command would resolve to a
nonexistent file, and the emitter-path search pattern names the advisor plugin,
which is not one of the drift tool's normalized token classes. It is allowlisted
with that reasoning recorded, alongside the same three vendored files heroku
already allowlists.

Deliberately NOT in this commit, per the sequence: Bicep/ARM/RDfA/live/billing/
app-code fragments (step 2), the mapping tables (step 3), real clustering and the
pattern catalog (step 4), rubric and knowledge content (step 5), the dual estimate
output and the generators (step 6), fixtures (step 7). The `ai/` vendored subtree
is also deferred to the AI-route work, because the promoted refs depend on a
per-skill `references/shared/pricing-cache.md` this skill does not ship yet.
… external oracle

Pulls the canonicalization table forward from step 3 to before step 2's discovery
breadth, because until it exists there is nothing to TEST.

The reason is the DSL's central limitation: `_assert` postconditions verify shape,
never correctness, and the model both produces the artifact and evaluates the
assertion against it. There is no independent oracle. So `discover-iac.md` with no
extraction rules at all still passes: a capable model reads
`azurerm_linux_web_app`, emits `Microsoft.Web/sites` from pretraining, inspects its
own well-formed output against "every entry has a canonical Microsoft.* type", and
emits HANDOFF_OK. That is a FALSE GREEN — structure satisfied, contract satisfied,
zero skill content consulted, output stamped `confidence: deterministic` when it is
a model prior, and unreproducible across runs. You would conclude discovery works.

Landing:

- `references/shared/arm-type-canonicalization.md` — the `azurerm_* → Microsoft.*`
  table across compute/data/networking/identity/messaging/observability, plus the
  six traps where a plausible guess is wrong: `Microsoft.Web/functionApps` does not
  exist (a function app is `sites` + `kind`), the plan is `serverfarms` lowercase
  while `serverFarmId` is the camelCase property, `Microsoft.Cache/Redis` carries a
  capital R, Cosmos DB's provider is still `DocumentDB`, Azure OpenAI is a
  Cognitive Services account, and a resource group's own ID has no `/providers/`
  segment. Also the `azure_id` reconstruction rules, including the
  `<subscription-unknown>` placeholder — inventing a GUID makes the ID unstable
  across runs and breaks drift comparison against a live capture, which is the one
  thing the ID exists to support.

- `references/shared/extract-terraform.md` — extraction, per-type attributes (only
  what a downstream table actually reads), the edge table, and the secret boundary.
  `count`/`for_each` become ONE entry with the expression recorded rather than being
  fanned out, because the count is usually a variable and fanning out invents
  resources. `tfstate` is never read. Registry modules absent from the workspace get
  a warning, that being the commonest reason a Terraform-only inventory is
  confidently incomplete.

- `discover-iac.md` rewritten with real dialect detection and canonicalization, and
  a **halt guard**: a dialect present whose `extract-*.md` ref is missing now emits
  GATE_FAIL instead of exiting cleanly. Previously "no dialect present" and "the
  dialect is present but this skill was never taught to read it" were
  indistinguishable, and the second is exactly the hole improvisation walks through.
  So a `.bicep` or ARM workspace stops today rather than half-discovering — a
  partial inventory presented as complete makes the estimate confidently wrong about
  the size of the estate.

- `fixtures/azure-iac-terraform/` — a 28-resource synthetic estate, `expected-*.json`,
  and a stdlib asserter, both trees. Synthetic rather than a real repo because a real
  repo will not contain the traps: this one deliberately puts five web apps on one S1
  plan, an idle plan with zero apps, an app in `rg-app` with its database in `rg-data`,
  a second resource group holding two unrelated workloads, a function app, an SMB file
  share, a Kafka-enabled Event Hubs namespace, a Mongo-API Cosmos account, a private
  endpoint, a cost-bearing type absent from the table, and an unresolvable module.

  The asserter pins ONLY facts where improvisation and correctness diverge; facts a
  model gets right by accident are deliberately not asserted, since they cost review
  attention and prove nothing. Verified non-vacuous against a hand-written inventory
  carrying `Microsoft.Web/functionApps` and `Microsoft.Cache/redis` — it fails and
  names both by the rule they violate. The secret sentinel is a greppable token
  rather than a realistic credential because a realistic one trips gitleaks.

  Smoke-only in `run-asserters.py`: the corpus is committed, the run tree is not,
  since producing one requires an agent run. CI runs it against an empty dir and
  requires a clean non-zero exit as a bitrot guard.

Gates: lint:frontmatter 7 phase files in both trees, shared:check, model-id lint,
fixtures:check (93 json, 9 asserters), fixtures:assert (4 golden, 5 smoke) all green.
drift:check clean for azure-to-aws; the two pre-existing gcp/heroku SKILL.md failures
from the telemetry work are untouched and still outstanding.
…ive defects it found

The corpus already existed; this actually runs it. Extraction was performed by hand
against `extract-terraform.md` + `arm-type-canonicalization.md`, the resulting
27-resource inventory is committed as a golden tree, and the asserter is promoted from
smoke-only to a GOLDEN CI gate.

Read this as a consistency test, not a capability test. Whoever authored the rules also
performed the extraction, so it proves the rules, the corpus and the oracle agree — not
that the skill teaches a model that has never seen them. It still found five real
defects, three of which would have made the oracle pass while testing nothing.

**1. Identity was not unique.** The asserter indexed resources by `tf_resource_name`.
Terraform namespaces local names PER TYPE, and this corpus collides on five of them:
`core`, `storefront`, `reporting`, `data`, `store`. So five of twenty-one type
assertions were silently pointed at whichever resource happened to be written last,
passing or failing by accident of iteration order. Identity is now the full Terraform
address (`azurerm_subnet.data`), recorded as `config.tf_address`, with the asserter
failing loudly on a missing or duplicate one. Six previously-unassertable resources
became assertable as a side effect.

**2. The horizontal-resource-group case was not actually expressed.** The corpus claimed
to exercise the merge of an app in `rg-app` with its database in `rg-data`, but the app
never referenced the database — `DATABASE_URL` was a literal. The only edges crossing a
group boundary were subnet ones, which are a far weaker signal and are not what the case
is for. The app now interpolates the server's `fqdn`, the edge taxonomy gained the
`data_ref` type it had no home for, and the assertion is a specific directed pair
(`storefront` → `store`) rather than "some edge crosses a boundary", which a broken
extractor could satisfy with the subnet edges alone.

**3. Child resources were being nulled out.** The rules said an absent
`resource_group_name` means `resource_group: null` plus a warning. Most child types
carry no `resource_group_name` — storage shares and containers, SQL databases, Event
Hubs, Service Bus queues, Cosmos databases — and reference their parent instead. Since
a resource with no group cannot be clustered, that rule collapsed the cluster seed for
most of a real estate. Children now inherit the parent's group via the parent reference,
and the fixture asserts it on the storage share.

**4. Redaction-in-place escaped the secret check.** Asserting on the sentinel's text
missed a truncated-sentinel-plus-`***` mutation, which the rules forbid but a substring
check cannot see. The oracle now asserts the VALUES CONTAINER does not exist at all —
no `app_settings`, `connection_string`, `value`, `admin_password`, or access-key field
on any resource's config.

**5. Nothing asserted the child resource group.** Covered by 3.

Verified by mutation: 18 realistic failures were injected into the passing inventory.
Sixteen were caught first time; the two that escaped are defects 4 and 5, and both are
now caught. Every mutation is named by the rule it violates rather than by a diff, so a
failure tells you which decision was made wrong: the function app typed as
`Microsoft.Web/functionApps`, `kind` dropped, Redis lower-cased, the plan camelCased, a
resource-group id given a `/providers/` segment, an invented subscription GUID, one and
then all five `hosted_on` edges dropped, the app-to-database edge dropped, `private_link`
dropped, a secret value recorded and then redacted in place, an unknown type guessed
rather than reported, a module silently omitted, `enabled_protocol` and `kafka_enabled`
dropped, a null child group, and `tf_address` absent.

One finding NOT fixed, recorded for review: nearly every `azure_id` here contains a
`tf:<local>` name segment, because every resource in the corpus names itself with a
`${var.prefix}` expression. That is honest and stable across runs, but it means
`azure_id` is not a real ARM ID for most Terraform-sourced resources and cannot be
matched against a live capture. `metadata.subscription_id_source` records the
subscription half of the problem; the name half has no equivalent flag yet. A
per-resource `azure_id_synthetic` marker is the obvious fix and belongs with the live
`az` path, where the drift comparison it protects actually happens.

Gates: lint:frontmatter 7 phase files both trees, shared:check, model-id lint,
fixtures:check (94 json), fixtures:assert (5 golden, 4 smoke) green. drift:check clean
for azure-to-aws; the two pre-existing gcp/heroku SKILL.md failures are untouched.
Untracked across two sessions; §11 (resolved open questions) and §12
(sequencing change: canonicalization before discovery breadth) only exist
here, so losing the file loses the reasoning behind the build order.
Build step 3 (plan §7a). Pass 1 of Design is real; pass 2 (the category
rubrics) lands in step 5, and until it does the phase HALTS rather than
mapping compute and databases from model priors.

The disposition table is knowledge/*.json rather than design-refs prose,
per plan §0 and following heroku's fast-path-addons.json. That choice buys
something specific: the asserter can validate the TABLE, so a design row
claiming `confidence: deterministic` is checked against a row that actually
exists — the one direction that catches an improvised label.

Three deltas from the plan, all flagged rather than quiet:

* The 10-row Direct Mappings table becomes 17. Four rows are additions of
  necessity: canonicalization emits subnets and the storage-service children
  as their own child-typed resources, so without a row each one reaches the
  unknown-type policy, matches "sits in a network/data provider namespace",
  and STOPS the design on every real estate. §7a.3 fixed the row set before
  the child-type vocabulary existed.

* Three rows carry a mechanical condition on one config field and are still
  architecture-invariant, because the condition reads a property of the
  resource itself: SMB/NFS share, kafka_enabled, and the Cosmos API. Owner
  decisions 11.4 and 11.5 say protocol is the WHOLE rubric for the first two,
  which means there is no rubric left — only a lookup. Same shape as gcp's
  google_sql_database_instance (SQL Server) row.

* Decision 11.5 disturbs no row. Canonicalization already makes a file share
  its own resources[] entry, so a share is never inside the account's mapping
  unit and cannot pull the account's target anywhere. The account keeps
  Always -> S3 for its blob surface; the only condition it needs is
  account_kind == FileStorage, where there is no blob surface at all.

An untranslated type is now cost-bearing BY DEFAULT and STOPs. §7a.4's
three-part test cannot clear it — no SKU, no consumption row, no provider
namespace — and the skill cannot demonstrate that a resource it could not
name is free. Design therefore reads iac_metadata.untranslated_types as a
separate input, because Discover drops those resources from resources[]
entirely and iterating resources[] cannot find them.

A STOP writes aws-design.json carrying everything determined plus a `halt`
object, then fails its gate. Discarding the work would make the user re-run
the phase to learn one missing table row, and would hide which resources
were already fine.

Oracle: after-design-halted/ + check_expected_design.py, registered golden.
30 injected mutations, 30 caught first pass, each naming the rule violated —
including the three famous wrong answers (DynamoDB for Mongo Cosmos, EFS for
an SMB share, Kinesis for a Kafka namespace) and the 5x App Service Plan
fan-out. Identity is the Terraform address resolved through the run's own
inventory, so the expectations survive the azure_id reconstruction change
that the live `az` path will bring.

Deliberately untested and said so in the fixture README: the five specialist
gate rows (the corpus fires none), and clusters[]/pattern_id (step 4).

Gates: lint:frontmatter 7 files both trees (and verified non-vacuous — a
broken _knowledge path fails it), shared:check, fixtures:check, fixtures:assert
6 golden / 4 smoke both trees, lint:model-ids all green. drift:check still RED
on the two pre-existing gcp/heroku SKILL.md files; not this branch's, not
allowlisted. lint:md and fmt:check UNVERIFIED — dprint and markdownlint are
not installed on this box.
A fresh-context agent ran Discover over the committed corpus and PASSED the
Terraform oracle — every casing trap, the fan-in edges, the cross-RG data_ref,
child resource-group inheritance, the secret boundary, all correct from the
refs alone. It then diverged from the golden tree on six things no ref stated,
and no existing assertion noticed any of them. The run was correct by its own
instructions; the instructions were incomplete.

1. `warnings[]` was mandated by three files and DEFINED BY NONE. Absent from
   schema-discover-azure.md entirely: no location, no entry shape, no code
   vocabulary. The run invented all three and picked `module_not_discovered`
   where the golden says `module_not_resolved`. Now defined once, with a closed
   seven-code vocabulary and a required subject (azure_id or identifier) —
   a warning nobody can attribute to anything is noise.

2. `subscription_id_source` was specified in BOTH `metadata` (canonicalization
   ref) and `iac_metadata` (discover-iac, the golden, the asserter). Resolved to
   `iac_metadata`: only Terraform needs the ID reconstructed, so only the IaC
   section has anything to say about where the subscription half came from.

3. `data_ref` — described as the app-to-data edge and the thing that makes the
   horizontal-RG merge work — was in extract-terraform.md and NOT in
   schema-discover-azure.md's typed-edge list. Step 4's clustering reads the
   canonical list, so the edge that matters most was missing from the file that
   consumes it. The list is now canonical and closed.

   Containment stays deliberately edge-less, and now says why: an ARM azure_id
   CONTAINS its parent's as a literal prefix, so a share's account is derivable
   by truncation for every source. Observability links (a component's
   workspace_id) live in config, not edges — an edge implies a dependency the
   architecture preserves, and that one does not survive the migration.

4. The `/fileServices/default/` implicit-singleton segment was stated nowhere.
   The run guessed it, and guessed RIGHT, which is worse than guessing wrong:
   the rule looked taught when it was not, for all four storage sub-services.
   Now in § Reconstructing azure_id, with the note that the segment belongs to
   the ID and never to the type string.

5. extract-terraform.md had no per-type row for Key Vault, Log Analytics,
   private endpoints, or App Insights, so a correct-by-the-ref extraction
   dropped sku_name, sku + retention_in_days, and subresource_names — all
   cost- or routing-bearing. Five rows added, plus two rules the table never
   stated: omit rather than null, and a type with no row still gets an entry.

6. discover.md's postcondition demanded "the edges[] set that justified the
   grouping" while discover-assemble.md instructs an EMPTY edges[]. A live
   contradiction that passed only because `_assert` has no teeth — the exact
   false-green class this project is organised around. Fixed by naming the real
   reason: clusters now carry `justification`, which is `seed:resource_group`
   today. Without it, "grouped by the seed" and "grouped for no recorded
   reason" produce identical output.

The oracle now enforces all of it, which is the point — this class of defect is
invisible to shape assertions by construction. New `contract_vocabulary` and
`required_config_fields` blocks check the warning codes, the edge types, the
no-null rule, and the 15 cost/routing fields per canonical type. Run against
the capability test's own output they produce 23 named failures; run against
the golden, PASS. Seven further mutations injected (invented code, missing
detail, unattributable warning, absent array, invented edge type, null config,
dropped field) — 7/7 caught, each naming its rule.

Golden inventory updated where the ref and the golden disagreed: App Insights
gains application_type + workspace_id, and warnings[] is rewritten to the
defined shape. The golden had been over-delivering relative to its own ref,
which is the mirror image of defect 5 and equally invisible.

Design gets its own closed warning vocabulary in design-infra.md, kept separate
from Discover's: Discover's codes say what could not be READ, Design's say what
was DECIDED, and a cost-bearing or untranslated type produces a `halt` rather
than a warning because a warning means the design continued.

Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check,
fixtures:assert 6 golden / 4 smoke both trees, lint:model-ids green.
drift:check still RED on the two pre-existing gcp/heroku SKILL.md files.
lint:md and fmt:check still UNVERIFIED — dprint and markdownlint absent.
…rinting the answers

A second capability run — fresh context, answer keys blocked — passed the Discover
oracle including all the new contract checks, so the previous commit's hygiene held.
It then failed the Design oracle in seven places, and every one was my
under-specification from bd398c7 rather than the model improvising.

FIRST, THE GOOD NEWS: the halt guard works under pressure.

Told to map resources whose rubric file is absent, the run produced NO mapping for
six of them — three App Service Plans, the function app, the Windows VM, the
Postgres server — and recorded each with its missing ref_file. Unprompted, it
reported being strongly tempted not to: index.md's "Typical AWS target" column
printed the answers and fast-path.md's Preferred-Target table printed the
tie-breaker, so a complete and plausible compute and database design was available
without opening compute.md at all, and every shape assertion would have passed it.

That is the whole false-green thesis, live: the guard is the only thing standing
between "no rubric needed" and "the rubric was never written", and it held. But it
held against a hazard I built, so:

* index.md's right-hand column is no longer an answer key. Rubric rows now carry an
  unordered CANDIDATE SET in braces — no defaults, no preference ordering — and the
  column note says plainly that being able to answer from the column alone means the
  row is wrong. Fast-path rows keep their target, because there the answer
  legitimately lives in the table.

SECOND: aws-design.json had no schema, and five of the seven failures were that.

Every other artifact has one. The run reverse-engineered the shape from twelve
postconditions and scattered prose and invented reasonable-but-different key names,
while the committed golden used a third set — so the oracle was asserting
hosted_app_azure_ids and sizing_source that NO skill file required. That is exactly
the golden-over-delivers-vs-ref defect I diagnosed for extract-terraform.md one
commit earlier, reintroduced in design-infra.md in the same session. A hand-authored
golden and a prose-only contract drift by construction.

references/shared/schema-design-aws.md is now the thing both have to agree with, and
it makes three things REQUIRED that were merely implied: fast_path_row on every
deterministic entry (so the label is auditable rather than an unverifiable claim),
and hosted_app_azure_ids + sizing_source on every compute unit — without which "five
apps correctly fanned in" and "four fanned in, one silently dropped" produce
identical artifacts.

THIRD: two contradictions the run walked straight into.

* index.md gave Microsoft.Web/sites two rows: one saying "none of its own — fans IN
  to its plan", the next giving kind=functionapp its own Lambda/Fargate target. It
  followed the second and emitted a standalone function-app entry, contradicting
  design.md's own postcondition. A Y1 consumption plan is exactly where the two
  readings collide. Resolved in favour of the plan: a site NEVER gets an entry, and a
  hosted site's kind narrows the plan's candidate set instead.

* design.md's cluster postcondition demanded target_architecture while patterns.md
  does not exist, so it was unsatisfiable. The run wrote "UNDETERMINED" and flagged
  that a green-seeking agent writes a plausible architecture string there instead,
  which nothing downstream could catch. Fixed the same way cluster `justification`
  was: pattern_status now distinguishes `recognized` from `unclassified` (catalog
  consulted, nothing matched) from `catalog_absent` (never attempted), and
  target_architecture MUST be null unless recognized. The honest state is now
  expressible AND asserted.

Also: postcondition 7 predated pending_rubric[] and so a correctly-halted design
could not pass it — pending_rubric is permanent contract, not scaffolding, since
adding an index.md row without its file can recur. And halt.blocking now requires one
missing_rubric_file entry per distinct missing ref, because pending_rubric[] with no
halt reads as complete to anything that only inspects services[].

Smaller: name_expression_unresolved is now one entry per run listing the affected
addresses. Per-resource it produced 21 of 25 Discover warnings and buried the four
actionable ones — the ref already granted that courtesy to
subscription_id_unresolved and not to this.

Oracle: 3 new assertion groups (cluster pattern_status, compute-unit required
fields, halt-covers-pending). 7 mutations injected — 7/7 caught, including the two
that matter most: clusters given a plausible recognized architecture, and a halt
naming compute.md but not database.md.

Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check, fixtures:assert
6 golden / 4 smoke both trees, lint:model-ids green. drift:check still RED on the two
pre-existing gcp/heroku SKILL.md files. lint:md and fmt:check still UNVERIFIED.
The plan existed in two places and forked: §11 and §12 landed in the repo copy
on 2026-09-04 while the Pippin artifact stayed at its 2026-09-03 content, so the
mirror still described a build order §12 inverts and an Azure SQL disposition
§11.3 overrules. Nothing enforces the sync, so the header says which way it
flows. Pippin is now published from this file (v4).
Keeps the handoff in the repo alongside the plan, for the same reason: the
Pippin copy forked once and nothing enforces the sync. Published to Pippin
from this file.

Corrects the previous handoff's claim that bedrock-quotas.md has zero inbound
references — it has two real load references (gcp design-ai.md:192 and
estimate-ai.md:108), and acting on the old claim would have deleted a live
file.
…d inventory

Two things: the pass-2 rubrics the corpus tests, and the inventory/cluster work
that has to be right before a rubric reading them means anything.

INVENTORY — three gaps that only show up on a real repo

* Canonicalization coverage 77 -> 137 azurerm types. The old set was sized for
  the synthetic corpus; a real repo routinely carries azurerm_firewall,
  azurerm_route_table, azurerm_virtual_network_peering, azurerm_bastion_host,
  azurerm_app_configuration, azurerm_key_vault_key and more, each of which was
  silently dropped as untranslated. Also added § Coverage is not completeness,
  because 137 is still not the provider surface and the honest framing matters:
  the STOP is the mechanism, one table row is the fix.

* `.terraform/modules/` is now READ, and this is the biggest practical fix in
  the commit. The blanket `.terraform/` ban was written to keep state files out
  and was correct for state — but `modules/` holds nothing except downloaded
  module SOURCE, exactly as safe as the local module path we already recurse
  into. Modern Azure Terraform leans hard on Azure Verified Modules, so under
  the old ban a repo could declare almost its whole estate through registry
  modules and get back a nearly empty inventory plus a few warnings, with every
  downstream phase then reasoning confidently about a fraction of the estate.
  Still banned: any *.tfstate, anywhere.

* Association-only resources now have a rule. `azurerm_subnet_*_association`
  and friends exist only in Terraform and have no ARM type at all. They were
  being reported as untranslated — i.e. as gaps in this skill — which is wrong
  and, worse, buries the real gaps: on an IaC-heavy repo the associations
  outnumber the genuinely-missing types. They emit an EDGE and no entry.

CLUSTERING — real, not a seed

references/clustering/{clustering-algorithm,typed-edges-strategy,classification-rules,tiering}.md.
Seed per resource group, split a candidate whose members have no edges between
them, merge candidates joined by a crossing edge, then tier + primary + roles.

The load-bearing judgement is that two edge types are AMBIENT and must not
merge: `network` (everything in a VNet shares subnets) and `secret_ref` (one
Key Vault serves the estate). Merging on either collapses the whole estate into
one cluster and destroys the partition. They are still recorded and still count
for internal connectivity in the split step — sharing a subnet is weak evidence
that two resources are related and good evidence that two already-related ones
belong together, and that asymmetry is the point.

Rejected a weight-and-threshold scheme: the threshold would have no defensible
source, would need re-tuning per estate shape, and its failures would be silent.
Binary merges/does-not per edge type is explainable in one sentence per row.

RUBRICS

compute.md — eliminators (App Runner is never a candidate; Lambda's 15-minute
ceiling, Durable Functions, Windows containers on EB, GPU on Fargate), the six
criteria in order, the fan-in, x86_64, and right-sizing as post-selection.

database.md — the availability override gate first, because it is the thing that
decides RDS vs Aurora and it is NOT inferable from configuration. When the
answer is absent the default is RDS single-AZ, explicitly not Aurora: a source
`high_availability { ZoneRedundant }` says what they bought, not what they need,
and inferring Aurora from silence inflates the estimate with no visible cause.
The source HA posture surfaces as a FINDING instead
(`availability_downgrade_from_source`), which is a genuinely useful thing to
tell someone.

FIXTURE — the untranslated-type case was fragile and is now durable

Broadening the canonicalization table added azurerm_dev_test_lab, which was the
corpus's untranslated-type fixture — so a coverage pass silently disarmed the
STOP test. Any fixture that depends on a type being ABSENT has that failure
mode. Swapped to azurerm_iothub: out of this skill's scope by design, so a
coverage pass will not absorb it, and it carries a real sku block so the corpus
comment ("it carries a sku") is now literally true rather than aspirational.

The golden design is regenerated with both passes applied: 14 services[], zero
pending_rubric[], one halt blocker instead of three. The pending_rubric
assertion INVERTED — at step 3 a resource mapped past a missing rubric was
improvisation; now a resource still parked as pending is a run that did not load
rubric files that exist.

Oracle: 4 new assertion groups (rubric outcomes with their tempting wrong
answers, the x86_64 default, the source-HA finding, empty pending_rubric).
12 mutations injected, 12 caught — including Postgres -> Aurora inferred from
the source's own HA setting, which is the single most valuable assertion here.

One oracle bug found and fixed by its own golden: the architecture check
substring-matched the whole aws_config, so a `graviton_reason` field explaining
why Graviton was NOT chosen tripped it. Narrowed to architecture-bearing keys —
a blunt match would have pushed authors to stop explaining themselves, which is
backwards, since that explanation is what stops the x86 default reading as a bug.

Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check,
fixtures:assert 6 golden / 4 smoke both trees, lint:model-ids green.
drift:check still RED on the two pre-existing gcp/heroku SKILL.md files.
lint:md and fmt:check still UNVERIFIED.
A third capability run — fresh context, answer keys blocked — confirmed the three
inventory fixes work end to end and then found that ONE of them had broken Design.
Both regressions were mine, from the previous commit.

WHAT WORKED (verified from a clean context, quoting the instruction it followed)

* Registry modules: read `.terraform/modules/modules.json`, resolved the key to a
  directory, discovered both module resources, tagged them `tf_module`. The
  `module.cdn` case whose source is NOT on disk still warned. The distinction held.
* Association-only resources: no entry, no untranslated warning, one `network` edge.
* Every halt honoured. On the route table and bastion host it reported being
  "strongly tempted" — warning-and-continuing would have made the gate 1-of-12
  instead of 2-of-12 and looked much better — and stopped anyway.
* The database went to RDS single-AZ with no availability answer, and surfaced the
  source's ZoneRedundant HA as a finding rather than reading it as the answer. That
  is the exact behaviour database.md §1 was written to produce.

REGRESSION 1 — the coverage pass left 53 canonical types with no disposition

Taking the canonicalization table 77 -> 137 types made the INVENTORY better and
DESIGN worse: each newly-translatable type is now discovered, matches no row, hits
the cost-bearing provider-namespace clause, and STOPS. On this run a route table and
a bastion host — both free-or-trivial in reality — halted the whole design.

Every oracle stayed green throughout, because the corpus contains a dozen types out
of 137. So the fix is not just the 53 rows (30 direct / 52 skip / 10 gate now) but
the invariant: check_expected_design.py reads both files and asserts ZERO orphans.
That check would have failed the moment the coverage pass landed.

Two dispositions worth flagging as judgement calls: Azure Bastion -> Session Manager
(no host at all, and free — an EC2 bastion would add a permanent cost line the target
does not have), and Traffic Manager -> a Route 53 routing policy rather than a new
service. Logic Apps, Batch, ML workspaces and Stream Analytics went to specialist
gates rather than being mapped, because in each case the work is a rewrite that the
resource does not describe.

REGRESSION 2 — clustering over-fragmented, and my own worked example proved it

The run produced 16 clusters where the file's example predicted 3. Cause: I had
deliberately made containment NOT an edge (an ARM id contains its parent's, so it is
derivable) and then wrote a split step that only looked at edges — so a VNet split
from its subnets, a storage account from its share, and every edgeless observability
resource became its own "workload". The example was written by hand instead of by
tracing the algorithm, so the two disagreed and the algorithm shipped.

Two fixes: containment counts as connectivity in the split step (derived from the id
prefix, still not stored as an edge), and a split only happens when two or more
components each contain a PRIMARY-ELIGIBLE resource — a component with nothing that
could be its primary is a fragment, not a workload, and attaches to the largest
component. Re-traced the example against the algorithm: 4 clusters.

Also `split:*` no longer requires a non-empty `edges[]`. It cannot have one — its
justification is an ABSENCE — and the old rule made a correct split unrepresentable,
so the run had to mislabel nine clusters `seed:resource_group` to pass the gate,
destroying the information the field exists to carry.

THE CORPUS NOW GATES THE FIXES PERMANENTLY

The module and association cases were tested once by a throwaway workspace. Both are
now in the committed corpus: `.terraform/modules/naming/` with two resources, a
`module "naming"` block, and an `azurerm_subnet_route_table_association`. 29 golden
resources from 31 declared. Verified `.terraform/` is not gitignored here, or the
fixture would have been silently empty.

Oracle: 5 new assertion groups. 10 mutations, 10 caught — including the blanket
`.terraform` ban being restored, module.naming warned as unresolved, the association
given an entry, the association's edge dropped, and the bastion mapped to an EC2
instance.

Known and NOT fixed: the coverage check verifies a disposition EXISTS, not that it is
correct. A wrong row for a type absent from the corpus escapes, and one injected
mutation demonstrates that. Closing it needs corpus breadth, not a cleverer asserter.

Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check, fixtures:assert
6 golden / 4 smoke both trees, lint:model-ids green. drift:check still RED on the two
pre-existing gcp/heroku SKILL.md files. lint:md and fmt:check still UNVERIFIED.
Build step 5, part 1 — the phase that was blocking everything. Clarify ran and then
failed four of its own postconditions because only clarify-global.md existed.

FOUR FRAGMENTS, following gcp's category model

clarify-compute.md   compute target, plan isolation, CPU architecture, traffic,
                     long-lived connections, VM cutover (ESSENTIAL)
clarify-database.md  availability (the override gate's input), DB cutover
                     (ESSENTIAL), traffic, storage I/O, Cosmos read/write split,
                     Redis modules
clarify-licensing.md CONDITIONAL — Windows model (ESSENTIAL), SQL model, AHUB
                     detection, the Azure Edition hard blocker
clarify-identity.md  always fires, one shallow row

ONE STRUCTURAL CORRECTION toward gcp's flow

clarify.md said each fragment "presents its sheet section", which would give the user
FIVE sheets and five interleaved rounds of essentials. gcp runs ONE sheet as a single
mandatory gate, then batches the essentials — and this phase's own postcondition says
"every assumption-sheet row the user was shown", singular.

So fragments now compute rows and ask nothing; the assembler owns the whole
conversation in three gates: the consolidated sheet (batched five at a time, with each
fragment's consequence line), then the ESSENTIAL questions with their context, then the
recap. One place knows the full row set.

THE TWO RULES THAT CARRY THE WEIGHT

`ESSENTIAL` + `value: null` IS the completion gate. An essential row has no default on
purpose and the phase must not complete while one is unanswered. It is the only way the
contract can say "shown and not answered".

A value taken from its default STAYS `PROPOSED`. `DETECTED` means read from the estate.
Design's rationale prints "you chose Elastic Beanstalk" differently from "we assumed
Elastic Beanstalk" — but only if this file recorded which happened.

Three rows have no default at all, because they select different RUNBOOKS rather than
different numbers: VM cutover (MGN vs rebuild), DB cutover (DMS vs dump/restore), and
the Cosmos read/write split (which moves the DynamoDB conversion by multiples).

`data.availability` becomes ESSENTIAL when the source is zone-redundant, rather than
PROPOSED. Silently downgrading resilience someone pays for today is the expensive
mistake, and reading the source's own HA setting as the answer is the other one — it
says what they bought, not what they need.

CLUSTERING WAS REAL AND COMPLETELY UNGATED

Found while wiring preferences.clusters[], which keys on cluster_id: there was no
golden clusters artifact at all, so the split/merge rules were reviewed prose and
nothing more. The design golden also still carried the old seed-only 3 clusters, not
the 4 the re-traced algorithm produces — so the two disagreed and nothing noticed.

after-discover/azure-resource-clusters.json now exists: 4 clusters, 29 members,
justifications merge:cross_group_edges and split:no_internal_edges. The asserter pins
the merge (rg-app + rg-data joined only by app-to-data edges — seed-only clustering
would cost the app without its database), the split inside one resource group, the
legal single-member idle-plan cluster, that a split MAY have empty edges[] while a
merge may not, and that a Microsoft.Web/sites is never a cluster primary.

THE CLARIFY TESTING GAP, ADDRESSED RATHER THAN DEFERRED

Clarify is _interactive: true, so the capability-test method cannot reach it — a
dispatched agent has no user. clarify-answers.json stands in for one, which makes the
sheet's BRANCHING testable. Stated plainly in both the asserter and the fixture: this
tests nothing about wording, batching or tone. Those need a human. It tests which rows
fire, which are ESSENTIAL vs PROPOSED, and which are N/A — which is where the defects
are, because a firing rule is a judgement and prose cannot check itself.

The golden is a BLOCKED clarify on purpose: the scripted user declines to state Azure
spend, so an ESSENTIAL row is unanswered and the phase must GATE_FAIL. Pinning the gate
matters more than pinning the happy path.

15 mutations, 15 caught, including: availability PROPOSED instead of ESSENTIAL;
availability DETECTED from the source's own HA; the Cosmos RU question asked for a
Mongo-API account; isolation asked for the 1-app plan; licensing N/A despite a Windows
VM; identity N/A because no managed identities were found; spend invented to get past
the gate; and clarify_status COMPLETED with an essential row still null.

The asserter also caught an omission in my own golden — `ahub_in_use` was DETECTED with
no stated source, which is the promoted-default signature. It was genuinely read from
the estate; I just had not recorded what was read. AHUB matters because it does NOT
travel: if in use, the Azure baseline must be compared at the UNREDUCED rate or AWS
looks worse than it is.

Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check (101 json, 11
asserters), fixtures:assert 7 golden / 4 smoke both trees, lint:model-ids green.
drift:check still RED on the two pre-existing gcp/heroku SKILL.md files. lint:md and
fmt:check still UNVERIFIED.
Prepares a session handoff, and brings the authoritative plan up to date — it
had not been touched since §11/§12 on 2026-09-04, so eight commits of decisions
lived only in commit messages.

Plan additions:

§13 — the 21 decisions taken 2026-09-05/06, grouped by what they touch, with a
reason per row. Includes the three corrections to earlier claims: 11.5 disturbs
no Direct Mappings row, §7a.3's 10-row table is superseded by 30, and
bedrock-quotas.md is NOT dead (two real load references — the earlier claim came
from a grep that ate its own evidence).

§14 — the sequencing change. Terraform end to end BEFORE source breadth,
superseding both the original build order and §12's amendment. The previous
order optimised for discovery breadth and produced a skill that discovers five
ways and cannot produce an answer. Records the counter-argument too, because it
is real: most startups have no azurerm_* Terraform, and SKILL.md commits to
live-first — so this trades reach for completeness on purpose.

§15 — how the skill is actually tested, which is the most transferable part.
The extensional/intensional distinction and why only building the first left
every contract disagreement invisible; the capability-test method and what
each of the three runs found; why two of the three goldens are FAILING states.

Handoff at .agents/scratchpad/azure-to-aws-handoff-2026-09-06.md, kept in the
repo for the same reason as the plan: the previous handoff existed only in
Pippin and drifted unnoticed until the plan had forked by 8k characters.
Verified against the design golden rather than asserted from memory: the corpus
routes ZERO resources to a missing rubric file. Design already emits 16 services[]
with pending_rubric[] empty, and halts on exactly one thing — the untranslated
azurerm_iothub.

So networking.md + messaging.md are needed before any REAL customer repo and block
nothing on the fixture; they were wrongly placed next. Estimate and Generate are on
the critical path for both definitions of the goal, networking.md for only one. And
those two rubrics need corpus EXTENSION to be testable at all, since nothing in the
fixture reaches them.

This also promotes plan §13.7 #1 from an open question to THE gate on the end-to-end
goal: estimate.md carries _check_phase_completed: design and Design GATE_FAILs on the
corpus, so neither Estimate nor Generate can run there until the untranslated-type
halt is resolved. Both paths written up, with the user override recommended and the
STOP kept as the default.
…ring

The summary box and §1 still named networking.md as next after §8 was corrected.
Also distinguishes the two halt situations, which is the thing that was conflated:
on the CORPUS the only blocker is the untranslated type (pending_rubric is empty,
16 services mapped); on an arbitrary REAL repo the 7 missing rubric files would
also halt.
Rule 2 of arm-type-canonicalization.md claimed ARM type strings are compared
case-sensitively, and named two casing "traps": a capital R in
Microsoft.Cache/Redis and all-lowercase Microsoft.Web/serverfarms. Both were
unsourced, and the second is unsourceable — Azure/bicep-types-az ships BOTH
Microsoft.Web/serverFarms and Microsoft.Web/serverfarms in one generated index.
For the cache type, that index and magodo/aztft (the mapping library behind
Microsoft's supported Azure/aztfexport) both render Microsoft.Cache/redis. The
capital-R form is what appears in azurerm RESOURCE IDs — a Terraform-provider
artifact, not an ARM type.

Nine of this file's 124 externally checkable rows differ from aztft by casing
alone, so the discipline it claimed to enforce was wrong about 7% of its own
content. Worse, the rule made a mis-cased type "silently fall through to the
unknown-type policy", so a run emitting the CORRECT Microsoft.Cache/redis would
have had its Redis treated as untranslated and STOPped the design. And
expected-iac-terraform.json listed both correct forms under forbidden_types,
failing them as "the signature of a guessed translation".

Separate the two things the file conflated:

- MATCHING folds case, everywhere (fast-path, Skip Mappings, index.md routing,
  rubric selection). A case-only difference must never reach the unknown-type
  policy.
- EMISSION follows this file's spelling, as a stated CONVENTION justified by
  azure_id strings being joined by exact match — cluster membership, cluster
  keys, and the future drift comparison against a live capture all need one
  resource to yield one string. Explicitly not a claim about ARM.

Deletes the Redis trap, reframes the serverfarms trap to keep its real content
(serverFarmId is the property pointing AT the plan, not the plan's type name),
drops four inline casing assertions that are all among the nine unsourceable
rows, and records the evidence in a new section so it does not get "fixed" back.

No golden churn and no strictness lost: the ~74 existing occurrences of the two
display strings are correct under the convention and are untouched, and all
three azure oracles still pass. Mutation-tested — sites->functionApps (8 fails),
Redis->redis (2 fails, now reported as a convention violation rather than a
guessed translation), DocumentDB->CosmosDB (2 fails).

Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check,
lint:model-ids, fixtures:assert (11 asserters) all green both trees.
drift:check still red on only the two pre-existing gcp/heroku SKILL.md files.
lint:md and fmt:check remain UNVERIFIED — dprint and markdownlint are absent.
…r branch

This branch is cut from main, where `hooks/telemetry/` does not exist — it is
created by feat/telemetry-hooks. So azure's SKILL.md declared PostToolUse and
Stop hooks invoking `${CLAUDE_PLUGIN_ROOT}/hooks/telemetry/emit.mjs`, a path
that cannot resolve, and carried a 60-line consent step driving a CLI that is
not present. A hook command pointing at a missing file is exactly the defect
that makes the two pre-existing gcp/heroku drift failures unfixable on the
telemetry branch; shipping azure with the same shape would add a third.

Removed here, not redesigned:

- the `hooks:` frontmatter block and its SessionEnd note (advisor tree; the
  migrate copy never had one, since migrate/plugins/migration-to-aws ships no
  hooks/ directory at all)
- the Telemetry consent section from SKILL.md § Execution, including the
  $EMIT resolution ladder
- the `AZURE_TO_AWS` SKILL_INVENTORY entry and the Cursor hook registration,
  which were hunks of 317c3ac against files that do not exist on main
- `.cursor-plugin/marketplace.json`, which was created by the telemetry branch
  (2d09d36) rather than by main, so it is a telemetry artifact and not azure's

Plan §3a is annotated as DEFERRED rather than deleted: the design stands, and it
lands as a follow-up on feat/telemetry-hooks (or after it merges), together with
the two emit.mjs defects §3a documents and this branch no longer fixes.

Consequence to be explicit about: azure-to-aws currently emits NO telemetry.
Plan §3a calls that a shipping requirement, so this is a known gap on this
branch, deliberately taken to keep azure independently mergeable from main.
Rebasing 940aadd onto current main left `skills/shared/ai/` stale. That commit
created the canonical copy as a pure ADD while MOVING the gcp copy to
`references/vendored/ai/`, so git could follow the rename and carry main's
subsequent edits into the vendored path, but had nothing to follow for the
canonical one. The canonical file therefore kept the pre-rebase content.

Only `ai-migration-guardrails.md` was actually affected, in both trees: main's
8c65c92 replaced the "shared 10,000 RPM account limit" risk table with the
GPT-5.6 per-model TPM-only quotas (no RPM quota, cached input tokens excluded),
and the canonical copy still carried the old table. `shared:check` caught it,
which is the whole point of that gate — a stale canonical is worse than a stale
vendored copy, because `shared:sync` would have propagated it outward and
silently reverted main's facts in every consuming skill.

The remaining seven promoted files were already identical.

shared:check and drift:check are both green: 334 identical, 28 allowlisted.
…and defects

Six things the plan did not carry, and the handoff either omitted or got wrong.

§16.1 The Rule 2 casing correction, with the evidence, so it is not reverted:
bicep-types-az ships BOTH serverFarms and serverfarms in one generated index, and
both it and aztft render Microsoft.Cache/redis. Nine of 124 checkable rows differed
by case alone. Records the two live consequences — a correct answer STOPped Design,
and the oracle failed correct answers as "the signature of a guessed translation".

§16.2 magodo/aztft as external prior art. It validates our mapping semantics on
123 of 124 rows, and settles the "1 to N mapping" question empirically: 1048 of
1089 azurerm types resolve to exactly ONE ARM type and ZERO resolve to more than
one, so Terraform→ARM is a function and there is nothing for a model to choose.
Also records that aztft does NOT cover 13.2d — 22 association entries, none
flagged property-not-resource — and that it was deliberately NOT adopted as an
intensional check.

§16.3 The `confirm` phase, replacing §13.7 #1's open question and both paths the
handoff proposed. A phase rather than post-Discover afterwork because resolving an
untranslated type ADDS a resource and Clarify's fragment triggers read the
inventory. Option (b), amend in place. Design's "untranslated_types is empty"
postcondition passes UNCHANGED, so the STOP stays absolute. Carries the obligation
that would otherwise have been silent: amending the inventory obliges re-deriving
the clusters, because 64 azure_id occurrences live there in four roles.

§16.4 Corrects the handoff's "ONE owner decision": TWO gates block the corpus,
since after-clarify is BLOCKED_ON_ESSENTIAL independently of the Design halt.

§16.5 Why TF→ARM→TF is not a lossy double translation — config.tf_address retains
the azurerm type, Generate never reads Azure HCL, and the real loss is the config
attribute allowlist, which is independent of ARM.

§16.6 Four defects found and not fixed: azapi_resource unhandled, config.tf_file
empty on 24 of 29 golden entries against a ref that mandates it, the missing
clusters→inventory pairing check, and the canonical-staleness shape where a
promoted file created as an ADD misses upstream edits its vendored twin receives.

§16.7 The branch restructure, and the fact that drift:check is now GREEN — the
blocker handoff §7 called four sessions old was caused by telemetry landing
advisor-only, and basing on main removes it. Also states plainly that
azure-to-aws currently emits no telemetry.
… AzAPI discovery

Step 5b's three rubric files, and the one Discover gap that made real repos
undiscoverable. Both trees.

## The three rubrics

Each follows compute.md/database.md's anatomy: eliminators, the six criteria in
order with first-match-wins, feature parity, cluster context, post-selection
sizing, output contract, status. Confidence is `inferred`, never `deterministic`.
App Runner appears in networking.md's eliminator table as ALWAYS eliminated.

The content is not a mapping table with prose around it. Each file carries the
trap that makes its category non-obvious:

- **networking.md** — the DOUBLE-BALANCER trap. An Elastic Beanstalk
  load-balanced environment provisions its own ALB, so mapping an App Gateway
  whose backend pool is a web app to a *second* ALB double-counts the balancer.
  This is the networking analogue of the App Service Plan fan-in and it fails the
  same way: quietly, and only in the estimate. Also: Azure NAT Gateway is
  REGIONAL and AWS NAT Gateway is ZONAL, so one Azure resource becomes N AWS
  resources where N is the AZ count from preferences.json — an estimate that
  prices one NAT for a multi-AZ design is wrong in the flattering direction. And
  APIM's policy XML has no single AWS counterpart; the service maps, the policy
  layer is a rewrite.

- **messaging.md** — Service Bus is one broker; AWS splits its capabilities across
  SQS, SNS, EventBridge and Amazon MQ, so a NAMESPACE maps to a SET derived from
  its children and carries no cost line of its own unless the answer is Amazon MQ.
  Eight eliminators are all readable from IaC (256 KB SQS ceiling vs Premium's
  100 MB, the 5-minute FIFO dedup window vs Service Bus's 7 days, 15-minute
  DelaySeconds vs scheduled messages, AMQP/JMS clients). Event Hubs consumer
  groups are free on Azure and priced on Kinesis enhanced fan-out — named,
  because nothing else would surface it.

- **analytics.md** — for AI Search, the index maps and the PIPELINE THAT FILLS IT
  does not: indexers and skillsets have no OpenSearch equivalent and become
  explicit ingestion infrastructure the source did not have. For Databricks the
  like-for-like default is Databricks on AWS, with EMR only when the workspace is
  confirmed plain Spark — EMR is Spark, not Databricks with a different bill.
  Says explicitly that Synapse/ADF/Stream Analytics/ML workspaces are specialist
  gates and must NOT get rubric rows here, since a gate may never be overridden.

Corrected while writing: messaging.md initially claimed Event Grid types are Skip
Mappings. They are DIRECT MAPPINGS to EventBridge. The status section now says so
and explains why a rubric row for them would be unreachable.

## AzAPI

`azapi_resource` had zero mentions anywhere in the skill (plan §16.6), yet it is
common in modern Azure Terraform and its `type` argument states the canonical ARM
type verbatim — `Microsoft.Consumption/budgets@2023-05-01`. So the easiest
extraction in the skill was the one that halted. `extract-terraform.md` § Step 2a
now covers it: split on `@`, use the left half AS GIVEN without consulting the
canonicalization table, keep the API version in config, derive azure_id from
parent_id, read `body` for routing attributes only and never for values, and stamp
`config.declared_via: "azapi"`.

`azapi_update_resource` and `azapi_resource_action` emit no entry — they mutate a
resource declared elsewhere, the AzAPI analogue of the association-only class.
And a cost-bearing AzAPI type with no DISPOSITION still STOPs: AzAPI clears the
type-vocabulary problem, not the unknown-type policy. discover-iac.md now states
that azapi is part of the terraform dialect, not a fourth one, so the halt guard
does not misreport an azapi-only file as unreadable.

## New intensional check

`check_index_reference_files_exist` — every .md named in an index.md ROUTING ROW
must be on disk or declared in `known_pending` with a reason. This was on the
handoff's list of unguarded pairs: index.md carries a HALT, so a deliberately
pending rubric and a typo'd filename were indistinguishable at runtime, and a typo
would have been diagnosed as "write that rubric" when the rubric already existed.

Checked in BOTH directions — a known_pending entry that now EXISTS also fails, so
the list cannot rot the way a fixture depending on an absent type does. Only
`ai.md` is pending, and it belongs to deferred step 6.

Mutation-tested: a typo'd routing row fails naming the file; a stale
known_pending entry fails naming it.

## Effect

Of five realistic example estates, Design now completes on three (was one).
`legacy-lift` and `data-platform` newly clear. The two that still halt do so ONLY
on untranslated types — azurerm_maps_account and azurerm_notification_hub_namespace
— which are deliberately left absent as the `confirm`-phase test cases. Adding
rows for them would silently disarm the only examples that exercise the STOP.

Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check,
lint:model-ids green; drift:check OK (337 identical, 28 allowlisted);
fixtures:assert PASS both trees. lint:md and fmt:check remain UNVERIFIED —
dprint and markdownlint absent.
… were wrong

Answering "for services we cannot map, are we not relying on the LLM's internal
knowledge?" — yes, and the largest instance of it was sizing.

Five of the six sizing tables the rubrics NAMED did not exist; only
fast-path-services.json was on disk. So every instance size in the committed
golden came from pretraining: S1 -> t3.medium, Standard_D4s_v5 -> m6i.xlarge,
GP_Standard_D2s_v3 -> db.t4g.medium. No skill file contained any of those
mappings.

Three compounding problems, all now closed:

1. The instruction was UNSATISFIABLE. compute.md, database.md and analytics.md
   each said sizing "states the dev-tier default and says the table is absent
   rather than inventing a number". With no table, the dev-tier default WAS the
   invented number. The rule forbade exactly what it required.

2. `sizing_source` made an unsourced number look sourced. It records the Azure
   INPUT ({"sku_name": "S1", "worker_count": 2}) and says nothing about where
   t3.medium came from, so a reviewer sees a provenance field and reasonably
   assumes provenance.

3. The oracle pinned NO sizes, so two runs could disagree on every number and
   both stay green — section 12's false green in the one dimension Estimate is
   built on.

## Eight tables, auditable by construction

Every row records vCPU and memory for BOTH sides, so a reviewer can check the
arithmetic without trusting the author. A row whose AWS side is smaller than its
Azure side on either axis is a bug by inspection.

  appservice-eb-sizing.json   27 plan SKUs, incl. Y1/FC1/EP* routing to Lambda
  vm-ec2-sizing.json          15 family classes + 24 explicit sizes
  flexible-server-rds-sizing.json  14 SKUs, Graviton default with x86 per row
  cosmos-dynamodb-conversion.json  RU -> ops/s -> WCU/RCU, with the multipliers
  azure-region-map.json       44 regions, same_country flagged per row
  disk-ebs-sizing.json        tier map + per-volume IOPS breakpoints
  aks-eks-sizing.json         only what DIFFERS between AKS and EKS
  rightsizing-thresholds.json P95 bands + the aggressiveness slider

Wired into design.md's _knowledge with _when guards so a run loads only the
tables its inventory reaches. lint:frontmatter validates those paths, so a typo
fails the gate.

## Two of the three golden sizes were wrong

Writing the tables immediately contradicted the golden:

  S1 (1 vCPU / 1.75 GiB)              t3.medium (2/4)      -> t3.small (2/2)
  GP_Standard_D2s_v3 (2 vCPU / 8 GiB) db.t4g.medium (2/4)  -> db.m6g.large (2/8)
  Standard_D4s_v5 (4 vCPU / 16 GiB)   m6i.xlarge (4/16)    unchanged, correct

The RDS one shipped at HALF the source memory. The EB one at double. The third
was right — which is the point: an unasserted number is not the same as a
correct one, and nothing distinguished them.

## sizing_provenance

New REQUIRED field on any services[] entry whose aws_config carries a size:
`table` | `measured` | `user_stated` | `model_prior`.

`model_prior` is a legal value and must be used honestly — writing `table` for a
row that was never consulted is the failure the field exists to prevent, and the
only one a reviewer cannot detect from the artifact alone. A model_prior entry
SHOULD also carry a warnings[] entry naming the type with no sizing row, so the
gap is visible in the report and fixable in one place.

analytics.md still has no sizing table, deliberately, so it now says outright
that its node counts are `model_prior` rather than implying otherwise.

Also corrected in passing: plan section 0 records the EBS breakpoint as "gp3 to
80K IOPS". A single gp3 volume caps at 16,000 IOPS; 80,000 is an instance-level
aggregate. disk-ebs-sizing.json carries the per-volume limits, which is what a
disk maps to.

## Checks

check_sizing_provenance pins the three corrected sizes and requires the
provenance enum on every sized entry. Mutation-tested four ways: reverting either
wrong size fails naming the table row and the source SKU; dropping the field
fails; an invented enum value fails.

Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check,
lint:model-ids green; drift:check OK (345 identical); fixtures:assert PASS both
trees; all three azure oracles PASS both trees. lint:md and fmt:check remain
UNVERIFIED — dprint and markdownlint absent.
… a STOP

"Should we not use the LLM to resolve what the json does not carry? We don't have
such limitations in gcp-to-aws." Half right, and the half that was wrong is the
more useful half.

gcp-to-aws does NOT have broader knowledge. It lists 28 google_* types total
(22 in index.md, 19 in fast-path.md) against azure's 65 index.md rows and 92
fast-path rows. What it has is a FALLBACK before the stop, at
gcp design-infra.md:62-65: substring-match the type NAME to a category, run that
category's rubric, and only STOP if no pattern matches.

So azure was STRICTER than gcp on more than double the per-type coverage: a type
we could name, in a namespace we understood, still halted Design. That is the
real defect.

But gcp's fallback is not the one to copy. Matching `"log" -> monitoring` against
a type name also matches google_diaLOGflow_agent, which is a silent
miscategorisation into the wrong rubric. Azure has a better signal for free: the
provider namespace is a structural, authoritative segment of an ARM type string,
not a substring guess — and the cost-bearing test already reads it.

## Two derived rules, ahead of the STOP

Both are data in fast-path-services.json; design-infra.md § 2 places them.

- **namespace_routing** — 54 provider namespaces -> 30 rubric routes, 16 skips,
  8 gates. Rules, not rows: 54 rules cover a provider surface past a thousand
  types.
- **child_type_rule** — a canonical type with 2+ segments after the provider is a
  child of the type formed by dropping its last segment; if the parent has a
  disposition, the child is a config_source of it. Derived from evidence, not
  invented: 27 of the 33 config_source rows in skip_mappings are child types and
  0 of the 19 noise rows are, so the disposition was always derivable from the
  type path. Only the description of WHAT it contributes needed authoring.

Three constraints on both: an explicit row always wins; the route is RECORDED;
and a derived route can never produce confidence: deterministic.

## routing_provenance

New REQUIRED field on every services[] entry: `table` | `index_md` |
`child_type_rule` | `namespace_rule`. Asserted, with the deterministic guard
asserted separately — that tier needs an authored direct_mappings row and a
fast_path_row naming it, so a derived route claiming it is a defect.

The field exists because the risk these rules introduce is not wrongness, it is
INVISIBILITY: without it, a namespace-routed mapping is indistinguishable from
curated judgement in the artifact. Same discipline as sizing_provenance.

## Two thin rubrics, because the rules made them load-bearing

namespace_routing sends Microsoft.Storage/* and Microsoft.NetApp/* to storage.md
and Microsoft.KeyVault/* and Microsoft.ManagedIdentity/* to identity.md, so
leaving them unwritten would have converted the unknown-type STOP into a
missing-rubric HALT — moving the failure, not fixing it. Written thin, per plan
§14:

- storage.md — the SMB/NFS protocol rule's reasoning, access tiers, and the
  replication finding that matters: GRS/GZRS becomes S3 Cross-Region Replication,
  a NEW cost line that Azure bundled into one SKU.
- identity.md — the secret/key/certificate split, and an explicit REFUSAL to
  translate Azure RBAC into IAM policy. A translated policy is plausible and
  wrong, and the failure is either privilege escalation or an outage, both silent
  until exercised. Read the assignment for its identity_grant edge, record the
  intent, let a human author the policy.

ai.md stays the one declared-pending route (build step 6).

## Also

index.md's closing paragraph sent readers straight from "no row" to the STOP. It
now names both derived rules, and explains why a speculative row is worse than no
row: an authored row SHADOWS the namespace rule that would have routed the type
correctly and recorded that the decision was derived.

_coverage_contract revised: a type may now resolve by an authored row, an index.md
row, child_type_rule, or namespace_routing. Requiring a per-type row made coverage
growth O(n) in authored judgement against a provider surface past a thousand.

## Checks

check_routing_provenance (enum + the deterministic guard) and
check_namespace_routing_targets (every namespace route resolves to a file on disk
or a declared-pending one). Mutation-tested three ways: a derived route claiming
deterministic fails; a missing field fails; a namespace rule pointing at an
unwritten rubric fails.

Corpus expectation recorded honestly: this corpus routes NOTHING through a derived
rule, so the derived counts are zero and the checks guard the contract rather than
exercise it. Extending the corpus with a type only a namespace rule can place is
the follow-up.

Effect on the five example estates: legacy-lift and data-platform already cleared
with the 5b rubrics; the two remaining halts are untranslated types only, which is
confirm's job. Nothing now halts for want of a disposition.

Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check,
lint:model-ids green; drift:check OK (347 identical); fixtures:assert PASS both
trees. lint:md and fmt:check remain UNVERIFIED — dprint and markdownlint absent.
…exception list

"Building this for all types is impossible and not a good design. Deterministic
mappings for what we are sure about, LLM for the rest." Measured, and the data
backs it completely.

Of the 132 rows in arm-type-canonicalization.md:

  105 (79%) have a resource segment a PATTERN derives
  104 of those also have a namespace already declared in namespace_routing
   27 (20%) carry information no pattern can produce

So four fifths of the file was restating a rule, at 13% coverage of the azurerm
surface, decaying every provider release. The 27 that matter are the traps it
already documents: linux_web_app -> sites, service_plan -> serverfarms,
public_ip -> publicIPAddresses, lb -> loadBalancers, api_management -> service,
application_insights -> components, stream_analytics_job -> streamingjobs.

## Derivation is now the default path

1. Table first - because that is what makes derivation SAFE: the cases where a
   guess goes wrong are enumerated there.
2. Not listed -> derive. Resource segment by camelCase-pluralise of the Terraform
   suffix (79% provable). Namespace from knowledge, then CROSS-CHECKED against
   fast-path-services.json namespace_routing, which declares 55 namespaces
   independently of this file.
3. Namespace recognised -> keep the resource with its full config,
   azure_type_provenance: "derived", type recorded in iac_metadata.derived_types.
4. Namespace NOT recognised -> STOP, recorded in untranslated_types. That field
   now means "no recognised namespace", a far stronger signal than "no row".

The cross-check is the guard, not a hope: a derived namespace that an independent
artefact also carries is corroborated; one nobody recognises is exactly where the
model is inventing, and that is the case that stops.

## An unlisted type is no longer DROPPED

The old behaviour discarded the resource entirely - type string only, no config.
That is what made 87% of the provider surface a hard stop AND what destroyed the
sku/tier evidence Design needs for the cost-bearing test, which is why 13.1d had
to assume cost-bearing for anything unnamed. A derived resource keeps its place in
resources[] with config intact, so namespace_routing can place it and the SKU test
can actually run.

## The growth cap, so this is mechanical rather than aspirational

check_canonicalization_governance:
  - max_derivable_rows 105 - a CEILING, not a target. Adding a derivable row means
    the file is chasing completeness again and fails the gate. It should FALL as
    redundant rows are pruned, never rise.
  - min_divergent_rows 27 - a floor, so pruning cannot delete the informative rows.
  - every namespace in the table must exist in namespace_routing, or the two
    artefacts contradict each other and the cross-check would reject a type the
    table asserts is real.

The derivable rows present today are a CLOSED CORE: verified, free at runtime, so
they stay. What changed is the admission test - a new row is admissible only where
derivation would be WRONG.

That third check found a real contradiction on its first run:
Microsoft.DevTestLab was in the table with no namespace rule. Added as a skip - a
DevTest Lab is a management wrapper around VMs that are inventoried in their own
right.

## One DELIBERATE test failure, and it is the point

check_halt_not_stale fails on after-design-halted, correctly.

Derivation resolves azurerm_iothub to Microsoft.Devices/iotHubs, and
Microsoft.Devices IS a namespace_routing gate - so the corpus's halt describes
behaviour the skill no longer produces, while every shape assertion still passes.
That is the hand-authored-golden-drifts-from-a-prose-ref trap in its dangerous
direction: green and wrong. A golden cannot notice that a prose rule changed
underneath it, so the check reads the rule and the artifact TOGETHER.

Left red on purpose. A knowingly-stale green is worse than a red with a precise
reason, and the message names the fix: regenerate the golden with a capability run.
Consequence worth stating plainly - Design should now COMPLETE on the corpus, which
makes Estimate reachable for the first time.

Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check,
lint:model-ids green; drift:check OK (347 identical). fixtures:assert has the one
intentional azure design failure above; the other 10 asserters pass in both trees.
lint:md and fmt:check remain UNVERIFIED - dprint and markdownlint absent.
…day's work

Ran the capability test on the corpus: fresh isolated agent, answer keys hard
prohibited, pointed only at SKILL.md. All three phases completed with HANDOFF_OK
and NO halt, so the derivation design works end to end and Design completes on
the corpus for the first time. It also found 10 file-level contradictions and 20
under-specifications. The worst were mine, from the same day.

## The one that mattered most (NOTES 2.8)

extract-terraform.md said a type with no per-type-attributes row carries no sizing
attributes. design-infra.md decides cost-bearing-ness by asking whether config has
a SKU/tier/capacity property. A DERIVED type has no row by definition, so its SKU
was dropped and it was then GUARANTEED to look benign — reintroducing, from the
other end, the exact silent under-report that decision 13.1d exists to prevent.

The run carried the sku anyway and recorded that no rule authorised it. It only
escaped mattering here because the namespace gate fired first; on an estate whose
derived type lands in a rubric namespace it would have understated a cost-bearing
resource.

Fixed by generalising the rule Step 2a already states for AzAPI: always extract
sku / sku_name / tier / capacity / size, for EVERY type, listed or not.

## Contradictions closed (all introduced yesterday)

- NOTES 1.1 — discover-iac.md still said an unlisted type is SKIPPED, against three
  files saying it is derived and retained. It is the FIRST file the fragment loads,
  so two readings of one estate gave a completed design versus a halt. Replaced with
  the derive-cross-check-or-stop rule in full.
- NOTES 1.2 — extract-terraform.md's own closing checklist required every azure_type
  to appear in the canonicalization table, which Step 2b exists to admit exceptions
  to. A correct derived entry failed the checklist. Now states both provenances, and
  gains a checklist item for the SKU rule above.
- NOTES 1.10 — two status blocks asserted files were missing that had landed hours
  earlier: clarify-global.md on azure-region-map.json (the run correctly used the
  table and suppressed the "absent" flag it was told to emit), and design.md listing
  five rubrics as "still to come". Both corrected, and design.md now says outright
  that the filesystem wins over the paragraph — a prose status list goes stale the
  moment a file lands, and the halt decision reads from disk.

## Skill files no longer cite their own oracle

Four places named check_expected_design.py as the authority on rules they did not
state. Three were mine from yesterday. The run reported being STRONGLY TEMPTED to
open it, and put the reason precisely: "a skill that cites its own oracle as a
normative reference is teaching by pointing at the answer key."

That is worse than untidy — reading an asserter invalidates a capability run, which
is how this skill is tested, so the citation actively degrades the test method. All
four replaced by the rule stated where it belongs. design-infra.md now says: if a
rule is only discoverable by reading test code, the rule is missing and the skill
file is the bug.

## service_id was unreproducible, and the oracle keyed on it

The schema specified it as "<stable slug, unique within the artifact>". Capability
run 4 produced 8 of 14 ids differently from the golden — every one a valid stable
slug — and three size assertions failed on a NAMING difference rather than a wrong
size.

An identifier nobody can reproduce cannot be referenced from a report, cross-checked
between artifacts, or asserted by a fixture, and the instability is invisible because
each run is internally consistent. Two fixes, both needed:

- schema-design-aws.md gains a derivation rule: <aws-service-slug>-<local-name>, no
  abbreviating, because an abbreviation is a choice and choices are what made this
  unreproducible.
- the asserter now pins by azure_id, which is built by a stated rule and is the same
  string in every artifact that mentions the resource. Nothing may key on service_id
  across artifacts.

## Still deliberately red, unchanged

check_halt_not_stale still fails on after-design-halted, and the capability run still
fails the no-halt and 29-resource expectations. Same single root cause: those
expectations describe pre-derivation behaviour. Regenerating the goldens from this
run — which also fixes the Cosmos clustering error the run exposed in the golden —
is the next step and was scoped out of this change.

Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check, lint:model-ids
green; drift:check OK (347 identical). lint:md and fmt:check UNVERIFIED.
…ing halts for

want of knowledge

gcp-to-aws serves customers today, and checking WHY it handles every case settled
this. It is not mapping breadth: gcp lists 28 google_* types against azure's 65
index rows and 92 fast-path rows. It is three properties, and only one is knowledge:

  1. complete CATEGORY coverage — 9 rubrics, all on disk
  2. a permissive router — unknown type -> category -> rubric answers
  3. NO veto — nothing can reject a type once a category is found

  and no missing-rubric halt rule at all.

Azure had double the per-type coverage and stopped more often, which is the wrong
trade. It had acquired two azure-only obstacles:

- a 55-entry namespace list that could VETO a derived type. That is the same failure
  as a 1089-row table, one level up: 14 of 15 sampled plausible namespaces were
  absent, so Microsoft.Maps/accounts halted a run despite being unambiguous.
- an explicit prohibition, written by me yesterday: "Do NOT guess a category from
  the type NAME."

## What changed

**The namespace cross-check is now a SIGNAL, not a veto.** Recognised ->
azure_type_provenance "derived". Unrecognised -> "derived_uncorroborated" plus a
type_derived_uncorroborated warning. Either way the resource is KEPT with its full
config. It records whether a second artefact agreed; it does not decide whether the
resource exists.

**An unrecognised namespace routes to a model-chosen category.** Pick the best-fit
rubric from those on disk, apply its six criteria like any other pass-2 resource,
record routing_provenance "model_category", warn with the namespace and the category
chosen, take confidence inferred. Not a stop.

**iac_metadata.untranslated_types now means one thing only: the skill cannot NAME
the service.** Not "no table row", not "namespace unrecognised" — both of those are
derived and retained. Expect it empty on almost every repo.

**A STOP survives only where the model genuinely cannot say what a service does.**
That is a real answer, and it is rare.

## Deviating from gcp in exactly one way, deliberately

gcp records NO provenance for a pattern-routed category — it uses the category and
sets confidence inferred, and nothing in the artifact says the category was guessed.
Azure keeps the provenance field, because that is what makes a derived decision
reviewable instead of indistinguishable from a curated one. Mimic the behaviour, not
the blind spot.

Azure's router is also better than gcp's on the same axis: gcp matches a SUBSTRING of
the type name ("log" -> monitoring), which also matches google_diaLOGflow_agent. The
provider namespace is structural and does not have that failure mode.

## The line this settles

Fall back to the model for FACTS about Azure — what a service is, what a SKU's vCPU
count is, which category a namespace belongs to. Never for this project's OPINIONS:
Elastic Beanstalk over Fargate for PaaS posture, x86_64 over Graviton, single-AZ plus
a visible finding rather than inferring Aurora from a zone-redundant source. A
model-chosen category is fact-shaped; the rubric it lands in still supplies the
opinion. That is why the missing-rubric halt stays and the namespace veto goes.

Recorded in design-infra.md so the reasoning survives, not just the rule.

## Effect

Nothing in the five example estates blocks Design any more — three-tier-saas's
azurerm_maps_account was the last one, and Microsoft.Maps not being in the list is now
irrelevant to whether it resolves.

Enums widened: azure_type_provenance gains derived_uncorroborated,
routing_provenance gains model_category, and the deterministic guard now names all
three derived routes explicitly.

Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check, lint:model-ids
green; drift:check OK (347 identical). The azure design asserter still fails on the
two pre-derivation expectations and the stale halt — same root cause, still scoped to
the golden regeneration. lint:md and fmt:check UNVERIFIED.
The whole suite is green again, and the red it clears was real: two goldens were
describing behaviour the skill no longer produces, while every shape assertion
still passed against them.

## after-discover — replaced, and it fixes an error the golden had

Adopted run 4's output verbatim: 30 resources (was 29) because a derived type is now
RETAINED, provenance per entry, and config.tf_file populated on 30 of 30 where the
hand-authored golden had it empty on 24 of 29 against a ref that mandates it.

Clusters 4 -> 5, and this one was a defect in the golden rather than a consequence of
my change. The golden put the Cosmos account inside the 21-member cluster justified
"merge:cross_group_edges" -- but that account has ZERO edges in either direction and
none of the cluster's three justifying edges touches it. Run 4 split it out, which is
what the algorithm produces: an edgeless primary-eligible resource in its own resource
group is its own workload (13.4b), and a split is justified by an ABSENCE (13.3d) so
its edges[] is legitimately empty.

The old number came from intuition -- "the app must use the database" -- and the corpus
declares no such reference. That is 13.4e firing a THIRD time: a golden or worked
example must be TRACED against the algorithm, never authored by hand.

## after-design-halted -> after-design

The halt became unreachable: azurerm_iothub derives to Microsoft.Devices/iothubs,
Microsoft.Devices is a namespace_routing gate, and the resource defers cleanly. A
golden for a state the skill cannot produce is worse than none, so the directory is
renamed and the assertion INVERTED -- the oracle now requires that halt NOT occur, and
requires the derived type in deferred[]. Mapped or deferred, never quietly skipped.

The benign-skip guard survives unchanged and is still the sharpest assertion here.

## Two more instances of the oracle asserting what no skill file taught

Run 4's clarify output failed three checks, and two were the oracle's fault:

- clarify-compute.md showed a SINGLE-element app_service_plans example and never said
  one row per plan. Run 4 emitted only the plan it questioned, correctly by the ref.
  Now stated: one row per plan ALWAYS, including plans you do not ask about, with
  disposition N/A and a reason for 0- and 1-app plans -- because design-infra.md's
  fan-in rule looks a plan up here and a plan with no row is indistinguishable from a
  plan the user declined to split.
- NOTHING required licensing._fired or _firing_reason. Now required, with the reason:
  "the category did not fire" and "it fired and the fragment forgot to write it down"
  otherwise produce the same artifact, and the second silently drops a licensing
  decision worth real money.

clarify_status was in the same category -- asserted by the oracle, defined nowhere.
Now defined in clarify-assemble.md as COMPLETE | BLOCKED_ON_ESSENTIAL, including the
rule run 4 had to invent: a user answer that conflicts with a hard_blocker does NOT
block. Record it as given with a conflict key, put the blocker in licensing.blockers[]
with severity blocker, and keep status COMPLETE -- the customer answered, and the
blocker is a prerequisite rather than a competing preference.

## What I deliberately did NOT do

after-clarify-complete/ is NOT shipped. Run 4 produced one, but it predates the two
rules above, and hand-editing it to comply would be authoring a golden by intuition --
the exact failure that has now bitten three times. clarify-answers-complete.json is
committed with the completing answer set and a _no_golden_yet note explaining that the
next capability run against the corrected files produces it properly.

after-clarify/ (BLOCKED) is kept: the scripted user there declines to state spend, and
that branch is what proves ESSENTIAL + value:null gates the phase. Its clusters[] list
gained the new cluster id -- mechanically forced, since preferences.clusters[] must key
exactly to the clusters artifact.

## Also fixed

The canonical-type coverage check read EVERY Microsoft.* string in
arm-type-canonicalization.md, so a type named in prose looked like a table row and
reported a false orphan (Microsoft.Maps/accounts, which I had used as an example). It
now reads table rows only.

Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check, lint:model-ids
green; drift:check OK (347 identical); fixtures:assert PASS in BOTH trees -- 11
asserters, 7 golden, 4 smoke, no deliberate failures left. lint:md and fmt:check
UNVERIFIED.
…missing its root volume

Two findings from capability run 4, both of which understate or misdirect a real
estate while leaving every test green.

## NOTES 1.6 — database.md read the availability answer from the wrong path

database.md's override-gate table keyed on `design_constraints.availability`. The
answer lives at **`data.availability`** — clarify-database.md writes it there,
schema-preferences.md documents it there, and the committed golden has it there.
database.md was the ONLY reader, and it was looking in the place gcp-to-aws keeps it.

The bug was invisible rather than loud, which is what makes it worth a commit
message: looking in the wrong place finds nothing, and the gate's own "if the answer
is absent, default to RDS single-AZ" rule then produces the SAME answer the corpus
expects. So every oracle stayed green. On an estate where the customer answers
`multi-az-ha` it silently produces RDS where Aurora was chosen — and records the
customer's answer faithfully, right next to a target that ignores it.

Fixed, with the reason it stayed hidden written into the file so the next person does
not re-port the gcp path.

## NOTES 2.10 — the VM's root volume was absent from the design entirely

A VM's OS disk is normally an INLINE BLOCK, not an azurerm_managed_disk:

    os_disk { caching = "ReadWrite", storage_account_type = "Premium_LRS" }

So no Microsoft.Compute/disks resource exists, a resource-block walk misses it, and
no EBS volume reaches the design. disk-ebs-sizing.json already carried the RULE (OS
disk becomes the instance's root volume in aws_config, never its own entry) — what
was missing was the EXTRACTION. Run 4 recorded a prose note because no rule let it do
better, and flagged that this systematically understates every VM estimate.

extract-terraform.md gains a section for inline blocks that are not resources, and
os_disk joins the VM and VMSS attribute rows. Design maps it to aws_config.root_volume
using the tier map: Premium_LRS -> gp3, a lookup rather than a judgement.

**When disk_size_gb is absent the size is the image default, and the rule now forbids
supplying one.** Carry size_gib null with size_source "image_default_unstated". A
remembered 127 GiB is indistinguishable from a measured one and Estimate would price
it as fact; a stated unknown is worth more than a plausible number.

## Goldens

Both edits to the goldens are mechanical reads of committed input, not judgements: the
inventory's VM gains config.os_disk verbatim from compute.tf (with disk_size_gb null
because the source omits it), and the design's EC2 entry trades its prose note for a
structured root_volume whose volume_type comes from the tier map.

check_inline_os_disk asserts it four ways, all mutation-tested: reverting to a prose
note fails; a wrong volume_type fails; INVENTING the image-default size fails; and
emitting the OS disk as its own services[] entry fails, because counting it twice is
the commonest way a VM estate's storage cost gets inflated.

Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check, lint:model-ids
green; drift:check OK; fixtures:assert PASS both trees. lint:md and fmt:check
UNVERIFIED.
… clarify branch

Run 5: fresh isolated agent, answer keys hard-prohibited, pointed only at SKILL.md.
All three phases HANDOFF_OK, no halt. Discover and Design PASSED their oracles
unmodified — the first time a capability run has cleared them on first contact.

## What run 5 corroborated

It independently produced exactly what I had hand-added yesterday: config.os_disk
with disk_size_gb null, and a root_volume of gp3 / size_gib null /
size_source "image_default_unstated" — plus the gp3 baseline_iops I had not
bothered with. Same 5 clusters, same 16 services.

So the goldens are now the RUN's output rather than my edits. Per 13.4e a golden
must be TRACED, not authored, and an earned artifact that happens to match a hand
edit is still worth more than the hand edit.

## The completing clarify branch finally exists

after-clarify-complete/ + expected-clarify-complete.json, from run 5. Both branches
of the phase are now asserted, and they are opposite states on purpose:

- after-clarify/ — the scripted user DECLINES to state spend, so an ESSENTIAL row is
  null and the phase must GATE. Proves 13.5b.
- after-clarify-complete/ — spend answered, so the phase must COMPLETE. This is what
  makes Design reachable, and it additionally pins the conflicting-answer rule.

check_conflicting_answer is new: vm_cutover mgn against the Azure-Edition hard
blocker must be recorded AS GIVEN with a conflict key, the blocker must appear in
licensing.blockers[], and clarify_status must stay COMPLETE. Run 5 got all three
right and cited clarify-assemble.md:106-111 as specifying it "completely".

check_expected_clarify.py takes an optional spec argument so one asserter serves both;
check_expected_clarify_complete.py is a thin wrapper only because run-asserters.py
keys its registry on a bare path and cannot pass arguments.

## The finding worth the most: nine names for one field

Run 5's clarify failed one assertion — 9 rows DETECTED with value == default and no
stated source. The content was entirely correct; the KEY NAMES were `note` and
`mapped_from` where the oracle accepted `source`, `forced_by` or `reason`. Nothing in
the skill named the key. Across the golden and one run there were NINE different
justification key names in play: source, reason, note, mapped_from, forced_by,
context, conflict, source_ha_context, residency_warning.

This is the third instance of the same defect class — service_id, then cluster_id,
now this. A field a fixture keys on has to be named by a rule.

schema-preferences.md now fixes the key BY DISPOSITION: DETECTED needs `source`
(what in the estate was read), N/A needs `reason`, ESSENTIAL needs `context`, and
`forced_by` takes precedence for a value a blocker removed the choice on. Extra keys
are allowed alongside; they never substitute.

The oracle narrowed to source|forced_by, and row_dispositions' requires_keys were
aligned to follow the disposition — several rows were demanding `reason` on a
DETECTED row, which is the N/A key, so the assertion was asking for the wrong field.
cpu_architecture requires forced_by, because x86_64 there is not a free reading of
the estate: a Windows workload removes Graviton as an option.

Also fixed a bug in my own new check — it used get(), which returns None for anything
that is not a dict, so a list-valued path read as empty and reported a blocker run 5
had recorded correctly.

## Adopted with two mechanical edits, declared

The run's DETECTED rows had `note`/`mapped_from` renamed to `source` (a key rename;
content untouched, and the name was unspecified when the run started), and 5 cluster
pattern_id rows gained a `source` whose content is fixed by 13.3c plus the fact that
patterns.md is absent — one possible reason, not a judgement. Both recorded in the
artifact.

Gates: lint:frontmatter 7 both trees, shared:check, fixtures:check, lint:model-ids
green; drift:check OK (347 identical); fixtures:assert PASS both trees — 12
asserters, 8 golden, 4 smoke. lint:md and fmt:check UNVERIFIED.
@icarthick

Copy link
Copy Markdown
Collaborator Author

Verification evidence — round-2 report finding fixed (+ AI-only companion gap)

Ran the skill end-to-end (headless claude -p, advisor plugin) on two real Azure estates and re-ran your validator. Both report findings now pass; details below.

1. The 35-failure check now returns REPORT_OK

Repo: az-enterprise-fintech (App Service Plan + PostgreSQL Flexible Server + Cosmos + Redis + Front Door) · run 0918-1426 · phase: complete, run_mode: decide_and_execute · migration-report.html = 45 KB.

The exact validator invocation from your finding:

$ python3 scripts/validate-migration-report.py migration-report.html \
    --estimation-infra estimation-infra.json --migration-dir .
REPORT_OK | structure=complete | sections=10/10 |
optional=exec-share,exec-architecture,exec-optimization,appendix-config,appendix-optimization,appendix-glossary

(Was: a 41-line, 4-section stub → 35 failures.)

2. Savings Plans / Reserved Instances section — the monetary omission — is present with product-specific rows

appendix-optimization (Appendix B.1) now renders, sourced from estimation-infra.json's 7 product-specific optimization_opportunities (compute_savings_plan, database_savings_plan, reserved_instances, dynamodb_reserved_capacity, elasticache_reserved_nodes, s3_intelligent_tiering, commitment_state_summary):

Optimization Target Est. savings Commitment
Compute Savings Plan Elastic Beanstalk / EC2 20–66% 1-yr or 3-yr
Database Savings Plan RDS PostgreSQL + DynamoDB ~20% (provisioned) 1-yr or 3-yr; mutually exclusive with RDS RIs
RDS Reserved Instances RDS PostgreSQL up to 69% (~$110/mo) 1-yr or 3-yr; reprice single-AZ first
DynamoDB reserved capacity DynamoDB up to 54% (1-yr) / 77% (3-yr) 1-yr or 3-yr

Not a single generic "commitment discounts" row, and it carries the contract caveats (incremental to Balanced on-demand, not additive on Optimized, RDS SP/RI mutual exclusivity).

3. Stable Appendix A–J lettering

<h2> and TOC render A, B, C, E, F, J. Appendix D (AI) is legitimately absent (infra-only repo) and E is not renumbered to D — conditional sections keep their reserved letters, per the Step 3.75 map.

4. Companion gap found + fixed: AI-only runs were producing no report at all

Testing surfaced a different instance of the same silent-degradation class: on an app-code-only AI repo (az-langgraph-research-agent), Generate completed (run_mode: decide_and_execute, phase: complete) but produced no migration-report.html — generate.md was authored as a strict infra-shaped DSL phase (hardcoded aws-design.json/terraform/main.tf in _produces/pre/postconditions), so the AI-only path that Discover/Clarify-ai-only/Design-ai/Estimate-ai support end-to-end never rendered a report and self-certified past it.

Fixed by making generate.md's _input/_produces/fragment-triggers conditional (_when infra / _when AI) and adding a --mode ai_only to the validator (AI-appropriate required-section set). After the fix, a fresh run of that repo:

$ python3 scripts/validate-migration-report.py migration-report.html \
    --mode ai_only --estimation-ai estimation-ai.json --migration-dir .
REPORT_OK | mode=ai_only | structure=complete | sections=7/7 |
optional=exec-share,exec-optimization,appendix-ai,appendix-config,appendix-glossary

Mode is discriminating, not a rubber stamp — the same report fails full mode (missing infra sections). 67 existing validator unit tests still pass.

Caveat (so the numbers aren't over-read)

These runs priced via a search fallback (awspricing MCP unreachable in this environment) and ran passed_degraded_offline (no terraform binary; policy gate ran → POLICY_OK), with is_floor: true on 9 lines. The report structure/completeness fixes are what this evidence validates — which is exactly the scope of the finding; the dollar figures are estimates, not authoritative.

Commits: 5ed068f9 (round-2 report fix) · 916a9762 (AI-only report-path gap).

@herosjourney herosjourney left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

User-facing surfaces: Azure is announced everywhere except the three plugin descriptions

Checked whether the PR notifies users of the new Azure→AWS capability across the README / marketplace / plugin-manifest surfaces (advisor tree only — confirmed the PR correctly touches zero migrate-tree files, which matches the plan to deprecate that tree).

Correctly updated — these announce Azure well:

  • .claude-plugin/marketplace.json — both plugin descriptions now lead with “Azure, GCP, or Heroku” and name the concrete mappings (App Service → EB, AKS → EKS, Azure OpenAI → Bedrock). ✅
  • top-level README.md — azure-to-aws listed in both plugin rows. ✅
  • advisor/README.md — full azure-to-aws skill bullet with mappings, the intent-routing table, and the MCP-server line. ✅
  • advisor/plugins/aws-startup-advisor/setup.md — install list + skill description. ✅
  • advisor/AGENTS.md. ✅
  • All three plugin manifests’ keywords arrays — azure, azure-to-aws, azure-sql, azure-openai. ✅

Gap — the human-readable description prose in all three plugin manifests still omits Azure, even though the keywords and the marketplace entry include it. This is the sentence a user actually reads in the Claude Code / Cursor / Codex plugin UI when deciding whether to install, so it now contradicts the marketplace entry:

File Field Current prose Fix
.claude-plugin/plugin.json (line 5) description “Migrate to AWS from GCP or Heroku—…” “Migrate to AWS from Azure, GCP, or Heroku—…”
.cursor-plugin/plugin.json (line 4) description “Migrate to AWS from GCP or Heroku—…” “Migrate to AWS from Azure, GCP, or Heroku—…”
.codex-plugin/plugin.json (line 94) longDescription “Migrate from GCP, Heroku, OpenAI, or Gemini…” “Migrate from Azure, GCP, Heroku, OpenAI, or Gemini…” (inline comment below)

Suggest also naming an Azure primitive or two in the claude/cursor description for parity with the marketplace copy (e.g. “…including Azure App Service → Elastic Beanstalk, AKS → EKS, and Azure OpenAI → Bedrock…”), so the in-agent install card matches what the marketplace promises.

None of this blocks the skill — it works and is discoverable via keywords — but a user reading the plugin’s own description in-agent won’t see Azure listed as a supported source, which undercuts the launch you’re trying to announce.

Verification note: read against the PR head (916a9762) diffed vs main @ b30cb39b; the claude/cursor description lines fall outside the PR’s diff hunks (only their keyword arrays changed), so those two are flagged here in the body rather than inline.

"displayName": "AWS Startup Advisor",
"shortDescription": "Build and migrate on AWS with startup-focused architecture, cost, and security guidance plus GCP/Heroku-to-AWS and AI-stack (Bedrock, agents) migration workflows.",
"longDescription": "Personalized architecture, cost, security, and migration guidance for startups. From day-one account setup and security baselines to production-ready infrastructure, cost optimization, and beyond. Includes AWS Activate Credits eligibility, 60+ exclusive startup offers, and multi-account multi-region support. Built on expertise from AWS Startup Solutions Architects and patterns from 350,000+ startups.\n\nFeatures\nVetted prompts for startups at every stage:\n\n\nDay-one account setup, security baseline, least-privilege roles\nScaffold a production architecture for your stack\nSet up GuardDuty, Security Hub, and a vulnerability scanner\nMigrate from GCP, Heroku, OpenAI, or Gemini — infrastructure, AI SDKs to Amazon Bedrock, and agentic systems\n\n\nAWS Activate Benefits\nCheck eligibility and explore offers available to startups on AWS:\n\n\nAWS Activate Credits: Check eligibility and learn how to apply for credits to offset costs across 200+ services, including infrastructure, data services, and AI/ML models on Bedrock.\nExclusive Startup Offers: Access discounts and extended trials on dozens of tools for payments, analytics, communications, and developer productivity.",
"longDescription": "Personalized architecture, cost, security, and migration guidance for startups. From day-one account setup and security baselines to production-ready infrastructure, cost optimization, and beyond. Includes AWS Activate Credits eligibility, 60+ exclusive startup offers, and multi-account multi-region support. Built on expertise from AWS Startup Solutions Architects and patterns from 350,000+ startups.\n\nFeatures\nVetted prompts for startups at every stage:\n\n\nDay-one account setup, security baseline, least-privilege roles\nScaffold a production architecture for your stack\nSet up GuardDuty, Security Hub, and a vulnerability scanner\nMigrate from GCP, Heroku, OpenAI, or Gemini \u2014 infrastructure, AI SDKs to Amazon Bedrock, and agentic systems\n\n\nAWS Activate Benefits\nCheck eligibility and explore offers available to startups on AWS:\n\n\nAWS Activate Credits: Check eligibility and learn how to apply for credits to offset costs across 200+ services, including infrastructure, data services, and AI/ML models on Bedrock.\nExclusive Startup Offers: Access discounts and extended trials on dozens of tools for payments, analytics, communications, and developer productivity.",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This longDescription still reads “Migrate from GCP, Heroku, OpenAI, or Gemini” — Azure isn’t listed, even though this PR adds azure/azure-to-aws/azure-sql/azure-openai to the keywords just below and the top-level marketplace.json now leads with “Azure, GCP, or Heroku.” This is the description a user reads in the Codex plugin UI, so it now contradicts the marketplace entry. Add Azure to the sentence, e.g. “Migrate from Azure, GCP, Heroku, OpenAI, or Gemini — infrastructure, AI SDKs to Amazon Bedrock, and agentic systems.” The same fix is needed in the description field of .claude-plugin/plugin.json and .cursor-plugin/plugin.json (both still say “GCP or Heroku”; flagged in the review body since those lines are outside this PR’s diff).

@herosjourney herosjourney left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Three correctness findings, all verified against this branch (head 916a9762) vs main @ b30cb39b. None is a style nit; each is a way the skill either ships a false statement or advertises a capability that halts.


1. This PR promotes a lifecycle file that reasserts the “6-month Legacy” error #307 is fixing — and they clobber each other silently

Confirmed: #304 creates the canonical skills/shared/ai/ai-model-lifecycle.md (plus vendored copies under each skill’s references/vendored/ai/), and line 11 reads “Legacy (minimum 6 months before EOL)” with no mention of the 45-day policy, the 2026-09-07 split, model-card governance, or the EOL’d rows. That is exactly the false statement open PR #307 exists to correct.

The important nuance: the two PRs touch different paths, so git will not raise a merge conflict.

  • #307 edits the old path skills/gcp-to-aws/references/**shared**/ai-model-lifecycle.md.
  • #304 deletes that path and relocates to skills/**shared**/ai/ + references/**vendored**/ai/.

So whichever lands second silently wins, with no conflict marker to force a human to reconcile: if #304 lands after #307, the stale content silently reappears in the new canonical copy; if #307 lands after #304, it patches a path that no longer exists. #307 is currently OPEN and main still says “minimum 6 months,” so the risk is live.

Fix: a plain rebase is not enough (the paths differ, so #307’s diff won’t apply to #304’s new files). Port #307’s corrected content into #304’s new canonical skills/shared/ai/ai-model-lifecycle.md, then re-sync the vendored copies from it (they are byte-identical to the canonical today — keep them so). Verify the canonical copy carries the 45-day/2026-09-07 policy before merge, since it is the source every vendored copy derives from and no conflict will protect it.


2. The skill advertises Bicep / ARM / live az in its triggers, but those paths halt or don’t exist

Confirmed. SKILL.md’s description triggers include “migrate Bicep to Terraform,” “migrate ARM templates to Terraform,” and a first-class “read-only consent-gated live az CLI capture.” But:

  • discover-iac.md:208 states plainly: “Until those two refs exist, a workspace containing .bicep or ARM templates halts.” And references/shared/extract-bicep.md / extract-arm.md do not exist on this branch.
  • There is no live-az capture fragment in references/phases/discover/ (only discover-iac.md, discover-app-code.md, discover-assemble.md) and no Azure live-capture fixture.

So a user who says “migrate my Bicep to Terraform” or expects the advertised live az capture gets matched, then hits a halt. Discover itself is honest — the problem is the marketing in the trigger list and the 7-phase description overpromising relative to what’s wired.

Fix: make the description Terraform-first and name Bicep / ARM / live-az as not-yet-supported (or remove them from the triggers) until extract-bicep.md, extract-arm.md, and the az fragment land. Inline comment on the trigger line below.


3. llm-to-bedrock still hardcodes a “gcp-to-aws” handoff with no Azure branch

Confirmed. skills/llm-to-bedrock/SKILL.md:127 reads: “I'm now invoking the gcp-to-aws Assess skill to discover your AI workloads and design…” — there is no Azure signal, azure-to-aws is not mentioned anywhere in that skill, and #304 does not touch llm-to-bedrock at all. Once the Azure skill ships, an Azure user routed into the AI path is told the tool is invoking “gcp-to-aws,” which is a user-visible inconsistency.

The PR body already identified this as the minimal fix and deferred it. That’s a defensible sequencing call (it pairs with #290, which is already merged) — but it should be an explicit decision, not a silent gap. Either add the one-line “Azure estate → azure-to-aws” branch/note to llm-to-bedrock/SKILL.md, or state in the PR that it is intentionally out of scope and tracked as a follow-up. Leaving the hardcoded “gcp-to-aws” sentence while shipping an Azure skill is the remaining user-visible footgun.


Verification note: findings 1 and 3 are flagged in this body rather than inline — finding 1’s file was relocated (the anchor path is ambiguous across the move) and finding 3’s file is untouched by this PR, so neither line sits in the diff. Finding 2 is anchored inline on the SKILL.md trigger line.

@@ -0,0 +1,221 @@
---
name: azure-to-aws
description: "Migrate workloads from Microsoft Azure to AWS. Triggers on: migrate from Azure, Azure to AWS, move off Azure, migrate Azure app to AWS, migrate AKS to EKS, migrate App Service to AWS, migrate App Service to Elastic Beanstalk, migrate Azure VMs to EC2, migrate Azure SQL to RDS, migrate Azure Database for PostgreSQL to RDS, migrate Azure Database for MySQL to RDS, migrate Cosmos DB to DynamoDB, migrate Azure Cache for Redis to ElastiCache, migrate Blob Storage to S3, migrate Azure Functions to Lambda, migrate Service Bus to SQS, migrate Event Hubs to Kinesis or MSK, migrate Azure OpenAI to Bedrock, migrate my Azure OpenAI app to AWS, move Azure AI workloads to AWS, migrate Azure-hosted agentic workloads to AWS, migrate Azure-hosted LangChain to Bedrock, migrate Azure-hosted LangGraph to AWS, migrate Bicep to Terraform, migrate ARM templates to Terraform, leave Azure, estimate AWS costs for my Azure infrastructure, what-if workshop, reprice Azure migration, compare migration scenarios, workshop mode. Runs a 7-phase process: discover Azure resources from Terraform/Bicep/ARM templates, a read-only consent-gated live `az` CLI capture, an optional Resource Discovery for Azure report, application code, and optional billing exports; clarify migration requirements via an assumption sheet; design AWS architecture; estimate costs (both a 1:1 lift and a right-sized target); optionally reprice what-if scenarios in a workshop sidebar; generate migration artifacts when the user opts in; and collect optional feedback. Clarify must finish before Design, Estimate, or Generate, and Generate is opt-in — it runs only after the user chooses Execute at the post-Estimate decision gate. Uses resource-group-seeded clustering refined by typed edges from ARM resource IDs, canonical `Microsoft.*` ARM resource types as the mapping key, a deterministic fast-path table for architecture-invariant primitives, and a pattern catalog so recommendations describe workloads rather than isolated resources. Do not use for: GCP migrations (see gcp-to-aws), Heroku migrations (see heroku-to-aws), general AWS architecture advice without migration intent (see architect-for-startups), AWS-to-Azure reverse migration, on-premises-into-Azure discovery (that is Azure Migrate's job, not this skill's), or Azure-to-Azure refactoring."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These triggers advertise “migrate Bicep to Terraform,” “migrate ARM templates to Terraform,” and a “read-only consent-gated live az CLI capture,” but none of those paths is wired yet: discover-iac.md:208 says a workspace with .bicep or ARM halts until extract-bicep.md / extract-arm.md exist (they don’t on this branch), and there is no live-az capture fragment or fixture in references/phases/discover/. A user whose ask matches one of these triggers gets routed in and then stops. Make the description Terraform-first and name Bicep / ARM / live-az as not-yet-supported (or drop them from the trigger list) until those extractors and the az fragment land. The 7-phase summary sentence (“discover … from Terraform/Bicep/ARM … a read-only consent-gated live az CLI capture”) should be softened the same way.

@herosjourney

Copy link
Copy Markdown
Contributor

Blocking checklist item: canonical AI lifecycle file must carry #307\u2019s 45-day policy before merge

Sharpening finding #1 from my earlier review into a concrete, blocking action now that the plan is settled: #307 is being folded into this PR and closed as superseded. #307\u2019s fix is correct and needs to land within ~10-14 days; this PR merges (~7 days) inside that window, so it should carry the fix rather than have #307 merge separately on a path this PR deletes.\n\nThe risk this guards against: this PR\u2019s canonical skills/shared/ai/ai-model-lifecycle.md currently still reads \u201cLegacy (minimum 6 months before EOL)\u201d with no mention of the 45-day policy, the 2026-09-07 split, model-card governance, or the EOL\u2019d rows \u2014 exactly the false statement #307 fixes. Because this PR relocates the file (deleting gcp-to-aws/references/shared/ai-model-lifecycle.md), there is no merge conflict to force reconciliation, and the stale canonical file is vendored byte-identically into 4 copies that shared:check will actively keep in sync with the wrong content.\n\nBlocking checklist before merge:\n- [ ] Canonical skills/shared/ai/ai-model-lifecycle.md carries #307\u2019s content: the 45-day / 2026-09-07 split policy, model-card governance language, and the EOL\u2019d rows (Claude 3 Haiku, Nova Premier v1, Nova Sonic v1, etc.) \u2014 not the stale \u201cminimum 6 months\u201d text.\n- [ ] shared:sync run so all vendored references/vendored/ai/ai-model-lifecycle.md copies match; shared:check green.\n- [ ] Applied in both trees if shared:check enforces cross-tree byte-identity (advisor + migrate), even though the migrate tree is being deprecated \u2014 otherwise CI fails.\n- [ ] #307 closed as superseded only after this content is committed on this branch (closing it earlier while this file stays stale loses the fix silently).\n\nThe content already exists verbatim in #307 (head d2473bb4); this is a copy-into-the-new-file + re-sync, not new authoring.\n

@herosjourney

Copy link
Copy Markdown
Contributor

Ready-to-apply: #307's lifecycle fix, merged into this PR's canonical file

Per the plan in the earlier comment — #307 is being folded into this PR and closed as superseded. I've done the merge and verified it; below is the corrected canonical file plus one command to regenerate the vendored copies.

This is a two-way merge, not a copy — I kept #307's facts and this PR's framing:

  • From fix(bedrock): refresh model lifecycle statuses and adopt the 45-day Legacy policy #307: the 45-day / 2026-09-07 split policy, the "don't apply the 6-month assumption post-2026-09-07" warning, both reference links, the EOL'd rows (Claude 3 Haiku, Nova Premier v1, Nova Sonic v1) moved into Removed, the restricted/Covered-Model handling, and the GetFoundationModel unverified-fallback guidance.
  • From this PR: the canonical-file blockquote header (source-agnostic + edit HERE, then shared:sync) and the shared-location refresh wording.

How to apply (edit the canonical, then sync)

  1. Replace advisor/plugins/aws-startup-advisor/skills/shared/ai/ai-model-lifecycle.md with the content in the block below.
  2. Regenerate the vendored copies and confirm byte-identity:
mise run shared:sync     # canonical -> the 4 vendored copies (both trees)
mise run shared:check    # must pass

shared:sync also writes the migrate-tree canonical + vendored copies, so all 6 lifecycle files end up identical. (If you'd rather apply a git patch than paste, I have one that applies cleanly to this branch — say the word and I'll attach it another way.)

Verified before posting

Against this branch (916a9762) with the merge applied: shared:check OK (4 skills/tree, byte-identical) · drift:check OK (397 identical) · lint:model-ids OK · fmt:check clean · lint:md 0 errors · only the 6 lifecycle files touched. Stale global "minimum 6 months before EOL" rule gone; 45-day policy present in every copy; Haiku/Premier/Sonic in Removed.

One judgment call: I kept #307's Sep-21 dated table. Since this PR is ~a week out, those days_to_eol numbers may be slightly stale at merge — the file says to recompute at run time (snapshot, not a runtime dependency), so it's correct as-is, but feel free to bump the date/counts at merge if you want.

Do not close #307 until this content is committed here — closing it early while this file stays stale loses the fix silently (no merge conflict fires, and shared:check would defend the wrong content across all 6 copies).

Merged canonical skills/shared/ai/ai-model-lifecycle.md (paste this)
# Bedrock Model Lifecycle Awareness

> Canonical Bedrock model Active/Legacy/EOL registry and the 90-day exclusion
> rule. Source-cloud agnostic: it is a property of Bedrock's model catalog.
> Vendored into each consuming skill as
> `references/vendored/ai/ai-model-lifecycle.md` and kept byte-identical by
> `shared:check`; edit HERE, then run `shared:sync`.

References:

- [Models launched on or after 2026-09-07](https://docs.aws.amazon.com/bedrock/latest/userguide/model-lifecycle.html)
- [Models launched before 2026-09-07](https://docs.aws.amazon.com/bedrock/latest/userguide/model-lifecycle-legacy.html)

Models on Bedrock move through three states: **Active** → **Legacy** → **End-of-Life (EOL)**. After EOL, the model is unavailable and requests fail.

The notice period a model gets depends on when it launched:

- Models launched **before 2026-09-07** follow the original policy: at least **6 months** in Legacy before EOL. For EOL dates after February 1, 2026, a **public extended access** period begins at least 3 months into the Legacy state, and pricing may increase at the model provider's discretion.
- Models launched **on or after 2026-09-07** are governed by their model card. Each card declares an `EOL no sooner than` date and a Legacy period that is either 6 months or **45 days**. The actual EOL date appears on the card once Legacy begins.

**Do not apply the 6-month assumption to a model launched on or after 2026-09-07** — a 45-day Legacy period leaves less time than a typical migration takes to reach production.

---

## Lifecycle States (Not the Same Thing)

| State      | What it means                                                                                                                                                                                                               | Usable?                |
| ---------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------- |
| **Active** | Provider is actively maintaining the model. Full feature access.                                                                                                                                                            | Yes                    |
| **Legacy** | Deprecated. Still works for existing users, but new Provisioned Throughput cannot be created, new customers cannot onboard, pricing may increase during public extended access, and the model is on a countdown to removal. | Yes, with restrictions |
| **EOL**    | Model is removed. All inference requests fail.                                                                                                                                                                              | **No**                 |

**Legacy does not mean unavailable** — it means the model still functions today but has a firm expiration date. EOL means unavailable.

---

## Selection Rules

### Rule 1: Active models only for new migrations

**New migrations must target Active models only.** Do not recommend a Legacy or EOL model as the primary selection for any new migration, even if it is cheaper.

### Rule 2: 90-day exclusion zone

**Models within 90 days of their EOL date must be excluded from all recommendation and comparison tables.** A migration takes weeks or months to plan, test, and deploy. Recommending a model that will be unavailable before the migration is production-ready is harmful.

- **Excluded** = do not list in "Best Bedrock Match" columns, tiered strategy tables, or `recommended_model` / `backup_model` fields.
- These models may still appear in the pricing cache (for reference by users already on them) but must be marked `excluded (EOL YYYY-MM-DD)` in the Status column.

### Rule 3: Legacy models outside the 90-day zone

Legacy models with >90 days until EOL may appear in comparison tables **with annotation** (`Legacy — EOL YYYY-MM-DD`), but never as `recommended_model` or "Best Bedrock Match" when an Active alternative exists.

### Applying the rules

On each run, compute `days_to_eol = EOL date − today` for every model in the Legacy/EOL table below. Then:

1. `days_to_eol ≤ 0` → EOL. Remove from all tables.
2. `0 < days_to_eol ≤ 90` → **Exclusion zone.** Remove from recommendation/comparison tables. Mark `excluded` in pricing cache.
3. `days_to_eol > 90` and Legacy → Annotate, never recommend as primary.
4. Active → No restrictions.

---

## Legacy / EOL Models (as of September 21, 2026)

For models launched before 2026-09-07, the [legacy lifecycle table](https://docs.aws.amazon.com/bedrock/latest/userguide/model-lifecycle-legacy.html) is authoritative. For models launched on or after that date, the model card and the runtime `modelLifecycle.status` field are authoritative — they will not appear in the table below. The table captures pre-policy-change models referenced elsewhere in this plugin. **Recompute the Status column on each run** using `days_to_eol = EOL date − today`.

| Model              | Model ID                                  | EOL Date   | Days to EOL | Status       | Active Replacement      |
| ------------------ | ----------------------------------------- | ---------- | ----------- | ------------ | ----------------------- |
| Nova Canvas v1     | `amazon.nova-canvas-v1:0`                 | 2026-09-30 | 9           | **excluded** | Stability AI (see note) |
| Nova Reel v1       | `amazon.nova-reel-v1:0` / `v1:1`          | 2026-09-30 | 9           | **excluded** | —                       |
| Claude Sonnet 4    | `anthropic.claude-sonnet-4-20250514-v1:0` | 2026-10-14 | 23          | **excluded** | Claude Sonnet 5 / 4.6   |
| Jamba 1.5 Large    | `ai21.jamba-1-5-large-v1:0`               | 2026-11-26 | 66          | **excluded** | —                       |
| Jamba 1.5 Mini     | `ai21.jamba-1-5-mini-v1:0`                | 2026-11-26 | 66          | **excluded** | —                       |
| Marengo Embed v2.7 | `twelvelabs.marengo-embed-2-7-v1:0`       | 2026-11-30 | 70          | **excluded** | Marengo Embed 3.0       |
| Claude Opus 4.1    | `anthropic.claude-opus-4-1-20250805-v1:0` | 2027-01-08 | 109         | legacy       | Claude Opus 4.8 / 4.6   |

**Notes (as of Sep 21, 2026):** Jamba 1.5 Large / Mini and Marengo Embed v2.7 are inside the 90-day exclusion zone (`excluded`, not `legacy`) — they must no longer appear in recommendation or comparison tables. Jamba 1.5 Large / Mini are also in public extended access, so provider pricing may increase. Claude Opus 4.1 is the only row still outside the exclusion zone.

**Removed (past EOL as of Sep 21, 2026):**

- Titan Image Generator v2 (`amazon.titan-image-generator-v2:0`) — EOL 2026-06-30
- Llama 3.2 all sizes (`meta.llama3-2-*-instruct-v1:0`) — EOL 2026-07-07
- Llama 3.1 405B Instruct (`meta.llama3-1-405b-instruct-v1:0`) — EOL 2026-07-07
- Claude 3 Sonnet (`anthropic.claude-3-sonnet-20240229-v1:0`) — EOL 2026-07-30
- Claude 3.5 Sonnet v1 (`anthropic.claude-3-5-sonnet-20240620-v1:0`) — EOL 2026-07-30
- Claude 3.5 Sonnet v2 (`anthropic.claude-3-5-sonnet-20241022-v2:0`) — EOL 2026-07-30
- Command R / R+ (`cohere.command-r-v1:0` / `cohere.command-r-plus-v1:0`) — EOL 2026-08-19
- Claude 3 Haiku (`anthropic.claude-3-haiku-20240307-v1:0`) — EOL 2026-09-10 (replacement: Claude Haiku 4.5)
- Nova Premier v1 (`amazon.nova-premier-v1:0`) — EOL 2026-09-14 (replacement: Nova 2 Pro)
- Nova Sonic v1 (`amazon.nova-sonic-v1:0`) — EOL 2026-09-14 (replacement: Nova 2 Sonic)

> **AWS page lag:** As of Sep 21, 2026, the legacy lifecycle page still lists rows whose published EOL date has already passed (Command R / R+ among them), even though the same page states that past-EOL rows are dropped. This file treats the **EOL date as authoritative** and keeps those models in Removed rather than the live table, so users already on them still see a warning. Never recommend or invoke a model listed in Removed.

**Status key:** `excluded` = ≤90 days to EOL, must not appear in any recommendation. `legacy` = >90 days to EOL, annotate but do not recommend as primary.

**⚠️ Image generation — Active successor is Stability AI:** Nova Canvas v1 is Legacy and now inside the exclusion zone (EOL 2026-09-30), so it must not appear in recommendation or comparison tables. The Active image generation models on Bedrock are **Stability AI** models:

| Model                      | Model ID                            | Pricing       | Tier     | Use case                              |
| -------------------------- | ----------------------------------- | ------------- | -------- | ------------------------------------- |
| Stable Image Ultra         | `stability.stable-image-ultra-v1:0` | ~$0.08/image  | premium  | Photorealistic, high-end visuals      |
| Stable Diffusion 3.5 Large | `stability.sd3-5-large-v1:0`        | ~$0.065/image | flagship | High volume creative assets           |
| Stable Image Core          | `stability.stable-image-core-v1:0`  | ~$0.04/image  | fast     | Rapid, affordable generation at scale |

When `image_generation` capability is detected:

1. Recommend **Stability AI** models as the primary Active target (not Nova Canvas).
2. Note the pricing model difference: Stability AI charges **per image**, not per token. Direct cost comparison with source provider (DALL-E, Imagen) requires converting to per-image equivalents.
3. If the user's source workload is DALL-E or Imagen, map to Stable Image Ultra (quality-first) or Stable Image Core (cost-first) based on `quality_vs_cost` preference in `preferences.json`.
4. Nova Canvas is inside the exclusion zone and must not appear at all — not as `recommended_model`, and not as a Legacy fallback annotation.

---

## Integration Points

### Design Phase (`design-ai.md`)

After selecting a Bedrock model for each workload:

1. Check the Legacy/EOL table above (or the lifecycle page).
2. If the model is in the **exclusion zone** (≤90 days to EOL) or EOL: reject it. Use the Active replacement.
3. If the model is Legacy but >90 days from EOL: replace with Active replacement if one exists. If no Active replacement exists, note the EOL date and recommend the user plan a follow-up migration.
4. If Active: proceed normally.
5. If `restricted` (Covered Model or gated preview — see the Status table below): do not select it as a default or as `recommended_model` / `backup_model`; name it only when the user explicitly asks for a frontier model, together with its access requirement.

### Estimate Phase (`estimate-ai.md`)

When building the model comparison table:

- **Exclusion zone models**: omit entirely from `model_comparison`. Do not include in `recommended_model` or `backup_model`.
- **Legacy (>90 days)**: include with `(Legacy — EOL YYYY-MM-DD)` annotation. Never use as `recommended_model` if an Active alternative exists.
- **Active**: no restrictions.
- **Restricted** (`restricted (…)` in the pricing-cache Status column): never `recommended_model` or `backup_model`, never a default in a mapping guide. Include in `model_comparison` only when the user explicitly asks about frontier / Covered Models, annotated with the access requirement.

### Pricing Cache (`pricing-cache.md`)

The multi-provider quick reference table includes a `Status` column:

| Status value                | Meaning                                                                                                                                                                                                                                                                                                                                     |
| --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `active`                    | No restrictions                                                                                                                                                                                                                                                                                                                             |
| `legacy (EOL YYYY-MM-DD)`   | Legacy, >90 days from EOL. Listed for reference, annotated.                                                                                                                                                                                                                                                                                 |
| `excluded (EOL YYYY-MM-DD)` | ≤90 days from EOL. Kept for existing users but must not be selected for new migrations.                                                                                                                                                                                                                                                     |
| `restricted (<reason>)`     | Access-restricted: a Covered Model that needs an account-level data-retention opt-in (`aws_review` or `provider_data_share`), or a gated preview. Never `recommended_model` / `backup_model`, never a default; offer only on explicit user request, with the requirement stated. Claude Fable 5 / 5.1 and the Mythos line carry this value. |

When refreshing the cache, recompute `days_to_eol` and refresh the `active` / `legacy` / `excluded` values from the [model lifecycle page](https://docs.aws.amazon.com/bedrock/latest/userguide/model-lifecycle.html). **Preserve an existing `restricted (…)` value** — that page publishes only active / legacy / EOL and does not track access gating, so it can never produce `restricted`; change or remove a `restricted` value only when the model card's access requirement itself changed (data-retention mode, gated preview status).

### Mapping Guides (`ai-openai-to-bedrock.md`, `ai-anthropic-to-bedrock.md`, `ai-gemini-to-bedrock.md`, and any source-cloud-specific guide the consuming skill ships)

- "Best Bedrock Match" columns must only contain Active models.
- Exclusion-zone models must not appear in any recommendation row.
- `restricted` models never appear as a default match; they may be named only as an explicit opt-in alternative with the access requirement stated.
- Legacy models (>90 days) may appear in notes or legacy-source mapping rows but never as the primary recommendation.

---

## Refresh Cadence

**On every design run:** The agent MUST recompute `days_to_eol = EOL date − today` for every row in the table above and apply the four rules in "Applying the rules" before making any model recommendation. The static Days to EOL column in this file is a snapshot only — do not use it directly without recomputing.

**Newer models are not covered by the table.** A model launched on or after 2026-09-07 will never appear in the Legacy/EOL table above, and its absence is not evidence that it is Active. When a candidate is not in the table, verify it by calling `GetFoundationModel` (or `ListFoundationModels`) and reading `modelLifecycle.status`: `LEGACY` and `EOL` are never valid targets for a new migration. If the model is Legacy, read its model card for the actual EOL date and whether the Legacy period is 6 months or 45 days. If neither the API nor the card is reachable, treat the model's lifecycle as **unverified** and say so in the output rather than inferring `active` from a `Status` column in a pricing cache.

**Periodic table refresh:** When the table itself needs updating (new models added, EOL dates changed by AWS, or past-EOL rows to remove), update this file and `pricing-cache.md` together. Edit the canonical `skills/shared/ai/ai-model-lifecycle.md` and run `shared:sync` so every vendored copy stays byte-identical.

**Past-EOL rows:** Once `days_to_eol ≤ 0`, move the model out of the live table into **Removed** and add its ID to `tools/model-id-lint.py` so CI catches any remaining reference to it as a target. Keep the Removed entry long enough that users already on the model still get a warning.

@herosjourney

Copy link
Copy Markdown
Contributor

Proposed: discover-live.md for azure-to-aws (live az capture) — live-validated

This fills the live-az discovery capability the SKILL.md advertises (and that Discover currently halts without — the gap flagged earlier). It's a port of gcp-to-aws/discover-live.md, adapted to this skill's conventions and corrected to match the execution model discover.md already prescribes.

Two things that make it a real port, not a paper one

  1. It honors the _interactive: false split. discover.md says the discover phase runs file-only and the live path "will NOT be a plain fragment … Live capture therefore becomes main-window pre-work invoked from _preconditions … the dispatched fragment merely parsing that directory." The draft implements exactly that: Part A (consent + az capture) runs in the main window; Part B (_fragment: live-parse) is the file-only parse half that reads live-capture/.
  2. No type-translation table. az emits canonical Microsoft.* ARM types natively — the same key the IaC path canonicalizes to — so the live path is a second producer of the identical inventory contract, not a new mapping surface.

Validated live against a real subscription (az-cli 2.90.0)

I ran the read-only commands against a real Azure subscription. Findings that changed the draft:

  • Fast-path is az resource list, NOT az graph query. Confirmed: on a fresh CLI az graph triggers an interactive dynamic-install prompt and hangs — fatal in a non-interactive worker. az resource list covers the same one-call inventory need, is built-in (no extension), and returned Microsoft.* types directly. Draft's fast-path + error table updated accordingly.
  • Added Azure Container Apps (rows 4a/4b), Container Registry (11a), Log Analytics (11b) — all present in the test subscription, all missing from the initial draft. Container Apps → Fargate/App Runner is a first-class startup compute service; its managedEnvironmentId → env hosted_on edge mirrors the App Service Plan fan-in.
  • AI-source keys off the deployment MODEL, not the account kind. The test account was kind: AIServices (umbrella) with mixed deployments: gpt-4.1-mini + text-embedding-3-small (Azure OpenAI) alongside Kimi-K2.6 (a partner model that is NOT). Draft now sets ai_source: "azure_openai" only when a deployment model.name is an OpenAI-family id, and classifies non-OpenAI Foundry models as other so Design can flag "no direct Bedrock equivalent."
  • Verified-working rows (returned projected fields intact, zero extensions): preflight, az resource list, containerapp (4a/4b), postgres flexible-server (7), storage (11), acr (11a), log-analytics (11b), cognitiveservices account+deployment (16).
  • Not yet live-tested (absent from the test subscription): webapp, aks, vm, mysql, cosmos, redis, service bus, event hubs, vnet, key vault, ML workspace, dns. Command shapes are ported from the pattern; they deserve one live pass on a subscription that has them before merge.

To wire it in

Drop the file at advisor/plugins/aws-startup-advisor/skills/azure-to-aws/references/phases/discover/discover-live.md (and the migrate-tree twin if shared:check requires it), then register it in discover.md's step-2 slot: the Part A pre-work invoked from _preconditions, and the Part B live-parse fragment in _fragments with a consent-based trigger. Clean under the repo markdownlint config (MD013 disabled).

Not committed here — this is a fork I can't push to. Handing it over as a proposal; happy to iterate on the field projections or the wiring.

Proposed discover-live.md (live-validated draft)
---
_fragment: live-parse
_of_phase: discover
_contributes:
  - azure-resource-inventory.json (resources[], live_metadata section)
  - azure-resource-clusters.json (simplified_live clustering, or merged with the IaC base)
  - ai-workload-profile.json (minimal iac_cognitive-style profile when a live Cognitive Services / Azure ML signal is present)
---
# Discover — Live Discovery (az CLI)

> **Two-part fragment — read the execution-model note first.** Per `discover.md`,
> the discover phase runs under `_exec: { _agent: rw }` with `_interactive: false`,
> so a dispatched worker CANNOT prompt for consent or run interactive `az`. Live
> discovery is therefore split, exactly as `discover.md` prescribes:
>
> 1. **Part A — main-window pre-work (Steps 0–2).** The consent gate and the actual
>    `az` capture commands run in the interactive MAIN window, invoked from this
>    phase's `_preconditions` prose (NOT inside the dispatched worker). They write
>    raw JSON to `$MIGRATION_DIR/live-capture/` and a `manifest.json`. If `az` is
>    missing or the user declines, Part A writes nothing and records the decline.
> 2. **Part B — dispatched parse fragment (Steps 3–7, this file's `_fragment`).**
>    File-only, non-interactive: reads the `live-capture/` directory Part A wrote and
>    maps it into `azure-resource-inventory.json` / `azure-resource-clusters.json`.
>    Runs no `az` command and never prompts. If `live-capture/` is absent or empty
>    (Part A skipped/declined), it contributes nothing and exits cleanly.
>
> This mirrors the RDfA split (RDfA needs no consent step — reading an archive the
> customer handed over is not interactive — but the parse-half is the same shape).
>
> **Status.** Net-new. Fills the live-`az` capability the SKILL.md description
> advertises and that `discover.md` already reserves a step-2 slot for ("the live
> `az` path — security contract, capture pre-work, parsing fragment, in that
> order"). Until it lands, a live-`az` ask has no producer.
>
> **Live-validated (az-cli 2.90.0, real subscription):** the preflight (`az version`,
> `az account show`), the fast-path (`az resource list`), and the enrichment rows for
> Container Apps (4a/4b), PostgreSQL Flexible Server (7), Storage (11), Container
> Registry (11a), Log Analytics (11b), and Cognitive Services accounts + deployments
> (16) were all run read-only and returned the projected fields intact. `az resource
> list`, `az acr`, `az monitor log-analytics`, and `az containerapp` all worked with
> NO extension installed. Rows for services not present in the test subscription
> (webapp, aks, vm, mysql, cosmos, redis, service bus, event hubs, vnet, key vault,
> ML workspace, dns) carry command shapes ported from the pattern and should get one
> live pass before merge.

Produces the SAME artifacts as `discover-iac.md`, keyed on canonical `Microsoft.*`
ARM types, so every downstream phase works identically. When IaC discovery also ran,
Part B merges live findings into the existing inventory and surfaces drift.

**Execute ALL steps in order. Do not skip or optimize.**

---

## Why the parse half needs no type-translation table

The `az` CLI emits canonical `Microsoft.*` ARM type strings natively (`az resource
list` returns `"type": "Microsoft.Web/sites"`), which is the SAME key
`discover-iac.md` canonicalizes to via `arm-type-canonicalization.md`. So the live
path is a second *producer* of the identical inventory contract — not a new mapping
surface. There is no "asset type → Terraform type" translation table (as the
gcp-to-aws live path needs) because the live output already speaks the inventory's
own type language.

Live discovery is offered when the workspace has no `azurerm_*`/Bicep/ARM IaC — the
common startup case — or as an accuracy upgrade alongside IaC. Part A's consent gate
is what "offers" it.

---

## Part A — Main-window pre-work (consent + capture)

> Runs in the interactive main window from the phase `_preconditions`, BEFORE the
> dispatched worker starts. NOT part of the `_fragment` body below.

---

### Security Contract (applies to every step)

1. **Exact-command allowlist.** Run ONLY commands that appear in Step 0 (preflight)
   or the Step 2 Capture Command Table. Never a mutating verb (`create`, `update`,
   `delete`, `set`, `add`, `remove`, `deploy`, `start`, `stop`, `restart`, `import`),
   never `az login` (interactive — hand off to the user), never
   `az account get-access-token` (prints a bearer token), never any
   `... show`/`list` variant that returns secret material (see rule 2).
2. **Never capture secret values.** Every capture command uses an explicit
   `--query` (JMESPath) projection that selects NAMES/metadata, never values.
   Specifically FORBIDDEN commands (they return secret material):
   - `az webapp config appsettings list` / `az functionapp config appsettings list`
     return app-setting **values** — use the projection in the table (names only) or
     skip; never write raw appsettings to a capture file.
   - `az webapp config connection-string list` — returns connection strings; skip.
   - `az keyvault secret show` / `... secret list --query "[].value"` — secret
     values; capture Key Vault **names** only (`az keyvault list`).
   - `az vm show ... osProfile.customData` — cloud-init payload; never project it.
   Additionally, apply a sensitive-key redaction pass (`password`, `secret`,
   `api_key`, `access_key`, `private_key`, `client_secret`, `connectionstring`,
   `token`, `credential`, `auth` — case-insensitive) to any config field before it
   is written into an artifact: replace matched values with `"[REDACTED]"`.
3. **Always explicit scope.** Every command passes `--subscription "$AZURE_SUBSCRIPTION"`
   explicitly. Never rely on the active `az` default subscription inside capture
   commands. One subscription per run — for multiple subscriptions, run the
   migration once per subscription.
4. **Capture to files, not context.** Redirect stdout to files under
   `$MIGRATION_DIR/live-capture/`. Process any capture file larger than ~100
   resources with a throwaway extraction script — do NOT Read large raw captures
   into context (see the Scale guard in Step 2).
5. **Consent first.** No `az` command from the Step 2 table runs before the user
   answers `[A]` in Step 1. Preflight commands in Step 0 are limited to
   version/account checks that touch no resource data.

---

### Step 0: Preflight

1. **CLI installed:** run `az version --output json` (read the `azure-cli` field).
   - Missing → tell the user: "The Azure CLI (`az`) isn't installed. Install it
     (<https://learn.microsoft.com/cli/azure/install-azure-cli>) and tell me to
     continue, or skip live discovery." Wait. If skipped → exit cleanly.
2. **Authenticated + subscription:** run `az account show --output json`
   (a local token/context read — no resource data).
   - Error / not logged in → tell the user: "Your Azure CLI isn't logged in. Run
     `az login` in your terminal — it needs a browser, so I can't run it for you —
     then tell me to continue." Wait. If declined → exit cleanly.
   - Success → show `name` + `id`, then ask: "Discover subscription
     `[name] ([id])`? [Y] Yes / [N] Use a different subscription (type its id or
     name)." Set `$AZURE_SUBSCRIPTION` to the chosen subscription **id**. If the
     user names a different one, resolve it with
     `az account show --subscription "<typed>" --output json` and use its `id`.

### Step 1: Consent Gate

Output exactly, then wait for the user's choice:

```
─── Live Azure Discovery (read-only) ───

I can inventory subscription [$AZURE_SUBSCRIPTION] directly using your
authenticated az CLI. This runs LIST/SHOW commands only:

  ✓ Captured: resource names, ARM types, regions, SKU/tier/
    capacity, container images, network topology, app-setting
    NAMES, Key Vault NAMES, and tags.
  ✗ Never captured: app-setting values, connection strings,
    Key Vault secret values, database contents, VM customData,
    or access tokens. No command that creates, changes, or
    deletes anything will run.

Output is written to .migration/<run>/live-capture/ (gitignored).

[A] Proceed with live discovery
[B] Skip — use workspace files only
```

- **[A]** → continue to Step 2.
- **[B]** → exit cleanly with no output (record the decline for the orchestrator).

### Step 2: Capture

Create `$MIGRATION_DIR/live-capture/`.

**2a. Fast path — subscription-wide inventory (`az resource list`, one call, no
extension):**

```
az resource list --subscription "$AZURE_SUBSCRIPTION" \
  --query "[].{name:name, type:type, location:location, resourceGroup:resourceGroup, sku:sku.name, kind:kind}" \
  --output json > $MIGRATION_DIR/live-capture/resources.json
```

`az resource list` returns every resource's canonical ARM `type`, name, location,
resourceGroup, and sku/kind in one call, is **built in (no extension)**, and needs
only the **Reader** role. This is the fast-path of choice — validated live against a
real subscription: it returned `Microsoft.*` types directly (e.g.
`Microsoft.App/containerApps`, `Microsoft.DBforPostgreSQL/flexibleServers`,
`Microsoft.CognitiveServices/accounts`) that Step 3 consumes without translation.

> **Do NOT use `az graph query` as the default fast-path.** Azure Resource Graph
> lives in the `resource-graph` CLI extension. On a fresh `az` install the command
> triggers an **interactive dynamic-install prompt** (confirmed live: it hangs
> waiting for input, it does not cleanly error) — fatal inside the non-interactive
> dispatched worker, and awkward even in the main window. `az resource list` covers
> the same fast-path need with no extension and no prompt. Only if a future need
> requires Resource Graph's KQL power, offer `az extension add --name resource-graph`
> explicitly in the main window (a local CLI install, not an Azure mutation) — never
> let dynamic-install prompt.

- Success → record `method: "resource_list"` in the manifest. `az resource list`
  gives types/names/locations/sku but thin per-service config, so then run only the
  **enrichment rows** (marked E) of the table below for the ARM types that were
  found — those add the config fields (app-setting names, container images, network
  wiring, versions) that edge inference and sizing need.
- Failure → classify and branch:
  - **Permission denied** (the identity lacks Reader on the subscription) → tell the
    user briefly that live discovery needs the **Reader** role on subscription
    `$AZURE_SUBSCRIPTION` (<https://learn.microsoft.com/azure/role-based-access-control/built-in-roles#reader>),
    then per-service fallthrough (the per-service `list` calls hit the same wall and
    record `failed`, which is honest).
  - **Other errors** → per-service fallthrough immediately.

  **Per-service fallthrough:** record `method: "per_service"` and run every applicable
  table row. Keep the failed `az resource list` entry in `captures[]` with
  `status: "failed"` and the stderr summary in `note`.

**2b. Capture Command Table.** Each row redirects to the named file, always with
`--subscription "$AZURE_SUBSCRIPTION"`. On permission/"not found" errors: record the
row as `failed`/`skipped` in the manifest and continue — a missing service is normal,
never a halt. Every row uses `--query` to project NAMES/metadata only (Security
Contract rule 2).

| #  | Command (always `--subscription "$AZURE_SUBSCRIPTION" --output json`)                                                                                                                                                                                          | Output file        | Mode | Canonical ARM type                                          |
| -- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------ | ---- | ----------------------------------------------------------- |
| 1  | `az webapp list --query "[].{name:name, location:location, kind:kind, sku:appServicePlanId, https:httpsOnly, appSettingNames:siteConfig.appSettings[].name, linuxFxVersion:siteConfig.linuxFxVersion, vnet:virtualNetworkSubnetId}"`                            | `webapp.json`      | E    | `Microsoft.Web/sites`                                       |
| 2  | `az appservice plan list --query "[].{name:name, location:location, sku:sku.name, tier:sku.tier, capacity:sku.capacity, reserved:reserved}"`                                                                                                                    | `plans.json`       | E    | `Microsoft.Web/serverfarms`                                 |
| 3  | `az functionapp list --query "[].{name:name, location:location, kind:kind, runtime:siteConfig.linuxFxVersion, appSettingNames:siteConfig.appSettings[].name, plan:appServicePlanId}"`                                                                          | `functionapp.json` | E    | `Microsoft.Web/sites` (kind `functionapp`)                  |
| 4  | `az aks list --query "[].{name:name, location:location, k8sVersion:kubernetesVersion, nodePools:agentPoolProfiles[].{name:name, vmSize:vmSize, count:count, mode:mode}, network:networkProfile.networkPlugin, vnetSubnet:agentPoolProfiles[0].vnetSubnetId}"`    | `aks.json`         | E    | `Microsoft.ContainerService/managedClusters`               |
| 4a | `az containerapp list --query "[].{name:name, location:location, env:properties.managedEnvironmentId, image:properties.template.containers[].image, cpu:properties.template.containers[].resources.cpu, memory:properties.template.containers[].resources.memory, minReplicas:properties.template.scale.minReplicas, maxReplicas:properties.template.scale.maxReplicas, ingress:properties.configuration.ingress.external}"` | `containerapp.json` | E   | `Microsoft.App/containerApps`                              |
| 4b | `az containerapp env list --query "[].{name:name, location:location}"` — the managed environment each Container App runs in (edge target for row 4a's `env`)                                                                                                     | `containerappenv.json` | E | `Microsoft.App/managedEnvironments`                        |
| 5  | `az vm list --query "[].{name:name, location:location, size:hardwareProfile.vmSize, os:storageProfile.osDisk.osType, image:storageProfile.imageReference, subnet:networkProfile.networkInterfaces[0].id, tags:tags}"`                                           | `vm.json`          | E    | `Microsoft.Compute/virtualMachines`                         |
| 6  | `az sql server list --query "[].{name:name, location:location, version:version}"` then `az sql db list --server <each> --query "[].{name:name, sku:currentSku.name, tier:currentSku.tier, capacity:currentSku.capacity, maxSizeBytes:maxSizeBytes}"`            | `sql.json`         | E    | `Microsoft.Sql/servers`, `Microsoft.Sql/servers/databases` |
| 7  | `az postgres flexible-server list --query "[].{name:name, location:location, sku:sku.name, tier:sku.tier, version:version, storageGb:storage.storageSizeGb, haMode:highAvailability.mode}"`                                                                      | `postgres.json`    | E    | `Microsoft.DBforPostgreSQL/flexibleServers`                 |
| 8  | `az mysql flexible-server list --query "[].{name:name, location:location, sku:sku.name, tier:sku.tier, version:version, storageGb:storage.storageSizeGb}"`                                                                                                       | `mysql.json`       | E    | `Microsoft.DBforMySQL/flexibleServers`                      |
| 9  | `az cosmosdb list --query "[].{name:name, location:location, kind:kind, capabilities:capabilities[].name, apiKind:apiProperties.serverVersion, multiRegion:enableMultipleWriteLocations, locations:locations[].locationName}"`                                   | `cosmos.json`      | E    | `Microsoft.DocumentDB/databaseAccounts`                     |
| 10 | `az redis list --query "[].{name:name, location:location, sku:sku.name, family:sku.family, capacity:sku.capacity, version:redisVersion, subnet:subnetId}"`                                                                                                       | `redis.json`       | E    | `Microsoft.Cache/redis`                                     |
| 11 | `az storage account list --query "[].{name:name, location:location, sku:sku.name, kind:kind, tier:accessTier, https:enableHttpsTrafficOnly}"`                                                                                                                    | `storage.json`     | E    | `Microsoft.Storage/storageAccounts`                         |
| 11a | `az acr list --query "[].{name:name, location:location, sku:sku.name, adminEnabled:adminUserEnabled}"` — container registry (→ Amazon ECR)                                                                                                                      | `acr.json`         | E    | `Microsoft.ContainerRegistry/registries`                    |
| 11b | `az monitor log-analytics workspace list --query "[].{name:name, location:location, sku:sku.name, retentionDays:retentionInDays}"` — Log Analytics (→ CloudWatch Logs)                                                                                          | `loganalytics.json` |     | `Microsoft.OperationalInsights/workspaces`                  |
| 12 | `az servicebus namespace list --query "[].{name:name, location:location, sku:sku.name, tier:sku.tier}"`                                                                                                                                                          | `servicebus.json`  |      | `Microsoft.ServiceBus/namespaces`                           |
| 13 | `az eventhubs namespace list --query "[].{name:name, location:location, sku:sku.name, capacity:sku.capacity, kafka:kafkaEnabled}"`                                                                                                                               | `eventhubs.json`   |      | `Microsoft.EventHub/namespaces`                             |
| 14 | `az network vnet list --query "[].{name:name, location:location, addressSpace:addressSpace.addressPrefixes, subnets:subnets[].{name:name, prefix:addressPrefix}}"`                                                                                               | `vnet.json`        | E    | `Microsoft.Network/virtualNetworks`                         |
| 15 | `az keyvault list --query "[].{name:name, location:location}"` — vault NAMES only, never `secret show`/`secret list --query "[].value"`                                                                                                                          | `keyvault.json`    | E    | `Microsoft.KeyVault/vaults`                                 |
| 16 | `az cognitiveservices account list --query "[].{name:name, location:location, kind:kind, sku:sku.name}"` then per account `az cognitiveservices account deployment list -n <name> -g <rg> --query "[].{name:name, model:properties.model.name, version:properties.model.version}"` | `cognitive.json`   | E    | `Microsoft.CognitiveServices/accounts`, `.../deployments`   |
| 17 | `az ml workspace list --query "[].{name:name, location:location}"` — only if the `ml` extension is present; else record `skipped`                                                                                                                                | `mlworkspace.json` | E    | `Microsoft.MachineLearningServices/workspaces`              |
| 18 | `az network private-dns zone list --query "[].{name:name}"` and `az network dns zone list --query "[].{name:name, records:numberOfRecordSets}"`                                                                                                                  | `dns.json`         |      | `Microsoft.Network/dnszones`, `privateDnsZones`             |

**Row 6 note (two-step):** SQL is server-then-database. List servers, then list
databases per server; skip the `master` system database. Elastic pools
(`az sql elastic-pool list`) are an enrichment row only when a server is found.

**Row 16 note (two-step):** list Cognitive Services accounts, then list deployments
per account — the deployment's `model.name`/`version` is the AI signal Design's
lifecycle check consumes. Never capture keys (`az cognitiveservices account keys
list` is FORBIDDEN).

**Sizing caveat:** SKU capacity / `storageGb` are PROVISIONED, not actual usage.
Downstream sizing must treat them as an upper bound. (Follow-up: actual utilization
and spend come from the billing-export / RDfA path, not live `az`.)

**Scale guard:** if `graph.json` (or any capture) exceeds ~100 resources, write a
throwaway extraction script to `$MIGRATION_DIR/_extract_live.py` that projects only
the fields Step 3 needs, run it, write its JSON output next to the raw file with a
`-extracted.json` suffix, and delete the script. Never Read the oversized raw file
directly.

**2c. Write the manifest** — `$MIGRATION_DIR/live-capture/manifest.json`:

```json
{
  "captured_at": "<ISO 8601 UTC>",
  "az_version": "<azure-cli from az version>",
  "account": "<user/servicePrincipal from az account show>",
  "subscription": "<$AZURE_SUBSCRIPTION>",
  "method": "resource_list|per_service",
  "captures": [
    { "command": "<row command>", "file": "<file>", "status": "ok|failed|skipped", "note": null }
  ]
}
```

Every attempted or deliberately skipped row gets an entry.

---

## Part B — Dispatched parse fragment (`_fragment: live-parse`)

> File-only, non-interactive. Runs `az` NEVER; only reads `$MIGRATION_DIR/live-capture/`.
> **Entry guard:** if `$MIGRATION_DIR/live-capture/manifest.json` is absent (Part A
> was skipped or the user declined), contribute nothing and exit cleanly — this is a
> normal outcome, not a failure.

### Step 3: Map Captures to Inventory Resources

The captured `type` is ALREADY a canonical `Microsoft.*` ARM string, so there is no
type-translation table — this is the key simplification over the gcp-to-aws live path.
For each captured resource, synthesize an inventory entry matching
`references/shared/schema-discover-iac.md`:

- `azure_type` = the captured ARM `type` (case-folded per
  `arm-type-canonicalization.md`; kind-qualify where the table notes it — a
  `Microsoft.Web/sites` with `kind` containing `functionapp` is a Function App).
- `azure_id` = the resource's ARM resource id (`/subscriptions/.../resourceGroups/.../providers/...`).
  One resource → one id string (exact-match join key downstream).
- `name` = the resource name.
- `resource_group` = the captured `resourceGroup`.
- `config` = the projected fields from the capture (redaction rules from the Security
  Contract apply; keep `sku`/`tier`/`capacity` — Design's cost-bearing test reads them).
- `source` = `"live"` on every entry.
- `azure_type_provenance` = `"live"`.

A captured `type` absent from `fast-path-services.json` is **derived, not skipped** —
follow `discover-iac.md` Step 3's derive-and-record rule (`azure_type_provenance:
"derived"` / `"derived_uncorroborated"`), keep the resource, and record it in
`live_metadata.derived_types`. Never silently drop a resource.

**Classification & clustering:** apply `discover-iac.md`'s PRIMARY/SECONDARY rules and
`{category}_{type}_{region}_{sequence}` cluster naming (see Step 5), `confidence: 0.99`.

**AI detection:** if any `Microsoft.CognitiveServices/accounts`,
`.../accounts/deployments`, or `Microsoft.MachineLearningServices/workspaces` resource
was captured, contribute the minimal `ai-workload-profile.json` exactly as
`discover-iac.md` Step 4.5 would (signal method `"live_az"`, `profile_source:
"iac_cognitive"`), so the AI track fires. `Microsoft.Search/searchServices` alone is
NOT a strong signal.

> **`ai_source` keys off the DEPLOYMENT MODEL, not the account `kind`** (validated
> live). Modern accounts report `kind: "AIServices"` — an umbrella that can host
> mixed families: a real subscription returned deployments `gpt-4.1-mini` +
> `text-embedding-3-small` (OpenAI family → `ai_source: "azure_openai"`) alongside
> `Kimi-K2.6` (a partner model that is NOT Azure OpenAI). Set `ai_source:
> "azure_openai"` when any deployment's `model.name` is an OpenAI-family id
> (`gpt-*`, `text-embedding-*`, `o1*`/`o3*`, `dall-e*`); classify non-OpenAI
> deployments as `ai_source: "other"` and list them so Design can flag that a
> non-OpenAI Foundry model has no direct Bedrock equivalent. Do NOT infer the source
> from `kind: AIServices` alone.

### Step 4: Infer Edges from Resolved Config

Live captures carry resolved ids, which beat IaC references. Build `edges[]` using
ONLY these deterministic rules (evidence = the config field path):

| Config field (captured)                                     | Edge                                                                                                  |
| ----------------------------------------------------------- | ----------------------------------------------------------------------------------------------------- |
| webapp/functionapp `virtualNetworkSubnetId` (row 1/3)       | site → vnet, `network_membership` (resolve subnet id → its parent vnet id)                            |
| webapp/functionapp `sku`/plan id → plan (row 1/3 → row 2)   | site → App Service Plan, `hosted_on` (the plan-cost fan-in edge; see `schema-discover-azure.md`)      |
| AKS `agentPoolProfiles[].vnetSubnetId` (row 4)              | cluster → vnet, `network_membership`                                                                  |
| Container App `managedEnvironmentId` (row 4a → 4b)          | container app → managed environment, `hosted_on` (the env is the shared compute host — parity with App Service Plan fan-in) |
| Container App `image` → registry login server (row 4a → 11a) | container app → registry, `data_dependency` when the image host matches a captured `acr` login server (`<name>.azurecr.io`) |
| VM `networkProfile.networkInterfaces[0].id` (row 5)         | vm → vnet, `network_membership` (resolve NIC → subnet → vnet only if the NIC/subnet was also captured; else record the NIC id in `config` and emit NO edge) |
| redis `subnetId` (row 10)                                   | redis → vnet, `network_membership`                                                                    |
| cosmos `locations[]` length > 1 or `enableMultipleWriteLocations` | annotate cosmos `config.multi_region: true` (drives DynamoDB Global Tables downstream) — not an edge |

No other inference — do not guess relationships from names, tags, or app-setting
names.

### Step 5: Cluster (Simplified Mode)

Apply `discover-iac.md`'s clustering rules regardless of resource count (networking
cluster at depth 0; one cluster per PRIMARY plus its `serves` secondaries at depth 1;
same `{category}_{type}_{region}_{sequence}` naming; region from the captured
`location`). Set metadata `"clustering_mode": "simplified_live"`. If more than 25
PRIMARY resources were captured, warn that clustering is coarse at this scale and
suggest narrowing to specific resource groups or regions — but continue.

**Live-specific clustering rules:**

- **Regionless / global resources** (some DNS zones, global storage): use `"global"`
  as the region component of `cluster_id`.
- **Shared secondaries** (e.g., one vnet serving multiple primaries): assign to the
  cluster of the FIRST primary in its `serves[]`; `serves[]` still lists all.
- **Evidence-less secondaries** (no Step 4 edge and empty `serves[]`, e.g., Key
  Vaults): group into their own cluster per category+region (e.g.,
  `security_keyvault_eastus_001`) at depth 1 — never attach to an unrelated primary.

### Step 6: Merge with IaC Discovery (only if `discover-iac.md` produced output)

If `azure-resource-inventory.json` does NOT already exist, skip to Step 7 (live is the
sole source).

Otherwise the IaC inventory + clusters are the BASE. Match live↔IaC entries by
**`azure_id`** when both carry one (exact match — the reliable join), falling back to
`azure_type` + `name` + `resource_group`. Then:

1. **Matched:** keep the IaC entry (its address/id, classification, cluster, depth).
   Overwrite `config` values where live disagrees — sizing, SKU, capacity, versions,
   images (live reflects reality). Record every OVERWRITTEN field in
   `live_metadata.drift.config_conflicts[]` as
   `{ "azure_id", "field", "iac_value", "live_value" }`. Set `source: "live+terraform"`.
2. **Live-only:** append with `unmanaged_by_iac: true`. Attach to an existing cluster
   of the same category+region when one exists; else append a new simplified cluster.
3. **IaC-only:** set `source: "terraform"` (or the IaC dialect) on every unmatched IaC
   entry. Set `not_found_live: true` ONLY if the capture covering that resource's
   service succeeded (manifest `ok`). If the relevant capture failed/was skipped, leave
   it untouched — absence of evidence is not drift.
4. **Drift summary:** `live_metadata.drift = { "resources_live_only": N,
   "resources_iac_only": M, "config_conflicts": [...] }`. `resources_iac_only` counts
   ONLY entries with `not_found_live: true`, never capture-failed unknowns.
5. **Merged metadata:** set `metadata.discovery_sources` to include both sources.

Never silently resolve a disagreement — every conflict lands in the drift record.

### Step 7: Write Output Files

Load `references/shared/schema-discover-iac.md` (if not already loaded) and
write/update:

1. `$MIGRATION_DIR/azure-resource-inventory.json` — exact schema; plus:
   - `metadata.discovery_sources`: `["live"]`, `["terraform", "live"]`, etc.
   - `metadata.discovery_timestamp`, `metadata.subscriptions_discovered: ["$AZURE_SUBSCRIPTION"]`
   - `metadata.clustering_mode`: `"simplified_live"` (live-only runs)
   - top-level `live_metadata`:

   ```json
   {
     "found": true,
     "captured_at": "<from manifest>",
     "subscription": "<$AZURE_SUBSCRIPTION>",
     "method": "resource_list|per_service",
     "capture_warnings": ["<failed/skipped manifest entries>"],
     "derived_types": { "<arm type>": 1 },
     "drift": { "resources_live_only": 0, "resources_iac_only": 0, "config_conflicts": [] }
   }
   ```

   (`drift` present only when Step 6 merged.)

2. `$MIGRATION_DIR/azure-resource-clusters.json` — exact schema (merged or fresh).
3. Validate per `discover-iac.md`'s output rules (every resource in exactly one
   cluster, ids consistent, valid JSON). Report: "Live discovery: X resources captured
   from subscription [name] (Y unmanaged by IaC, Z config conflicts)."

The parent `discover.md` owns the phase status update — do not touch
`.phase-status.json` here.

---

### Error Handling

| Error                                              | Part | Behavior                                                                                                                     |
| -------------------------------------------------- | ---- | --------------------------------------------------------------------------------------------------------------------------- |
| `az` missing / not logged in / user declines       | A    | Write no `live-capture/`; record the decline for the orchestrator. Part B then no-ops on the missing manifest.              |
| `az resource list` fails (permission denied)       | A    | State the **Reader** role requirement + docs link; fall to per-service (agent never grants roles)                           |
| `az resource list` fails (other)                   | A    | Fall to per-service immediately                                                                                             |
| `az graph` prompts to install an extension         | A    | Do NOT use `az graph` as the fast-path (it hangs on the interactive dynamic-install prompt); `az resource list` is the fast-path |
| Individual row fails (permission, not found)       | A    | Record `failed`/`skipped` in the manifest, continue — never a halt                                                         |
| Token expired mid-capture                          | A    | Stop capturing; hand off ("run `az login`, then tell me to continue"); on resume re-run Part A Step 2 (captures overwrite)  |
| `live-capture/manifest.json` absent                | B    | Contribute nothing, exit cleanly (Part A skipped/declined — normal)                                                        |
| Capture file unparseable                           | B    | Record warning in `live_metadata.capture_warnings`, skip that file, continue                                               |
| Every capture failed (manifest all `failed`)       | B    | Contribute nothing; the assembler surfaces that live discovery yielded no resources and names the **Reader** role gap      |

**Key principle:** partial results are better than no results. Record what failed;
never fabricate what wasn't captured.

### Scope Boundary

**This fragment covers live Azure discovery ONLY.**

FORBIDDEN — Do NOT include ANY of:

- AWS service names, recommendations, or equivalents
- Migration strategies, phases, timelines, cost estimates, or effort estimates
- Any mutating `az` command, `az login`, or token printing
- App-setting values, connection strings, Key Vault secret values, VM customData, or
  unredacted sensitive config anywhere

**Your ONLY job: inventory what exists in Azure. Nothing else.**

…l ai-model-lifecycle.md (PR awslabs#304 merge blocker)

herosjourney's blocking checklist (PR awslabs#304): awslabs#307 is being folded into awslabs#304 and
closed as superseded. Because this PR relocates ai-model-lifecycle.md into the
canonical skills/shared/ai/ tree (deleting gcp-to-aws/references/shared/...), no
merge conflict fires — so if awslabs#304 merged with the stale canonical, shared:check
would keep the WRONG content ("Legacy minimum 6 months before EOL", no 45-day
policy) byte-identical across all 6 copies. This carries awslabs#307's fix in.

Two-way merge (herosjourney's ready-to-apply content, comment 5770573792):
kept awslabs#307's FACTS + this PR's canonical FRAMING.
- From awslabs#307: the 45-day / 2026-09-07 model-card policy split and the
  "don't apply the 6-month assumption post-2026-09-07" warning; both reference
  links; EOL'd rows (Claude 3 Haiku, Nova Premier v1, Nova Sonic v1) moved from
  the live table into Removed; restricted / Covered-Model handling; the
  GetFoundationModel unverified-fallback guidance; dates refreshed to Sep 21.
- From this PR: the canonical blockquote header (source-agnostic + edit-here-
  then-shared:sync) and the shared-location wording.

Applied to BOTH tree canonicals (advisor + migrate), then shared:sync propagated
to all vendored copies. All 6 ai-model-lifecycle.md files are byte-identical.

Verified: shared:check OK (both trees), drift:check OK (397 identical),
lint:model-ids OK, fmt:check clean, lint:md 0 errors. Only the 6 lifecycle files
touched. Stale "minimum 6 months before EOL" gone; 45-day policy present in every
copy; Haiku/Premier/Sonic in Removed.

awslabs#307 to be closed as superseded only after this lands.
@icarthick

Copy link
Copy Markdown
Collaborator Author

Blocking lifecycle item — done ✅ (31c1f179)

Applied your ready-to-apply merged canonical skills/shared/ai/ai-model-lifecycle.md and re-synced. The two-way merge is in: #307's facts (45-day / 2026-09-07 model-card split policy, the "don't apply the 6-month assumption post-2026-09-07" warning, both reference links, Claude 3 Haiku / Nova Premier v1 / Nova Sonic v1 moved into Removed, restricted/Covered-Model handling, the GetFoundationModel unverified-fallback guidance, Sep-21 dates) + this PR's canonical framing (blockquote header + shared-location wording).

Applied to both tree canonicals (advisor + migrate), then shared:sync propagated to all vendored copies.

Blocking checklist:

Verified on 31c1f179: shared:check OK · drift:check OK (397 identical) · lint:model-ids OK · fmt:check clean · lint:md 0 errors. Only the 6 lifecycle files touched.

One note on your judgment call: I kept your Sep-21 dated table verbatim, including the days_to_eol snapshot. The file recomputes at run time, so it's correct as-is; happy to bump the date/counts at merge if you'd prefer fresher numbers then.

Thanks for doing the merge legwork and pasting it ready-to-apply — made this a clean carry-in.

…ng-cache + phase-status conflicts)

main advanced 39 commits (incl. awslabs#307 Bedrock lifecycle refresh merged as 39ca3d4,
awslabs#291 heroku decision gate, awslabs#310/awslabs#314 report + plan-writer). Conflict resolution:

- ai-model-lifecycle.md (gcp vendored, both trees): took OURS — same awslabs#307 facts,
  but at this PR's correct vendored/canonical path with the canonical header
  (main edited the old gcp-to-aws/references/shared/ path this PR deletes).
- pricing-cache.md (gcp shared, both trees): hybrid — kept main's FRESHER lifecycle
  facts (Nova Canvas/Reel excluded; Nova Sonic past-EOL) but rewrote the path
  references from shared/ai-model-lifecycle.md -> vendored/ai/ai-model-lifecycle.md
  (this PR relocated the file; shared/ path would dangle).
- phase-status.schema.json (shared + heroku vendored, both trees): took THEIRS
  (main's newer run_mode description prose); run_mode key preserved.
- model-id-lint.py (both trees): updated the Haiku/Premier/Sonic allowlist paths
  from the deleted skills/gcp-to-aws/references/shared/ai-model-lifecycle.md to the
  canonical skills/shared/ai/ai-model-lifecycle.md (vendored copies inherit via
  canonicalize()) — main's linter didn't know this PR moved the file.

Post-merge shared:sync re-propagated the shared phase-status schema into the azure
vendored copy (canonical-staleness trap).

Verified green both trees: shared:check OK, drift:check OK (398 identical),
frontmatter 7 files, asserters 15 PASS, 92 validator unit tests pass, mise run build
all 16 tasks clean (lint:model-ids OK, fmt clean, lint:md 0 errors, all security
scanners clean). All six ai-model-lifecycle.md copies byte-identical.
… skill-list conflict)

main advanced another 13 commits (awslabs#313, awslabs#316, cask-1555 cost-routes, pricing/plan.json
fixes). Single conflict: advisor/README.md — main reworded the MCP-servers sentence
('live AWS data' / 'AWS documentation lookups') while this PR added azure-to-aws to the
skill list. Resolved as a union: kept azure-to-aws in the list AND main's improved
wording.

Post-merge shared:sync clean (0 updates). Verified: shared:check OK, drift OK
(398 identical), asserters 15 PASS, lint:model-ids OK, mise run build all 16 tasks clean.
…onsolidation conflicts)

main awslabs#318 consolidated pricing onto a cache-only model (removed the live pricing MCP),
touching pricing-mode.md, pricing-fallback.md, aws-infra-pricing.json, design-ai.md, and
the README across both trees — 13 conflicts. Resolution:

- pricing-mode.md / pricing-fallback.md (canonical + heroku vendored): took THEIRS
  (awslabs#318's cache-only rewrite; our PR does not touch these).
- aws-infra-pricing.json (_instances_note): UNION — kept our azure-specific note
  (x86 families as the azure default for Windows/.NET fleets, flexible-server-rds-sizing.json
  ref) and adopted awslabs#318's cache-only correction ('verify against the public AWS pricing
  pages' instead of 'via the AWS Pricing MCP').
- design-ai.md: took awslabs#318's 'If MCP call fails' phrasing (no retry count — cache-only)
  but kept our vendored path references/vendored/ai/ai-migration-guardrails.md.
- advisor/README.md: union — azure-to-aws in the skill list + awslabs#318's wording.

Post-merge shared:sync reconciled vendored copies (6/tree). Verified: shared:check OK,
drift OK (399 identical), frontmatter 7, asserters 15 PASS, 92 validator tests pass,
mise run build all 16 tasks clean.

@herosjourney herosjourney left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: add azure-to-aws migration skill

Reviewed at the design/behavioral level (393 files, +56k — not read line-by-line); findings below are verified against head 1df5cb095ed0. This is a strong foundation. Nothing here is architectural — the blockers are truthfulness of the SKILL.md triggers and the per-file Status blocks, plus one dangling reference.

Genuinely good (verified)

  • Cost/estimate honesty is excellent. Estimate asserts every design service appears in the breakdown exactly once — priced, or carrying an exclusion_reason with excluded_from_total (nothing silently dropped); totals are explicit floors; traditional-AI workloads (Document Intelligence→Textract, etc.) land in services_not_estimated[] rather than faked into token cost. The estimate golden fixture is deliberately "not clean" (5/16 unpriced) — it tests the states an implementation is most likely to fake tidy.
  • Pricing is cache-first and honest — split infra/AI caches with declared staleness windows and accuracy bands (±5-10% infra, ±15-25% AI), a pricing:coverage gate, and an explicit us-east-1 region warning.
  • AI track reuses the shared canonical skills/shared/ai/ files (vendored byte-identical) instead of duplicating, and its ai-model-lifecycle.md carries the corrected 45-day policy (the #307 clash is not reintroduced).
  • Drift, pricing-coverage, and 5 golden asserters (~3k lines of independent oracles) pass live. The clarify golden is a BLOCKED run — it pins the completion gate, not just the happy path.

Findings

[P1 — honesty] SKILL.md triggers over-promise the discovery surface. The description advertises "migrate Bicep to Terraform," "migrate ARM templates to Terraform," "discover from Terraform/Bicep/ARM templates," and "live az CLI capture" — but on this head only Terraform is built: extract-bicep.md/extract-arm.md are absent and a .bicep/ARM workspace halts (discover-iac.md §Status), and the live-az path is "lands in step 2" (discover.md §Status). The halt-don't-fake behavior is the right mitigation, but the trigger list matches user intents the skill can't fulfill — a Bicep user is matched, then stopped. Soften the triggers to Terraform-first, naming Bicep/ARM/live-az as not-yet. (A separate PR already adds live-az and softens exactly these triggers; once it lands this narrows to Bicep/ARM only.)

[P2 — consistency] Two stale Status blocks misrepresent completeness in opposite directions.

  • references/phases/generate/generate.md:123 calls itself "skeleton (build step 1)… emitters land in step 6," but the emitter fragment generate-artifacts-infra.md:236 says "implemented (build step: Generate)." The emitters are real — the orchestrator's Status under-reports.
  • references/phases/design/design.md:146 lists ai.md as "still to come," but references/design-refs/ai.md exists as a real ~8KB rubric — over-reports the gap.
    Refresh both; a reviewer trusting Status blocks misjudges completeness both ways.

[P2 — dangling reference] graviton.md is referenced but absent from the azure tree. SKILL.md:21 (and ~6 other spots: compute.md, database.md, clarify-compute.md, clarify-global.md, two knowledge JSONs) reference references/shared/graviton.md's escape path as firing "routinely," but that file is absent from both azure trees — it exists only in the gcp-to-aws sibling. The vendored README declares the skill must be self-contained (runnable lifted-out). These are prose mentions, not hard Load directives, so likely not a runtime halt — but it's a real broken path against the self-contained guarantee. Either vendor graviton.md into the azure tree or inline the escape criteria.

[P3 — disclosed, non-blocking] The run_mode: decide_and_execute Generate consent gate is an _assert that CI binds but never evaluates — honestly disclosed, but the opt-in-to-execute gate has no mechanical teeth and depends on the interpreter honoring it. Worth tracking, since Generate emits Terraform.

[P3 — deferred, confirmed] Cross-skill Azure routing isn't wired: llm-to-bedrock hardcodes gcp-to-aws for Assess and agent-advisor inline-executes gcp phase files via $GCP_BASE, neither with an Azure path. Named as deferred cross-skill work — fine to leave; flagging that a pure-Azure-AI user still touches gcp paths until it lands.

Net

The architecture is sound and the cost-honesty engineering is above the bar. Fix the SKILL.md trigger truthfulness (P1), the two Status blocks (P2), and the dangling graviton.md (P2), and the Terraform path is merge-ready — Bicep/ARM/live-az/AI-route can honestly land as sequenced follow-ups.

… + dangling graviton.md (PR awslabs#304 review)

herosjourney's merge-readiness review (verified vs head 1df5cb0):

P1 (honesty) — SKILL.md triggers over-promised the discovery surface. Only Terraform
is built; Bicep/ARM extractors and the live-az capture fragment are not on this branch
(a Bicep/ARM/live-only workspace halts). Fixes:
- dropped 'migrate Bicep to Terraform' / 'migrate ARM templates to Terraform' triggers;
- 7-phase summary now says discovery is from Terraform, naming Bicep/ARM/live-az as
  advertised triggers whose extractors are not yet on this branch (sequenced follow-ups);
- softened the 'Live-first discovery' philosophy bullet to '(planned)'.

P2 (consistency) — two stale Status blocks:
- generate.md said 'skeleton (build step 1)… emitters land in step 6' but the emitters
  are implemented — refreshed to 'implemented (build step: Generate)', describing the
  infra/AI-only/both tracks and the validated report path.
- design.md listed ai.md as 'still to come' but references/design-refs/ai.md exists
  (~8KB) — corrected to 'on disk'; only licensing.md/gpu-hpc.md/patterns.md remain.

P2 (dangling reference) — graviton.md (referenced from SKILL.md + ~6 spots as firing
'routinely') was absent from the azure tree, breaking the self-contained guarantee.
Vendored gcp's references/shared/graviton.md + its schema-graviton.md dependency into
azure's references/shared/ (both trees, byte-identical) — the azure-private placement
matching pricing-cache.md / report-decision-core.md.

P3 items (run_mode _assert-no-teeth; cross-skill llm-to-bedrock/agent-advisor gcp-
hardcoding) left as disclosed/deferred per the review.

Verified both trees: drift OK (401 identical), shared:check OK, frontmatter 7,
asserters 15 PASS, lint:model-ids OK, mise run build all 16 tasks clean.
@icarthick

Copy link
Copy Markdown
Collaborator Author

Review findings addressed — 9d57d393 (P1 + both P2s)

Thanks for the verified pass. All three merge-readiness findings fixed, both trees:

[P1 — SKILL.md trigger truthfulness] Made the discovery surface Terraform-first:

  • Dropped the migrate Bicep to Terraform / migrate ARM templates to Terraform triggers.
  • The 7-phase summary now says discovery is from Terraform, and names Bicep / ARM / live-az as advertised triggers whose extractors are not yet on this branch — a Bicep/ARM/live-only workspace halts rather than guessing; they land as sequenced follow-ups.
  • Softened the "Live-first discovery" philosophy bullet to (planned) with the same not-yet-on-this-branch note.

[P2 — stale Status blocks]

  • generate.md — was "skeleton (build step 1)… emitters land in step 6"; refreshed to implemented, describing the infra / AI-only / both tracks and the validated report path (matches generate-artifacts-infra.md's own Status).
  • design.md — dropped ai.md from "still to come"; it's on disk (~8KB rubric). Only licensing.md / gpu-hpc.md / patterns.md remain.

[P2 — dangling graviton.md] Vendored gcp's references/shared/graviton.md (and its schema-graviton.md dependency) into the azure tree's references/shared/ — the azure-private placement matching pricing-cache.md / report-decision-core.md, so the skill is self-contained again. No unit hard-Loads a still-missing file.

[P3] Left as you called them: the run_mode _assert-no-teeth is disclosed; the cross-skill llm-to-bedrock / agent-advisor gcp-hardcoding stays deferred.

Verified both trees: drift:check OK (401 identical), shared:check OK, frontmatter 7, asserters 15 PASS, lint:model-ids OK, mise run build all 16 tasks clean.

The Bicep/ARM extractors, the live-az fragment, and the cross-skill Azure route will land as the sequenced follow-ups you outlined.

@leon1418 leon1418 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[🤖 AI review 🤖]

Reviewed 9d57d393154aaa7c32a9987750c00299b3cfa5f0 against 6b48f8ad84fdd3928de41bf116e1e0f3a181cc6c with OCR delegation, an isolated general review, a separate repository review, and candidate verification. Requesting changes for the behavioral contracts below.

The two perspectives accounted for all 413 additions/modifications/deletions (397 GitHub file entries with rename compression). The current head contains changes in both plugin trees. The correction list is consolidated before publication; it includes the earlier fixes' remaining integration gaps.

Validation: 323 Agent Advisor tests and 152 report/policy tests pass per plugin; 15 fixture asserters pass per plugin; frontmatter, vendored/shared, drift, model-ID and pricing-vocabulary checks pass. All 25 source-contract/Heroku tests pass with Terraform 1.13 (the first run selected ambient 1.15.5). These validate source contracts and stored fixtures; I did not run a live Azure migration or deploy resources.

Already present: the IaC Cognitive AI producer, current lifecycle-policy content, refreshed Design/Generate Status blocks, and the missing Graviton references. The deferred Bicep/ARM/live discovery and cross-skill routing do not need to be implemented for this review.

Correction paths: plugin-relative paths below apply to both advisor/plugins/aws-startup-advisor/ and migrate/plugins/migration-to-aws/. Root documentation paths are stated separately. Change the deficient contract sites; preserve producers/consumers already correct.

F1 [P1] Make the declared AI-only route pass the backbone gates

An app-code-only Azure AI project completes Discover and the standalone AI-only Clarify flow. Discover intentionally leaves infrastructure artifacts absent, but Design requires inventory and clusters, Estimate requires infrastructure design and inventory, and Clarify still checks infrastructure-only fields. The documented AI-only chain cannot satisfy its own gates. Confirmed reentry also resets statuses while negative producers leave old optional files in place; artifact-presence routing can retain the previous track or fail the absence gate.

Correction: Carry track conditions through Clarify, Design, Estimate, their fragments/assemblers, the decision gate and Generate validation directory. Gate or adapt the workshop offer for AI-only runs. Preserve mandatory Clarify and execute consent; test fresh and resumed AI-only and mixed runs. On confirmed Discover reruns, retire obsolete phase-owned optional outputs and derived descendants based on current contributions, preserving a currently valid IaC AI producer.

Paths: skills/azure-to-aws/references/phases/clarify/clarify-ai-only.md, skills/azure-to-aws/references/phases/clarify/clarify-assemble.md, skills/azure-to-aws/references/phases/clarify/clarify.md, skills/azure-to-aws/references/phases/design/design-assemble.md, skills/azure-to-aws/references/phases/design/design-infra.md, skills/azure-to-aws/references/phases/design/design.md, skills/azure-to-aws/references/phases/discover/discover-app-code.md, skills/azure-to-aws/references/phases/discover/discover-assemble.md, skills/azure-to-aws/references/phases/discover/discover.md, skills/azure-to-aws/references/phases/estimate/estimate-assemble.md, skills/azure-to-aws/references/phases/estimate/estimate-infra.md, skills/azure-to-aws/references/phases/estimate/estimate.md, skills/azure-to-aws/references/phases/generate/generate.md, skills/azure-to-aws/references/phases/workshop/workshop.md, skills/azure-to-aws/references/shared/schema-discover-ai.md

Acceptance: A fresh AI-only project reaches a decision and, with execute consent, its report without fabricated infra artifacts. Mixed and infra runs retain their gates; decide-complete resume and workshop offer do not require absent files. AI-only validation runs against ai-migration rather than nonexistent terraform/. Confirmed mixed-to-infra-only and mixed-to-AI-only reruns leave the current route's artifact set; a valid current IaC AI contribution is retained.

F2 [P1] Do not prescribe unsupported Redis cluster encryption arguments

The new policy checks a standalone Redis aws_elasticache_cluster and suggests adding both encryption flags. at_rest_encryption_enabled is not in that Terraform resource schema. The new 'good' fixture passes POLICY_OK while carrying an unsupported argument; following the prescribed fix produces invalid Terraform.

Correction: Require an encrypted aws_elasticache_replication_group for Redis rather than inventing cluster arguments. Update checker, fix_hint, authoring posture, skill text, fixtures and tests in both trees; validate the accepted fixture against the supported provider schema.

Paths: skills/tf-best-practices/SKILL.md, skills/tf-best-practices/fixtures/terraform-policy/bad-elasticache-cluster-redis-unencrypted/cache.tf, skills/tf-best-practices/fixtures/terraform-policy/good-elasticache-cluster-redis-encrypted/cache.tf, skills/tf-best-practices/references/security-posture-rules.md, skills/tf-best-practices/scripts/test_validate_terraform_policy.py, skills/tf-best-practices/scripts/validate-terraform-policy.py

Acceptance: The accepted encrypted Redis fixture is valid for the supported AWS provider. The policy's repair guidance cannot recommend an unsupported cluster argument. Encrypted replication groups pass and unencrypted Redis is still rejected in both copies.

F3 [P2] Use one preference envelope across the Azure AI consumers

A mixed infrastructure/AI run confirms a European region and high AI token volume. Mixed Clarify writes global.target_region and top-level AI rows, while AI Design reads design_constraints.target_region with a us-east-1 fallback and Estimate reads ai_constraints.ai_token_volume.value. Recorded region/usage choices can be lost. The AI-only funding field is also nested where Generate reads a top-level field.

Correction: Normalize mixed and AI-only preferences to one documented shape or add explicit adapters. Trace region, all AI rows, agentic fields and startup program status through Design, Estimate, Generate and report; verify user choices survive.

Paths: skills/azure-to-aws/references/phases/clarify/clarify-ai-only.md, skills/azure-to-aws/references/phases/clarify/clarify-ai.md, skills/azure-to-aws/references/phases/clarify/clarify-assemble.md, skills/azure-to-aws/references/phases/design/design-ai.md, skills/azure-to-aws/references/phases/estimate/estimate-ai.md, skills/azure-to-aws/references/phases/generate/generate-artifacts-ai.md, skills/azure-to-aws/references/shared/schema-preferences.md

Acceptance: A mixed run's eu-west-1 region and high token volume are exactly what AI Design and Estimate consume. AI-only and mixed program-status and agentic preferences survive through the generated report/artifacts. Defaults apply only when the canonical field is absent, never because its producer wrote another shape.

F4 [P1] Run Generate dependencies and validation in their owning context

Generate is dispatched to generic-phase-worker-rw through its declared _exec tier. The worker has Read/Grep/Glob/Write/Edit only, but the infra fragment requires a skill invocation and the report fragment requires Python execution. The main-window finish step runs only Terraform validation. The report additionally consumes AI artifacts and an accounting ledger produced later in the declared order.

Correction: Load authoring posture before dispatch, complete emitters/accounting before report rendering, and run mandatory report and Terraform validators in the main window before postconditions. Preserve fail-closed behavior and both execution tracks.

Paths: skills/azure-to-aws/references/phases/generate/generate-artifacts-infra.md, skills/azure-to-aws/references/phases/generate/generate-artifacts-report.md, skills/azure-to-aws/references/phases/generate/generate-assemble.md, skills/azure-to-aws/references/phases/generate/generate.md

Acceptance: A strict rw worker finishes its assigned work without shell or skill calls. Authoring posture is supplied before emission; accounting/AI outputs exist before dependent report content. The main window runs both required validators and refuses completion on REPORT_FAIL or POLICY_FAIL.

F5 [P2] Admit the Terraform forms the extractor actually supports

The workspace contains only azapi_resource resources, JSON Terraform, or downloaded module resources referenced by root module blocks. Entry detection requires an azurerm resource and excludes the downloaded module directory before the extractor is loaded. A valid supported source can be treated as no IaC and stop without reaching the AzAPI/module extraction rules.

Correction: Align source admission and dialect detection with AzAPI, Terraform JSON and resolvable modules, while preserving state/secret exclusions and explicit halts for deferred Bicep/ARM sources. Add cold-start cases for each supported form.

Paths: skills/azure-to-aws/references/phases/discover/discover.md, skills/azure-to-aws/references/phases/discover/discover-iac.md, skills/azure-to-aws/references/shared/extract-terraform.md

Acceptance: AzAPI-only, JSON Terraform and downloaded-module-only supported inputs reach extraction. State files remain unread and Bicep/ARM remain explicitly deferred.

F6 [P2] Accept the provenance values emitted by the fallback routes

A named Azure resource uses an unlisted namespace, such as the documented Microsoft.Maps/accounts example. The producer emits derived_uncorroborated, but the Discover gate rejects it. Even after passing that gate, the model_category Design route is omitted from the Design schema checklist. The advertised non-halting fallback cannot complete honestly.

Correction: Synchronize producer, schemas, validation checklists and fixture vocabulary for derived_uncorroborated and model_category; retain warning and inferred-confidence requirements.

Paths: skills/azure-to-aws/references/phases/design/design-infra.md, skills/azure-to-aws/references/phases/discover/discover.md, skills/azure-to-aws/references/shared/extract-terraform.md, skills/azure-to-aws/references/shared/schema-design-aws.md, skills/azure-to-aws/references/shared/schema-discover-azure.md

Acceptance: The documented Maps/unlisted-namespace path can retain derived_uncorroborated and model_category without failing its own gates. Derived routes remain inferred and keep their provenance warnings.

F7 [P2] Preserve the routing fields consumed by the new rubrics

Terraform declares a Service Bus queue with requires_session=true, or uses separately declared AKS node pools or App Service container/runtime settings. The extraction allowlist has no Service Bus queue or child node-pool row and does not carry several explicit web runtime fields. Only generic SKU/tier/capacity survives for unlisted types. Downstream routing and sizing therefore lose session requirements, pool capacity and runtime evidence despite a structurally valid inventory.

Correction: Audit each currently supported rubric and sizing consumer against the extractor allowlist. Preserve their concrete routing/sizing fields and canonical shapes, including Service Bus entity settings, child node pools, web runtime settings and load-balancer configuration; test representative downstream decisions.

Paths: fixtures/azure-iac-terraform/check_expected_iac_terraform.py, fixtures/azure-iac-terraform/expected-iac-terraform.json, skills/azure-to-aws/references/design-refs/messaging.md, skills/azure-to-aws/references/shared/extract-terraform.md, skills/azure-to-aws/references/shared/schema-discover-azure.md

Acceptance: requires_session=true reaches the messaging rubric and produces the intended FIFO decision. A separately declared node pool retains VM size/count/spot/taints in its parent EKS design. All current rubric and sizing inputs have an extraction source or an explicit unavailable-data treatment.

F8 [P2] Honor an explicit App Service isolation split through Generate

The user selects separate environments per app in Clarify Q-C2. Design's phase manifest permits the split, but its schema checklist unconditionally prohibits site entries and Generate again requires exactly one compute target per plan. The requested isolation either fails validation or is folded away.

Correction: Make the no-split default conditional throughout schema, design, generation/accounting and estimation. With isolation_split=true, emit and price the intended separate environments; retain one plan target when false.

Paths: skills/azure-to-aws/references/phases/design/design-infra.md, skills/azure-to-aws/references/design-refs/index.md, skills/azure-to-aws/references/design-refs/compute.md, skills/azure-to-aws/references/shared/schema-design-aws.md, skills/azure-to-aws/references/phases/estimate/estimate-infra.md, skills/azure-to-aws/references/phases/generate/generate-artifacts-infra.md, skills/azure-to-aws/references/phases/generate/generate-assemble.md

Acceptance: An explicit two-app isolation split produces two corresponding environments and matching cost/accounting. The default unsplit plan still produces one compute target and one cost unit.

F9 [P2] Generate AI artifacts from each workload's source and target

An Azure-hosted application uses direct OpenAI or Anthropic, both explicitly supported discovery/design sources. Generate always emits an AzureOpenAI source branch, defaults AI_PROVIDER to azure_openai and runs comparison against Azure OpenAI. Such users lack that source endpoint/credential, so the default adapter and rollback do not call their existing provider. A traditional-AI target such as Textract is also ignored: the generator never reads target_aws_service and defaults to Bedrock scaffolding instead.

Correction: Derive source adapter, comparison client, credentials and rollback default from per-workload source provenance. Cover azure_openai, direct openai, anthropic and mixed sources without changing the working pre-cutover provider. Dispatch on each design block's target class; implement the selected non-Bedrock service or explicitly defer it in accounting and the report. Do not count unrelated Bedrock scaffolding as completion.

Paths: skills/azure-to-aws/references/phases/generate/generate-artifacts-ai.md, skills/azure-to-aws/references/phases/generate/generate-artifacts-report.md, skills/azure-to-aws/references/phases/generate/generate-assemble.md, skills/azure-to-aws/references/phases/generate/generate.md

Acceptance: Pre-cutover and rollback calls use the actually detected Azure OpenAI, OpenAI or Anthropic source. Mixed-provider workloads do not share an incorrect forced Azure client or credential. A Textract-only or mixed text/document design produces the selected implementation or an explicit visible deferral, with complete per-workload accounting.

F10 [P2] Ship the report validator with the supported skill-only install

The user installs the documented skill directories with npx skills add, including the sibling tf-best-practices skill. The mandatory report validator lives at plugin root and is absent from that install. The standalone skill therefore cannot satisfy the Generate report gate even though its own references are vendored.

Correction: Bundle/resolve the mandatory validator within the distributable skill dependency layout, or explicitly require and validate a complete plugin installation before execution. Test the documented installation layout through the report gate.

Paths: skills/azure-to-aws/references/phases/generate/generate-artifacts-report.md, skills/azure-to-aws/references/phases/generate/generate.md, skills/azure-to-aws/references/vendored/README.md, scripts/validate-migration-report.py

Acceptance: The documented install layout can resolve and execute the mandatory report validator. Missing dependencies produce an explicit preflight error; no skipped validator is reported as a pass.

F11 [P2] Finish the discovery-scope correction on consumer entry surfaces

A reader follows the README's Bicep, ARM, live az or billing-based Azure discovery promise. The README and AGENTS entry still promise those sources, and migrate/README.md describes a live discovery procedure as available. The implemented Discover phase still contains only Terraform and app-code producers and intentionally halts other paths.

Correction: Make all active Azure entry documentation Terraform/app-code first, and consistently label deferred sources. Update advisor/README.md, advisor/AGENTS.md, migrate/README.md and both Azure SKILL.md Defaults; do not implement the deferred capabilities merely to satisfy documentation.

Paths: skills/azure-to-aws/SKILL.md, advisor/README.md, advisor/AGENTS.md, migrate/README.md

Acceptance: Every active Azure source-discovery entrypoint agrees that Terraform/app code are implemented and the other sources are deferred.

F12 [P2] Use the current SQS message-size limit in the eliminator

A Service Bus workload needs messages larger than 256 KiB but no larger than 1 MiB. The hard eliminator incorrectly excludes SQS and forces a broker or S3 claim-check architecture for a payload SQS accepts directly.

Correction: Use the supported SQS limit and byte units in both rubric copies, preserving claim-check/broker handling above that limit; check the boundary without adding an unnecessary component.

Paths: skills/azure-to-aws/references/design-refs/messaging.md

Acceptance: A payload above 256 KiB and at or below 1 MiB does not eliminate SQS solely for size; larger payload handling remains explicit.

F13 [P2] Do not default New Zealand estates abroad because of a nonexistent-region claim

An estate in Azure newzealandnorth accepts the mapped default AWS region. The map sends it to Sydney and claims AWS has no New Zealand region, contrary to both the current region list and this table's same-country preference. This needlessly changes the country of the proposed deployment.

Correction: Use the available New Zealand region as the geographic default, subject to normal service availability and account opt-in checks; retain an explicit justified fallback where a needed service is unavailable.

Paths: skills/azure-to-aws/knowledge/design/azure-region-map.json

Acceptance: The New Zealand row no longer claims an absent AWS region and follows its same-country geographic default, with explicit availability/opt-in fallback.

F14 [P2] Construct complete and collision-safe ARM resource identities

Two invocations of the same local/downloaded module each declare azurerm_service_plan.this with name=var.name, different module inputs, and the same resource group. A top-level AzAPI resource also uses a resource-group parent. The prescribed unresolved name is tf:this in both instances. tf_address is also specified without the module prefix; only an auxiliary tf_module differs. The assembler deduplicates by the resulting identical azure_id, collapsing two real plans into one and understating the estate and cost. The AzAPI recipe separately omits /providers/ for that scope, producing an invalid ARM ID.

Correction: Use a stable fully qualified Terraform/module identity in synthetic unresolved ARM-name segments or as a separate collision-safe identity key. Include module instance addresses in tf_address, and resolve references through the same identity mapping without pretending unresolved names are real Azure names. Use the full ARM type when constructing top-level/extension IDs; only append remaining child segments under an already matching provider-qualified parent.

Paths: skills/azure-to-aws/references/phases/discover/discover-assemble.md, skills/azure-to-aws/references/shared/arm-type-canonicalization.md, skills/azure-to-aws/references/shared/extract-terraform.md, skills/azure-to-aws/references/shared/schema-discover-azure.md

Acceptance: Two instances of the same parameterized module retain two inventory and design entries. References into each module resolve to its own resource and stay stable across reruns. The documented budget example yields a provider-qualified ARM ID. Nested AzAPI children keep exactly one appropriate provider segment and join their parent.

F15 [P2] Apply the endpoint contract per GPT family rather than banning runtime

A GPT-5.6 workload needs the supported runtime/CRIS endpoint, for example Converse with Guardrails and permitted cross-region inference. The frozen fact reference explicitly supports GPT-5.6 runtime/CRIS and Converse Guardrails, but the new Azure design, schema and Generate gates classify every proprietary openai.gpt-* ID as Mantle-only. They require a model change or reject a supported same-model configuration, defeating the stated same-model selection policy.

Correction: Propagate the model/endpoint/feature matrix into Azure selection, schemas, pricing and generation. Permit GPT-5.6 runtime with the proper inference-profile IDs and IAM when its feature and residency constraints match; keep the actual Mantle-only restrictions for GPT-5.5/5.4.

Paths: skills/azure-to-aws/references/phases/design/design-ai.md, skills/azure-to-aws/references/phases/design/design.md, skills/azure-to-aws/references/phases/estimate/estimate-ai.md, skills/azure-to-aws/references/phases/generate/generate-artifacts-ai.md, skills/azure-to-aws/references/phases/generate/generate.md, skills/azure-to-aws/references/shared/pricing-cache.md, skills/azure-to-aws/references/shared/schema-design-aws-ai.md

Acceptance: A permitted GPT-5.6 Converse/Guardrails case keeps the same model and valid CRIS ID. GPT-5.5/5.4 runtime requests remain rejected, and Mantle-specific features still select Mantle. Generated credentials, IAM and endpoint agree with the chosen per-model path.

Provider/API facts were checked against HashiCorp's aws_elasticache_cluster schema (including AWS provider v5.80.0), AWS's current SQS message quotas and Region list, and the GPT-5.6 Terra/Sol model cards. The latter explicitly support runtime/CRIS and Converse; the blanket Mantle-only rule is not the current contract.

Approval remains for the human reviewer; this review does not approve or clear another review.

_on_failure: _halt_and_inform
- _check_single_active_phase: true
_on_failure: _halt_and_inform
- _check_file_exists: [azure-resource-inventory.json, azure-resource-clusters.json, preferences.json]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[🤖 AI review 🤖] F1 [P1] Make the declared AI-only route pass the backbone gates

An app-code-only Azure AI project completes Discover and the standalone AI-only Clarify flow. Discover intentionally leaves infrastructure artifacts absent, but Design requires inventory and clusters, Estimate requires infrastructure design and inventory, and Clarify still checks infrastructure-only fields. The documented AI-only chain cannot satisfy its own gates. Confirmed reentry also resets statuses while negative producers leave old optional files in place; artifact-presence routing can retain the previous track or fail the absence gate.

Carry track conditions through Clarify, Design, Estimate, their fragments/assemblers, the decision gate and Generate validation directory. Gate or adapt the workshop offer for AI-only runs. Preserve mandatory Clarify and execute consent; test fresh and resumed AI-only and mixed runs. On confirmed Discover reruns, retire obsolete phase-owned optional outputs and derived descendants based on current contributions, preserving a currently valid IaC AI producer.

The review summary lists the full affected paths and acceptance checks for both plugin copies.

continue # memcached / variable-driven / absent engine → exempt
if _has_own_attr(body, "replication_group_id"):
continue # node of a replication group — encryption enforced there
for attr in ("at_rest_encryption_enabled", "transit_encryption_enabled"):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[🤖 AI review 🤖] F2 [P1] Do not prescribe unsupported Redis cluster encryption arguments

The new policy checks a standalone Redis aws_elasticache_cluster and suggests adding both encryption flags. at_rest_encryption_enabled is not in that Terraform resource schema. The new 'good' fixture passes POLICY_OK while carrying an unsupported argument; following the prescribed fix produces invalid Terraform.

Require an encrypted aws_elasticache_replication_group for Redis rather than inventing cluster arguments. Update checker, fix_hint, authoring posture, skill text, fixtures and tests in both trees; validate the accepted fixture against the supported provider schema.

The review summary lists the full affected paths and acceptance checks for both plugin copies.

"ai_monthly_spend": "$500-$2K", // DETECTED | PROPOSED
"ai_priority": "balanced", // PROPOSED
"ai_critical_feature": null, // PROPOSED
"ai_token_volume": "low", // PROPOSED

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[🤖 AI review 🤖] F3 [P2] Use one preference envelope across the Azure AI consumers

A mixed infrastructure/AI run confirms a European region and high AI token volume. Mixed Clarify writes global.target_region and top-level AI rows, while AI Design reads design_constraints.target_region with a us-east-1 fallback and Estimate reads ai_constraints.ai_token_volume.value. Recorded region/usage choices can be lost. The AI-only funding field is also nested where Generate reads a top-level field.

Normalize mixed and AI-only preferences to one documented shape or add explicit adapters. Trace region, all AI rows, agentic fields and startup program status through Design, Estimate, Generate and report; verify user choices survive.

The review summary lists the full affected paths and acceptance checks for both plugin copies.


> **Plan-share links are GATED OFF** (landing page not live). Do NOT emit a share link.

## Step 5: Validate the rendered report (REQUIRED — mandatory gate)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[🤖 AI review 🤖] F4 [P1] Run Generate dependencies and validation in their owning context

Generate is dispatched to generic-phase-worker-rw through its declared _exec tier. The worker has Read/Grep/Glob/Write/Edit only, but the infra fragment requires a skill invocation and the report fragment requires Python execution. The main-window finish step runs only Terraform validation. The report additionally consumes AI artifacts and an accounting ledger produced later in the declared order.

Load authoring posture before dispatch, complete emitters/accounting before report rendering, and run mandatory report and Terraform validators in the main window before postconditions. Preserve fail-closed behavior and both execution tracks.

The review summary lists the full affected paths and acceptance checks for both plugin copies.


| Dialect | Detection |
| ----------- | -------------------------------------------------------------------------- |
| `terraform` | a `**/*.tf` or `**/*.tf.json` file containing a `resource "azurerm_` block |

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[🤖 AI review 🤖] F5 [P2] Admit the Terraform forms the extractor actually supports

The workspace contains only azapi_resource resources, JSON Terraform, or downloaded module resources referenced by root module blocks. Entry detection requires an azurerm resource and excludes the downloaded module directory before the extractor is loaded. A valid supported source can be treated as no IaC and stop without reaching the AzAPI/module extraction rules.

Align source admission and dialect detection with AzAPI, Terraform JSON and resolvable modules, while preserving state/secret exclusions and explicit halts for deferred Bicep/ARM sources. Add cold-start cases for each supported form.

The review summary lists the full affected paths and acceptance checks for both plugin copies.

Comment thread advisor/README.md
**Migrate to AWS**

- **`gcp-to-aws`** — Plan a migration from Google Cloud Platform — and OpenAI/Gemini AI workloads — to AWS, directly in your IDE. Runs a guided, multi-phase workflow: discover resources from Terraform/IaC, app code, and GCP billing exports; design an AWS architecture; estimate costs; and generate migration artifacts. AI-provider migration maps OpenAI/Gemini usage to closest-fit Amazon Bedrock model families.
- **`azure-to-aws`** — Plan a migration from Microsoft Azure — and Azure OpenAI / agentic AI workloads — to AWS. Runs a 7-phase workflow: discover resources from Terraform/Bicep/ARM, a consent-gated read-only `az` CLI capture, app code, and Azure billing exports; design an AWS architecture; estimate costs (1:1 lift and right-sized); optionally reprice what-if scenarios; and generate migration artifacts when the user opts in. App Service → Elastic Beanstalk, AKS → EKS, Cosmos DB → DynamoDB, Azure OpenAI → Bedrock.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[🤖 AI review 🤖] F11 [P2] Finish the discovery-scope correction on consumer entry surfaces

A reader follows the README's Bicep, ARM, live az or billing-based Azure discovery promise. The README and AGENTS entry still promise those sources, and migrate/README.md describes a live discovery procedure as available. The implemented Discover phase still contains only Terraform and app-code producers and intentionally halts other paths.

Make all active Azure entry documentation Terraform/app-code first, and consistently label deferred sources. Update advisor/README.md, advisor/AGENTS.md, migrate/README.md and both Azure SKILL.md Defaults; do not implement the deferred capabilities merely to satisfy documentation.

The review summary lists the full affected paths and acceptance checks for both plugin copies.


| Candidate | Eliminated when |
| -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **SQS** (standard or FIFO) | the entity's `max_message_size_in_kilobytes` exceeds **256 KB**. SQS caps at 256 KB; Service Bus Premium allows 100 MB. Route to Amazon MQ, or to SQS with an S3 claim-check — and say which |

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[🤖 AI review 🤖] F12 [P2] Use the current SQS message-size limit in the eliminator

A Service Bus workload needs messages larger than 256 KiB but no larger than 1 MiB. The hard eliminator incorrectly excludes SQS and forces a broker or S3 claim-check architecture for a payload SQS accepts directly.

Use the supported SQS limit and byte units in both rubric copies, preserving claim-check/broker handling above that limit; check the boundary without adding an unnecessary component.

The review summary lists the full affected paths and acceptance checks for both plugin copies.

"australiasoutheast": { "aws": "ap-southeast-2", "same_country": true, "geo": "Melbourne -> Sydney, Australia" },
"australiacentral": { "aws": "ap-southeast-2", "same_country": true, "geo": "Canberra -> Sydney, Australia" },
"indonesiacentral": { "aws": "ap-southeast-3", "same_country": true, "geo": "Jakarta, Indonesia" },
"newzealandnorth": {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[🤖 AI review 🤖] F13 [P2] Do not default New Zealand estates abroad because of a nonexistent-region claim

An estate in Azure newzealandnorth accepts the mapped default AWS region. The map sends it to Sydney and claims AWS has no New Zealand region, contrary to both the current region list and this table's same-country preference. This needlessly changes the country of the proposed deployment.

Use the available New Zealand region as the geographic default, subject to normal service availability and account opt-in checks; retain an explicit justified fallback where a needed service is unavailable.

The review summary lists the full affected paths and acceptance checks for both plugin copies.

needs to decide whether the resource costs money. Derivation is the design, not a
fallback — but the table is consulted FIRST, because the cases where a guess goes wrong
are enumerated there.
2. **Resolve `name`** from the block's `name` attribute. When it is an expression

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[🤖 AI review 🤖] F14 [P2] Construct complete and collision-safe ARM resource identities

Two invocations of the same local/downloaded module each declare azurerm_service_plan.this with name=var.name, different module inputs, and the same resource group. A top-level AzAPI resource also uses a resource-group parent. The prescribed unresolved name is tf:this in both instances. tf_address is also specified without the module prefix; only an auxiliary tf_module differs. The assembler deduplicates by the resulting identical azure_id, collapsing two real plans into one and understating the estate and cost. The AzAPI recipe separately omits /providers/ for that scope, producing an invalid ARM ID.

Use a stable fully qualified Terraform/module identity in synthetic unresolved ARM-name segments or as a separate collision-safe identity key. Include module instance addresses in tf_address, and resolve references through the same identity mapping without pretending unresolved names are real Azure names. Use the full ARM type when constructing top-level/extension IDs; only append remaining child segments under an already matching provider-qualified parent.

The review summary lists the full affected paths and acceptance checks for both plugin copies.

keeps the OpenAI SDK; only the base URL (`.../openai/v1/responses`), credential (a Bedrock API
key/token provider), model ID, and IAM (`bedrock-mantle:*`) change. Read
`references/shared/openai-on-bedrock.md` for exact values. If the source uses Chat Completions,
plan a reshape to `responses.create`. No Converse fallback exists for proprietary GPT models

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[🤖 AI review 🤖] F15 [P2] Apply the endpoint contract per GPT family rather than banning runtime

A GPT-5.6 workload needs the supported runtime/CRIS endpoint, for example Converse with Guardrails and permitted cross-region inference. The frozen fact reference explicitly supports GPT-5.6 runtime/CRIS and Converse Guardrails, but the new Azure design, schema and Generate gates classify every proprietary openai.gpt-* ID as Mantle-only. They require a model change or reject a supported same-model configuration, defeating the stated same-model selection policy.

Propagate the model/endpoint/feature matrix into Azure selection, schemas, pricing and generation. Permit GPT-5.6 runtime with the proper inference-profile IDs and IAM when its feature and residency constraints match; keep the actual Mantle-only restrictions for GPT-5.5/5.4.

The review summary lists the full affected paths and acceptance checks for both plugin copies.

…xed; false positives noted)

Triaged the 15-finding automated review (leon1418, verified vs head 9d57d39). Fixed
the genuine bugs; documented push-backs on false positives / deferred items.

REAL bugs fixed (both trees):

F2 [P1] Redis cluster encryption arg — verified against the AWS provider schema
(terraform providers schema -json): at_rest_encryption_enabled is NOT a valid
argument on aws_elasticache_cluster (only transit_encryption_enabled is; at-rest
requires aws_elasticache_replication_group). Our checker required it and the 'good'
fixture set it, so following our guidance produced un-appliable Terraform. Fixed the
checker to require transit_encryption_enabled on the standalone Redis cluster and
direct at-rest to a replication group; corrected the fixture, fix_hint, header
docstring, security-posture-rules.md, and the test comment. 60 policy tests pass.

F13 [P2] New Zealand region — the region map claimed "No AWS region in New Zealand"
and defaulted newzealandnorth to Sydney. AWS Asia Pacific (New Zealand), ap-southeast-6,
has been GA since 2025-09-30 (confirmed via the internal region registry). Fixed to
same-country ap-southeast-6 with a service-availability/opt-in fallback note.

F6 [P2] derived_uncorroborated provenance — producers emit
azure_type_provenance: "derived_uncorroborated" for an unrecognized namespace (the
advertised non-halting fallback), but the discover.md gate _assert only allowed
{table, derived, user_confirmed} — so that resource failed the gate. Added
derived_uncorroborated to the enum. (The model_category half of the finding was already
correct in the design schema — no change.)

F5 [P2] Terraform-form admission — extract-terraform.md handles azapi_resource and
downloaded modules, but discover-iac.md dialect detection only fired 'terraform' on a
resource "azurerm_ block, so an AzAPI-only or module-only repo was seen as "no IaC" and
exited before the extractor. Broadened the detection trigger to azurerm_/azapi_/module
blocks.

F7 [P2] Routing fields dropped by the extractor allowlist — messaging.md routes
ServiceBus queues with requires_session=true to SQS FIFO, but the allowlist had no
ServiceBus rows at all, and no separate agentPools (child AKS node pool) row. Added
Microsoft.ServiceBus/namespaces{,/queues,/topics} (requires_session, size/TTL/dedup)
and managedClusters/agentPools rows.

F14 [P2] Collision-safe ARM identity — tf_address was <type>.<local_name> without the
module prefix, and the tf: unresolved-name segment was the bare local name, so two
invocations of the same module (each azurerm_service_plan.this, same RG) reconstructed
the SAME azure_id and the assembler collapsed two real resources into one. Made
tf_address the full module-qualified address and module-qualified the tf: segment.

False positives (no change; will note on the PR):
- F12 [P2] SQS message size — our rubric correctly caps SQS at 256 KB and routes larger
  payloads to claim-check/Amazon MQ. The finding's "SQS accepts up to 1 MiB directly" is
  incorrect (that is Service Bus's limit).

Deferred / disclosed (out of this PR's scope per the human reviewer's sequencing):
F1, F4, F8, F9, F15, F3 — AI-only backbone-gate carry-through, Generate worker-context
ordering, App Service isolation split, per-workload AI source branching, GPT-5.6
CRIS/Mantle classification, and preference-envelope normalization. These are the deferred
AI-route / cross-skill follow-ups; F11's doc-scope tail (advisor/AGENTS.md +
migrate/README.md Terraform-first) is also included in this commit.

Verified both trees: drift OK (401 identical), frontmatter 7, asserters 15 PASS,
lint:model-ids OK, 60 policy tests + 92 report tests pass, mise run build all 16 tasks
clean, dprint clean.
@icarthick

Copy link
Copy Markdown
Collaborator Author

Response to the automated (open-code-review) findings — f1aa1bdf

Triaged all 15. Fixed the genuine bugs, with one verification worth calling out; noted the false positive and the deferred set.

Fixed (real bugs, both trees)

F2 [P1] Redis cluster encryption arg. Verified against the actual AWS provider schema (terraform providers schema -json): at_rest_encryption_enabled is not a valid argument on aws_elasticache_cluster (only transit_encryption_enabled is; at-rest requires aws_elasticache_replication_group). Our checker required it and the "good" fixture set it, so following our own fix_hint produced un-appliable Terraform. Fixed the checker to require transit_encryption_enabled on the standalone Redis cluster and route at-rest to a replication group; corrected the fixture, fix_hint, header docstring, security-posture-rules.md, and the test. 60 policy tests pass. Good catch.

F13 [P2] New Zealand region. Confirmed AWS Asia Pacific (New Zealand), ap-southeast-6, GA 2025-09-30. The map's "No AWS region in New Zealand → Sydney" was stale; fixed to same-country ap-southeast-6 with a service-availability/opt-in fallback note.

F6 [P2] derived_uncorroborated provenance. The Discover gate _assert only allowed {table, derived, user_confirmed}, so the value the unrecognized-namespace fallback actually emits failed the gate. Added it to the enum. (The model_category half was already correct in the design schema.)

F5 [P2] Terraform-form admission. extract-terraform.md handles azapi_resource and downloaded modules, but discover-iac.md only detected terraform on a resource "azurerm_ block — so an AzAPI-only or module-only repo exited as "no IaC". Broadened the trigger.

F7 [P2] Routing fields dropped. messaging.md routes ServiceBus queues with requires_session=true to SQS FIFO, but the extractor allowlist had no ServiceBus rows and no child agentPools row. Added Microsoft.ServiceBus/namespaces{,/queues,/topics} (with requires_session, size/TTL/dedup) and managedClusters/agentPools.

F14 [P2] Collision-safe ARM identity. tf_address and the tf: unresolved-name segment omitted the module prefix, so two invocations of the same module (azurerm_service_plan.this, same RG) reconstructed the same azure_id and collapsed into one. Made both module-qualified.

F11 [P2] doc-scope tail — finished the Terraform-first correction on advisor/AGENTS.md and migrate/README.md (SKILL.md + advisor/README were done in 9d57d393).

False positive (no change)

F12 [P2] SQS message size. The rubric correctly caps SQS at 256 KB and routes larger payloads to a claim-check / Amazon MQ. The finding's premise ("SQS accepts up to 1 MiB directly") is incorrect — 1 MiB is Service Bus's limit; SQS's hard cap is 256 KB, which is exactly the boundary the eliminator uses.

Deferred (out of this PR's scope, per the human review's sequencing)

F1, F4, F8, F9, F15, F3 — AI-only backbone-gate carry-through, Generate worker-context ordering, App Service isolation split, per-workload AI source branching, GPT-5.6 CRIS/Mantle classification, preference-envelope normalization. These are the deferred AI-route / cross-skill follow-ups already scoped as sequenced work; the run_mode _assert-no-teeth (F4-adjacent) remains disclosed.

Verified both trees: drift:check OK (401 identical), frontmatter 7, asserters 15 PASS, lint:model-ids OK, 60 policy + 92 report tests pass, mise run build all 16 tasks clean.

…, not a false positive)

Correcting my earlier triage: F12 from the open-code-review is REAL. Amazon SQS raised
its max message payload from 256 KiB to 1 MiB on 2025-08-04 (standard AND FIFO, all
commercial Regions + GovCloud; Lambda event source mapping updated to match) —
https://aws.amazon.com/about-aws/whats-new/2025/08/amazon-sqs-max-payload-size-1mib/

messaging.md's hard eliminator still capped SQS at 256 KB, so a Service Bus workload
with payloads between 256 KiB and 1 MiB was wrongly excluded from SQS and pushed to an
unnecessary Amazon MQ broker or S3 claim-check. Fixed the boundary to 1 MiB (1024 KiB)
in both rubric copies: <=1 MiB maps to SQS directly (no claim-check); only >1 MiB routes
to Amazon MQ / S3 claim-check.

Verified both trees: drift OK (401 identical), asserters 15 PASS, fmt clean,
mise run build all 16 tasks clean.
@icarthick

Copy link
Copy Markdown
Collaborator Author

Correction on F12 — it's a real bug, now fixed ('"$(git rev-parse --short HEAD)"')

I got F12 wrong in my previous response — apologies. Amazon SQS did raise its max message payload from 256 KiB to 1 MiB on 2025-08-04 (standard and FIFO, all commercial Regions + GovCloud, with Lambda event source mapping updated to match). My "false positive" call was based on the stale 256 KiB limit — the finding is correct.

messaging.md's eliminator was capping SQS at 256 KB, so a Service Bus workload with payloads between 256 KiB and 1 MiB was wrongly excluded from SQS and forced onto an unnecessary Amazon MQ broker or S3 claim-check. Fixed the boundary to 1 MiB (1024 KiB) in both rubric copies: a payload at or under 1 MiB maps to SQS directly (no claim-check); only above 1 MiB routes to Amazon MQ / S3 claim-check.

So the tally on the automated review is now 7 real bugs fixed (F2, F5, F6, F7, F11, F13, F14, F12) — no remaining false positives. The deferred set (F1, F3, F4, F8, F9, F15) stands as the sequenced AI-route / cross-skill follow-ups.

Verified both trees: drift:check OK (401 identical), asserters 15 PASS, fmt:check clean, mise run build all 16 tasks clean.

@icarthick
icarthick merged commit f476698 into awslabs:main Sep 23, 2026
9 checks passed
pull Bot pushed a commit to Spencerx/agent-toolkit-for-aws that referenced this pull request Sep 30, 2026
…ws#352)

* feat(aws-startup-advisor): add live az CLI discovery to azure-to-aws

azure-to-aws discovered only from Terraform/app code; a workspace with no
azurerm_* IaC (the common startup case) halted. This adds the live `az` capture
path the SKILL.md advertised and discover.md reserved a slot for — a second
producer of the same inventory contract, keyed on canonical Microsoft.* ARM
types (no translation table; `az` emits them natively).

Two-part design, honoring discover.md's execution model (_interactive: false,
_exec: { _agent: rw } — the dispatched worker is file-only and cannot prompt):

- Part A (main-window pre-work, invoked from discover.md _preconditions):
  preflight `az`, consent gate, read-only `az resource list` fast-path with
  per-service enrichment fallthrough, writing $MIGRATION_DIR/live-capture/ +
  manifest.json. Read-only allowlist; never captures app-setting/connection-
  string/Key-Vault values or mints a token; explicit --subscription scope.
- Part B (dispatched `live` fragment, this file): file-only parse of
  live-capture/ into azure-resource-inventory.json / clusters, edge inference
  from resolved ids, simplified clustering, and IaC-merge-with-drift when
  Terraform also ran. Entry-guards on the manifest — absent ⇒ no-op.

Wiring: registered the `live` fragment (triggers on live-capture/manifest.json),
added the _live_capture_prework precondition step, and relaxed the
_unrecoverable assertion + two _postconditions so a live-only run is a valid
producer (a no-IaC/no-code workspace with a reachable subscription now discovers
instead of halting). Updated the discover.md Status/live-note and three SKILL.md
claims that said live `az` was "not yet implemented."

Corrections vs the source draft (awslabs/startups#304 comment): schema
references point to the real schema-discover-azure.md (schema-discover-iac.md
does not exist here); azure_type_provenance uses the closed set
{table,derived,derived_uncorroborated,user_confirmed} and ai_source the closed
set {azure_openai,openai,anthropic,both,other}; warnings use the closed
schema-discover-azure.md vocabulary — so it satisfies discover.md's postconditions.

Provenance: capture commands live-validated against a real subscription
(az-cli 2.90.0) for the preflight, `az resource list` fast-path, Container Apps,
PostgreSQL Flexible Server, Storage, ACR, Log Analytics, and Cognitive Services
rows. Rows for services absent from the test subscription (webapp, aks, vm,
mysql, cosmos, redis, service bus, event hubs, vnet, key vault, ML, dns) carry
ported command shapes and should get one live pass before being relied on.

Verified: mise run build exit 0 (markdownlint, manifests, v1.0.0 spec,
skill-parity, gitleaks all green; description 1012 chars, within the 1024 limit).

* fix(aws-startup-advisor): declare the live fragment in the discover assembler _reads

Self-review of aws#352: discover.md registers `live` as a producing fragment, but
discover-assemble.md's _reads block listed only iac + app-code. The assembler
BODY already handles live `az` (it sits highest in the source-precedence rule
and in the drift logic), so this was a declaration gap, not a logic gap — add
`live` to _reads so the assembler's inputs match the phase's registered
fragments.

* fix(aws-startup-advisor): align discover-live.md to the real azure-to-aws schema

Addresses the review of aws#352 — I had ported the source draft's vocabulary
without validating it against this skill's schema/postconditions. Each of these
would have failed the exact no-Terraform run the PR unlocks:

Blocking:
- Edge types: the draft's `network_membership` / `data_dependency` are not in
  schema-discover-azure.md § Typed edges, and the phase postcondition rejects any
  edge type not in that table. Use `network` and `data_ref` (kept `hosted_on`).
- azure_id: the fast-path `az resource list --query` (and per-service rows)
  omitted the ARM id, so a live-only inventory could not fill the required
  `azure_id`. Project `id:id`; Step 3 copies it into `azure_id`.
- Live-only vs Terraform-shaped postconditions: made the `iac_metadata`
  assertion conditional on an IaC dialect contributing, and split the
  producer-agreement assertion so a live-sourced Cognitive Services/ML resource
  satisfies it via detection_signals[].method live_az, infrastructure[] keyed by
  azure_id, sources_analyzed.terraform false, inferred_from_iac false, and no
  tf_address requirement. (models[].detected_via has no live value in the schema
  — the live signal rides on detection_signals[], not detected_via.)

Should-fix:
- Warnings: `reader_role_missing` is not in the closed vocabulary; a capture
  failure is recorded in the manifest note + live_metadata.capture_warnings, not
  warnings[]. Drift: written as resources[].drift `{field, values:[{source,
  value}], won:live}` per the schema, not a `live_metadata.drift.config_conflicts`
  shape with iac_value/live_value.
- SQL row: `az sql db list` requires `--resource-group` as well as `--server`.
  Log Analytics row 11b marked mode E. Added `sql` to the not-yet-live-tested list.

Nit:
- confidence is the enum `deterministic|measured|inferred|billing_inferred`;
  live `az` without metrics is `inferred`, not `0.99`.

Also (self-review, prior commit): declared the `live` fragment in the assembler
_reads.

Verified: markdownlint 0 issues, mise run build exit 0.

* fix(aws-startup-advisor): id:id in every live capture command + add live_metadata to the schema

Follow-up review of aws#352, two should-fixes:

1. The prior commit added id:id only to the fast path and the SQL row, with a
   prose note saying "each per-service --query must project id:id, stated once."
   But the cells ARE the commands, and rows that run only on fallthrough (service
   bus, event hubs, dns) never inherit the fast-path id — so a per_service run
   captured no ARM id and the azure_id postcondition would fail. Put id:id inside
   every --query in the capture table, including both steps of the SQL and
   Cognitive Services (account + deployment) rows.

2. The fragment wrote a top-level live_metadata object plus metadata.clustering_mode
   and resources[].unmanaged_by_iac / not_found_live, none of which the inventory
   schema (schema-discover-azure.md) defined — and the assembler owns the contract
   from that schema. Added them to the schema: a `## live_metadata` section (the
   live counterpart of iac_metadata, documenting capture_warnings/derived_types/
   drift-rollup), clustering_mode under metadata, and the two optional resource
   flags on resources[]. So the live fragment's outputs are now schema-sanctioned
   rather than off-contract.

Verified: markdownlint 0 issues, mise run build exit 0.

* fix(aws-startup-advisor): project resourceGroup on every live capture command

Follow-up review of aws#352 (same class as the id:id fix). Step 3 fills the
inventory's required resource_group from the captured resourceGroup field, and
the phase postcondition requires resource_group on every entry — but only the
fast path and the SQL server list projected it. On a per_service fallthrough the
capture had an ARM id but no resourceGroup, so Step 3 could not fill
resource_group and the postcondition would fail the run.

Add resourceGroup:resourceGroup to every top-level --query in the capture table
(including az sql db list and the Cognitive Services deployment list). Nested
sub-projections (AKS agentPoolProfiles[], vnet subnets[]) are left as name-only —
they are not top-level resources and carry no id/resourceGroup of their own.
Dropped the now-stale "stated once rather than in every cell" note since every
row now carries both required fields explicitly.

Verified: markdownlint 0 issues, mise run build exit 0.

* fix(aws-startup-advisor): address 6 P2 review findings on live az discovery

All six confirmed valid against the actual DSL contract and schema, not
just plausible-sounding review prose:

1. discover.md's _preconditions named a _live_capture_prework check kind
   that does not exist in INTERPRETER.md's closed vocabulary
   (_check_phase_completed/_check_single_active_phase/_check_file_exists/
   _validate_json/_assert) and had no _on_failure. Replaced with a real
   _assert predicate; moved the actual main-window action into prose in
   the phase body (the DSL has no check kind for 'perform an action,'
   only for verifying one).

2. A confirmed re-entry reuses  (per INTERPRETER.md's
   single-run-directory rule), so declining live capture on a re-entry
   left the prior attempt's live-capture/manifest.json in place, and the
   live fragment's existence-based trigger fired on it anyway. Added an
   explicit re-entry check that invalidates/deletes the stale manifest
   on decline.

3. Several capture rows project field names/units that diverge from the
   canonical config keys extract-terraform.md defines for the same ARM
   type (sku/capacity -> sku_name/worker_count for App Service Plans;
   storageGb -> storage_mb, GB->MB, for Postgres/MySQL; etc.), so sizing
   and drift comparisons against an IaC-sourced entry would silently
   read those fields as absent. Added a rename/reshape table. Also fixed
   row 16's deployment capture: the --query projection already flattens
   properties.model.name into a plain string under the key 'model', but
   the prose read 'model.name' as if model were still an object.

4. The fast-path/enrichment merge joined on name+type, which is not
   unique within a subscription, even though every row already
   projects the full ARM id needed for an exact join. Changed the join
   key to azure_id.

5-6. schema-discover-ai.md's infrastructure[] only defined a
   Terraform-shaped entry, and discover.md's producer-agreement
   postcondition had two clauses that required
   sources_analyzed.terraform/inferred_from_iac to be both true (IaC
   resource present) and false (live resource present) in the same
   profile - unsatisfiable by any run with both. Added a live-sourced
   infrastructure[] entry shape (azure_id-keyed, no address/file),
   documented profile_source/sources_analyzed as OR'd across every
   qualifying resource rather than assigned exclusively per producer,
   and rewrote the postcondition plus discover-assemble.md's/
   discover-app-code.md's merge rules to match.

* fix(aws-startup-advisor): check ml extension presence non-interactively before row 17

Live-tested row 17 (az ml workspace list) against a real subscription
without the ml extension installed: it hangs on an interactive Y/n
dynamic-install prompt with no clean error, inside a 120s timeout with
no way to answer - the exact same failure mode the file already warns
about for az graph query. The doc said 'only if the ml extension is
present' but never said how to check that without running the command
first. Added an 'az extension list' pre-check (safe, non-interactive,
names no resource) before row 17, matching the graph-query guardrail
already in this file.

* docs(aws-startup-advisor): README catalog row still said no live az capture

SKILL.md's frontmatter description, top build-status blockquote, and
Philosophy bullet were already updated (in the earlier merge-conflict
resolution against aws#351) to say the live-az capture path is
implemented, but the plugin README's azure-to-aws catalog row was never
touched and still said discovery does not read 'a live Azure CLI
capture.' Swept the rest of the plugin (SKILL.md itself, plugin.json
keywords, other skills' cross-references, the top-level repo README)
for the same stale claim - README.md's catalog row was the only
remaining gap.

* fix(aws-startup-advisor): address 4 P2 review findings on live az discovery

- Reconcile IaC-declared and live-observed AI infrastructure entries that
  describe the same deployed resource: IaC-sourced infrastructure[] entries
  now carry the resource's reconstructed azure_id when resolvable, and the
  assembler merges an IaC entry with a live entry sharing that azure_id into
  one entry (live wins on config disagreement) instead of always treating
  them as two distinct resources.
- Add a closed-vocabulary warning code (enrichment_id_unmatched) for the case
  where a live-capture enrichment row's id has no matching az resource list
  fast-path entry, so discover-live.md's diagnostic instruction can actually
  satisfy the phase's closed-warnings-vocabulary postcondition.
- Mark Event Hubs (row 13) as an enrichment (E) row so kafka_enabled is
  captured on the successful az resource list fast path, not only in the
  per-service fallthrough — otherwise a fast-path Event Hubs namespace always
  routed to Kinesis regardless of its actual Kafka protocol setting.
- Complete the live-capture field normalization table: rename sku->sku_name
  and haMode->high_availability for Postgres/MySQL flexible servers (rows
  7/8), and kind->account_kind for storage accounts (row 11) — without the
  latter, fast-path-services.json's FileStorage exception could never match
  a live-captured storage account.

* fix(aws-startup-advisor): resolve _init/_preconditions ordering cycle, finish live capture field normalization

- discover.md: the pre-dispatch main-window action (live az consent + capture)
  runs BEFORE _preconditions, per this phase's own design -- but this phase
  carries _init: true, and INTERPRETER.md's generic _exec contract runs _init
  state setup AFTER _preconditions. On a cold start this is a dependency
  cycle: capture needs $MIGRATION_DIR, which only _init creates, but _init
  normally doesn't run until after the gate that requires capture to have
  already happened. Fixed by performing _init state setup as the pre-dispatch
  action's new Step 0, before the re-entry check -- _init is idempotent
  (a resumed run just reads the existing .phase-status.json), so this doesn't
  double-initialize or re-prompt resume-vs-fresh; the later "Step: Run the
  phase" _init call is now a no-op re-confirmation for this phase, documented
  as such.
- discover-live.md: completed the field-normalization table with the
  remaining gaps a prior partial fix left uncovered -- rows 1/3 (plan/runtime/
  app-setting/https aliases), 4 (Kubernetes version/node-pool fields,
  preserving every captured pool), 4a (replica bounds), 6 (SQL database
  sku/size), 7 (haMode must populate the nested high_availability.mode shape,
  not replace the object with a scalar), 11 (sku needs SPLITTING into
  account_tier + account_replication_type, not a straight rename -- and 11's
  captured accessTier 'tier' field is a different Azure concept from the
  canonical account_tier, despite the similar name), and 11b (retention_in_days).
  Also updated Step 4's edge-inference table to read the correct field names
  post-normalization (e.g. config.service_plan_id, not the raw sku/plan
  capture key) since edges are inferred after normalization runs.

* fix(aws-startup-advisor): Step 4 edge table must read post-normalization field names for Redis and VM rows

Redis row 10's edge-inference entry named the raw --query source path
(subnetId) instead of the flat capture key that path is projected into
(config.subnet) -- Step 3 only ever writes config.subnet, so a literal
subnetId lookup in Step 4 can never find the value and the redis->vnet
network edge could never be emitted. Fixed that row and proactively fixed
the same latent mismatch on the VM row (row 5), which named
networkProfile.networkInterfaces[0].id instead of its flat capture key
config.subnet (note: despite the key's name it holds a NIC id, not a
subnet id -- the existing resolution note on that row is unchanged).

---------

Co-authored-by: Logan Kleier <lkleier@amazon.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants