Skip to content

feat(cdk): extract an optional network stack and enforce deployment budgets - #912

Merged
isadeks merged 22 commits into
mainfrom
docs/852-stack-decomposition-research
Oct 7, 2026
Merged

isadeks merged 22 commits into
mainfrom
docs/852-stack-decomposition-research

Conversation

@krokoko

@krokoko krokoko commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Wide inline deployments exceed the 490-resource application budget. This change adds an optional network stack for new installations, enforces the resource ceiling during production synthesis, and supports ordered compute-backend lists while preserving legacy deployments' AgentCore runtime.

Area

  • cdk — network boundaries, compute wiring and synthesis audits
  • agent — optional MicroVM configuration contract documentation/tests
  • cli — compute selection, onboarding and status
  • docs — architecture, deployment guidance and generated mirrors
  • tooling — offline synthesis census

Related

Partial implementation of approved issue #852. Related work: #851, #681, #306 and #735. Issues #852 and #735 remain open after this PR.

The populated refactor/import and rollback rehearsal required by #852 remains open. It has not been waived or satisfied by local synthesis. ADR-023 remains proposed until its implementing PR merges, following the repository's documented ADR lifecycle.

Changes

  • Add opt-in networkTopology=split for new installations. Networking owns the VPC, subnets, endpoints, flow logs and DNS firewall; the application consumes stable exports with one-way dependencies. Inline remains the default. Keep API Gateway integrations, routes, CORS, permissions and deployment together.
  • Enforce the 490-resource limit on every parent/nested template during ordinary synthesis. Operators may tighten the limit but cannot raise it. Share a 122-profile product between the census and normal build, covering all backend sets, optional services, MicroVM image states, legacy selectors and explicit three-AZ pins. Ten over-budget inline profiles must be rejected by the production guard; unrelated errors cannot satisfy those assertions.
  • Cover bedrockModels expansion with six additional profiles. SessionRole audit exceptions now follow CDK's lazily generated overflow policies and accept only tenant object prefixes, InvokeModel* and region wildcards on literal foundation-model IDs. Regression tests reject unrelated wildcard actions, resources, model patterns and partitions. The exact-model regression inspects inline and managed policy documents without treating audit metadata as grants.
  • Incorporate the reviewer's compute_types implementation. Comma lists or arrays select one or more backends; the first is the repository default. Legacy compute_type=ecs|lambda-microvm keeps AgentCore alongside the optional backend. The complete ordered list is published in ComputeTypes and ComputeSubstrate, and enforced by the orchestrator. Linear and Jira OAuth reads reach every deployed compute role.
  • Complete CLI support for ComputeTypes, including non-AgentCore defaults and lists that omit AgentCore. Onboarding rejects incompatible stored/input pins before writes or backend probes. repo show and runtime status report unavailable pins with corrective errors and omit incompatible probes. Malformed or conflicting deployment outputs fail explicitly.
  • Remove the broad retention aspect, Blueprint migration controller/ledger, custom guardrail-version binding and Docker context changes. Original resource lifecycle policies and providers remain in use. Regression tests cover cleanup policies for all four fixed-name log groups flagged by review. Explicit removed prototype settings are rejected before AWS lookups; previously deployed experimental stacks still require a separate recovery plan.
  • Incorporate the reviewer's Git fixture isolation fix and prove that fixture creation under inherited GIT_* hook variables cannot alter another repository or its index.
  • Preserve Gateway/vault wiring and scoped IAM audit fixes needed by supported combinations. Linear vault exceptions now reject wildcard partitions, regions and accounts, including grants in overflow policies; regression tests cover both concrete and unresolved deployment environments. Retain networkReservedAzs and tests for releasing an AZ export before removing a trailing subnet without shifting the remaining subnet CIDRs.
  • Integrate current main while retaining its library-owned log-delivery IDs and preflight migration tooling. Main supplies the Object.assign fix identified by review and fix(deps): clear the http-cache-semantics osv advisory #941's http-cache-semantics resolution floor of ^4.3.0; the lockfile and installed tree resolve 4.3.0. The dependency is used by Astro's documentation image build, and it is absent from the standalone Jira Forge dependency tree.
  • Correct stale registry lifecycle guidance: disabling an existing registry still deletes its records. The network split does not install the deferred retention behavior.

Measurements

Fresh structural measurements for the source committed as 61a1818f, with two AZs. The census source fingerprint matches that commit. Counts below are for the application template:

Backends Default inline Default split Widest inline Widest split
AgentCore 464 409 483 428
ECS 466 411 486 431
Lambda MicroVMs 472 417 492 — rejected 437
AgentCore + MicroVM 482 427 502 — rejected 447
AgentCore + ECS 476 421 496 — rejected 441
ECS + MicroVM 485 430 504 — rejected 449
All three 495 — rejected 440 514 — rejected 459

“Widest” enables Gateway, Registry, vault, alert email and a fork Blueprint; MicroVM profiles use managed images. The largest split application uses 459 resources and 716,391 bytes, within the 490-resource/800,000-byte budgets. Each two-AZ network has 58 resources and four exports.

The widest single-backend inline configurations with explicit three-AZ pins reach 491/494/500 resources and are rejected. Their split applications remain at 428/431/437 resources; each three-AZ network has 66 resources and five exports. All parent and nested templates are audited individually.

The additional model-expansion probes (platform defaults plus eight synthetic model IDs) use 485/487/493 resources for AgentCore/ECS/MicroVM inline; the 493-resource MicroVM case is rejected. Their split application templates use 430/432/438 resources and report no IAM audit errors.

Reproduce the offline, unbundled census:

MISE_EXPERIMENTAL=1 mise //cdk:census -- --output /tmp/stack-census

The independent-process --check-stability option preserves timestamps, logical IDs and asset hashes. With the original Blueprint and guardrail implementations restored, it reports existing synthesis churn; budget success is not a deterministic-build claim.

Deferred work and deployment limits

Retention and Blueprint ownership migration need separate reviews/releases. ADR-023 records the concrete follow-ups: failed-create cleanup, fixed-name resource recovery, protection against unsafe managed-to-legacy downgrade, preservation of repository overrides, ownership release/re-onboarding, ledger cleanup and bootstrap coverage.

Do not flip an existing inline stack to split with an ordinary deploy. Physical-resource preservation, live refactor/import eligibility and rollback are still unvalidated. This revision has not been deployed to AWS. The deployment guide now documents the reviewer's AgentCore ENIs remaining attached for more than five hours, DELETE_FAILED recovery using exact retained logical IDs, and later cleanup.

Validation

  • MISE_EXPERIMENTAL=1 MISE_JOBS=2 JEST_MAX_WORKERS=2 mise run build: passed — CDK 233 suites / 5,698 tests, CLI 64 suites / 1,041 tests, agent 1,800 tests, plus compilation, lint, bundled synthesis, docs build/link checks, Jira Forge tests and drift checks.
  • The normal build exercises all 122 profiles: 112 successful syntheses and ten required resource-budget rejections.
  • Final standalone boundary census: 44 profiles, 34 successful syntheses, ten expected rejections, zero audit failures. Source inputs remained unchanged during the run, and the source fingerprint was verified against 61a1818f.
  • Comparing those 34 successful application templates against the prior revision shows identical IAM resource properties. The Linear change narrows audit metadata without changing deployed IAM grants.
  • CDK and CLI lint, agent quality, generated-doc sync, contract/type/coverage/pin drift checks and staged/PR-range secret scans passed.
  • All four checks used by the required PR security job pass locally: dependency scan, PR-range secret scan, masking regression scan and GitHub Actions analysis. Grype also reports no vulnerabilities.
  • Full standard SAST, Retire.js and Bandit passed. The masking scan relative to origin/main found no introduced findings.
  • The dependency update clears GHSA-ch52-4w7c-c8xp: 4.3.0 is outside the advisory's flagged range (<=4.2.0). The upstream maintainer disputes the original report; the advisory still has no first-patched-version entry. This validation records a clean dependency gate without claiming a confirmed upstream vulnerability patch.

Remaining validation limits:

  • The full masking scan has three unsuppressed existing findings in agent/src/server.py, linear-oauth-resolver.ts and orchestration-store.ts. All affected files match main; the PR-range scan passes.
  • The container image scan could not run because the local Docker daemon is unavailable.
  • No AWS deployment, populated migration or rollback rehearsal was performed for this revision.

Acknowledgment

By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of the project license.

bgagent and others added 10 commits September 17, 2026 14:17
Protect data stores and destructive cleanup helpers on deletion and replacement.
Preserve resource identities and properties so this prerequisite can be applied
separately to the existing deployment topology before any compute or stack move.

Validate root and nested storage, S3 cleanup retention, and unchanged properties.

Refs #852
Co-Authored-By: Codex <noreply@openai.com>
)

Finish backend selection across CDK, orchestration, CLI defaults, and MicroVM
optional-service configuration. Reject repository pins for undeployed backends.

Check all 43 deployment profiles, including nested templates, during the normal
build. Share quota and retention checks with the offline census and document
the staged migration prerequisites and intended network boundary.

Validation: mise run build passed (4970 CDK, 1001 CLI, 1782 agent tests).
All 43 census profiles passed two independent syntheses with no differences.
Staged secret and diff-based masking scans passed; the full masking scan has
three pre-existing findings in unchanged files.

Refs #852
Co-Authored-By: Codex <noreply@openai.com>
Keep the application and shared Task API together while placing VPC and DNS
resources in a separate stack with networkTopology=split. Preserve generated
network names and maintain the complete export set across backend changes.
Resolve Blueprint configuration before constructing either stack and carry
retention, solution attribution and provenance tags across the boundary.

Cover both layouts with 90 deployment profiles and verify the full product
twice without staging dependency archives. The widest application drops from
489 to 434 resources. Document local evidence and the unvalidated AWS
ownership-transfer procedure; no live rehearsal or deployment was performed.
@laithalsaadoon

laithalsaadoon commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

The census reproduces exactly, so I went looking for a dimension it does not
cover. Measured at head e2836248 with node_modules from an unchanged
yarn.lock, and against main 9d4beda8 in a separate worktree.

Your maxima table reproduces per template, to the resource. AgentCore 482,
ECS 483, Lambda MicroVM 489 on backgroundagent-dev.template.json, and
--list prints exactly 90 profiles. That puts the widest inline profile one
resource under DEFAULT_BUDGETS.resources = 490, whose comment reads "Leave room
for the next change instead of waiting for CloudFormation's hard limit."

The gap: agentcore:availabilityZones appears zero times in
cdk/src/synthesis/.
It is a documented operator override, and its own source calls
it the escape hatch: "operators can always pin explicitly via the
AGENTCORE_AZS_CONTEXT_KEY override". It validates a floor
(MIN_AGENTCORE_AZS = 2, distinct entries, zone IDs rejected, unsupported zones
rejected) with no ceiling. AGENTCORE_SUPPORTED_AZ_IDS declares three
supported zones for us-east-1, which is FIXTURE's own region, while
FIXTURE.zones declares two.

To be clear about what is not wrong: auto-pin is correct and deliberately capped
at AUTO_PIN_AZ_COUNT = 2, kept equal to AgentVpc's maxAzs, with a docstring
naming this exact concern ("3 supported zones would otherwise mean 6 subnets
instead of 4") and a test asserting the coupling. The two-zone fixture is right for
auto-pin. It is only wrong as a bound on the override.

On the widest inline profile:

Configuration Parent resources Subnets Against the 490 budget
auto-pin (fixture) 489 4 1 spare
operator pins 2 AZs 489 4 1 spare
operator pins 3 AZs 497 6 7 over, 3 under CloudFormation's 500
split, 2 AZs 434 0 56 spare
split, 3 AZs 434 0 56 spare

Split is immune, because the subnets move to ${stackName}-network. Your own
recommended topology is the remedy; the default one is what breaches.

This is newly reachable rather than pre-existing, which is why I think it is
worth a profile rather than a note.
On main the widest cell you allow
(ecs + gateway + vault + alertEmail + forkBlueprintRepo) is 496 at two zones,
and a three-zone pin throws at synth:

Number of resources in stack 'backgroundagent-dev': 504 is greater than allowed maximum of 500

and lambda-microvm + vault is refused outright by its own fail-fast. So today a
three-AZ operator gets a loud refusal. After this change the same intent
synthesizes at 497 and deploys, three resources from the hard ceiling, and the new
budget guard reports PASS for all 90 profiles.

The guard itself works, and I confirmed that by making it fire. Adding one supplemental
profile that pins three AZs, plus the third us-east-1 zone in FIXTURE.zones:

  • mise //cdk:census exits 1 with
    backgroundagent-dev.template.json: 497 resources exceeds 490
  • test/synthesis/deployment.test.ts, which gates build, fails
    keeps every template within budget and protects its stateful resources on the
    inline profile and passes on its split twin
  • every pre-existing profile I measured is unchanged: the widest inline baseline
    still reports 489 and PASS, and both other backends still PASS

So it is about six lines of profile definition and one fixture zone, and it
disturbs none of the existing 90. The alternative, if you would rather not widen
the product, is an explicit statement in ADR-023 and the deployment guide that a
pin above AUTO_PIN_AZ_COUNT requires networkTopology=split, but then the
budget is enforced by documentation on the path that is still the default.

One thing your description does not claim, in your favor. On main today,
enableToolGateway=true alone emits 1 error-level AwsSolutions-IAM5
annotation; with the vault, 2; ecs + gateway + vault, 4, three of them on
OverflowPolicy1 paths where the suppression did not follow the spilled statement.
On this branch every one of those configurations reports 0. The shipped
cdk.json default is clean on both revisions at 464 resources, which is why CI
never surfaced it. Worth a line in the description, since it is a real
improvement nobody will otherwise notice.

And one hypothesis that did not survive, recorded so nobody repeats it.
bedrockModels is the other unbounded documented knob missing from the product,
and it never breaches the budget: the widest inline profile holds at 489 through
six extra models and reaches 490 at seven. What it does instead is spill
AgentSessionRole into an OverflowPolicy1 whose cdk-nag suppression does not
follow, emitting 2 error-level AwsSolutions-IAM5 annotations at seven extra
models and 3 at eight. Real, but it needs eleven models in total, and a census
that varied the dimension would catch it.

drafted by Bonk

Bring the latest workflow convergence changes into PR #912.

Co-Authored-By: Codex <noreply@openai.com>
bgagent and others added 2 commits September 22, 2026 15:27
Enforce the 490-resource ceiling in production synthesis, add mode-aware three-AZ boundary profiles and rejection checks, and document the supported configurations and migration requirements.

Co-Authored-By: Codex <noreply@openai.com>
Keep named AgentCore logs owned across backend switches, reserve AZ address slots for staged network reductions, and flag unavailable repository compute pins in CLI status output. Update resource-budget expectations and migration guidance for #852.

@isadeks isadeks left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tested this live on separate stacks in account, at head 00dcf8b1 and with main at 8cd87fc4 as the baseline. I also prototyped the compute_types idea we discussed. Everything below was run against AWS unless it says otherwise.

What worked

Scenario Result
Fresh install: networkTopology=split + blueprintProvisioning=managed + MicroVM ✅ network stack, then app stack. Smoke task completed
Upgrade from main with default flags (inline, legacy) ✅ 259 s. AgentCore removed as documented, repo rows intact, smoke task completed
prepare → adopt → managed on a stack with real repo rows ✅ every check in the deployment guide passed: retention, inert legacy CR, onboarded_at and the CLI override kept, ledger owner moved to managed. As far as I know this is the first live run of the handoff

What broke

1. B2 is data loss, not only a tombstone. After managed, one deploy without -c blueprintProvisioning goes green with no warning. The legacy PutItem overwrites the Blueprint's repo row, which loses the CLI-set max_turns, compute_type and the original onboarded_at. Then the managed Delete tombstones the row. Going straight back to managed is refused ("Existing repository requires the prepare/adopt blueprint handoff"). prepare → adopt → managed revives the row, but the lost fields don't come back.
Suggestion: persist the mode in cdk.json, or refuse to synth a legacy/prepare mode when the deployed stack already has a Custom::BlueprintRepoConfig.

2. B1 confirmed. Destroying and then redeploying under the same name fails early validation on 4 retained log groups with fixed names:

  • /aws/lambda-microvms/<stack>-abca-agent
  • /aws/bedrock/model-invocation-logs/<stack>
  • /aws/vendedlogs/bedrock-agentcore/runtime/APPLICATION_LOGS/<stack>
  • /aws/vendedlogs/bedrock-agentcore/runtime/USAGE_LOGS/<stack>

The redeploy never reached the Agent Registry, so I couldn't confirm whether that would collide too.

3. New: a failed first install can't be retried. The retention aspect also applies when a create rolls back. One failed create (it hit my VPC quota) left 50 resources behind, including the same 4 log groups, so every retry fails until someone cleans up by hand. RemovalPolicy.RETAIN_ON_UPDATE_OR_DELETE (CloudFormation's RetainExceptOnCreate) would avoid this.

4. New: one orphaned ledger table per mode change or destroy. Each orphaned table has PITR enabled, so it keeps billing. The upgraded test stack accumulated 3.

5. New, and also on main: AgentCore network interfaces outlive the Runtime. After the Runtime is deleted, its agentic_ai ENIs (attached by amazon-aws, so we can't detach them) stay in use for 5+ hours. They block deleting the RuntimeSG, the private subnets and the VPC, so the destroy ends in DELETE_FAILED. This #912 makes it more likely, because the upgrade deletes AgentCore, but plain destroys hit it too. Worth a note in the deployment guide's teardown section: delete-stack --retain-resources, then delete the VPC later.

6. Naïve inline → split flip (cdk diff only, not deployed). As the guide warns: 63 network resources replaced, 2 log groups orphaned, and two VPCs exist during the cut-over. The populated migration and cdk refactor from #852 still haven't been tested.

compute_types prototype

Branch: https://github.com/isadeks/sample-autonomous-cloud-coding-agents/tree/feat/852-compute-types. It's two commits on top of 00dcf8b1, so it merges cleanly into this branch. The full CDK suite passes (237 suites, 5,655 tests).

  • compute_types=agentcore,lambda-microvm (a comma list or an array) replaces the exclusive compute_type. The first entry listed is the default for repositories.
  • If only the old compute_type=ecs|lambda-microvm is set, you get AgentCore plus that backend, as on main today. So nobody loses AgentCore on upgrade unless they opt in with compute_types=<one backend>.
  • The orchestrator gets the full list in DEPLOYED_COMPUTE_TYPE and checks repositories against it.
  • Outputs:
    • ComputeTypes lists every deployed backend.
    • ComputeDeploymentMode is exclusive or additive.
    • ComputeSubstrate is the comma list on additive stacks, which is the format both main's CLI and this branch's CLI already parse.
  • The Linear and Jira OAuth reads are granted to every compute role.
  • Tests:
    • 12 additive census profiles were added. The 4 that go over budget with the inline network are marked as expected failures.
    • Unit and stack tests were added for both the list and legacy selectors.
    • The exclusive stack tests now use compute_types=<backend>.

Census (yarn census, legacy provisioning, against the 490 cap). Counts are for the app stack; split also has the 58-resource network stack.

Backends Default, inline Default, split Widest, inline Widest, split
agentcore only (this PR) 464 409 483 428
ecs only (this PR) 466 411 486 431
lambda-microvm only (this PR) 472 417 over (expected) 437
agentcore + lambda-microvm 482 427 ✗ 502 447
agentcore + ecs 476 421 ✗ 496 441
agentcore + ecs + lambda-microvm ✗ 495 440 ✗ 514 459

"Widest" means gateway, registry, vault, alert email and a forked Blueprint repo all enabled. Keeping AgentCore costs about 10 resources. With the split network every combination fits.

Live test on stack backgroundagent-dev-pr912m:

  1. Deployed main with AgentCore + MicroVM.
  2. Seeded one AgentCore repo row and one MicroVM repo row with max_turns=77.
  3. Upgraded to the prototype with the same contexts. The diff has no Runtime, role or log-delivery removals; your upgrade deletes all of them. Result ✅ 360 s: the Runtime ARN is unchanged, every table and bucket kept its physical ID, and both rows are intact.
  4. Ran a smoke task on each backend in the same stack. AgentCore ✅ and Lambda MicroVM ✅. The task records show compute_type=agentcore and compute_type=lambda-microvm, and the VM terminated afterwards.

Gotchas I hit:

  • ComputeSubstrate: my first version set it to the default backend only. main's CLI then refused repo onboard --compute-type lambda-microvm, so additive stacks have to keep the comma list there.
  • Default backend: existing CLIs assume agentcore is the default on any non-exclusive stack. So in an additive list AgentCore needs to come first, unless the CLI learns to read ComputeTypes. I didn't change the CLI.

Not done: CLI support for ComputeTypes, and updates to ADR-021/ADR-023 and the compute and deployment docs.

Two more things I hit while pushing the branch

  • cdk/test/synthesis/workspace.test.ts leaks into the developer's repo under git hooks. Its fixture helper runs git init/add/commit without removing GIT_DIR and related variables, which git sets for hooks. The pre-push hook therefore fails from a worktree. When I reproduced it with GIT_DIR set, the "fixture" commit was made on my real branch, and the index was replaced. src/synthesis/workspace.ts already strips GIT_*; the second commit on my branch does the same in the test.
  • The pre-push Semgrep scan fails on insecure-object-assign at cdk/src/handlers/linear-webhook-processor.ts:928. That line is on main too (from #863), so this is presumably a --config auto rule update, not this PR. It still blocks every push until someone adds a nosemgrep comment with a reason or changes the code. I pushed with --no-verify for that reason only.

Still open from the code review (not testable live)

  • B3: adopt never releases a repository, and a removed repository blocks re-onboarding for 30+ days.
  • B4: Custom::BlueprintRepoConfig has no entry in resource-action-map.ts and no coverage test.
  • Scope: this PR combines exclusive compute, retention, the network split, guardrail versioning, the Blueprint controller, and .dockerignore plus dependency bumps. DEVELOPER_GUIDE.md:139 asks for retention to ship as its own release. Every failure above came from retention or Blueprint provisioning; nothing failed in the network split or the budgets. So I'd suggest landing the network split + budgets + compute_types first, and retention and Blueprint provisioning separately.
  • Governance: ADR-023 is still proposed. #735 isn't labelled approved. #852's cdk refactor requirement has only been waived by its author.

Test limits

  • Only MicroVM and AgentCore were deployed. ECS was measured in the census only.
  • The account runs bootstrap 1.9.0, not the 1.6.0 this branch expects, so missing bootstrap permissions wouldn't show up.

Full deploy logs are available if you want them.

bgagent and others added 5 commits October 2, 2026 17:22
Resolve AgentCore log-delivery conflict by retaining the library-owned IDs and the main preflight migration. Refs #852.

Co-Authored-By: Codex <noreply@openai.com>
Replace the exclusive `compute_type` selector with a `compute_types` list
(comma list or array). The first listed backend is the repository default.
Without `compute_types`, a legacy `compute_type=ecs|lambda-microvm` keeps the
additive shape `main` deploys today (AgentCore plus that backend), so an
upgrade no longer deletes a co-deployed AgentCore runtime.

- orchestrator receives the full list in DEPLOYED_COMPUTE_TYPE and checks
  repositories against it
- new ComputeTypes output; ComputeDeploymentMode is exclusive|additive;
  ComputeSubstrate stays the comma list existing CLIs parse on additive stacks
- shared OAuth secret reads are granted to every compute role
- 12 additive census profiles; inline profiles over the 490 budget are
  expected failures, every split combination fits
- #912's exclusive tests and profiles now select with compute_types

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Git hooks export GIT_DIR (and related variables) to their children. The
fixture helper in workspace.test.ts inherited them, so under the pre-push
hook its `git init/add/commit` ran against the developer's repository: the
tests failed, and a manual run with GIT_DIR set committed the fixture onto
the current branch. Strip GIT_* the same way src/synthesis/workspace.ts
already does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Remove the retention and Blueprint migration prototypes, restore existing
resource lifecycle behavior, and reject obsolete migration settings.
Complete ordered compute_types support in the CLI, preserve legacy
AgentCore deployments, and cover all backend combinations and hook-safe
Git fixtures. Document deferred migrations and AgentCore ENI cleanup.

Refs #852

Co-authored-by: Codex <codex@openai.com>
Carry SessionRole IAM audit exceptions into lazily generated policies and
restrict them to the intended tenant prefixes and literal model grants.
Exercise six expanded-model deployment profiles and reject unrelated
wildcard permissions. Inspect actual policy documents in the model-set
regression, and update the documented 122-profile census.

Refs #852

Co-authored-by: Codex <codex@openai.com>
bgagent and others added 2 commits October 5, 2026 15:04
Reject wildcard partitions, regions, and accounts even when an IAM grant
matches the Linear provider prefix. Verify both concrete and unresolved
environments after CDK creates overflow policies.

Correct deployment guidance to state that disabling the registry still
deletes its records; retention is deferred.

Refs #852

Co-authored-by: Codex <codex@openai.com>
@krokoko
krokoko requested a review from isadeks October 6, 2026 04:02
@krokoko

krokoko commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

Review follow-up at 61a1818f: the PR is now scoped to optional network extraction, deployment budgets, and compatible compute selection. Here is how the review findings were addressed:

  • Retention, Blueprint handoff, and scope (B1–B4, failed-create retries, orphaned ledgers): Removed the broad retention aspect, Blueprint migration controller/ledger, custom guardrail-version binding, and Docker-context changes. The existing resource lifecycle policies and Blueprint provider remain in use. Regression tests cover the destroy/failed-create cleanup policies of all four fixed-name log groups called out in the live review. Explicit blueprintProvisioning and guardrailVersionMigration settings now fail synthesis with a recovery warning. Stacks deployed with the earlier experimental implementation still need a separate recovery plan; dropping those settings is not a safe downgrade.
  • Compute compatibility and the reviewer’s prototype: Incorporated the compute_types implementation while preserving legacy compute_type=ecs|lambda-microvm behavior, including AgentCore. Explicit ordered lists select the deployed backends and repository default. Completed CLI support for ComputeTypes, including defaults other than AgentCore and lists that omit it. Onboarding rejects incompatible input or stored pins before writes/backend probes; repo show and runtime status report unavailable pins. The orchestrator enforces membership, and every deployed compute role receives the required Linear/Jira OAuth reads.
  • Three-AZ census gap: Added explicit three-AZ profiles for both topologies. The 490-resource ceiling is also enforced by production synthesis for parent and nested stacks, including operator configurations outside the census. Over-budget inline configurations must fail on that exact guard; their split counterparts are checked independently.
  • Expanded model grants and IAM overflow: Added six profiles with eight supplemental model IDs. SessionRole audit exceptions now follow CDK’s lazily generated overflow policies and permit only the required tenant-prefix/model wildcard forms. Negative tests keep unrelated grants visible. The follow-up review also narrowed Linear vault exceptions so wildcard partitions, regions, or accounts cannot hide behind the Linear provider prefix. These tests cover concrete and unresolved deployment environments. IAM resource properties are identical to the prior revision across all 34 successfully synthesized boundary configurations.
  • Git fixture isolation: Incorporated the reviewer’s GIT_* environment isolation fix and regression coverage proving fixture creation cannot modify the calling repository or its index when invoked under Git hooks.
  • Operational guidance: Documented the observed AgentCore ENI delay of more than five hours, inspection of DELETE_FAILED resources, retention of exact logical IDs during a deletion retry, and subsequent manual cleanup. Corrected a stale registry paragraph: disabling an existing registry still deletes its records; retention is deferred.
  • Main integration and dependency failure: Merged main through c0998765, preserving its library-owned log-delivery IDs, migration preflight, and Object.assign fix. This also incorporates fix(deps): clear the http-cache-semantics osv advisory #941’s http-cache-semantics resolution floor of ^4.3.0; both the lockfile and installed tree resolve 4.3.0. OSV and Grype now pass. The package is absent from the standalone Jira Forge dependency tree, and the transitive-pin drift check passes.

Validation at this head:

  • mise run build passed: 5,698 CDK tests, 1,041 CLI tests, and 1,800 agent tests, plus compilation, lint, bundled synthesis, docs, and drift checks.
  • The normal build covers 122 deployment profiles: 112 successful syntheses and ten required budget rejections.
  • The independent boundary census passed 44 profiles: 34 successful syntheses and ten expected rejections, with no audit failures or source changes during the run. The largest split application remains at 459 resources / 716,391 bytes.
  • GitHub’s build and required security workflow are green, as are CodeQL and the other reported checks. The branch has no merge conflict.

The populated inline-to-split refactor/import and rollback rehearsal remains outstanding. No AWS deployment or migration was performed for these follow-ups, and local synthesis does not establish physical-resource preservation. Retention and Blueprint ownership migration remain separate work; issues #852 and #735 remain open, and ADR-023 remains proposed until merge.

The full repository masking scan still reports three existing findings in files identical to main; the required PR-range masking scan passes. The local container image scan remains unverified because Docker was unavailable.

@isadeks isadeks left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. Live test of 61a1818f on two throwaway stacks (same account as the 2026-10-02 run; main at c0998765). Code review alongside. Nothing below is theoretical unless marked offline.

Live results

Scenario Result
Fresh networkTopology=split with compute_types=lambda-microvm,agentcore ✅ network then app. ComputeTypes=lambda-microvm,agentcore, mode additive. repo onboard without --compute-type inherits lambda-microvm (first listed); repo show lists both backends. Smoke task on the default → ran on MicroVM; pinned to agentcore → ran on AgentCore. Both completed
Upgrade main → this head, inline, legacy compute_type=lambda-microvm, contexts unchanged ✅ 329 s. The diff removes only the GuardrailVersion, API deployments and a Lambda version. Runtime ARN and all 42 stateful physical IDs (tables, buckets, secrets, pool, KMS) unchanged; repo rows intact including a CLI max_turns override. ComputeTypes=agentcore,lambda-microvm. Smoke tasks completed on AgentCore and on MicroVM with the guest image built from main, so an unchanged-context upgrade doesn't require repackaging
blueprintProvisioning / guardrailVersionMigration ✅ synth throws the "deferred migration prototype … needs a separate recovery plan" error before any AWS lookup
Destroy, then check the 4 fixed-name log groups ✅ B1 fixed: all 8 (2 stacks × 4) were deleted. The Agent Registries were deleted by the provider too. Leftovers went from 67 (previous head) to 18, and the 18 are the same set main leaves: the API Gateway CloudWatch roles and RegistryApi access log groups (RETAIN by design) plus the VPC pieces below
Destroy with an AgentCore Runtime ⚠️ still DELETE_FAILED on RuntimeSG → subnets → VPC because AgentCore's ENIs outlive the Runtime (hours). Pre-existing on main; the PR now documents it, which is the right outcome for this PR

Code review

Non-blocking (suggestion inline): when both compute_types and the legacy compute_type are set and disagree, compute_types silently wins. For example, cdk.json with compute_type=lambda-microvm plus -c compute_types=agentcore synthesizes a change set that deletes the MicroVM backend without any warning. Nothing deployed today sets both, so this doesn't break anything. I left a small suggestion on compute-backend.ts that throws when the legacy value isn't in the list. Locally it passes the compute-selection and profile tests; the one compute-backend.test.ts row that expects ['lambda-microvm','ecs'] → ['lambda-microvm'] would need to flip to toThrow. Fine here or in a follow-up.

Nits

  • cdk/test/stacks/compute-selection.test.ts:136-194: the additive cases don't assert the cancel Lambda's env/grants for every deployed backend. The code is right (agent.ts:475-485), but nothing pins it.
  • cdk/test/stacks/agent.test.ts:1138: the suite titled compute_type=ecs now synthesizes exclusive compute_types: 'ecs', so the legacy AgentCore+ECS shape has no stack-level test.
  • agent.ts:1133 selectedComputeRole is now only used in a throw guard.
  • Upgrade churn, harmless: the compute_type tag becomes agentcore+lambda-microvm on every taggable resource, and bedrock-alpha re-hashes the GuardrailVersion logical ID.

Checked offline, not an upgrade risk from main: two scenarios looked blocking on paper. (1) Guests built from main reject the new tool_gateway_url/linear_vault_* platform-config keys, so a MicroVM+gateway stack upgraded without repackaging would fail every task. (2) The 490 ceiling rejects some wide inline configs. Neither is reachable: main at c0998765 cannot synthesize enableToolGateway=true with any compute type (cdk-nag AwsSolutions-IAM5 on ToolGatewayRepoConfigFn), refuses lambda-microvm+vault, and fails ecs+vault on cdk-nag as well (this PR synthesizes it at 486). Worth one line in the deployment guide: repackage the MicroVM artifact before enabling the gateway or vault on an existing MicroVM stack.

Looks correct: default backend = first listed, lists without AgentCore, legacy selector → AgentCore first; DEPLOYED_COMPUTE_TYPE membership check fails closed; cancel wiring for all three backends (ECS cancel was never wired on main, so this fixes it); every compute role in the session-role trust, gateway invoke and OAuth secret reads; CLI against old stacks (single-value ComputeSubstrate) and new ones; network exports byte-identical across backends with one-way dependencies; no new resource types for the bootstrap map; GIT_* stripped in the census and its tests.

Still outstanding, as the PR says: the populated inline→split migration has not been rehearsed, and ADR-023 stays proposed until merge.

Comment thread cdk/src/handlers/shared/compute-backend.ts
@isadeks
isadeks added this pull request to the merge queue Oct 6, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Oct 6, 2026
@isadeks
isadeks added this pull request to the merge queue Oct 6, 2026
@isadeks
isadeks removed this pull request from the merge queue due to a manual request Oct 6, 2026
@isadeks
isadeks added this pull request to the merge queue Oct 7, 2026
Merged via the queue into main with commit cd9bc54 Oct 7, 2026
9 checks passed
@isadeks
isadeks deleted the docs/852-stack-decomposition-research branch October 7, 2026 02:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants