Repository navigation
feat(cdk): extract an optional network stack and enforce deployment budgets - #912
Conversation
Protect data stores and destructive cleanup helpers on deletion and replacement. Preserve resource identities and properties so this prerequisite can be applied separately to the existing deployment topology before any compute or stack move. Validate root and nested storage, S3 cleanup retention, and unchanged properties. Refs #852 Co-Authored-By: Codex <noreply@openai.com>
) Finish backend selection across CDK, orchestration, CLI defaults, and MicroVM optional-service configuration. Reject repository pins for undeployed backends. Check all 43 deployment profiles, including nested templates, during the normal build. Share quota and retention checks with the offline census and document the staged migration prerequisites and intended network boundary. Validation: mise run build passed (4970 CDK, 1001 CLI, 1782 agent tests). All 43 census profiles passed two independent syntheses with no differences. Staged secret and diff-based masking scans passed; the full masking scan has three pre-existing findings in unchanged files. Refs #852 Co-Authored-By: Codex <noreply@openai.com>
Keep the application and shared Task API together while placing VPC and DNS resources in a separate stack with networkTopology=split. Preserve generated network names and maintain the complete export set across backend changes. Resolve Blueprint configuration before constructing either stack and carry retention, solution attribution and provenance tags across the boundary. Cover both layouts with 90 deployment profiles and verify the full product twice without staging dependency archives. The widest application drops from 489 to 434 resources. Document local evidence and the unvalidated AWS ownership-transfer procedure; no live rehearsal or deployment was performed.
|
The census reproduces exactly, so I went looking for a dimension it does not Your maxima table reproduces per template, to the resource. AgentCore 482, The gap: To be clear about what is not wrong: auto-pin is correct and deliberately capped On the widest inline profile:
Split is immune, because the subnets move to This is newly reachable rather than pre-existing, which is why I think it is and The guard itself works, and I confirmed that by making it fire. Adding one supplemental
So it is about six lines of profile definition and one fixture zone, and it One thing your description does not claim, in your favor. On And one hypothesis that did not survive, recorded so nobody repeats it. drafted by Bonk |
Bring the latest workflow convergence changes into PR #912. Co-Authored-By: Codex <noreply@openai.com>
Enforce the 490-resource ceiling in production synthesis, add mode-aware three-AZ boundary profiles and rejection checks, and document the supported configurations and migration requirements. Co-Authored-By: Codex <noreply@openai.com>
Keep named AgentCore logs owned across backend switches, reserve AZ address slots for staged network reductions, and flag unavailable repository compute pins in CLI status output. Update resource-budget expectations and migration guidance for #852.
There was a problem hiding this comment.
I tested this live on separate stacks in account, at head 00dcf8b1 and with main at 8cd87fc4 as the baseline. I also prototyped the compute_types idea we discussed. Everything below was run against AWS unless it says otherwise.
What worked
| Scenario | Result |
|---|---|
Fresh install: networkTopology=split + blueprintProvisioning=managed + MicroVM |
✅ network stack, then app stack. Smoke task completed |
Upgrade from main with default flags (inline, legacy) |
✅ 259 s. AgentCore removed as documented, repo rows intact, smoke task completed |
prepare → adopt → managed on a stack with real repo rows |
✅ every check in the deployment guide passed: retention, inert legacy CR, onboarded_at and the CLI override kept, ledger owner moved to managed. As far as I know this is the first live run of the handoff |
What broke
1. B2 is data loss, not only a tombstone. After managed, one deploy without -c blueprintProvisioning goes green with no warning. The legacy PutItem overwrites the Blueprint's repo row, which loses the CLI-set max_turns, compute_type and the original onboarded_at. Then the managed Delete tombstones the row. Going straight back to managed is refused ("Existing repository requires the prepare/adopt blueprint handoff"). prepare → adopt → managed revives the row, but the lost fields don't come back.
Suggestion: persist the mode in cdk.json, or refuse to synth a legacy/prepare mode when the deployed stack already has a Custom::BlueprintRepoConfig.
2. B1 confirmed. Destroying and then redeploying under the same name fails early validation on 4 retained log groups with fixed names:
/aws/lambda-microvms/<stack>-abca-agent/aws/bedrock/model-invocation-logs/<stack>/aws/vendedlogs/bedrock-agentcore/runtime/APPLICATION_LOGS/<stack>/aws/vendedlogs/bedrock-agentcore/runtime/USAGE_LOGS/<stack>
The redeploy never reached the Agent Registry, so I couldn't confirm whether that would collide too.
3. New: a failed first install can't be retried. The retention aspect also applies when a create rolls back. One failed create (it hit my VPC quota) left 50 resources behind, including the same 4 log groups, so every retry fails until someone cleans up by hand. RemovalPolicy.RETAIN_ON_UPDATE_OR_DELETE (CloudFormation's RetainExceptOnCreate) would avoid this.
4. New: one orphaned ledger table per mode change or destroy. Each orphaned table has PITR enabled, so it keeps billing. The upgraded test stack accumulated 3.
5. New, and also on main: AgentCore network interfaces outlive the Runtime. After the Runtime is deleted, its agentic_ai ENIs (attached by amazon-aws, so we can't detach them) stay in use for 5+ hours. They block deleting the RuntimeSG, the private subnets and the VPC, so the destroy ends in DELETE_FAILED. This #912 makes it more likely, because the upgrade deletes AgentCore, but plain destroys hit it too. Worth a note in the deployment guide's teardown section: delete-stack --retain-resources, then delete the VPC later.
6. Naïve inline → split flip (cdk diff only, not deployed). As the guide warns: 63 network resources replaced, 2 log groups orphaned, and two VPCs exist during the cut-over. The populated migration and cdk refactor from #852 still haven't been tested.
compute_types prototype
Branch: https://github.com/isadeks/sample-autonomous-cloud-coding-agents/tree/feat/852-compute-types. It's two commits on top of 00dcf8b1, so it merges cleanly into this branch. The full CDK suite passes (237 suites, 5,655 tests).
compute_types=agentcore,lambda-microvm(a comma list or an array) replaces the exclusivecompute_type. The first entry listed is the default for repositories.- If only the old
compute_type=ecs|lambda-microvmis set, you get AgentCore plus that backend, as onmaintoday. So nobody loses AgentCore on upgrade unless they opt in withcompute_types=<one backend>. - The orchestrator gets the full list in
DEPLOYED_COMPUTE_TYPEand checks repositories against it. - Outputs:
ComputeTypeslists every deployed backend.ComputeDeploymentModeisexclusiveoradditive.ComputeSubstrateis the comma list on additive stacks, which is the format bothmain's CLI and this branch's CLI already parse.
- The Linear and Jira OAuth reads are granted to every compute role.
- Tests:
- 12 additive census profiles were added. The 4 that go over budget with the inline network are marked as expected failures.
- Unit and stack tests were added for both the list and legacy selectors.
- The exclusive stack tests now use
compute_types=<backend>.
Census (yarn census, legacy provisioning, against the 490 cap). Counts are for the app stack; split also has the 58-resource network stack.
| Backends | Default, inline | Default, split | Widest, inline | Widest, split |
|---|---|---|---|---|
| agentcore only (this PR) | 464 | 409 | 483 | 428 |
| ecs only (this PR) | 466 | 411 | 486 | 431 |
| lambda-microvm only (this PR) | 472 | 417 | over (expected) | 437 |
| agentcore + lambda-microvm | 482 | 427 | ✗ 502 | 447 |
| agentcore + ecs | 476 | 421 | ✗ 496 | 441 |
| agentcore + ecs + lambda-microvm | ✗ 495 | 440 | ✗ 514 | 459 |
"Widest" means gateway, registry, vault, alert email and a forked Blueprint repo all enabled. Keeping AgentCore costs about 10 resources. With the split network every combination fits.
Live test on stack backgroundagent-dev-pr912m:
- Deployed
mainwith AgentCore + MicroVM. - Seeded one AgentCore repo row and one MicroVM repo row with
max_turns=77. - Upgraded to the prototype with the same contexts. The diff has no Runtime, role or log-delivery removals; your upgrade deletes all of them. Result ✅ 360 s: the Runtime ARN is unchanged, every table and bucket kept its physical ID, and both rows are intact.
- Ran a smoke task on each backend in the same stack. AgentCore ✅ and Lambda MicroVM ✅. The task records show
compute_type=agentcoreandcompute_type=lambda-microvm, and the VM terminated afterwards.
Gotchas I hit:
ComputeSubstrate: my first version set it to the default backend only.main's CLI then refusedrepo onboard --compute-type lambda-microvm, so additive stacks have to keep the comma list there.- Default backend: existing CLIs assume
agentcoreis the default on any non-exclusive stack. So in an additive list AgentCore needs to come first, unless the CLI learns to readComputeTypes. I didn't change the CLI.
Not done: CLI support for ComputeTypes, and updates to ADR-021/ADR-023 and the compute and deployment docs.
Two more things I hit while pushing the branch
cdk/test/synthesis/workspace.test.tsleaks into the developer's repo under git hooks. Its fixture helper runsgit init/add/commitwithout removingGIT_DIRand related variables, which git sets for hooks. The pre-push hook therefore fails from a worktree. When I reproduced it withGIT_DIRset, the "fixture" commit was made on my real branch, and the index was replaced.src/synthesis/workspace.tsalready stripsGIT_*; the second commit on my branch does the same in the test.- The pre-push Semgrep scan fails on
insecure-object-assignatcdk/src/handlers/linear-webhook-processor.ts:928. That line is onmaintoo (from #863), so this is presumably a--config autorule update, not this PR. It still blocks every push until someone adds anosemgrepcomment with a reason or changes the code. I pushed with--no-verifyfor that reason only.
Still open from the code review (not testable live)
- B3:
adoptnever releases a repository, and a removed repository blocks re-onboarding for 30+ days. - B4:
Custom::BlueprintRepoConfighas no entry inresource-action-map.tsand no coverage test. - Scope: this PR combines exclusive compute, retention, the network split, guardrail versioning, the Blueprint controller, and
.dockerignoreplus dependency bumps.DEVELOPER_GUIDE.md:139asks for retention to ship as its own release. Every failure above came from retention or Blueprint provisioning; nothing failed in the network split or the budgets. So I'd suggest landing the network split + budgets +compute_typesfirst, and retention and Blueprint provisioning separately. - Governance: ADR-023 is still
proposed. #735 isn't labelledapproved. #852'scdk refactorrequirement has only been waived by its author.
Test limits
- Only MicroVM and AgentCore were deployed. ECS was measured in the census only.
- The account runs bootstrap 1.9.0, not the 1.6.0 this branch expects, so missing bootstrap permissions wouldn't show up.
Full deploy logs are available if you want them.
Resolve AgentCore log-delivery conflict by retaining the library-owned IDs and the main preflight migration. Refs #852. Co-Authored-By: Codex <noreply@openai.com>
Replace the exclusive `compute_type` selector with a `compute_types` list (comma list or array). The first listed backend is the repository default. Without `compute_types`, a legacy `compute_type=ecs|lambda-microvm` keeps the additive shape `main` deploys today (AgentCore plus that backend), so an upgrade no longer deletes a co-deployed AgentCore runtime. - orchestrator receives the full list in DEPLOYED_COMPUTE_TYPE and checks repositories against it - new ComputeTypes output; ComputeDeploymentMode is exclusive|additive; ComputeSubstrate stays the comma list existing CLIs parse on additive stacks - shared OAuth secret reads are granted to every compute role - 12 additive census profiles; inline profiles over the 490 budget are expected failures, every split combination fits - #912's exclusive tests and profiles now select with compute_types Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Git hooks export GIT_DIR (and related variables) to their children. The fixture helper in workspace.test.ts inherited them, so under the pre-push hook its `git init/add/commit` ran against the developer's repository: the tests failed, and a manual run with GIT_DIR set committed the fixture onto the current branch. Strip GIT_* the same way src/synthesis/workspace.ts already does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Remove the retention and Blueprint migration prototypes, restore existing resource lifecycle behavior, and reject obsolete migration settings. Complete ordered compute_types support in the CLI, preserve legacy AgentCore deployments, and cover all backend combinations and hook-safe Git fixtures. Document deferred migrations and AgentCore ENI cleanup. Refs #852 Co-authored-by: Codex <codex@openai.com>
Carry SessionRole IAM audit exceptions into lazily generated policies and restrict them to the intended tenant prefixes and literal model grants. Exercise six expanded-model deployment profiles and reject unrelated wildcard permissions. Inspect actual policy documents in the model-set regression, and update the documented 122-profile census. Refs #852 Co-authored-by: Codex <codex@openai.com>
Reject wildcard partitions, regions, and accounts even when an IAM grant matches the Linear provider prefix. Verify both concrete and unresolved environments after CDK creates overflow policies. Correct deployment guidance to state that disabling the registry still deletes its records; retention is deferred. Refs #852 Co-authored-by: Codex <codex@openai.com>
|
Review follow-up at
Validation at this head:
The populated inline-to-split refactor/import and rollback rehearsal remains outstanding. No AWS deployment or migration was performed for these follow-ups, and local synthesis does not establish physical-resource preservation. Retention and Blueprint ownership migration remain separate work; issues #852 and #735 remain open, and ADR-023 remains proposed until merge. The full repository masking scan still reports three existing findings in files identical to main; the required PR-range masking scan passes. The local container image scan remains unverified because Docker was unavailable. |
isadeks
left a comment
There was a problem hiding this comment.
Approving. Live test of 61a1818f on two throwaway stacks (same account as the 2026-10-02 run; main at c0998765). Code review alongside. Nothing below is theoretical unless marked offline.
Live results
| Scenario | Result |
|---|---|
Fresh networkTopology=split with compute_types=lambda-microvm,agentcore |
✅ network then app. ComputeTypes=lambda-microvm,agentcore, mode additive. repo onboard without --compute-type inherits lambda-microvm (first listed); repo show lists both backends. Smoke task on the default → ran on MicroVM; pinned to agentcore → ran on AgentCore. Both completed |
Upgrade main → this head, inline, legacy compute_type=lambda-microvm, contexts unchanged |
✅ 329 s. The diff removes only the GuardrailVersion, API deployments and a Lambda version. Runtime ARN and all 42 stateful physical IDs (tables, buckets, secrets, pool, KMS) unchanged; repo rows intact including a CLI max_turns override. ComputeTypes=agentcore,lambda-microvm. Smoke tasks completed on AgentCore and on MicroVM with the guest image built from main, so an unchanged-context upgrade doesn't require repackaging |
blueprintProvisioning / guardrailVersionMigration |
✅ synth throws the "deferred migration prototype … needs a separate recovery plan" error before any AWS lookup |
| Destroy, then check the 4 fixed-name log groups | ✅ B1 fixed: all 8 (2 stacks × 4) were deleted. The Agent Registries were deleted by the provider too. Leftovers went from 67 (previous head) to 18, and the 18 are the same set main leaves: the API Gateway CloudWatch roles and RegistryApi access log groups (RETAIN by design) plus the VPC pieces below |
| Destroy with an AgentCore Runtime | DELETE_FAILED on RuntimeSG → subnets → VPC because AgentCore's ENIs outlive the Runtime (hours). Pre-existing on main; the PR now documents it, which is the right outcome for this PR |
Code review
Non-blocking (suggestion inline): when both compute_types and the legacy compute_type are set and disagree, compute_types silently wins. For example, cdk.json with compute_type=lambda-microvm plus -c compute_types=agentcore synthesizes a change set that deletes the MicroVM backend without any warning. Nothing deployed today sets both, so this doesn't break anything. I left a small suggestion on compute-backend.ts that throws when the legacy value isn't in the list. Locally it passes the compute-selection and profile tests; the one compute-backend.test.ts row that expects ['lambda-microvm','ecs'] → ['lambda-microvm'] would need to flip to toThrow. Fine here or in a follow-up.
Nits
cdk/test/stacks/compute-selection.test.ts:136-194: the additive cases don't assert the cancel Lambda's env/grants for every deployed backend. The code is right (agent.ts:475-485), but nothing pins it.cdk/test/stacks/agent.test.ts:1138: the suite titledcompute_type=ecsnow synthesizes exclusivecompute_types: 'ecs', so the legacy AgentCore+ECS shape has no stack-level test.agent.ts:1133selectedComputeRoleis now only used in a throw guard.- Upgrade churn, harmless: the
compute_typetag becomesagentcore+lambda-microvmon every taggable resource, and bedrock-alpha re-hashes the GuardrailVersion logical ID.
Checked offline, not an upgrade risk from main: two scenarios looked blocking on paper. (1) Guests built from main reject the new tool_gateway_url/linear_vault_* platform-config keys, so a MicroVM+gateway stack upgraded without repackaging would fail every task. (2) The 490 ceiling rejects some wide inline configs. Neither is reachable: main at c0998765 cannot synthesize enableToolGateway=true with any compute type (cdk-nag AwsSolutions-IAM5 on ToolGatewayRepoConfigFn), refuses lambda-microvm+vault, and fails ecs+vault on cdk-nag as well (this PR synthesizes it at 486). Worth one line in the deployment guide: repackage the MicroVM artifact before enabling the gateway or vault on an existing MicroVM stack.
Looks correct: default backend = first listed, lists without AgentCore, legacy selector → AgentCore first; DEPLOYED_COMPUTE_TYPE membership check fails closed; cancel wiring for all three backends (ECS cancel was never wired on main, so this fixes it); every compute role in the session-role trust, gateway invoke and OAuth secret reads; CLI against old stacks (single-value ComputeSubstrate) and new ones; network exports byte-identical across backends with one-way dependencies; no new resource types for the bootstrap map; GIT_* stripped in the census and its tests.
Still outstanding, as the PR says: the populated inline→split migration has not been rehearsed, and ADR-023 stays proposed until merge.
Wide inline deployments exceed the 490-resource application budget. This change adds an optional network stack for new installations, enforces the resource ceiling during production synthesis, and supports ordered compute-backend lists while preserving legacy deployments' AgentCore runtime.
Area
cdk— network boundaries, compute wiring and synthesis auditsagent— optional MicroVM configuration contract documentation/testscli— compute selection, onboarding and statusdocs— architecture, deployment guidance and generated mirrorstooling— offline synthesis censusRelated
Partial implementation of approved issue #852. Related work: #851, #681, #306 and #735. Issues #852 and #735 remain open after this PR.
The populated refactor/import and rollback rehearsal required by #852 remains open. It has not been waived or satisfied by local synthesis. ADR-023 remains proposed until its implementing PR merges, following the repository's documented ADR lifecycle.
Changes
networkTopology=splitfor new installations. Networking owns the VPC, subnets, endpoints, flow logs and DNS firewall; the application consumes stable exports with one-way dependencies. Inline remains the default. Keep API Gateway integrations, routes, CORS, permissions and deployment together.bedrockModelsexpansion with six additional profiles. SessionRole audit exceptions now follow CDK's lazily generated overflow policies and accept only tenant object prefixes,InvokeModel*and region wildcards on literal foundation-model IDs. Regression tests reject unrelated wildcard actions, resources, model patterns and partitions. The exact-model regression inspects inline and managed policy documents without treating audit metadata as grants.compute_typesimplementation. Comma lists or arrays select one or more backends; the first is the repository default. Legacycompute_type=ecs|lambda-microvmkeeps AgentCore alongside the optional backend. The complete ordered list is published inComputeTypesandComputeSubstrate, and enforced by the orchestrator. Linear and Jira OAuth reads reach every deployed compute role.ComputeTypes, including non-AgentCore defaults and lists that omit AgentCore. Onboarding rejects incompatible stored/input pins before writes or backend probes.repo showandruntime statusreport unavailable pins with corrective errors and omit incompatible probes. Malformed or conflicting deployment outputs fail explicitly.GIT_*hook variables cannot alter another repository or its index.networkReservedAzsand tests for releasing an AZ export before removing a trailing subnet without shifting the remaining subnet CIDRs.Object.assignfix identified by review and fix(deps): clear the http-cache-semantics osv advisory #941'shttp-cache-semanticsresolution floor of^4.3.0; the lockfile and installed tree resolve 4.3.0. The dependency is used by Astro's documentation image build, and it is absent from the standalone Jira Forge dependency tree.Measurements
Fresh structural measurements for the source committed as
61a1818f, with two AZs. The census source fingerprint matches that commit. Counts below are for the application template:“Widest” enables Gateway, Registry, vault, alert email and a fork Blueprint; MicroVM profiles use managed images. The largest split application uses 459 resources and 716,391 bytes, within the 490-resource/800,000-byte budgets. Each two-AZ network has 58 resources and four exports.
The widest single-backend inline configurations with explicit three-AZ pins reach 491/494/500 resources and are rejected. Their split applications remain at 428/431/437 resources; each three-AZ network has 66 resources and five exports. All parent and nested templates are audited individually.
The additional model-expansion probes (platform defaults plus eight synthetic model IDs) use 485/487/493 resources for AgentCore/ECS/MicroVM inline; the 493-resource MicroVM case is rejected. Their split application templates use 430/432/438 resources and report no IAM audit errors.
Reproduce the offline, unbundled census:
The independent-process
--check-stabilityoption preserves timestamps, logical IDs and asset hashes. With the original Blueprint and guardrail implementations restored, it reports existing synthesis churn; budget success is not a deterministic-build claim.Deferred work and deployment limits
Retention and Blueprint ownership migration need separate reviews/releases. ADR-023 records the concrete follow-ups: failed-create cleanup, fixed-name resource recovery, protection against unsafe managed-to-legacy downgrade, preservation of repository overrides, ownership release/re-onboarding, ledger cleanup and bootstrap coverage.
Do not flip an existing inline stack to split with an ordinary deploy. Physical-resource preservation, live refactor/import eligibility and rollback are still unvalidated. This revision has not been deployed to AWS. The deployment guide now documents the reviewer's AgentCore ENIs remaining attached for more than five hours,
DELETE_FAILEDrecovery using exact retained logical IDs, and later cleanup.Validation
MISE_EXPERIMENTAL=1 MISE_JOBS=2 JEST_MAX_WORKERS=2 mise run build: passed — CDK 233 suites / 5,698 tests, CLI 64 suites / 1,041 tests, agent 1,800 tests, plus compilation, lint, bundled synthesis, docs build/link checks, Jira Forge tests and drift checks.61a1818f.origin/mainfound no introduced findings.<=4.2.0). The upstream maintainer disputes the original report; the advisory still has no first-patched-version entry. This validation records a clean dependency gate without claiming a confirmed upstream vulnerability patch.Remaining validation limits:
agent/src/server.py,linear-oauth-resolver.tsandorchestration-store.ts. All affected files match main; the PR-range scan passes.Acknowledgment
By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of the project license.