diff --git a/.dockerignore b/.dockerignore index 32526d373..eb755aa97 100644 --- a/.dockerignore +++ b/.dockerignore @@ -47,6 +47,10 @@ abca-worktrees/ # Test/coverage output coverage/ **/coverage/ +.jest-cache/ +**/.jest-cache/ +test-reports/ +**/test-reports/ .pytest_cache/ **/.pytest_cache/ # Coverage data FILES, not just the directories above. pytest-cov writes diff --git a/.github/workflows/build.yml b/.github/workflows/build.yml index a3ee7a2f8..303061520 100644 --- a/.github/workflows/build.yml +++ b/.github/workflows/build.yml @@ -52,8 +52,15 @@ jobs: compute_type: [agentcore] outputs: self_mutation_happened: ${{ steps.self_mutation.outputs.self_mutation_happened }} + services: + dynamodb: + # Pin the multi-architecture image used for real transaction assertions. + image: amazon/dynamodb-local@sha256:ff89bd48ff32cd8d9be5fee8873b65b8854dc408f1afe881be6eb00247bc0dab + ports: + - 8000/tcp env: CI: "true" + ABCA_TEST_SDK_CONTINUATION: "1" MISE_EXPERIMENTAL: "1" GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }} GITHUB_API_TOKEN: ${{ secrets.GITHUB_TOKEN }} @@ -293,6 +300,7 @@ jobs: - name: build env: TMPDIR: ${{ runner.temp }} + ABCA_DDB_LOCAL_ENDPOINT: http://127.0.0.1:${{ job.services.dynamodb.ports[8000] }} run: | echo "::notice::Runner: $(nproc) cores, $(free -h | awk '/Mem:/{print $2}') RAM" SECONDS=0 diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 672bdc1a9..8d2bffe1c 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -98,6 +98,19 @@ PRs labeled `auto-approve` are approved automatically by the `auto-approve` work If `prek install` fails with "refusing to install hooks with `core.hooksPath` set", another tool owns your hooks. Either unset it (`git config --unset-all core.hooksPath`) or integrate these checks into your hook manager. +The build's transaction tests use DynamoDB Local through +`ABCA_DDB_LOCAL_ENDPOINT` (a loopback HTTP endpoint). CI starts a Docker service +pinned by image digest and supplies its mapped port to both CDK and Python tests; +those tests fail if `CI=true` without an endpoint. To reproduce locally, start +DynamoDB Local with `-inMemory -sharedDb`, bind port 8000 to loopback, and run +`ABCA_DDB_LOCAL_ENDPOINT=http://127.0.0.1:8000 mise run build`. + +The pinned SDK continuation probe also runs in CI. Locally enable it with +`ABCA_TEST_SDK_CONTINUATION=1` when running +`agent/tests/test_continuation_sdk_probe.py`. It uses a local simulated Bedrock +endpoint and synthetic credentials; it does not require an AWS account or model +billing. + ## Versioning The project uses semantic versioning based on [Conventional Commits](https://www.conventionalcommits.org/en/v1.0.0/): diff --git a/agent/.vulture_allowlist.py b/agent/.vulture_allowlist.py index 7f59bc591..e4a3a38c7 100644 --- a/agent/.vulture_allowlist.py +++ b/agent/.vulture_allowlist.py @@ -21,3 +21,8 @@ hook_context # unused variable (src/hooks.py:132) hook_context # unused variable (src/hooks.py:918) hook_context # unused variable (src/hooks.py:1258) + +# urllib's HTTPRedirectHandler supplies this named parameter to redirect_request. +# The bootstrap override intentionally rejects every redirect without inspecting +# its destination; retain the standard-library callback signature. +newurl # unused variable (src/payload_bootstrap.py:_NoRedirect.redirect_request) diff --git a/agent/AGENTS.md b/agent/AGENTS.md index bb13fa087..2f24f83c4 100644 --- a/agent/AGENTS.md +++ b/agent/AGENTS.md @@ -27,6 +27,9 @@ Root `mise run build` includes `//agent:quality` in parallel with `//cdk:build`. | `progress_writer.py` | `agent/tests/test_progress_writer.py` | | `hooks.py`, `policy.py` | `agent/tests/test_hooks.py`, `test_policy.py` | | `pipeline.py`, `runner.py` | `agent/tests/test_pipeline.py`, etc. | +| `microvm_lifecycle.py`, `microvm_checkpoint.py` | `test_microvm_lifecycle.py`, `test_microvm_checkpoint.py`; checkpoint transaction tests require `ABCA_DDB_LOCAL_ENDPOINT` in CI | +| `approval_requests.py`, `task_state.py` | `test_approval_requests.py`, `test_task_state.py`; SigV4 writer protocol and persistence outcomes | +| `continuation_*.py` | Matching `test_continuation_*.py`: capture, storage, safe restore, SDK session and usage recovery | Use `@pytest.fixture(autouse=True)` to reset shared module state between tests when handlers use circuit breakers or caches. @@ -91,4 +94,5 @@ def test_a(): - **Cedar parity** — `cedarpy==4.8.4` (agent) and `@cedar-policy/cedar-wasm` 4.8.2 (cdk) must move together. See [cdk/AGENTS.md](../cdk/AGENTS.md) and `docs/design/CEDAR_HITL_GATES.md` §15.6. - **Forgotten consumer** — Progress event schema changes need `cli/src/commands/watch.ts` and `test_progress_writer.py` updates. - **Image bundle** — CDK deploys this tree; root `mise run build` always runs agent quality. +- **Continuation compatibility** — the SDK pin must match `contracts/constants.json` → `microvm_continuation.verified_sdk_version`. Verify real SDK session recovery before updating both. Persisted checkpoint identities are versioned wire data; renaming a Python field can invalidate existing checkpoints. - **Un-attributed AWS SDK client (#319)** — build clients via `aws_session.tenant_client`/`tenant_resource` (tenant-scoped) or `aws_session.platform_client` (unscoped, still attributed); a naked `boto3.client(...)` silently drops solution attribution. diff --git a/agent/Dockerfile b/agent/Dockerfile index cc280daf1..100fc5611 100644 --- a/agent/Dockerfile +++ b/agent/Dockerfile @@ -128,7 +128,9 @@ COPY agent/prepare-commit-msg.sh /app/ # Claude Code managed settings (#215). The highest-precedence settings layer — # loaded regardless of setting_sources and unoverridable by the untrusted cloned # repo's project .claude/settings.json. Carries awsCredentialExport so Bedrock -# calls use session-tagged, refreshable credentials for cost attribution. +# calls use session-tagged credentials for cost attribution. For MicroVM workers, +# the helper returns no exported keys: the parent's scoped container provider +# handles renewal without Claude's stale awsCredentialExport cache. # Placing awsCredentialExport (an arbitrary command) anywhere the target repo # can influence would be RCE with the compute role, so it lives ONLY here. COPY agent/managed-settings.json /etc/claude-code/managed-settings.json diff --git a/agent/README.md b/agent/README.md index a5f98fcc3..0a48ff11c 100644 --- a/agent/README.md +++ b/agent/README.md @@ -139,6 +139,14 @@ tenant's OAuth or Forge credential to the next task. † You need valid Bedrock credentials in the container: export keys (Option A), let `run.sh` inject keys from the AWS CLI after `aws sso login` or similar (Option B), or mount `~/.aws` (Option C). `run.sh` also sets `CLAUDE_CODE_USE_BEDROCK=1` so Claude Code uses Bedrock. +MicroVM workers configure Claude's AWS provider internally at runtime. The parent +sets `ABCA_MICROVM_CREDENTIAL_BROKER=1` **only in the Claude child**, points +`AWS_CONTAINER_CREDENTIALS_FULL_URI` at an authenticated loopback endpoint, and +clears alternate credential sources in that child. Operators should not set this +internal flag themselves. The endpoint serves the current task's scoped session; +the parent's runtime credentials and other backends' attribution path remain +separate. See [recorded credential-renewal acceptance](../docs/verification/README.md#recorded-acceptance). + ### Examples ```bash @@ -219,7 +227,7 @@ Immediate response (acceptance): Final metrics (PR URL, cost, turns, build status, etc.) appear in **container logs**, in **DynamoDB** when configured, and in the **REST API** for deployed tasks (`GET /v1/tasks/{task_id}` via the `bgagent` CLI or HTTP client). -### AWS Lambda MicroVMs lifecycle hooks (ADR-021 P1 + P2) +### AWS Lambda MicroVMs lifecycle hooks (ADR-021 P1–P3) The same uvicorn process also serves the **Lambda MicroVMs** lifecycle hooks, on the same port (8080 — the port declared in the image's `hooks.port`). On that backend there is no `InvokeAgentRuntime` and no orchestrator→agent HTTP path at all: the task payload arrives as the `/run` hook body and nothing else dials in. @@ -236,31 +244,33 @@ The warm-up is the *primary* fix; the probe that failed is also now non-fatal. ` **`POST /aws/lambda-microvms/runtime/v1/validate`** — Build hook (P2). A **shallow self-check only**: server alive, every hook route registered, interpreter floor, `platform_config` contract loaded. Returns 200 with the individual check results, or 503 while still initialising (which fails the build if it never clears — the right outcome for a broken snapshot). That 503 branch is a **refactor tripwire**, not a state you can reach today: `_module_initialized` is set as the module's last statement and uvicorn accepts no request until the import completes, so it only becomes reachable once someone moves warm-up work behind the bind — and a hook with only a 200 path would then report a still-initialising snapshot as valid. -It runs under the **build role**, which deliberately holds no Bedrock / Secrets Manager / DynamoDB grants, so it **makes zero AWS API calls and must keep making zero** — including its own logging (both build hooks log to stdout via `_build_hook_log`, never through the CloudWatch writer). Two reasons: a Logs write under the build role can only fail, and each failure pollutes the shared `_debug_cw_failures` alarm signal; and `boto3.client(...)` populates `boto3.DEFAULT_SESSION`, a module global holding a resolved credential chain plus the build-time region, which the snapshot would then freeze in for every MicroVM launched from that image version. "Deeper warm-up assertions" (Bedrock reachability, Memory access, tool availability) are therefore *not* implementable here. The one member of that list that turned out to be partly implementable — proving the local `claude` binary execs — lives on `/ready` instead, as a side effect of warming it (above): a local `exec` is not an AWS call, and it belongs to the hook whose 200 gates the snapshot. +It runs under the **build role**, which has no Bedrock, Secrets Manager or DynamoDB runtime grants. Both build hooks make **zero AWS API calls**, including logging: `_build_hook_log` writes to stdout. The build role can write within its MicroVM log namespace, but an application log group may be outside that grant. More importantly, resolving credentials before taking the snapshot can preserve build-role credential state. Boto3 caches resolved credentials; environment-derived region settings are re-read for each new client. Runtime debug/warn writer failures emit a `cloudwatch_write_failed` stdout record containing `writer`, `task_id` and `error_type` ([#810](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/810)). The unused counter was removed. The fallback performs no AWS call and includes neither the failed log body nor exception message. It is a structured log, not a metric or configured alarm; its visibility depends on guest stdout collection, which AgentCore APPLICATION_LOGS does not automatically provide. Local binary warm-up belongs in `/ready`; runtime AWS access must be tested on a real task. Baked secrets are **reported, not enforced**: `warnings` lists the names (never values) of any credential-shaped env var present in the snapshot, because the build environment's own credentials may legitimately be in that env and failing here would fail every build. -**`POST /aws/lambda-microvms/runtime/v1/terminate`** — Runtime hook (P2). Best-effort: emits one final structured log line and returns 200 — always, inside the hook budget, even with nothing running, and for **any body**: malformed JSON, a wrong content-type, an empty body or no body at all. That is why the handler takes the raw request instead of a typed body model — FastAPI validates a typed body *before* the handler runs, so a truncated body would answer 422 and report a hook failure for a teardown that actually succeeded. It does **not** join the pipeline thread (that is `lifespan`'s job on graceful shutdown) and it **never writes terminal task status**: the orchestrator finalizes the task and *then* calls `TerminateMicrovm`, so a status write here would race that finalization. Nothing is buffered to flush — `ProgressWriter` does a synchronous `put_item` per event, so progress is already durable at call time. +**`POST /aws/lambda-microvms/runtime/v1/terminate`** — Closes the local coding barrier, logs teardown and acknowledges any request body within the hook budget. Raw request parsing avoids a premature FastAPI validation error for malformed input. Each step is best-effort. The hook neither joins the pipeline thread nor writes terminal task status: termination can interrupt active work or retire a worker whose approval remains pending. Task finalization belongs to the coordinator. Acknowledged checkpointing belongs to `/suspend`; ordinary progress logging is not a durability guarantee. `microvmId` is parsed defensively and **arrives empty in practice**: the service sends `""` here, unlike `/run` where it is populated (live-verified, ADR-021 P2-F8). So an empty id is expected-normal, not a degraded read — and this hook therefore **cannot** join the guest's record to the control-plane one. `/run`'s `hook accepted task_id=… microvm_id=…` line carries that correlation; `/terminate`'s value is the pipeline-state snapshot it reports. -**`POST /aws/lambda-microvms/runtime/v1/run`** — Payload delivery. Validates the body, installs `platform_config` (below), starts the pipeline in a background thread (the same `_extract_invocation_params` → `_spawn_background` path `/invocations` uses), and returns 200 inside the 1–60 s hook budget. Body: +**`POST /aws/lambda-microvms/runtime/v1/run`** — Authenticate and download a task, install `platform_config` (below), start the pipeline in a background thread, and return 200 inside the hook budget. See the [payload contract and upgrade checks](../docs/verification/645-payload-bootstrap.md) and [recorded live acceptance](../docs/verification/README.md#recorded-acceptance). + +`runHookPayload` is a JSON **string** passed through by `RunMicrovm`, containing: ```json { - "microvmId": "microvm-b44b69d9-…", - "runHookPayload": "{\"agent_payload_s3_uri\": \"s3://bucket//payload.json\", \"platform_config\": {…}}" + "version": 2, + "task_id": "TASK001", + "bootstrap_s3_uri": "s3://deployment-bucket/bootstrap/.json", + "payload_url": "", + "expires_at": 1789312500000 } ``` -`runHookPayload` is an opaque **string** the service passes through from `RunMicrovm`. ABCA's contract for it is one of two shapes, mirroring the ECS container env contract (`AGENT_PAYLOAD` / `AGENT_PAYLOAD_S3_URI`): +The worker reads the deployment manifest with its ambient AWS role, which explicitly denies object reads outside that bucket's `bootstrap/*` prefix. It downloads the task using the coordinator's short-lived signed URL, verifies the task identity, and requires the task's configuration to equal the manifest. The digest in the manifest filename checks its bytes; IAM authenticates its origin. The entire reference must fit **4,096 bytes**. Manifests and task payloads are capped at **16 KiB** and **8 MiB**. Redirects, environment proxies and hosts/paths other than the task's regional S3 object are rejected. -| Envelope | When | -|---|---| -| `{"agent_payload": {…}, "platform_config": {…}}` | the whole orchestrator payload inline — only when it fits | -| `{"agent_payload_s3_uri": "s3://bucket/key", "platform_config": {…}}` | pointer to the payload in the platform payload bucket | +ECS uses the same reference in `AGENT_PAYLOAD_REF`, with an empty manifest config because deployment settings already come from its task definition/overrides. `load_ecs_payload()` removes the capability environment variable before importing the pipeline. Neither backend accepts the old unsigned `AGENT_PAYLOAD`, `AGENT_PAYLOAD_S3_URI`, inline hook or S3-pointer formats. Roll out the matching coordinator, worker images and policies with admissions paused and old tasks drained; the runbook records upgrade and rollback steps. -The service caps `runHookPayload` at **4 096 bytes**, so the **pointer form is the normal one** — a hydrated payload is essentially always larger. Fetching it needs no new env var: the MicroVM execution role holds read-only access to that bucket and the URI carries bucket + key. +A presigned URL is a temporary download permission: **never log it**. Its requested lifetime is at most 900 seconds, shortened by known signer credential expiry; initial creation requires at least 300 seconds. The coordinator privately saves the exact URL for retries and deletes the payload and saved launch reference at finalization. An expired saved reference fails rather than being silently re-signed. #### `platform_config` — the agent's env, delivered per task (P2) @@ -270,12 +280,13 @@ On AgentCore and ECS the agent's non-secret platform env arrives as runtime env { "platform_config": { "task_table_name": "…", "github_token_secret_arn": "arn:…" } } ``` -Each snake_case key installs into its UPPER_SNAKE env var, and a payload value **wins** over any image/pre-existing value (the payload describes the live deployment; the snapshot describes a past one). Installation happens **before** any credential or pipeline initialisation — the very next step resolves the GitHub token from `GITHUB_TOKEN_SECRET_ARN`. Everything the hook logs before that point goes to stdout only (`[server/run-pre-config]`), for the same reason the build hooks do: the CloudWatch writer resolves AWS credentials and pins `boto3.DEFAULT_SESSION` (region included), and until the install has run the only environment available is whatever the snapshot baked. The single AWS call allowed before the install is the S3 payload fetch, because the config is inside the object being fetched. The allowlist lives in `contracts/constants.json` → `microvm_platform_config` (produced by the orchestrator, consumed here; shape enforced by `mise run check:constants-sync`): +Each snake_case key installs into its UPPER_SNAKE env var, and a payload value **wins** over any image/pre-existing value (the payload describes the live deployment; the snapshot describes a past one). Installation happens **before** task credential, secret or pipeline initialisation — the very next step resolves the GitHub token from `GITHUB_TOKEN_SECRET_ARN`. Everything the hook logs before that point goes to stdout only (`[server/run-pre-config]`), for the same reason the build hooks do: the CloudWatch writer resolves AWS credentials and can cache credential state in `boto3.DEFAULT_SESSION` (environment-derived region is re-read for new clients), and until the install has run the only environment available is whatever the snapshot baked. Before installation, bootstrap reads the deployment manifest through the attributed S3 client and downloads the single task object over signed HTTPS. Other pre-install diagnostics remain stdout-only. The allowlist lives in `contracts/constants.json` → `microvm_platform_config` (produced by the orchestrator, consumed here; shape enforced by `mise run check:constants-sync`): | Key | Env var | Required | |---|---|---| | `task_table_name` | `TASK_TABLE_NAME` | ✅ | | `task_events_table_name` | `TASK_EVENTS_TABLE_NAME` | ✅ | +| `approval_requests_api_url` | `APPROVAL_REQUESTS_API_URL` | ✅ | | `github_token_secret_arn` | `GITHUB_TOKEN_SECRET_ARN` | ✅ | | `agent_session_role_arn` | `AGENT_SESSION_ROLE_ARN` | ✅ | | `task_approvals_table_name` | `TASK_APPROVALS_TABLE_NAME` | | @@ -283,16 +294,67 @@ Each snake_case key installs into its UPPER_SNAKE env var, and a payload value * | `log_group_name` | `LOG_GROUP_NAME` | | | `artifacts_bucket_name` | `ARTIFACTS_BUCKET_NAME` | | | `trace_artifacts_bucket_name` | `TRACE_ARTIFACTS_BUCKET_NAME` | | +| `continuation_bucket_name` | `CONTINUATION_BUCKET_NAME` | | | `linear_oauth_secret_arn` | `LINEAR_OAUTH_SECRET_ARN` | | +| `linear_vault_enabled` | `LINEAR_VAULT_ENABLED` | | +| `linear_workload_identity_name` | `LINEAR_WORKLOAD_IDENTITY_NAME` | | | `jira_oauth_secret_arn` | `JIRA_OAUTH_SECRET_ARN` | | | `aws_sdk_ua_app_id` | `AWS_SDK_UA_APP_ID` | | | `anthropic_default_haiku_model` | `ANTHROPIC_DEFAULT_HAIKU_MODEL` | | +| `anthropic_model` | `ANTHROPIC_MODEL` | Main model profile for the deployment's configured geography | + +`APPROVAL_REQUESTS_API_URL` identifies the IAM-authenticated approval writer. All cloud backends receive it from CDK; workers create pending requests through this service and have read-only approval-table access. A MicroVM manifest missing the URL is rejected before execution. Deploy matching worker images and infrastructure together. + +Values are **non-secret configuration only**. Credentials are fetched at task startup from Secrets Manager or AgentCore Identity. Linear vault use requires `linear_vault_enabled="true"` and the workload identity name; the task's channel metadata identifies the workspace grant. The image carries no task credentials. The allowlist **fails closed**: an unrecognised key rejects the whole block. Blank optional values are skipped; blank required values, control characters and inconsistent ARN account/partition fields are rejected. Deployment authentication comes from the IAM-read manifest and exact configuration comparison, including same-account workspace identifiers ([#817](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/817)). + +Rejections are structured so they are readable in the MicroVM log group: `400 MICROVM_RUN_PAYLOAD_INVALID` (unusable envelope — retrying the same body cannot help), `500 MICROVM_RUN_PAYLOAD_UNREADABLE` (manifest/payload read or stored bytes failed), `400 MICROVM_RUN_PLATFORM_CONFIG_INVALID` (key off the allowlist, non-object block, or non-string value — fix the producer), `400 MICROVM_RUN_PLATFORM_CONFIG_INCOMPLETE` (a required key missing or blank — fix the deployment wiring), `400 TASK_RECORD_INCOMPLETE` (same validator and vocabulary as `/invocations`). + +`/suspend` and `/resume` are served by `microvm_http.py` and declared in managed images with 30-second service timeouts and a shared non-secret protocol marker. Pause requires an active original approval gate, matching coordinator intent, drained activity and an acknowledged checkpoint; wake renews credentials and atomically rechecks the original task/gate before allowing coding. Each handler has a 20-second total budget. `/validate` rejects a supplied incompatible image marker without contacting AWS. The coordinator checks the actual launched image version and persists support on that worker; missing support disables new suspension. See the [acceptance checklist](../docs/verification/README.md#live-acceptance-for-an-installation) and [nested deployment prerequisites](../docs/verification/645-p3-nested-stack.md). + +#### Conversation and workspace continuation -Values are **non-secret identifiers only** — secrets are still fetched at `/run` time from Secrets Manager using the ARNs delivered here, so the snapshot stays secret-free. The allowlist **fails closed**: these values land in `os.environ` of the process that spawns the agent's tool subprocesses, so an unrecognised key is an env-injection attempt (`LD_PRELOAD`, `AWS_ENDPOINT_URL`, …) and the whole run is rejected with nothing installed. Blank/`null` values for optional keys are skipped rather than clobbering an image value; blank required keys are rejected. An envelope with no `platform_config` at all is accepted with a loud warning (the image and the orchestrator deploy on independent cadences). +`src/continuation_runtime.py` connects the SDK conversation store, workspace +archive, approval hooks and production runner. Before retiring a worker, it +saves the exact pending tool call, conversation, workflow state, cumulative usage +and workspace in the versioned continuation bucket. A replacement restores those +objects using the coordinator-owned assignment; task payloads cannot choose +arbitrary checkpoint keys. -Rejections are structured so they are readable in the MicroVM log group: `400 MICROVM_RUN_PAYLOAD_INVALID` (unusable envelope — retrying the same body cannot help), `500 MICROVM_RUN_PAYLOAD_UNREADABLE` (the S3 fetch failed), `400 MICROVM_RUN_PLATFORM_CONFIG_INVALID` (key off the allowlist, non-object block, or non-string value — fix the producer), `400 MICROVM_RUN_PLATFORM_CONFIG_INCOMPLETE` (a required key missing or blank — fix the deployment wiring), `400 TASK_RECORD_INCOMPLETE` (same validator and vocabulary as `/invocations`). +`CheckpointSessionStore` implements the pinned SDK's public `SessionStore` +contract with `session_store_flush="eager"`. `checkpoint_pending()` must +acknowledge the exact assistant tool call: enabling transcript mirroring alone +is insufficient because SDK writes are asynchronous. The conversation envelope +has a 16 MiB / 50,000-entry limit. -`/suspend` and `/resume` are deliberately **not** served — declaring a hook nothing answers fails the corresponding lifecycle transition, so the CDK construct declares exactly the hooks the agent serves. They land in P3 with the ComputeStrategy interface widening. +`src/continuation_workspace.py` preserves Git history, staged/unstaged changes, +and untracked/ignored files, with a 1 GiB / 100,000-entry default limit. Restore +rebuilds Git configuration and refuses to replace an existing destination. +Unsupported filesystem/Git states or detected concurrent writes prevent capture. +Repository-free tasks use a private workspace with a local Git baseline. +MicroVM tasks default `UV_LINK_MODE=copy` before repository setup so `uv` does not +hardlink installed packages to its cache. Explicit overrides or commands using +hardlinks can still make the workspace ineligible for capture. + +`S3ContinuationStorage` uploads and verifies version-pinned, checksummed objects +using task-scoped credentials. The continuation bucket and SessionRole grants +are provisioned by CDK. A failed or incomplete save prevents planned retirement; +capacity is released only after the coordinator confirms the old worker stopped. +Checkpoints contain private task data, including unredacted conversation and +workspace contents. They must not be published as diagnostic attachments. + +The replacement consumes the recorded decision and keeps the original cost/turn +allowance minus accumulated usage. See the [retained approval protocol](../docs/design/ORCHESTRATOR.md#retained-microvm-approvals) +for ownership, admission and cleanup. + +The opt-in test uses the actual pinned SDK/CLI with a deterministic loopback +model. It kills the original process, deletes its configuration and workspace, +restores the conversation and files, and verifies approve and deny through a +fresh tool hook: + +```bash +cd agent +ABCA_TEST_SDK_CONTINUATION=1 uv run pytest tests/test_continuation_sdk_probe.py --no-cov +``` ### Testing Server Mode Locally @@ -451,7 +513,7 @@ agent/ │ ├── repo.py Repository setup: clone, branch, git auth, mise trust/install/build/lint │ ├── shell.py Shell utilities: log(), run_cmd(), redact_secrets(), slugify(), truncate() │ ├── telemetry.py Metrics, disk usage, trajectory writer (_TrajectoryWriter with write_policy_decision) -│ ├── server.py FastAPI — async /invocations (background thread), /ping health check, MicroVM /ready + /run lifecycle hooks, heartbeat daemon; OTEL session correlation +│ ├── server.py FastAPI — async /invocations (background thread), /ping health check, MicroVM build/run/terminate hooks (suspend/resume in microvm_http.py), heartbeat daemon; OTEL session correlation │ ├── task_state.py Best-effort DynamoDB task status and heartbeat writes (no-op if TASK_TABLE_NAME unset) │ ├── observability.py OpenTelemetry helpers (e.g. AgentCore session id) │ ├── memory.py Optional memory / episode integration for the agent diff --git a/agent/policies/soft_deny.cedar b/agent/policies/soft_deny.cedar index fd04ae7f2..8104baded 100644 --- a/agent/policies/soft_deny.cedar +++ b/agent/policies/soft_deny.cedar @@ -16,20 +16,20 @@ permit (principal, action, resource); // Every rule in this file MUST carry: // @tier("soft") // @rule_id("...") — stable ID for --pre-approve rule:X -// @approval_timeout_s — integer seconds >= 30 (<120 emits WARN per IMPL-25) // @severity — "low" | "medium" | "high" // @category — optional free-form UX grouping +// An optional @approval_timeout_s sets a positive explicit deadline (>= 30s). +// Omitting it uses the task setting, whose default is no automatic expiry. // // Blueprints may OPT OUT of specific rules here via // `security.cedarPolicies.disable: [rule_id]`. They may NOT disable any // rule in hard_deny.cedar (blueprint loader rejects those at task start). -// Gate any git --force / -f push. 300s default approval window, medium severity. +// Gate any git --force / -f push, with medium severity. // Covers both long-form (--force) and short-form (-f) variants, including // the bare `git push -f` invocation with no branch argument. @tier("soft") @rule_id("force_push_any") -@approval_timeout_s("300") @severity("medium") @category("destructive") forbid (principal, action == Agent::Action::"execute_bash", resource) @@ -37,12 +37,11 @@ when { context.command like "*git push --force*" || context.command like "*git push -f *" || context.command like "*git push -f" }; -// Force-push to main/prod specifically — longer window, higher severity. +// Force-push to main/prod specifically — higher severity. // Multi-match with force_push_any is expected: the engine's annotation -// merging picks min(300, 600)=300s and max(medium, high)=high. +// merging picks max(medium, high)=high. There is no implicit approval expiry. @tier("soft") @rule_id("force_push_main") -@approval_timeout_s("600") @severity("high") @category("destructive") forbid (principal, action == Agent::Action::"execute_bash", resource) @@ -55,7 +54,6 @@ when { context.command like "*git push --force origin main*" // agent bypasses PR workflow by pushing directly. @tier("soft") @rule_id("push_to_protected_branch") -@approval_timeout_s("300") @severity("medium") @category("destructive") forbid (principal, action == Agent::Action::"execute_bash", resource) @@ -64,20 +62,18 @@ when { context.command like "*git push origin main*" || context.command like "*git push origin prod*" || context.command like "*git push origin release/*" }; -// Writes to `.env` files typically contain secrets. 600s window, high severity. +// Writes to `.env` files typically contain secrets. High severity. @tier("soft") @rule_id("write_env_files") -@approval_timeout_s("600") @severity("high") @category("filesystem") forbid (principal, action == Agent::Action::"write_file", resource) when { context.file_path like "*.env" }; // Writes to any path containing "credentials" — SSH keys, AWS creds, -// service-account JSON, etc. 300s window, high severity. +// service-account JSON, etc. High severity. @tier("soft") @rule_id("write_credentials") -@approval_timeout_s("300") @severity("high") @category("auth") forbid (principal, action == Agent::Action::"write_file", resource) diff --git a/agent/scripts/verify_microvm_credentials.py b/agent/scripts/verify_microvm_credentials.py new file mode 100644 index 000000000..941285a78 --- /dev/null +++ b/agent/scripts/verify_microvm_credentials.py @@ -0,0 +1,395 @@ +#!/usr/bin/env python3 +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Opt-in pinned Claude SDK/CLI credential probe; model and credentials are fake. + +Run with agent/.venv/bin/python. This uses a loopback Bedrock event-stream server, +a temporary AWS profile and synthetic keys. It makes no paid model invocation. +It does not simulate a real VM snapshot or establish the AWS runtime provider's +refresh behavior. Keep it out of the ordinary unit suite. +""" + +from __future__ import annotations + +import argparse +import asyncio +import base64 +import importlib.metadata +import json +import os +import re +import shlex +import struct +import subprocess +import sys +import tempfile +import threading +import time +import zlib +from datetime import UTC, datetime +from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer +from pathlib import Path +from typing import TypedDict + +MODEL = "us.anthropic.claude-sonnet-4-20250514-v1:0" + + +class ProbeState(TypedDict): + key: str + expiry: float + fail: bool + + +def event_frame(event: dict) -> bytes: + """Encode one AWS event-stream chunk containing an Anthropic streaming event.""" + headers = bytearray() + for name, value in { + ":event-type": "chunk", + ":content-type": "application/json", + ":message-type": "event", + }.items(): + key, val = name.encode(), value.encode() + headers += bytes([len(key)]) + key + b"\x07" + struct.pack(">H", len(val)) + val + body = json.dumps({"bytes": base64.b64encode(json.dumps(event).encode()).decode()}).encode() + prelude = struct.pack(">II", 16 + len(headers) + len(body), len(headers)) + frame = prelude + struct.pack(">I", zlib.crc32(prelude)) + headers + body + return frame + struct.pack(">I", zlib.crc32(frame)) + + +def model_response() -> bytes: + events = [ + { + "type": "message_start", + "message": { + "id": "msg_offline", + "type": "message", + "role": "assistant", + "model": MODEL, + "content": [], + "stop_reason": None, + "stop_sequence": None, + "usage": {"input_tokens": 1, "output_tokens": 0}, + }, + }, + {"type": "content_block_start", "index": 0, "content_block": {"type": "text", "text": ""}}, + { + "type": "content_block_delta", + "index": 0, + "delta": {"type": "text_delta", "text": "Offline credential probe."}, + }, + {"type": "content_block_stop", "index": 0}, + { + "type": "message_delta", + "delta": {"stop_reason": "end_turn", "stop_sequence": None}, + "usage": {"output_tokens": 1}, + }, + {"type": "message_stop"}, + ] + return b"".join(event_frame(event) for event in events) + + +HELPER = """\ +import json, sys +from datetime import UTC, datetime +from pathlib import Path +state = json.loads(Path(sys.argv[1]).read_text()) +if state["fail"]: + print("synthetic credential renewal failure", file=sys.stderr) + raise SystemExit(7) +creds = { + "AccessKeyId": state["key"], + "SecretAccessKey": "synthetic-secret-not-valid-in-aws", + "SessionToken": "synthetic-session-not-valid-in-aws", + "Expiration": datetime.fromtimestamp(state["expiry"], UTC).isoformat(), +} +with open(sys.argv[2], "a") as log: + log.write(json.dumps({"key": state["key"]}) + "\\n") +print(json.dumps({"Credentials": creds} if sys.argv[3] == "export" else {"Version": 1, **creds})) +""" + + +async def probe(mode: str, *, ambient_fallback: bool) -> dict: + import claude_agent_sdk + from claude_agent_sdk import ClaudeAgentOptions, ClaudeSDKClient, ResultMessage + + sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "src")) + from microvm_credentials import ScopedCredentialBroker + from microvm_lifecycle import MicrovmLifecycle + + requests: list[dict] = [] + initial_expiry = int(time.time()) + 12 + state: ProbeState = {"key": "SYNTHETIC_INITIAL", "expiry": initial_expiry, "fail": False} + stream = model_response() + sdk_version = importlib.metadata.version("claude-agent-sdk") + cli = Path(claude_agent_sdk.__file__).parent / "_bundled" / "claude" + cli_version = subprocess.run( + [str(cli), "--version"], capture_output=True, text=True, check=True, timeout=5 + ).stdout.strip() + if sdk_version != "0.2.110" or cli_version != "2.1.191 (Claude Code)": + raise RuntimeError("Probe expectations require reviewed SDK 0.2.110 / Claude 2.1.191 pins") + + class Handler(BaseHTTPRequestHandler): + def log_message(self, format, *args): + del format, args + + def do_GET(self): + if self.path != "/credentials": + self.send_error(404) + return + body = json.dumps( + { + "AccessKeyId": "SYNTHETIC_AMBIENT", + "SecretAccessKey": "synthetic-ambient-secret", + "Token": "synthetic-ambient-token", + "SessionToken": "synthetic-ambient-token", + "Expiration": datetime.fromtimestamp(time.time() + 3600, UTC).strftime( + "%Y-%m-%dT%H:%M:%SZ" + ), + } + ).encode() + requests.append({"kind": "ambient-provider"}) + self.send_response(200) + self.send_header("Content-Type", "application/json") + self.send_header("Content-Length", str(len(body))) + self.end_headers() + self.wfile.write(body) + + def do_POST(self): + self.rfile.read(int(self.headers.get("Content-Length", "0"))) + key = re.search(r"Credential=([^/]+)", self.headers.get("Authorization", "")) + requests.append( + { + "kind": "model", + "key": key.group(1) if key else "missing", + "phase": state["key"], + "at": time.time(), + "path": self.path, + } + ) + self.send_response(200) + self.send_header("Content-Type", "application/vnd.amazon.eventstream") + self.send_header("Content-Length", str(len(stream))) + self.end_headers() + self.wfile.write(stream) + + server = ThreadingHTTPServer(("127.0.0.1", 0), Handler) + server_thread = threading.Thread(target=server.serve_forever, daemon=True) + server_thread.start() + endpoint = f"http://127.0.0.1:{server.server_port}" + broker = None + try: + with tempfile.TemporaryDirectory(prefix="abca-645-credentials-") as directory: + temp = Path(directory) + helper, state_file, audit = ( + temp / name for name in ("helper.py", "state.json", "audit") + ) + helper.write_text(HELPER) + state_file.write_text(json.dumps(state)) + settings = temp / "settings.json" + provider = ( + "export" + if mode == "export" + else "none" + if mode == "ambient" or mode.startswith("broker") + else "process" + ) + command = shlex.join( + [sys.executable, str(helper), str(state_file), str(audit), provider] + ) + export_command = ( + shlex.join( + [ + sys.executable, + str(Path(__file__).resolve().parents[1] / "src/bedrock_creds_helper.py"), + ] + ) + if mode.startswith("broker") + else command + ) + settings.write_text( + json.dumps( + {"awsCredentialExport": export_command} + if mode == "export" or mode.startswith("broker") + else {} + ) + ) + config = temp / "aws-config" + config.write_text( + "[profile probe]\nregion = us-west-2\n" + + (f"credential_process = {command}\n" if provider == "process" else "") + ) + credentials = temp / "aws-credentials" + credentials.write_text("") + config_dir = temp / "claude-config" + config_dir.mkdir() + child_env = { + "CLAUDE_CONFIG_DIR": str(config_dir), + "CLAUDE_CODE_USE_BEDROCK": "1", + "CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC": "1", + "CLAUDE_CODE_MAX_RETRIES": "0", + "DISABLE_TELEMETRY": "1", + "DISABLE_ERROR_REPORTING": "1", + "DISABLE_AUTOUPDATER": "1", + "ANTHROPIC_BEDROCK_BASE_URL": endpoint, + "AWS_ENDPOINT_URL": endpoint, + "AWS_REGION": "us-west-2", + "AWS_DEFAULT_REGION": "us-west-2", + "AWS_PROFILE": "probe", + "AWS_CONFIG_FILE": str(config), + "AWS_SHARED_CREDENTIALS_FILE": str(credentials), + "AWS_EC2_METADATA_DISABLED": "true", + "AWS_CONTAINER_CREDENTIALS_FULL_URI": endpoint + "/credentials" + if ambient_fallback or mode.startswith("broker") + else "", + "ABCA_MICROVM_CREDENTIAL_BROKER": "1" if mode.startswith("broker") else "", + "AWS_METADATA_SERVICE_TIMEOUT": "1", + "AWS_METADATA_SERVICE_NUM_ATTEMPTS": "1", + "AWS_MAX_ATTEMPTS": "1", + } + if mode.startswith("broker"): + + def scoped_provider(): + requests.append( + {"kind": "broker-failure" if state["fail"] else "broker-provider"} + ) + if state["fail"]: + raise RuntimeError("synthetic broker renewal failure") + return { + "AccessKeyId": state["key"], + "SecretAccessKey": "synthetic-scoped-secret", + "Token": "synthetic-scoped-token", + "Expiration": datetime.fromtimestamp(state["expiry"], UTC).strftime( + "%Y-%m-%dT%H:%M:%SZ" + ), + } + + broker = ScopedCredentialBroker( + MicrovmLifecycle("probe", "probe-vm"), provider=scoped_provider + ) + child_env.update(broker.environment) + errors: list[str] = [] + options = ClaudeAgentOptions( + model=MODEL, + max_turns=1, + cwd=directory, + setting_sources=[], + settings=str(settings), + env=child_env, + stderr=errors.append, + ) + results = [] + async with ClaudeSDKClient(options=options) as client: + for phase in ("initial", "after-expiry"): + if phase == "after-expiry": + await asyncio.sleep(max(0, state["expiry"] - time.time()) + 0.3) + state.update( + key="SYNTHETIC_RENEWED", + expiry=time.time() + 3600, + fail=mode in {"process-failure", "broker-failure"}, + ) + pending = state_file.with_suffix(".pending") + pending.write_text(json.dumps(state)) + pending.replace(state_file) + await client.query("Return one short sentence. Do not call tools.") + async for message in client.receive_response(): + if isinstance(message, ResultMessage): + results.append( + { + "phase": phase, + "error": message.is_error, + **({"detail": message.result} if message.is_error else {}), + } + ) + issued = ( + [json.loads(line) for line in audit.read_text().splitlines()] + if audit.exists() + else [] + ) + return { + "mode": mode, + "sdk_version": sdk_version, + "cli_version": cli_version, + "ambient_fallback": ambient_fallback, + "initial_expiry": initial_expiry, + "requests": requests, + "issued": issued, + "results": results, + "stderr_lines": len(errors), + } + finally: + if broker is not None: + broker.close() + server.shutdown() + server.server_close() + server_thread.join(timeout=2) + + +def verify(result: dict) -> None: + """Positive controls prevent a broken fake endpoint from proving safety.""" + model = [ + request + for request in result["requests"] + if request.get("path") == f"/model/{MODEL}/invoke-with-response-stream" + ] + initial = [request for request in model if request["phase"] == "SYNTHETIC_INITIAL"] + after = [request for request in model if request["phase"] == "SYNTHETIC_RENEWED"] + mode = result["mode"] + expected_initial = "SYNTHETIC_AMBIENT" if mode == "ambient" else "SYNTHETIC_INITIAL" + if ( + len(initial) != 1 + or initial[0]["key"] != expected_initial + or initial[0]["at"] >= result["initial_expiry"] + ): + raise RuntimeError("Initial query must succeed with the expected key before its expiry") + failed = mode == "broker-failure" or ( + mode == "process-failure" and not result["ambient_fallback"] + ) + if [item["error"] for item in result["results"]] != [False, failed]: + raise RuntimeError("Unexpected SDK result; credential-path proof failed") + if failed: + if after: + raise RuntimeError("Renewal failure allowed a model request") + else: + expected = ( + "SYNTHETIC_INITIAL" + if mode == "export" + else "SYNTHETIC_AMBIENT" + if mode in {"ambient", "process-failure"} + else "SYNTHETIC_RENEWED" + ) + if ( + len(after) != 1 + or after[0]["key"] != expected + or after[0]["at"] <= result["initial_expiry"] + ): + raise RuntimeError("Post-expiry query did not use the expected credential path") + result["verified"] = True + + +def main() -> int: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument( + "--mode", + choices=["export", "process", "process-failure", "ambient", "broker", "broker-failure"], + required=True, + ) + parser.add_argument("--ambient-fallback", action="store_true") + args = parser.parse_args() + if args.mode == "ambient" and not args.ambient_fallback: + parser.error("ambient positive control requires --ambient-fallback") + # This process exists only for the probe. Do not inherit operator credentials + # or auth/telemetry settings into the real CLI. HOME is never reassigned. + for key in tuple(os.environ): + if key.startswith(("AWS_", "ANTHROPIC_", "CLAUDE_", "OTEL_", "BEDROCK_")): + del os.environ[key] + result = asyncio.run( + asyncio.wait_for(probe(args.mode, ambient_fallback=args.ambient_fallback), 45) + ) + verify(result) + print(json.dumps(result, indent=2)) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/agent/src/approval_requests.py b/agent/src/approval_requests.py new file mode 100644 index 000000000..31f7e2683 --- /dev/null +++ b/agent/src/approval_requests.py @@ -0,0 +1,134 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""IAM-signed requests to the control-plane approval writer.""" + +import json +import os +import re +from http import HTTPStatus +from urllib.parse import urlsplit + +import requests +from botocore.auth import SigV4Auth +from botocore.awsrequest import AWSRequest +from botocore.exceptions import ClientError + +from aws_session import get_session +from microvm_lifecycle import get_context +from ua import sanitize_ua_value, static_user_agent_extra + +API_ENV = "APPROVAL_REQUESTS_API_URL" + + +def configured() -> bool: + available = bool(os.environ.get(API_ENV)) + if not available and os.environ.get("AGENT_SESSION_ROLE_ARN"): + raise RuntimeError( + "APPROVAL_REQUESTS_API_URL is required for cloud approval requests; " + "deploy matching CDK and worker images" + ) + return available + + +def record_request( + operation: str, + task_id: str, + request_id: str, + *, + approval: dict | None = None, + reason: str | None = None, +) -> None: + """Sign the task path with scoped credentials; never follow redirects.""" + endpoint = os.environ.get(API_ENV, "").rstrip("/") + parsed = urlsplit(endpoint) + match = re.fullmatch( + r"[a-z0-9]+\.execute-api\.([a-z0-9-]+)\.amazonaws\.com(?:\.cn)?", parsed.hostname or "" + ) + if ( + parsed.scheme != "https" + or not match + or parsed.username + or parsed.password + or parsed.port not in (None, 443) + or parsed.query + or parsed.fragment + or not re.fullmatch(r"/[A-Za-z0-9_-]+", parsed.path) + or not re.fullmatch(r"[A-Za-z0-9_-]{1,128}", task_id) + or operation not in {"create", "timeout"} + ): + raise ValueError("Invalid approval request endpoint, task id or operation") + payload = {"operation": operation, "task_id": task_id, "request_id": request_id} + if approval is not None: + payload["approval"] = approval + if reason is not None: + payload["reason"] = reason + lifecycle = get_context(task_id) + if lifecycle is not None: + payload["worker_attempt_id"] = lifecycle.attempt_id + body = json.dumps(payload, separators=(",", ":")).encode() + user_agent = static_user_agent_extra() + app_id = os.environ.get("AWS_SDK_UA_APP_ID") + if app_id: + user_agent += f" app/{sanitize_ua_value(app_id)}" + url = f"{endpoint}/tasks/{task_id}" + request = AWSRequest( + method="POST", + url=url, + data=body, + headers={"Content-Type": "application/json", "User-Agent": user_agent}, + ) + credentials = get_session().get_credentials() + if credentials is None: + raise RuntimeError("Approval request credentials unavailable") + SigV4Auth(credentials.get_frozen_credentials(), "execute-api", match.group(1)).add_auth(request) + response = requests.post( + url, + data=body, + headers=dict(request.headers), + timeout=(5, 20), + allow_redirects=False, + ) + try: + result = response.json() + except ValueError as exc: + raise RuntimeError( + f"Approval service returned invalid JSON (HTTP {response.status_code})" + ) from exc + data = result.get("data") if isinstance(result, dict) else None + if response.status_code == HTTPStatus.OK and isinstance(data, dict) and data.get("ok") is True: + return + error = result.get("error") if isinstance(result, dict) else None + code = ( + error.get("code", "ApprovalServiceUnavailable") + if isinstance(error, dict) + else "ApprovalServiceUnavailable" + ) + details = error.get("details") if isinstance(error, dict) else None + reasons = details.get("cancellation_reasons", []) if isinstance(details, dict) else [] + lease_failed = False + # The third transaction item authorizes the MicroVM worker lease. + match reasons: + case [_, _, {"Code": "ConditionalCheckFailed"}]: + lease_failed = code == "TransactionCanceledException" + message = ( + "MicroVM approval worker lease is missing or inactive. Check worker ownership; " + "pre-upgrade tasks must be drained and resubmitted with matching worker images " + "and coordinator versions." + if lease_failed + else "Control plane did not acknowledge the approval write" + ) + raise ClientError( + { + "Error": { + "Code": code, + "Message": message, + }, + "CancellationReasons": reasons, + "ResponseMetadata": { + "HTTPStatusCode": response.status_code, + "RequestId": error.get("request_id") if isinstance(error, dict) else None, + }, + }, + "RecordApprovalRequest", + ) diff --git a/agent/src/aws_session.py b/agent/src/aws_session.py index 0f2e8f079..1f6a3e0ba 100644 --- a/agent/src/aws_session.py +++ b/agent/src/aws_session.py @@ -8,8 +8,11 @@ uses the resulting short-lived, tag-scoped credentials. The SessionRole's IAM policy self-constrains via ``aws:PrincipalTag/*`` conditions (``dynamodb:LeadingKeys`` on ``task_id`` for the task tables, an S3 prefix -condition on ``user_id`` for the trace bucket), so a compromised session can -only reach its own task's data — not other tenants'. +condition on ``user_id`` for the trace bucket). Existing scoped credentials +can only reach their tagged task. The compute role chooses those tags: a +compromised worker with ambient credentials is a separate trust boundary. +Task updates are additionally restricted to reporting/approval attributes; +whole-row replacement and coordinator metadata writes are not granted. Two properties matter for correctness: @@ -45,6 +48,7 @@ import os import threading +from datetime import UTC from typing import Any # Env var holding the per-task SessionRole ARN. Set by the orchestrator on the @@ -61,6 +65,8 @@ _lock = threading.Lock() _session: Any = None # cached boto3.Session (scoped or plain) _scoped: bool | None = None # None until first resolution; True if tag-scoped +_ambient_lock = threading.Lock() +_ambient_credentials: dict[int, Any] = {} # Session-tag values, set once at startup by ``configure_session`` from the # resolved TaskConfig. Kept in private module state — NOT os.environ — so the @@ -102,11 +108,15 @@ def configure_session(user_id: str, repo: str, task_id: str) -> None: spawned subprocesses. """ global _tags - _tags = { + tags = { key: value for key, value in (("user_id", user_id), ("repo", repo), ("task_id", task_id)) if value } + with _lock: + if _session is not None and _tags != tags: + raise SessionScopingError("Cannot change identity after creating the task session") + _tags = tags def reset_session_cache() -> None: @@ -116,6 +126,8 @@ def reset_session_cache() -> None: _session = None _scoped = None _tags = {} + with _ambient_lock: + _ambient_credentials.clear() def _session_tags() -> list[dict[str, str]]: @@ -141,12 +153,14 @@ def _build_scoped_session(role_arn: str) -> Any: running past the 1-hour role-chaining cap keeps working. """ import boto3 + from botocore.config import Config from botocore.credentials import ( DeferredRefreshableCredentials, ) from botocore.session import get_session as get_botocore_session import ua + from microvm_lifecycle import get_context region = os.environ.get("AWS_REGION") or os.environ.get("AWS_DEFAULT_REGION") task_id = _tags.get("task_id", "") @@ -160,14 +174,21 @@ def _build_scoped_session(role_arn: str) -> Any: # This is the role-chaining caller; the assumed SessionRole credentials it # returns must NOT be used to build it, or refresh would recurse. Carries # the static md/ UA segment so the assume-role call is attributed too. - sts_client = boto3.client("sts", region_name=region, config=ua.client_config()) + sts_config = ( + Config(connect_timeout=2, read_timeout=2, retries={"total_max_attempts": 1}) + if get_context(task_id) is not None + else Config() + ) + sts_client = platform_client("sts", region_name=region, config=sts_config) + # A retained client's refresh must never pick up another task's identity. + tags = _session_tags() def _refresh() -> dict[str, str]: resp = sts_client.assume_role( RoleArn=role_arn, RoleSessionName=session_name, DurationSeconds=_CHAINED_SESSION_DURATION_S, - Tags=_session_tags(), + Tags=tags, ) creds = resp["Credentials"] return { @@ -331,4 +352,83 @@ def platform_client(service_name: str, **kwargs: Any) -> Any: """ import boto3 - return boto3.client(service_name, **_merge_ua_config(kwargs)) + client = boto3.client(service_name, **_merge_ua_config(kwargs)) + credentials = client._request_signer._credentials + with _ambient_lock: + _ambient_credentials[id(credentials)] = credentials + return client + + +def _microvm_scoped_credentials(task_id: str) -> Any: + """Require the existing task identity; never reset/rebuild a cached session.""" + session = get_session() + with _lock: + if not _scoped or not task_id or _tags.get("task_id") != task_id: + raise SessionScopingError("MicroVM credentials require the original scoped task") + return session.get_credentials() + + +def _locked_refresh(credentials: Any, *, force: bool) -> dict[str, str]: + """Refresh and export one coherent key/expiry pair on the retained object. + + Botocore has no public forced-refresh operation. Keep its lock/private API + use here and exercise it against the installed botocore in regression tests. + A mandatory refresh propagates errors even while the old keys remain valid. + The enclosing lifecycle callback supplies the overall wall-clock budget. + """ + from botocore.credentials import RefreshableCredentials + + if not isinstance(credentials, RefreshableCredentials): + # Static env keys cannot prove renewal after sleep. Do not invent a TTL + # or claim that rereading the same environment refreshed the runtime role. + raise SessionScopingError("MicroVM resume requires a refreshable credential provider") + if not credentials._refresh_lock.acquire(timeout=2): + raise TimeoutError("Credential refresh lock did not become available") + try: + if force or credentials.refresh_needed(): + credentials._protected_refresh( + is_mandatory=force + or credentials.refresh_needed(credentials._mandatory_refresh_timeout) + ) + frozen = credentials._frozen_credentials + expiry = credentials._expiry_time + if ( + frozen is None + or not frozen.access_key + or not frozen.secret_key + or not frozen.token + or expiry is None + or credentials._is_expired() + ): + raise SessionScopingError("Credential provider did not return usable temporary keys") + return { + "AccessKeyId": frozen.access_key, + "SecretAccessKey": frozen.secret_key, + "Token": frozen.token, + "Expiration": expiry.astimezone(UTC).strftime("%Y-%m-%dT%H:%M:%SZ"), + } + finally: + credentials._refresh_lock.release() + + +def export_microvm_credentials(task_id: str) -> dict[str, str]: + """Serve only the original task's temporary keys to the local Claude broker.""" + return _locked_refresh(_microvm_scoped_credentials(task_id), force=False) + + +def refresh_microvm_credentials(task_id: str) -> None: + """Refresh ambient callers first, then the same tag-scoped tenant object. + + Call only behind the closed guest resume barrier. Cached DynamoDB/S3 and + platform clients keep their credential references; replacing a boto3 session + would leave those references stale. Unknown/static ambient providers fail + closed until their runtime renewal path has been established. + """ + tenant = _microvm_scoped_credentials(task_id) + with _ambient_lock: + ambient = tuple(_ambient_credentials.values()) + if not ambient: + raise SessionScopingError("No runtime credential provider was recorded") + for credentials in ambient: + _locked_refresh(credentials, force=True) + _locked_refresh(tenant, force=True) diff --git a/agent/src/bedrock_creds_helper.py b/agent/src/bedrock_creds_helper.py index 5f5e4d022..b7f0b9f01 100644 --- a/agent/src/bedrock_creds_helper.py +++ b/agent/src/bedrock_creds_helper.py @@ -7,7 +7,11 @@ ``awsCredentialExport`` setting (in the image's managed-settings layer) runs this script, captures its JSON stdout, and signs Bedrock requests with the returned credentials. With a real ``Expiration`` it re-runs ~5 min before -expiry, so an 8 h task survives the 1 h role-chaining cap. +expiry during active use. Pinned Claude 2.1.191 returns its previous cached +value while that refresh runs in the background: this is NOT an acknowledged +refresh barrier after MicroVM sleep. MicroVM workers instead leave this export +empty and use the scoped container-credential provider installed by the parent. +That provider renews synchronously before signing an expired-key request. Goal: assume the per-task SessionRole with ``{user_id, repo, task_id}`` STS session tags so Bedrock spend is attributable per user/repo in AWS Cost @@ -15,7 +19,7 @@ the cost-allocation tags). The same role already carries the tenant-data grants; Track-1 only adds ``bedrock:InvokeModel*`` to it (see ``agent-session-role.ts``). -**Fails OPEN.** Bedrock attribution is a billing/observability control, not a +**Other backends fail OPEN.** Bedrock attribution is a billing/observability control, not a tenant-isolation one (contrast ``aws_session.py``, which fails closed). If the attribution config is absent or the assume-role fails, this helper emits the **ambient** compute-role credentials so Bedrock keeps working untagged — losing @@ -122,7 +126,13 @@ def _ambient_credentials() -> dict[str, str]: def resolve_credentials() -> dict[str, str]: - """Return tagged assumed-role creds, or ambient creds on any failure.""" + """Use the MicroVM default provider, otherwise retain attribution fallback.""" + # The parent supplies a single scoped loopback provider and removes other + # credential sources from the MicroVM Claude child. Exporting even those + # scoped credentials here would reintroduce Claude's stale export cache. + # Do not read the attribution file or initialize boto3 in this branch. + if os.environ.get("ABCA_MICROVM_CREDENTIAL_BROKER") == "1": + return {} path = attribution_file_path() try: with open(path) as fh: diff --git a/agent/src/config.py b/agent/src/config.py index 3e1c635e1..518bfd858 100644 --- a/agent/src/config.py +++ b/agent/src/config.py @@ -93,10 +93,8 @@ def _resolve_linear_token_via_vault( applicable / unavailable for any reason — the caller then falls back to the Secrets-Manager token, so a vault hiccup must never raise. - Self-mints through boto3 (``get_workload_access_token_for_user_id`` then - ``get_resource_oauth2_token``) rather than reading the AgentCore-injected - ``WorkloadAccessToken`` header, so it behaves identically on the AgentCore - and ECS substrates (both authorize against the agent session role's IAM). + Uses the compute execution role through ``platform_client`` on AgentCore, + ECS and MicroVM. The task-scoped SessionRole does not mint vault tokens. The grant must already be consented (done at ``bgagent linear setup`` time); if the vault returns an ``authorizationUrl`` instead of a token, that means consent is required and we fall back rather than block on a browser. @@ -186,43 +184,21 @@ def _resolve_linear_token_via_vault( def resolve_linear_api_token(channel_metadata: dict[str, str] | None = None) -> str: - """Resolve the Linear OAuth access token from Secrets Manager. - - Phase 2.0b-O2: the orchestrator stamps ``linear_oauth_secret_arn`` - into the task record's ``channel_metadata`` at task-creation time. - Pass that dict in via ``channel_metadata`` (the pipeline does this - automatically). We fetch the per-workspace secret, parse the token - JSON, refresh if expiring, and cache the access_token in - ``LINEAR_API_TOKEN`` so the one remaining consumer — - ``linear_reactions.py``'s direct-GraphQL Authorization header (reactions + - state transitions) — keeps working. (ADR-016: there is no Linear MCP; this - token no longer feeds an ``.mcp.json`` placeholder.) - - For local development, a pre-set ``LINEAR_API_TOKEN`` env var - short-circuits the lookup so the agent can run outside the runtime. - - Returns an empty string when the credential is absent — ``linear_reactions`` - then skips its reactions/state calls (best-effort, logged). This function is - only called when ``channel_source == 'linear'``. - - RFC #249 Phase 1 re-introduces AgentCore Identity as the PREFERRED path - (``_resolve_linear_token_via_vault``, tried first above) for workspaces - onboarded through the vault; Secrets Manager remains the fallback and the - only path for SM-only installs. The Phase-0 spike confirmed the - USER_FEDERATION service-side issue that parked Phase 2.0a no longer - reproduces and that the flow forwards the ``actor=app`` agent install. + """Resolve the Linear token for direct GraphQL reactions and state updates. + + Reuses ``LINEAR_API_TOKEN`` when set, otherwise tries the configured vault + grant before the workspace's Secrets Manager fallback. Task channel metadata + supplies the provider, vault user and fallback secret identifiers. + + Successful resolution caches the token in ``LINEAR_API_TOKEN``. Missing or + unavailable credentials return an empty string; reactions then skip their + calls without stopping the coding task. """ cached = os.environ.get("LINEAR_API_TOKEN", "") if cached: return cached - # RFC #249 Phase 1: when the vault is enabled AND this task carries a - # provider name (vault-onboarded workspace), mint the token through the - # AgentCore Identity Token Vault first. Any failure returns "" here, so we - # fall through to the Secrets-Manager path below — the vault never blocks a - # task. The Phase-0 spike proved USER_FEDERATION forwards the actor=app - # install and re-enables the path that was parked in 2.0a (the service-side - # USER_FEDERATION bug it cited no longer reproduces). + # Vault-managed workspaces may no longer have a usable fallback secret. vault_token = _resolve_linear_token_via_vault(channel_metadata) if vault_token: os.environ["LINEAR_API_TOKEN"] = vault_token @@ -250,7 +226,7 @@ def resolve_linear_api_token(channel_metadata: dict[str, str] | None = None) -> # boto3 is imported here (not just via platform_client, which imports it # lazily at call time) so a missing SDK still degrades gracefully — skip - # Linear MCP — instead of raising an uncaught ImportError. (#319) + # Linear reactions — instead of raising an uncaught ImportError. (#319) import boto3 # noqa: F401 -- availability probe for the graceful skip below from botocore.exceptions import BotoCoreError, ClientError @@ -273,7 +249,10 @@ def _fetch_token() -> dict | None: """ resp = sm.get_secret_value(SecretId=secret_arn) try: - return json.loads(resp["SecretString"]) + payload = json.loads(resp["SecretString"]) + if not isinstance(payload, dict): + raise TypeError("expected a JSON object") + return payload except (json.JSONDecodeError, KeyError, TypeError) as e: log( "ERROR", @@ -314,6 +293,19 @@ def _try_refresh_once(current: dict) -> tuple[str, dict | None]: except ImportError: return ("failure", None) + missing = [ + key + for key in ("refresh_token", "client_id", "client_secret") + if not isinstance(current.get(key), str) or not current[key].strip() + ] + if missing: + log( + "WARN", + "linear_oauth_refresh_unavailable: fallback secret lacks " + f"{', '.join(missing)}; skipping refresh", + ) + return ("failure", None) + body = urllib.parse.urlencode( { "grant_type": "refresh_token", @@ -349,10 +341,6 @@ def _try_refresh_once(current: dict) -> tuple[str, dict | None]: return ("invalid_grant", None) return ("failure", None) except (urllib.error.URLError, OSError) as e: - # Genuine network failures (DNS, timeout, TCP reset). Other - # exceptions (KeyError on missing field, TypeError on bad - # JSON shape) are programmer errors and should propagate - # with a clear stack trace rather than being swallowed. log("WARN", f"resolve_linear_api_token refresh failed: {type(e).__name__}: {e}") return ("failure", None) @@ -373,7 +361,7 @@ def _try_refresh_once(current: dict) -> tuple[str, dict | None]: "access_token": payload["access_token"], "refresh_token": payload.get("refresh_token", current["refresh_token"]), "expires_at": expires_at_iso, - "scope": payload.get("scope", current["scope"]), + "scope": payload.get("scope", current.get("scope", "")), "updated_at": now.isoformat().replace("+00:00", "Z"), } diff --git a/agent/src/continuation_runtime.py b/agent/src/continuation_runtime.py new file mode 100644 index 000000000..fab65ee58 --- /dev/null +++ b/agent/src/continuation_runtime.py @@ -0,0 +1,427 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Compose conversation, workflow and workspace recovery at an approval barrier.""" + +from __future__ import annotations + +import asyncio +import os +import subprocess +import tempfile +import time +from contextlib import contextmanager +from contextvars import ContextVar +from dataclasses import replace +from pathlib import Path +from typing import TYPE_CHECKING, Any + +from claude_agent_sdk import project_key_for_directory + +from continuation_session import ( + CheckpointIdentity, + CheckpointSessionStore, + ContinuationCheckpointError, + decode_checkpoint, +) +from continuation_storage import ( + ContinuationContext, + ContinuationManifest, + FileReceipt, + S3ContinuationStorage, +) +from continuation_usage import TOKEN_FIELDS, read_usage +from continuation_workspace import WorkspaceLimits, capture_workspace, restore_workspace + +if TYPE_CHECKING: + from microvm_lifecycle import ApprovalPark, MicrovmLifecycle + from models import TaskConfig + from policy import PolicyEngine + + +_current: ContextVar[ContinuationRuntime | None] = ContextVar("continuation_runtime", default=None) + + +def current_runtime() -> ContinuationRuntime | None: + return _current.get() + + +@contextmanager +def bind_runtime(runtime: ContinuationRuntime | None): + token = _current.set(runtime) + try: + yield + finally: + _current.reset(token) + + +def create_runtime(context: ContinuationContext, task_id: str) -> ContinuationRuntime | None: + from microvm_lifecycle import get_context + + bucket = os.environ.get("CONTINUATION_BUCKET_NAME", "") + if not bucket or get_context(task_id) is None: + return None + return ContinuationRuntime(context, S3ContinuationStorage(bucket)) + + +def restore_for_task(config: TaskConfig) -> ContinuationRuntime | None: + """Read a coordinator-owned replacement assignment; payloads do not choose S3 keys.""" + import task_state + from config import AGENT_WORKSPACE + from microvm_lifecycle import get_context + + bucket = os.environ.get("CONTINUATION_BUCKET_NAME", "") + lifecycle = get_context(config.task_id) + if not bucket or lifecycle is None: + return None + registration_deadline = time.monotonic() + 30 + while True: + task = task_state.get_task(config.task_id, consistent_read=True) + record = task.get("continuation") if task else None + if not isinstance(record, dict) or record.get("state") != "STARTING": + break + # /run can start the pipeline before RunMicrovm returns its handle to + # the coordinator. Never turn that registration race into a fresh clone. + if time.monotonic() >= registration_deadline: + raise ContinuationCheckpointError( + "Replacement worker registration was not acknowledged" + ) + time.sleep(0.2) + if ( + task is None + or task.get("user_id") != config.user_id + or (task.get("repo") or "") != config.repo_url + ): + raise ContinuationCheckpointError("Continuation task identity is unavailable") + record = task.get("continuation") + if record is None: + return None + if not isinstance(record, dict) or record.get("state") != "RESTORING": + raise ContinuationCheckpointError("Existing continuation cannot start as a fresh task") + identity = CheckpointIdentity(**record["identity"]) + if ( + record.get("version") != 1 + or record.get("worker_id") != lifecycle.microvm_id + or task.get("session_id") != lifecycle.microvm_id + or task.get("status") != "AWAITING_APPROVAL" + or task.get("awaiting_approval_request_id") != identity.request_id + or (identity.task_id, identity.user_id, identity.repo) + != (config.task_id, config.user_id, config.repo_url) + ): + raise ContinuationCheckpointError("Replacement worker does not own this continuation") + workflow = config.resolved_workflow or {} + runtime = ContinuationRuntime.restore( + S3ContinuationStorage(bucket), + FileReceipt.from_record(record["manifest"], identity), + identity, + expected_workspace=Path(AGENT_WORKSPACE).resolve() / config.task_id, + workflow_id=workflow.get("id", ""), + workflow_version=workflow.get("version", ""), + ) + approval = task_state.get_approval_row( + config.task_id, identity.request_id, consistent_read=True + ) + if ( + approval is None + or approval.get("user_id") != config.user_id + or (approval.get("repo") or "") != config.repo_url + or runtime.restored is None + or approval.get("tool_input_sha256") != runtime.restored["action"]["tool_input_sha256"] + or approval.get("status") not in {"APPROVED", "DENIED", "TIMED_OUT"} + ): + raise ContinuationCheckpointError("Continuation human decision is unavailable") + # No SDK process exists yet. Only after restoration and this conditional + # claim may its freshly gated tools start. + task_state.consume_restored_continuation(config.task_id, lifecycle.microvm_id, record) + runtime.resume_prompt = runtime.decision_prompt( + decision=approval["status"], reason=approval.get("deny_reason") or "" + ) + runtime.human_decision = approval + config.initial_approval_gate_count = max( + int(task.get("approval_gate_count", 0)), runtime.context.approval_gate_count + ) + return runtime + + +def prepare_repoless_runtime( + config: TaskConfig, + *, + user_prompt: str, + system_prompt: str, + workflow_id: str, + workflow_version: str, +) -> ContinuationRuntime | None: + """Give a MicroVM task a private scratch workspace with a local-only Git baseline.""" + from config import AGENT_WORKSPACE + from microvm_lifecycle import get_context + from models import RepoSetup + + if not os.environ.get("CONTINUATION_BUCKET_NAME") or get_context(config.task_id) is None: + return None + if config.repo_url: + raise ContinuationCheckpointError( + "Scratch recovery is only for a task without a repository" + ) + restored = restore_for_task(config) + if restored is not None: + return restored + # Validate the task component before joining it to a filesystem path. + CheckpointIdentity(config.task_id, "initial", "initial", config.user_id, "") + workspace = Path(AGENT_WORKSPACE).resolve() / config.task_id + workspace.parent.mkdir(parents=True, exist_ok=True) + workspace.mkdir(mode=0o700, exist_ok=False) + env = {key: value for key, value in os.environ.items() if not key.startswith("GIT_")} + env.update(GIT_CONFIG_GLOBAL=os.devnull, GIT_CONFIG_NOSYSTEM="1", GIT_TERMINAL_PROMPT="0") + + def git(*args: str) -> str: + return subprocess.run( + [ + "git", + "-C", + str(workspace), + "-c", + "core.hooksPath=/dev/null", + "-c", + "user.name=ABCA", + "-c", + "user.email=workspace@invalid", + *args, + ], + env=env, + capture_output=True, + text=True, + check=True, + timeout=30, + ).stdout.strip() + + git("init", "-b", "scratch") + git("commit", "--allow-empty", "--no-gpg-sign", "-m", "Private task workspace") + return create_runtime( + ContinuationContext( + RepoSetup( + repo_dir=str(workspace), branch="scratch", head_sha_before=git("rev-parse", "HEAD") + ), + user_prompt, + system_prompt, + workflow_id, + workflow_version, + ), + config.task_id, + ) + + +class ContinuationRuntime: + """A single runner's state; SDK append/checkpoint calls share its event loop.""" + + def __init__( + self, + context: ContinuationContext, + storage: S3ContinuationStorage, + *, + store: CheckpointSessionStore | None = None, + restored: dict | None = None, + ) -> None: + self.context = context + self.storage = storage + self.store = store or CheckpointSessionStore( + project_key_for_directory(context.setup.repo_dir) + ) + self.restored = restored + self.resume_prompt: str | None = None + self.human_decision: dict | None = None + self._approval_consumed = False + self.client: Any = None + # These remain fixed for this SDK process. Repeated captures must not + # add its cumulative counters to an earlier capture of the same process. + self.prior_cost_usd = context.cost_usd if restored is not None else 0.0 + self.prior_token_usage = dict(context.token_usage) if restored is not None else {} + + def seed_policy(self, engine: PolicyEngine) -> None: + """Carry session grants and the recorded denial across worker replacement.""" + for scope in self.context.approval_scopes: + engine.allowlist.add(scope) + decision = self.human_decision + if self.restored is None or decision is None: + return + action = self.restored["action"] + status = decision["status"] + if status == "APPROVED": + scope = decision.get("scope") or "this_call" + if scope != "this_call": + engine.allowlist.add(scope) + elif status in {"DENIED", "TIMED_OUT"}: + reason = decision.get("deny_reason") or "" + engine.recent_decisions.record( + action["tool_name"], + action["tool_input_sha256"], + decision=status, + reason=reason, + original_decision_ts=decision.get("decided_at"), + ) + if status == "DENIED": + for rule_id in decision.get("matching_rule_ids", []): + engine.recent_decisions.record_rule_decision( + action["tool_name"], + rule_id, + decision="DENIED", + reason=reason, + original_decision_ts=decision.get("decided_at"), + ) + + def consume_approved_action(self, tool_name: str, tool_input_sha256: str) -> dict | None: + """Consume one exact saved proposal, only after the normal policy check.""" + decision = self.human_decision + if ( + self.restored is None + or decision is None + or decision.get("status") != "APPROVED" + or self._approval_consumed + or self.restored["action"]["tool_name"] != tool_name + or self.restored["action"]["tool_input_sha256"] != tool_input_sha256 + ): + return None + self._approval_consumed = True + return decision + + async def capture( + self, + lifecycle: MicrovmLifecycle, + park: ApprovalPark, + *, + session_id: str, + tool_name: str, + tool_input: dict, + approval_scopes: tuple[str, ...] = (), + approval_gate_count: int = 0, + ) -> tuple[CheckpointIdentity, FileReceipt]: + if park.record is None: + raise ContinuationCheckpointError("Approval identity is unavailable for continuation") + identity = CheckpointIdentity( + park.task_id, + park.microvm_id, + park.request_id, + park.record.user_id, + park.record.repo, + ) + async with lifecycle.continuation_checkpoint(park): + body = await self.store.checkpoint_pending( + identity, + session_id=session_id, + tool_use_id=park.tool_use_id, + tool_name=tool_name, + tool_input=tool_input, + timeout_s=10, + ) + entries = decode_checkpoint(body, identity)["entries"] + usage = await read_usage(self.client) + turns = set() + for index, entry in enumerate(entries): + message = entry.get("message") + if entry.get("type") != "assistant" or not isinstance(message, dict): + continue + message_id = message.get("id") + turns.add( + message_id + if isinstance(message_id, str) and message_id + else entry.get("uuid") or f"entry-{index}" + ) + self.context = replace( + self.context, + approval_scopes=approval_scopes, + approval_gate_count=max(self.context.approval_gate_count, approval_gate_count), + turns_used=max(self.context.turns_used, len(turns)), + cost_usd=self.prior_cost_usd + usage.cost_usd, + token_usage={ + key: self.prior_token_usage.get(key, 0) + usage.tokens.get(key, 0) + for key in TOKEN_FIELDS + }, + ) + receipt = await asyncio.to_thread(self._save, body, identity) + return identity, receipt + + def _save(self, body: bytes, identity: CheckpointIdentity) -> FileReceipt: + # Staging is outside the workspace so it cannot recursively archive + # itself. TemporaryDirectory is private and is removed on every exit. + with tempfile.TemporaryDirectory(prefix="abca-continuation-") as directory: + archive = Path(directory).resolve() / "workspace.tar" + capture_workspace( + Path(self.context.setup.repo_dir), + archive, + identity, + limits=WorkspaceLimits(max_bytes=self.storage.limits.max_workspace_bytes), + ) + workspace = self.storage.save_workspace(archive, identity) + conversation = self.storage.conversations.save(body, identity) + return self.storage.save_manifest( + ContinuationManifest(identity, conversation, workspace, self.context) + ) + + @classmethod + def restore( + cls, + storage: S3ContinuationStorage, + receipt: FileReceipt, + identity: CheckpointIdentity, + *, + expected_workspace: Path, + workflow_id: str, + workflow_version: str, + ) -> ContinuationRuntime: + manifest = storage.load_manifest(receipt, identity) + context = manifest.context + if ( + context.setup.repo_dir != str(expected_workspace) + or context.workflow_id != workflow_id + or context.workflow_version != workflow_version + ): + raise ContinuationCheckpointError("Continuation workspace or workflow changed") + body = storage.conversations.load(manifest.conversation, identity) + restored = decode_checkpoint(body, identity) + if restored["project_key"] != project_key_for_directory(str(expected_workspace)): + raise ContinuationCheckpointError( + "Continuation SDK project does not match the workspace" + ) + with tempfile.TemporaryDirectory(prefix="abca-restore-") as directory: + archive = Path(directory).resolve() / "workspace.tar" + storage.download_workspace(manifest.workspace, identity, archive) + restore_workspace( + archive, + expected_workspace, + identity, + expected_sha256=manifest.workspace.sha256, + limits=WorkspaceLimits(max_bytes=storage.limits.max_workspace_bytes), + ) + return cls( + context, + storage, + store=CheckpointSessionStore.restore(body, identity), + restored=restored, + ) + + def decision_prompt(self, *, decision: str, reason: str = "") -> str: + """Supply the saved action as context; every newly proposed tool is gated.""" + import json + + if self.restored is None or decision not in {"APPROVED", "DENIED", "TIMED_OUT"}: + raise ContinuationCheckpointError("A resolved continuation decision is required") + action = self.restored["action"] + return ( + "Continue the saved task from the restored workspace and conversation. " + "The previous worker stopped while waiting for a human decision; its pending " + "tool call was not executed. The recorded decision for that request is " + + decision + + ". Human feedback: " + + json.dumps(reason) + + ". Saved tool proposal: " + + json.dumps({"tool_name": action["tool_name"], "tool_input": action["tool_input"]}) + + ( + ". If you perform this approved proposal, use exactly the saved tool name " + "and every saved input field and value, including descriptions and timeouts. " + "Do not paraphrase the description or add optional arguments: any input change " + "is a new proposal and requires its own permission check" + if decision == "APPROVED" + else "" + ) + + ". Decide how to continue using this context. A denial must not be worked around. " + "Any new tool proposal still goes through the normal permission checks." + ) diff --git a/agent/src/continuation_session.py b/agent/src/continuation_session.py new file mode 100644 index 000000000..78d25dd86 --- /dev/null +++ b/agent/src/continuation_session.py @@ -0,0 +1,459 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Acknowledged SDK conversation checkpoints for worker continuation. + +This module does not release workers or change approval deadlines. A caller must +hold the lifecycle barrier, checkpoint the workspace, and conditionally publish +the returned receipt for the same task/attempt/request before releasing compute. +Only SDK transcript entries are captured; CLI authentication files and process +environment are never read. Transcript/action contents remain private task data. +""" + +from __future__ import annotations + +import asyncio +import base64 +import hashlib +import json +import math +import re +from dataclasses import asdict, dataclass +from typing import TYPE_CHECKING, Any +from uuid import UUID + +from claude_agent_sdk import SessionStore + +from shared_constants import SHARED_CONSTANTS + +if TYPE_CHECKING: + from claude_agent_sdk import SessionKey, SessionStoreEntry + +CHECKPOINT_VERSION = 1 +MAX_CHECKPOINT_BYTES = SHARED_CONSTANTS["microvm_continuation"]["max_conversation_bytes"] +MAX_CHECKPOINT_ENTRIES = 50_000 +MAX_ID_LENGTH = 128 +_ID = re.compile(r"[A-Za-z0-9][A-Za-z0-9_-]{0,127}\Z") +_HASH = re.compile(r"[0-9a-f]{64}\Z") + + +class ContinuationCheckpointError(RuntimeError): + """A checkpoint cannot be acknowledged or safely restored.""" + + def __init__(self, message: str, *, code: str = "checkpoint_failed") -> None: + super().__init__(message) + self.code = code + + +def _encode(value: Any) -> bytes: + try: + return json.dumps(value, sort_keys=True, separators=(",", ":"), allow_nan=False).encode() + except (TypeError, ValueError, RecursionError) as exc: + raise ContinuationCheckpointError( + "Checkpoint contains invalid JSON data", code="checkpoint_invalid_json" + ) from exc + + +def _copy(value: Any) -> Any: + return json.loads(_encode(value)) + + +def _identifier(value: Any) -> bool: + return isinstance(value, str) and _ID.fullmatch(value) is not None + + +def _session_id(value: Any) -> bool: + if not isinstance(value, str): + return False + try: + return str(UUID(value)) == value + except ValueError: + return False + + +def _action_hash(tool_input: dict) -> str: + # Match the approval row / RecentDecisionCache encoding, including spaces + # and ASCII escaping. Invalid input is rejected rather than stringified. + _encode(tool_input) + return hashlib.sha256(json.dumps(tool_input, sort_keys=True).encode()).hexdigest() + + +@dataclass(frozen=True) +class CheckpointIdentity: + """Version-1 wire identity: attempt_id is the source physical MicroVM ID. + + The coordinator's continuation.attempt_id is a separate logical launch token. + Renaming this serialized field requires a checkpoint-version migration. + """ + + task_id: str + attempt_id: str + request_id: str + user_id: str + repo: str + + def __post_init__(self) -> None: + if ( + not all( + _identifier(value) for value in (self.task_id, self.attempt_id, self.request_id) + ) + or not isinstance(self.user_id, str) + or not self.user_id + or len(self.user_id) > MAX_ID_LENGTH + or not isinstance(self.repo, str) + ): + raise ContinuationCheckpointError("Checkpoint identity is invalid") + + @property + def prefix(self) -> str: + prefix = SHARED_CONSTANTS["microvm_continuation"]["object_key_prefix"] + return f"{prefix}{self.task_id}/{self.attempt_id}/{self.request_id}/" + + +@dataclass(frozen=True) +class CheckpointReceipt: + """Pin a verified object version, never an overwriteable current key.""" + + key: str + version_id: str + sha256: str + size_bytes: int + + def validate(self, identity: CheckpointIdentity) -> None: + if ( + not isinstance(self.sha256, str) + or not _HASH.fullmatch(self.sha256) + or self.key != identity.prefix + self.sha256 + ".json" + or not isinstance(self.version_id, str) + or not self.version_id + or self.version_id == "null" + or type(self.size_bytes) is not int + or not 0 < self.size_bytes <= MAX_CHECKPOINT_BYTES + ): + raise ContinuationCheckpointError("Checkpoint receipt is invalid or outside this task") + + +def _validate_entries(entries: Any) -> list[dict]: + if not isinstance(entries, list) or len(entries) > MAX_CHECKPOINT_ENTRIES: + raise ContinuationCheckpointError("Checkpoint transcript entry limit exceeded") + for entry in entries: + if not isinstance(entry, dict) or not isinstance(entry.get("type"), str): + raise ContinuationCheckpointError("Checkpoint transcript entry is invalid") + if "uuid" in entry and (not isinstance(entry["uuid"], str) or not entry["uuid"]): + raise ContinuationCheckpointError("Checkpoint transcript entry UUID is invalid") + return entries + + +def _pending_action_present(entries: list[dict], action: dict) -> bool: + found = False + for entry in entries: + message = entry.get("message") + content = message.get("content") if isinstance(message, dict) else None + if not isinstance(content, list): + continue + for block in content: + if not isinstance(block, dict): + continue + if ( + block.get("type") == "tool_result" + and block.get("tool_use_id") == action["tool_use_id"] + ): + raise ContinuationCheckpointError("Pending action already has a tool result") + if block.get("type") == "tool_use" and block.get("id") == action["tool_use_id"]: + if found: + raise ContinuationCheckpointError("Pending action has duplicate transcript IDs") + if ( + entry.get("type") != "assistant" + or block.get("name") != action["tool_name"] + or _encode(block.get("input")) != _encode(action["tool_input"]) + ): + raise ContinuationCheckpointError( + "Pending action disagrees with the transcript" + ) + found = True + return found + + +def decode_checkpoint(body: bytes, identity: CheckpointIdentity) -> dict: + """Validate the complete private envelope before using any restored data.""" + if not body or len(body) > MAX_CHECKPOINT_BYTES: + raise ContinuationCheckpointError("Checkpoint byte limit exceeded") + try: + envelope = json.loads(body) + except (ValueError, UnicodeError, RecursionError) as exc: + raise ContinuationCheckpointError( + "Checkpoint is not valid JSON", code="checkpoint_invalid_json" + ) from exc + if ( + not isinstance(envelope, dict) + or set(envelope) + != {"version", "identity", "project_key", "session_id", "action", "entries"} + or type(envelope.get("version")) is not int + or envelope["version"] != CHECKPOINT_VERSION + or envelope.get("identity") != asdict(identity) + or not _session_id(envelope.get("session_id")) + or not isinstance(envelope.get("project_key"), str) + or not envelope["project_key"] + or not isinstance(envelope.get("action"), dict) + ): + raise ContinuationCheckpointError("Checkpoint envelope or identity is invalid") + action = envelope["action"] + if ( + set(action) != {"tool_use_id", "tool_name", "tool_input", "tool_input_sha256"} + or not _identifier(action.get("tool_use_id")) + or not _identifier(action.get("tool_name")) + or not isinstance(action.get("tool_input"), dict) + or action.get("tool_input_sha256") != _action_hash(action["tool_input"]) + ): + raise ContinuationCheckpointError("Checkpoint pending action is invalid") + entries = _validate_entries(envelope.get("entries")) + if any( + "sessionId" in entry and entry["sessionId"] != envelope["session_id"] for entry in entries + ): + raise ContinuationCheckpointError("Transcript contains another SDK session") + if not _pending_action_present(entries, action): + raise ContinuationCheckpointError("Pending action is missing from the transcript") + # Reject non-finite numbers even though Python's JSON decoder accepts them. + _encode(envelope) + return envelope + + +class CheckpointSessionStore(SessionStore): + """Single SDK session buffer implementing its public append/load protocol. + + Eager mirroring is still asynchronous. ``checkpoint_pending`` waits for the + exact assistant action to reach this buffer; enabling mirroring alone is not + an acknowledgement. Buffers belong to the SDK runner's event loop. + """ + + def __init__(self, project_key: str) -> None: + if not isinstance(project_key, str) or not project_key: + raise ContinuationCheckpointError("SDK project key is unavailable") + self.project_key = project_key + self._session: str | None = None + self._entries: list[dict] = [] + self._changed = asyncio.Condition() + self._failure: str | None = None + + def _check_key(self, key: SessionKey) -> str: + if ( + not isinstance(key, dict) + or key.get("project_key") != self.project_key + or not _session_id(key.get("session_id")) + or "subpath" in key + or (self._session is not None and key["session_id"] != self._session) + ): + raise ContinuationCheckpointError("Unexpected SDK session or subagent transcript") + return key["session_id"] + + async def append(self, key: SessionKey, entries: list[SessionStoreEntry]) -> None: + async with self._changed: + try: + session = self._check_key(key) + incoming = _validate_entries(_copy(entries)) + if any( + "sessionId" in entry and entry["sessionId"] != session for entry in incoming + ): + raise ContinuationCheckpointError("Transcript contains another SDK session") + updated = _copy(self._entries) + positions = {entry["uuid"]: i for i, entry in enumerate(updated) if "uuid" in entry} + for entry in incoming: + entry_id = entry.get("uuid") + if entry_id in positions: + updated[positions[entry_id]] = entry + else: + if entry_id is not None: + positions[entry_id] = len(updated) + updated.append(entry) + if ( + len(updated) > MAX_CHECKPOINT_ENTRIES + or len(_encode(updated)) > MAX_CHECKPOINT_BYTES + ): + raise ContinuationCheckpointError("SDK transcript exceeds checkpoint limits") + self._session, self._entries = session, updated + except ContinuationCheckpointError as exc: + # The SDK can continue after a dropped mirror batch. Once a gap + # is possible, no later matching action may certify completeness. + self._failure = str(exc) + raise + finally: + self._changed.notify_all() + + async def load(self, key: SessionKey) -> list[SessionStoreEntry] | None: + async with self._changed: + self._check_key(key) + if self._failure: + raise ContinuationCheckpointError(self._failure) + return _copy(self._entries) if self._session else None + + async def checkpoint_pending( + self, + identity: CheckpointIdentity, + *, + session_id: str, + tool_use_id: str, + tool_name: str, + tool_input: dict, + timeout_s: float = 5, + ) -> bytes: + """Require mirrored action coverage; workspace/barrier checks are separate.""" + if ( + not _session_id(session_id) + or not _identifier(tool_use_id) + or not _identifier(tool_name) + or not isinstance(tool_input, dict) + or isinstance(timeout_s, bool) + or not isinstance(timeout_s, (float, int)) + or not math.isfinite(timeout_s) + or timeout_s <= 0 + ): + raise ContinuationCheckpointError("Pending SDK action identity is invalid") + action = { + "tool_use_id": tool_use_id, + "tool_name": tool_name, + "tool_input": _copy(tool_input), + "tool_input_sha256": _action_hash(tool_input), + } + try: + async with asyncio.timeout(timeout_s), self._changed: + while True: + if self._failure: + raise ContinuationCheckpointError(self._failure) + if self._session and self._session != session_id: + raise ContinuationCheckpointError("Pending SDK session identity changed") + if _pending_action_present(self._entries, action): + body = _encode( + { + "version": CHECKPOINT_VERSION, + "identity": asdict(identity), + "project_key": self.project_key, + "session_id": session_id, + "action": action, + "entries": self._entries, + } + ) + decode_checkpoint(body, identity) + return body + await self._changed.wait() + except TimeoutError as exc: + raise ContinuationCheckpointError( + "SDK mirror did not acknowledge the pending action", code="checkpoint_sdk_timeout" + ) from exc + + @classmethod + def restore(cls, body: bytes, identity: CheckpointIdentity) -> CheckpointSessionStore: + envelope = decode_checkpoint(body, identity) + store = cls(envelope["project_key"]) + store._session = envelope["session_id"] + store._entries = envelope["entries"] + return store + + +class S3ContinuationCheckpoints: + """Store immutable, encrypted checkpoints under a task/attempt/request key. + + The bucket must have versioning enabled. The task role needs PutObject, + GetObject and GetObjectVersion for its own continuations// prefix. + This adapter never lists buckets, deletes data or falls back to ambient AWS + credentials. The caller must arrange retention after the owning request + closes; this adapter does not configure object expiration. + """ + + def __init__(self, bucket: str, *, client: Any = None) -> None: + if not isinstance(bucket, str) or not bucket or "/" in bucket: + raise ContinuationCheckpointError( + "Checkpoint bucket is unavailable", code="checkpoint_storage_unavailable" + ) + if client is None: + from botocore.config import Config + + from aws_session import is_scoped, tenant_client + + if not is_scoped(): + raise ContinuationCheckpointError( + "Checkpoint storage requires task-scoped credentials" + ) + client = tenant_client( + "s3", + config=Config(connect_timeout=2, read_timeout=5, retries={"total_max_attempts": 1}), + ) + self.bucket, self.client = bucket, client + + def _read(self, key: str, *, version_id: str | None = None) -> tuple[bytes, str]: + kwargs = {"Bucket": self.bucket, "Key": key, "ChecksumMode": "ENABLED"} + if version_id is not None: + kwargs["VersionId"] = version_id + response = self.client.get_object(**kwargs) + stream = response["Body"] + try: + size = response.get("ContentLength") + if type(size) is not int or not 0 < size <= MAX_CHECKPOINT_BYTES: + raise ContinuationCheckpointError("Stored checkpoint byte limit exceeded") + body = stream.read(MAX_CHECKPOINT_BYTES + 1) + if len(body) != size: + raise ContinuationCheckpointError("Stored checkpoint is incomplete") + finally: + stream.close() + actual_version = response.get("VersionId") + if ( + not isinstance(actual_version, str) + or not actual_version + or actual_version == "null" + or (version_id is not None and actual_version != version_id) + ): + raise ContinuationCheckpointError("Checkpoint storage requires an exact object version") + checksum = base64.b64encode(hashlib.sha256(body).digest()).decode() + if response.get("ChecksumSHA256") != checksum: + raise ContinuationCheckpointError( + "Stored checkpoint checksum is unavailable or incorrect" + ) + return body, actual_version + + def save(self, body: bytes, identity: CheckpointIdentity) -> CheckpointReceipt: + """Read back and verify even after an ambiguous write response.""" + decode_checkpoint(body, identity) + digest = hashlib.sha256(body).hexdigest() + key = identity.prefix + digest + ".json" + write_error: Exception | None = None + try: + self.client.put_object( + Bucket=self.bucket, + Key=key, + Body=body, + ContentType="application/json", + ServerSideEncryption="AES256", + ChecksumSHA256=base64.b64encode(hashlib.sha256(body).digest()).decode(), + IfNoneMatch="*", + ) + except Exception as exc: + # A timeout can follow a committed write; a repeated save can also + # return 412. Only exact read-back permits an acknowledgement. + write_error = exc + try: + stored, version = self._read(key) + if stored != body: + raise ContinuationCheckpointError( + "Stored checkpoint differs from the prepared data" + ) + except Exception as exc: + cause = write_error if write_error is not None else exc + raise ContinuationCheckpointError( + "Checkpoint could not be verified; keep the current worker available", + code="checkpoint_storage_unverified", + ) from cause + receipt = CheckpointReceipt(key, version, digest, len(body)) + receipt.validate(identity) + return receipt + + def load(self, receipt: CheckpointReceipt, identity: CheckpointIdentity) -> bytes: + receipt.validate(identity) + try: + body, _ = self._read(receipt.key, version_id=receipt.version_id) + except Exception as exc: + raise ContinuationCheckpointError( + "Saved checkpoint could not be read", code="checkpoint_storage_unavailable" + ) from exc + if len(body) != receipt.size_bytes or hashlib.sha256(body).hexdigest() != receipt.sha256: + raise ContinuationCheckpointError("Saved checkpoint does not match its receipt") + decode_checkpoint(body, identity) + return body diff --git a/agent/src/continuation_storage.py b/agent/src/continuation_storage.py new file mode 100644 index 000000000..b4bc4dc2e --- /dev/null +++ b/agent/src/continuation_storage.py @@ -0,0 +1,476 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Bounded, version-pinned storage for complete continuation checkpoints. + +The small manifest is the commit record: it is written only after the workspace +and SDK conversation have both been read back successfully. Publishing that +manifest to the task's conditional state record is a separate operation. An S3 +upload alone never grants permission to stop a worker. +""" + +from __future__ import annotations + +import base64 +import hashlib +import json +import os +import re +import shutil +import stat +import tempfile +import time +from dataclasses import asdict, dataclass, field +from decimal import Decimal +from pathlib import Path +from typing import Any, BinaryIO + +from continuation_session import ( + CheckpointIdentity, + CheckpointReceipt, + ContinuationCheckpointError, + S3ContinuationCheckpoints, +) +from models import RepoSetup +from shared_constants import SHARED_CONSTANTS + +_CHUNK = 1024 * 1024 +_MAX_MANIFEST_BYTES = SHARED_CONSTANTS["microvm_continuation"]["max_manifest_bytes"] +_SHA256 = re.compile(r"[0-9a-f]{64}\Z") +_KINDS = {"workspace": "tar", "manifest": "json"} +_VERSION = SHARED_CONSTANTS["microvm_continuation"]["version"] +_MAX_VERSION_ID_LENGTH = 1024 + + +class ContinuationStorageError(ContinuationCheckpointError): + """Content-free failure classification for lifecycle diagnostics.""" + + def __init__(self, code: str, message: str) -> None: + super().__init__(message, code=code) + + +@dataclass(frozen=True) +class StorageLimits: + max_workspace_bytes: int = SHARED_CONSTANTS["microvm_continuation"]["max_workspace_bytes"] + disk_reserve_bytes: int = 128 * _CHUNK + transfer_timeout_s: int = 300 + + def __post_init__(self) -> None: + if any( + type(value) is not int or value <= 0 + for value in ( + self.max_workspace_bytes, + self.disk_reserve_bytes, + self.transfer_timeout_s, + ) + ): + raise ValueError("Continuation storage limits must be positive integers") + + +_DEFAULT_LIMITS = StorageLimits() + + +@dataclass(frozen=True) +class FileReceipt: + kind: str + key: str + version_id: str + sha256: str + size_bytes: int + + @classmethod + def from_record( + cls, value: Any, identity: CheckpointIdentity, limits: StorageLimits = _DEFAULT_LIMITS + ) -> FileReceipt: + """Accept DynamoDB's integral Decimal without coercing strings/floats/bools.""" + if not isinstance(value, dict) or set(value) != { + "kind", + "key", + "version_id", + "sha256", + "size_bytes", + }: + raise ContinuationStorageError( + "invalid_receipt", "Continuation file receipt is invalid" + ) + data = dict(value) + size = data["size_bytes"] + if isinstance(size, Decimal) and size.is_finite() and size == size.to_integral_value(): + data["size_bytes"] = int(size) + receipt = cls(**data) + receipt.validate(identity, limits) + return receipt + + def validate(self, identity: CheckpointIdentity, limits: StorageLimits) -> None: + suffix = _KINDS.get(self.kind) + bound = limits.max_workspace_bytes if self.kind == "workspace" else _MAX_MANIFEST_BYTES + if ( + suffix is None + or not isinstance(self.sha256, str) + or not _SHA256.fullmatch(self.sha256) + or self.key != f"{identity.prefix}{self.kind}/{self.sha256}.{suffix}" + or not isinstance(self.version_id, str) + or not self.version_id + or self.version_id == "null" + or len(self.version_id) > _MAX_VERSION_ID_LENGTH + or type(self.size_bytes) is not int + or not 0 < self.size_bytes <= bound + ): + raise ContinuationStorageError( + "invalid_receipt", "Continuation file receipt is invalid" + ) + + +@dataclass(frozen=True) +class ContinuationContext: + """Workflow products needed after recovery; never serialize TaskConfig secrets.""" + + setup: RepoSetup + user_prompt: str + system_prompt: str + workflow_id: str + workflow_version: str + approval_scopes: tuple[str, ...] = () + approval_gate_count: int = 0 + turns_used: int = 0 + cost_usd: float = 0.0 + token_usage: dict[str, int] = field(default_factory=dict) + started_reaction_id: str | None = None + + def to_dict(self) -> dict: + value = asdict(self) + value["setup"] = self.setup.model_dump(mode="json") + value["approval_scopes"] = list(self.approval_scopes) + return value + + @classmethod + def from_dict(cls, value: Any) -> ContinuationContext: + from continuation_usage import valid_cost, valid_tokens + + text_fields = {"user_prompt", "system_prompt", "workflow_id", "workflow_version"} + fields = text_fields | { + "setup", + "approval_scopes", + "approval_gate_count", + "turns_used", + "cost_usd", + "token_usage", + "started_reaction_id", + } + if ( + not isinstance(value, dict) + or set(value) != fields + or any(not isinstance(value[name], str) for name in text_fields) + or not value["workflow_id"] + or not value["workflow_version"] + or not isinstance(value["setup"], dict) + or set(value["setup"]) != set(RepoSetup.model_fields) + or not isinstance(value["approval_scopes"], list) + or any(not isinstance(scope, str) for scope in value["approval_scopes"]) + or any( + type(value[name]) is not int or value[name] < 0 + for name in ("approval_gate_count", "turns_used") + ) + or not valid_cost(value["cost_usd"]) + or not valid_tokens(value["token_usage"]) + or ( + value["started_reaction_id"] is not None + and not isinstance(value["started_reaction_id"], str) + ) + ): + raise ContinuationStorageError( + "invalid_context", "Continuation workflow context is invalid" + ) + try: + setup = RepoSetup.model_validate(value["setup"], strict=True) + except ValueError as exc: + raise ContinuationStorageError( + "invalid_context", "Continuation repository context is invalid" + ) from exc + if not Path(setup.repo_dir).is_absolute(): + raise ContinuationStorageError( + "invalid_context", "Continuation workspace must be absolute" + ) + return cls( + setup, + value["user_prompt"], + value["system_prompt"], + value["workflow_id"], + value["workflow_version"], + tuple(value["approval_scopes"]), + value["approval_gate_count"], + value["turns_used"], + float(value["cost_usd"]), + value["token_usage"], + value["started_reaction_id"], + ) + + +@dataclass(frozen=True) +class ContinuationManifest: + identity: CheckpointIdentity + conversation: CheckpointReceipt + workspace: FileReceipt + context: ContinuationContext + + def encode(self, limits: StorageLimits) -> bytes: + self.conversation.validate(self.identity) + self.workspace.validate(self.identity, limits) + if self.workspace.kind != "workspace": + raise ContinuationStorageError("invalid_manifest", "Manifest workspace kind is invalid") + context = self.context.to_dict() + ContinuationContext.from_dict(context) + try: + body = json.dumps( + { + "version": _VERSION, + "identity": asdict(self.identity), + "conversation": asdict(self.conversation), + "workspace": asdict(self.workspace), + "context": context, + }, + sort_keys=True, + separators=(",", ":"), + allow_nan=False, + ).encode() + except (TypeError, ValueError, RecursionError) as exc: + raise ContinuationStorageError("invalid_manifest", "Manifest data is invalid") from exc + if len(body) > _MAX_MANIFEST_BYTES: + raise ContinuationStorageError( + "size_limit", "Continuation manifest exceeds its byte limit" + ) + return body + + @classmethod + def decode( + cls, body: bytes, identity: CheckpointIdentity, limits: StorageLimits + ) -> ContinuationManifest: + if not body or len(body) > _MAX_MANIFEST_BYTES: + raise ContinuationStorageError( + "size_limit", "Continuation manifest exceeds its byte limit" + ) + try: + value = json.loads(body) + if ( + not isinstance(value, dict) + or set(value) != {"version", "identity", "conversation", "workspace", "context"} + or type(value["version"]) is not int + or value["version"] != _VERSION + or value["identity"] != asdict(identity) + ): + raise ValueError("Invalid envelope") + result = cls( + identity, + CheckpointReceipt(**value["conversation"]), + FileReceipt(**value["workspace"]), + ContinuationContext.from_dict(value["context"]), + ) + result.encode(limits) + return result + except (TypeError, ValueError, KeyError, RecursionError) as exc: + raise ContinuationStorageError( + "invalid_manifest", "Continuation manifest is invalid" + ) from exc + + +def _stamp(info: os.stat_result) -> tuple[int, ...]: + return (info.st_dev, info.st_ino, info.st_size, info.st_mtime_ns, info.st_ctime_ns) + + +def _time_left(deadline: float) -> None: + if time.monotonic() >= deadline: + raise ContinuationStorageError( + "transfer_timeout", "Continuation transfer exceeded its deadline" + ) + + +def _space(path: Path, required: int, limits: StorageLimits) -> None: + if shutil.disk_usage(path).free < required + limits.disk_reserve_bytes: + raise ContinuationStorageError( + "disk_pressure", "Continuation transfer has insufficient disk space" + ) + + +class S3ContinuationStorage: + def __init__( + self, bucket: str, *, client: Any = None, limits: StorageLimits = _DEFAULT_LIMITS + ) -> None: + # Reuse the mandatory task-scoped credential check and conservative SDK + # timeouts. An explicitly injected client is only for tests/operators. + self.conversations = S3ContinuationCheckpoints(bucket, client=client) + self.bucket = self.conversations.bucket + self.client = self.conversations.client + self.limits = limits + + def _read( + self, + receipt: FileReceipt, + identity: CheckpointIdentity, + destination: BinaryIO | None, + *, + deadline: float, + current: bool = False, + disk_path: Path | None = None, + ) -> FileReceipt: + receipt.validate(identity, self.limits) + _time_left(deadline) + response = self.client.get_object( + Bucket=self.bucket, + Key=receipt.key, + ChecksumMode="ENABLED", + **({} if current else {"VersionId": receipt.version_id}), + ) + stream = response["Body"] + try: + version = response.get("VersionId") + if ( + not isinstance(version, str) + or not version + or version == "null" + or (not current and version != receipt.version_id) + or type(response.get("ContentLength")) is not int + or response["ContentLength"] != receipt.size_bytes + or response.get("ChecksumSHA256") + != base64.b64encode(bytes.fromhex(receipt.sha256)).decode() + ): + raise ContinuationStorageError( + "integrity_failed", "Stored continuation metadata differs" + ) + digest = hashlib.sha256() + count = 0 + while True: + _time_left(deadline) + chunk = stream.read(min(_CHUNK, receipt.size_bytes - count + 1)) + if not chunk: + break + count += len(chunk) + if count > receipt.size_bytes: + raise ContinuationStorageError( + "size_limit", "Stored continuation exceeds its receipt" + ) + digest.update(chunk) + if destination is not None: + if disk_path is not None: + _space(disk_path, len(chunk), self.limits) + destination.write(chunk) + if count != receipt.size_bytes or digest.hexdigest() != receipt.sha256: + raise ContinuationStorageError( + "integrity_failed", "Stored continuation content differs" + ) + verified = FileReceipt(receipt.kind, receipt.key, version, receipt.sha256, count) + verified.validate(identity, self.limits) + return verified + finally: + stream.close() + + def _save( + self, source: BinaryIO, identity: CheckpointIdentity, kind: str, size: int + ) -> FileReceipt: + deadline = time.monotonic() + self.limits.transfer_timeout_s + digest = hashlib.sha256() + count = 0 + while chunk := source.read(_CHUNK): + _time_left(deadline) + count += len(chunk) + if count > size: + raise ContinuationStorageError( + "source_changed", "Prepared continuation changed size" + ) + digest.update(chunk) + if count != size: + raise ContinuationStorageError("source_changed", "Prepared continuation is incomplete") + sha256 = digest.hexdigest() + key = f"{identity.prefix}{kind}/{sha256}.{_KINDS[kind]}" + # The provisional version is never returned: _read must supply and + # validate a real S3 version, including after a lost PutObject reply. + provisional = FileReceipt(kind, key, "unverified", sha256, size) + provisional.validate(identity, self.limits) + source.seek(0) + write_error = None + try: + _time_left(deadline) + self.client.put_object( + Bucket=self.bucket, + Key=key, + Body=source, + ContentLength=size, + ContentType="application/x-tar" if kind == "workspace" else "application/json", + ServerSideEncryption="AES256", + ChecksumSHA256=base64.b64encode(digest.digest()).decode(), + IfNoneMatch="*", + ) + except Exception as exc: + write_error = exc + try: + return self._read(provisional, identity, None, deadline=deadline, current=True) + except Exception as exc: + raise ContinuationStorageError( + "unverified_upload", "Continuation upload could not be verified; retain the worker" + ) from (write_error if write_error is not None else exc) + + def save_workspace(self, archive: Path, identity: CheckpointIdentity) -> FileReceipt: + with os.fdopen( + os.open(archive, os.O_RDONLY | os.O_NOFOLLOW | os.O_NONBLOCK), "rb" + ) as source: + before = os.fstat(source.fileno()) + if ( + not stat.S_ISREG(before.st_mode) + or before.st_nlink != 1 + or not 0 < before.st_size <= self.limits.max_workspace_bytes + ): + raise ContinuationStorageError( + "invalid_source", "Prepared workspace archive is invalid" + ) + receipt = self._save(source, identity, "workspace", before.st_size) + if _stamp(os.fstat(source.fileno())) != _stamp(before): + raise ContinuationStorageError( + "source_changed", "Workspace archive changed during upload" + ) + return receipt + + def save_manifest(self, manifest: ContinuationManifest) -> FileReceipt: + import io + + body = manifest.encode(self.limits) + return self._save(io.BytesIO(body), manifest.identity, "manifest", len(body)) + + def load_manifest( + self, receipt: FileReceipt, identity: CheckpointIdentity + ) -> ContinuationManifest: + import io + + if receipt.kind != "manifest": + raise ContinuationStorageError("invalid_receipt", "Expected a continuation manifest") + target = io.BytesIO() + self._read( + receipt, + identity, + target, + deadline=time.monotonic() + self.limits.transfer_timeout_s, + ) + return ContinuationManifest.decode(target.getvalue(), identity, self.limits) + + def download_workspace( + self, receipt: FileReceipt, identity: CheckpointIdentity, destination: Path + ) -> None: + if receipt.kind != "workspace": + raise ContinuationStorageError("invalid_receipt", "Expected a workspace archive") + receipt.validate(identity, self.limits) + _space(destination.parent, receipt.size_bytes, self.limits) + # The caller supplies a private staging directory. Never truncate an + # existing destination; publish the verified file with a no-replace link. + fd, name = tempfile.mkstemp(prefix=".continuation-", dir=destination.parent) + try: + with os.fdopen(fd, "wb") as target: + self._read( + receipt, + identity, + target, + deadline=time.monotonic() + self.limits.transfer_timeout_s, + disk_path=destination.parent, + ) + target.flush() + os.fsync(target.fileno()) + os.link(name, destination) + finally: + os.unlink(name) diff --git a/agent/src/continuation_usage.py b/agent/src/continuation_usage.py new file mode 100644 index 000000000..cab40900e --- /dev/null +++ b/agent/src/continuation_usage.py @@ -0,0 +1,94 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Version-checked access to exact CLI spending at a paused approval hook.""" + +from __future__ import annotations + +import asyncio +import importlib.metadata +import math +from dataclasses import dataclass +from typing import Any + +from continuation_session import ContinuationCheckpointError +from shared_constants import SHARED_CONSTANTS + +TOKEN_FIELDS = { + "input_tokens": "inputTokens", + "output_tokens": "outputTokens", + "cache_read_input_tokens": "cacheReadInputTokens", + "cache_creation_input_tokens": "cacheCreationInputTokens", +} + + +@dataclass(frozen=True) +class UsageSnapshot: + cost_usd: float + tokens: dict[str, int] + + +def valid_cost(value: Any) -> bool: + return type(value) in (int, float) and math.isfinite(value) and value >= 0 + + +def valid_tokens(value: Any) -> bool: + return ( + isinstance(value, dict) + and set(value) <= set(TOKEN_FIELDS) + and all(type(count) is int and count >= 0 for count in value.values()) + ) + + +async def read_usage(client: Any) -> UsageSnapshot: + """Read cumulative spending without interrupting or issuing another query. + + SDK 0.2.110 bundles CLI 2.1.191, whose experimental ``get_usage`` control + request exposes exact session dollars and per-model token counters. Python + has no public wrapper yet. Keep this dependency isolated and covered by the + real SDK probe in CI; an upgrade must verify it before publishing checkpoints. + """ + if ( + importlib.metadata.version("claude-agent-sdk") + != SHARED_CONSTANTS["microvm_continuation"]["verified_sdk_version"] + ): + raise ContinuationCheckpointError( + "Continuation accounting requires the verified SDK version", + code="checkpoint_sdk_unverified", + ) + query = getattr(client, "_query", None) + send = getattr(query, "_send_control_request", None) + if not callable(send): + raise ContinuationCheckpointError( + "Continuation accounting client is unavailable", code="checkpoint_sdk_unverified" + ) + try: + response = await asyncio.wait_for(send({"subtype": "get_usage"}), timeout=5) + except TimeoutError as exc: + raise ContinuationCheckpointError( + "Continuation accounting request timed out", code="checkpoint_sdk_timeout" + ) from exc + session = response.get("session") if isinstance(response, dict) else None + if not isinstance(session, dict) or not valid_cost(session.get("total_cost_usd")): + raise ContinuationCheckpointError("Continuation accounting response has no valid cost") + models = session.get("model_usage") + if not isinstance(models, dict): + raise ContinuationCheckpointError("Continuation accounting response has no model usage") + tokens = dict.fromkeys(TOKEN_FIELDS, 0) + model_cost = 0.0 + for usage in models.values(): + if ( + not isinstance(usage, dict) + or not valid_cost(usage.get("costUSD")) + or any( + type(usage.get(key)) is not int or usage[key] < 0 for key in TOKEN_FIELDS.values() + ) + ): + raise ContinuationCheckpointError("Continuation model usage is invalid") + model_cost += usage["costUSD"] + for target, source in TOKEN_FIELDS.items(): + tokens[target] += usage[source] + cost = float(session["total_cost_usd"]) + if not math.isclose(model_cost, cost, rel_tol=1e-9, abs_tol=1e-12): + raise ContinuationCheckpointError("Continuation model costs do not match session spending") + return UsageSnapshot(cost, tokens) diff --git a/agent/src/continuation_workspace.py b/agent/src/continuation_workspace.py new file mode 100644 index 000000000..d86614e5f --- /dev/null +++ b/agent/src/continuation_workspace.py @@ -0,0 +1,780 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Offline workspace archives for worker continuation. + +The caller must stop workspace writers with the lifecycle barrier. Capture is +local only: durable upload/read-back and conditional publication alongside the +conversation receipt are separate requirements before releasing a worker. + +Preserve all working files, including ignored files, rather than guessing which +can be regenerated. Rebuild Git administration from a bundle and staged patch; +never copy the original Git config, hooks, home directory or process environment. +Repository contents themselves remain private and may contain sensitive data. +""" + +from __future__ import annotations + +import base64 +import hashlib +import io +import json +import os +import re +import selectors +import shutil +import signal +import stat +import subprocess +import tarfile +import tempfile +import time +from dataclasses import asdict, dataclass +from pathlib import Path, PurePosixPath +from typing import BinaryIO, NoReturn + +from continuation_session import CheckpointIdentity, ContinuationCheckpointError +from shared_constants import SHARED_CONSTANTS + +_OID = re.compile(r"[0-9a-f]{40}\Z") +_SHA256 = re.compile(r"[0-9a-f]{64}\Z") +_REPO = re.compile(r"[A-Za-z0-9_.-]+/[A-Za-z0-9_.-]+\Z") +_CHUNK = 1024 * 1024 +_ADMIN = {"git.bundle", "index.patch", "manifest.json"} +_MAX_PATH_BYTES = 4096 +_PERMISSION_BITS = 0o777 +_MAX_EXCLUDE_BYTES = 1024 * 1024 + + +class WorkspaceCheckpointError(ContinuationCheckpointError): + """A content-free stage code for lifecycle feedback.""" + + def __init__(self, code: str, message: str) -> None: + super().__init__(message, code=code) + + +@dataclass(frozen=True) +class WorkspaceLimits: + max_bytes: int = SHARED_CONSTANTS["microvm_continuation"]["max_workspace_bytes"] + max_entries: int = 100_000 + git_timeout_s: int = 60 + + def __post_init__(self) -> None: + if any( + type(value) is not int or value <= 0 + for value in (self.max_bytes, self.max_entries, self.git_timeout_s) + ): + raise ValueError("Workspace limits must be positive integers") + + +@dataclass(frozen=True) +class WorkspaceArchive: + sha256: str + size_bytes: int + head: str + branch: str | None + entries: int + + +_DEFAULT_LIMITS = WorkspaceLimits() + + +def _fail(code: str, message: str) -> NoReturn: + raise WorkspaceCheckpointError(code, message) + + +def _git( + root: Path, + args: list[str], + limits: WorkspaceLimits, + *, + output: BinaryIO | None = None, + max_bytes: int | None = None, + source: BinaryIO | None = None, +) -> bytes: + """Bound Git output/time, disable hooks/helpers, and never log repository data.""" + env = {key: value for key, value in os.environ.items() if not key.startswith("GIT_")} + env.update( + GIT_CONFIG_NOSYSTEM="1", + GIT_CONFIG_GLOBAL=os.devnull, + GIT_ATTR_NOSYSTEM="1", + GIT_NO_REPLACE_OBJECTS="1", + GIT_TERMINAL_PROMPT="0", + GIT_OPTIONAL_LOCKS="0", + LC_ALL="C", + ) + command = [ + "git", + "-c", + f"core.hooksPath={os.devnull}", + "-c", + "core.fsmonitor=false", + "-c", + "gc.auto=0", + "-c", + "protocol.allow=never", + "-c", + f"safe.directory={root}", + *args, + ] + target = output if output is not None else io.BytesIO() + bound = max_bytes if max_bytes is not None else min(limits.max_bytes, 16 * 1024 * 1024) + count = 0 + # stderr may contain repository contents/configuration; only the stage and + # exit code cross this boundary. The caller retains its original workspace. + with ( + subprocess.Popen( + command, + cwd=root, + env=env, + stdin=source if source is not None else subprocess.DEVNULL, + stdout=subprocess.PIPE, + stderr=subprocess.DEVNULL, + start_new_session=True, + ) as process, + selectors.DefaultSelector() as selector, + ): + if process.stdout is None: + _fail("git_failed", "Workspace Git output pipe is unavailable") + selector.register(process.stdout, selectors.EVENT_READ) + deadline = time.monotonic() + limits.git_timeout_s + try: + while selector.get_map(): + remaining = deadline - time.monotonic() + if remaining <= 0: + _fail("git_timeout", f"Workspace Git {args[0]} timed out") + for key, _ in selector.select(min(remaining, 0.1)): + block = os.read(key.fd, _CHUNK) + if not block: + selector.unregister(key.fd) + continue + count += len(block) + if count > bound: + _fail("size_limit", f"Workspace Git {args[0]} exceeds the byte limit") + target.write(block) + try: + code = process.wait(timeout=max(0.01, deadline - time.monotonic())) + except subprocess.TimeoutExpired as exc: + raise WorkspaceCheckpointError("git_timeout", "Workspace Git timed out") from exc + if code: + _fail("git_failed", f"Workspace Git {args[0]} failed (exit {code})") + finally: + if process.poll() is None: + process.stdout.close() + try: + process.wait(timeout=0.1) + except subprocess.TimeoutExpired: + try: + os.killpg(process.pid, signal.SIGKILL) + except (ProcessLookupError, PermissionError): + # Git may exit between wait and killpg. An error is + # ignorable only after wait confirms that it exited. + process.wait(timeout=0.1) + process.wait() + return target.getvalue() if isinstance(target, io.BytesIO) else b"" + + +def _path(value: str) -> PurePosixPath: + if not isinstance(value, str) or not value or len(value.encode()) > _MAX_PATH_BYTES: + _fail("invalid_path", "Workspace archive path is invalid") + path = PurePosixPath(value) + if ( + path.is_absolute() + or str(path) != value + or "\\" in value + or "\0" in value + or any(part in {".", ".."} or part.lower() == ".git" for part in path.parts) + ): + _fail("invalid_path", "Workspace archive path escapes the working tree") + return path + + +def _stamp(info: os.stat_result) -> tuple[int, ...]: + return ( + info.st_mode, + info.st_dev, + info.st_ino, + info.st_size, + info.st_mtime_ns, + info.st_ctime_ns, + info.st_nlink, + ) + + +def _scan(root: Path, limits: WorkspaceLimits) -> dict[str, tuple[int, ...]]: + entries = {} + device = root.stat().st_dev + for directory, dirs, files, fd in os.fwalk(root, follow_symlinks=False): + relative = Path(directory).relative_to(root) + if relative == Path("."): + dirs[:] = [name for name in dirs if name != ".git"] + for name in sorted(dirs + files): + path = str(relative / name) + _path(path) + info = os.stat(name, dir_fd=fd, follow_symlinks=False) + if ( + not ( + stat.S_ISREG(info.st_mode) + or stat.S_ISDIR(info.st_mode) + or stat.S_ISLNK(info.st_mode) + ) + or info.st_dev != device + or (stat.S_ISREG(info.st_mode) and info.st_nlink != 1) + or info.st_mode & (stat.S_ISUID | stat.S_ISGID | stat.S_ISVTX) + ): + _fail("unsupported_file", "Workspace has a special, linked or mounted file") + entries[path] = _stamp(info) + if len(entries) > limits.max_entries: + _fail("entry_limit", "Workspace exceeds the entry limit") + return dict(sorted(entries.items())) + + +def _open_file(root: Path, path: str) -> BinaryIO: + """Open through directory descriptors; never traverse a substituted symlink.""" + parts = _path(path).parts + fd = os.open(root, os.O_RDONLY | os.O_DIRECTORY | os.O_NOFOLLOW) + try: + for part in parts[:-1]: + child = os.open(part, os.O_RDONLY | os.O_DIRECTORY | os.O_NOFOLLOW, dir_fd=fd) + os.close(fd) + fd = child + result = os.open(parts[-1], os.O_RDONLY | os.O_NOFOLLOW | os.O_NONBLOCK, dir_fd=fd) + return os.fdopen(result, "rb") + finally: + os.close(fd) + + +def _git_state(root: Path, limits: WorkspaceLimits) -> dict: + gitdir = root / ".git" + if gitdir.is_symlink() or not gitdir.is_dir(): + _fail("unsupported_git", "Workspace requires a standalone Git checkout") + for name in ( + "MERGE_HEAD", + "CHERRY_PICK_HEAD", + "REVERT_HEAD", + "rebase-apply", + "rebase-merge", + "sequencer", + "index.lock", + "shallow", + "objects/info/alternates", + "info/grafts", + "info/sparse-checkout", + ): + if os.path.lexists(gitdir / name): + _fail("unsupported_git", "Workspace has an active or unsupported Git operation") + if ( + _git(root, ["rev-parse", "--show-toplevel"], limits).decode().strip() != str(root) + or Path(_git(root, ["rev-parse", "--absolute-git-dir"], limits).decode().strip()) != gitdir + or _git(root, ["rev-parse", "--show-object-format"], limits).strip() != b"sha1" + ): + _fail("unsupported_git", "Workspace Git layout or object format is unsupported") + head = _git(root, ["rev-parse", "HEAD"], limits).decode().strip() + branch_value = _git(root, ["rev-parse", "--abbrev-ref", "HEAD"], limits).decode().strip() + branch = None if branch_value == "HEAD" else branch_value + refs = {} + for line in _git( + root, ["for-each-ref", "--format=%(objectname) %(refname)"], limits + ).splitlines(): + oid, ref = line.decode().split(" ", 1) + if ref.startswith("refs/replace/"): + _fail("unsupported_git", "Workspace replacement refs cannot be checkpointed") + refs[ref] = oid + index = _git(root, ["ls-files", "--stage", "-z"], limits) + for line in index.split(b"\0"): + if not line: + continue + details, path = line.split(b"\t", 1) + mode, _, stage = details.split() + _path(path.decode()) + if mode == b"160000" or stage != b"0": + _fail( + "unsupported_git", + "Workspace submodules or unresolved index entries are unsupported", + ) + flags = _git(root, ["ls-files", "-v", "-z"], limits).split(b"\0") + if any(entry and entry[:1] != b"H" for entry in flags): + _fail("unsupported_git", "Workspace has sparse or assume-unchanged index flags") + visible = _git( + root, + ["diff", "--cached", "--raw", "--no-ext-diff", "--no-textconv", "--ita-visible-in-index"], + limits, + ) + invisible = _git( + root, + ["diff", "--cached", "--raw", "--no-ext-diff", "--no-textconv", "--ita-invisible-in-index"], + limits, + ) + if visible != invisible: + _fail("unsupported_git", "Workspace intent-to-add entries cannot be checkpointed") + exclude_b64 = None + if os.path.lexists(gitdir / "info/exclude"): + with _open_file(gitdir, "info/exclude") as source: + if not stat.S_ISREG(os.fstat(source.fileno()).st_mode): + _fail("unsupported_git", "Workspace Git ignore file must be regular") + exclude = source.read(_MAX_EXCLUDE_BYTES + 1) + if len(exclude) > _MAX_EXCLUDE_BYTES: + _fail("size_limit", "Workspace Git ignore file exceeds the byte limit") + exclude_b64 = base64.b64encode(exclude).decode() + return { + "head": head, + "branch": branch, + "refs": refs, + "index_sha256": hashlib.sha256(index).hexdigest(), + "exclude_b64": exclude_b64, + } + + +class _DigestReader: + def __init__(self, stream: BinaryIO) -> None: + self.stream = stream + self.digest = hashlib.sha256() + + def read(self, size: int = -1) -> bytes: + block = self.stream.read(size) + self.digest.update(block) + return block + + +def _file_digest(path: Path) -> str: + with path.open("rb") as stream: + return hashlib.file_digest(stream, "sha256").hexdigest() + + +def _repository_identity(identity: CheckpointIdentity) -> None: + if identity.repo == "": + return # Repository-free tasks use a private Git baseline with no remote. + if not _REPO.fullmatch(identity.repo) or any( + part in {".", ".."} for part in identity.repo.split("/") + ): + _fail("invalid_identity", "Workspace repository identity is invalid") + + +def capture_workspace( + workspace: Path, + destination: Path, + identity: CheckpointIdentity, + *, + limits: WorkspaceLimits = _DEFAULT_LIMITS, +) -> WorkspaceArchive: + """Create an owned local archive; never overwrite a previously saved file.""" + try: + return _capture_workspace(workspace, destination, identity, limits) + except WorkspaceCheckpointError: + raise + except ( + OSError, + ValueError, + TypeError, + RecursionError, + subprocess.SubprocessError, + tarfile.TarError, + ) as exc: + raise WorkspaceCheckpointError( + "capture_failed", "Workspace capture failed; keep the worker available" + ) from exc + + +def _capture_workspace( + workspace: Path, destination: Path, identity: CheckpointIdentity, limits: WorkspaceLimits +) -> WorkspaceArchive: + _repository_identity(identity) + root = workspace.absolute() + destination = destination.absolute() + if root.resolve() != root or not root.is_dir(): + _fail("invalid_path", "Workspace must be an existing canonical directory") + if destination.is_relative_to(root) or destination.parent.resolve() != destination.parent: + _fail("invalid_path", "Workspace archive must be outside the working tree") + if os.path.lexists(destination): + _fail("destination_exists", "Workspace archive destination already exists") + original = _git_state(root, limits) + before = _scan(root, limits) + with tempfile.TemporaryDirectory( + prefix=".workspace-capture-", dir=destination.parent + ) as scratch: + temp = Path(scratch) + bundle, patch = temp / "git.bundle", temp / "index.patch" + with bundle.open("wb") as output: + _git( + root, + ["bundle", "create", "-", "--all", "HEAD"], + limits, + output=output, + max_bytes=limits.max_bytes, + ) + with patch.open("wb") as output: + _git( + root, + [ + "diff", + "--cached", + "--binary", + "--full-index", + "--no-ext-diff", + "--no-textconv", + "HEAD", + "--", + ], + limits, + output=output, + max_bytes=limits.max_bytes - bundle.stat().st_size, + ) + records = [] + total = bundle.stat().st_size + patch.stat().st_size + archive = temp / "workspace.tar" + with tarfile.open(archive, "w", format=tarfile.PAX_FORMAT) as tar: + for path, stamp in before.items(): + mode, _, _, size, *_ = stamp + member = tarfile.TarInfo("files/" + path) + member.mode = stat.S_IMODE(mode) + record = {"path": path, "mode": member.mode} + if stat.S_ISDIR(mode): + member.type = tarfile.DIRTYPE + record["kind"] = "directory" + tar.addfile(member) + elif stat.S_ISLNK(mode): + member.type = tarfile.SYMTYPE + member.linkname = os.readlink(root / path) + record.update(kind="symlink", target=member.linkname) + tar.addfile(member) + else: + total += size + if total > limits.max_bytes: + _fail("size_limit", "Workspace exceeds the byte limit") + member.size = size + with _open_file(root, path) as stream: + if _stamp(os.fstat(stream.fileno())) != stamp: + _fail("workspace_changed", "Workspace changed during capture") + reader = _DigestReader(stream) + tar.addfile(member, reader) + if _stamp(os.fstat(stream.fileno())) != stamp: + _fail("workspace_changed", "Workspace changed during capture") + record.update(kind="file", size=size, sha256=reader.digest.hexdigest()) + records.append(record) + metadata = { + "version": 1, + "identity": asdict(identity), + "workspace": str(root), + "git": original, + "files": records, + "bundle_sha256": _file_digest(bundle), + "patch_sha256": _file_digest(patch), + } + for path in (bundle, patch): + member = tarfile.TarInfo(path.name) + member.mode, member.size = 0o600, path.stat().st_size + with path.open("rb") as stream: + tar.addfile(member, stream) + body = json.dumps(metadata, sort_keys=True, separators=(",", ":")).encode() + member = tarfile.TarInfo("manifest.json") + member.mode, member.size = 0o600, len(body) + tar.addfile(member, io.BytesIO(body)) + if archive.stat().st_size > limits.max_bytes: + _fail("size_limit", "Workspace archive exceeds the byte limit") + if _scan(root, limits) != before or _git_state(root, limits) != original: + _fail("workspace_changed", "Workspace or Git state changed during capture") + with tarfile.open(archive, "r:") as tar: + _validated_manifest(tar, identity, root, limits) + digest = _file_digest(archive) + size = archive.stat().st_size + # A hard link publishes the completed file without replacing a racing + # writer's destination. The temporary name is removed by its owned scope. + archive.chmod(0o600) + os.link(archive, destination) + return WorkspaceArchive(digest, size, original["head"], original["branch"], len(records)) + + +def _validated_manifest( + tar: tarfile.TarFile, identity: CheckpointIdentity, root: Path, limits: WorkspaceLimits +) -> tuple[dict, dict[str, tarfile.TarInfo]]: + members = {} + for member in tar: + name = member.name.rstrip("/") if member.isdir() else member.name + if ( + name in members + or len(members) >= limits.max_entries + len(_ADMIN) + or member.size < 0 + or member.size > limits.max_bytes + or member.mode < 0 + or member.mode > _PERMISSION_BITS + or member.sparse is not None + or set(member.pax_headers) - {"path", "linkpath"} + ): + _fail("invalid_archive", "Workspace archive has duplicate or unsupported entries") + if member.name in _ADMIN: + if not member.isfile(): + _fail("invalid_archive", "Workspace archive metadata is not a regular file") + elif member.name.startswith("files/"): + _path( + member.name.removeprefix("files/").removesuffix("/") + if member.isdir() + else member.name.removeprefix("files/") + ) + if member.type not in {tarfile.REGTYPE, tarfile.DIRTYPE, tarfile.SYMTYPE}: + _fail("invalid_archive", "Workspace archive file type is unsupported") + else: + _fail("invalid_archive", "Workspace archive has an unknown entry") + members[name] = member + if not _ADMIN.issubset(members) or members["manifest.json"].size > 32 * 1024 * 1024: + _fail("invalid_archive", "Workspace archive manifest is unavailable") + stream = tar.extractfile(members["manifest.json"]) + if stream is None: + _fail("invalid_archive", "Workspace archive manifest has no data") + with stream: + manifest = json.load(stream) + if ( + not isinstance(manifest, dict) + or set(manifest) + != {"version", "identity", "workspace", "git", "files", "bundle_sha256", "patch_sha256"} + or type(manifest["version"]) is not int + or manifest["version"] != 1 + or manifest["identity"] != asdict(identity) + or manifest["workspace"] != str(root) + or not isinstance(manifest["files"], list) + or len(manifest["files"]) > limits.max_entries + ): + _fail("invalid_archive", "Workspace checkpoint identity, path or version differs") + state = manifest["git"] + if ( + not isinstance(state, dict) + or set(state) != {"head", "branch", "refs", "index_sha256", "exclude_b64"} + or not isinstance(state["head"], str) + or not _OID.fullmatch(state["head"]) + or (state["branch"] is not None and not isinstance(state["branch"], str)) + or not isinstance(state["refs"], dict) + or not isinstance(state["index_sha256"], str) + or not _SHA256.fullmatch(state["index_sha256"]) + ): + _fail("invalid_archive", "Workspace Git metadata is invalid") + exclude = state["exclude_b64"] + if exclude is not None: + if not isinstance(exclude, str) or len(exclude) > 4 * (_MAX_EXCLUDE_BYTES // 3 + 1): + _fail("invalid_archive", "Workspace Git ignore data is invalid") + if len(base64.b64decode(exclude, validate=True)) > _MAX_EXCLUDE_BYTES: + _fail("invalid_archive", "Workspace Git ignore data exceeds its limit") + paths = {} + folded = set() + for record in manifest["files"]: + if not isinstance(record, dict): + _fail("invalid_archive", "Workspace file manifest is invalid") + path = str(_path(record.get("path"))) + if path.casefold() in folded: + _fail("invalid_archive", "Workspace file paths collide") + folded.add(path.casefold()) + kind = record.get("kind") + fields = {"path", "mode", "kind"} | ( + {"sha256", "size"} if kind == "file" else {"target"} if kind == "symlink" else set() + ) + member = members.get("files/" + path) + if ( + set(record) != fields + or kind not in {"file", "directory", "symlink"} + or member is None + or type(record["mode"]) is not int + or member.mode != record["mode"] + or member.type + != {"file": tarfile.REGTYPE, "directory": tarfile.DIRTYPE, "symlink": tarfile.SYMTYPE}[ + kind + ] + ): + _fail("invalid_archive", "Workspace file manifest disagrees with its archive") + if kind == "file": + if type(record["size"]) is not int or member.size != record["size"]: + _fail("invalid_archive", "Workspace file size disagrees with its archive") + _verify_member(tar, member, record["sha256"]) + elif kind == "symlink": + if ( + not isinstance(record["target"], str) + or not record["target"] + or "\0" in record["target"] + or len(record["target"].encode()) > _MAX_PATH_BYTES + or member.linkname != record["target"] + or member.size != 0 + ): + _fail("invalid_archive", "Workspace symlink metadata is invalid") + elif member.size != 0: + _fail("invalid_archive", "Workspace directory contains data") + paths[path] = kind + if set(members) != _ADMIN | {"files/" + path for path in paths}: + _fail("invalid_archive", "Workspace archive has unlisted files") + for path in paths: + for parent in PurePosixPath(path).parents: + if str(parent) != "." and paths.get(str(parent)) != "directory": + _fail( + "invalid_archive", "Workspace archive traverses a symlink or missing directory" + ) + for name, key in (("git.bundle", "bundle_sha256"), ("index.patch", "patch_sha256")): + _verify_member(tar, members[name], manifest[key]) + return manifest, members + + +def _verify_member(tar: tarfile.TarFile, member: tarfile.TarInfo, digest: str) -> None: + if not isinstance(digest, str) or not _SHA256.fullmatch(digest): + _fail("invalid_archive", "Workspace file checksum is invalid") + stream = tar.extractfile(member) + if stream is None: + _fail("invalid_archive", "Workspace archive file has no data") + with stream: + actual = hashlib.sha256() + while block := stream.read(_CHUNK): + actual.update(block) + if actual.hexdigest() != digest: + _fail("checksum_mismatch", "Workspace file checksum did not match") + + +def restore_workspace( + archive: Path, + workspace: Path, + identity: CheckpointIdentity, + *, + expected_sha256: str, + limits: WorkspaceLimits = _DEFAULT_LIMITS, +) -> WorkspaceArchive: + """Restore only into a nonexistent stable path; validate before publication.""" + try: + return _restore_workspace(archive, workspace, identity, expected_sha256, limits) + except WorkspaceCheckpointError: + raise + except ( + OSError, + ValueError, + TypeError, + RecursionError, + subprocess.SubprocessError, + tarfile.TarError, + ) as exc: + raise WorkspaceCheckpointError( + "restore_failed", "Workspace restoration failed; continuation cannot start" + ) from exc + + +def _restore_workspace( + archive: Path, + workspace: Path, + identity: CheckpointIdentity, + expected_sha256: str, + limits: WorkspaceLimits, +) -> WorkspaceArchive: + root = workspace.absolute() + if root.parent.resolve() != root.parent or os.path.lexists(root): + _fail("destination_exists", "Workspace restore requires a new canonical destination") + _repository_identity(identity) + size = archive.stat().st_size + if not 0 < size <= limits.max_bytes: + _fail("size_limit", "Workspace archive exceeds the byte limit") + if not isinstance(expected_sha256, str) or not _SHA256.fullmatch(expected_sha256): + _fail("checksum_mismatch", "Workspace archive receipt checksum is invalid") + # The only tree removed on error is this newly-created owned staging tree. + # Never use extractall: paths/types/parents are validated, and links are leaves. + fd = os.open(archive, os.O_RDONLY | os.O_NOFOLLOW | os.O_NONBLOCK) + with ( + os.fdopen(fd, "rb") as archive_stream, + tempfile.TemporaryDirectory(prefix=".workspace-restore-", dir=root.parent) as scratch, + ): + archive_stamp = _stamp(os.fstat(archive_stream.fileno())) + if not stat.S_ISREG(archive_stamp[0]) or archive_stamp[3] != size: + _fail("invalid_archive", "Workspace archive is not a stable regular file") + if hashlib.file_digest(archive_stream, "sha256").hexdigest() != expected_sha256: + _fail("checksum_mismatch", "Workspace archive differs from its receipt") + archive_stream.seek(0) + staging = Path(scratch) / "tree" + staging.mkdir(mode=0o700) + with tarfile.open(fileobj=archive_stream, mode="r:") as tar: + manifest, members = _validated_manifest(tar, identity, root, limits) + for name in ("git.bundle", "index.patch"): + source = tar.extractfile(members[name]) + if source is None: + _fail("invalid_archive", "Workspace Git archive has no data") + with source, (Path(scratch) / name).open("xb") as output: + shutil.copyfileobj(source, output, _CHUNK) + records = manifest["files"] + for record in sorted( + records, key=lambda record: len(PurePosixPath(record["path"]).parts) + ): + target = staging / record["path"] + if record["kind"] == "directory": + target.mkdir(mode=0o700) + elif record["kind"] == "symlink": + os.symlink(record["target"], target) + else: + source = tar.extractfile(members["files/" + record["path"]]) + if source is None: + _fail("invalid_archive", "Workspace archive file has no data") + with source, target.open("xb") as output: + shutil.copyfileobj(source, output, _CHUNK) + target.chmod(record["mode"]) + state = manifest["git"] + _git(staging, ["init", "--template=", "--quiet"], limits) + for ref, oid in state["refs"].items(): + if ( + not isinstance(ref, str) + or not ref.startswith("refs/") + or ref.startswith("refs/replace/") + or not isinstance(oid, str) + or not _OID.fullmatch(oid) + ): + _fail("invalid_archive", "Workspace saved Git ref is invalid") + _git(staging, ["check-ref-format", ref], limits) + bundle = str(Path(scratch) / "git.bundle") + _git(staging, ["bundle", "verify", bundle], limits) + advertised = _git(staging, ["bundle", "unbundle", bundle], limits) + bundle_refs = dict(line.decode().split(" ", 1)[::-1] for line in advertised.splitlines()) + if bundle_refs != {**state["refs"], "HEAD": state["head"]}: + _fail("invalid_archive", "Workspace Git bundle refs disagree with the manifest") + ref_updates = Path(scratch) / "refs.txt" + ref_updates.write_text( + "".join(f"update {ref} {oid}\n" for ref, oid in state["refs"].items()) + ) + with ref_updates.open("rb") as source: + _git(staging, ["update-ref", "--stdin"], limits, source=source) + branch = state["branch"] + if branch is None: + _git(staging, ["update-ref", "--no-deref", "HEAD", state["head"]], limits) + else: + _git(staging, ["check-ref-format", "refs/heads/" + branch], limits) + if state["refs"].get("refs/heads/" + branch) != state["head"]: + _fail("invalid_archive", "Workspace branch disagrees with saved HEAD") + _git(staging, ["symbolic-ref", "HEAD", "refs/heads/" + branch], limits) + _git(staging, ["read-tree", "HEAD"], limits) + patch = Path(scratch) / "index.patch" + if patch.stat().st_size: + _git( + staging, + ["apply", "--cached", "--binary", "--whitespace=nowarn", str(patch)], + limits, + ) + index = _git(staging, ["ls-files", "--stage", "-z"], limits) + if hashlib.sha256(index).hexdigest() != state["index_sha256"]: + _fail("invalid_archive", "Workspace staged state was not restored exactly") + if identity.repo: + _git( + staging, + ["remote", "add", "origin", f"https://github.com/{identity.repo}.git"], + limits, + ) + _git( + staging, + ["config", "--local", "credential.helper", "!gh auth git-credential"], + limits, + ) + if state["exclude_b64"] is not None: + ignore_file = staging / ".git/info/exclude" + ignore_file.parent.mkdir(exist_ok=True) + ignore_file.write_bytes(base64.b64decode(state["exclude_b64"], validate=True)) + for record in sorted( + records, key=lambda record: len(PurePosixPath(record["path"]).parts), reverse=True + ): + if record["kind"] == "directory": + (staging / record["path"]).chmod(record["mode"]) + if _stamp(os.fstat(archive_stream.fileno())) != archive_stamp: + _fail("workspace_changed", "Workspace archive changed during restoration") + # Reserve the final directory exclusively, then move the owned contents. + # A partial move is rolled back; an existing destination is never cleared. + root.mkdir(mode=0o700) + try: + for entry in staging.iterdir(): + os.rename(entry, root / entry.name) + except BaseException: + shutil.rmtree(root) + raise + return WorkspaceArchive(expected_sha256, size, state["head"], branch, len(records)) diff --git a/agent/src/hooks.py b/agent/src/hooks.py index 99a1252cd..5575ae514 100644 --- a/agent/src/hooks.py +++ b/agent/src/hooks.py @@ -1,8 +1,9 @@ """PreToolUse, PostToolUse, and Stop hook callbacks. - PreToolUse: three-outcome Cedar policy enforcement (ALLOW / DENY / - REQUIRE_APPROVAL). The REQUIRE_APPROVAL path writes a pending approval - row + transitions the task to AWAITING_APPROVAL atomically, polls for a + REQUIRE_APPROVAL). The REQUIRE_APPROVAL path asks the trusted approval + service to create a pending row and atomically transition the task to + AWAITING_APPROVAL, polls for a human decision, then resumes / denies per the user's input. See ``docs/design/CEDAR_HITL_GATES.md``. - PostToolUse: output scanner for secrets/PII. @@ -24,15 +25,19 @@ import re import time from collections.abc import Callable -from datetime import UTC +from contextlib import nullcontext +from dataclasses import dataclass +from datetime import UTC, datetime from typing import TYPE_CHECKING, Any import nudge_reader import task_state +from microvm_lifecycle import ApprovalRecord, get_context from nudge_reader import _xml_escape from output_scanner import scan_tool_output from policy import APPROVAL_RATE_LIMIT, FLOOR_TIMEOUT_S, Outcome from progress_writer import _generate_ulid +from shared_constants import SHARED_CONSTANTS from shell import log, log_error_cw from stuck_guard import StuckGuard @@ -57,6 +62,38 @@ TOOL_INPUT_PREVIEW_MAX: int = 256 # cap for the strip-ANSI, truncated input preview ELLIPSIS_LEN: int = 3 # chars reserved for the "..." truncation marker + +@dataclass(frozen=True) +class _ApprovalDeadline: + """One gate's deadline, retained across polling and a possible VM sleep. + + UTC counts elapsed time even if the guest's monotonic clock freezes. The monotonic + cap prevents a backward UTC correction from extending the original window. + Create this with the approval row, before database writes or notifications; + never recreate it on resume. + """ + + wall_deadline: float + monotonic_deadline: float + + @classmethod + def from_recorded(cls, created_at: str, timeout_s: int) -> _ApprovalDeadline: + if timeout_s == 0: + return cls(float("inf"), float("inf")) + created_epoch = ( + datetime.strptime(created_at, "%Y-%m-%dT%H:%M:%SZ").replace(tzinfo=UTC).timestamp() + ) + wall_deadline = created_epoch + timeout_s + remaining = max(0.0, min(timeout_s, wall_deadline - time.time())) + return cls(wall_deadline, time.monotonic() + remaining) + + def remaining_s(self) -> float: + return max( + 0.0, + min(self.monotonic_deadline - time.monotonic(), self.wall_deadline - time.time()), + ) + + # ANSI CSI / OSC escape sequence stripper for ``tool_input_preview`` + # ``permissionDecisionReason`` fields, so persisted/logged reasons can't carry # terminal-escape injection. Kept local to avoid a cross-module dependency for @@ -381,8 +418,9 @@ async def pre_tool_use_hook( - permissionDecision: "allow" or "deny" - permissionDecisionReason: explanation string - The REQUIRE_APPROVAL path pauses here: writes a pending approval row - + transitions the task to AWAITING_APPROVAL atomically, polls for a + The REQUIRE_APPROVAL path pauses here: the trusted service creates a pending + approval row and atomically transitions the task to AWAITING_APPROVAL, then + the worker polls for a human decision with 2s→5s backoff, then returns allow / deny based on the decision. On TIMED_OUT a ConditionCheckFailed from the best-effort status write triggers a re-read — if the user's decision landed between @@ -498,6 +536,25 @@ async def pre_tool_use_hook( return _deny_response(decision.reason) # -- REQUIRE_APPROVAL path ---------------------------------------------- + from continuation_runtime import current_runtime + + continuation = current_runtime() + saved_decision = ( + continuation.consume_approved_action(tool_name, _sha256_tool_input_for_row(tool_input)) + if continuation is not None + else None + ) + if saved_decision is not None: + if progress is not None: + _try_progress( + progress, + "write_approval_granted", + request_id=saved_decision["request_id"], + scope=saved_decision.get("scope") or "this_call", + decided_at=saved_decision.get("decided_at"), + created_at=saved_decision.get("created_at"), + ) + return _allow_response("User approved this saved action before worker continuation") return await _handle_require_approval( decision=decision, tool_name=tool_name, @@ -507,6 +564,8 @@ async def pre_tool_use_hook( user_id=user_id, progress=progress, ts=ts_module, + tool_use_id=tool_use_id, + session_id=hook_input.get("session_id", ""), ) @@ -520,6 +579,8 @@ async def _handle_require_approval( user_id: str | None, progress: Any, ts: Any, + tool_use_id: str | None = None, + session_id: str = "", ) -> dict: """REQUIRE_APPROVAL branch of ``pre_tool_use_hook``. @@ -565,7 +626,11 @@ async def _handle_require_approval( # Step 3 — effective timeout with floor/ceiling math. Emit # ``approval_timeout_capped`` when the caller's ask is clipped so the # user can see why. - remaining_lifetime = _remaining_maxlifetime_s() + # A durable checkpoint can move to a replacement worker. Its request's + # deadline is independent of this process's remaining service lifetime. + from continuation_runtime import current_runtime + + remaining_lifetime = None if current_runtime() is not None else _remaining_maxlifetime_s() effective_timeout, clip_reason, requested_timeout = _compute_effective_timeout( decision_timeout_s=decision.timeout_s, task_default_timeout_s=engine.task_default_timeout_s, @@ -623,10 +688,19 @@ async def _handle_require_approval( "status": "PENDING", "created_at": _iso_now(), "timeout_s": effective_timeout, - "ttl": int(time.time()) + effective_timeout + CLEANUP_MARGIN_120S, "user_id": user_id or "", "repo": engine.repo, } + if effective_timeout > 0: + row["deadline_epoch"] = ( + int( + datetime.strptime(row["created_at"], "%Y-%m-%dT%H:%M:%SZ") + .replace(tzinfo=UTC) + .timestamp() + ) + + effective_timeout + ) + deadline = _ApprovalDeadline.from_recorded(row["created_at"], effective_timeout) # Step 6 — bump counters BEFORE the write so cap/rate checks on # subsequent gates reflect the attempt even if the DDB write itself @@ -696,14 +770,87 @@ async def _handle_require_approval( matching_rule_ids=list(decision.matching_rule_ids), ) - # Step 9 — poll for a decision. - outcome = await _poll_for_decision( - task_id=task_id, - request_id=request_id, - timeout_s=effective_timeout, - progress=progress, - ts=ts, + # Step 9 — register the exact persisted gate and its ORIGINAL deadline. + # The SDK wrapper already tracks this tool; direct/legacy callers without a + # lifecycle context keep their existing approval behavior. + lifecycle = get_context(task_id) + park = ( + lifecycle.park_approval( + request_id, + tool_use_id, + deadline, + record=ApprovalRecord( + user_id=row["user_id"], + repo=row["repo"], + created_at=row["created_at"], + timeout_s=effective_timeout, + ), + ) + if lifecycle + else None ) + from continuation_runtime import current_runtime + + continuation = current_runtime() + checkpoint_held = False + if continuation is not None and lifecycle is not None and park is not None: + try: + identity, receipt = await continuation.capture( + lifecycle, + park, + session_id=session_id, + tool_name=tool_name, + tool_input=tool_input, + approval_scopes=engine.allowlist.snapshot_scopes(), + approval_gate_count=engine.approval_gate_count, + ) + checkpoint_held = True + await asyncio.to_thread( + ts.publish_continuation_checkpoint, + identity, + receipt, + tool_input_sha256=tool_input_sha256, + cost_usd=continuation.context.cost_usd, + turns_used=continuation.context.turns_used, + ) + log( + "TASK", + f"Continuation checkpoint ready task_id={task_id} request_id={request_id} " + f"attempt_id={identity.attempt_id}", + ) + except Exception as exc: + # A failed or uncertain publish cannot authorize worker release. + # If a receipt was saved, keep tools held until the conditional + # RUNNING claim resolves any race with the coordinator. + log( + "WARN", + f"Continuation checkpoint unavailable task_id={task_id} request_id={request_id} " + f"code={getattr(exc, 'code', 'checkpoint_failed')} error_type={type(exc).__name__}", + ) + _try_progress( + progress, + "write_agent_milestone", + milestone="continuation_unavailable", + details="Could not save recovery state. This worker must remain available " + f"until the decision is received. code={getattr(exc, 'code', 'checkpoint_failed')}", + ) + try: + outcome = await _poll_for_decision( + task_id=task_id, + request_id=request_id, + deadline=deadline, + progress=progress, + ts=ts, + ) + except BaseException: + if checkpoint_held and lifecycle: + lifecycle.close() + raise + finally: + if lifecycle and park and not checkpoint_held: + # Clear the safe point before ANY approval/task-state mutation or + # hook return. A concurrent suspend owns the barrier until resume. + await lifecycle.leave_approval(park) # Step 10 — VM-throttle + late-approval race. Best-effort flip to # TIMED_OUT; if ConditionCheckFailed, the user beat us — read and honor. @@ -741,6 +888,13 @@ async def _handle_require_approval( row_reread = None outcome = _reconcile_late_decision(outcome, row_reread, progress, request_id) + if outcome.get("status") == "CANCELLED": + # The platform closed the approval in the same transaction as task + # cancellation. Do not try to restore RUNNING or cache a human denial. + if checkpoint_held and lifecycle: + lifecycle.close() + return _deny_response("Task was cancelled; this approval is closed.") + # Step 11 — resume transition (RUNNING). The ``awaiting_approval_request_id`` # condition prevents resuming a cancelled task or racing with another # approval. @@ -754,6 +908,8 @@ async def _handle_require_approval( request_id=request_id, error=f"cancelled: {exc.cancellation_reasons}", ) + if checkpoint_held and lifecycle: + lifecycle.close() return _deny_response("task no longer awaiting approval") except Exception as exc: log("WARN", f"approval resume raised: {type(exc).__name__}: {exc}") @@ -764,8 +920,13 @@ async def _handle_require_approval( request_id=request_id, error=f"{type(exc).__name__}: {exc}", ) + if checkpoint_held and lifecycle: + lifecycle.close() return _deny_response("approval resume failed") + if checkpoint_held and lifecycle and park: + await lifecycle.release_continuation(park) + # Step 12 — terminal branches. status = outcome.get("status") if status == "APPROVED": @@ -892,6 +1053,7 @@ def _reconcile_late_decision( - ``row["status"] == "APPROVED"`` → rebuild as APPROVED (allow flow). - ``row["status"] == "DENIED"`` → rebuild as DENIED (deny flow). + - ``row["status"] == "CANCELLED"`` → close the wait without resuming the task. - Anything else (row gone, still PENDING) → fall through with the original TIMED_OUT (the fail-closed branch). @@ -916,6 +1078,8 @@ def _reconcile_late_decision( "decided_at": row.get("decided_at"), "decided_by": row.get("user_id"), } + if status == "CANCELLED": + return {"status": "CANCELLED", "reason": row.get("cancellation_reason")} if status == "DENIED": if progress is not None: _try_progress( @@ -937,7 +1101,7 @@ async def _poll_for_decision( *, task_id: str, request_id: str, - timeout_s: int, + deadline: _ApprovalDeadline, progress: Any, ts: Any, ) -> dict: @@ -949,70 +1113,77 @@ async def _poll_for_decision( ``approval_poll_degraded``; at ``POLL_MAX_CONSECUTIVE_FAILS`` we fall through as TIMED_OUT with a distinct reason. - Returns an outcome dict mirroring the approval row's terminal fields. + Uses the deadline captured with the original approval row, including time + spent writing/notifying or suspended. Returns an outcome dict mirroring + the approval row's terminal fields. """ - deadline = time.monotonic() + timeout_s start = time.monotonic() consecutive_fails = 0 degraded_emitted = False + lifecycle = get_context(task_id) while True: - now = time.monotonic() - if now >= deadline: - return {"status": "TIMED_OUT", "reason": None} + async with lifecycle.approval_poll() if lifecycle else nullcontext(): + if deadline.remaining_s() <= 0: + return {"status": "TIMED_OUT", "reason": None} - try: - row = await asyncio.to_thread( - ts.get_approval_row, - task_id, - request_id, - consistent_read=True, - ) - consecutive_fails = 0 - except Exception as exc: - consecutive_fails += 1 - log( - "WARN", - f"approval poll get_item raised ({consecutive_fails}/" - f"{POLL_MAX_CONSECUTIVE_FAILS}): {type(exc).__name__}: {exc}", - ) - if consecutive_fails >= POLL_DEGRADED_FAILS and not degraded_emitted: - if progress is not None: - _try_progress( - progress, - "write_approval_poll_degraded", - request_id=request_id, - consecutive_failures=consecutive_fails, - ) - degraded_emitted = True - if consecutive_fails >= POLL_MAX_CONSECUTIVE_FAILS: - return { - "status": "TIMED_OUT", - "reason": f"poll failed {consecutive_fails} consecutive times", - } - row = None # force sleep below - - if row is not None: - status = row.get("status") - if status == "APPROVED": - return { - "status": "APPROVED", - "scope": row.get("scope"), - "decided_at": row.get("decided_at"), - "decided_by": row.get("user_id"), - } - if status == "DENIED": - return { - "status": "DENIED", - "reason": row.get("deny_reason") or "denied", - "decided_at": row.get("decided_at"), - } + try: + row = await asyncio.to_thread( + ts.get_approval_row, + task_id, + request_id, + consistent_read=True, + ) + consecutive_fails = 0 + except Exception as exc: + consecutive_fails += 1 + log( + "WARN", + f"approval poll get_item raised ({consecutive_fails}/" + f"{POLL_MAX_CONSECUTIVE_FAILS}): {type(exc).__name__}: {exc}", + ) + if consecutive_fails >= POLL_DEGRADED_FAILS and not degraded_emitted: + if progress is not None: + _try_progress( + progress, + "write_approval_poll_degraded", + request_id=request_id, + consecutive_failures=consecutive_fails, + ) + degraded_emitted = True + if consecutive_fails >= POLL_MAX_CONSECUTIVE_FAILS: + return { + "status": "TIMED_OUT", + "reason": f"poll failed {consecutive_fails} consecutive times", + } + row = None # force sleep below + + if row is not None: + status = row.get("status") + if status == "APPROVED": + return { + "status": "APPROVED", + "scope": row.get("scope"), + "decided_at": row.get("decided_at"), + "decided_by": row.get("user_id"), + } + if status == "DENIED": + return { + "status": "DENIED", + "reason": row.get("deny_reason") or "denied", + "decided_at": row.get("decided_at"), + } + if status == "CANCELLED": + return { + "status": "CANCELLED", + "reason": row.get("cancellation_reason"), + } # Compute sleep interval based on elapsed since poll started. elapsed = time.monotonic() - start interval = POLL_FAST_INTERVAL_S if elapsed < POLL_FAST_DURATION_S else POLL_SLOW_INTERVAL_S # Clamp sleep against remaining deadline so we don't oversleep. - sleep_for = min(interval, max(0.0, deadline - time.monotonic())) + sleep_for = min(interval, deadline.remaining_s()) if sleep_for <= 0: return {"status": "TIMED_OUT", "reason": None} await asyncio.sleep(sleep_for) @@ -1026,14 +1197,10 @@ def _compute_effective_timeout( ) -> tuple[int, str | None, int]: """Compute the effective approval timeout. - ``min(rule-annotation timeout, task default, remaining lifetime - - cleanup margin)``, floored at FLOOR_30S. The engine's - ``_merge_annotations`` already applies ``min(rule_annotation, - task_default)`` — decision.timeout_s reaches us pre-clipped against - those two. Here we apply the remaining-lifetime ceiling and report - whichever source pulled the effective timeout below the task - default, so the user sees "your gate was clipped because ..." rather - than silent clipping. + The engine has already merged positive rule/task deadlines. Zero means + no deadline and returns unchanged. For a positive result, apply an optional + remaining-lifetime ceiling and the 30-second floor. Continuation runtimes + omit the lifetime ceiling because the request can outlive this worker. Returns ``(effective, clip_reason, requested)``: - ``requested`` — the user-visible "would have liked" value (task @@ -1047,6 +1214,8 @@ def _compute_effective_timeout( decision_value = ( decision_timeout_s if decision_timeout_s is not None else task_default_timeout_s ) + if decision_value == 0: + return 0, None, requested # Start with the decision value (already clipped by rule vs task default # in the engine) and apply the remaining-lifetime ceiling here. @@ -1096,8 +1265,6 @@ def _remaining_maxlifetime_s() -> int | None: if started_at.isdigit(): started_epoch = int(started_at) else: - from datetime import datetime - # The trailing Z means UTC; strptime returns a naive datetime whose # .timestamp() would otherwise be interpreted in the container's # local TZ, skewing remaining-lifetime math by the UTC offset. @@ -1730,8 +1897,9 @@ def build_hook_matchers( Returns a dict mapping HookEvent strings to lists of HookMatcher instances, ready to pass as ``hooks=...`` to ClaudeAgentOptions. - The SDK expects ``dict[HookEvent, list[HookMatcher]]`` where HookMatcher - has ``matcher: str | None`` and ``hooks: list[HookCallback]``. + The SDK expects ``dict[HookEvent, list[HookMatcher]]``. PreToolUse also needs + an explicit callback timeout: the CLI can otherwise cancel a live approval + wait and return a generic denial without completing its approval record. ``progress`` is forwarded to both the PreToolUse hook (approval gate milestones) and the Stop hook (nudge/denial acks). ``user_id`` is @@ -1751,6 +1919,9 @@ def build_hook_matchers( # PostToolUse closure feeds it every tool result; the Stop closure reads it # between turns to steer / bail on a repeating failing command. _stuck_guard = StuckGuard() + # Retain the controller even after registry removal. A late callback must + # see its CLOSED barrier, not fall back to the non-MicroVM path. + lifecycle = get_context(task_id) # Closure-based wrapper matches the HookCallback signature exactly: # (HookInput, str | None, HookContext) -> Awaitable[HookJSONOutput] @@ -1758,13 +1929,24 @@ async def _pre( hook_input: HookInput, tool_use_id: str | None, ctx: HookContext ) -> HookJSONOutput: # Fail-closed wrapper (mirrors _post and _stop). If the inner hook - # or its dispatch path raises an unexpected exception (asyncio - # cancellation, TypeError from a malformed payload, etc.), the + # or its dispatch path raises an unexpected exception (for example, + # TypeError from a malformed payload), the # SDK's default behaviour for an unhandled hook exception is # undefined — we MUST NOT trust it to fail closed. Mapping every # uncaught exception to a DENY here makes the security posture - # explicit at the SDK boundary. + # explicit at the SDK boundary. SDK cancellation propagates and still + # releases the tool registration in finally. + allowed = False try: + if lifecycle: + await lifecycle.tool_started(tool_use_id) + if isinstance(hook_input, dict): + tool_input = hook_input.get("tool_input") + # Legacy serialized inputs are normalized deeper in the + # policy hook. Keep their tool behavior, but do not infer a + # safe foreground-only execution from an opaque value here. + if not isinstance(tool_input, dict) or tool_input.get("run_in_background"): + lifecycle.disable_suspend() result = await pre_tool_use_hook( hook_input, tool_use_id, @@ -1776,6 +1958,10 @@ async def _pre( progress=progress, repo_url=repo_url or None, ) + if result.get("hookSpecificOutput", {}).get("permissionDecision") == "allow": + if lifecycle: + await lifecycle.wait_until_open() + allowed = True except Exception as exc: log( "ERROR", @@ -1788,6 +1974,9 @@ async def _pre( return SyncHookJSONOutput( **_deny_response("Hook error — fail-closed deny"), ) + finally: + if lifecycle and not allowed: + lifecycle.tool_finished(tool_use_id) return SyncHookJSONOutput(**result) async def _post( @@ -1810,6 +1999,25 @@ async def _post( "updatedMCPToolOutput": "[Output redacted: hook error — fail-closed]", } return SyncHookJSONOutput(hookSpecificOutput=fail_closed) + finally: + if lifecycle: + # Background Bash may be promoted after invocation. A returned + # background identifier means the tool's post hook is not a + # reliable indication that its child stopped. + if isinstance(hook_input, dict): + response = hook_input.get("tool_response") + if isinstance(response, dict) and ( + response.get("backgroundTaskId") or response.get("background_task_id") + ): + lifecycle.disable_suspend() + lifecycle.tool_finished(tool_use_id) + + async def _post_failure( + hook_input: HookInput, tool_use_id: str | None, ctx: HookContext + ) -> HookJSONOutput: + if lifecycle: + lifecycle.tool_finished(tool_use_id) + return SyncHookJSONOutput() async def _stop( hook_input: HookInput, tool_use_id: str | None, ctx: HookContext @@ -1836,8 +2044,22 @@ async def _stop( # Empty dict == allow stop. SyncHookJSONOutput(**{}) is fine. return SyncHookJSONOutput(**result) - return { - "PreToolUse": [HookMatcher(matcher=None, hooks=[_pre])], + # All backends share this bounded SDK callback window (eight hours plus + # cleanup). It is separate from the human decision deadline and also covers + # MicroVM freeze/recovery. A retained request can outlive a callback only + # through a verified continuation; this does not extend compute lifetime. + callback_window_s = SHARED_CONSTANTS["microvm_lifecycle"]["maximum_duration_seconds"] + matchers = { + "PreToolUse": [ + HookMatcher( + matcher=None, + hooks=[_pre], + timeout=callback_window_s + CLEANUP_MARGIN_120S, + ) + ], "PostToolUse": [HookMatcher(matcher=None, hooks=[_post])], "Stop": [HookMatcher(matcher=None, hooks=[_stop])], } + if lifecycle: + matchers["PostToolUseFailure"] = [HookMatcher(matcher=None, hooks=[_post_failure])] + return matchers diff --git a/agent/src/microvm_checkpoint.py b/agent/src/microvm_checkpoint.py new file mode 100644 index 000000000..c51c81a08 --- /dev/null +++ b/agent/src/microvm_checkpoint.py @@ -0,0 +1,251 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Durable guest pause checks; the coordinator retains lifecycle intent ownership.""" + +from __future__ import annotations + +import os +from datetime import UTC, datetime +from decimal import Decimal +from typing import TYPE_CHECKING, Any, Literal + +from microvm_diagnostics import lifecycle_stage +from microvm_lifecycle import ApprovalPark, LifecycleUnavailable +from progress_writer import _ProgressWriter + +if TYPE_CHECKING: + from microvm_lifecycle import ApprovalRecord + + +def _integer(value: Any) -> bool: + return ( + isinstance(value, (int, Decimal)) + and not isinstance(value, bool) + and value == int(value) + and 0 <= value <= 2**53 - 1 + ) + + +def _record(park: ApprovalPark) -> tuple[ApprovalRecord, int | None]: + record = park.record + if ( + record is None + or not record.user_id + or not isinstance(record.repo, str) + or not _integer(record.timeout_s) + ): + raise LifecycleUnavailable("Original approval identity is unavailable") + try: + created = datetime.strptime(record.created_at, "%Y-%m-%dT%H:%M:%SZ").replace(tzinfo=UTC) + except (TypeError, ValueError) as exc: + raise LifecycleUnavailable("Original approval timestamp is invalid") from exc + if created.strftime("%Y-%m-%dT%H:%M:%SZ") != record.created_at: + raise LifecycleUnavailable("Original approval timestamp is not canonical") + deadline_ms = ( + int(created.timestamp() * 1000) + record.timeout_s * 1000 if record.timeout_s > 0 else None + ) + return record, deadline_ms + + +def _read( + park: ApprovalPark, action: Literal["suspend", "resume"] +) -> tuple[Any, str, str, dict, int | None]: + """Read current task/gate identity strongly; absence and mismatch fail closed.""" + from boto3.dynamodb.types import TypeDeserializer + from botocore.config import Config + + from aws_session import tenant_client + + record, deadline_ms = _record(park) + task_table = os.environ.get("TASK_TABLE_NAME", "").strip() + approvals_table = os.environ.get("TASK_APPROVALS_TABLE_NAME", "").strip() + if not task_table or not approvals_table: + raise RuntimeError("Lifecycle task/approval tables are unavailable") + client = tenant_client( + "dynamodb", + region_name=os.environ.get("AWS_REGION") or os.environ.get("AWS_DEFAULT_REGION"), + config=Config(connect_timeout=2, read_timeout=2, retries={"total_max_attempts": 1}), + ) + deserialize = TypeDeserializer().deserialize + + def read_item(table: str, key: dict) -> dict: + item = client.get_item(TableName=table, Key=key, ConsistentRead=True).get("Item") + if not isinstance(item, dict): + raise LifecycleUnavailable("Lifecycle task or approval is missing") + return {key: deserialize(value) for key, value in item.items()} + + task = read_item(task_table, {"task_id": {"S": park.task_id}}) + metadata = task.get("compute_metadata") + intent = task.get("microvm_lifecycle") + if ( + task.get("task_id") != park.task_id + or task.get("user_id") != record.user_id + or task.get("repo", "") != record.repo + or task.get("compute_type") != "lambda-microvm" + or task.get("session_id") != park.microvm_id + or not isinstance(metadata, dict) + or metadata.get("microvmId") != park.microvm_id + or task.get("status") != "AWAITING_APPROVAL" + or task.get("awaiting_approval_request_id") != park.request_id + ): + raise LifecycleUnavailable("Lifecycle task identity or approval state changed") + if ( + not isinstance(intent, dict) + or not _integer(intent.get("version")) + or intent["version"] != 1 + or not isinstance(intent.get("generation"), str) + or not intent["generation"].strip() + or intent.get("microvm_id") != park.microvm_id + or intent.get("request_id") != park.request_id + or intent.get("action") != action + or not _integer(intent.get("requested_at_ms")) + or "deadline_ms" not in intent + or (intent["deadline_ms"] is not None and not _integer(intent["deadline_ms"])) + or intent["deadline_ms"] != deadline_ms + ): + raise LifecycleUnavailable("Coordinator lifecycle intent does not match this approval") + + approval = read_item( + approvals_table, + {"task_id": {"S": park.task_id}, "request_id": {"S": park.request_id}}, + ) + statuses = ( + {"PENDING"} + if action == "suspend" + else {"PENDING", "APPROVED", "DENIED", "TIMED_OUT", "STRANDED"} + ) + if ( + approval.get("task_id") != park.task_id + or approval.get("request_id") != park.request_id + or approval.get("user_id") != record.user_id + or approval.get("repo") != record.repo + or approval.get("created_at") != record.created_at + or not _integer(approval.get("timeout_s")) + or approval["timeout_s"] != record.timeout_s + or approval.get("status") not in statuses + ): + raise LifecycleUnavailable("Original approval changed or cannot be reconciled") + return client, task_table, approvals_table, intent, deadline_ms + + +def _transaction_checks( + park: ApprovalPark, task_table: str, approvals_table: str, intent: dict +) -> list[dict]: + """Recheck both rows atomically after validating their complete read shapes.""" + from boto3.dynamodb.types import TypeSerializer + + record, _ = _record(park) + serialize = TypeSerializer().serialize + values = { + ":task": park.task_id, + ":user": record.user_id, + ":repo": record.repo, + ":vm": park.microvm_id, + ":request": park.request_id, + ":awaiting": "AWAITING_APPROVAL", + ":backend": "lambda-microvm", + ":intent": intent, + } + task_check = { + "TableName": task_table, + "Key": {"task_id": {"S": park.task_id}}, + "ConditionExpression": ( + "task_id = :task AND user_id = :user AND " + + ( + "(attribute_not_exists(repo) OR repo = :repo)" + if not record.repo + else "repo = :repo" + ) + + " AND compute_type = :backend AND session_id = :vm" + " AND compute_metadata.microvmId = :vm AND #status = :awaiting" + " AND awaiting_approval_request_id = :request AND microvm_lifecycle = :intent" + ), + "ExpressionAttributeNames": {"#status": "status"}, + "ExpressionAttributeValues": {key: serialize(value) for key, value in values.items()}, + } + approval_check = { + "TableName": approvals_table, + "Key": {"task_id": {"S": park.task_id}, "request_id": {"S": park.request_id}}, + "ConditionExpression": ( + "task_id = :task AND request_id = :request AND user_id = :user AND repo = :repo" + " AND created_at = :created AND timeout_s = :timeout AND " + + ( + "#status = :pending" + if intent["action"] == "suspend" + else "#status IN (:pending, :approved, :denied, :timed_out, :stranded)" + ) + ), + "ExpressionAttributeNames": {"#status": "status"}, + "ExpressionAttributeValues": { + key: serialize(value) + for key, value in { + ":task": park.task_id, + ":request": park.request_id, + ":user": record.user_id, + ":repo": record.repo, + ":created": record.created_at, + ":timeout": record.timeout_s, + ":pending": "PENDING", + **( + { + ":approved": "APPROVED", + ":denied": "DENIED", + ":timed_out": "TIMED_OUT", + ":stranded": "STRANDED", + } + if intent["action"] == "resume" + else {} + ), + }.items() + }, + } + return [task_check, approval_check] + + +def checkpoint_before_suspend(park: ApprovalPark) -> None: + """Commit a marker only if the same task, sleep intent and pending gate hold.""" + record, _ = _record(park) + with lifecycle_stage("checkpoint-identity-read"): + client, task_table, approvals_table, intent, deadline_ms = _read(park, "suspend") + if park.deadline.remaining_s() <= 0: + raise LifecycleUnavailable("Approval deadline elapsed before checkpoint") + # Never reuse an old transaction client token across HTTP requests: cached + # success must not bypass conditions after a concurrent approval/cancellation. + with lifecycle_stage("checkpoint-transaction"): + _ProgressWriter( + park.task_id, user_id=record.user_id, repo=record.repo + ).write_microvm_checkpoint( + client=client, + condition_checks=_transaction_checks(park, task_table, approvals_table, intent), + metadata={ + "request_id": park.request_id, + "microvm_id": park.microvm_id, + "generation": intent["generation"], + "approval_deadline_ms": deadline_ms, + }, + ) + + +def refresh_and_reconcile_after_resume(park: ApprovalPark) -> None: + """Refresh before reading AWS; the original approval loop owns any decision.""" + from aws_session import refresh_microvm_credentials + + _record(park) + with lifecycle_stage("credential-refresh"): + refresh_microvm_credentials(park.task_id) + with lifecycle_stage("resume-identity-read"): + client, task_table, approvals_table, intent, _ = _read(park, "resume") + # This transaction writes no task/approval state. It acknowledges that both + # identities still hold, including cancellation or intent changes after reads. + # A concurrent valid decision is allowed; the original loop observes it. + with lifecycle_stage("resume-identity-transaction"): + client.transact_write_items( + TransactItems=[ + {"ConditionCheck": check} + for check in _transaction_checks(park, task_table, approvals_table, intent) + ] + ) + # The existing approval loop checks its original stopwatch/UTC cap as soon + # as the barrier opens. Expiry enters its timeout/late-decision path; a timely + # approval already recorded must still win that conditional race. diff --git a/agent/src/microvm_credentials.py b/agent/src/microvm_credentials.py new file mode 100644 index 000000000..d6ed1c06e --- /dev/null +++ b/agent/src/microvm_credentials.py @@ -0,0 +1,137 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Runtime-only scoped credential endpoint for the MicroVM Claude subprocess. + +AWS's container provider waits for replacement keys before signing after expiry. +The managed export helper deliberately returns no keys in this mode. The child +has no alternate AWS provider; the parent's runtime credential chain is untouched. +This protects provider selection, not isolation from code running as the same OS +user. Never start this server during image warm-up or snapshot validation. +""" + +from __future__ import annotations + +import hmac +import json +import os +import secrets +import threading +from http.server import BaseHTTPRequestHandler, HTTPServer +from typing import TYPE_CHECKING + +from aws_session import export_microvm_credentials + +if TYPE_CHECKING: + from collections.abc import Callable + + from microvm_lifecycle import MicrovmLifecycle + + +class ScopedCredentialBroker: + """One task's loopback endpoint, owned and closed by its Claude session.""" + + def __init__( + self, + lifecycle: MicrovmLifecycle, + *, + provider: Callable[[], dict[str, str]] | None = None, + ) -> None: + self._closed = threading.Event() + token = secrets.token_urlsafe(32) + resolve = provider or (lambda: export_microvm_credentials(lifecycle.task_id)) + closed = self._closed + + class Handler(BaseHTTPRequestHandler): + def log_message(self, format, *args): + # Authorization headers and response bodies contain credentials. + del format, args + + def do_GET(self): + if self.path != "/credentials": + self.send_error(404) + return + if not hmac.compare_digest(self.headers.get("Authorization", ""), token): + self.send_error(403) + return + try: + # Count the entire response in the suspend drain. Paused or + # failed lifecycle controllers reject before consulting AWS. + with lifecycle.activity(): + if closed.is_set(): + raise RuntimeError("Credential broker is closed") + body = json.dumps(resolve()).encode() + if closed.is_set(): + raise RuntimeError("Credential broker closed during renewal") + self.send_response(200) + self.send_header("Content-Type", "application/json") + self.send_header("Cache-Control", "no-store") + self.send_header("Content-Length", str(len(body))) + self.end_headers() + self.wfile.write(body) + except Exception: + # Deliberately omit provider exception text/credentials. + # The caller fails closed; no ambient fallback is available. + self.send_error(503, "Scoped credentials unavailable") + + class Server(HTTPServer): + def get_request(self): + connection, address = super().get_request() + connection.settimeout(2) + return connection, address + + def handle_error(self, request, client_address): + # A disconnected client may interrupt the generic 503 response. + # No request/exception dumping from this credential endpoint. + del request, client_address + + self._server = Server(("127.0.0.1", 0), Handler) + self._thread = threading.Thread( + target=self._server.serve_forever, + kwargs={"poll_interval": 0.05}, + name="microvm-scoped-credentials", + daemon=True, + ) + # ClaudeAgentOptions.env overlays the parent environment; empty strings + # suppress inherited providers. Controlled empty files also suppress + # ~/.aws profile/SSO/process credentials without changing HOME. + self.environment = { + key: "" + for key in os.environ + if key.startswith("AWS_") + and key not in {"AWS_REGION", "AWS_DEFAULT_REGION", "AWS_SDK_UA_APP_ID"} + } + self.environment.update( + { + "AWS_ACCESS_KEY_ID": "", + "AWS_SECRET_ACCESS_KEY": "", + "AWS_SESSION_TOKEN": "", + "AWS_PROFILE": "", + "AWS_DEFAULT_PROFILE": "", + "AWS_CONFIG_FILE": os.devnull, + "AWS_SHARED_CREDENTIALS_FILE": os.devnull, + "AWS_WEB_IDENTITY_TOKEN_FILE": "", + "AWS_ROLE_ARN": "", + "AWS_CONTAINER_CREDENTIALS_RELATIVE_URI": "", + "AWS_CONTAINER_CREDENTIALS_FULL_URI": ( + f"http://127.0.0.1:{self._server.server_port}/credentials" + ), + "AWS_CONTAINER_AUTHORIZATION_TOKEN": token, + "AWS_CONTAINER_AUTHORIZATION_TOKEN_FILE": "", + "AWS_EC2_METADATA_DISABLED": "true", + "AWS_BEARER_TOKEN_BEDROCK": "", + "ANTHROPIC_API_KEY": "", + "ANTHROPIC_AUTH_TOKEN": "", + "ABCA_MICROVM_CREDENTIAL_BROKER": "1", + } + ) + self._thread.start() + + def close(self) -> None: + """Stop serving keys, then stop the loop before the task is torn down.""" + if self._closed.is_set(): + return + self._closed.set() + self._server.shutdown() + self._server.server_close() + self._thread.join(timeout=2) diff --git a/agent/src/microvm_diagnostics.py b/agent/src/microvm_diagnostics.py new file mode 100644 index 000000000..20aa19ebe --- /dev/null +++ b/agent/src/microvm_diagnostics.py @@ -0,0 +1,139 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Bounded, payload-free lifecycle diagnostics shared with callback threads.""" + +from __future__ import annotations + +import json +import os +import re +import threading +import time +import uuid +from contextlib import contextmanager +from contextvars import ContextVar +from typing import TYPE_CHECKING, Any + +if TYPE_CHECKING: + from collections.abc import Callable + +_ACTIVE: ContextVar[HookDiagnostics | None] = ContextVar("microvm_hook_diagnostics", default=None) +_IDENTIFIER = re.compile(r"[A-Za-z0-9_-]{1,128}\Z") +_OUTPUT_LOCK = threading.Lock() + + +def _error_identity(error: BaseException | None) -> dict[str, str]: + if error is None: + return {} + name = type(error).__name__ + identity = {"error_type": name if _IDENTIFIER.fullmatch(name) else "Error"} + response = getattr(error, "response", None) + if isinstance(response, dict): + for source, field, target in [ + ("Error", "Code", "aws_error_code"), + ("ResponseMetadata", "RequestId", "aws_request_id"), + ]: + section = response.get(source) + value = section.get(field) if isinstance(section, dict) else None + if isinstance(value, str) and _IDENTIFIER.fullmatch(value): + identity[target] = value + return identity + + +class HookDiagnostics: + """One HTTP invocation; copied thread contexts retain this same correlation id. + + Logging cannot authorize a transition. Callback threads may outlive a timeout; + their subsequent records are marked late, never a successful HTTP acknowledgment. + """ + + def __init__(self, action: str, snapshot: Callable[[], dict[str, Any]]) -> None: + self.action = action + self.hook_id = str(uuid.uuid4()) + self._snapshot = snapshot + self._started = time.monotonic() + self._stage = "body-read" + self._closed = False + self._lock = threading.Lock() + + def emit(self, event: str, **fields: Any) -> None: + # Snapshot only safe local facts. Never include request bodies, SDK messages, + # approval records, tool arguments, credentials, or exception tracebacks. + try: + snapshot = self._snapshot() + safe = { + key: value + for key, value in snapshot.items() + if not isinstance(value, str) or _IDENTIFIER.fullmatch(value) + } + with self._lock: + record = { + **safe, + "event": event, + "level": "WARN" if event.endswith("_failed") else "INFO", + "action": self.action, + "hook_id": self.hook_id, + "pid": os.getpid(), + "timestamp_ms": int(time.time() * 1000), + "elapsed_ms": round((time.monotonic() - self._started) * 1000), + "stage": self._stage, + "late": self._closed, + **fields, + } + # Concurrent hooks and their callback threads must not interleave + # the JSON and newline writes of these records. + with _OUTPUT_LOCK: + print(json.dumps(record), flush=True) + except Exception: # noqa: S110 — failure in the logging sink cannot safely log to itself. + # Diagnostics are best effort; a broken stdout must not change the + # response or bypass the controller's generation/barrier checks. + pass + + def stage(self, name: str) -> None: + with self._lock: + if not self._closed: + self._stage = name + self.emit("microvm_hook_stage", callback_stage=name) + + def finish(self, status: int | None, code: str, error: BaseException | None = None) -> None: + with self._lock: + self._closed = True + self.emit( + "microvm_hook_finished", + http_status=status, + code=code, + late=False, + level="INFO" if code == "acknowledged" else "WARN", + **_error_identity(error), + ) + + +@contextmanager +def hook_diagnostics(action: str, snapshot: Callable[[], dict[str, Any]]): + diagnostics = HookDiagnostics(action, snapshot) + token = _ACTIVE.set(diagnostics) + try: + diagnostics.emit("microvm_hook_started") + yield diagnostics + finally: + _ACTIVE.reset(token) + + +@contextmanager +def lifecycle_stage(name: str): + """Record entry before potentially blocking work, including inside to_thread.""" + diagnostics = _ACTIVE.get() + if diagnostics: + diagnostics.stage(name) + try: + yield + except BaseException as error: + if diagnostics: + diagnostics.emit( + "microvm_hook_stage_failed", callback_stage=name, **_error_identity(error) + ) + raise + else: + if diagnostics: + diagnostics.emit("microvm_hook_stage_finished", callback_stage=name) diff --git a/agent/src/microvm_http.py b/agent/src/microvm_http.py new file mode 100644 index 000000000..ba7f526c3 --- /dev/null +++ b/agent/src/microvm_http.py @@ -0,0 +1,126 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Bounded service-owned MicroVM pause/wake routes; no public ingress is added.""" + +from __future__ import annotations + +import asyncio +import json +import time +from typing import Literal + +from fastapi import Request # noqa: TC002 — FastAPI resolves this annotation at registration. +from fastapi.responses import JSONResponse + +from microvm_checkpoint import checkpoint_before_suspend, refresh_and_reconcile_after_resume +from microvm_diagnostics import hook_diagnostics, lifecycle_stage +from microvm_lifecycle import LifecycleUnavailable, get_registered_context +from shared_constants import SHARED_CONSTANTS + +_BUDGETS = SHARED_CONSTANTS["microvm_hook_budgets"] +LIFECYCLE_HANDLER_BUDGET_S: float = _BUDGETS["lifecycle_handler_budget_seconds"] +LIFECYCLE_HOOK_TIMEOUT_S: float = _BUDGETS["lifecycle_hook_timeout_seconds"] +if not 0 < LIFECYCLE_HANDLER_BUDGET_S < LIFECYCLE_HOOK_TIMEOUT_S: + raise ValueError("Lifecycle handler must leave time for the service hook response") +_MAX_BODY_BYTES = 4096 + + +class _InvalidBody(ValueError): + def __init__(self, status: int): + self.status = status + + +async def _microvm_id(request: Request) -> str: + body = bytearray() + async for chunk in request.stream(): + body.extend(chunk) + if len(body) > _MAX_BODY_BYTES: + raise _InvalidBody(413) + try: + payload = json.loads(body) if body else {} + except (ValueError, UnicodeError) as exc: + raise _InvalidBody(400) from exc + if not isinstance(payload, dict): + raise _InvalidBody(400) + microvm_id = payload.get("microvmId", "") + if not isinstance(microvm_id, str) or microvm_id != microvm_id.strip(): + raise _InvalidBody(400) + # The service sends an empty id on /terminate; tolerate an absent/empty id + # here too, using only the local /run registration. A supplied id must match. + return microvm_id + + +async def _transition(request: Request, action: Literal["suspend", "resume"]) -> JSONResponse: + lifecycle = get_registered_context() + with hook_diagnostics( + action, lifecycle.diagnostic_snapshot if lifecycle else dict + ) as diagnostics: + try: + response = await _handle_transition(request, action) + except asyncio.CancelledError as exc: + diagnostics.finish(None, "MICROVM_LIFECYCLE_CANCELLED", exc) + raise + # A frozen VM retains this connection and its idle timer. Close it with + # the response so the next hook uses a fresh connection after restoration. + response.headers["Connection"] = "close" + diagnostics.finish( + response.status_code, json.loads(bytes(response.body)).get("code", "acknowledged") + ) + return response + + +async def _handle_transition( + request: Request, action: Literal["suspend", "resume"] +) -> JSONResponse: + try: + end = time.monotonic() + LIFECYCLE_HANDLER_BUDGET_S + async with asyncio.timeout(LIFECYCLE_HANDLER_BUDGET_S): + with lifecycle_stage("body-read"): + microvm_id = await _microvm_id(request) + with lifecycle_stage("identity-check"): + lifecycle = get_registered_context() + if lifecycle is None or (microvm_id and microvm_id != lifecycle.microvm_id): + raise LifecycleUnavailable("No matching MicroVM task is registered") + remaining = end - time.monotonic() + if remaining <= 0: + raise TimeoutError("Lifecycle body consumed its budget") + with lifecycle_stage("controller"): + if action == "suspend": + park = await lifecycle.suspend(checkpoint_before_suspend, budget_s=remaining) + else: + park = await lifecycle.resume( + refresh_and_reconcile_after_resume, budget_s=remaining + ) + return JSONResponse( + content={ + "status": "acknowledged", + "action": action, + "task_id": park.task_id, + "microvm_id": park.microvm_id, + "request_id": park.request_id, + } + ) + except _InvalidBody as exc: + return JSONResponse( + status_code=exc.status, content={"code": "MICROVM_LIFECYCLE_BODY_INVALID"} + ) + except LifecycleUnavailable: + return JSONResponse(status_code=409, content={"code": "MICROVM_LIFECYCLE_UNAVAILABLE"}) + except TimeoutError: + return JSONResponse(status_code=503, content={"code": "MICROVM_LIFECYCLE_TIMEOUT"}) + except Exception as exc: + # Neither service-hook responses nor logs may echo AWS exception details. + # A failed/uncertain wake stays behind the controller's closed barrier. + return JSONResponse( + status_code=503, + content={"code": "MICROVM_LIFECYCLE_FAILED", "error_type": type(exc).__name__}, + ) + + +async def microvm_suspend(request: Request) -> JSONResponse: + return await _transition(request, "suspend") + + +async def microvm_resume(request: Request) -> JSONResponse: + return await _transition(request, "resume") diff --git a/agent/src/microvm_lifecycle.py b/agent/src/microvm_lifecycle.py new file mode 100644 index 000000000..a519d33f3 --- /dev/null +++ b/agent/src/microvm_lifecycle.py @@ -0,0 +1,492 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Guest-side pause barrier for ADR-021. + +This controller does not call AWS or decide approvals. The server owns one +controller per running MicroVM task; tool hooks register work and the *original* +approval deadline. HTTP lifecycle handlers supply the durable checkpoint and +credential-refresh operations; this controller owns permission to release work. + +Returning from a suspend callback is not proof that AWS actually froze the VM. +Once acknowledged, the barrier opens only after a successful resume callback. +There is deliberately no timer that releases coding after an ambiguous suspend: +the supervisor must wake or terminate within its bounded recovery window. +""" + +from __future__ import annotations + +import asyncio +import math +import os +import random +import threading +import time +from contextlib import asynccontextmanager, contextmanager +from dataclasses import dataclass +from typing import TYPE_CHECKING, Protocol + +from microvm_diagnostics import lifecycle_stage + +if TYPE_CHECKING: + from collections.abc import Callable + + +class ApprovalDeadline(Protocol): + def remaining_s(self) -> float: ... + + +class LifecycleUnavailable(RuntimeError): + """The guest cannot establish a safe lifecycle boundary.""" + + +def reseed_random() -> None: + """Give each run/wake fresh application PRNG state; secrets still use OS RNG.""" + random.seed(os.urandom(32)) + + +@dataclass(frozen=True) +class ApprovalRecord: + """Original durable gate fields, captured when its request is written.""" + + user_id: str + repo: str + created_at: str + timeout_s: int + + +@dataclass(frozen=True) +class ApprovalPark: + task_id: str + microvm_id: str + request_id: str + tool_use_id: str + deadline: ApprovalDeadline + record: ApprovalRecord | None = None + + +class MicrovmLifecycle: + """Synchronize tool execution, approval exit, and bounded lifecycle work. + + Locks protect only local state; never hold one across an await or AWS call. + Timed-out callbacks can finish in a worker thread, but only this controller + can commit a transition, and its generation check rejects late completion. + """ + + def __init__(self, task_id: str, microvm_id: str, *, attempt_id: str | None = None) -> None: + if not task_id or not microvm_id: + raise ValueError("Lifecycle requires task and MicroVM identity") + self.task_id = task_id + self.microvm_id = microvm_id + self.attempt_id = attempt_id or task_id + self._lock = threading.Lock() + self._tools: set[str] = set() + self._park: ApprovalPark | None = None + self._phase = "active" + self._generation = 0 + self._activities = 0 + self._progress_failed = False + self._suspend_ineligible = False + self._slept_request_id: str | None = None + self._last_resume_park: ApprovalPark | None = None + self._continuation_park: ApprovalPark | None = None + + def _check_open(self) -> bool: + if self._phase in {"closed", "failed"}: + raise LifecycleUnavailable("Lifecycle barrier is closed") + return self._phase in {"active", "parked"} + + def diagnostic_snapshot(self) -> dict: + """Read only local state; never include approval contents or tool inputs.""" + with self._lock: + park = self._park or self._last_resume_park + return { + "task_id": self.task_id, + "microvm_id": self.microvm_id, + "request_id": park.request_id if park else None, + "phase": self._phase, + "local_generation": self._generation, + "active_tools": len(self._tools), + "active_activities": self._activities, + "progress_failed": self._progress_failed, + "suspend_ineligible": self._suspend_ineligible, + } + + async def wait_until_open(self) -> None: + # A thread-safe local predicate works across the server and pipeline's + # separate event loops. No AWS calls or new approval timeout are needed. + while True: + with self._lock: + if self._check_open(): + return + await asyncio.sleep(0.02) + + async def tool_started(self, tool_use_id: str | None) -> None: + while True: + await self.wait_until_open() + with self._lock: + if not self._check_open(): + continue + if not tool_use_id or tool_use_id in self._tools: + # Preserve existing tool behavior, but unknown/duplicate + # identities cannot prove that every parallel tool stopped. + self._suspend_ineligible = True + else: + self._tools.add(tool_use_id) + return + + def tool_finished(self, tool_use_id: str | None) -> None: + with self._lock: + self._tools.discard(tool_use_id or "") + + def disable_suspend(self) -> None: + """Retain normal execution when detached work cannot be accounted for.""" + with self._lock: + self._suspend_ineligible = True + + def park_approval( + self, + request_id: str, + tool_use_id: str | None, + deadline: ApprovalDeadline, + *, + record: ApprovalRecord | None = None, + ) -> ApprovalPark | None: + with self._lock: + if ( + not request_id + or not tool_use_id + or tool_use_id not in self._tools + or self._park is not None + or self._phase != "active" + ): + self._suspend_ineligible = True + return None + park = ApprovalPark( + self.task_id, self.microvm_id, request_id, tool_use_id, deadline, record + ) + self._park = park + # A new gate cannot acknowledge a wake using the previous gate's + # cached result. Its own suspend must establish a fresh safe point. + self._last_resume_park = None + # Local "parked" keeps this worker alive at a safe tool boundary. + # The coordinator's persisted PARKED continuation means source retirement. + self._phase = "parked" + return park + + async def leave_approval(self, park: ApprovalPark) -> None: + """Remove the safe point *before* the hook changes task state/returns.""" + with self._lock: + if self._continuation_park is park: + raise LifecycleUnavailable( + "Continuation must claim task ownership before releasing tools" + ) + while True: + await self.wait_until_open() + with self._lock: + if not self._check_open(): + continue + if self._park is not park: + raise LifecycleUnavailable("Approval park changed") + self._park = None + self._phase = "active" + return + + @contextmanager + def activity(self, *, checkpoint_safe: bool = False): + """Drain synchronous progress and heartbeat work before suspension. + + Best-effort writers must explicitly report missing/failed acknowledgments + with progress_write_failed(). A later successful event cannot recover a + dropped earlier event, so that failure remains latched for the task. + Progress events may finish during a continuation upload: they do not + modify the workspace, and capture drains their acknowledgments again + before publishing. Other activity remains paused during that upload. + """ + with self._lock: + capture_progress = checkpoint_safe and self._phase == "checkpointing" + if ( + not self._check_open() + and self._phase != "checkpoint-ready" + and not capture_progress + ): + # Do not silently write with pre-wake credentials. The progress + # writer catches this like its existing best-effort failures. + raise LifecycleUnavailable("Guest activity is paused") + self._activities += 1 + try: + yield + finally: + with self._lock: + self._activities -= 1 + + @asynccontextmanager + async def approval_poll(self): + """Pause new approval reads and drain an already-running read safely.""" + while True: + with self._lock: + if self._check_open() or self._phase == "checkpoint-ready": + self._activities += 1 + break + await asyncio.sleep(0.02) + try: + yield + finally: + with self._lock: + self._activities -= 1 + + def progress_write_failed(self) -> None: + with self._lock: + self._progress_failed = True + + @staticmethod + def _budget(seconds: float) -> float: + if not math.isfinite(seconds) or seconds <= 0: + raise ValueError("Lifecycle budget must be finite and positive") + return time.monotonic() + seconds + + @asynccontextmanager + async def continuation_checkpoint(self, park: ApprovalPark, *, drain_budget_s: float = 5): + """Hold the sole pending tool while saving a replacement-worker checkpoint. + + This runs on the SDK loop before decision polling starts. It is separate + from the short HTTP suspend hook: a full workspace transfer may take + minutes. Cancellation keeps the barrier closed because a cancelled + ``to_thread`` capture can still be reading the workspace. + """ + end = self._budget(drain_budget_s) + with self._lock: + if ( + self._park is not park + or self._phase != "parked" + or self._tools != {park.tool_use_id} + or self._suspend_ineligible + or self._progress_failed + ): + raise LifecycleUnavailable( + "Task is not safely parked for a continuation checkpoint" + ) + self._phase = "checkpointing" + self._generation += 1 + generation = self._generation + cancelled = False + complete = False + try: + while True: + with self._lock: + self._assert_transition(generation, "checkpointing") + if self._progress_failed: + raise LifecycleUnavailable("Progress was not acknowledged") + drained = self._activities == 0 + if drained: + break + if time.monotonic() >= end: + raise TimeoutError("Progress did not drain before continuation capture") + await asyncio.sleep(0.02) + yield + # SDK progress can arrive while the full workspace upload awaits + # a thread. Do not drop it or publish before its write is known. + # The upload can take minutes, so this drain gets its own budget. + end = self._budget(drain_budget_s) + while True: + with self._lock: + self._assert_transition(generation, "checkpointing") + if self._progress_failed: + raise LifecycleUnavailable("Progress was not acknowledged") + if self._activities == 0: + self._continuation_park = park + complete = True + break + if time.monotonic() >= end: + raise TimeoutError("Progress did not drain after continuation capture") + await asyncio.sleep(0.02) + except BaseException as exc: + cancelled = not isinstance(exc, Exception) + raise + finally: + with self._lock: + if self._generation == generation and self._phase == "checkpointing": + self._phase = ( + "failed" if cancelled else ("checkpoint-ready" if complete else "parked") + ) + self._generation += 1 + + async def release_continuation(self, park: ApprovalPark) -> None: + """Open tools only after a conditional RUNNING claim by this same worker.""" + while True: + with self._lock: + self._check_open() + if self._continuation_park is not park or self._park is not park: + raise LifecycleUnavailable("Continuation approval changed") + if self._phase == "checkpoint-ready": + self._continuation_park = None + self._park = None + self._phase = "active" + return + await asyncio.sleep(0.02) + + async def suspend( + self, checkpoint: Callable[[ApprovalPark], None], *, budget_s: float + ) -> ApprovalPark: + """Drain progress, then require an acknowledged gate/checkpoint check. + + ``checkpoint`` must synchronously verify durable task/gate identity and + deadline and raise on any failed or uncertain write/read. Returning None + means acknowledged success, never a best-effort event method. + """ + end = self._budget(budget_s) + with self._lock: + park = self._park + if ( + self._phase == "suspend-ready" + and park is not None + and not self._progress_failed + and park.deadline.remaining_s() > 0 + ): + # An already-acknowledged HTTP retry neither checkpoints again + # nor opens the barrier. Expired/unsafe retries remain closed. + return park + if ( + park is None + or self._phase not in {"parked", "checkpoint-ready"} + or self._suspend_ineligible + or park.request_id == self._slept_request_id + or self._progress_failed + or self._tools != {park.tool_use_id} + or park.deadline.remaining_s() <= 0 + ): + raise LifecycleUnavailable("Task is not safely parked for suspend") + self._phase = "suspending" + self._generation += 1 + generation = self._generation + try: + with lifecycle_stage("activity-drain"): + while True: + with self._lock: + self._assert_transition(generation, "suspending") + if self._progress_failed: + raise LifecycleUnavailable("Progress was not acknowledged") + drained = self._activities == 0 + if drained: + break + if time.monotonic() >= end: + raise TimeoutError("Progress did not drain within lifecycle budget") + await asyncio.sleep(min(0.02, max(0, end - time.monotonic()))) + with lifecycle_stage("checkpoint"): + await self._run_bounded(checkpoint, park, end) + with self._lock: + self._assert_transition(generation, "suspending") + if self._progress_failed or park.deadline.remaining_s() <= 0: + raise LifecycleUnavailable("Suspend checkpoint is no longer safe") + self._phase = "suspend-ready" + return park + except BaseException: + # An unacknowledged checkpoint never authorizes a freeze. In-flight + # progress may finish later; new suspension stays off after failure. + with self._lock: + if self._generation == generation and self._phase == "suspending": + self._phase = ( + "checkpoint-ready" if self._continuation_park is park else "parked" + ) + self._suspend_ineligible = True + self._generation += 1 + raise + + async def resume( + self, refresh_and_reconcile: Callable[[ApprovalPark], None], *, budget_s: float + ) -> ApprovalPark: + """Release the original wait only after credentials and state are safe. + + Callback ownership is intentionally narrow: refresh credentials and read + the existing gate, without deciding it, changing its deadline, or calling + leave_approval. Timeout/cancellation closes the barrier permanently; a + late thread must not release coding. The supervisor owns termination. + """ + end = self._budget(budget_s) + with self._lock: + park = self._park + if ( + self._phase in {"active", "parked", "checkpoint-ready"} + and self._last_resume_park is not None + ): + # A duplicate wake acknowledgment cannot renew the approval + # timeout or re-run credential refresh on an executing task. + return self._last_resume_park + if self._phase != "suspend-ready" or park is None: + raise LifecycleUnavailable("No acknowledged suspend to resume") + self._phase = "resuming" + self._generation += 1 + generation = self._generation + try: + with lifecycle_stage("refresh-and-reconcile"): + await self._run_bounded(refresh_and_reconcile, park, end) + reseed_random() + with self._lock: + self._assert_transition(generation, "resuming") + self._phase = "checkpoint-ready" if self._continuation_park is park else "parked" + # One sleep per approval gate. A duplicate suspend must not + # race the newly released decision loop. + self._slept_request_id = park.request_id + self._last_resume_park = park + return park + except BaseException: + with self._lock: + if self._generation == generation and self._phase == "resuming": + self._phase = "failed" + self._generation += 1 + raise + + def _assert_transition(self, generation: int, phase: str) -> None: + if self._generation != generation or self._phase != phase: + raise LifecycleUnavailable("Lifecycle transition was superseded") + + @staticmethod + async def _run_bounded( + callback: Callable[[ApprovalPark], None], park: ApprovalPark, end: float + ) -> None: + remaining = end - time.monotonic() + if remaining <= 0: + raise TimeoutError("Lifecycle budget expired") + await asyncio.wait_for(asyncio.to_thread(callback, park), timeout=remaining) + + def close(self) -> None: + """Invalidate in-flight callbacks before pipeline teardown.""" + with self._lock: + self._phase = "closed" + self._generation += 1 + self._park = None + self._continuation_park = None + self._tools.clear() + + +_registry_lock = threading.Lock() +_contexts: dict[str, MicrovmLifecycle] = {} + + +def register_task( + task_id: str, microvm_id: str, *, attempt_id: str | None = None +) -> MicrovmLifecycle: + with _registry_lock: + if _contexts: + raise LifecycleUnavailable("A MicroVM pipeline is already registered") + context = MicrovmLifecycle(task_id, microvm_id, attempt_id=attempt_id) + _contexts[task_id] = context + return context + + +def get_context(task_id: str | None) -> MicrovmLifecycle | None: + with _registry_lock: + return _contexts.get(task_id or "") + + +def get_registered_context() -> MicrovmLifecycle | None: + """The service hook belongs to the sole task registered by this VM's /run.""" + with _registry_lock: + return next(iter(_contexts.values()), None) + + +def unregister_task(context: MicrovmLifecycle) -> None: + with _registry_lock: + context.close() + if _contexts.get(context.task_id) is context: + del _contexts[context.task_id] diff --git a/agent/src/models.py b/agent/src/models.py index 8c170f97c..4dd2dee6d 100644 --- a/agent/src/models.py +++ b/agent/src/models.py @@ -185,24 +185,21 @@ class TaskConfig(BaseModel): # scheme is unchanged; read-only enforcement no longer keys off this # principal — it keys off ``read_only`` below. policy_principal: str = "new_task" - # Whether the resolved workflow is read-only (may not mutate the working - # tree). Threaded into the Cedar request ``context.read_only`` so the - # hard-deny Write/Edit rules fire for *any* read-only workflow, and drives the - # runner's allowed_tools tightening. + # The workflow's read-only policy flag. Threaded into Cedar's + # ``context.read_only`` to hard-deny Write/Edit, and used by the runner to + # remove those tools from SDK auto-approval. Other tools still depend on + # their own policy rules; this flag does not remove Bash from the surface. read_only: bool = False - # The SDK tool surface for this task, from the resolved workflow's - # ``agent_config.allowed_tools``. This is the second enforcement layer - # alongside ``read_only``: ``run_agent`` passes it to - # ``ClaudeAgentOptions.allowed_tools`` verbatim, and drops ``Write``/``Edit`` - # when ``read_only`` is true. Empty list means "fall back to the built-in - # full surface" so legacy/batch callers that never resolved a workflow keep - # working unchanged; a workflow that wants to restrict tools MUST declare a - # non-empty list (every shipped workflow does). + # SDK auto-approval list from ``agent_config.allowed_tools``. The runner + # drops Write/Edit when read_only is true; an empty list falls back to the + # built-in list for legacy/batch callers. This does not restrict available + # tools: unlisted tools fall through to the SDK permission mode. Actual + # restrictions require disallowed_tools or applicable Cedar forbid rules. allowed_tools: list[str] = Field(default_factory=list) - # Whether the resolved workflow requires a repo. False for repo-less - # knowledge workflows: the pipeline skips clone/build/PR and drives the agent - # + deliver_artifact steps through the workflow runner. Defaults True so - # coding tasks (and any caller that omits it) keep the repo-bound path. + # Whether the workflow requires a repo. False makes the repo optional: + # the pipeline skips repository setup only when no repo was supplied. + # A supplied repo still takes the repository-bound path. Defaults True for + # coding tasks and callers that omit the field. requires_repo: bool = True # True when the resolved workflow operates on an existing PR (pr_* coding # workflows) — gates the "resume existing branch / resolve PR" behavior that diff --git a/agent/src/payload_bootstrap.py b/agent/src/payload_bootstrap.py new file mode 100644 index 000000000..84b39e999 --- /dev/null +++ b/agent/src/payload_bootstrap.py @@ -0,0 +1,205 @@ +"""Authenticate deployment settings before reading one task's payload. + +The ambient worker role can GetObject only under its deployment's bootstrap/ +prefix, with an explicit deny outside that prefix (including public buckets). +Only the coordinator can write those manifests. The payload itself is fetched +without worker credentials, using a single-object capability. Never log it. +""" + +from __future__ import annotations + +import hashlib +import json +import os +import re +from typing import Any +from urllib.error import HTTPError +from urllib.parse import parse_qs, urlsplit +from urllib.request import HTTPRedirectHandler, ProxyHandler, build_opener + +from shared_constants import SHARED_CONSTANTS + +CONTRACT = SHARED_CONSTANTS["payload_bootstrap"] + + +class PayloadFetchError(RuntimeError): + """Authenticated bootstrap or payload bytes could not be read; contains no URL.""" + + +class _NoRedirect(HTTPRedirectHandler): + def redirect_request(self, req, fp, code, msg, headers, newurl): + return None + + +def _object(raw: bytes, label: str) -> dict: + try: + value = json.loads(raw) + except (UnicodeError, ValueError): + raise PayloadFetchError(f"{label} is not valid JSON") from None + if not isinstance(value, dict): + raise PayloadFetchError(f"{label} must contain an object") + return value + + +def _manifest(uri: str, backend: str) -> tuple[str, dict]: + parsed = urlsplit(uri) + if ( + parsed.scheme != "s3" + or not re.fullmatch(r"[a-z0-9][a-z0-9.-]{1,61}[a-z0-9]", parsed.netloc) + or parsed.query + or parsed.fragment + ): + raise ValueError("bootstrap_s3_uri must identify a deployment manifest") + key = parsed.path.removeprefix("/") + expected = re.fullmatch(re.escape(CONTRACT["manifest_prefix"]) + r"([a-f0-9]{64})\.json", key) + if not expected: + raise ValueError("bootstrap_s3_uri has an invalid manifest key") + + # Do not use a caller-supplied endpoint or default cached boto3 session. + from botocore.config import Config + + from aws_session import platform_client + + try: + client = platform_client( + "s3", + config=Config(connect_timeout=5, read_timeout=10, retries={"max_attempts": 2}), + ) + response = client.get_object(Bucket=parsed.netloc, Key=key) + body = response["Body"] + try: + length = response.get("ContentLength") + if not isinstance(length, int) or not 0 < length <= CONTRACT["max_manifest_bytes"]: + raise PayloadFetchError("deployment manifest has an invalid length") + raw = body.read(CONTRACT["max_manifest_bytes"] + 1) + if len(raw) != length: + raise PayloadFetchError("deployment manifest body is incomplete") + finally: + body.close() + except PayloadFetchError: + raise + except Exception as exc: + # SDK/HTTP exception text can contain URLs. Expose only the class. + raise PayloadFetchError(f"deployment manifest read failed ({type(exc).__name__})") from None + if hashlib.sha256(raw).hexdigest() != expected[1]: + raise PayloadFetchError("deployment manifest digest does not match its key") + manifest = _object(raw, "deployment manifest") + if manifest.get("version") != CONTRACT["version"] or manifest.get("backend") != backend: + raise ValueError("deployment manifest version/backend does not match this worker") + if not isinstance(manifest.get("platform_config"), dict): + raise ValueError("deployment manifest has no platform configuration") + return parsed.netloc, manifest["platform_config"] + + +def _payload_url(url: str, bucket: str, task_id: str, attempt_id: str | None = None) -> None: + """Permit only the exact object's regional S3 HTTPS endpoint, without redirects.""" + try: + parsed = urlsplit(url) + query = parse_qs(parsed.query, keep_blank_values=True, strict_parsing=True) + if any(len(values) != 1 for values in query.values()): + raise ValueError + _access_key, _date, region, service, terminator = query["X-Amz-Credential"][0].split("/") + if ( + service != "s3" + or terminator != "aws4_request" + or not re.fullmatch(r"[a-z]{2}(?:-[a-z]+)+-\d", region) + or query["X-Amz-Algorithm"] != ["AWS4-HMAC-SHA256"] + or query["X-Amz-SignedHeaders"] != ["host"] + or not re.fullmatch(r"[a-f0-9]{64}", query["X-Amz-Signature"][0]) + or not 0 < int(query["X-Amz-Expires"][0]) <= CONTRACT["url_ttl_seconds"] + or parsed.scheme != "https" + or parsed.username is not None + or parsed.password is not None + or parsed.port is not None + or parsed.fragment + ): + raise ValueError + suffix = "amazonaws.com.cn" if region.startswith("cn-") else "amazonaws.com" + host = f"s3.{region}.{suffix}" + prefix = f"{task_id}/{attempt_id}" if attempt_id is not None else task_id + valid = ( + parsed.netloc == f"{bucket}.{host}" and parsed.path == f"/{prefix}/payload.json" + ) or (parsed.netloc == host and parsed.path == f"/{bucket}/{prefix}/payload.json") + if not valid: + raise ValueError + except (KeyError, IndexError, TypeError, ValueError): + raise ValueError("payload reference must sign this task's exact S3 object") from None + + +def _download(url: str) -> dict: + # Disable environment proxies and redirects. No AWS credential provider is + # involved; S3 authorizes the coordinator's signature on this one object. + opener = build_opener(ProxyHandler({}), _NoRedirect()) + try: + with opener.open(url, timeout=10) as response: + length = int(response.headers.get("Content-Length", "0")) + if not 0 < length <= CONTRACT["max_payload_bytes"]: + raise PayloadFetchError("task payload has an invalid length") + raw = response.read(CONTRACT["max_payload_bytes"] + 1) + if len(raw) != length: + raise PayloadFetchError("task payload body is incomplete") + return _object(raw, "task payload") + except PayloadFetchError: + raise + except HTTPError as exc: + raise PayloadFetchError(f"task payload download returned HTTP {exc.code}") from None + except Exception as exc: + raise PayloadFetchError(f"task payload download failed ({type(exc).__name__})") from None + + +def resolve_payload_reference(reference: Any, backend: str) -> tuple[dict, dict]: + if not isinstance(reference, dict) or reference.get("version") != CONTRACT["version"]: + raise ValueError( + "payload bootstrap v2 is required; deploy a matching coordinator and image" + ) + task_id = reference.get("task_id") + attempt_id = reference.get("attempt_id") + uri = reference.get("bootstrap_s3_uri") + url = reference.get("payload_url") + if ( + not isinstance(task_id, str) + or not re.fullmatch(r"[A-Za-z0-9_-]{1,128}", task_id) + or not isinstance(uri, str) + or not isinstance(url, str) + or ( + attempt_id is not None + and ( + backend != "lambda-microvm" + or not isinstance(attempt_id, str) + or not re.fullmatch(r"[A-Za-z0-9_-]{1,128}", attempt_id) + ) + ) + ): + raise ValueError("payload reference is missing task identity or download coordinates") + bucket, config = _manifest(uri, backend) + _payload_url(url, bucket, task_id, attempt_id) + document = _download(url) + payload = document.get("agent_payload") + if ( + document.get("version") != CONTRACT["version"] + or document.get("task_id") != task_id + or not isinstance(payload, dict) + or payload.get("task_id") != task_id + or (attempt_id is not None and payload.get("attempt_id") != attempt_id) + ): + raise ValueError("downloaded payload does not belong to the referenced task") + if document.get("platform_config") != config: + raise ValueError( + "payload configuration does not match the authenticated deployment manifest" + ) + return payload, config + + +def load_ecs_payload() -> dict: + """Consume the capability before task subprocesses can inherit the environment.""" + raw = os.environ.pop("AGENT_PAYLOAD_REF", "") + if not raw: + raise ValueError("AGENT_PAYLOAD_REF is required; deploy a matching coordinator and image") + try: + reference = json.loads(raw) + except (UnicodeError, ValueError) as exc: + raise ValueError("AGENT_PAYLOAD_REF is not valid JSON") from exc + if not isinstance(reference, dict) or reference.get("task_id") != os.environ.get("TASK_ID"): + raise ValueError("payload reference does not match the ECS task identity") + payload, _ = resolve_payload_reference(reference, "ecs") + return payload diff --git a/agent/src/pipeline.py b/agent/src/pipeline.py index 709301d7c..b5bfdc39c 100644 --- a/agent/src/pipeline.py +++ b/agent/src/pipeline.py @@ -214,6 +214,9 @@ def _execute_agent_step( hydrated, trajectory, progress, + *, + continuation=None, + started_reaction_id=None, ): """Run the agentic step through the workflow step runner. @@ -249,7 +252,22 @@ def _execute_agent_step( # and post-hooks stay on the inline path. only_kinds keeps the runner from # re-running the deterministic steps the pipeline already owns (double clone # / double PR). - result = run_workflow(wf, ctx, only_kinds={"run_agent"}) + from continuation_runtime import bind_runtime, create_runtime + from continuation_storage import ContinuationContext + + runtime = continuation or create_runtime( + ContinuationContext( + setup, + prompt, + system_prompt, + wf.id, + wf.version, + started_reaction_id=started_reaction_id, + ), + config.task_id, + ) + with bind_runtime(runtime): + result = run_workflow(wf, ctx, only_kinds={"run_agent"}) if ctx.agent_result is None: # The run_agent step did not produce a result — i.e. its handler raised @@ -299,6 +317,21 @@ def _run_repoless_task( workflow_id = (config.resolved_workflow or {}).get("id", "default/agent-v1") wf = load_workflow(workflow_id) system_prompt = build_repoless_system_prompt(config, hc, system_prompt_overrides) + from continuation_runtime import bind_runtime, prepare_repoless_runtime + from microvm_lifecycle import get_context as get_microvm_context + + continuation = prepare_repoless_runtime( + config, + user_prompt=prompt, + system_prompt=system_prompt, + workflow_id=wf.id, + workflow_version=wf.version, + ) + if continuation is not None and continuation.restored is not None: + if continuation.resume_prompt is None: + raise RuntimeError("Restored continuation has no recorded decision prompt") + prompt = continuation.resume_prompt + system_prompt = continuation.context.system_prompt ctx = StepContext( workflow=wf, @@ -306,7 +339,7 @@ def _run_repoless_task( hydrated=hc, progress=progress, trajectory=trajectory, - setup=None, # repo-less: no RepoSetup + setup=continuation.context.setup if continuation is not None else None, system_prompt=system_prompt, user_prompt=prompt, ) @@ -314,8 +347,12 @@ def _run_repoless_task( # deliver_artifact. The deliverer uploads the agent's result text to # artifacts/{task_id}/ (and/or surfaces it as a comment), so the declared # terminal outcome is actually produced (#248 Phase 3). - with task_span("task.agent_execution"): + with bind_runtime(continuation), task_span("task.agent_execution"): wf_result = run_workflow(wf, ctx) + lifecycle = get_microvm_context(config.task_id) + if lifecycle and lifecycle.diagnostic_snapshot()["phase"] in {"closed", "failed"}: + log("TASK", "Worker execution is closed; coordinator owns continuation or cleanup.") + return {"task_id": config.task_id, "status": "parked"} agent_result = ctx.agent_result if agent_result is None: @@ -429,10 +466,21 @@ def _run_repoless_task( print_metrics(result_dict) terminal_status = "COMPLETED" if overall_status == "success" else "FAILED" - task_state.write_terminal(config.task_id, terminal_status, result_dict) + _persist_finished_task(config.task_id, terminal_status, result_dict) return result_dict +def _persist_finished_task(task_id: str, status: str, result: dict) -> None: + outcome = task_state.write_terminal(task_id, status, result) + if outcome == task_state.TerminalWriteOutcome.FAILED: + raise task_state.TerminalWriteError( + f"Task result was not committed ({outcome.value}); " + "inspect the task record and worker lease" + ) + # A cancel or another terminal writer won the status race. Its result stays + # authoritative; this is not a worker crash and must not emit a failure reaction. + + def _apply_post_hook_gates( workflow: Workflow | None, *, @@ -903,6 +951,7 @@ def run_task( repo=config.repo_url, task_id=config.task_id, ) + task_state.verify_worker_lease(config.task_id) # Surface the credential-scoping posture once per task so every task's logs # state plainly whether tenant-data isolation was active. is_scoped() # resolves the session; if scoping was requested but unbuildable it raises @@ -1058,6 +1107,9 @@ def _on_trace_truncated(max_bytes: int, first_dropped: int) -> None: os.environ["GIT_COMMITTER_EMAIL"] = "bgagent@noreply.github.com" os.environ["GITHUB_TOKEN"] = config.github_token os.environ["GH_TOKEN"] = config.github_token + from continuation_runtime import restore_for_task + + continuation = restore_for_task(config) # Set env vars for the prepare-commit-msg hook BEFORE setup_repo() # so the hook has access to TASK_ID/PROMPT_VERSION from the start. @@ -1094,17 +1146,23 @@ def _on_trace_truncated(max_bytes: int, first_dropped: int) -> None: # pr-review) never transition — the orchestration panel owns the # parent's state, and a planning run shouldn't advance the issue. linear_transition_state = not config.read_only - linear_eyes_reaction_id = react_task_started( - config.channel_source, - config.channel_metadata, - transition_state=linear_transition_state, + linear_eyes_reaction_id = ( + continuation.context.started_reaction_id + if continuation is not None + else react_task_started( + config.channel_source, + config.channel_metadata, + transition_state=linear_transition_state, + ) ) # "Starting" comment on the Jira issue through the Forge app actor # (or legacy OAuth fallback). No-op for non-Jira tasks. # Best-effort; failures are logged, never block. workflow_id = (config.resolved_workflow or {}).get("id", "coding/new-task-v1") - if _should_post_start_comment(config.channel_source, workflow_id): + if continuation is None and _should_post_start_comment( + config.channel_source, workflow_id + ): comment_task_started( config.channel_source, config.channel_metadata, @@ -1116,10 +1174,11 @@ def _on_trace_truncated(max_bytes: int, first_dropped: int) -> None: # Part of the Early-ACK block (moved before setup_repo with the 👀 # and start comment) so board state updates immediately, not after # the multi-minute baseline build. - transition_task_started( - config.channel_source, - config.channel_metadata, - ) + if continuation is None: + transition_task_started( + config.channel_source, + config.channel_metadata, + ) # Setup repo (deterministic pre-hooks). A failure/timeout/OOM in the # pre-agent baseline build raises here; it needs no local handler — @@ -1130,7 +1189,13 @@ def _on_trace_truncated(max_bytes: int, first_dropped: int) -> None: # with no visible signal; posting the 👀 earlier is what makes the # outer handler's ❌-swap actually visible for setup failures. with task_span("task.repo_setup") as setup_span: - setup = setup_repo(config, progress=progress) + if continuation is None: + setup = setup_repo(config, progress=progress) + else: + from repo import prepare_restored_repo + + setup = continuation.context.setup + prepare_restored_repo(setup.repo_dir) setup_span.set_attribute("build.before", setup.build_before) progress.write_agent_milestone( "repo_setup_complete", @@ -1178,9 +1243,11 @@ def _on_trace_truncated(max_bytes: int, first_dropped: int) -> None: "asset (ADR-016 — the agent must have no Linear tools)", ) - # Download attachments from S3 (version-pinned, integrity-verified) + # Recovery already restored .attachments and the prompt's exact + # references. Re-downloading would depend on the original attachment + # bucket's shorter retention and could overwrite saved local work. prepared_attachments: list = [] - if config.attachments: + if config.attachments and continuation is None: from attachments import download_attachments try: @@ -1221,6 +1288,17 @@ def _on_trace_truncated(max_bytes: int, first_dropped: int) -> None: if prepared_attachments: prompt = _inject_attachment_context(prompt, prepared_attachments) + if continuation is not None: + if continuation.resume_prompt is None: + raise RuntimeError("Restored continuation has no recorded decision prompt") + prompt = continuation.resume_prompt + system_prompt = continuation.context.system_prompt + progress.write_agent_milestone( + "continuation_restored", + "Saved conversation and workspace restored; " + "continuing the recorded human decision.", + ) + # Run agent disk_before = get_disk_usage(AGENT_WORKSPACE) start_time = time.time() @@ -1247,6 +1325,8 @@ def _on_trace_truncated(max_bytes: int, first_dropped: int) -> None: hc, trajectory, progress, + continuation=continuation, + started_reaction_id=linear_eyes_reaction_id, ) except Exception as e: # Fatal agent error: mirror to APPLICATION_LOGS so @@ -1259,6 +1339,15 @@ def _on_trace_truncated(max_bytes: int, first_dropped: int) -> None: agent_span.set_status(StatusCode.ERROR, str(e)) agent_span.record_exception(e) agent_result = AgentResult(status="error", error=str(e)) + from microvm_lifecycle import get_context as get_microvm_context + + lifecycle = get_microvm_context(config.task_id) + if lifecycle and lifecycle.diagnostic_snapshot()["phase"] in {"closed", "failed"}: + # A transferred/cancelled worker may finish unwinding its SDK. + # It must not run deterministic commit/PR/channel post-hooks. + log("TASK", "Worker execution is closed; coordinator owns continuation or cleanup.") + return {"task_id": config.task_id, "status": "parked"} + progress.write_agent_milestone( "agent_execution_complete", f"status={agent_result.status} turns={agent_result.turns}", @@ -1746,7 +1835,7 @@ def _on_trace_truncated(max_bytes: int, first_dropped: int) -> None: # Persist terminal state to DynamoDB terminal_status = "COMPLETED" if overall_status == "success" else "FAILED" - task_state.write_terminal(config.task_id, terminal_status, result_dict) + _persist_finished_task(config.task_id, terminal_status, result_dict) return result_dict diff --git a/agent/src/policy.py b/agent/src/policy.py index a18aa3842..3d6abd8bd 100644 --- a/agent/src/policy.py +++ b/agent/src/policy.py @@ -12,9 +12,9 @@ 2. Approval allowlist fast-path (tool_type, tool_group, bash_pattern, write_path, all_session scopes). Skips human prompt for pre-approved patterns. - 2.5 Recent-decision cache: same (tool_name, input_sha256) within 60s of a - DENIED/TIMED_OUT outcome auto-denies. Session-scoped, cleared on - container restart (§12.8). + 2.5 Recent-decision cache: the same (tool_name, input_sha256) auto-denies + for 60s after inserting a DENIED/TIMED_OUT outcome. Ordinary restarts + clear it; continuation restores the saved decision before SDK startup. 3. Soft-deny Cedar eval (agent/policies/soft_deny.cedar + blueprint soft). Match → REQUIRE_APPROVAL with merged annotations; rule-scope allowlist match → ALLOW; no match → fall through to step 4. @@ -26,10 +26,11 @@ ``context.file_path``, because Cedar entity UIDs cannot contain arbitrary characters. -**Annotations** expected on every rule in hard_deny/soft_deny files -(§5.2): ``@rule_id`` (globally unique, kebab/snake_case), ``@tier`` -("hard"|"soft"), ``@approval_timeout_s`` (int seconds ≥ 30; soft-deny -only), ``@severity`` ("low"|"medium"|"high"; soft-deny only), ``@category`` +**Annotations** used by rules in hard_deny/soft_deny files (§5.2): +``@rule_id`` (globally unique, kebab/snake_case), ``@tier`` +("hard"|"soft"), optional ``@approval_timeout_s`` (int seconds ≥ 30; +soft-deny only; omission uses the task setting), ``@severity`` +("low"|"medium"|"high"; soft-deny only), ``@category`` (free-form; UX grouping). Annotation recovery goes through cedarpy's ``policies_to_json_str()``; the round-trip contract is locked by ``tests/test_cedarpy_annotations_contract.py``. @@ -95,7 +96,7 @@ def _validate_constants() -> None: raise ValueError( f"contracts/constants.json: approval_timeout_s.min must be > 0, got {FLOOR_TIMEOUT_S}" ) - if DEFAULT_TASK_TIMEOUT_S < FLOOR_TIMEOUT_S: + if DEFAULT_TASK_TIMEOUT_S != 0 and DEFAULT_TASK_TIMEOUT_S < FLOOR_TIMEOUT_S: raise ValueError( f"contracts/constants.json: approval_timeout_s.default ({DEFAULT_TASK_TIMEOUT_S}) " f"must be >= min ({FLOOR_TIMEOUT_S})" @@ -399,6 +400,19 @@ def rule_ids(self) -> frozenset[str]: """Snapshot of rule-ID scopes, checked post-soft-deny in the engine.""" return frozenset(self._rule_ids) + def snapshot_scopes(self) -> tuple[str, ...]: + """Export normalized session grants for an acknowledged continuation.""" + scopes = ["all_session"] if self._all_session else [] + for prefix, values in ( + ("tool_type", self._tool_types), + ("tool_group", self._tool_groups), + ("rule", self._rule_ids), + ("bash_pattern", self._bash_patterns), + ("write_path", self._write_path_patterns), + ): + scopes.extend(f"{prefix}:{value}" for value in sorted(set(values))) + return tuple(scopes) + def matches(self, tool_name: str, tool_input: dict) -> bool: """Return True if a non-rule scope pre-approves this tool call.""" if self._all_session: @@ -438,8 +452,9 @@ class RecentDecisionCache: DENIED/TIMED_OUT — NEVER on APPROVED (so a just-approved call does not auto-deny on the next identical invocation). - **Session-scoped**: cleared on container restart. Documented caveat in - §12.8 — not a bug. Persistent cache is §17.5 future work. + **Session-scoped**: ordinary process restarts clear this cache. A saved + MicroVM continuation seeds its recorded denial or timeout during restoration; + it does not serialize the entire cache. """ def __init__( @@ -684,7 +699,8 @@ def _merge_annotations( ) -> tuple[list[str], int, str]: """Merge annotations across multiple matching soft-deny policies (§6.3). - Timeout: min across rules (clamped by FLOOR_TIMEOUT_S). Severity: max. + Timeout: shortest positive rule/task value (clamped by FLOOR_TIMEOUT_S), + or zero when no deadline is configured. Severity: max. rule_ids preserved in order of match. If a matching rule has no annotation data (shouldn't happen post-validation), falls back to the policy ID. @@ -698,15 +714,16 @@ def _merge_annotations( if rule is None: continue rule_ids.append(rule.rule_id or pid) - if rule.approval_timeout_s is not None: + if rule.approval_timeout_s is not None and rule.approval_timeout_s > 0: timeouts.append(rule.approval_timeout_s) severities.append(rule.severity or _DEFAULT_SEVERITY) - timeouts.append(task_default_timeout_s) + if task_default_timeout_s > 0: + timeouts.append(task_default_timeout_s) # Defensive only: load-time validation already rejects below-floor values, # but the clamp costs nothing and protects against a future caller that # bypasses validation (e.g. programmatic rule injection). - effective_timeout = max(FLOOR_TIMEOUT_S, min(timeouts)) + effective_timeout = max(FLOOR_TIMEOUT_S, min(timeouts)) if timeouts else 0 if severities: effective_severity = max(severities, key=lambda s: _SEVERITY_ORDER.get(s, 0)) @@ -1220,7 +1237,10 @@ def evaluate_tool_use(self, tool_name: str, tool_input: dict) -> PolicyDecision: # engine pure — policy.py never calls the progress writer. return PolicyDecision( outcome=Outcome.DENY, - reason=f"Recent {cached.decision} within {int(CACHE_TTL_S)}s: {cached.reason}", + reason=( + f"Recorded {cached.decision} at {cached.original_decision_ts}: " + f"{cached.reason}" + ), duration_ms=(time.monotonic() - start) * 1000, cache_hit_metadata={ "tool_name": tool_name, @@ -1273,8 +1293,8 @@ def evaluate_tool_use(self, tool_name: str, tool_input: dict) -> PolicyDecision: return PolicyDecision( outcome=Outcome.DENY, reason=( - f"Recent {cached.decision} on rule {matched_rule_id!r} " - f"within {int(CACHE_TTL_S)}s: {cached.reason}" + f"Recorded {cached.decision} on rule {matched_rule_id!r} " + f"at {cached.original_decision_ts}: {cached.reason}" ), duration_ms=(time.monotonic() - start) * 1000, cache_hit_metadata={ diff --git a/agent/src/progress_writer.py b/agent/src/progress_writer.py index f7dbea692..13a651724 100644 --- a/agent/src/progress_writer.py +++ b/agent/src/progress_writer.py @@ -360,7 +360,7 @@ def _reset_circuit_breakers() -> None: class _ProgressWriter: """Write AG-UI-style progress events to the existing DynamoDB TaskEventsTable. - Fail-open: a DDB write failure is logged but never raises. After + Ordinary event methods fail open: a DDB write failure is logged but never raises. After ``_MAX_FAILURES`` consecutive *transient* failures the task's stream is permanently disabled (circuit breaker). Permanent errors (``ValidationException`` et al.) drop the individual event without @@ -454,7 +454,71 @@ def _ensure_table(self): # -- core write ------------------------------------------------------------ + def write_microvm_checkpoint( + self, *, metadata: dict, condition_checks: list[dict], client + ) -> None: + """Atomically acknowledge a pause marker and its task/gate preconditions. + + Called only by the lifecycle checkpoint callback after activity drains. + This deliberately bypasses the best-effort event path: absent tables, + disabled progress, failed conditions and uncertain writes must raise. + A saved marker records a safe point, not proof that AWS froze the VM. + """ + if not self._table_name or self._disabled: + raise RuntimeError("Checkpoint progress table is unavailable") + from boto3.dynamodb.types import TypeSerializer + + now = datetime.now(UTC) + item = { + "task_id": self._task_id, + "event_id": _generate_ulid(), + "event_type": "agent_milestone", + "metadata": {"milestone": "microvm_suspend_checkpoint", **metadata}, + "timestamp": now.isoformat(), + "ttl": int(now.timestamp()) + _TTL_SECONDS, + "user_id": self._user_id, + } + if self._repo: + item["repo"] = self._repo + serializer = TypeSerializer() + client.transact_write_items( + TransactItems=[ + *({"ConditionCheck": check} for check in condition_checks), + { + "Put": { + "TableName": self._table_name, + "Item": {key: serializer.serialize(value) for key, value in item.items()}, + "ConditionExpression": "attribute_not_exists(task_id)", + } + }, + ] + ) + def _put_event(self, event_type: str, metadata: dict) -> None: + from microvm_lifecycle import LifecycleUnavailable, get_context + + lifecycle = get_context(self._task_id) + if lifecycle is None: + self._put_event_best_effort(event_type, metadata) + return + try: + with lifecycle.activity(checkpoint_safe=True): + acknowledged = False + try: + acknowledged = self._put_event_best_effort(event_type, metadata) + finally: + if not acknowledged: + lifecycle.progress_write_failed() + except LifecycleUnavailable: + lifecycle.progress_write_failed() + print( + f"[progress] lifecycle barrier closed — event not acknowledged " + f"task_id={self._task_id} event_type={event_type} " + f"phase={lifecycle.diagnostic_snapshot()['phase']}", + flush=True, + ) + + def _put_event_best_effort(self, event_type: str, metadata: dict) -> bool: """Write a single progress event item to DynamoDB. Error handling splits three ways: @@ -472,12 +536,12 @@ def _put_event(self, event_type: str, metadata: dict) -> None: louder ERROR level so unexpected codes surface in reviews. """ if not self._table_name or self._disabled: - return + return False try: self._ensure_table() if self._table is None: self._disabled = True - return + return False now = datetime.now(UTC) # Correlation envelope (#245): trace_id is read per-event from the @@ -509,6 +573,7 @@ def _put_event(self, event_type: str, metadata: dict) -> None: # for the rest of the task (see ``_SharedCircuitBreaker`` # docstring). _CIRCUIT_BREAKERS.record_success(self._task_id) + return True except ImportError: self._disabled = True @@ -556,7 +621,7 @@ def _put_event(self, event_type: str, metadata: dict) -> None: f"({exc_type}: {code}); breaker NOT incremented: {e}", flush=True, ) - return + return False if classification == "transient": new_count, now_disabled = _CIRCUIT_BREAKERS.record_failure( @@ -575,7 +640,7 @@ def _put_event(self, event_type: str, metadata: dict) -> None: f"{self._MAX_FAILURES}, transient): {exc_type}: {e}", flush=True, ) - return + return False # Unknown: count like transient but flag loudly so operators # can add the new code to the classifier next release. @@ -596,6 +661,7 @@ def _put_event(self, event_type: str, metadata: dict) -> None: f"adding {exc_type} to the classifier: {e}", flush=True, ) + return False # -- public event methods -------------------------------------------------- diff --git a/agent/src/repo.py b/agent/src/repo.py index 19f2aa7ec..6b8010202 100644 --- a/agent/src/repo.py +++ b/agent/src/repo.py @@ -672,6 +672,20 @@ def _merge_predecessor_branch(repo_dir: str, pred_branch: str, notes: list[str]) log("SETUP", f"Predecessor merge conflicted, aborted: {pred_branch}") +def prepare_restored_repo(repo_dir: str) -> None: + """Recreate trusted tools/hooks without cloning or replacing saved build baselines.""" + run_cmd( + ["git", "config", "--global", "--add", "safe.directory", repo_dir], + label="safe-directory", + ) + for cfg in [repo_dir, *_find_mise_configs(repo_dir)]: + run_cmd(["mise", "trust", cfg], label="mise-trust-restored", cwd=repo_dir, check=False) + result = run_cmd(["mise", "install"], label="mise-install-restored", cwd=repo_dir, check=False) + if result.returncode != 0: + log("WARN", f"Restored workspace mise install failed (exit {result.returncode})") + _install_commit_hook(repo_dir) + + def _install_commit_hook(repo_dir: str) -> None: """Install the prepare-commit-msg git hook for Task-Id/Prompt-Version trailers.""" try: diff --git a/agent/src/runner.py b/agent/src/runner.py index d4d83c8d3..ac6138b42 100644 --- a/agent/src/runner.py +++ b/agent/src/runner.py @@ -73,10 +73,13 @@ def _setup_bedrock_cost_attribution(config: TaskConfig) -> None: 1. **Per-user/repo chargeback (CUR 2.0 / Cost Explorer).** Write the SessionRole ARN + ``{user_id, repo, task_id}`` STS tags to a 0600 file - that ``bedrock_creds_helper.py`` reads. Claude Code's managed-settings - ``awsCredentialExport`` runs that helper and signs Bedrock requests with - the tagged assumed-role credentials. Skipped when ``AGENT_SESSION_ROLE_ARN`` - is unset (local/dev) — the helper then fails open to ambient creds. + that ``bedrock_creds_helper.py`` reads on AgentCore/ECS. Claude Code's + managed-settings ``awsCredentialExport`` runs that helper and signs + Bedrock requests with the tagged assumed-role credentials. MicroVM + instead uses the parent's scoped container-credential provider; its + export helper returns no credentials and cannot supply a wake barrier. + File writing is skipped when ``AGENT_SESSION_ROLE_ARN`` is unset + (local/dev); the non-MicroVM helper can fall back to ambient credentials. 2. **Per-call forensics (model-invocation logs).** Set ``X-Amzn-Bedrock-Request-Metadata`` via ``ANTHROPIC_CUSTOM_HEADERS`` on the @@ -290,6 +293,11 @@ def _initialize_policy_engine_and_hooks( extra_policies=cedar_policies if cedar_policies else None, **engine_kwargs, ) + from continuation_runtime import current_runtime + + continuation = current_runtime() + if continuation is not None: + continuation.seed_policy(policy_engine) # Surface the resolved cap + its source so operators can distinguish a # blueprint-threaded value from the engine's compile-time default on a # container restart. Mirrors the ``approval_gate_cap_source`` field on the @@ -631,19 +639,52 @@ def _on_stderr(line: str) -> None: **({"mcp_servers": mcp_servers} if mcp_servers else {}), ) - result = AgentResult() + from continuation_runtime import current_runtime + + continuation = current_runtime() + if continuation is not None: + options.session_store = continuation.store + options.session_store_flush = "eager" + if continuation.restored is not None: + options.resume = continuation.restored["session_id"] + options.max_turns = config.max_turns - continuation.context.turns_used + if options.max_turns <= 0: + raise RuntimeError("Task turn limit was reached before the saved continuation") + if config.max_budget_usd is not None: + options.max_budget_usd = config.max_budget_usd - continuation.prior_cost_usd + if options.max_budget_usd <= 0: + raise RuntimeError( + "Task dollar budget was reached before the saved continuation" + ) + + prior_turns = ( + continuation.context.turns_used + if continuation is not None and continuation.restored is not None + else 0 + ) + result = AgentResult(turns=prior_turns) message_counts = {"system": 0, "assistant": 0, "result": 0, "other": 0} # Use ClaudeSDKClient (connect/query/receive_response) instead of the # standalone query() function. This matches the official AWS sample: # https://github.com/aws-samples/sample-deploy-ClaudeAgentSDK-based-agents-to-AgentCore-Runtime client = ClaudeSDKClient(options=options) - log("AGENT", "Connecting to Claude Code CLI subprocess...") - await client.connect() - log("AGENT", "Connected. Sending prompt...") - await client.query(prompt=prompt) - log("AGENT", "Prompt sent. Receiving messages...") + if continuation is not None: + continuation.client = client + from microvm_credentials import ScopedCredentialBroker + from microvm_lifecycle import get_context + + broker = None try: + lifecycle = get_context(config.task_id or "") + if lifecycle is not None: + broker = ScopedCredentialBroker(lifecycle) + options.env.update(broker.environment) + log("AGENT", "Connecting to Claude Code CLI subprocess...") + await client.connect() + log("AGENT", "Connected. Sending prompt...") + await client.query(prompt=prompt) + log("AGENT", "Prompt sent. Receiving messages...") async for message in client.receive_response(): if isinstance(message, SystemMessage): message_counts["system"] += 1 @@ -747,7 +788,9 @@ def _on_stderr(line: str) -> None: subtype = getattr(message, "subtype", "unknown") result.status = subtype result.cost_usd = getattr(message, "total_cost_usd", None) - result.num_turns = getattr(message, "num_turns", 0) + if result.cost_usd is not None and continuation is not None: + result.cost_usd += continuation.prior_cost_usd + result.num_turns = getattr(message, "num_turns", 0) + prior_turns result.duration_ms = getattr(message, "duration_ms", 0) result.duration_api_ms = getattr(message, "duration_api_ms", 0) result.session_id = getattr(message, "session_id", "") or "" @@ -779,6 +822,13 @@ def _on_stderr(line: str) -> None: if raw_usage is not None: # Handle both object (dataclass) and dict forms usage = _parse_token_usage(raw_usage) + if continuation is not None: + usage = TokenUsage( + **{ + key: count + continuation.prior_token_usage.get(key, 0) + for key, count in usage.model_dump().items() + } + ) result.usage = usage if all(v == 0 for v in usage.model_dump().values()): log( @@ -797,8 +847,8 @@ def _on_stderr(line: str) -> None: log( "DONE", - f"status={result.status} turns={message.num_turns} " - f"cost=${message.total_cost_usd or 0:.4f} " + f"status={result.status} turns={result.num_turns} " + f"cost=${result.cost_usd or 0:.4f} " f"duration={message.duration_ms / 1000:.1f}s", ) if message.is_error and message.result: @@ -811,8 +861,8 @@ def _on_stderr(line: str) -> None: # Write trajectory result summary (use effective status after is_error remap) trajectory.write_result( subtype=result.status, - num_turns=getattr(message, "num_turns", 0), - cost_usd=getattr(message, "total_cost_usd", None), + num_turns=result.num_turns, + cost_usd=result.cost_usd, duration_ms=getattr(message, "duration_ms", 0), duration_api_ms=getattr(message, "duration_api_ms", 0), session_id=getattr(message, "session_id", ""), @@ -823,10 +873,10 @@ def _on_stderr(line: str) -> None: input_toks = usage.input_tokens if usage else 0 output_toks = usage.output_tokens if usage else 0 progress.write_agent_cost_update( - cost_usd=getattr(message, "total_cost_usd", None), + cost_usd=result.cost_usd, input_tokens=input_toks, output_tokens=output_toks, - turn=getattr(message, "num_turns", 0), + turn=result.num_turns, ) elif isinstance(message, UserMessage): @@ -837,6 +887,12 @@ def _on_stderr(line: str) -> None: if isinstance(message.content, list): for block in message.content: if isinstance(block, ToolResultBlock): + # A later project hook can deny a call our pre-hook + # allowed without emitting either post-tool hook. + # The CLI's error result retires only that call; + # unrelated active/background work remains fenced. + if lifecycle is not None and block.is_error: + lifecycle.tool_finished(block.tool_use_id) status, content = _format_tool_result(block) log("RESULT", f"[{status}] {truncate(content)}") tool_name = tool_use_id_to_name.get( @@ -864,13 +920,21 @@ def _on_stderr(line: str) -> None: # see it on the dashboard widget + ``bgagent status`` and not # just on the runtime-DEFAULT stream. log_error_cw( - f"Exception during receive_response(): {type(e).__name__}: {e}", + f"Exception during Claude session: {type(e).__name__}: {e}", task_id=config.task_id or None, ) progress.write_agent_error(error_type=type(e).__name__, message=str(e)) if result.status == "unknown": result.status = "error" - result.error = f"receive_response() failed: {e}" + result.error = f"Claude session failed: {e}" + finally: + # Also cover startup/query failure and cancellation, before pipeline + # teardown removes the task's lifecycle registry entry. + try: + await client.disconnect() + finally: + if broker is not None: + broker.close() log("AGENT", f"Generator finished. Messages received: {message_counts}") log("AGENT", f"CLI stderr lines received: {stderr_line_count}") diff --git a/agent/src/server.py b/agent/src/server.py index 50942d8f8..e62b0ab3d 100644 --- a/agent/src/server.py +++ b/agent/src/server.py @@ -30,24 +30,20 @@ import task_state from config import resolve_github_token +from microvm_http import microvm_resume, microvm_suspend +from microvm_lifecycle import ( + LifecycleUnavailable, + get_context, + get_registered_context, + register_task, + reseed_random, + unregister_task, +) from models import TaskResult from observability import propagate_correlation_context from pipeline import run_task from shared_constants import SHARED_CONSTANTS -# --- _debug_cw / _warn_cw failure counter ------------------------------- -# Shared counter for BOTH the debug and warn CloudWatch writers. AgentCore -# doesn't forward container stdout to APPLICATION_LOGS, so a broken writer -# is invisible except for this metric. Single counter = single alarm -# surface — the trade-off is that the alarm can't distinguish which writer -# is broken (see Chunk 7c review notes). Defined BEFORE any function that -# references it (including ``_debug_cw`` / ``_warn_cw``) so the ordering is -# import-time safe: a daemon thread spawned from a write-blocking function -# can never race with module-level globals still being assigned. -_debug_cw_failures = 0 -_debug_cw_failures_lock = threading.Lock() -_DEBUG_CW_FAILURE_EMIT_EVERY = 5 - # Only redact secrets at least this long — replacing very short strings # would mangle unrelated text that happens to contain them. _MIN_REDACTABLE_SECRET_LEN = 12 @@ -86,6 +82,25 @@ def _emit_stdout_line(stamped: str) -> None: pass +def _report_cloudwatch_failure(writer: str, task_id: str | None, exc: Exception) -> None: + """Emit a stdout fallback without another AWS call or the failed message. + + This is a structured log, not a metric or configured alarm. Its availability + depends on the backend collecting guest stdout; AgentCore APPLICATION_LOGS + does not automatically forward it. + """ + _emit_stdout_line( + json.dumps( + { + "event": "cloudwatch_write_failed", + "writer": writer, + "task_id": task_id, + "error_type": type(exc).__name__, + } + ) + ) + + def _debug_cw(msg: str, *, task_id: str | None = None) -> None: """Write a debug line to a CloudWatch stream in a background thread. @@ -141,9 +156,9 @@ def _warn_cw(msg: str, *, task_id: str | None = None) -> None: The stdout emission is preserved so local ``docker-compose`` runs and the ``capfd``-based unit tests still observe the line. - CloudWatch delivery is fire-and-forget — failures bump the - shared ``_debug_cw_failures`` counter via ``_warn_cw_write_blocking`` - so a silently broken writer still surfaces via that single metric. + CloudWatch delivery is fire-and-forget. Failures emit a structured + ``cloudwatch_write_failed`` stdout record without retrying through the + failed logging path. No metric or alarm is installed by this helper. """ # Redact cached credentials and emit via the same os.write path as # ``_debug_cw``: warn messages can embed payload fragments, so they @@ -172,9 +187,8 @@ def _warn_cw_write_blocking(log_group: str, task_id: str | None, stamped: str) - Mirrors ``_debug_cw_write_blocking`` but writes to the ``server_warn/`` stream so warn-level traffic is easy to - alarm on independently of debug breadcrumbs. Failures bump the - shared ``_debug_cw_failures`` counter — a single alarm surface - covers both writers. + filter independently of debug breadcrumbs. Failures emit the shared + structured stdout fallback without another AWS call. """ try: from aws_session import platform_client @@ -191,14 +205,8 @@ def _warn_cw_write_blocking(log_group: str, task_id: str | None, stamped: str) - logStreamName=stream, logEvents=[{"timestamp": int(_time_for_debug.time() * 1000), "message": stamped}], ) - except Exception as _exc: - global _debug_cw_failures - with _debug_cw_failures_lock: - _debug_cw_failures += 1 - print( - f"[server/warn/self] CloudWatch write failed: {type(_exc).__name__}: {_exc}", - flush=True, - ) + except Exception as exc: + _report_cloudwatch_failure("warn", task_id, exc) def _debug_cw_write_blocking(log_group: str, task_id: str | None, stamped: str) -> None: @@ -218,16 +226,9 @@ def _debug_cw_write_blocking(log_group: str, task_id: str | None, stamped: str) logStreamName=stream, logEvents=[{"timestamp": int(_time_for_debug.time() * 1000), "message": stamped}], ) - except Exception as _exc: - # Never let debug logging break the request path. Bump the failure - # counter so operators can alarm on a blind debug path. - global _debug_cw_failures - with _debug_cw_failures_lock: - _debug_cw_failures += 1 - print( - f"[server/debug/self] CloudWatch write failed: {type(_exc).__name__}: {_exc}", - flush=True, - ) + except Exception as exc: + # Logging failures must not break the request or recursively log to AWS. + _report_cloudwatch_failure("debug", task_id, exc) # Log the active event loop policy at import time. @@ -264,8 +265,8 @@ def filter(self, record: logging.LogRecord) -> bool: _last_ping_status: str = "" # Heartbeat cadence for the TaskTable ``agent_heartbeat_at`` writer thread. -# Each live pipeline bumps the heartbeat every N seconds so operators can -# distinguish a stuck pipeline from a healthy long-running one. +# The independent worker reports process/writer liveness. A stuck pipeline can +# leave this thread running, so a fresh heartbeat does not prove work is progressing. _HEARTBEAT_INTERVAL_SECONDS = 45 @@ -273,7 +274,16 @@ def _heartbeat_worker(task_id: str, stop: threading.Event) -> None: """Periodically refresh ``agent_heartbeat_at`` so the orchestrator can detect crashes.""" while not stop.wait(timeout=_HEARTBEAT_INTERVAL_SECONDS): try: - task_state.write_heartbeat(task_id) + lifecycle = get_context(task_id) + if lifecycle: + with lifecycle.activity(): + task_state.write_heartbeat(task_id) + else: + task_state.write_heartbeat(task_id) + except LifecycleUnavailable: + # An approval waiter has no liveness obligation while frozen. + # Resume must refresh credentials before another heartbeat writes. + continue except Exception as e: print( f"[heartbeat] write_heartbeat error (will retry): {type(e).__name__}: {e}", @@ -417,6 +427,8 @@ def _run_task_background( workload_access_token: str = "", attachments: list[dict] | None = None, resolved_assets: list[dict] | None = None, + microvm_id: str = "", + attempt_id: str = "", ) -> None: """Run the agent task in a background thread.""" global _background_pipeline_failed @@ -456,18 +468,24 @@ def _run_task_background( task_id=task_id, ) + lifecycle = ( + register_task(task_id, microvm_id, attempt_id=attempt_id or task_id) if microvm_id else None + ) + if lifecycle is not None: + # Linux uv defaults to cache hardlinks, which checkpoint capture rejects + # because another path can mutate the same inode outside the workspace. + os.environ.setdefault("UV_LINK_MODE", "copy") stop_heartbeat = threading.Event() hb_thread: threading.Thread | None = None - if task_id: - hb_thread = threading.Thread( - target=_heartbeat_worker, - args=(task_id, stop_heartbeat), - name=f"heartbeat-{task_id}", - daemon=True, - ) - hb_thread.start() - try: + if task_id: + hb_thread = threading.Thread( + target=_heartbeat_worker, + args=(task_id, stop_heartbeat), + name=f"heartbeat-{task_id}", + daemon=True, + ) + hb_thread.start() # Propagate the correlation envelope into this thread's OTEL context # so spans are correlated with the AgentCore session and the platform # identity in CloudWatch (#245). Runs whenever any field is present — @@ -521,6 +539,8 @@ def _run_task_background( ) task_state.write_terminal(task_id, "FAILED", backup.model_dump()) finally: + if lifecycle: + unregister_task(lifecycle) stop_heartbeat.set() if hb_thread is not None and hb_thread.is_alive(): hb_thread.join(timeout=3) @@ -829,35 +849,17 @@ async def invoke_agent(request: Request, body: InvocationRequest): # -------------------------------------------------------------------------- -# AWS Lambda MicroVMs lifecycle hooks (ADR-021 P1 + P2) +# AWS Lambda MicroVMs lifecycle hooks (ADR-021 P1 through P3) # -------------------------------------------------------------------------- -# The MicroVM backend has NO orchestrator→agent HTTP path: the task payload -# arrives as the ``/run`` hook's request body and nothing else dials in. The -# service calls these routes on the port declared in the image's ``hooks.port`` -# (8080 — the same uvicorn process that serves /invocations and /ping), so the -# hooks live here rather than in a sidecar. -# -# Four hooks are served; ``/suspend`` + ``/resume`` are still P3: -# * ``/ready`` (build, P1) is MANDATORY. ``CreateMicrovmImage`` refuses an image -# that enables ANY lifecycle hook without it ("The ready (/ready) MicroVM -# image hook must be enabled when any MicroVM lifecycle hook … is enabled"), -# and an image with no hooks at all cannot receive a ``runHookPayload``. So -# ADR-021's original "declare /run in P1, serve it in P2" split was not a -# reachable service state. -# * ``/run`` (runtime, P1) is the payload-delivery channel — and, since P2, the -# platform-configuration channel (see ``platform_config`` below). -# * ``/validate`` (build, P2) is the snapshot self-check. It runs under the -# BUILD role and makes ZERO AWS calls — see ``microvm_validate``. -# * ``/terminate`` (runtime, P2) is a best-effort final flush. It never writes -# terminal task status — the orchestrator owns terminal state. -# ``/suspend`` and ``/resume`` are P3 (they need the ComputeStrategy interface -# widening). Declaring a hook the agent does not answer fails the corresponding -# build or lifecycle transition, which is why the construct declares exactly the -# hooks served here. +# The service calls six hooks on the image listener. Build hooks warm/check the +# snapshot without AWS access; /run authenticates configuration and starts work. +# /terminate closes the coding barrier and acknowledges teardown. /suspend and +# /resume checkpoint and reconcile credentials/gates. Automatic suspension also +# requires coordinator approval of the launched image version and live settings. MICROVM_HOOK_PREFIX = "/aws/lambda-microvms/runtime/v1" -#: ``s3://`` scheme prefix for the out-of-band payload pointer. -_S3_URI_SCHEME = "s3://" +app.add_api_route(f"{MICROVM_HOOK_PREFIX}/suspend", microvm_suspend, methods=["POST"]) +app.add_api_route(f"{MICROVM_HOOK_PREFIX}/resume", microvm_resume, methods=["POST"]) # --- platform_config allowlist (ADR-021 P2) -------------------------------- # WHY the agent's platform env arrives in the ``/run`` payload at all, instead of @@ -909,70 +911,15 @@ async def invoke_agent(request: Request, body: InvocationRequest): #: WHAT THIS BUYS, STATED PRECISELY — because the honest answer is narrower than #: "stops secret exfiltration", and overstating it would hide the residual gap. #: -#: The key allowlist above stops a payload setting ``LD_PRELOAD``; it does not stop -#: a payload pointing an *allowlisted* key at a different resource. The value that -#: matters most is ``github_token_secret_arn``: ``config.resolve_github_token`` -#: fetches whatever ARN it names using the UNSCOPED execution role and caches the -#: raw ``SecretString`` into ``os.environ["GITHUB_TOKEN"]``, from which ``shell.py`` -#: hands the environment to every repo subprocess — i.e. into the model's tool -#: surface. So the value is worth validating. -#: -#: This check is DEFENCE IN DEPTH AND FAIL-FAST, not the primary control: -#: -#: * What actually stops a cross-account read today is IAM. Every grant on the -#: execution role is account-scoped by construction — ``grantRead`` on the GitHub -#: PAT secret, and the ``bgagent-linear-oauth-*`` / ``bgagent-jira-oauth-*`` -#: prefix grants built with ``stack.formatArn`` (``lambda-microvm-compute.ts``). -#: A foreign-account ARN therefore AccessDenies with or without this check. What -#: this adds is a structured 400 at the door instead of an opaque -#: ``AccessDeniedException`` mid-startup, and a guard that still holds if a -#: future grant is ever widened. -#: * What this check does NOT stop is an IN-ACCOUNT redirect. The channel-OAuth -#: grants are prefix grants (unavoidable: the CLI mints ``bgagent-*-oauth-`` -#: at setup, so the names are unknown at synth), so a block whose anchor and -#: whose ``github_token_secret_arn`` name the SAME account but a DIFFERENT -#: workspace's OAuth secret is accepted here. Partition + account is the only -#: boundary this check enforces; it is not a per-workspace authorization check. -#: * That residual gap is currently unreachable from the guest, which is why it is -#: left open. ``platform_config`` is produced by the orchestrator Lambda and the -#: MicroVM execution role holds ``grantRead`` ONLY on the payload bucket -#: (``payloadBucket.grantRead``), so a running MicroVM can read another task's -#: payload but cannot write one. Reaching the in-account case requires already -#: controlling the orchestrator's environment or the bucket's write path. -#: -#: ESCALATION: if the payload path ever becomes less trusted — a third-party -#: producer, an operator-editable envelope, or any grant that lets the guest write -#: the payload bucket — this must grow into a name-shape check that ties -#: ``github_token_secret_arn`` to the task's own channel/workspace, because -#: partition+account pinning provably does not cover that case. -#: -#: The prefix grant itself is at ECS parity. The ASYMMETRY that makes value -#: validation worth doing here at all is new to this backend: on ECS these ARNs -#: arrive as deploy-time container env; here they arrive in a network payload. +#: The allowlist restricts environment-variable names; ARN shape/account checks +#: below are defense in depth. Startup settings come from an IAM-authenticated +#: bootstrap manifest read before the task payload. That independent manifest +#: pins the exact GitHub secret and other deployment identifiers, including valid +#: cross-region secrets. A same-payload account anchor alone cannot do that. MICROVM_PLATFORM_CONFIG_ARN_KEYS: frozenset[str] = frozenset(_PLATFORM_CONFIG_CONTRACT["arn_keys"]) -#: The key whose ARN supplies the partition/account every other ARN is checked -#: against. -#: -#: Deliberately a payload key rather than ``os.environ`` or an STS call. -#: ``os.environ`` is empty here by construction (nothing is baked into the -#: snapshot — see ``imageEnvironmentVariables``), so anchoring on the environment -#: would silently degrade this whole check to shape-only in the intended -#: deployment. An ``sts:GetCallerIdentity`` is not available either: this runs -#: BEFORE ``platform_config`` is installed, on the path that must make zero AWS -#: calls beyond the S3 payload fetch. -#: -#: CONSEQUENCE, stated plainly: because the anchor travels in the same block as the -#: values it validates, this check enforces INTERNAL CONSISTENCY of the block, not -#: agreement with the account the guest is actually running in. A block that names -#: one foreign account throughout is self-consistent and passes here — it then -#: fails at IAM, which is the control that really holds (see -#: :data:`MICROVM_PLATFORM_CONFIG_ARN_KEYS`). ``agent_session_role_arn`` is still -#: the best available anchor: it is REQUIRED (so always present when this check -#: runs, which is what stops disarm-by-omission) and it is the one value whose -#: misdirection costs the attacker the run rather than gaining them anything — a -#: foreign session role fails closed at ``sts:AssumeRole`` -#: (``SessionScopingError``). +#: Consistency anchor after manifest authentication, not deployment identity. +#: Required membership is enforced by the shared contract. MICROVM_PLATFORM_CONFIG_ACCOUNT_ANCHOR_KEY: str = _PLATFORM_CONFIG_CONTRACT["account_anchor_key"] _PLATFORM_CONFIG_KEY_RE = re.compile(r"^[a-z][a-z0-9_]*$") @@ -1056,6 +1003,11 @@ def _validate_platform_config_contract() -> None: unknown_arn = sorted(MICROVM_PLATFORM_CONFIG_ARN_KEYS - set(MICROVM_PLATFORM_CONFIG_ENV_BY_KEY)) if unknown_arn: raise ValueError(f"{where}.arn_keys names key(s) absent from env_by_key: {unknown_arn}") + for key, env_name in MICROVM_PLATFORM_CONFIG_ENV_BY_KEY.items(): + if ( + key.endswith("_arn") or env_name.endswith("_ARN") + ) and key not in MICROVM_PLATFORM_CONFIG_ARN_KEYS: + raise ValueError(f"{where}: ARN-shaped key {key!r} is missing from arn_keys") anchor = MICROVM_PLATFORM_CONFIG_ACCOUNT_ANCHOR_KEY if anchor not in MICROVM_PLATFORM_CONFIG_ARN_KEYS: raise ValueError(f"{where}.account_anchor_key {anchor!r} must be one of arn_keys") @@ -1135,25 +1087,6 @@ def __init__(self, code: str, message: str) -> None: self.code = code -def _absent_required_platform_env() -> list[str]: - """Required ``platform_config`` env vars that are unset in the LIVE environment. - - The no-``platform_config`` path's audit. Checks the ENV VAR names rather than - the contract keys, because on that path the only possible source is whatever - the image snapshot baked — so the effective environment is the thing to - interrogate, and a value that arrived by any route counts. - - Blank/whitespace-only counts as absent, matching - :func:`_install_platform_config`'s own rule: CloudFormation renders an - unresolved value as ``""``, and an empty table name is not a table name. - """ - return sorted( - MICROVM_PLATFORM_CONFIG_ENV_BY_KEY[key] - for key in MICROVM_PLATFORM_CONFIG_REQUIRED_KEYS - if not os.environ.get(MICROVM_PLATFORM_CONFIG_ENV_BY_KEY[key], "").strip() - ) - - def _reject_foreign_arns(resolved: dict[str, str]) -> None: """Require every ARN-shaped value to agree with the anchor's partition + account. @@ -1172,9 +1105,10 @@ def _reject_foreign_arns(resolved: dict[str, str]) -> None: Region is deliberately NOT compared. Secrets Manager and IAM ARNs legitimately differ on that axis in this system — IAM is global (empty region field), and a cross-Region secret is a supported deployment shape — so requiring agreement - would reject valid configurations while adding nothing: the execution role's - grants are account-scoped, so an in-account cross-Region ARN reaches nothing - the in-Region one does not. Partition + account is the boundary this checks. + would reject valid configurations. IAM grants separately constrain accessible + resources. Partition + account consistency does not bind these identifiers + to this deployment or prevent selecting another workspace's same-account + secret where IAM permits it (#817). """ anchor_value = resolved.get(MICROVM_PLATFORM_CONFIG_ACCOUNT_ANCHOR_KEY, "") # The anchor is contract-guaranteed REQUIRED (asserted at import), so by the @@ -1227,10 +1161,10 @@ def _install_platform_config(raw: Any) -> list[str]: Returns the sorted env var names actually installed. Rules, all deliberate: - * ``None`` / absent → install nothing and return ``[]``. This is the P1 - envelope (no ``platform_config`` sibling), where the snapshot's own env is - all there is; a MicroVM image can be launched by an orchestrator that - predates Stage B, and the two deploy on independent cadences. + * ``None`` → install nothing and return ``[]`` for direct helper callers. + The v2 ``/run`` path requires a dictionary from the authenticated manifest + before calling this helper. This no-op does not accept legacy unsigned + envelopes or permit an independent coordinator/image contract rollout. * present but not an object, or carrying ANY key outside the allowlist, or carrying a non-string value, or carrying a control character in a value → reject the run (``…_INVALID``). Unknown keys are an env-injection attempt, @@ -1238,13 +1172,12 @@ def _install_platform_config(raw: Any) -> list[str]: block is refused rather than filtered. Control characters are refused for the reason in ``_PLATFORM_CONFIG_FORBIDDEN_VALUE_CHARS``. * ``None`` / blank / whitespace-only values are treated as ABSENT, not as an - instruction to clear the variable: the natural TypeScript producer - (``process.env.X ?? ''``) emits an empty string for a resource the - deployment does not have, and clobbering an image value with ``""`` would - turn "not configured over there" into "unconfigured here". + instruction to clear the variable. The current TypeScript producer omits + unconfigured values; this helper also filters explicit empty values + instead of installing them into the environment. * every required key must survive that filter, else reject - (``…_INCOMPLETE``). An explicitly-sent-but-empty ``{}`` therefore fails — - a producer with nothing to say must omit the key entirely. + (``…_INCOMPLETE``). An explicitly-sent-but-empty ``{}`` therefore fails; + the live v2 boot path always requires the complete required subset. * every ARN-shaped value must agree with the anchor ARN's partition + account, else reject (``…_INVALID``) — see :func:`_reject_foreign_arns`, which is explicit that this is internal consistency plus fail-fast, not an @@ -1368,9 +1301,8 @@ def _build_hook_log(msg: str) -> None: 1. The build role's Logs grant is scoped to the service's own ``/aws/lambda-microvms/*`` namespace, so a write to any OTHER ``LOG_GROUP_NAME`` — e.g. an APPLICATION_LOGS group baked into a legacy or - hand-built image — can only FAIL. That failure bumps the shared - ``_debug_cw_failures`` counter, i.e. such an image build would poison the - "debug path is blind" signal with a false positive. + hand-built image — can fail. The counter incremented by such a failure + is currently unexported (#810), so it provides no alarm signal. 2. ``boto3.DEFAULT_SESSION`` created during ``/ready`` freezes the BUILD role's CREDENTIALS into the snapshot, where every launched MicroVM would inherit them (region is re-resolved per client; credentials are not — see @@ -1389,11 +1321,11 @@ def _pre_config_log(msg: str) -> None: ``_install_platform_config`` has run, ``LOG_GROUP_NAME`` is whatever the snapshot happens to carry — normally nothing, but a legacy or hand-built image could bake it, and then a ``_debug_cw`` on this path would resolve credentials - and pin ``boto3.DEFAULT_SESSION`` *before* ``AGENT_SESSION_ROLE_ARN`` is in the - environment — memoizing the UNSCOPED compute-role credentials for the life of - the process, where the whole point of that variable is that every later client - is tenant-scoped. (Region and ``AWS_SDK_UA_APP_ID`` are re-resolved per client - and so are NOT at risk here; the credentials are the exposure.) The one AWS + and create ``boto3.DEFAULT_SESSION`` before configuration is installed. + Tenant-data clients use the separate, tag-scoped session in ``aws_session``; + setting ``AGENT_SESSION_ROLE_ARN`` does not scope boto3's default session. + Platform clients intentionally retain compute-role credentials. Region and + ``AWS_SDK_UA_APP_ID`` are re-resolved per client. The one AWS call this phase is allowed to make is the S3 payload fetch, because ``platform_config`` is inside the object it fetches. @@ -1456,176 +1388,27 @@ def _parse_terminate_microvm_id(raw: bytes) -> str: class MicrovmRunHookRequest(BaseModel): - """Body the MicroVM service POSTs to the ``/run`` hook. - - ``runHookPayload`` is the opaque STRING the orchestrator passed to - ``RunMicrovm`` — the service does not parse it. ABCA's contract for that - string (``lambda-microvm-strategy.ts``) is one of two shapes, mirroring the - ECS container env contract (``AGENT_PAYLOAD`` / ``AGENT_PAYLOAD_S3_URI``): - - * ``{"agent_payload": {...}, "platform_config": {...}}`` — inline. - * ``{"agent_payload_s3_uri": "s3://bucket/key", "platform_config": {...}}`` — - a pointer; the object at the URI carries the task payload (and a - ``platform_config`` copy, so either end of the fetch yields it). - - The pointer form is the DOMINANT one: the service caps ``runHookPayload`` at - 4 096 bytes and a hydrated payload is essentially always larger. - - ``platform_config`` (ADR-021 P2) is a SIBLING of ``agent_payload``, not a - field inside it: it configures the agent's *process*, whereas - ``agent_payload`` describes the *task* (``memory_id`` and friends stay - inside ``agent_payload``, unchanged). See ``_install_platform_config``. - - Both fields default to empty so a malformed call produces this module's - structured 400 rather than FastAPI's 422 — the service surfaces a 4xx as a - generic "client error" hook failure either way, and our own body is what ends - up in the MicroVM log group. + """Service request containing a serialized v2 payload reference. + + The reference identifies an IAM-authenticated deployment manifest and one + task's signed download URL. Unsigned inline and legacy S3 envelopes are + refused: this protocol requires a matching coordinator and agent image. + Missing fields reach our structured 400 instead of FastAPI's 422. """ microvmId: str = "" # service field name; camelCase on the wire runHookPayload: str = "" # service field name; camelCase on the wire -class _PayloadFetchError(Exception): - """A ``/run`` payload the agent could not READ, as opposed to could not PARSE. - - Exists purely to be *not* a ``ValueError``, because the ``/run`` handler - discriminates its 400 from its 500 on exactly that type and the two answers - make opposite promises to the operator: - - * 400 ``MICROVM_RUN_PAYLOAD_INVALID`` — "the orchestrator built a bad - envelope; retrying an identical body cannot help." - * 500 ``MICROVM_RUN_PAYLOAD_UNREADABLE`` — "the payload could not be read; - retrying CAN help." - - A truncated, racing, or half-written S3 object is the SECOND kind, but its - natural exception is ``json.JSONDecodeError`` — a ``ValueError`` subclass — so - it landed in the 400 branch and told the operator the orchestrator was at - fault when the orchestrator was fine and the object was bad. Only the - *pre-fetch* URI-shape check legitimately raises ``ValueError`` on this path, - which is why a blanket ``except ValueError`` is the wrong discriminator and - this type exists. - """ - - -def _fetch_microvm_payload_from_s3(uri: str) -> dict: - """Read and parse the out-of-band ``/run`` payload from S3. - - Same fetch the ECS boot command performs for ``AGENT_PAYLOAD_S3_URI``; the - MicroVM **execution role** holds the read grant, scoped to the platform - payload bucket. Errors propagate to the caller, which turns them into a - structured 400/500 — silently starting a pipeline with no payload would - produce a task that runs with an empty prompt. - - The URI-SHAPE check raises ``ValueError`` (the orchestrator's envelope is - wrong → 400). Everything AFTER the fetch raises :class:`_PayloadFetchError` - (the object is wrong → 500, retryable). See that class. - - Built through ``aws_session.platform_client`` so the call carries the ABCA - ``md/`` solution-attribution segment (#319). Platform, not tenant: the bucket - is platform-owned and — decisively — this is the ONE call that must happen - BEFORE ``platform_config`` is installed (the config is inside the object - being fetched), so ``AGENT_SESSION_ROLE_ARN`` may not be set yet and a - tenant-scoped client could not be built. ``platform_client`` does not touch - the cached session, so this call also cannot pin an unscoped session for the - rest of the task. The ``app/`` UA segment (native, from ``AWS_SDK_UA_APP_ID``) - is the one attribution field this single call can miss for the same - chicken-and-egg reason. - """ - remainder = uri[len(_S3_URI_SCHEME) :] - bucket, _, key = remainder.partition("/") - if not bucket or not key: - raise ValueError(f"agent_payload_s3_uri is not a bucket/key URI: {uri!r}") - - from aws_session import platform_client +def _resolve_microvm_run_payload(run_hook_payload: str) -> tuple[dict, dict]: + """Authenticate deployment settings and resolve the v2 task reference.""" + from payload_bootstrap import resolve_payload_reference - region = os.environ.get("AWS_REGION") or os.environ.get("AWS_DEFAULT_REGION") - client = platform_client("s3", region_name=region) - body = client.get_object(Bucket=bucket, Key=key)["Body"].read() try: - payload = json.loads(body) - except ValueError as exc: - # ``JSONDecodeError`` IS a ``ValueError``, so without this re-raise a - # truncated or half-written object would be reported as an orchestrator - # envelope bug and marked non-retryable. Re-raised as the type the handler - # routes to its retryable 500. - raise _PayloadFetchError( - f"S3 payload at {uri!r} is not valid JSON ({exc}); the object may be " - "truncated or still being written" - ) from exc - if not isinstance(payload, dict): - raise _PayloadFetchError( - f"S3 payload at {uri!r} is {type(payload).__name__}, expected an object" - ) - return payload - - -def _resolve_microvm_run_payload(run_hook_payload: str) -> tuple[dict, Any]: - """Split the ``runHookPayload`` string into (agent payload, platform config). - - The second element is returned RAW (unvalidated) — ``_install_platform_config`` - owns its allowlist checks so the two failure classes get distinct wire codes. - ``None`` means the envelope carried no ``platform_config`` at all. - - Raises ``ValueError`` for every ENVELOPE shape the agent cannot act on — the - caller maps that onto its 400 ("the orchestrator built this; a retry cannot - help"). Problems with the CONTENT of a fetched S3 object raise - :class:`_PayloadFetchError` instead, which the caller maps onto its retryable - 500: the orchestrator's envelope was fine and the object was not. - """ - if not run_hook_payload.strip(): - raise ValueError("runHookPayload is empty") - - try: - envelope = json.loads(run_hook_payload) - except json.JSONDecodeError as exc: - raise ValueError(f"runHookPayload is not valid JSON: {exc}") from exc - - if not isinstance(envelope, dict): - raise ValueError(f"runHookPayload must be a JSON object, got {type(envelope).__name__}") - - inline = envelope.get("agent_payload") - if inline is not None: - if not isinstance(inline, dict): - raise ValueError(f"agent_payload must be an object, got {type(inline).__name__}") - return inline, envelope.get("platform_config") - - uri = envelope.get("agent_payload_s3_uri") - if isinstance(uri, str) and uri.startswith(_S3_URI_SCHEME): - fetched = _fetch_microvm_payload_from_s3(uri) - # ``platform_config`` may sit beside the pointer (outer envelope) or - # inside the fetched object — the producer writes it in BOTH places on - # this path deliberately, so the agent gets it whichever end it reads. - # Inner first, outer as the fallback. - platform_config = fetched.get("platform_config") - if platform_config is None: - platform_config = envelope.get("platform_config") - # The fetched object is EITHER the same envelope shape as the inline form - # ({"agent_payload": …}) or the task payload itself with ``platform_config`` - # merged in at the top level (what the strategy writes today, and what P1 - # wrote without the config). Both are accepted because the image snapshot - # and the orchestrator Lambda deploy on independent cadences — a new image - # must not require a same-instant orchestrator. Discriminating on the - # ``agent_payload`` key is unambiguous: no orchestrator task payload has a - # field by that name. A stray ``platform_config`` key left in the bare - # form is inert — ``_extract_invocation_params`` reads named fields only. - nested = fetched.get("agent_payload") - if nested is None: - return fetched, platform_config - if not isinstance(nested, dict): - # Content of the FETCHED OBJECT, not of the envelope — so this is the - # retryable class, same as a truncated body. See ``_PayloadFetchError``. - raise _PayloadFetchError( - f"agent_payload in the S3 payload must be an object, got {type(nested).__name__}" - ) - return nested, platform_config - if uri is not None: - raise ValueError(f"agent_payload_s3_uri must be an s3:// URI, got {uri!r}") - - raise ValueError( - "runHookPayload envelope has neither agent_payload nor agent_payload_s3_uri " - f"(keys: {sorted(envelope)})" - ) + reference = json.loads(run_hook_payload) + except (UnicodeError, ValueError) as exc: + raise ValueError("runHookPayload is not valid JSON") from exc + return resolve_payload_reference(reference, "lambda-microvm") #: The ONE executable whose warm-up gates the snapshot, exec'd FIRST. @@ -1895,8 +1678,8 @@ def microvm_validate(): It must also not touch credential resolution: ``platform_config`` has not arrived yet (it comes with ``/run``), and any client built here would leave a - resolved boto3 session — with the build role's credentials and the build - region — frozen in the snapshot for every MicroVM launched from it. Hence + resolved boto3 session with cached build-role credentials in the snapshot + for every MicroVM launched from it (region is re-resolved per client). Hence ``_build_hook_log`` instead of ``_debug_cw``, and no import of ``aws_session``. @@ -1929,7 +1712,8 @@ def microvm_validate(): would fail every build. Names only — never values. """ expected_routes = { - f"{MICROVM_HOOK_PREFIX}/{hook}" for hook in ("ready", "validate", "run", "terminate") + f"{MICROVM_HOOK_PREFIX}/{hook}" + for hook in ("ready", "validate", "run", "terminate", "suspend", "resume") } registered = {getattr(route, "path", None) for route in app.routes} missing_routes = sorted(expected_routes - registered) @@ -1939,6 +1723,12 @@ def microvm_validate(): "hook_routes_registered": not missing_routes, "python_version_supported": sys.version_info[:2] >= _MIN_PYTHON_VERSION, "platform_config_contract_loaded": bool(MICROVM_PLATFORM_CONFIG_ENV_BY_KEY), + # Absent markers support ordinary/legacy images. A marked image must + # actually contain the protocol it advertises, before AWS snapshots it. + "image_lifecycle_protocol_supported": os.environ.get( + SHARED_CONSTANTS["microvm_lifecycle"]["image_protocol_env"] + ) + in (None, str(SHARED_CONSTANTS["microvm_lifecycle"]["protocol_version"])), } failed = sorted(name for name, ok in checks.items() if not ok) @@ -1983,44 +1773,33 @@ def microvm_validate(): return body +# The terminate body is only optional correlation data. Leave ample headroom +# inside the image's 15-second hook timeout if the service stream stalls. +_TERMINATE_BODY_BUDGET_SECONDS = 1.0 + + @app.post(f"{MICROVM_HOOK_PREFIX}/terminate") async def microvm_terminate(request: Request): - """MicroVM ``/terminate`` runtime hook — best-effort flush, always 200. - - Called as the MicroVM is torn down. Three hard constraints: - - * **It must not write terminal task status.** The orchestrator owns terminal - state: it finalizes the task and THEN calls ``TerminateMicrovm``, so a - terminate hook that wrote ``FAILED``/``COMPLETED`` would race the - finalization it follows and could clobber the real outcome with a - substrate-shutdown artifact. The pipeline thread's own crash path - (``_run_task_background``) remains the only in-guest terminal writer. - * **It must return 200 inside the hook budget, even with nothing running.** - So it never joins the pipeline thread (a drain could take minutes — that is - ``lifespan``'s job on a graceful shutdown) and every best-effort step is - wrapped: a failure here must not turn a clean teardown into a hook failure. - * **It must return 200 for any BODY too.** That is why this handler takes the - raw ``Request`` instead of a Pydantic body model: FastAPI validates a typed - body BEFORE the handler runs, so malformed JSON, a wrong content-type, or a - missing body would produce a 422 this function never gets a chance to - prevent — a reported hook failure on a successful teardown. Parsing is - deferred to ``_parse_terminate_microvm_id``, which degrades to ``""``. - - ``async def`` (unlike ``/ready`` and ``/run``) because reading the raw body - requires awaiting it. Safe on the event loop: the work is a JSON parse, a - thread-count read and a fire-and-forget log — no blocking AWS call. - - On flushing: there is nothing buffered to flush. ``_ProgressWriter`` performs - a synchronous DynamoDB ``put_item`` per event, and ``task_state`` writes - inline, so every progress/status write is already durable at call time — this - hook has no queue to drain, which is why it is a log-and-acknowledge rather - than a flush loop. (ADR-021 sub-decision 2's "flush progress events before - returning 200" applies to ``/suspend`` in P3 for the same reason: durability - is per-write, so the hook only has to observe it.) + """Close the coding barrier, log teardown and acknowledge any request body. + + This hook never joins the pipeline or writes terminal task status. Termination + can interrupt active work, so it cannot assume finalization already finished. + Raw Request parsing avoids FastAPI rejecting malformed bodies before entry. + Each best-effort step is guarded; acknowledged checkpointing belongs to + /suspend, not this hook. """ + # Close the local barrier before reading the body or emitting diagnostics. + # A slow checkpoint/refresh thread must not release coding during teardown. + try: + lifecycle = get_registered_context() + if lifecycle is not None: + lifecycle.close() + except Exception as exc: + _emit_stdout_line(f"[server/warn] /terminate barrier close failed: {type(exc).__name__}") raw = b"" try: - raw = await request.body() + async with asyncio.timeout(_TERMINATE_BODY_BUDGET_SECONDS): + raw = await request.body() except Exception as exc: # A truncated/aborted body must not become a 5xx: the VM is going away and # the id is only a correlation string. Logged, not swallowed. @@ -2079,7 +1858,7 @@ def microvm_run(request: Request, body: MicrovmRunHookRequest): mechanism ``/invocations`` uses — ``_extract_invocation_params`` → ``_validate_required_params`` → ``_spawn_background`` — rather than a second, drifting payload mapper. The orchestrator payload is byte-identical across - substrates (AgentCore receives it as ``input``, ECS as ``AGENT_PAYLOAD``, + substrates (AgentCore receives it as ``input``, ECS through the authenticated payload reference, MicroVMs inside this envelope), which is what makes that reuse correct. ``platform_config`` (P2) is installed into ``os.environ`` FIRST — before @@ -2096,15 +1875,15 @@ def microvm_run(request: Request, body: MicrovmRunHookRequest): has, per ADR-021 sub-decision 3's identity delta. Sync ``def`` for the same reason as ``/ready``, and additionally because the - S3 payload fetch is a blocking boto3 call: in a threadpool it cannot stall + manifest read and signed download are blocking calls: in a threadpool it cannot stall the event loop. **Every log line before the install goes through ``_pre_config_log``** (stdout only). Until ``platform_config`` is in the environment, a ``_debug_cw`` here would resolve AWS credentials and pin ``boto3.DEFAULT_SESSION`` off whatever the snapshot happens to carry — the same defect the build hooks avoid, one - phase later. The single AWS call this phase is allowed to make is the S3 - payload fetch, because the config is inside the object being fetched. + phase later. Pre-install operations are the IAM-authenticated manifest read + and the single-object HTTPS download. Build hooks remain AWS-silent. """ _pre_config_log( f"/run hook received: microvm_id={body.microvmId!r} bytes={len(body.runHookPayload)}" @@ -2125,16 +1904,14 @@ def microvm_run(request: Request, body: MicrovmRunHookRequest): ) except Exception as exc: # Payload could not be READ: S3 AccessDenied / NoSuchKey / transient, or a - # `_PayloadFetchError` for an object that fetched but was truncated, + # `PayloadFetchError` for an object that fetched but was truncated, # non-JSON, or not an object. 500 so the failure is distinguishable from a - # malformed ENVELOPE (the 400 above) and is correctly reported as - # retryable, and loud enough to find in the MicroVM log group — via the - # response body, since the CloudWatch writer is off-limits until the config - # is installed. - _pre_config_log( - f"/run hook payload fetch FAILED [{type(exc).__name__}: {exc}]\n" - f"{traceback.format_exc()}" - ) + # malformed reference (the 400 above). The response body preserves the + # distinction while CloudWatch is off-limits before config installation. + # Corrupt stored bytes require repair, not an assumption that retry helps. + # A chained HTTP error may contain the bearer URL. Log only the + # bootstrap reader's sanitized message, never its exception chain. + _pre_config_log(f"/run hook payload fetch FAILED [{type(exc).__name__}: {exc}]") return JSONResponse( status_code=500, content={ @@ -2146,6 +1923,11 @@ def microvm_run(request: Request, body: MicrovmRunHookRequest): task_id_log = str(payload.get("task_id", "")) try: + if platform_config is None: + raise _PlatformConfigError( + "MICROVM_RUN_PLATFORM_CONFIG_INVALID", + "Authenticated platform configuration is required", + ) installed_env = _install_platform_config(platform_config) except _PlatformConfigError as exc: _pre_config_log(f"/run hook rejected: {exc}") @@ -2163,55 +1945,6 @@ def microvm_run(request: Request, body: MicrovmRunHookRequest): f"/run hook installed platform_config env: {installed_env}", task_id=task_id_log or None, ) - else: - # No `platform_config` — the legacy P1 envelope. This branch must NOT simply - # shrug: the required keys exist because without them the agent cannot write - # status/progress, resolve the GitHub token, or (decisively) - # tenant-scope its credentials — `aws_session.get_session` falls back to the - # ambient compute role with scoping silently OFF when - # `AGENT_SESSION_ROLE_ARN` is unset. So the check is re-run against the - # EFFECTIVE environment: a legacy or hand-built image that bakes those - # values still runs (that is the compatibility this branch is for), while - # version skew — a pre-Stage-B orchestrator launching a P2 image, which - # bakes nothing — is REJECTED instead of running unscoped. - # - # STILL pre-install, so the line is stdout only. A `_warn_cw` here would - # spawn the CloudWatch writer thread and pin `boto3.DEFAULT_SESSION` off the - # snapshot's baked env, which is the very defect this branch is reporting. - # Nothing is lost: on the intended deployment (no baked `LOG_GROUP_NAME`) - # `_warn_cw` would have degraded to this same stdout line, on a legacy image - # the log group would be the wrong one anyway, and the rejection reason also - # travels in the structured response body the service surfaces. - absent = _absent_required_platform_env() - if absent: - _pre_config_log( - f"/run hook REJECTED: no platform_config and the image snapshot does " - f"not supply required value(s) either: {absent}" - + (f" task_id={task_id_log!r}" if task_id_log else "") - ) - return JSONResponse( - status_code=400, - content={ - "code": "MICROVM_RUN_PLATFORM_CONFIG_INCOMPLETE", - "message": ( - "The /run envelope carried no platform_config and the image " - "snapshot does not carry the required values either, so this " - f"MicroVM cannot run a task: {absent} are unset. This is a " - "version skew — an orchestrator predating ADR-021 P2 launching " - "a P2 image, which bakes no environment by design. Refusing " - "rather than running with tenant scoping disabled. Redeploy the " - "orchestrator so it sends platform_config." - ), - "missing_env": absent, - }, - ) - _pre_config_log( - "/run hook received no platform_config; running on the image snapshot's " - "own environment, which is frozen at build time and which DOES supply " - "every required value. Expected only from an orchestrator that predates " - "ADR-021 P2 paired with an image that bakes its own configuration." - + (f" task_id={task_id_log!r}" if task_id_log else "") - ) try: params = _extract_invocation_params(payload, request) @@ -2239,8 +1972,17 @@ def microvm_run(request: Request, body: MicrovmRunHookRequest): }, ) - _spawn_background(params) + # The attempt token comes only from the authenticated bootstrap payload. + # It identifies coordinator authority before RunMicrovm returns a VM ID. task_id = params["task_id"] + attempt_id = payload.get("attempt_id", task_id) + if not isinstance(attempt_id, str) or not re.fullmatch(r"[A-Za-z0-9_-]{1,128}", attempt_id): + return JSONResponse( + status_code=400, + content={"code": "MICROVM_ATTEMPT_ID_INVALID", "message": "Invalid worker attempt"}, + ) + reseed_random() + _spawn_background({**params, "microvm_id": body.microvmId, "attempt_id": attempt_id}) # Carries microvm_id as well as task_id: the "/run hook received" line that # used to correlate the two is stdout-only now (pre-install), so this is the # first line that reaches the task's log group and it has to join the CloudWatch diff --git a/agent/src/task_state.py b/agent/src/task_state.py index bf40d712f..caaea02d4 100644 --- a/agent/src/task_state.py +++ b/agent/src/task_state.py @@ -1,13 +1,16 @@ """Best-effort task state persistence to DynamoDB. -All writes are wrapped in try/except so a DynamoDB outage never breaks the -agent pipeline. When the TASK_TABLE_NAME environment variable is unset, all -operations are no-ops. +Progress/status writes are best-effort; approval transactions fail closed. +The coordinator creates task records and owns compute identity and capacity +reservations. This module only reads tasks and updates reporting/approval fields; +its allowed attributes are pinned by the CDK agent-task-write-attributes contract. """ import os +import random import time -from typing import TypedDict +from enum import StrEnum +from typing import NotRequired, TypedDict from shell import log, log_error_cw @@ -34,7 +37,8 @@ class ApprovalRow(TypedDict): status: str # always 'PENDING' on initial write. created_at: str timeout_s: int - ttl: int + deadline_epoch: NotRequired[int] + ttl: NotRequired[int] user_id: str repo: str @@ -64,6 +68,88 @@ def _now_iso() -> str: return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()) +def _lease_check(task_id: str, *, low_level: bool, identity=None) -> dict | None: + """Read coordinator authority without granting workers permission to rewrite it.""" + from microvm_lifecycle import get_context + from shared_constants import SHARED_CONSTANTS + + lifecycle = get_context(task_id) + if lifecycle is None or not os.environ.get("CONTINUATION_BUCKET_NAME"): + return None + values = {":attempt": lifecycle.attempt_id, ":active": "ACTIVE"} + condition = "lease_attempt_id = :attempt AND lease_state = :active" + if identity is not None: + values.update( + {":user": identity.user_id, ":repo": identity.repo, ":vm": lifecycle.microvm_id} + ) + condition += " AND lease_user_id = :user AND lease_repo = :repo AND lease_microvm_id = :vm" + task_table, _ = _require_tables() + key = {"task_id": SHARED_CONSTANTS["microvm_continuation"]["lease_key_prefix"] + task_id} + if low_level: + key = {k: _py_to_ddb_attr(v) for k, v in key.items()} + values = {k: _py_to_ddb_attr(v) for k, v in values.items()} + return { + "ConditionCheck": { + "TableName": task_table, + "Key": key, + "ConditionExpression": condition, + "ExpressionAttributeValues": values, + } + } + + +def _transact_with_conflict_retry(client, items: list): + """Retry only uncommitted transaction conflicts, never a failed ownership check.""" + from botocore.exceptions import ClientError + + max_attempts = 3 + for attempt in range(max_attempts): + try: + return client.transact_write_items(TransactItems=items) + except ClientError as error: + code = error.response.get("Error", {}).get("Code") + reasons = _extract_cancellation_reasons(error) + codes = {reason.get("Code") for reason in reasons} + conflict = code == "TransactionConflictException" or ( + code == "TransactionCanceledException" + and "TransactionConflict" in codes + and codes <= {"None", "TransactionConflict"} + ) + if not conflict or attempt == max_attempts - 1: + raise + time.sleep(random.SystemRandom().uniform(0.025, 0.075) * (2**attempt)) + + +def _transact_task(client, task_id: str, *, TransactItems: list, identity=None): + lease = _lease_check(task_id, low_level=True, identity=identity) + return _transact_with_conflict_retry(client, [*TransactItems, *([lease] if lease else [])]) + + +def _update_task(table, task_id: str, *, low_level: bool = False, **operation): + lease = _lease_check(task_id, low_level=low_level) + if lease is None: + return table.update_item(**operation) + if not low_level: + operation["TableName"] = table.name + client = table if low_level else table.meta.client + return _transact_with_conflict_retry(client, [{"Update": operation}, lease]) + + +def _task_status_conflict(error: Exception) -> bool: + """An expected status race is benign only if worker ownership still passed.""" + from botocore.exceptions import ClientError + + if not isinstance(error, ClientError): + return False + code = error.response.get("Error", {}).get("Code") + if code == "ConditionalCheckFailedException": + return True + reasons = error.response.get("CancellationReasons") or [] + return code == "TransactionCanceledException" and [ + reason.get("Code") for reason in reasons + ] == ["ConditionalCheckFailed", "None"] + + def _build_logs_url(task_id: str) -> str | None: """Build a CloudWatch Logs console URL filtered to this task_id.""" region = os.environ.get("AWS_REGION") or os.environ.get("AWS_DEFAULT_REGION") @@ -79,28 +165,32 @@ def _build_logs_url(task_id: str) -> str | None: ) -def write_submitted( - task_id: str, repo_url: str = "", issue_number: str = "", task_description: str = "" -) -> None: - """Record a task as SUBMITTED (called from the invoke script or server).""" - try: - table = _get_table() - if table is None: - return - item = { - "task_id": task_id, - "status": "SUBMITTED", - "created_at": _now_iso(), - } - if repo_url: - item["repo_url"] = repo_url - if issue_number: - item["issue_number"] = issue_number - if task_description: - item["task_description"] = task_description - table.put_item(Item=item) - except Exception as e: - log("WARN", f"[task_state] write_submitted failed (best-effort): {e}") +def verify_worker_lease(task_id: str, *, client=None) -> None: + """Reject a superseded launch before repository code or SDK tools can execute.""" + from boto3.dynamodb.types import TypeDeserializer + + from microvm_lifecycle import get_context + from shared_constants import SHARED_CONSTANTS + + lifecycle = get_context(task_id) + if lifecycle is None or not os.environ.get("CONTINUATION_BUCKET_NAME"): + return + table, _ = _require_tables() + result = _get_ddb_client(client=client).get_item( + TableName=table, + Key={ + "task_id": {"S": SHARED_CONSTANTS["microvm_continuation"]["lease_key_prefix"] + task_id} + }, + ConsistentRead=True, + ) + raw = result.get("Item") or {} + deserialize = TypeDeserializer().deserialize + lease = {key: deserialize(value) for key, value in raw.items()} + if ( + lease.get("lease_state") != "ACTIVE" + or lease.get("lease_attempt_id") != lifecycle.attempt_id + ): + raise RuntimeError("MICROVM_LEASE_LOST: this launch no longer owns task execution") def write_heartbeat(task_id: str) -> None: @@ -109,7 +199,9 @@ def write_heartbeat(task_id: str) -> None: table = _get_table() if table is None: return - table.update_item( + _update_task( + table, + task_id, Key={"task_id": task_id}, UpdateExpression="SET agent_heartbeat_at = :t", ConditionExpression="#s = :running", @@ -117,73 +209,11 @@ def write_heartbeat(task_id: str) -> None: ExpressionAttributeValues={":t": _now_iso(), ":running": "RUNNING"}, ) except Exception as e: - from botocore.exceptions import ClientError - - if ( - isinstance(e, ClientError) - and e.response.get("Error", {}).get("Code") == "ConditionalCheckFailedException" - ): + if _task_status_conflict(e): return log("WARN", f"[task_state] write_heartbeat failed (best-effort): {type(e).__name__}: {e}") -def write_session_info(task_id: str, session_id: str, agent_runtime_arn: str) -> None: - """Record session_id + agent_runtime_arn on a pre-RUNNING task. - - The orchestrator Lambda writes these fields on the HYDRATING → RUNNING - transition so ``cancel-task`` can ``StopRuntimeSession`` on the right - runtime and operators can correlate a stuck task to a specific AgentCore - session. Currently only the orchestrator calls this; the agent-side - invocation path inherits the fields from the orchestrator's payload. - - Idempotent + best-effort. Skips silently if the task is already - past SUBMITTED/HYDRATING (concurrent transition winning is fine). - """ - if not task_id or (not session_id and not agent_runtime_arn): - return - try: - table = _get_table() - if table is None: - return - set_parts: list[str] = [] - expr_values: dict = { - ":submitted": "SUBMITTED", - ":hydrating": "HYDRATING", - } - if session_id: - set_parts.append("session_id = :sid") - expr_values[":sid"] = session_id - if agent_runtime_arn: - set_parts.append("agent_runtime_arn = :arn") - set_parts.append("compute_type = :ct") - set_parts.append("compute_metadata = :cm") - expr_values[":arn"] = agent_runtime_arn - expr_values[":ct"] = "agentcore" - expr_values[":cm"] = {"runtimeArn": agent_runtime_arn} - if not set_parts: - return - table.update_item( - Key={"task_id": task_id}, - UpdateExpression="SET " + ", ".join(set_parts), - ConditionExpression="#s IN (:submitted, :hydrating)", - ExpressionAttributeNames={"#s": "status"}, - ExpressionAttributeValues=expr_values, - ) - except Exception as e: - from botocore.exceptions import ClientError - - if ( - isinstance(e, ClientError) - and e.response.get("Error", {}).get("Code") == "ConditionalCheckFailedException" - ): - # Task already advanced — concurrent legitimate transition wins. - return - log( - "WARN", - f"[task_state] write_session_info failed (best-effort): {type(e).__name__}: {e}", - ) - - def write_running(task_id: str) -> None: """Transition a task to RUNNING (called at agent start). @@ -214,7 +244,9 @@ def write_running(task_id: str) -> None: update_parts.append("logs_url = :logs") expr_values[":logs"] = logs_url - table.update_item( + _update_task( + table, + task_id, Key={"task_id": task_id}, UpdateExpression="SET " + ", ".join(update_parts), ConditionExpression="#s IN (:submitted, :hydrating)", @@ -222,27 +254,35 @@ def write_running(task_id: str) -> None: ExpressionAttributeValues=expr_values, ) except Exception as e: - from botocore.exceptions import ClientError - - if ( - isinstance(e, ClientError) - and e.response.get("Error", {}).get("Code") == "ConditionalCheckFailedException" - ): + if _task_status_conflict(e): log("INFO", "[task_state] write_running skipped: status precondition not met") return log("WARN", f"[task_state] write_running failed (best-effort): {type(e).__name__}") -def write_terminal(task_id: str, status: str, result: dict | None = None) -> None: +class TerminalWriteOutcome(StrEnum): + WRITTEN = "written" + SUPERSEDED = "superseded" + FAILED = "failed" + DISABLED = "disabled" + + +class TerminalWriteError(RuntimeError): + """The worker finished but could not commit its result to the task record.""" + + +def write_terminal(task_id: str, status: str, result: dict | None = None) -> TerminalWriteOutcome: """Transition a task to a terminal state (COMPLETED or FAILED). Updates ``status_created_at`` alongside ``status`` — see - :func:`write_running` for why. + :func:`write_running` for why. Callers must not report successful completion + when the result is FAILED or SUPERSEDED. DISABLED supports local runs without + a configured task table; crash-path callers already report failure. """ try: table = _get_table() if table is None: - return + return TerminalWriteOutcome.DISABLED now = _now_iso() expr_names = {"#s": "status"} # Mixed value types: most are strings, but build_passed/lint_passed are @@ -344,20 +384,18 @@ def write_terminal(task_id: str, status: str, result: dict | None = None) -> Non update_parts.append("artifact_uri = :au") expr_values[":au"] = result["artifact_uri"] - table.update_item( + _update_task( + table, + task_id, Key={"task_id": task_id}, UpdateExpression="SET " + ", ".join(update_parts), ConditionExpression="#s IN (:running, :hydrating, :finalizing, :awaiting_approval)", ExpressionAttributeNames=expr_names, ExpressionAttributeValues=expr_values, ) + return TerminalWriteOutcome.WRITTEN except Exception as e: - from botocore.exceptions import ClientError - - if ( - isinstance(e, ClientError) - and e.response.get("Error", {}).get("Code") == "ConditionalCheckFailedException" - ): + if _task_status_conflict(e): log( "INFO", "[task_state] write_terminal skipped: " @@ -400,8 +438,15 @@ def write_terminal(task_id: str, status: str, result: dict | None = None) -> Non f"task_id={task_id!r} after ConditionalCheckFailed " f"(terminal-state race).", ) - return - log("WARN", f"[task_state] write_terminal failed (best-effort): {type(e).__name__}") + return TerminalWriteOutcome.SUPERSEDED + # Include DynamoDB's cancellation reasons: losing worker ownership is + # not a benign status race and must never trigger trace self-healing. + log_error_cw( + f"[task_state] write_terminal failed (best-effort): {type(e).__name__}: {e}; " + f"CancellationReasons={_extract_cancellation_reasons(e)}", + task_id=task_id, + ) + return TerminalWriteOutcome.FAILED def write_trace_uri_conditional(task_id: str, uri: str) -> bool: @@ -421,7 +466,9 @@ def write_trace_uri_conditional(task_id: str, uri: str) -> bool: table = _get_table() if table is None: return False - table.update_item( + _update_task( + table, + task_id, Key={"task_id": task_id}, UpdateExpression="SET trace_s3_uri = :ts3", ConditionExpression=( @@ -439,12 +486,7 @@ def write_trace_uri_conditional(task_id: str, uri: str) -> bool: ) return True except Exception as e: - from botocore.exceptions import ClientError - - if ( - isinstance(e, ClientError) - and e.response.get("Error", {}).get("Code") == "ConditionalCheckFailedException" - ): + if _task_status_conflict(e): # Benign: URI was already persisted, or status isn't terminal yet. log( "INFO", @@ -456,7 +498,8 @@ def write_trace_uri_conditional(task_id: str, uri: str) -> bool: log( "WARN", f"[task_state] write_trace_uri_conditional failed for " - f"task_id={task_id!r}: {type(e).__name__}: {e}", + f"task_id={task_id!r}: {type(e).__name__}: {e}; " + f"CancellationReasons={_extract_cancellation_reasons(e)}", ) return False @@ -469,7 +512,7 @@ class TaskFetchError(Exception): """ -def get_task(task_id: str) -> dict | None: +def get_task(task_id: str, *, consistent_read: bool = False) -> dict | None: """Fetch a task record by ID. Returns: @@ -492,7 +535,9 @@ def get_task(task_id: str) -> dict | None: if table is None: return None try: - resp = table.get_item(Key={"task_id": task_id}) + resp = table.get_item( + Key={"task_id": task_id}, **({"ConsistentRead": True} if consistent_read else {}) + ) except Exception as e: log("WARN", f"[task_state] get_task failed: {type(e).__name__}: {e}") raise TaskFetchError(f"{type(e).__name__}: {e}") from e @@ -504,11 +549,9 @@ def get_task(task_id: str) -> dict | None: # --------------------------------------------------------------------------- # # ``TaskApprovalsTable`` and the AWAITING_APPROVAL status transitions are -# provisioned by the CDK stack. The agent-side helpers below are written to -# that contract and exposed so the ``pre_tool_use_hook`` can be implemented + -# unit-tested (via mocked boto3 clients); once the stack sets -# ``TASK_APPROVALS_TABLE_NAME`` + grants IAM, the same helpers start making -# real DDB calls with no further code change on the agent side. +# provisioned by the CDK stack. The ``pre_tool_use_hook`` uses these helpers +# with task-scoped credentials. Transactions authorize each item separately: +# Put on the approvals table and attribute-restricted Update on TaskTable. # # Primitives exposed: # - ``transact_write_approval_request`` — atomic Put(TaskApprovals) + @@ -525,10 +568,9 @@ def get_task(task_id: str) -> dict | None: # - ``get_approval_row`` — strongly-consistent GetItem; default # ``consistent_read=True`` because the race fix relies on it. # -# Errors beyond the structural conditions (unreachable DDB, IAM drift, -# missing env var) raise ``ApprovalTablesUnavailable`` so the hook can -# fail CLOSED without guessing. The hook maps that to DENY so a deploy -# without the approvals table cannot silently bypass gates. +# New deployments route creation and timeout writes through the trusted approval +# service. Missing table configuration raises ``ApprovalTablesUnavailable``; +# transport/IAM errors propagate. The hook fails closed in either case. TASK_APPROVALS_TABLE_ENV = "TASK_APPROVALS_TABLE_NAME" TASK_TABLE_ENV = "TASK_TABLE_NAME" @@ -664,6 +706,9 @@ def transact_write_approval_request( ) -> None: """Atomically record a pending approval + transition the task to AWAITING_APPROVAL. + The configured approval service performs the transaction. Direct DynamoDB + is retained only for older deployments with the legacy permission model. + Two items: 1. Put on ``TaskApprovalsTable`` with ``ConditionExpression: attribute_not_exists(request_id)`` — guards against ULID collisions @@ -680,6 +725,23 @@ def transact_write_approval_request( DDB-layer exceptions propagate so the hook's outer try/except can fail-closed with a specific reason. """ + import approval_requests + + if approval_requests.configured(): + try: + approval_requests.record_request( + "create", task_id, request_id, approval=dict(approval_row) + ) + return + except Exception as exc: + if _extract_error_code(exc) == "TransactionCanceledException": + reasons = _extract_cancellation_reasons(exc) + raise ApprovalWriteError( + f"approval write cancelled: reasons={reasons}", cancellation_reasons=reasons + ) from exc + raise + # Compatibility with unscoped legacy/local deployments. configured() rejects + # a missing service URL for cloud workers that use the session role. task_table, approvals_table = _require_tables() ddb = _get_ddb_client(client=client) @@ -690,7 +752,9 @@ def transact_write_approval_request( approval_item.setdefault("status", {"S": "PENDING"}) try: - ddb.transact_write_items( + _transact_task( + ddb, + task_id, TransactItems=[ { "Put": { @@ -715,7 +779,7 @@ def transact_write_approval_request( }, } }, - ] + ], ) except Exception as exc: # TransactionCanceledException carries per-item reasons. Keep the @@ -746,6 +810,10 @@ def transact_resume_from_approval( - resuming with a stale request_id after a race with the reconciler / a concurrent approval. + Refresh the heartbeat in the same update: writes pause during approval waits, + so restoring RUNNING with the old timestamp could let an orchestrator poll + mark this healthy task lost before the next periodic heartbeat. + Raises ``ApprovalResumeError`` on ``TransactionCanceledException`` so the hook can emit ``approval_resume_failed`` + DENY. """ @@ -753,14 +821,17 @@ def transact_resume_from_approval( ddb = _get_ddb_client(client=client) try: - ddb.transact_write_items( + _transact_task( + ddb, + task_id, TransactItems=[ { "Update": { "TableName": task_table, "Key": {"task_id": {"S": task_id}}, "UpdateExpression": ( - "SET #s = :running REMOVE awaiting_approval_request_id" + "SET #s = :running, agent_heartbeat_at = :heartbeat " + "REMOVE awaiting_approval_request_id, continuation" ), "ConditionExpression": ( "#s = :awaiting AND awaiting_approval_request_id = :rid" @@ -770,10 +841,11 @@ def transact_resume_from_approval( ":running": {"S": _STATUS_RUNNING}, ":awaiting": {"S": _STATUS_AWAITING_APPROVAL}, ":rid": {"S": request_id}, + ":heartbeat": {"S": _now_iso()}, }, } } - ] + ], ) except Exception as exc: reasons = _extract_cancellation_reasons(exc) @@ -786,6 +858,187 @@ def transact_resume_from_approval( raise +def publish_continuation_checkpoint( + identity, + receipt, + *, + tool_input_sha256: str, + cost_usd: float | None = None, + turns_used: int | None = None, + client=None, +) -> dict: + """Publish only the same worker's still-pending, fully saved approval checkpoint.""" + from dataclasses import asdict + from decimal import Decimal + + from boto3.dynamodb.types import TypeSerializer + + from continuation_storage import StorageLimits + from continuation_usage import valid_cost + from shared_constants import SHARED_CONSTANTS + + serialize = TypeSerializer().serialize + receipt.validate(identity, StorageLimits()) + if receipt.kind != "manifest": + raise ValueError("Only a complete continuation manifest may be published") + task_table, approvals_table = _require_tables() + ddb = _get_ddb_client(client=client) + record = { + "version": SHARED_CONSTANTS["microvm_continuation"]["version"], + "state": "READY", + "identity": asdict(identity), + "manifest": asdict(receipt), + } + values = { + ":request": identity.request_id, + ":awaiting": _STATUS_AWAITING_APPROVAL, + ":record": record, + ":consumed": "CONSUMED", + } + update = "SET continuation = :record" + if cost_usd is not None: + if not valid_cost(cost_usd): + raise ValueError("Continuation cost must be finite and nonnegative") + values[":cost"] = Decimal(str(cost_usd)) + update += ", cost_usd = :cost" + if turns_used is not None: + if type(turns_used) is not int or turns_used < 0: + raise ValueError("Continuation turns must be a nonnegative integer") + values[":turns"] = turns_used + update += ", turns = :turns" + _transact_task( + ddb, + identity.task_id, + identity=identity, + TransactItems=[ + { + "Update": { + "TableName": task_table, + "Key": {"task_id": {"S": identity.task_id}}, + "UpdateExpression": update, + "ConditionExpression": ( + "#status = :awaiting AND " + "awaiting_approval_request_id = :request AND " + "(attribute_not_exists(continuation) OR continuation = :record OR " + "continuation.#state = :consumed)" + ), + "ExpressionAttributeNames": {"#status": "status", "#state": "state"}, + "ExpressionAttributeValues": {k: serialize(v) for k, v in values.items()}, + } + }, + { + "ConditionCheck": { + "TableName": approvals_table, + "Key": { + "task_id": {"S": identity.task_id}, + "request_id": {"S": identity.request_id}, + }, + "ConditionExpression": ( + "user_id = :user AND repo = :repo AND #status = :pending " + "AND tool_input_sha256 = :hash" + ), + "ExpressionAttributeNames": {"#status": "status"}, + "ExpressionAttributeValues": { + ":user": {"S": identity.user_id}, + ":repo": {"S": identity.repo}, + ":pending": {"S": "PENDING"}, + ":hash": {"S": tool_input_sha256}, + }, + } + }, + ], + ) + return record + + +def consume_restored_continuation( + task_id: str, worker_id: str, record: dict, *, client=None +) -> None: + """Claim RUNNING after restoration, without deciding or revalidating the human action.""" + from boto3.dynamodb.types import TypeSerializer + + serialize = TypeSerializer().serialize + from continuation_session import CheckpointIdentity + + task_table, approvals_table = _require_tables() + identity = record["identity"] + if record.get("state") != "RESTORING" or record.get("worker_id") != worker_id: + raise ValueError("Continuation is not owned by this replacement worker") + values = { + ":request": identity["request_id"], + ":record": record, + ":awaiting": _STATUS_AWAITING_APPROVAL, + ":running": _STATUS_RUNNING, + ":consumed": "CONSUMED", + ":heartbeat": _now_iso(), + } + ddb = _get_ddb_client(client=client) + items = [ + { + "Update": { + "TableName": task_table, + "Key": {"task_id": {"S": task_id}}, + "UpdateExpression": ( + "SET #status = :running, agent_heartbeat_at = :heartbeat, " + "continuation.#state = :consumed REMOVE awaiting_approval_request_id" + ), + "ConditionExpression": ( + "continuation = :record " + "AND #status = :awaiting AND awaiting_approval_request_id = :request" + ), + "ExpressionAttributeNames": {"#status": "status", "#state": "state"}, + "ExpressionAttributeValues": {k: serialize(v) for k, v in values.items()}, + } + }, + { + "ConditionCheck": { + "TableName": approvals_table, + "Key": { + "task_id": {"S": task_id}, + "request_id": {"S": identity["request_id"]}, + }, + "ConditionExpression": ( + "user_id = :user AND #status IN (:approved, :denied, :timedout)" + ), + "ExpressionAttributeNames": {"#status": "status"}, + "ExpressionAttributeValues": { + ":user": {"S": identity["user_id"]}, + ":approved": {"S": "APPROVED"}, + ":denied": {"S": "DENIED"}, + ":timedout": {"S": "TIMED_OUT"}, + }, + } + }, + ] + try: + _transact_task(ddb, task_id, identity=CheckpointIdentity(**identity), TransactItems=items) + except Exception: + # A transport error can lose the response after DynamoDB commits. Only + # acknowledge that exact assignment, still RUNNING under this lease. + # Cancellation or a newer assignment must never become a successful retry. + from boto3.dynamodb.types import TypeDeserializer + + deserialize = TypeDeserializer().deserialize + try: + response = ddb.get_item( + TableName=task_table, + Key={"task_id": {"S": task_id}}, + ConsistentRead=True, + ) + task = {key: deserialize(value) for key, value in response.get("Item", {}).items()} + expected = {**record, "state": "CONSUMED"} + if ( + task.get("status") != _STATUS_RUNNING + or task.get("session_id") != worker_id + or task.get("continuation") != expected + or task.get("awaiting_approval_request_id") is not None + ): + raise RuntimeError("Continuation claim was not acknowledged") + verify_worker_lease(task_id, client=ddb) + except Exception: + raise + + def best_effort_update_approval_status( task_id: str, request_id: str, @@ -804,6 +1057,23 @@ def best_effort_update_approval_status( Returns ``True`` on successful write, ``False`` on ``ConditionalCheckFailedException``. All other errors propagate. """ + import approval_requests + + if new_status != "TIMED_OUT": + raise ValueError("Workers may record only non-human timeouts") + if approval_requests.configured(): + try: + approval_requests.record_request("timeout", task_id, request_id, reason=reason) + return True + except Exception as exc: + reasons = _extract_cancellation_reasons(exc) + if _extract_error_code(exc) == "TransactionCanceledException" and ( + reasons + and reasons[0].get("Code") == "ConditionalCheckFailed" + and all(reason.get("Code") == "None" for reason in reasons[1:]) + ): + return False + raise _, approvals_table = _require_tables() ddb = _get_ddb_client(client=client) @@ -903,7 +1173,10 @@ def increment_approval_gate_count_in_ddb( try: ddb = _get_ddb_client(client=client) - ddb.update_item( + _update_task( + ddb, + task_id, + low_level=True, TableName=task_table, Key={"task_id": {"S": task_id}}, UpdateExpression="ADD approval_gate_count :one", diff --git a/agent/src/workflow/runner.py b/agent/src/workflow/runner.py index ba7a7bc31..5fffc5b61 100644 --- a/agent/src/workflow/runner.py +++ b/agent/src/workflow/runner.py @@ -407,9 +407,11 @@ def _handle_hydrate_context(step: Step, ctx: StepContext) -> StepOutcome: Hydration is largely orchestrator-side today (WORKFLOWS.md open question #4 leans "orchestrator hydrates, the agent step only consumes"); this handler is - that consumer. It sets BOTH prompts the ``run_agent`` step needs: + that consumer. It fills missing prompts for the ``run_agent`` step: - - ``ctx.user_prompt`` — the hydrated ``user_prompt`` when present. + - ``ctx.user_prompt`` — the hydrated ``user_prompt`` when the caller has + not already prepared one. A restored human decision or attachment context + must survive this step. - ``ctx.system_prompt`` — built via the existing ``build_system_prompt`` so the workflow path produces the same system prompt as ``pipeline.run_task`` (repo_url/branch/workspace/max_turns/setup_notes/memory_context + overrides @@ -418,7 +420,7 @@ def _handle_hydrate_context(step: Step, ctx: StepContext) -> StepOutcome: the ``RepoSetup``; when absent (repo-less workflows) the system prompt is left to the caller, since ``build_system_prompt`` is repo-shaped today. """ - if ctx.hydrated is not None: + if ctx.hydrated is not None and not ctx.user_prompt: ctx.user_prompt = ctx.hydrated.user_prompt built_system_prompt = False @@ -476,6 +478,11 @@ def _handle_run_agent(step: Step, ctx: StepContext) -> StepOutcome: ) ) ctx.agent_result = result + from microvm_lifecycle import get_context + + lifecycle = get_context(ctx.config.task_id) + if lifecycle and lifecycle.diagnostic_snapshot()["phase"] in {"closed", "failed"}: + raise RuntimeError("Worker execution is closed; remaining workflow steps must not run") # The agent loop "failing" is not a step failure here: pipeline's # _resolve_overall_task_status owns success inference. The step succeeds if # the SDK ran; downstream steps and the terminal-outcome check decide done. diff --git a/agent/tests/test_approval_requests.py b/agent/tests/test_approval_requests.py new file mode 100644 index 000000000..1b649ddcf --- /dev/null +++ b/agent/tests/test_approval_requests.py @@ -0,0 +1,129 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +import json +from types import SimpleNamespace +from unittest.mock import MagicMock + +import pytest +from botocore.credentials import Credentials +from botocore.exceptions import ClientError + +import approval_requests as broker +import task_state + + +def pending_request() -> task_state.ApprovalRow: + return { + "task_id": "task", + "request_id": "request", + "tool_name": "Bash", + "tool_input_preview": '{"command":"git push"}', + "tool_input_sha256": "a" * 64, + "reason": "Protected operation", + "severity": "high", + "matching_rule_ids": ["protected"], + "status": "PENDING", + "created_at": "2026-09-18T00:00:00Z", + "timeout_s": 0, + "user_id": "owner", + "repo": "owner/repo", + } + + +@pytest.fixture +def transport(monkeypatch): + monkeypatch.setenv(broker.API_ENV, "https://fixture.execute-api.us-east-1.amazonaws.com/v1/") + session = MagicMock() + session.get_credentials.return_value = Credentials( + "testing", "testing-secret", "testing-session" + ) + monkeypatch.setattr(broker, "get_session", lambda: session) + post = MagicMock( + return_value=SimpleNamespace(status_code=200, json=lambda: {"data": {"ok": True}}) + ) + monkeypatch.setattr(broker.requests, "post", post) + return post + + +def test_scoped_signature_binds_the_task_path_and_does_not_redirect(transport): + broker.record_request("create", "task", "request", approval={"status": "PENDING"}) + args, kwargs = transport.call_args + assert args == ("https://fixture.execute-api.us-east-1.amazonaws.com/v1/tasks/task",) + assert "execute-api/aws4_request" in kwargs["headers"]["Authorization"] + assert kwargs["headers"]["X-Amz-Security-Token"] == "testing-session" + assert "md/uksb-wt64nei4u6#agent" in kwargs["headers"]["User-Agent"] + assert kwargs["allow_redirects"] is False + assert json.loads(kwargs["data"])["task_id"] == "task" + + +@pytest.mark.parametrize( + "url", + [ + "http://fixture.execute-api.us-east-1.amazonaws.com/v1", + "https://attacker.example/v1", + "https://fixture.execute-api.us-east-1.amazonaws.com/v1?redirect=elsewhere", + ], +) +def test_rejects_untrusted_endpoints_before_signing(transport, monkeypatch, url): + monkeypatch.setenv(broker.API_ENV, url) + with pytest.raises(ValueError): + broker.record_request("create", "task", "request") + transport.assert_not_called() + + +def test_new_writers_use_service_instead_of_direct_dynamodb(transport, monkeypatch): + direct = MagicMock() + monkeypatch.setattr(task_state, "_get_ddb_client", direct) + task_state.transact_write_approval_request("task", "request", pending_request()) + assert task_state.best_effort_update_approval_status("task", "request", "TIMED_OUT") + assert transport.call_count == 2 + direct.assert_not_called() + + +@pytest.mark.parametrize("operation", ["create", "timeout"]) +def test_cloud_worker_missing_endpoint_reports_configuration_error(monkeypatch, operation): + monkeypatch.delenv(broker.API_ENV, raising=False) + monkeypatch.setenv("AGENT_SESSION_ROLE_ARN", "arn:aws:iam::123456789012:role/session") + direct = MagicMock() + monkeypatch.setattr(task_state, "_get_ddb_client", direct) + with pytest.raises(RuntimeError, match=r"APPROVAL_REQUESTS_API_URL.*matching CDK"): + if operation == "create": + task_state.transact_write_approval_request("task", "request", pending_request()) + else: + task_state.best_effort_update_approval_status("task", "request", "TIMED_OUT") + direct.assert_not_called() + + +@pytest.mark.parametrize("status", ["APPROVED", "DENIED", "PENDING", "CANCELLED"]) +def test_worker_api_cannot_record_a_human_decision(transport, status): + with pytest.raises(ValueError, match="non-human timeouts"): + task_state.best_effort_update_approval_status("task", "request", status) + transport.assert_not_called() + + +def test_uncertain_service_write_does_not_fall_back_to_direct_dynamodb(transport, monkeypatch): + direct = MagicMock() + monkeypatch.setattr(task_state, "_get_ddb_client", direct) + transport.side_effect = TimeoutError("reply lost") + with pytest.raises(TimeoutError): + task_state.transact_write_approval_request("task", "request", pending_request()) + direct.assert_not_called() + + +def test_timeout_race_preserves_human_winner_but_lease_loss_is_not_benign(transport): + reasons = [{"Code": "ConditionalCheckFailed"}, {"Code": "None"}, {"Code": "None"}] + transport.return_value = SimpleNamespace( + status_code=409, + json=lambda: { + "error": { + "code": "TransactionCanceledException", + "details": {"cancellation_reasons": reasons}, + }, + }, + ) + assert not task_state.best_effort_update_approval_status("task", "request", "TIMED_OUT") + reasons[-1] = {"Code": "ConditionalCheckFailed"} + with pytest.raises(ClientError, match="pre-upgrade tasks must be drained") as failure: + task_state.best_effort_update_approval_status("task", "request", "TIMED_OUT") + assert failure.value.response["Error"]["Code"] == "TransactionCanceledException" diff --git a/agent/tests/test_approval_retention.py b/agent/tests/test_approval_retention.py new file mode 100644 index 000000000..ba872ab6f --- /dev/null +++ b/agent/tests/test_approval_retention.py @@ -0,0 +1,37 @@ +"""Approval deadlines are explicit; compute lifetime never creates one.""" + +import math + +from hooks import _ApprovalDeadline, _compute_effective_timeout +from policy import PolicyEngine + + +def test_default_builtin_gate_has_no_automatic_deadline(): + engine = PolicyEngine(task_type="new_task", repo="owner/repo") + decision = engine.evaluate_tool_use("Bash", {"command": "git push --force origin feature"}) + assert decision.outcome == "require_approval" + assert decision.timeout_s == 0 + + +def test_zero_deadline_survives_a_long_pause(monkeypatch): + deadline = _ApprovalDeadline.from_recorded("2026-01-01T00:00:00Z", 0) + monkeypatch.setattr("hooks.time.time", lambda: 9_999_999_999) + monkeypatch.setattr("hooks.time.monotonic", lambda: 9_999_999_999) + assert math.isinf(deadline.remaining_s()) + assert _compute_effective_timeout( + decision_timeout_s=0, task_default_timeout_s=0, remaining_lifetime_s=10 + ) == (0, None, 0) + + +def test_explicit_rule_deadline_survives_a_no_deadline_task_default(): + engine = PolicyEngine( + task_type="new_task", + repo="owner/repo", + blueprint_soft_policies=( + '@tier("soft") @rule_id("explicit") @approval_timeout_s("120") ' + 'forbid (principal, action == Agent::Action::"execute_bash", resource) ' + 'when { context.command like "*explicit-tool*" };' + ), + ) + decision = engine.evaluate_tool_use("Bash", {"command": "explicit-tool"}) + assert decision.timeout_s == 120 diff --git a/agent/tests/test_bedrock_creds_helper.py b/agent/tests/test_bedrock_creds_helper.py index 23f4c2a7d..e2f242935 100644 --- a/agent/tests/test_bedrock_creds_helper.py +++ b/agent/tests/test_bedrock_creds_helper.py @@ -32,6 +32,19 @@ def attr_file(tmp_path, monkeypatch): return path +def test_microvm_export_is_empty_without_reading_files_or_resolving_ambient(monkeypatch, capsys): + monkeypatch.setenv("ABCA_MICROVM_CREDENTIAL_BROKER", "1") + with ( + patch("builtins.open", side_effect=AssertionError("must not read attribution")), + patch.object( + helper, "_ambient_credentials", side_effect=AssertionError("must not resolve") + ), + patch("boto3.client", side_effect=AssertionError("must not call STS")), + ): + assert helper.main() == 0 + assert json.loads(capsys.readouterr().out) == {"Credentials": {}} + + def test_write_attribution_file_is_0600(attr_file): tags = build_session_tags("u1", "owner/repo", "task123") written = helper.write_attribution_file("arn:aws:iam::1:role/SR", tags, attr_file) diff --git a/agent/tests/test_config.py b/agent/tests/test_config.py index 12b42233c..5fd0ec036 100644 --- a/agent/tests/test_config.py +++ b/agent/tests/test_config.py @@ -201,7 +201,7 @@ class TestResolveLinearApiToken: The orchestrator stamps `linear_oauth_secret_arn` into the task's channel_metadata at creation time. resolve_linear_api_token reads the secret JSON via boto3, refreshes it if expiring, and caches the - access_token in `LINEAR_API_TOKEN` for the Linear MCP placeholder. + access_token in `LINEAR_API_TOKEN` for direct Linear API calls. """ def test_returns_cached_value_without_calling_secrets_manager(self, monkeypatch): @@ -228,6 +228,44 @@ def test_returns_empty_when_region_missing(self, monkeypatch): assert resolve_linear_api_token({"linear_oauth_secret_arn": "arn:test"}) == "" mock_boto.assert_not_called() + @pytest.mark.parametrize( + "payload", + [ + {"workspace_id": "ws", "provider_name": "bgagent-linear-oauth-acme"}, + {"refresh_token": None, "client_id": "cid", "client_secret": "secret"}, + {"refresh_token": "rt", "client_id": "", "client_secret": "secret"}, + {"refresh_token": "rt", "client_id": "cid"}, + ], + ) + def test_incomplete_fallback_skips_refresh_without_crashing(self, monkeypatch, payload): + monkeypatch.delenv("LINEAR_API_TOKEN", raising=False) + monkeypatch.delenv("LINEAR_VAULT_ENABLED", raising=False) + monkeypatch.setenv("AWS_REGION", "us-east-1") + mock_sm = MagicMock() + mock_sm.get_secret_value.return_value = { + "SecretString": __import__("json").dumps(payload), + } + with ( + patch("boto3.client", return_value=mock_sm), + patch("urllib.request.urlopen") as post, + patch("config.log") as log, + ): + assert resolve_linear_api_token({"linear_oauth_secret_arn": "arn:test"}) == "" + post.assert_not_called() + assert any( + "linear_oauth_refresh_unavailable" in call.args[1] for call in log.call_args_list + ) + + @pytest.mark.parametrize("payload", ["null", "[]", '"text"']) + def test_non_object_fallback_returns_empty(self, monkeypatch, payload): + monkeypatch.delenv("LINEAR_API_TOKEN", raising=False) + monkeypatch.delenv("LINEAR_VAULT_ENABLED", raising=False) + monkeypatch.setenv("AWS_REGION", "us-east-1") + mock_sm = MagicMock() + mock_sm.get_secret_value.return_value = {"SecretString": payload} + with patch("boto3.client", return_value=mock_sm): + assert resolve_linear_api_token({"linear_oauth_secret_arn": "arn:test"}) == "" + def test_resolves_from_secrets_manager_and_caches_in_env(self, monkeypatch): """Happy path: channel_metadata carries the ARN, secret has access_token + future expiry.""" from datetime import datetime, timedelta diff --git a/agent/tests/test_continuation_runtime.py b/agent/tests/test_continuation_runtime.py new file mode 100644 index 000000000..a3f059516 --- /dev/null +++ b/agent/tests/test_continuation_runtime.py @@ -0,0 +1,476 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""A complete checkpoint must recover after losing both process and workspace.""" + +from __future__ import annotations + +import asyncio +import shutil +import subprocess +from dataclasses import asdict +from typing import Any +from unittest.mock import MagicMock + +import pytest + +from continuation_runtime import ContinuationRuntime, restore_for_task +from continuation_storage import ContinuationContext, S3ContinuationStorage +from microvm_lifecycle import ApprovalRecord, LifecycleUnavailable, MicrovmLifecycle +from models import RepoSetup, TaskConfig +from tests.test_continuation_storage import VersionedS3 + +SESSION = "11111111-1111-4111-8111-aaaaaaaaaaaa" + + +class Deadline: + def remaining_s(self) -> float: + return 5000 + + +def git(workspace, *args): + return subprocess.run( + ["git", "-C", str(workspace), *args], + check=True, + capture_output=True, + text=True, + ).stdout.strip() + + +@pytest.fixture +def ready(tmp_path, monkeypatch): + from continuation_usage import UsageSnapshot + + async def usage(_): + return UsageSnapshot(0.02, {"input_tokens": 100, "output_tokens": 10}) + + monkeypatch.setattr("continuation_runtime.read_usage", usage) + workspace = tmp_path / "task" + workspace.mkdir() + git(workspace, "init", "-b", "main") + git(workspace, "config", "user.name", "Test") + git(workspace, "config", "user.email", "test@example.test") + (workspace / "tracked.txt").write_text("committed\n") + git(workspace, "add", ".") + git(workspace, "commit", "-m", "initial") + original_head = git(workspace, "rev-parse", "HEAD") + (workspace / "tracked.txt").write_text("staged edit\n") + git(workspace, "add", ".") + (workspace / "tracked.txt").write_text("unstaged edit\n") + (workspace / "untracked.txt").write_text("keep this work\n") + context = ContinuationContext( + RepoSetup( + repo_dir=str(workspace), + branch="main", + build_before=False, + head_sha_before=original_head, + ), + "Original user prompt", + "Original system prompt", + "coding/new-task-v1", + "1", + ) + client = VersionedS3() + runtime = ContinuationRuntime(context, S3ContinuationStorage("owned-bucket", client=client)) + return workspace, runtime + + +async def park_runtime(runtime, repo="owner/repo"): + lifecycle = MicrovmLifecycle("task", "microvm-one") + await lifecycle.tool_started("toolu_pending") + park = lifecycle.park_approval( + "request", + "toolu_pending", + Deadline(), + record=ApprovalRecord("user", repo, "2026-09-17T00:00:00Z", 5000), + ) + assert park is not None + tool_input = {"file_path": runtime.context.setup.repo_dir + "/untracked.txt"} + entries: Any = [ + { + "type": "assistant", + "uuid": "entry-one", + "sessionId": SESSION, + "message": { + "content": [ + { + "type": "tool_use", + "id": park.tool_use_id, + "name": "Read", + "input": tool_input, + } + ] + }, + } + ] + await runtime.store.append( + {"project_key": runtime.store.project_key, "session_id": SESSION}, entries + ) + return lifecycle, park, tool_input + + +class TestCompleteRecovery: + def test_restore_preserves_git_baseline_conversation_and_pending_action(self, ready): + workspace, runtime = ready + before_status = git(workspace, "status", "--porcelain=v1") + before_index = git(workspace, "diff", "--cached") + + async def save(): + lifecycle, park, tool_input = await park_runtime(runtime) + identity, receipt = await runtime.capture( + lifecycle, park, session_id=SESSION, tool_name="Read", tool_input=tool_input + ) + assert lifecycle.diagnostic_snapshot()["phase"] == "checkpoint-ready" + lifecycle.close() + return identity, receipt + + identity, receipt = asyncio.run(save()) + shutil.rmtree(workspace) + restored = ContinuationRuntime.restore( + runtime.storage, + receipt, + identity, + expected_workspace=workspace, + workflow_id="coding/new-task-v1", + workflow_version="1", + ) + assert restored.context == runtime.context + assert git(workspace, "status", "--porcelain=v1") == before_status + assert git(workspace, "diff", "--cached") == before_index + assert (workspace / "untracked.txt").read_text() == "keep this work\n" + assert restored.restored is not None + assert restored.restored["session_id"] == SESSION + prompt = restored.decision_prompt(decision="DENIED", reason="Use another approach") + assert "DENIED" in prompt and "Use another approach" in prompt + assert str(workspace / "untracked.txt") in prompt + assert "normal permission checks" in prompt + + def test_later_parallel_tool_cannot_modify_a_published_workspace(self, ready): + _, runtime = ready + + async def scenario(): + lifecycle, park, tool_input = await park_runtime(runtime) + await runtime.capture( + lifecycle, park, session_id=SESSION, tool_name="Read", tool_input=tool_input + ) + later = asyncio.create_task(lifecycle.tool_started("toolu_later")) + await asyncio.sleep(0.03) + assert not later.done() + # Reads/feedback remain available while tools are held. + async with lifecycle.approval_poll(): + pass + with lifecycle.activity(): + pass + with pytest.raises(LifecycleUnavailable, match="claim task ownership"): + await lifecycle.leave_approval(park) + await lifecycle.release_continuation(park) + await asyncio.wait_for(later, 1) + assert lifecycle.diagnostic_snapshot()["phase"] == "active" + + asyncio.run(scenario()) + + def test_sleep_and_wake_do_not_release_continuation_tools(self, ready): + _, runtime = ready + + async def scenario(): + lifecycle, park, tool_input = await park_runtime(runtime) + await runtime.capture( + lifecycle, park, session_id=SESSION, tool_name="Read", tool_input=tool_input + ) + await lifecycle.suspend(lambda _: None, budget_s=1) + await lifecycle.resume(lambda _: None, budget_s=1) + assert lifecycle.diagnostic_snapshot()["phase"] == "checkpoint-ready" + await lifecycle.release_continuation(park) + assert lifecycle.diagnostic_snapshot()["phase"] == "active" + + asyncio.run(scenario()) + + def test_failed_capture_keeps_original_approval_usable(self, ready, monkeypatch): + _, runtime = ready + + def fail(*_): + raise RuntimeError("owned storage failure") + + monkeypatch.setattr(runtime, "_save", fail) + + async def scenario(): + lifecycle, park, tool_input = await park_runtime(runtime) + with pytest.raises(RuntimeError, match="storage failure"): + await runtime.capture( + lifecycle, park, session_id=SESSION, tool_name="Read", tool_input=tool_input + ) + assert lifecycle.diagnostic_snapshot()["phase"] == "parked" + await lifecycle.leave_approval(park) + + asyncio.run(scenario()) + + def test_cancellation_never_opens_tools_while_capture_thread_may_still_run(self, ready): + _, runtime = ready + + async def scenario(): + lifecycle, park, _ = await park_runtime(runtime) + entered = asyncio.Event() + + async def capture(): + async with lifecycle.continuation_checkpoint(park): + entered.set() + await asyncio.Event().wait() + + task = asyncio.create_task(capture()) + await entered.wait() + task.cancel() + with pytest.raises(asyncio.CancelledError): + await task + assert lifecycle.diagnostic_snapshot()["phase"] == "failed" + with pytest.raises(LifecycleUnavailable): + await lifecycle.tool_started("toolu_later") + + asyncio.run(scenario()) + + @pytest.mark.parametrize("mismatch", ["workspace", "workflow", "version"]) + def test_wrong_restore_context_is_rejected_before_creating_files( + self, ready, mismatch, tmp_path + ): + workspace, runtime = ready + + async def save(): + lifecycle, park, tool_input = await park_runtime(runtime) + return await runtime.capture( + lifecycle, park, session_id=SESSION, tool_name="Read", tool_input=tool_input + ) + + identity, receipt = asyncio.run(save()) + shutil.rmtree(workspace) + with pytest.raises(RuntimeError, match="workspace or workflow"): + ContinuationRuntime.restore( + runtime.storage, + receipt, + identity, + expected_workspace=tmp_path / "another" if mismatch == "workspace" else workspace, + workflow_id="another" if mismatch == "workflow" else "coding/new-task-v1", + workflow_version="2" if mismatch == "version" else "1", + ) + assert not workspace.exists() + assert not (tmp_path / "another").exists() + + +@pytest.fixture +def replacement(ready, monkeypatch): + from continuation_session import decode_checkpoint + + workspace, runtime = ready + + async def save(): + lifecycle, park, tool_input = await park_runtime(runtime) + result = await runtime.capture( + lifecycle, park, session_id=SESSION, tool_name="Read", tool_input=tool_input + ) + lifecycle.close() + return result + + identity, receipt = asyncio.run(save()) + manifest = runtime.storage.load_manifest(receipt, identity) + checkpoint = runtime.storage.conversations.load(manifest.conversation, identity) + action = decode_checkpoint(checkpoint, identity)["action"] + task = { + "task_id": "task", + "user_id": "user", + "repo": "owner/repo", + "status": "AWAITING_APPROVAL", + "session_id": "microvm-new", + "awaiting_approval_request_id": "request", + "approval_gate_count": 4, + "continuation": { + "version": 1, + "state": "RESTORING", + "worker_id": "microvm-new", + "attempt_id": "attempt-new", + "identity": asdict(identity), + "manifest": asdict(receipt), + }, + } + approval = { + "user_id": "user", + "repo": "owner/repo", + "status": "APPROVED", + "tool_input_sha256": action["tool_input_sha256"], + } + lifecycle = MicrovmLifecycle("task", "microvm-new") + monkeypatch.setenv("CONTINUATION_BUCKET_NAME", "owned-bucket") + monkeypatch.setattr("config.AGENT_WORKSPACE", str(workspace.parent)) + monkeypatch.setattr("microvm_lifecycle.get_context", lambda _: lifecycle) + monkeypatch.setattr("continuation_runtime.S3ContinuationStorage", lambda _: runtime.storage) + get_task = MagicMock(return_value=task) + consume = MagicMock() + monkeypatch.setattr("task_state.get_task", get_task) + monkeypatch.setattr("task_state.get_approval_row", lambda *_args, **_kw: approval) + monkeypatch.setattr("task_state.consume_restored_continuation", consume) + config = TaskConfig( + task_id="task", + user_id="user", + repo_url="owner/repo", + github_token="test", + aws_region="us-west-2", + resolved_workflow={"id": "coding/new-task-v1", "version": "1"}, + ) + shutil.rmtree(workspace) + yield config, task, approval, consume, get_task, workspace + lifecycle.close() + + +@pytest.mark.parametrize("decision", ["APPROVED", "DENIED", "TIMED_OUT"]) +def test_production_restore_loads_saved_files_then_claims_recorded_decision(replacement, decision): + config, task, approval, consume, get_task, workspace = replacement + approval["status"] = decision + approval["deny_reason"] = "Choose another approach" + runtime = restore_for_task(config) + assert runtime is not None and runtime.restored is not None + assert (workspace / "untracked.txt").read_text() == "keep this work\n" + assert runtime.human_decision == approval + assert runtime.resume_prompt is not None and decision in runtime.resume_prompt + assert runtime.prior_cost_usd == pytest.approx(0.02) + assert config.initial_approval_gate_count == 4 + get_task.assert_called_once_with("task", consistent_read=True) + consume.assert_called_once_with("task", "microvm-new", task["continuation"]) + + +@pytest.mark.parametrize( + "field,value", + [ + ("status", "PENDING"), + ("status", "CANCELLED"), + ("user_id", "other"), + ("repo", "other/repo"), + ("tool_input_sha256", "different"), + ], +) +def test_production_restore_never_claims_an_unavailable_or_mismatched_decision( + replacement, field, value +): + from continuation_session import ContinuationCheckpointError + + config, _, approval, consume, _, _ = replacement + approval[field] = value + with pytest.raises(ContinuationCheckpointError, match="human decision"): + restore_for_task(config) + consume.assert_not_called() + + +def test_production_restore_waits_for_start_registration_without_fresh_clone( + replacement, monkeypatch +): + from copy import deepcopy + + config, task, _, consume, get_task, _ = replacement + starting = deepcopy(task) + starting["continuation"]["state"] = "STARTING" + get_task.side_effect = [starting, task] + monkeypatch.setattr("continuation_runtime.time.sleep", lambda _: None) + assert restore_for_task(config) is not None + assert get_task.call_count == 2 + consume.assert_called_once() + + +def test_production_restore_rejects_a_different_physical_worker_before_writing_files(replacement): + from continuation_session import ContinuationCheckpointError + + config, task, _, consume, _, workspace = replacement + task["continuation"]["worker_id"] = "other-worker" + with pytest.raises(ContinuationCheckpointError, match="does not own"): + restore_for_task(config) + assert not workspace.exists() + consume.assert_not_called() + + +def test_repository_free_microvm_restores_its_private_scratch_files_without_a_remote( + ready, tmp_path, monkeypatch +): + from continuation_runtime import prepare_repoless_runtime + from continuation_session import decode_checkpoint + + _, reference = ready + base = tmp_path / "scratch-root" + monkeypatch.setattr("config.AGENT_WORKSPACE", str(base)) + monkeypatch.setenv("CONTINUATION_BUCKET_NAME", "owned-bucket") + monkeypatch.setattr("continuation_runtime.S3ContinuationStorage", lambda _: reference.storage) + lifecycle = MicrovmLifecycle("task", "microvm-one") + monkeypatch.setattr("microvm_lifecycle.get_context", lambda _: lifecycle) + task = {"task_id": "task", "user_id": "user", "status": "RUNNING"} + monkeypatch.setattr("task_state.get_task", lambda *_args, **_kw: task) + config = TaskConfig( + task_id="task", + user_id="user", + repo_url="", + github_token="", + requires_repo=False, + aws_region="us-west-2", + resolved_workflow={"id": "default/agent-v1", "version": "1"}, + ) + runtime = prepare_repoless_runtime( + config, + user_prompt="Keep these scratch files", + system_prompt="Saved instructions", + workflow_id="default/agent-v1", + workflow_version="1", + ) + assert runtime is not None + workspace = base / "task" + assert runtime.context.setup.repo_dir == str(workspace) + assert git(workspace, "remote") == "" + (workspace / "untracked.txt").write_text("private draft\n") + + async def capture(): + active, park, tool_input = await park_runtime(runtime, repo="") + result = await runtime.capture( + active, park, session_id=SESSION, tool_name="Read", tool_input=tool_input + ) + active.close() + return result + + identity, receipt = asyncio.run(capture()) + manifest = runtime.storage.load_manifest(receipt, identity) + action = decode_checkpoint( + runtime.storage.conversations.load(manifest.conversation, identity), identity + )["action"] + lifecycle.close() + lifecycle = MicrovmLifecycle("task", "microvm-new") + task.update( + status="AWAITING_APPROVAL", + session_id="microvm-new", + awaiting_approval_request_id="request", + continuation={ + "version": 1, + "state": "RESTORING", + "worker_id": "microvm-new", + "identity": asdict(identity), + "manifest": asdict(receipt), + }, + ) + monkeypatch.setattr( + "task_state.get_approval_row", + lambda *_args, **_kw: { + "user_id": "user", + "repo": "", + "status": "APPROVED", + "tool_input_sha256": action["tool_input_sha256"], + }, + ) + consume = MagicMock() + monkeypatch.setattr("task_state.consume_restored_continuation", consume) + shutil.rmtree(workspace) + try: + restored = prepare_repoless_runtime( + config, + user_prompt="replacement input", + system_prompt="replacement input", + workflow_id="default/agent-v1", + workflow_version="1", + ) + assert restored is not None and restored.restored is not None + assert restored.context.system_prompt == "Saved instructions" + assert (workspace / "untracked.txt").read_text() == "private draft\n" + assert git(workspace, "remote") == "" + assert "credential" not in (workspace / ".git/config").read_text() + consume.assert_called_once() + finally: + lifecycle.close() diff --git a/agent/tests/test_continuation_sdk_probe.py b/agent/tests/test_continuation_sdk_probe.py new file mode 100644 index 000000000..a09b53f1b --- /dev/null +++ b/agent/tests/test_continuation_sdk_probe.py @@ -0,0 +1,396 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Opt-in real SDK/CLI continuation test with a synthetic loopback model. + +Run ABCA_TEST_SDK_CONTINUATION=1 uv run pytest tests/test_continuation_sdk_probe.py +--no-cov. No model service or real AWS credentials are used. +""" + +from __future__ import annotations + +import argparse +import asyncio +import importlib.metadata +import json +import os +import shutil +import signal +import subprocess +import sys +import threading +import time +from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer +from pathlib import Path + +import pytest + +AGENT_ROOT = Path(__file__).resolve().parents[1] +sys.path.insert(0, str(AGENT_ROOT)) +sys.path.insert(0, str(AGENT_ROOT / "src")) + +from continuation_session import CheckpointIdentity, CheckpointSessionStore, decode_checkpoint +from continuation_usage import read_usage +from continuation_workspace import capture_workspace, restore_workspace +from scripts.verify_microvm_credentials import MODEL, event_frame, model_response + +IDENTITY = CheckpointIdentity("sdk-probe", "attempt-1", "request-1", "owner", "probe/owned") + + +def _tool_response(target: Path, tool_id: str) -> bytes: + events = [ + { + "type": "message_start", + "message": { + "id": "msg_" + tool_id, + "type": "message", + "role": "assistant", + "model": MODEL, + "content": [], + "stop_reason": None, + "stop_sequence": None, + "usage": {"input_tokens": 1, "output_tokens": 0}, + }, + }, + { + "type": "content_block_start", + "index": 0, + "content_block": {"type": "tool_use", "id": tool_id, "name": "Read", "input": {}}, + }, + { + "type": "content_block_delta", + "index": 0, + "delta": { + "type": "input_json_delta", + "partial_json": json.dumps({"file_path": str(target)}), + }, + }, + {"type": "content_block_stop", "index": 0}, + { + "type": "message_delta", + "delta": {"stop_reason": "tool_use", "stop_sequence": None}, + "usage": {"output_tokens": 1}, + }, + {"type": "message_stop"}, + ] + return b"".join(event_frame(event) for event in events) + + +async def _worker(directory: Path, endpoint: str, phase: str, decision: str) -> None: + from claude_agent_sdk import ( + ClaudeAgentOptions, + ClaudeSDKClient, + ResultMessage, + project_key_for_directory, + ) + from claude_agent_sdk.types import HookMatcher + + workspace = directory / "workspace" + target = workspace / "owned.txt" + original = phase == "original" + store = ( + CheckpointSessionStore(project_key_for_directory(str(workspace))) + if original + else CheckpointSessionStore.restore((directory / "checkpoint.json").read_bytes(), IDENTITY) + ) + saved = ( + None + if original + else decode_checkpoint((directory / "checkpoint.json").read_bytes(), IDENTITY) + ) + audit: list[dict] = [] + + def record(kind: str, **data) -> None: + audit.append({"kind": kind, **data}) + (directory / f"{phase}-audit.json").write_text(json.dumps(audit)) + + async def pre(data, tool_id, context): + assert data["tool_name"] == "Read" + assert data["tool_input"] == {"file_path": str(target)} + record("pre", session_id=data["session_id"], tool_id=tool_id) + usage = await read_usage(client) + record("usage", cost_usd=usage.cost_usd, tokens=usage.tokens) + if original: + body = await store.checkpoint_pending( + IDENTITY, + session_id=data["session_id"], + tool_use_id=tool_id, + tool_name=data["tool_name"], + tool_input=data["tool_input"], + ) + pending = directory / "checkpoint.tmp" + pending.write_bytes(body) + pending.chmod(0o600) + pending.replace(directory / "checkpoint.json") + await asyncio.sleep(100) + raise RuntimeError("Original process was not stopped") + return { + "hookSpecificOutput": { + "hookEventName": "PreToolUse", + "permissionDecision": "allow" if decision == "approve" else "deny", + "permissionDecisionReason": "Owned continuation diagnostic decision", + } + } + + async def post(data, tool_id, context): + record("post", tool_id=tool_id) + return {} + + env = { + "CLAUDE_CONFIG_DIR": str(directory / f"{phase}-config"), + "CLAUDE_CODE_USE_BEDROCK": "1", + "CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC": "1", + "CLAUDE_CODE_MAX_RETRIES": "0", + "DISABLE_TELEMETRY": "1", + "DISABLE_ERROR_REPORTING": "1", + "DISABLE_AUTOUPDATER": "1", + "ANTHROPIC_BEDROCK_BASE_URL": endpoint, + "AWS_ENDPOINT_URL": endpoint, + "AWS_REGION": "us-west-2", + "AWS_DEFAULT_REGION": "us-west-2", + "AWS_CONFIG_FILE": str(directory / "aws-config"), + "AWS_SHARED_CREDENTIALS_FILE": str(directory / "aws-credentials"), + "AWS_ACCESS_KEY_ID": "SYNTHETIC_NOT_VALID", + "AWS_SECRET_ACCESS_KEY": "synthetic-not-valid-in-aws", + "AWS_SESSION_TOKEN": "synthetic-not-valid-in-aws", + "AWS_EC2_METADATA_DISABLED": "true", + } + options = ClaudeAgentOptions( + model=MODEL, + max_turns=3, + max_budget_usd=0.00006 + - ( + next( + item["cost_usd"] + for item in json.loads((directory / "original-audit.json").read_text()) + if item["kind"] == "usage" + ) + if not original + else 0 + ), + cwd=str(workspace), + tools=["Read"], + permission_mode="bypassPermissions", + setting_sources=[], + settings=str(directory / "settings.json"), + env=env, + resume=saved["session_id"] if saved else None, + session_store=store, + session_store_flush="eager", + stderr=lambda line: record("stderr", text=line), + hooks={ + "PreToolUse": [HookMatcher(hooks=[pre], timeout=100)], + "PostToolUse": [HookMatcher(hooks=[post])], + }, + ) + async with ClaudeSDKClient(options=options) as client: + prompt = ( + "Continue the saved task. The human decision was " + + decision + + ". The saved pending action was " + + json.dumps(saved["action"]) + if saved + else "Read the owned marker once, then stop." + ) + await client.query(prompt) + async for message in client.receive_response(): + if isinstance(message, ResultMessage): + record( + "result", + session_id=message.session_id, + is_error=message.is_error, + cost_usd=message.total_cost_usd, + ) + + +@pytest.mark.skipif( + os.environ.get("ABCA_TEST_SDK_CONTINUATION") != "1" and os.environ.get("CI") != "true", + reason="Pinned SDK/CLI subprocess diagnostic runs in CI or by local opt-in", +) +@pytest.mark.parametrize("decision", ["approve", "deny"]) +def test_sdk_resumes_from_checkpoint_without_original_config(tmp_path, decision): + import claude_agent_sdk + + assert importlib.metadata.version("claude-agent-sdk") == "0.2.110" + cli = Path(claude_agent_sdk.__file__).parent / "_bundled/claude" + version = subprocess.check_output([str(cli), "--version"], text=True, timeout=10).strip() + assert version == "2.1.191 (Claude Code)" + workspace = tmp_path / "workspace" + workspace.mkdir() + target = workspace / "owned.txt" + target.write_text("OWNED_CONTINUATION_MARKER\n") + git_env = {key: value for key, value in os.environ.items() if not key.startswith("GIT_")} + git_env.update(GIT_CONFIG_GLOBAL=os.devnull, GIT_CONFIG_NOSYSTEM="1") + for command in ( + ["init", "--quiet", "--template=", "--initial-branch=work/probe"], + ["add", "owned.txt"], + [ + "-c", + "user.name=Fixture", + "-c", + "user.email=fixture@example.invalid", + "commit", + "--quiet", + "-m", + "owned", + ], + ): + subprocess.run( + ["git", "-c", f"core.hooksPath={os.devnull}", *command], + cwd=workspace, + env=git_env, + check=True, + capture_output=True, + timeout=10, + ) + (workspace / "untracked.txt").write_text("UNTRACKED_CONTINUATION_MARKER\n") + for phase in ("original", "restored"): + (tmp_path / f"{phase}-config").mkdir() + auth_sentinel = "SYNTHETIC_AUTH_FILE_MUST_NOT_BE_COPIED" + (tmp_path / "original-config/.credentials.json").write_text( + json.dumps({"sentinel": auth_sentinel}) + ) + (tmp_path / "settings.json").write_text("{}") + (tmp_path / "aws-config").write_text("[default]\nregion = us-west-2\n") + (tmp_path / "aws-credentials").write_text("") + phase = ["original"] + requests = [] + + class Handler(BaseHTTPRequestHandler): + def log_message(self, format, *args): + del format, args + + def do_POST(self): + body = json.loads(self.rfile.read(int(self.headers["Content-Length"]))) + requests.append({"phase": phase[0], "body": body}) + results = [ + item + for message in body.get("messages", []) + for item in ( + message.get("content", []) if isinstance(message.get("content"), list) else [] + ) + if isinstance(item, dict) and item.get("type") == "tool_result" + ] + response = model_response() if results else _tool_response(target, "toolu_" + phase[0]) + self.send_response(200) + self.send_header("Content-Type", "application/vnd.amazon.eventstream") + self.send_header("Content-Length", str(len(response))) + self.end_headers() + self.wfile.write(response) + + server = ThreadingHTTPServer(("127.0.0.1", 0), Handler) + thread = threading.Thread(target=server.serve_forever, daemon=True) + thread.start() + endpoint = f"http://127.0.0.1:{server.server_port}" + children = [] + logs = [] + + def start(name): + log = (tmp_path / f"{name}-process.log").open("w") + logs.append(log) + child = subprocess.Popen( + [ + sys.executable, + str(__file__), + "--worker", + name, + "--directory", + str(tmp_path), + "--endpoint", + endpoint, + "--decision", + decision, + ], + stdout=log, + stderr=log, + start_new_session=True, + ) + children.append(child) + return child + + try: + first = start("original") + limit = time.monotonic() + 30 + while not (tmp_path / "checkpoint.json").exists(): + assert first.poll() is None, (tmp_path / "original-process.log").read_text() + assert time.monotonic() < limit, (tmp_path / "original-audit.json").read_text() + time.sleep(0.05) + body = (tmp_path / "checkpoint.json").read_bytes() + saved = decode_checkpoint(body, IDENTITY) + assert auth_sentinel.encode() not in body + workspace_archive = tmp_path / "workspace.tar" + workspace_receipt = capture_workspace(workspace, workspace_archive, IDENTITY) + os.killpg(first.pid, signal.SIGKILL) + assert first.wait(timeout=10) == -signal.SIGKILL + original_audit = json.loads((tmp_path / "original-audit.json").read_text()) + assert not any(event["kind"] == "post" for event in original_audit) + shutil.rmtree(tmp_path / "original-config") + shutil.rmtree(workspace) + restore_workspace( + workspace_archive, workspace, IDENTITY, expected_sha256=workspace_receipt.sha256 + ) + phase[0] = "restored" + second = start("restored") + assert second.wait(timeout=35) == 0, (tmp_path / "restored-process.log").read_text() + audit = json.loads((tmp_path / "restored-audit.json").read_text()) + assert [event["tool_id"] for event in audit if event["kind"] == "pre"] == ["toolu_restored"] + posts = [event["tool_id"] for event in audit if event["kind"] == "post"] + assert posts == (["toolu_restored"] if decision == "approve" else []) + result = next(event for event in audit if event["kind"] == "result") + assert not result["is_error"] and result["session_id"] == saved["session_id"] + original_usage = next(event for event in original_audit if event["kind"] == "usage") + restored_usage = next(event for event in audit if event["kind"] == "usage") + assert original_usage["cost_usd"] == pytest.approx(0.000018) + assert restored_usage["cost_usd"] == pytest.approx(0.000018) + assert result["cost_usd"] + original_usage["cost_usd"] == pytest.approx(0.000054) + assert target.read_text() == "OWNED_CONTINUATION_MARKER\n" + assert (workspace / "untracked.txt").read_text() == "UNTRACKED_CONTINUATION_MARKER\n" + restored_requests = [request for request in requests if request["phase"] == "restored"] + assert any( + saved["action"]["tool_use_id"] in json.dumps(r["body"]) for r in restored_requests + ) + proof = { + "sdk": "0.2.110", + "cli": version, + "decision": decision, + "session_id": saved["session_id"], + "original_config_deleted": True, + "pending_action_acknowledged": True, + "auth_file_excluded": True, + "workspace_archive_sha256": workspace_receipt.sha256, + "workspace_head": workspace_receipt.head, + "workspace_restored_from_archive": True, + "untracked_file_preserved": True, + "restored_posts": posts, + "only_synthetic_loopback": True, + "prior_cost_usd": original_usage["cost_usd"], + "restored_budget_usd": 0.00006 - original_usage["cost_usd"], + "total_cost_usd": result["cost_usd"] + original_usage["cost_usd"], + } + (tmp_path / "verification.json").write_text(json.dumps(proof, indent=2)) + finally: + for child in children: + if child.poll() is None: + os.killpg(child.pid, signal.SIGKILL) + child.wait(timeout=10) + server.shutdown() + server.server_close() + thread.join(timeout=2) + for log in logs: + log.close() + (tmp_path / "model-requests.json").write_text(json.dumps(requests, indent=2)) + + +if __name__ == "__main__": + parser = argparse.ArgumentParser() + parser.add_argument("--worker", required=True) + parser.add_argument("--directory", type=Path, required=True) + parser.add_argument("--endpoint", required=True) + parser.add_argument("--decision", choices=["approve", "deny"], required=True) + args = parser.parse_args() + for key in tuple(os.environ): + if key.startswith(("AWS_", "ANTHROPIC_", "CLAUDE_", "OTEL_", "BEDROCK_")): + del os.environ[key] + asyncio.run( + asyncio.wait_for(_worker(args.directory, args.endpoint, args.worker, args.decision), 110) + ) diff --git a/agent/tests/test_continuation_session.py b/agent/tests/test_continuation_session.py new file mode 100644 index 000000000..2350f8281 --- /dev/null +++ b/agent/tests/test_continuation_session.py @@ -0,0 +1,413 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Conversation recovery must acknowledge exact data, not merely a write attempt.""" + +from __future__ import annotations + +import asyncio +import base64 +import hashlib +import io +import json +from dataclasses import replace +from typing import TYPE_CHECKING, Any +from unittest.mock import Mock + +import pytest +from botocore.exceptions import ClientError, EndpointConnectionError + +import continuation_session as checkpoint +from hooks import _sha256_tool_input_for_row + +if TYPE_CHECKING: + from claude_agent_sdk import SessionKey + +SESSION_ID = "11111111-1111-4111-8111-aaaaaaaaaaaa" +PROJECT = "-workspace-task" +KEY: SessionKey = {"project_key": PROJECT, "session_id": SESSION_ID} +IDENTITY = checkpoint.CheckpointIdentity("task", "attempt", "request", "user", "owner/repo") +TOOL_INPUT = {"file_path": "/workspace/task/雪.txt", "offset": 1} + + +def assistant(**overrides) -> Any: + entry = { + "type": "assistant", + "uuid": "entry-1", + "sessionId": SESSION_ID, + "message": { + "role": "assistant", + "content": [ + {"type": "tool_use", "id": "toolu_owned", "name": "Read", "input": TOOL_INPUT} + ], + }, + } + entry.update(overrides) + return entry + + +async def capture(store, **kwargs): + return await store.checkpoint_pending( + IDENTITY, + session_id=SESSION_ID, + tool_use_id="toolu_owned", + tool_name="Read", + tool_input=TOOL_INPUT, + timeout_s=0.1, + **kwargs, + ) + + +@pytest.fixture +def body(): + async def prepare(): + store = checkpoint.CheckpointSessionStore(PROJECT) + await store.append(KEY, [assistant()]) + return await capture(store) + + return asyncio.run(prepare()) + + +class TestSessionMirror: + def test_checkpoint_waits_for_exact_action_and_preserves_full_input(self): + async def scenario(): + store = checkpoint.CheckpointSessionStore(PROJECT) + pending = asyncio.create_task(capture(store)) + await asyncio.sleep(0) + assert not pending.done() + await store.append(KEY, [assistant()]) + result = checkpoint.decode_checkpoint(await pending, IDENTITY) + assert result["action"]["tool_input"] == TOOL_INPUT + assert result["action"]["tool_input_sha256"] == _sha256_tool_input_for_row(TOOL_INPUT) + assert result["session_id"] == SESSION_ID + + asyncio.run(scenario()) + + def test_uuid_upserts_keep_order_and_opaque_fields_without_deduplicating_markers(self): + async def scenario(): + store = checkpoint.CheckpointSessionStore(PROJECT) + user: Any = {"type": "user", "uuid": "user-1", "opaque": {"future": [1, 2]}} + marker: Any = {"type": "mode", "data": "plan"} + await store.append(KEY, [user, assistant(), marker]) + replacement = assistant(opaque={"changed": True}) + await store.append(KEY, [replacement, marker]) + expected = [user, replacement, marker, marker] + assert await store.load(KEY) == expected + replacement["opaque"]["changed"] = False + loaded: Any = await store.load(KEY) + loaded[0]["opaque"]["future"].append(3) + unchanged: Any = await store.load(KEY) + assert unchanged[1]["opaque"] == {"changed": True} + assert unchanged[0]["opaque"]["future"] == [1, 2] + + asyncio.run(scenario()) + + def test_restore_round_trip_through_public_store_contract(self, body): + async def scenario(): + restored = checkpoint.CheckpointSessionStore.restore(body, IDENTITY) + assert await restored.load(KEY) == [assistant()] + reply: Any = {"type": "user", "uuid": "reply", "decision": "deny"} + await restored.append(KEY, [reply]) + loaded = await restored.load(KEY) + assert loaded is not None and len(loaded) == 2 + + asyncio.run(scenario()) + + def test_missing_mirror_never_certifies_the_checkpoint(self): + async def scenario(): + with pytest.raises( + checkpoint.ContinuationCheckpointError, match="did not acknowledge" + ) as error: + await capture(checkpoint.CheckpointSessionStore(PROJECT)) + assert error.value.code == "checkpoint_sdk_timeout" + + asyncio.run(scenario()) + + @pytest.mark.parametrize( + "bad_key", + [ + {**KEY, "project_key": "other"}, + {**KEY, "session_id": "../../escape"}, + {**KEY, "subpath": "subagents/agent-other"}, + ], + ) + def test_a_mirror_gap_prevents_later_checkpoint_acknowledgement(self, bad_key): + async def scenario(): + store = checkpoint.CheckpointSessionStore(PROJECT) + with pytest.raises(checkpoint.ContinuationCheckpointError): + await store.append(bad_key, [assistant()]) + await store.append(KEY, [assistant()]) + with pytest.raises(checkpoint.ContinuationCheckpointError, match="Unexpected SDK"): + await capture(store) + + asyncio.run(scenario()) + + def test_mismatched_action_fails_instead_of_waiting_for_a_later_match(self): + async def scenario(): + store = checkpoint.CheckpointSessionStore(PROJECT) + entry = assistant() + entry["message"]["content"][0]["input"] = {"file_path": "/different"} + await store.append(KEY, [entry]) + with pytest.raises(checkpoint.ContinuationCheckpointError, match="disagrees"): + await capture(store) + + asyncio.run(scenario()) + + def test_completed_tool_cannot_be_recorded_as_pending(self): + async def scenario(): + store = checkpoint.CheckpointSessionStore(PROJECT) + result: Any = { + "type": "user", + "message": {"content": [{"type": "tool_result", "tool_use_id": "toolu_owned"}]}, + } + await store.append( + KEY, + [assistant(), result], + ) + with pytest.raises(checkpoint.ContinuationCheckpointError, match="already has"): + await capture(store) + + asyncio.run(scenario()) + + def test_buffer_limit_failure_does_not_silently_drop_entries(self, monkeypatch): + async def scenario(): + store = checkpoint.CheckpointSessionStore(PROJECT) + monkeypatch.setattr(checkpoint, "MAX_CHECKPOINT_ENTRIES", 1) + with pytest.raises(checkpoint.ContinuationCheckpointError, match="limit"): + await store.append(KEY, [assistant(), {"type": "mode"}]) + with pytest.raises(checkpoint.ContinuationCheckpointError, match="limit"): + await capture(store) + + asyncio.run(scenario()) + + def test_cancelled_checkpoint_wait_propagates_cancellation(self): + async def scenario(): + task = asyncio.create_task(capture(checkpoint.CheckpointSessionStore(PROJECT))) + await asyncio.sleep(0) + task.cancel() + with pytest.raises(asyncio.CancelledError): + await task + + asyncio.run(scenario()) + + def test_unwritten_session_load_returns_none(self): + assert asyncio.run(checkpoint.CheckpointSessionStore(PROJECT).load(KEY)) is None + + def test_transcript_cannot_mix_another_session_into_the_saved_conversation(self): + async def scenario(): + store = checkpoint.CheckpointSessionStore(PROJECT) + with pytest.raises(checkpoint.ContinuationCheckpointError, match="another SDK session"): + await store.append( + KEY, [assistant(sessionId="22222222-2222-4222-8222-bbbbbbbbbbbb")] + ) + + asyncio.run(scenario()) + + def test_duplicate_tool_identity_is_ambiguous(self): + async def scenario(): + store = checkpoint.CheckpointSessionStore(PROJECT) + await store.append(KEY, [assistant(), assistant(uuid="different-entry")]) + with pytest.raises(checkpoint.ContinuationCheckpointError, match="duplicate"): + await capture(store) + + asyncio.run(scenario()) + + +class VersionedS3: + """A versioned object service with commit-then-disconnect failure injection.""" + + def __init__(self): + self.objects = {} + self.calls = [] + self.write_error = None + self.lose_write_reply = False + self.versioned = True + self.bad_checksum = False + self.bad_length = False + self.streams = [] + + def put_object(self, **kwargs): + self.calls.append(("put", kwargs)) + if self.write_error: + raise self.write_error + object_key = (kwargs["Bucket"], kwargs["Key"]) + if kwargs.get("IfNoneMatch") == "*" and object_key in self.objects: + raise ClientError({"Error": {"Code": "PreconditionFailed"}}, "PutObject") + versions = self.objects.setdefault(object_key, []) + version = f"version-{len(versions) + 1}" if self.versioned else "null" + versions.append((version, kwargs["Body"])) + if self.lose_write_reply: + raise EndpointConnectionError(endpoint_url="https://s3.invalid") + return {"VersionId": version} + + def get_object(self, **kwargs): + self.calls.append(("get", kwargs)) + versions = self.objects.get((kwargs["Bucket"], kwargs["Key"])) + if not versions: + raise ClientError({"Error": {"Code": "NoSuchKey"}}, "GetObject") + requested = kwargs.get("VersionId") + version, body = ( + next(v for v in versions if v[0] == requested) if requested else versions[-1] + ) + stream = io.BytesIO(body) + self.streams.append(stream) + return { + "Body": stream, + "VersionId": version, + "ContentLength": len(body) + int(self.bad_length), + "ChecksumSHA256": "wrong" + if self.bad_checksum + else base64.b64encode(hashlib.sha256(body).digest()).decode(), + } + + +class TestImmutableStorage: + def test_missing_bucket_reports_storage_configuration_failure(self): + with pytest.raises(checkpoint.ContinuationCheckpointError) as error: + checkpoint.S3ContinuationCheckpoints("") + assert error.value.code == "checkpoint_storage_unavailable" + + @pytest.fixture + def storage(self): + client = VersionedS3() + return checkpoint.S3ContinuationCheckpoints("private-checkpoints", client=client), client + + def test_default_client_refuses_ambient_credentials(self, monkeypatch): + import aws_session + + client_factory = Mock() + monkeypatch.setattr(aws_session, "is_scoped", lambda: False) + monkeypatch.setattr(aws_session, "tenant_client", client_factory) + with pytest.raises(checkpoint.ContinuationCheckpointError, match="task-scoped credentials"): + checkpoint.S3ContinuationCheckpoints("private-checkpoints") + client_factory.assert_not_called() + + def test_default_client_uses_attributed_scoped_factory(self, monkeypatch): + import aws_session + + client_factory = Mock(return_value=VersionedS3()) + monkeypatch.setattr(aws_session, "is_scoped", lambda: True) + monkeypatch.setattr(aws_session, "tenant_client", client_factory) + store = checkpoint.S3ContinuationCheckpoints("private-checkpoints") + assert store.client is client_factory.return_value + assert client_factory.call_args.args == ("s3",) + assert client_factory.call_args.kwargs["config"].retries == {"total_max_attempts": 1} + + def test_save_requires_read_back_and_returns_a_version_pinned_receipt(self, storage, body): + store, client = storage + receipt = store.save(body, IDENTITY) + assert [call[0] for call in client.calls] == ["put", "get"] + assert receipt.key == IDENTITY.prefix + hashlib.sha256(body).hexdigest() + ".json" + assert client.calls[0][1]["IfNoneMatch"] == "*" + assert client.calls[0][1]["ServerSideEncryption"] == "AES256" + assert store.load(receipt, IDENTITY) == body + assert client.calls[-1][1]["VersionId"] == receipt.version_id + assert all(stream.closed for stream in client.streams) + + def test_unreadable_saved_version_reports_storage_failure(self, storage, body): + store, client = storage + receipt = store.save(body, IDENTITY) + client.get_object = Mock( + side_effect=ClientError({"Error": {"Code": "AccessDenied"}}, "GetObject") + ) + with pytest.raises(checkpoint.ContinuationCheckpointError) as error: + store.load(receipt, IDENTITY) + assert error.value.code == "checkpoint_storage_unavailable" + + def test_lost_write_reply_recovers_from_exact_read_back(self, storage, body): + store, client = storage + client.lose_write_reply = True + receipt = store.save(body, IDENTITY) + assert receipt.version_id == "version-1" + assert store.save(body, IDENTITY) == receipt + assert len(client.objects[(store.bucket, receipt.key)]) == 1 + + def test_uncommitted_write_failure_never_returns_a_receipt(self, storage, body): + store, client = storage + client.write_error = ClientError({"Error": {"Code": "AccessDenied"}}, "PutObject") + with pytest.raises(checkpoint.ContinuationCheckpointError, match="keep the current worker"): + store.save(body, IDENTITY) + assert not client.objects + + @pytest.mark.parametrize("failure", ["versioned", "bad_checksum", "bad_length"]) + def test_unverifiable_storage_never_acknowledges(self, storage, body, failure): + store, client = storage + setattr(client, failure, failure != "versioned") + with pytest.raises( + checkpoint.ContinuationCheckpointError, match="could not be verified" + ) as error: + store.save(body, IDENTITY) + assert error.value.code == "checkpoint_storage_unverified" + assert all(stream.closed for stream in client.streams) + + def test_load_keeps_the_original_version_even_if_current_key_changes(self, storage, body): + store, client = storage + receipt = store.save(body, IDENTITY) + client.put_object(Bucket=store.bucket, Key=receipt.key, Body=b"replacement") + assert store.load(receipt, IDENTITY) == body + + @pytest.mark.parametrize("field", ["task_id", "attempt_id", "request_id", "user_id", "repo"]) + def test_cross_identity_data_is_rejected_before_any_io(self, storage, body, field): + store, client = storage + wrong = replace(IDENTITY, **{field: "other"}) + with pytest.raises(checkpoint.ContinuationCheckpointError, match="identity"): + store.save(body, wrong) + assert not client.calls + + def test_receipt_cannot_redirect_a_read_outside_its_task(self, storage, body): + store, client = storage + receipt = store.save(body, IDENTITY) + count = len(client.calls) + with pytest.raises(checkpoint.ContinuationCheckpointError, match="outside this task"): + store.load(replace(receipt, key="continuations/other/data"), IDENTITY) + assert len(client.calls) == count + + def test_valid_transport_checksum_does_not_replace_receipt_integrity(self, storage, body): + store, client = storage + receipt = store.save(body, IDENTITY) + client.objects[(store.bucket, receipt.key)][0] = (receipt.version_id, b"x" * len(body)) + with pytest.raises(checkpoint.ContinuationCheckpointError, match="does not match"): + store.load(receipt, IDENTITY) + + +class TestEnvelope: + @pytest.mark.parametrize( + "change", + [ + lambda data: data.update(version=2), + lambda data: data.update(version=True), + lambda data: data.update(session_id="../../other"), + lambda data: data["action"].update(tool_input_sha256="0" * 64), + lambda data: data["action"].update(tool_input={"file_path": "changed"}), + lambda data: data.update(entries=[]), + lambda data: data.update(environment={"AWS_SECRET_ACCESS_KEY": "synthetic"}), + ], + ) + def test_corrupt_or_incompatible_envelope_fails(self, body, change): + data = json.loads(body) + change(data) + with pytest.raises(checkpoint.ContinuationCheckpointError): + checkpoint.decode_checkpoint(json.dumps(data).encode(), IDENTITY) + + @pytest.mark.parametrize("body", [b"{", b"\xff", b"null", b""]) + def test_invalid_json_fails(self, body): + with pytest.raises(checkpoint.ContinuationCheckpointError) as error: + checkpoint.decode_checkpoint(body, IDENTITY) + assert error.value.code == ( + "checkpoint_failed" if body in (b"null", b"") else "checkpoint_invalid_json" + ) + + @pytest.mark.parametrize("value", ["../task", "", "task/other", "x" * 129]) + def test_path_components_cannot_escape_task_prefix(self, value): + with pytest.raises(checkpoint.ContinuationCheckpointError): + replace(IDENTITY, task_id=value) + + +def test_invalid_json_reports_a_specific_checkpoint_code(): + with pytest.raises(checkpoint.ContinuationCheckpointError) as error: + checkpoint._encode({"not_json": object()}) + assert error.value.code == "checkpoint_invalid_json" + + +def test_unspecified_checkpoint_failure_does_not_claim_invalid_data(): + assert checkpoint.ContinuationCheckpointError("Unknown failure").code == "checkpoint_failed" diff --git a/agent/tests/test_continuation_storage.py b/agent/tests/test_continuation_storage.py new file mode 100644 index 000000000..7cf38d134 --- /dev/null +++ b/agent/tests/test_continuation_storage.py @@ -0,0 +1,283 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Storage acknowledgement must survive retries without accepting partial state.""" + +from __future__ import annotations + +import base64 +import hashlib +import io +import json +from dataclasses import asdict, replace +from decimal import Decimal +from types import SimpleNamespace + +import pytest + +import continuation_storage as storage +from continuation_session import CheckpointIdentity, CheckpointReceipt +from models import RepoSetup + +IDENTITY = CheckpointIdentity("task", "attempt", "request", "user", "owner/repo") + + +class RecordingStream(io.BytesIO): + def __init__(self, body, max_read=1024 * 1024): + super().__init__(body) + self.read_sizes = [] + self.max_read = max_read + + def read(self, size=-1): + assert 0 < size <= self.max_read + self.read_sizes.append(size) + return super().read(size) + + +class VersionedS3: + def __init__(self): + self.versions = {} + self.calls = [] + self.streams = [] + self.lost_reply = False + self.read_override = {} + + def put_object(self, **kwargs): + self.calls.append(("put", kwargs)) + body = kwargs["Body"] if isinstance(kwargs["Body"], bytes) else kwargs["Body"].read() + assert len(body) == kwargs.get("ContentLength", len(body)) + assert kwargs["IfNoneMatch"] == "*" + assert kwargs["ServerSideEncryption"] == "AES256" + assert base64.b64encode(hashlib.sha256(body).digest()).decode() == kwargs["ChecksumSHA256"] + key = kwargs["Key"] + if key in self.versions: + raise RuntimeError("PreconditionFailed") + self.versions[key] = [body] + if self.lost_reply: + raise TimeoutError("reply lost after commit") + return {"VersionId": "1"} + + def get_object(self, **kwargs): + self.calls.append(("get", kwargs)) + bodies = self.versions[kwargs["Key"]] + version = kwargs.get("VersionId", str(len(bodies))) + body = bodies[int(version) - 1] + stream = RecordingStream( + body, 16 * 1024 * 1024 + 1 if kwargs["Key"].endswith(".json") else 1024 * 1024 + ) + self.streams.append(stream) + return { + "Body": stream, + "VersionId": version, + "ContentLength": len(body), + "ChecksumSHA256": base64.b64encode(hashlib.sha256(body).digest()).decode(), + **self.read_override, + } + + +@pytest.fixture +def prepared(tmp_path): + source = tmp_path / "workspace.tar" + source.write_bytes(b"saved workspace\n" * 180_000) + client = VersionedS3() + adapter = storage.S3ContinuationStorage("owned-bucket", client=client) + return source, client, adapter + + +def manifest(receipt): + digest = "a" * 64 + return storage.ContinuationManifest( + IDENTITY, + CheckpointReceipt(IDENTITY.prefix + digest + ".json", "1", digest, 10), + receipt, + storage.ContinuationContext( + RepoSetup(repo_dir="/workspace/task", branch="agent/test", build_before=False), + "Original user prompt", + "Original system prompt", + "coding/new-task-v1", + "1", + ), + ) + + +class TestPinnedFiles: + @pytest.mark.parametrize("size", [Decimal("42.5"), Decimal("NaN"), "42", 42.0, True]) + def test_ddb_receipt_does_not_coerce_invalid_sizes(self, size): + digest = "a" * 64 + value = asdict( + storage.FileReceipt( + "manifest", IDENTITY.prefix + "manifest/" + digest + ".json", "1", digest, 42 + ) + ) + value["size_bytes"] = size + with pytest.raises(storage.ContinuationStorageError): + storage.FileReceipt.from_record(value, IDENTITY) + + def test_ddb_integral_receipt_roundtrip(self): + digest = "a" * 64 + receipt = storage.FileReceipt( + "manifest", IDENTITY.prefix + "manifest/" + digest + ".json", "1", digest, 42 + ) + assert ( + storage.FileReceipt.from_record( + {**asdict(receipt), "size_bytes": Decimal("42")}, IDENTITY + ) + == receipt + ) + + def test_streamed_roundtrip_pins_version_even_after_current_object_changes( + self, prepared, tmp_path + ): + source, client, adapter = prepared + receipt = adapter.save_workspace(source, IDENTITY) + client.versions[receipt.key].append(b"later unrelated contents") + destination = tmp_path / "restored.tar" + adapter.download_workspace(receipt, IDENTITY, destination) + assert destination.read_bytes() == source.read_bytes() + assert client.calls[-1][1]["VersionId"] == "1" + assert all(stream.closed for stream in client.streams) + assert len(client.streams[-1].read_sizes) > 2 + assert destination.stat().st_mode & 0o777 == 0o600 + assert not list(tmp_path.glob(".continuation-*")) + + def test_lost_put_reply_and_repeated_put_require_successful_readback(self, prepared): + source, client, adapter = prepared + client.lost_reply = True + first = adapter.save_workspace(source, IDENTITY) + assert adapter.save_workspace(source, IDENTITY) == first + assert len(client.versions[first.key]) == 1 + + @pytest.mark.parametrize( + "override", + [ + {"VersionId": "null"}, + {"VersionId": ""}, + {"ContentLength": 1}, + {"ContentLength": True}, + {"ChecksumSHA256": None}, + ], + ) + def test_bad_readback_never_acknowledges_worker_release(self, prepared, override): + source, client, adapter = prepared + client.read_override = override + with pytest.raises(storage.ContinuationStorageError, match="retain the worker"): + adapter.save_workspace(source, IDENTITY) + assert client.streams[-1].closed + + @pytest.mark.parametrize("payload", [b"short", b"x" * 2_880_000]) + def test_corrupt_or_truncated_stream_leaves_no_destination(self, prepared, tmp_path, payload): + source, client, adapter = prepared + receipt = adapter.save_workspace(source, IDENTITY) + bad_stream = RecordingStream(payload) + client.read_override["Body"] = bad_stream + destination = tmp_path / "restored.tar" + with pytest.raises(storage.ContinuationStorageError): + adapter.download_workspace(receipt, IDENTITY, destination) + assert not destination.exists() + assert not list(tmp_path.glob(".continuation-*")) + assert bad_stream.closed + + def test_existing_destination_is_never_overwritten(self, prepared, tmp_path): + source, _, adapter = prepared + receipt = adapter.save_workspace(source, IDENTITY) + destination = tmp_path / "restored.tar" + destination.write_text("keep me") + with pytest.raises(FileExistsError): + adapter.download_workspace(receipt, IDENTITY, destination) + assert destination.read_text() == "keep me" + assert not list(tmp_path.glob(".continuation-*")) + + def test_task_and_attempt_receipts_cannot_cross_boundaries(self, prepared, tmp_path): + source, client, adapter = prepared + receipt = adapter.save_workspace(source, IDENTITY) + count = len(client.calls) + for identity in (replace(IDENTITY, task_id="other"), replace(IDENTITY, attempt_id="other")): + with pytest.raises(storage.ContinuationStorageError, match="receipt"): + adapter.download_workspace(receipt, identity, tmp_path / "bad.tar") + assert len(client.calls) == count + + def test_disk_pressure_is_distinct_and_detected_before_network( + self, prepared, tmp_path, monkeypatch + ): + source, client, adapter = prepared + receipt = adapter.save_workspace(source, IDENTITY) + count = len(client.calls) + monkeypatch.setattr(storage.shutil, "disk_usage", lambda _: SimpleNamespace(free=0)) + with pytest.raises(storage.ContinuationStorageError) as raised: + adapter.download_workspace(receipt, IDENTITY, tmp_path / "bad.tar") + assert raised.value.code == "disk_pressure" + assert len(client.calls) == count + + def test_disk_pressure_during_download_removes_partial_file( + self, prepared, tmp_path, monkeypatch + ): + source, _, adapter = prepared + receipt = adapter.save_workspace(source, IDENTITY) + readings = iter([10**12, 10**12, 0]) + monkeypatch.setattr( + storage.shutil, "disk_usage", lambda _: SimpleNamespace(free=next(readings)) + ) + with pytest.raises(storage.ContinuationStorageError) as raised: + adapter.download_workspace(receipt, IDENTITY, tmp_path / "bad.tar") + assert raised.value.code == "disk_pressure" + assert not list(tmp_path.glob(".continuation-*")) + assert not (tmp_path / "bad.tar").exists() + + def test_upload_rejects_symlinks_and_hardlinks(self, prepared, tmp_path): + source, client, adapter = prepared + link = tmp_path / "link" + link.symlink_to(source) + with pytest.raises(OSError): + adapter.save_workspace(link, IDENTITY) + link.unlink() + link.hardlink_to(source) + with pytest.raises(storage.ContinuationStorageError, match="invalid"): + adapter.save_workspace(source, IDENTITY) + assert not client.calls + + def test_transfer_deadline_closes_stream_and_discards_partial_file( + self, prepared, tmp_path, monkeypatch + ): + source, _, adapter = prepared + receipt = adapter.save_workspace(source, IDENTITY) + times = iter([0, 0, 0, 301]) + monkeypatch.setattr(storage.time, "monotonic", lambda: next(times)) + with pytest.raises(storage.ContinuationStorageError) as raised: + adapter.download_workspace(receipt, IDENTITY, tmp_path / "bad.tar") + assert raised.value.code == "transfer_timeout" + assert not (tmp_path / "bad.tar").exists() + + +class TestManifest: + def test_complete_manifest_preserves_baseline_and_exact_receipts(self, prepared): + source, _, adapter = prepared + record = manifest(adapter.save_workspace(source, IDENTITY)) + receipt = adapter.save_manifest(record) + restored = adapter.load_manifest(receipt, IDENTITY) + assert restored == record + assert restored.context.setup.build_before is False + assert restored.workspace.version_id == "1" + + @pytest.mark.parametrize( + "change", + [ + lambda value: value.update(version=True), + lambda value: value["identity"].update(request_id="another"), + lambda value: value["context"].update(github_token="must not be serialized"), + lambda value: value["context"]["setup"].update(build_before="false"), + lambda value: value["context"]["setup"].update(repo_dir="relative"), + lambda value: value["workspace"].update(kind="manifest"), + lambda value: value["conversation"].update(version_id="null"), + ], + ) + def test_manifest_rejects_mixed_identity_unknown_fields_and_invalid_context( + self, prepared, change + ): + source, _, adapter = prepared + body = manifest(adapter.save_workspace(source, IDENTITY)).encode(adapter.limits) + value = json.loads(body) + change(value) + with pytest.raises(storage.ContinuationCheckpointError): + storage.ContinuationManifest.decode( + json.dumps(value).encode(), IDENTITY, adapter.limits + ) diff --git a/agent/tests/test_continuation_usage.py b/agent/tests/test_continuation_usage.py new file mode 100644 index 000000000..81eb9b046 --- /dev/null +++ b/agent/tests/test_continuation_usage.py @@ -0,0 +1,96 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Accounting cannot silently reset or estimate spend after replacement.""" + +import asyncio +from copy import deepcopy +from types import SimpleNamespace +from typing import Any +from unittest.mock import AsyncMock + +import pytest + +from continuation_session import ContinuationCheckpointError +from continuation_usage import read_usage + +RESPONSE: dict[str, Any] = { + "session": { + "total_cost_usd": 0.03, + "model_usage": { + "sonnet": { + "costUSD": 0.01, + "inputTokens": 100, + "outputTokens": 20, + "cacheReadInputTokens": 30, + "cacheCreationInputTokens": 40, + }, + "haiku": { + "costUSD": 0.02, + "inputTokens": 5, + "outputTokens": 6, + "cacheReadInputTokens": 7, + "cacheCreationInputTokens": 8, + }, + }, + } +} + + +def test_reads_exact_current_process_spending_across_models(): + send = AsyncMock(return_value=RESPONSE) + snapshot = asyncio.run( + read_usage(SimpleNamespace(_query=SimpleNamespace(_send_control_request=send))) + ) + assert snapshot.cost_usd == 0.03 + assert snapshot.tokens == { + "input_tokens": 105, + "output_tokens": 26, + "cache_read_input_tokens": 37, + "cache_creation_input_tokens": 48, + } + send.assert_awaited_once_with({"subtype": "get_usage"}) + + +@pytest.mark.parametrize( + "mutation", ["negative", "nan", "missing", "empty_models", "bad_tokens", "different_cost"] +) +def test_incomplete_or_invalid_accounting_cannot_publish_a_checkpoint(mutation): + response = deepcopy(RESPONSE) + if mutation in {"negative", "nan"}: + response["session"]["total_cost_usd"] = -1 if mutation == "negative" else float("nan") + elif mutation == "missing": + response = {} + elif mutation == "empty_models": + response["session"]["model_usage"] = {} + elif mutation == "bad_tokens": + response["session"]["model_usage"]["haiku"]["inputTokens"] = "5" + else: + response["session"]["model_usage"]["haiku"]["costUSD"] = 0.5 + client = SimpleNamespace( + _query=SimpleNamespace(_send_control_request=AsyncMock(return_value=response)) + ) + with pytest.raises(ContinuationCheckpointError): + asyncio.run(read_usage(client)) + + +def test_unverified_sdk_upgrade_requires_explicit_accounting_validation(monkeypatch): + monkeypatch.setattr("continuation_usage.importlib.metadata.version", lambda _: "0.3.0") + with pytest.raises(ContinuationCheckpointError, match="verified SDK") as error: + asyncio.run(read_usage(None)) + assert error.value.code == "checkpoint_sdk_unverified" + + +def test_missing_sdk_accounting_interface_is_not_invalid_checkpoint_data(): + with pytest.raises(ContinuationCheckpointError) as error: + asyncio.run(read_usage(SimpleNamespace())) + assert error.value.code == "checkpoint_sdk_unverified" + + +def test_sdk_accounting_timeout_reports_transport_failure(): + client = SimpleNamespace( + _query=SimpleNamespace(_send_control_request=AsyncMock(side_effect=TimeoutError)) + ) + with pytest.raises(ContinuationCheckpointError) as error: + asyncio.run(read_usage(client)) + assert error.value.code == "checkpoint_sdk_timeout" diff --git a/agent/tests/test_continuation_workspace.py b/agent/tests/test_continuation_workspace.py new file mode 100644 index 000000000..e234b6c52 --- /dev/null +++ b/agent/tests/test_continuation_workspace.py @@ -0,0 +1,502 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Real offline Git/working-tree recovery and hostile archive boundary checks.""" + +from __future__ import annotations + +import hashlib +import io +import json +import os +import stat +import subprocess +import tarfile +import time +from dataclasses import replace +from pathlib import Path + +import pytest + +import continuation_workspace as workspace +from continuation_session import CheckpointIdentity + +IDENTITY = CheckpointIdentity("task", "attempt", "approval", "owner", "example/repository") + + +def git(root, *args): + env = {key: value for key, value in os.environ.items() if not key.startswith("GIT_")} + env.update( + GIT_CONFIG_NOSYSTEM="1", + GIT_CONFIG_GLOBAL=os.devnull, + GIT_AUTHOR_NAME="Fixture", + GIT_AUTHOR_EMAIL="fixture@example.invalid", + GIT_COMMITTER_NAME="Fixture", + GIT_COMMITTER_EMAIL="fixture@example.invalid", + LC_ALL="C", + ) + return subprocess.check_output( + ["git", "-c", f"core.hooksPath={os.devnull}", *args], + cwd=root, + env=env, + stderr=subprocess.PIPE, + timeout=10, + ) + + +@pytest.fixture +def repo(tmp_path): + root = tmp_path / "repo" + root.mkdir() + git(root, "init", "--quiet", "--template=", "--initial-branch=work/task") + (root / "tracked.txt").write_text("committed\n") + (root / "binary.dat").write_bytes(b"\0original\xff") + (root / "deleted.txt").write_text("delete from working tree") + (root / "renamed.txt").write_text("rename in index") + (root / "run.sh").write_text("#!/bin/sh\nexit 0\n") + (root / ".gitignore").write_text("ignored/\n") + git(root, "add", ".") + git(root, "commit", "--quiet", "-m", "base") + (root / "local-commit.txt").write_text("unpushed work") + git(root, "add", ".") + git(root, "commit", "--quiet", "-m", "unpushed commit") + git(root, "tag", "local-tag") + git(root, "branch", "another-branch") + git(root, "remote", "add", "origin", "https://github.com/example/repository.git") + (root / "tracked.txt").write_text("staged\n") + (root / "binary.dat").write_bytes(b"\0staged\xfe") + git(root, "add", "tracked.txt", "binary.dat") + git(root, "mv", "renamed.txt", "renamed-new.txt") + (root / "tracked.txt").write_text("unstaged\n") + (root / "binary.dat").write_bytes(b"\0unstaged\xfd") + (root / "deleted.txt").unlink() + (root / "run.sh").chmod(0o755) + (root / "untracked 雪\nfile.txt").write_text("untracked") + (root / "ignored").mkdir() + (root / "ignored/important.bin").write_bytes(b"\0ignored but required") + (root / ".git/info").mkdir(exist_ok=True) + (root / ".git/info/exclude").write_text("local-ignored.txt\n") + (root / "local-ignored.txt").write_text("local ignore rules must survive") + (root / "empty").mkdir() + os.symlink("tracked.txt", root / "relative-link") + git(root, "config", "--local", "http.extraHeader", "SYNTHETIC_AUTH_MUST_NOT_BE_COPIED") + hooks = root / ".git/hooks" + hooks.mkdir(exist_ok=True) + marker = tmp_path / "hook-ran" + hook = hooks / "post-checkout" + hook.write_text(f"#!/bin/sh\ntouch '{marker}'\n") + hook.chmod(0o755) + return root + + +def saved(repo): + archive = repo.parent / "workspace.tar" + receipt = workspace.capture_workspace(repo, archive, IDENTITY) + return archive, receipt + + +def remove_original(repo): + # Preserve the original owned test tree for comparison without leaving it at + # the stable path that the replacement worker must use. + original = repo.with_name("original") + repo.rename(original) + return original + + +def inventory(root): + result = {} + for directory, dirs, files in os.walk(root, followlinks=False): + if Path(directory) == root: + dirs.remove(".git") + for name in dirs + files: + path = Path(directory) / name + info = path.lstat() + value = ( + os.readlink(path) + if path.is_symlink() + else path.read_bytes() + if path.is_file() + else None + ) + result[str(path.relative_to(root))] = ( + stat.S_IFMT(info.st_mode), + info.st_mode & 0o777, + value, + ) + return result + + +def rewrite(archive, *, mutate_manifest=None, mutate_member=None, extras=()): + records = [] + with tarfile.open(archive, "r:") as tar: + for member in tar: + stream = tar.extractfile(member) if member.isfile() else None + data = stream.read() if stream else None + if stream: + stream.close() + if member.name == "manifest.json" and mutate_manifest: + assert data is not None + manifest = json.loads(data) + mutate_manifest(manifest) + data = json.dumps(manifest).encode() + member.size = len(data) + if mutate_member: + member, data = mutate_member(member, data) + records.append((member, data)) + with tarfile.open(archive, "w", format=tarfile.PAX_FORMAT) as tar: + for member, data in [*records, *extras]: + tar.addfile(member, io.BytesIO(data) if data is not None else None) + return hashlib.sha256(archive.read_bytes()).hexdigest() + + +class TestWorkspaceRoundTrip: + def test_preserves_commits_staged_unstaged_binary_deleted_untracked_ignored_and_modes( + self, repo + ): + expected_files = inventory(repo) + expected_index = git(repo, "ls-files", "--stage", "-z") + expected_diff = git(repo, "diff", "--binary", "--no-ext-diff", "--no-textconv") + expected_refs = git(repo, "for-each-ref", "--format=%(objectname) %(refname)") + expected_ignored = git( + repo, "ls-files", "--others", "--ignored", "--exclude-standard", "-z" + ) + archive, receipt = saved(repo) + assert receipt.head == git(repo, "rev-parse", "HEAD").decode().strip() + assert receipt.branch == "work/task" + assert archive.stat().st_mode & 0o777 == 0o600 + assert b"SYNTHETIC_AUTH_MUST_NOT_BE_COPIED" not in archive.read_bytes() + original = remove_original(repo) + restored = workspace.restore_workspace( + archive, repo, IDENTITY, expected_sha256=receipt.sha256 + ) + assert restored == receipt + assert inventory(repo) == expected_files + assert git(repo, "ls-files", "--stage", "-z") == expected_index + assert git(repo, "diff", "--binary", "--no-ext-diff", "--no-textconv") == expected_diff + assert git(repo, "for-each-ref", "--format=%(objectname) %(refname)") == expected_refs + assert ( + git(repo, "ls-files", "--others", "--ignored", "--exclude-standard", "-z") + == expected_ignored + ) + assert git(repo, "log", "--format=%s").splitlines() == [b"unpushed commit", b"base"] + assert ( + git(repo, "remote", "get-url", "origin").strip() + == b"https://github.com/example/repository.git" + ) + assert ( + git(repo, "config", "--local", "credential.helper").strip() + == b"!gh auth git-credential" + ) + assert "SYNTHETIC_AUTH" not in (repo / ".git/config").read_text() + assert not (repo / ".git/hooks/post-checkout").exists() + assert not (repo.parent / "hook-ran").exists() + assert inventory(original) == expected_files + assert not list(repo.parent.glob(".workspace-*")) + + def test_detached_head_is_preserved(self, repo): + git(repo, "checkout", "--detach") + archive, receipt = saved(repo) + remove_original(repo) + workspace.restore_workspace(archive, repo, IDENTITY, expected_sha256=receipt.sha256) + assert receipt.branch is None + assert git(repo, "rev-parse", "--abbrev-ref", "HEAD").strip() == b"HEAD" + + def test_external_symlink_is_saved_as_a_leaf_without_reading_its_target(self, repo): + secret = repo.parent / "outside" + secret.write_text("SYNTHETIC_OUTSIDE_DATA_NOT_IN_ARCHIVE") + os.symlink(secret, repo / "outside-link") + archive, receipt = saved(repo) + assert secret.read_bytes() not in archive.read_bytes() + remove_original(repo) + workspace.restore_workspace(archive, repo, IDENTITY, expected_sha256=receipt.sha256) + assert os.readlink(repo / "outside-link") == str(secret) + assert secret.read_text() == "SYNTHETIC_OUTSIDE_DATA_NOT_IN_ARCHIVE" + + def test_external_diff_configuration_does_not_execute_during_capture(self, repo): + script = repo.parent / "external-diff" + marker = repo.parent / "diff-ran" + script.write_text(f"#!/bin/sh\ntouch '{marker}'\nexit 1\n") + script.chmod(0o755) + git(repo, "config", "diff.external", str(script)) + saved(repo) + assert not marker.exists() + + +class TestCaptureFailures: + def test_git_deadline_stops_the_process_group(self, repo): + git(repo, "config", "alias.checkpoint-wait", "!sleep 10") + started = time.monotonic() + with pytest.raises(workspace.WorkspaceCheckpointError) as error: + workspace._git(repo, ["checkpoint-wait"], workspace.WorkspaceLimits(git_timeout_s=1)) + assert error.value.code == "git_timeout" + assert time.monotonic() - started < 5 + + @pytest.mark.parametrize("repo_name", ["../repo", "owner/..", "https://github.com/owner/repo"]) + def test_repository_identity_must_also_be_restorable(self, repo, repo_name): + destination = repo.parent / "workspace.tar" + with pytest.raises(workspace.WorkspaceCheckpointError) as error: + workspace.capture_workspace(repo, destination, replace(IDENTITY, repo=repo_name)) + assert error.value.code == "invalid_identity" + assert not destination.exists() + + def test_submodule_index_entry_fails_without_fetching_it(self, repo): + head = git(repo, "rev-parse", "HEAD").decode().strip() + git(repo, "update-index", "--add", "--cacheinfo", f"160000,{head},submodule") + with pytest.raises(workspace.WorkspaceCheckpointError) as error: + saved(repo) + assert error.value.code == "unsupported_git" + + @pytest.mark.parametrize( + "limits,code", + [ + (workspace.WorkspaceLimits(max_bytes=32), "size_limit"), + (workspace.WorkspaceLimits(max_entries=1), "entry_limit"), + ], + ) + def test_limits_do_not_publish_partial_archive(self, repo, limits, code): + expected = inventory(repo) + destination = repo.parent / "limited.tar" + with pytest.raises(workspace.WorkspaceCheckpointError) as error: + workspace.capture_workspace(repo, destination, IDENTITY, limits=limits) + assert error.value.code == code + assert not destination.exists() + assert inventory(repo) == expected + assert not list(repo.parent.glob(".workspace-capture-*")) + + @pytest.mark.parametrize( + "path", ["MERGE_HEAD", "index.lock", "shallow", "objects/info/alternates"] + ) + def test_active_or_nonportable_git_layout_fails(self, repo, path): + target = repo / ".git" / path + target.parent.mkdir(exist_ok=True, parents=True) + target.touch() + with pytest.raises(workspace.WorkspaceCheckpointError) as error: + saved(repo) + assert error.value.code == "unsupported_git" + + @pytest.mark.parametrize("flag", ["--assume-unchanged", "--skip-worktree", "intent-to-add"]) + def test_nonportable_index_flags_fail_instead_of_losing_state(self, repo, flag): + if flag == "intent-to-add": + (repo / "intent.txt").write_text("not staged yet") + git(repo, "add", "-N", "intent.txt") + else: + git(repo, "update-index", flag, "tracked.txt") + with pytest.raises(workspace.WorkspaceCheckpointError) as error: + saved(repo) + assert error.value.code == "unsupported_git" + + @pytest.mark.parametrize("kind", ["fifo", "hardlink", "nested-git"]) + def test_special_files_are_rejected(self, repo, kind): + target = repo / "unsupported" + if kind == "fifo": + os.mkfifo(target) + elif kind == "hardlink": + os.link(repo / "tracked.txt", target) + else: + target.mkdir() + (target / ".git").mkdir() + with pytest.raises(workspace.WorkspaceCheckpointError): + saved(repo) + assert not (repo.parent / "workspace.tar").exists() + + def test_existing_archive_is_not_overwritten(self, repo): + destination = repo.parent / "workspace.tar" + destination.write_bytes(b"keep") + with pytest.raises(workspace.WorkspaceCheckpointError) as error: + saved(repo) + assert error.value.code == "destination_exists" + assert destination.read_bytes() == b"keep" + + def test_archive_cannot_be_written_inside_the_workspace(self, repo): + with pytest.raises(workspace.WorkspaceCheckpointError) as error: + workspace.capture_workspace(repo, repo / "backup.tar", IDENTITY) + assert error.value.code == "invalid_path" + + def test_worktree_mutation_prevents_acknowledgement(self, repo, monkeypatch): + real_read = workspace._DigestReader.read + changed = False + + def mutate(reader, size=-1): + nonlocal changed + block = real_read(reader, size) + if not changed: + changed = True + (repo / "created-during-capture").write_text("concurrent writer") + return block + + monkeypatch.setattr(workspace._DigestReader, "read", mutate) + with pytest.raises(workspace.WorkspaceCheckpointError) as error: + saved(repo) + assert error.value.code == "workspace_changed" + assert not (repo.parent / "workspace.tar").exists() + + def test_racing_destination_is_not_replaced(self, repo, monkeypatch): + real_link = os.link + destination = repo.parent / "workspace.tar" + + def race(source, target): + destination.write_bytes(b"other writer") + real_link(source, target) + + monkeypatch.setattr(workspace.os, "link", race) + with pytest.raises(workspace.WorkspaceCheckpointError): + saved(repo) + assert destination.read_bytes() == b"other writer" + + +class TestRestoreBoundaries: + def test_partial_publish_is_removed_after_a_move_failure(self, repo, monkeypatch): + archive, receipt = saved(repo) + remove_original(repo) + real_rename = os.rename + moved = [] + + def fail_second_move(source, target): + if Path(target).parent == repo: + if moved: + raise OSError("injected publish failure") + moved.append(Path(target)) + return real_rename(source, target) + + monkeypatch.setattr(workspace.os, "rename", fail_second_move) + with pytest.raises(workspace.WorkspaceCheckpointError) as error: + workspace.restore_workspace(archive, repo, IDENTITY, expected_sha256=receipt.sha256) + assert error.value.code == "restore_failed" + assert isinstance(error.value.__cause__, OSError) + assert "injected publish failure" in str(error.value.__cause__) + assert len(moved) == 1 + assert not repo.exists() + assert not list(repo.parent.glob(".workspace-restore-*")) + assert archive.exists() + + def test_existing_workspace_is_never_cleared(self, repo): + archive, receipt = saved(repo) + before = inventory(repo) + with pytest.raises(workspace.WorkspaceCheckpointError) as error: + workspace.restore_workspace(archive, repo, IDENTITY, expected_sha256=receipt.sha256) + assert error.value.code == "destination_exists" + assert inventory(repo) == before + + def test_receipt_checksum_is_required(self, repo): + archive, _ = saved(repo) + remove_original(repo) + with pytest.raises(workspace.WorkspaceCheckpointError) as error: + workspace.restore_workspace(archive, repo, IDENTITY, expected_sha256="0" * 64) + assert error.value.code == "checksum_mismatch" + assert not repo.exists() + + @pytest.mark.parametrize("field", ["task_id", "attempt_id", "request_id", "user_id", "repo"]) + def test_cross_identity_restore_is_rejected(self, repo, field): + archive, receipt = saved(repo) + remove_original(repo) + wrong = replace(IDENTITY, **{field: "other/repo" if field == "repo" else "other"}) + with pytest.raises(workspace.WorkspaceCheckpointError): + workspace.restore_workspace(archive, repo, wrong, expected_sha256=receipt.sha256) + assert not repo.exists() + + def test_workspace_path_must_remain_stable(self, repo): + archive, receipt = saved(repo) + target = repo.parent / "different-workspace" + with pytest.raises(workspace.WorkspaceCheckpointError): + workspace.restore_workspace(archive, target, IDENTITY, expected_sha256=receipt.sha256) + assert not target.exists() + + @pytest.mark.parametrize( + "change", + [ + lambda m: m.update(version=True), + lambda m: m.update(environment={"SYNTHETIC": "not allowed"}), + lambda m: m["git"].update(index_sha256="0" * 64), + lambda m: m["git"]["refs"].update({"refs/heads/../escape": "a" * 40}), + lambda m: m["git"].update(branch="different"), + lambda m: m["git"].update(exclude_b64="not valid base64"), + lambda m: m["files"][0].update(path="../escape"), + ], + ) + def test_invalid_manifest_does_not_publish_a_workspace(self, repo, change): + archive, _ = saved(repo) + remove_original(repo) + digest = rewrite(archive, mutate_manifest=change) + with pytest.raises(workspace.WorkspaceCheckpointError): + workspace.restore_workspace(archive, repo, IDENTITY, expected_sha256=digest) + assert not repo.exists() + assert not list(repo.parent.glob(".workspace-restore-*")) + + def test_changed_file_bytes_fail_inner_checksum_even_with_valid_outer_receipt(self, repo): + archive, _ = saved(repo) + remove_original(repo) + + def corrupt(member, data): + if member.name == "files/tracked.txt": + data = b"x" * len(data) + return member, data + + digest = rewrite(archive, mutate_member=corrupt) + with pytest.raises(workspace.WorkspaceCheckpointError) as error: + workspace.restore_workspace(archive, repo, IDENTITY, expected_sha256=digest) + assert error.value.code == "checksum_mismatch" + assert not repo.exists() + + @pytest.mark.parametrize( + "name,kind", + [ + ("files/../escape", tarfile.REGTYPE), + ("/absolute", tarfile.REGTYPE), + ("files/.git/config", tarfile.REGTYPE), + ("files/unlisted", tarfile.REGTYPE), + ("files/empty/", tarfile.DIRTYPE), + ("files/link", tarfile.LNKTYPE), + ], + ) + def test_extra_traversal_duplicate_and_hardlink_members_are_rejected(self, repo, name, kind): + archive, _ = saved(repo) + remove_original(repo) + extra = tarfile.TarInfo(name) + extra.type = kind + extra.linkname = "files/tracked.txt" if kind == tarfile.LNKTYPE else "" + digest = rewrite(archive, extras=[(extra, b"" if kind == tarfile.REGTYPE else None)]) + with pytest.raises(workspace.WorkspaceCheckpointError): + workspace.restore_workspace(archive, repo, IDENTITY, expected_sha256=digest) + assert not repo.exists() + assert not (repo.parent / "escape").exists() + + def test_archive_cannot_write_through_a_symlink(self, repo): + archive, _ = saved(repo) + remove_original(repo) + extra = tarfile.TarInfo("files/relative-link/escape") + extra.mode = 0o644 + extra.size = 1 + + def add_child(manifest): + manifest["files"].append( + { + "path": "relative-link/escape", + "kind": "file", + "mode": 0o644, + "size": 1, + "sha256": hashlib.sha256(b"x").hexdigest(), + } + ) + + digest = rewrite(archive, mutate_manifest=add_child, extras=[(extra, b"x")]) + with pytest.raises(workspace.WorkspaceCheckpointError) as error: + workspace.restore_workspace(archive, repo, IDENTITY, expected_sha256=digest) + assert error.value.code == "invalid_archive" + assert not repo.exists() + + def test_archive_change_during_restore_does_not_publish(self, repo, monkeypatch): + archive, receipt = saved(repo) + remove_original(repo) + original_git = workspace._git + + def change(root, args, *rest, **kwargs): + if args[0] == "init": + with archive.open("ab") as stream: + stream.write(b"changed") + return original_git(root, args, *rest, **kwargs) + + monkeypatch.setattr(workspace, "_git", change) + with pytest.raises(workspace.WorkspaceCheckpointError) as error: + workspace.restore_workspace(archive, repo, IDENTITY, expected_sha256=receipt.sha256) + assert error.value.code == "workspace_changed" + assert not repo.exists() diff --git a/agent/tests/test_hooks.py b/agent/tests/test_hooks.py index f78f0588e..3ebd3fa9d 100644 --- a/agent/tests/test_hooks.py +++ b/agent/tests/test_hooks.py @@ -1,6 +1,7 @@ """Unit tests for hooks.py — Cedar policy SDK hook callbacks.""" import asyncio +from datetime import UTC, datetime from unittest.mock import MagicMock, patch import pytest @@ -10,6 +11,7 @@ from hooks import ( _is_self_reclone, _reset_blocker_reason_for_tests, + _sha256_tool_input_for_row, _stuck_guard_between_turns_hook, build_hook_matchers, detect_egress_denial, @@ -19,7 +21,7 @@ pre_tool_use_hook, reset_stuck_summary, ) -from policy import PolicyEngine +from policy import Outcome, PolicyEngine @pytest.fixture(autouse=True) @@ -725,6 +727,31 @@ def test_post_hook_matcher_structure(self): assert post_matcher.matcher is None assert len(post_matcher.hooks) == 1 + @pytest.mark.parametrize("microvm", [False, True]) + def test_callback_budget_outlives_approval_without_changing_its_deadline(self, microvm): + from hooks import _ApprovalDeadline + from microvm_lifecycle import register_task, unregister_task + from shared_constants import SHARED_CONSTANTS + + context = register_task("callback-budget", "owned-vm") if microvm else None + engine = PolicyEngine(task_type="new_task", repo="owner/repo", task_default_timeout_s=30) + original = _ApprovalDeadline.from_recorded("2026-09-15T23:00:00Z", 30) + try: + matchers = build_hook_matchers(engine=engine, task_id="callback-budget") + protected_window = ( + SHARED_CONSTANTS["microvm_lifecycle"]["maximum_duration_seconds"] + if microvm + else SHARED_CONSTANTS["approval_timeout_s"]["max"] + ) + assert matchers["PreToolUse"][0].timeout > protected_window + assert engine.task_default_timeout_s == 30 + assert original.wall_deadline == 1789513230 + assert matchers["PostToolUse"][0].timeout is None + assert matchers["Stop"][0].timeout is None + finally: + if context: + unregister_task(context) + def test_matchers_with_trajectory(self): engine = PolicyEngine(task_type="new_task", repo="owner/repo") # Pass None for trajectory — should still work @@ -755,7 +782,6 @@ def test_matchers_with_trajectory(self): import hashlib import json as _json from collections import deque -from datetime import UTC from typing import Any import hooks @@ -950,6 +976,152 @@ def _hook_input(tool_name: str = "Bash", command: str = "echo foo") -> dict: } +@pytest.fixture() +def restored_approval_runtime(): + from unittest.mock import MagicMock + + from continuation_runtime import ContinuationRuntime + from continuation_storage import ContinuationContext + from models import RepoSetup + + runtime = ContinuationRuntime( + ContinuationContext( + RepoSetup(repo_dir="/workspace/task", branch="main", build_before=False), + "Original request", + "System prompt", + "coding/new-task-v1", + "1", + ), + MagicMock(), + restored={ + "action": { + "tool_name": "Bash", + "tool_input_sha256": _sha256_tool_input_for_row({"command": "echo foo"}), + } + }, + ) + runtime.human_decision = { + "request_id": "original-request", + "status": "APPROVED", + "scope": "this_call", + "decided_at": "2026-09-17T17:00:00Z", + "created_at": "2026-09-16T17:00:00Z", + "matching_rule_ids": ["test_bash_foo"], + } + return runtime + + +class TestRestoredApproval: + def test_exact_saved_action_is_approved_once_without_another_request( + self, restored_approval_runtime, engine_with_soft_gate, fake_task_state, progress + ): + from continuation_runtime import bind_runtime + + async def call(command="echo foo"): + return await pre_tool_use_hook( + _hook_input(command=command), + "new-tool-id", + {}, + engine=engine_with_soft_gate, + progress=progress, + task_state_module=fake_task_state, + ) + + with bind_runtime(restored_approval_runtime): + # A different proposal does not spend the single-action grant. + changed = asyncio.run(call("echo other-foo")) + assert changed["hookSpecificOutput"]["permissionDecision"] == "deny" + approved = asyncio.run(call()) + assert approved["hookSpecificOutput"]["permissionDecision"] == "allow" + repeated = asyncio.run(call()) + assert repeated["hookSpecificOutput"]["permissionDecision"] == "deny" + assert fake_task_state.write_calls == [] + assert fake_task_state.gate_count_calls == [] + assert engine_with_soft_gate.approval_gate_count == 0 + grants = [kwargs for name, kwargs in progress.calls if name == "write_approval_granted"] + assert len(grants) == 1 + assert grants[0]["request_id"] == "original-request" + assert grants[0]["decided_at"] == "2026-09-17T17:00:00Z" + + def test_hard_denial_wins_without_consuming_saved_approval( + self, restored_approval_runtime, fake_task_state + ): + from continuation_runtime import bind_runtime + + engine = PolicyEngine( + task_type="new_task", + repo="owner/repo", + blueprint_hard_policies=( + '@tier("hard") @rule_id("never_foo") ' + 'forbid (principal, action == Agent::Action::"execute_bash", resource) ' + 'when { context.command like "*foo*" };' + ), + ) + with bind_runtime(restored_approval_runtime): + result = asyncio.run( + pre_tool_use_hook( + _hook_input(), + "new-tool-id", + {}, + engine=engine, + task_state_module=fake_task_state, + ) + ) + assert result["hookSpecificOutput"]["permissionDecision"] == "deny" + assert ( + restored_approval_runtime.consume_approved_action( + "Bash", _sha256_tool_input_for_row({"command": "echo foo"}) + ) + is not None + ) + + @pytest.mark.parametrize("status", ["DENIED", "TIMED_OUT"]) + def test_recorded_denial_seeds_normal_policy_cache( + self, restored_approval_runtime, engine_with_soft_gate, status + ): + restored_approval_runtime.human_decision.update( + status=status, deny_reason="Use another approach" + ) + restored_approval_runtime.seed_policy(engine_with_soft_gate) + decision = engine_with_soft_gate.evaluate_tool_use("Bash", {"command": "echo foo"}) + assert decision.outcome == Outcome.DENY + assert decision.cache_hit_metadata["original_decision_ts"] == "2026-09-17T17:00:00Z" + # Explicit denials also cover cosmetic changes under the same rule. + changed = engine_with_soft_gate.evaluate_tool_use("Bash", {"command": "echo other-foo"}) + assert changed.outcome == (Outcome.DENY if status == "DENIED" else Outcome.REQUIRE_APPROVAL) + + def test_grants_survive_a_replacement(self, restored_approval_runtime, engine_with_soft_gate): + from dataclasses import replace + + from policy import ApprovalAllowlist + + scopes = [ + "tool_type: Read", + "tool_group:file_write", + "rule:other_rule", + "bash_pattern:git status*", + "write_path:/workspace/docs/*", + ] + original = ApprovalAllowlist(scopes) + restored_approval_runtime.context = replace( + restored_approval_runtime.context, approval_scopes=original.snapshot_scopes() + ) + restored_approval_runtime.human_decision["scope"] = "rule:test_bash_foo" + restored_approval_runtime.seed_policy(engine_with_soft_gate) + for tool, tool_input in [ + ("Read", {}), + ("Write", {"file_path": "/workspace/a"}), + ("Bash", {"command": "git status --short"}), + ("Bash", {"command": "echo foo"}), + ]: + assert ( + engine_with_soft_gate.evaluate_tool_use(tool, tool_input).outcome == Outcome.ALLOW + ) + assert engine_with_soft_gate.allowlist.rule_ids == {"other_rule", "test_bash_foo"} + original.add("all_session") + assert ApprovalAllowlist(list(original.snapshot_scopes())).matches("AnyTool", {}) + + def _prime_approval(fake: _FakeTaskState, terminal_row: dict) -> None: """Queue an ``APPROVED``/``DENIED`` row on the second poll iteration. @@ -971,10 +1143,9 @@ def _fast_poll(monkeypatch): """Collapse poll intervals so tests run instantly. Swaps ``asyncio.sleep`` for a no-op AND advances ``hooks.time.monotonic`` - by the requested sleep duration each call so the poll's wall-clock - deadline actually trips. Without the monotonic advance the poll spins - forever when the script runs out of rows (deque empty → default - PENDING row → never terminal). + by the requested sleep duration each call so the monotonic deadline + trips without waiting for real UTC time to pass. When scripted rows + run out, the fake returns PENDING until that deadline. """ fake_clock = {"now": 0.0} @@ -988,10 +1159,329 @@ async def _zero_sleep(seconds): monkeypatch.setattr(hooks.asyncio, "sleep", _zero_sleep) +@pytest.fixture() +def approval_clock(monkeypatch): + """Control elapsed and UTC time independently, including a frozen guest clock.""" + clock = {"wall": 1_800_000_000.0, "monotonic": 100.0} + monkeypatch.setattr(hooks.time, "time", lambda: clock["wall"]) + monkeypatch.setattr(hooks.time, "monotonic", lambda: clock["monotonic"]) + monkeypatch.setattr( + hooks, + "_iso_now", + lambda: datetime.fromtimestamp(clock["wall"], UTC).strftime("%Y-%m-%dT%H:%M:%SZ"), + ) + return clock + + +class TestApprovalDeadline: + def test_frozen_monotonic_clock_still_expires_after_sleep( + self, fake_task_state, progress, engine_with_soft_gate, monkeypatch, approval_clock + ): + engine_with_soft_gate._task_default_timeout_s = 30 + sleeps = [] + + async def frozen_sleep(seconds): + sleeps.append(seconds) + assert len(sleeps) == 1, "The expired gate must not restart polling after waking" + approval_clock["wall"] += 600 + + monkeypatch.setattr(hooks.asyncio, "sleep", frozen_sleep) + result = _run( + pre_tool_use_hook( + _hook_input(), + "tu-1", + {}, + engine=engine_with_soft_gate, + task_id="01KTASK", + progress=progress, + task_state_module=fake_task_state, + ) + ) + + assert result["hookSpecificOutput"]["permissionDecision"] == "deny" + assert fake_task_state.update_calls[-1][2] == "TIMED_OUT" + assert len(fake_task_state.get_calls) == 1 + + @pytest.mark.parametrize("wall_change", [12, -120]) + def test_database_write_time_counts_even_if_wall_clock_moves_back( + self, + fake_task_state, + progress, + engine_with_soft_gate, + monkeypatch, + approval_clock, + wall_change, + ): + engine_with_soft_gate._task_default_timeout_s = 30 + original_write = fake_task_state.transact_write_approval_request + sleeps = [] + + def delayed_write(*args, **kwargs): + original_write(*args, **kwargs) + approval_clock["wall"] += wall_change + approval_clock["monotonic"] += 12 + + async def advance(seconds): + sleeps.append(seconds) + approval_clock["wall"] += seconds + approval_clock["monotonic"] += seconds + + monkeypatch.setattr(fake_task_state, "transact_write_approval_request", delayed_write) + monkeypatch.setattr(hooks.asyncio, "sleep", advance) + result = _run( + pre_tool_use_hook( + _hook_input(), + "tu-1", + {}, + engine=engine_with_soft_gate, + task_id="01KTASK", + progress=progress, + task_state_module=fake_task_state, + ) + ) + + assert result["hookSpecificOutput"]["permissionDecision"] == "deny" + assert sum(sleeps) == 18 + assert approval_clock["monotonic"] == 130 + + @pytest.mark.parametrize("wall_jump,expected_wait", [(27, 3), (-3600, 30)]) + def test_clock_changes_clamp_next_sleep_and_never_extend_window( + self, + fake_task_state, + progress, + engine_with_soft_gate, + monkeypatch, + approval_clock, + wall_jump, + expected_wait, + ): + engine_with_soft_gate._task_default_timeout_s = 30 + sleeps = [] + + async def advance(seconds): + sleeps.append(seconds) + approval_clock["monotonic"] += seconds + approval_clock["wall"] += seconds + (wall_jump if len(sleeps) == 1 else 0) + + monkeypatch.setattr(hooks.asyncio, "sleep", advance) + result = _run( + pre_tool_use_hook( + _hook_input(), + "tu-1", + {}, + engine=engine_with_soft_gate, + task_id="01KTASK", + progress=progress, + task_state_module=fake_task_state, + ) + ) + + assert result["hookSpecificOutput"]["permissionDecision"] == "deny" + assert sum(sleeps) == expected_wait + if wall_jump > 0: + assert sleeps == [2, 1] + + @pytest.mark.parametrize( + "reread_row,cancelled,expected", + [ + ({"status": "APPROVED", "scope": "this_call"}, False, "allow"), + ({"status": "DENIED", "deny_reason": "no"}, False, "deny"), + ({"status": "PENDING"}, False, "deny"), + (None, False, "deny"), + ({"status": "APPROVED", "scope": "this_call"}, True, "deny"), + ], + ) + def test_waking_after_deadline_preserves_decision_race_and_cancellation( + self, + fake_task_state, + progress, + engine_with_soft_gate, + monkeypatch, + approval_clock, + reread_row, + cancelled, + expected, + ): + engine_with_soft_gate._task_default_timeout_s = 30 + fake_task_state.best_effort_return = False + fake_task_state.reread_row = reread_row + if reread_row is None: + # A missing/TTL-reaped row during the poll also stays fail-closed. + fake_task_state.get_row_script.append(None) + if cancelled: + fake_task_state.resume_raises = _FakeApprovalResumeError("cancelled") + + async def frozen_sleep(_seconds): + approval_clock["wall"] += 600 + + monkeypatch.setattr(hooks.asyncio, "sleep", frozen_sleep) + result = _run( + pre_tool_use_hook( + _hook_input(), + "tu-1", + {}, + engine=engine_with_soft_gate, + task_id="01KTASK", + progress=progress, + task_state_module=fake_task_state, + ) + ) + + assert result["hookSpecificOutput"]["permissionDecision"] == expected + assert fake_task_state.update_calls[-1][2] == "TIMED_OUT" + assert len(fake_task_state.get_calls) == 2 + assert all(consistent for _, _, consistent in fake_task_state.get_calls) + if cancelled: + assert "write_approval_granted" not in progress.milestones() + + def test_decision_committed_during_slow_read_is_honored( + self, fake_task_state, progress, engine_with_soft_gate, monkeypatch, approval_clock + ): + engine_with_soft_gate._task_default_timeout_s = 30 + + def slow_read(*_args, **kwargs): + assert kwargs["consistent_read"] + approval_clock["wall"] += 600 + return {"status": "APPROVED", "scope": "this_call"} + + monkeypatch.setattr(fake_task_state, "get_approval_row", slow_read) + result = _run( + pre_tool_use_hook( + _hook_input(), + "tu-1", + {}, + engine=engine_with_soft_gate, + task_id="01KTASK", + progress=progress, + task_state_module=fake_task_state, + ) + ) + + assert result["hookSpecificOutput"]["permissionDecision"] == "allow" + assert not fake_task_state.update_calls + + def test_coroutine_cancellation_does_not_become_a_timeout_or_allow( + self, fake_task_state, progress, engine_with_soft_gate, monkeypatch, approval_clock + ): + async def cancelled_sleep(_seconds): + raise asyncio.CancelledError + + monkeypatch.setattr(hooks.asyncio, "sleep", cancelled_sleep) + with pytest.raises(asyncio.CancelledError): + _run( + pre_tool_use_hook( + _hook_input(), + "tu-1", + {}, + engine=engine_with_soft_gate, + task_id="01KTASK", + progress=progress, + task_state_module=fake_task_state, + ) + ) + assert not fake_task_state.update_calls + assert not fake_task_state.resume_calls + + # --- Happy paths ---------------------------------------------------------- class TestApprovedPath: + @pytest.mark.parametrize("transferred", [False, True]) + @pytest.mark.parametrize("publish_failed", [False, True]) + def test_checkpoint_holds_tools_until_same_worker_claims_decision( + self, + fake_task_state, + progress, + engine_with_soft_gate, + monkeypatch, + transferred, + publish_failed, + ): + from types import SimpleNamespace + from typing import Any + + from continuation_runtime import bind_runtime + from continuation_session import CheckpointIdentity + from microvm_lifecycle import register_task, unregister_task + + _fast_poll(monkeypatch) + _prime_approval(fake_task_state, {"status": "APPROVED", "scope": "this_call"}) + lifecycle = register_task("checkpoint-hook-task", "microvm-hook") + phases = [] + original_resume = fake_task_state.transact_resume_from_approval + + async def capture(context, park, **kwargs): + assert kwargs["session_id"] == "sdk-session" + async with context.continuation_checkpoint(park): + phases.append(context.diagnostic_snapshot()["phase"]) + return CheckpointIdentity( + lifecycle.task_id, lifecycle.microvm_id, park.request_id, "user", "owner/repo" + ), "owned-receipt" + + def publish(*args, **kwargs): + phases.append(lifecycle.diagnostic_snapshot()["phase"]) + if publish_failed: + raise TimeoutError("publication acknowledgement lost") + + def resume(*args, **kwargs): + phases.append(lifecycle.diagnostic_snapshot()["phase"]) + if transferred: + raise _FakeApprovalResumeError("coordinator now owns the checkpoint") + return original_resume(*args, **kwargs) + + runtime: Any = SimpleNamespace( + capture=capture, + consume_approved_action=lambda *_: None, + context=SimpleNamespace(cost_usd=0.02, turns_used=1), + ) + monkeypatch.setattr( + fake_task_state, "publish_continuation_checkpoint", publish, raising=False + ) + monkeypatch.setattr(fake_task_state, "transact_resume_from_approval", resume) + monkeypatch.setattr(hooks, "task_state", fake_task_state) + matchers = build_hook_matchers( + engine=engine_with_soft_gate, + task_id=lifecycle.task_id, + user_id="user", + progress=progress, + ) + try: + with bind_runtime(runtime): + result = _run( + matchers["PreToolUse"][0].hooks[0]( + {**_hook_input(), "session_id": "sdk-session"}, + "tu-1", + {}, + ) + ) + assert phases == ["checkpointing", "checkpoint-ready", "checkpoint-ready"] + unavailable = [ + data + for method, data in progress.calls + if method == "write_agent_milestone" + and data.get("milestone") == "continuation_unavailable" + ] + assert len(unavailable) == int(publish_failed) + if publish_failed: + # Use the real signature so an incorrectly named keyword cannot + # pass through the permissive recording double unnoticed. + from progress_writer import _ProgressWriter + + writer = MagicMock(spec=_ProgressWriter) + import inspect + + inspect.signature(_ProgressWriter.write_agent_milestone).bind( + writer, **unavailable[0] + ) + expected = "deny" if transferred else "allow" + assert result["hookSpecificOutput"]["permissionDecision"] == expected + assert lifecycle.diagnostic_snapshot()["phase"] == ( + "closed" if transferred else "active" + ) + finally: + unregister_task(lifecycle) + def test_approved_returns_allow_and_propagates_scope( self, fake_task_state, progress, engine_with_soft_gate, monkeypatch ): @@ -1084,10 +1574,48 @@ def test_denied_returns_deny_queues_injection_and_caches( # Recent-decision cache populated — identical next call auto-denies. follow_up = engine_with_soft_gate.evaluate_tool_use("Bash", {"command": "echo foo"}) assert follow_up.outcome.value == "deny" - assert "Recent DENIED" in follow_up.reason + assert "Recorded DENIED at t1" in follow_up.reason assert "write_approval_denied" in progress.milestones() +class TestCancelledPath: + @pytest.mark.parametrize("during_timeout", [False, True]) + def test_cancellation_closes_wait_without_resuming_or_caching_a_human_denial( + self, fake_task_state, progress, engine_with_soft_gate, monkeypatch, during_timeout + ): + _fast_poll(monkeypatch) + cancelled = {"status": "CANCELLED", "cancellation_reason": "Task cancelled by its owner"} + if during_timeout: + fake_task_state.best_effort_return = False + fake_task_state.reread_row = cancelled + else: + _prime_approval(fake_task_state, cancelled) + + result = _run( + pre_tool_use_hook( + _hook_input(), + "tu-1", + {}, + engine=engine_with_soft_gate, + task_id="01KTASK", + user_id="u-1", + progress=progress, + task_state_module=fake_task_state, + ) + ) + + assert result["hookSpecificOutput"]["permissionDecision"] == "deny" + assert "cancelled" in result["hookSpecificOutput"]["permissionDecisionReason"] + assert not fake_task_state.resume_calls + if not during_timeout: + assert not fake_task_state.update_calls + assert not engine_with_soft_gate.drain_denial_injections() + follow_up = engine_with_soft_gate.evaluate_tool_use("Bash", {"command": "echo foo"}) + assert "Recent DENIED" not in follow_up.reason + assert "write_approval_denied" not in progress.milestones() + assert "write_approval_timed_out" not in progress.milestones() + + class TestPersistentGateCount: """Chunk 7 (§13.6): REQUIRE_APPROVAL path must bump BOTH the session counter and the TaskTable-persisted counter so a container restart @@ -1617,6 +2145,7 @@ def test_rule_annotation_clip_emits_milestone(self, fake_task_state, progress, m task_type="new_task", repo="owner/repo", blueprint_soft_policies=blueprint_soft, + task_default_timeout_s=300, ) _prime_approval( fake_task_state, @@ -2013,3 +2542,66 @@ def test_build_hook_matchers_creates_a_guard_without_crashing(self): engine = PolicyEngine(task_type="new_task", repo="owner/repo") matchers = build_hook_matchers(engine, task_id="t") assert "PostToolUse" in matchers and "Stop" in matchers + + +@pytest.mark.parametrize("cancelled", [False, True]) +def test_microvm_approval_cannot_return_until_resume_barrier_opens( + monkeypatch, engine_with_soft_gate, fake_task_state, progress, cancelled +): + import threading + + from microvm_lifecycle import register_task, unregister_task + + context = register_task("lifecycle-hook-task", "microvm-hook") + entered = threading.Event() + answer = threading.Event() + original_read = fake_task_state.get_approval_row + + def read(*args, **kwargs): + entered.set() + return {"status": "APPROVED", "scope": "once"} if answer.is_set() else {"status": "PENDING"} + + fake_task_state.get_approval_row = read + monkeypatch.setattr(hooks, "task_state", fake_task_state) + monkeypatch.setattr(hooks, "POLL_FAST_INTERVAL_S", 0.01) + if cancelled: + fake_task_state.resume_raises = _FakeApprovalResumeError("cancel won") + matchers = build_hook_matchers( + engine=engine_with_soft_gate, + task_id=context.task_id, + progress=progress, + ) + + async def scenario(): + pending = asyncio.create_task(matchers["PreToolUse"][0].hooks[0](_hook_input(), "tu-1", {})) + try: + assert await asyncio.to_thread(entered.wait, 1) + checkpoints = [] + await context.suspend(checkpoints.append, budget_s=1) + park = checkpoints[0] + assert park.request_id == fake_task_state.write_calls[0][1] + assert park.deadline.remaining_s() > 0 + answer.set() + await asyncio.sleep(0.04) + assert not pending.done() + assert fake_task_state.resume_calls == [] + refreshed = [] + await context.resume(refreshed.append, budget_s=1) + assert refreshed == [park] + result = await asyncio.wait_for(pending, 1) + expected = "deny" if cancelled else "allow" + assert result["hookSpecificOutput"]["permissionDecision"] == expected + if not cancelled: + await matchers["PostToolUse"][0].hooks[0]( + {"tool_name": "Bash", "tool_response": "ok"}, "tu-1", {} + ) + finally: + answer.set() + context.close() + await asyncio.gather(pending, return_exceptions=True) + + try: + _run(scenario()) + finally: + fake_task_state.get_approval_row = original_read + unregister_task(context) diff --git a/agent/tests/test_microvm_checkpoint.py b/agent/tests/test_microvm_checkpoint.py new file mode 100644 index 000000000..12c0e7d4d --- /dev/null +++ b/agent/tests/test_microvm_checkpoint.py @@ -0,0 +1,471 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Guest durable-state checks; optional DynamoDB Local cases execute conditions.""" + +from __future__ import annotations + +import asyncio +import copy +import os +import uuid +from dataclasses import dataclass, replace +from datetime import UTC, datetime +from unittest.mock import MagicMock +from urllib.parse import urlsplit + +import pytest +from boto3.dynamodb.types import TypeDeserializer, TypeSerializer +from botocore.exceptions import ClientError + +import microvm_checkpoint as checkpoint +from microvm_lifecycle import ApprovalPark, ApprovalRecord, LifecycleUnavailable +from progress_writer import _reset_circuit_breakers + + +@dataclass +class Deadline: + remaining: float = 60 + + def remaining_s(self) -> float: + return self.remaining + + +def make_park() -> ApprovalPark: + return ApprovalPark( + task_id="task", + microvm_id="vm", + request_id="request", + tool_use_id="tool", + deadline=Deadline(), + record=ApprovalRecord( + "user", "owner/repo", datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%SZ"), 60 + ), + ) + + +def make_rows(park: ApprovalPark) -> tuple[dict, dict]: + record, deadline_ms = checkpoint._record(park) + task = { + "task_id": park.task_id, + "user_id": record.user_id, + "repo": record.repo, + "status": "AWAITING_APPROVAL", + "compute_type": "lambda-microvm", + "session_id": park.microvm_id, + "compute_metadata": {"microvmId": park.microvm_id, "endpoint": "https://synthetic.invalid"}, + "awaiting_approval_request_id": park.request_id, + "microvm_lifecycle": { + "version": 1, + "generation": "suspend-generation", + "microvm_id": park.microvm_id, + "request_id": park.request_id, + "action": "suspend", + "requested_at_ms": int(datetime.now(UTC).timestamp() * 1000), + "deadline_ms": deadline_ms, + }, + } + approval = { + "task_id": park.task_id, + "request_id": park.request_id, + "user_id": record.user_id, + "repo": record.repo, + "created_at": record.created_at, + "timeout_s": record.timeout_s, + "status": "PENDING", + } + return task, approval + + +def serialize(item: dict) -> dict: + return {key: TypeSerializer().serialize(value) for key, value in item.items()} + + +@pytest.fixture(autouse=True) +def _progress(): + _reset_circuit_breakers() + yield + _reset_circuit_breakers() + + +@pytest.fixture +def state(monkeypatch): + monkeypatch.setenv("TASK_TABLE_NAME", "tasks") + monkeypatch.setenv("TASK_APPROVALS_TABLE_NAME", "approvals") + monkeypatch.setenv("TASK_EVENTS_TABLE_NAME", "events") + park = make_park() + task, approval = make_rows(park) + client = MagicMock() + client.get_item.side_effect = lambda **kw: { + "Item": serialize(task if kw["TableName"] == "tasks" else approval) + } + monkeypatch.setattr("aws_session.tenant_client", lambda *_a, **_kw: client) + monkeypatch.setattr("aws_session.refresh_microvm_credentials", MagicMock()) + return park, task, approval, client + + +class TestCheckpoint: + @pytest.mark.parametrize("repo", ["owner/repo", ""]) + def test_retained_request_can_suspend_and_resume_with_explicit_null_deadline(self, state, repo): + from hooks import _ApprovalDeadline + + original, task, approval, client = state + assert original.record is not None + record = replace(original.record, timeout_s=0, repo=repo) + park = replace( + original, + record=record, + deadline=_ApprovalDeadline.from_recorded(record.created_at, 0), + ) + new_task, new_approval = make_rows(park) + task.update(new_task) + if not repo: + task.pop("repo") + approval.update(new_approval) + checkpoint.checkpoint_before_suspend(park) + operations = client.transact_write_items.call_args.kwargs["TransactItems"] + marker = { + key: TypeDeserializer().deserialize(value) + for key, value in operations[2]["Put"]["Item"].items() + } + assert marker["metadata"]["approval_deadline_ms"] is None + assert task["microvm_lifecycle"]["deadline_ms"] is None + task["microvm_lifecycle"].update(action="resume", generation="resume-generation") + approval["status"] = "APPROVED" + checkpoint.refresh_and_reconcile_after_resume(park) + assert park.deadline.remaining_s() == float("inf") + assert approval["status"] == "APPROVED" + assert len(client.transact_write_items.call_args.kwargs["TransactItems"]) == 2 + + @pytest.mark.parametrize("deadline_value", [0, True, "null", "missing"]) + def test_retained_request_rejects_non_null_or_missing_intent_deadline( + self, state, deadline_value + ): + original, task, approval, client = state + assert original.record is not None + park = replace(original, record=replace(original.record, timeout_s=0)) + new_task, new_approval = make_rows(park) + task.update(new_task) + approval.update(new_approval) + if deadline_value == "missing": + task["microvm_lifecycle"].pop("deadline_ms") + else: + task["microvm_lifecycle"]["deadline_ms"] = deadline_value + with pytest.raises(LifecycleUnavailable, match="intent"): + checkpoint.checkpoint_before_suspend(park) + client.transact_write_items.assert_not_called() + + def test_checkpoint_is_an_acknowledged_cross_table_transaction(self, state): + park, task, approval, client = state + checkpoint.checkpoint_before_suspend(park) + assert all(call.kwargs["ConsistentRead"] for call in client.get_item.call_args_list) + operations = client.transact_write_items.call_args.kwargs["TransactItems"] + assert len(operations) == 3 + assert [op["ConditionCheck"]["TableName"] for op in operations[:2]] == [ + "tasks", + "approvals", + ] + assert operations[0]["ConditionCheck"]["ExpressionAttributeValues"][":intent"] == ( + TypeSerializer().serialize(task["microvm_lifecycle"]) + ) + assert ":pending" in operations[1]["ConditionCheck"]["ConditionExpression"] + put = operations[2]["Put"] + assert put["TableName"] == "events" + item = {key: TypeDeserializer().deserialize(value) for key, value in put["Item"].items()} + assert item["task_id"] == park.task_id + assert item["metadata"]["milestone"] == "microvm_suspend_checkpoint" + assert item["metadata"]["generation"] == task["microvm_lifecycle"]["generation"] + assert item["metadata"]["request_id"] == approval["request_id"] + assert "ClientRequestToken" not in client.transact_write_items.call_args.kwargs + assert task["status"] == "AWAITING_APPROVAL" + assert approval["status"] == "PENDING" + + @pytest.mark.parametrize( + ("row", "field", "value"), + [ + ("task", "task_id", "other"), + ("task", "user_id", "other"), + ("task", "repo", "other/repo"), + ("task", "status", "CANCELLED"), + ("task", "compute_type", "agentcore"), + ("task", "session_id", "other-vm"), + ("task", "compute_metadata", {"microvmId": "other-vm"}), + ("task", "awaiting_approval_request_id", "other-request"), + ("intent", "version", 2), + ("intent", "version", True), + ("intent", "generation", ""), + ("intent", "microvm_id", "other-vm"), + ("intent", "request_id", "other-request"), + ("intent", "action", "resume"), + ("intent", "deadline_ms", 1), + ("intent", "deadline_ms", None), + ("intent", "requested_at_ms", -1), + ("approval", "task_id", "other"), + ("approval", "request_id", "other-request"), + ("approval", "user_id", "other"), + ("approval", "repo", "other/repo"), + ("approval", "created_at", "2020-01-01T00:00:00Z"), + ("approval", "timeout_s", 61), + ("approval", "timeout_s", True), + ("approval", "status", "APPROVED"), + ], + ) + def test_changed_identity_or_deadline_never_writes(self, state, row, field, value): + park, task, approval, client = state + target = {"task": task, "approval": approval, "intent": task["microvm_lifecycle"]}[row] + target[field] = value + with pytest.raises(LifecycleUnavailable): + checkpoint.checkpoint_before_suspend(park) + client.transact_write_items.assert_not_called() + + @pytest.mark.parametrize("stage", ["task", "approval", "write"]) + def test_missing_rows_and_uncertain_write_do_not_acknowledge(self, state, stage): + park, task, _, client = state + if stage == "write": + client.transact_write_items.side_effect = TimeoutError("lost response") + expected = TimeoutError + else: + client.get_item.side_effect = ( + [{}] if stage == "task" else [{"Item": serialize(task)}, {}] + ) + expected = LifecycleUnavailable + with pytest.raises(expected): + checkpoint.checkpoint_before_suspend(park) + + def test_unconfigured_or_disabled_progress_cannot_checkpoint(self, state, monkeypatch): + from progress_writer import _ProgressWriter + + park, _, _, client = state + monkeypatch.delenv("TASK_EVENTS_TABLE_NAME") + with pytest.raises(RuntimeError, match="progress table"): + checkpoint.checkpoint_before_suspend(park) + monkeypatch.setenv("TASK_EVENTS_TABLE_NAME", "events") + _ProgressWriter(park.task_id)._disabled = True + with pytest.raises(RuntimeError, match="progress table"): + checkpoint.checkpoint_before_suspend(park) + client.transact_write_items.assert_not_called() + + def test_expired_original_deadline_cannot_suspend(self, state): + park, _, _, client = state + assert isinstance(park.deadline, Deadline) + park.deadline.remaining = 0 + with pytest.raises(LifecycleUnavailable, match="deadline elapsed"): + checkpoint.checkpoint_before_suspend(park) + client.transact_write_items.assert_not_called() + + @pytest.mark.parametrize("status", ["PENDING", "APPROVED", "DENIED", "TIMED_OUT", "STRANDED"]) + def test_expired_wake_retains_original_gate_for_existing_decision_loop( + self, state, status, monkeypatch + ): + park, task, approval, client = state + task["microvm_lifecycle"].update(action="resume", generation="resume-generation") + approval["status"] = status + before = copy.deepcopy(approval) + assert isinstance(park.deadline, Deadline) + park.deadline.remaining = 0 + original_deadline = park.deadline + refresh = MagicMock() + monkeypatch.setattr("aws_session.refresh_microvm_credentials", refresh) + client.get_item.side_effect = lambda **kw: ( + {"Item": serialize(task if kw["TableName"] == "tasks" else approval)} + if refresh.called + else pytest.fail("AWS read occurred before credential refresh") + ) + checkpoint.refresh_and_reconcile_after_resume(park) + refresh.assert_called_once_with(park.task_id) + assert park.deadline is original_deadline + assert park.deadline.remaining_s() == 0 + assert approval == before + operations = client.transact_write_items.call_args.kwargs["TransactItems"] + assert len(operations) == 2 + assert all(set(operation) == {"ConditionCheck"} for operation in operations) + assert "IN (:pending, :approved" in operations[1]["ConditionCheck"]["ConditionExpression"] + + def test_failed_refresh_prevents_any_aws_reconciliation(self, state, monkeypatch): + park, _, _, client = state + monkeypatch.setattr( + "aws_session.refresh_microvm_credentials", MagicMock(side_effect=RuntimeError("denied")) + ) + with pytest.raises(RuntimeError, match="denied"): + checkpoint.refresh_and_reconcile_after_resume(park) + client.get_item.assert_not_called() + client.transact_write_items.assert_not_called() + + +LOCAL_ENDPOINT = os.environ.get("ABCA_DDB_LOCAL_ENDPOINT", "") +if os.environ.get("CI") == "true" and not LOCAL_ENDPOINT: + raise RuntimeError( + "CI requires ABCA_DDB_LOCAL_ENDPOINT; checkpoint transaction tests must not skip" + ) + + +@pytest.fixture +def local_tables(monkeypatch): + import boto3 + from botocore.config import Config + + # Never let an accidentally set endpoint create tables in an AWS account. + endpoint = urlsplit(LOCAL_ENDPOINT) + assert endpoint.scheme == "http" and endpoint.hostname == "127.0.0.1" + client = boto3.client( + "dynamodb", + endpoint_url=LOCAL_ENDPOINT, + region_name="us-west-2", + aws_access_key_id="SYNTHETIC", + aws_secret_access_key="synthetic", + config=Config(connect_timeout=1, read_timeout=2, retries={"total_max_attempts": 1}), + ) + names = {} + try: + for kind, sort_key, env in [ + ("tasks", None, "TASK_TABLE_NAME"), + ("approvals", "request_id", "TASK_APPROVALS_TABLE_NAME"), + ("events", "event_id", "TASK_EVENTS_TABLE_NAME"), + ]: + name = f"p3-{kind}-{uuid.uuid4().hex}" + keys = ["task_id", *([sort_key] if sort_key else [])] + client.create_table( + TableName=name, + KeySchema=[ + {"AttributeName": key, "KeyType": "HASH" if index == 0 else "RANGE"} + for index, key in enumerate(keys) + ], + AttributeDefinitions=[{"AttributeName": key, "AttributeType": "S"} for key in keys], + BillingMode="PAY_PER_REQUEST", + ) + names[kind] = name + monkeypatch.setenv(env, name) + monkeypatch.setattr("aws_session.tenant_client", lambda *_a, **_kw: client) + monkeypatch.setattr("aws_session.refresh_microvm_credentials", MagicMock()) + yield client, names + finally: + for name in names.values(): + client.delete_table(TableName=name) + + +@pytest.mark.skipif(not LOCAL_ENDPOINT, reason="opt-in DynamoDB Local condition verification") +class TestDynamoLocalCheckpoint: + @pytest.mark.parametrize("race", [None, "cancel", "approve", "deadline", "intent"]) + def test_real_suspend_transaction_guards_races(self, local_tables, monkeypatch, race): + client, names = local_tables + park = make_park() + task, approval = make_rows(park) + client.put_item(TableName=names["tasks"], Item=serialize(task)) + client.put_item(TableName=names["approvals"], Item=serialize(approval)) + transact = client.transact_write_items + + def race_then_transact(**kwargs): + if race == "cancel": + task["status"] = "CANCELLED" + elif race == "approve": + approval["status"] = "APPROVED" + elif race == "deadline": + approval["timeout_s"] += 1 + elif race == "intent": + task["microvm_lifecycle"].update(action="resume", generation="new-generation") + client.put_item(TableName=names["tasks"], Item=serialize(task)) + client.put_item(TableName=names["approvals"], Item=serialize(approval)) + return transact(**kwargs) + + monkeypatch.setattr(client, "transact_write_items", race_then_transact) + if race: + with pytest.raises(ClientError) as error: + checkpoint.checkpoint_before_suspend(park) + assert error.value.response["Error"]["Code"] == "TransactionCanceledException" + else: + checkpoint.checkpoint_before_suspend(park) + assert client.scan(TableName=names["events"], ConsistentRead=True)["Count"] == ( + 0 if race else 1 + ) + + @pytest.mark.parametrize("race", ["approve", "cancel"]) + def test_real_resume_transaction_accepts_decision_but_rejects_cancellation( + self, local_tables, monkeypatch, race + ): + client, names = local_tables + park = make_park() + task, approval = make_rows(park) + task["microvm_lifecycle"].update(action="resume", generation="wake") + client.put_item(TableName=names["tasks"], Item=serialize(task)) + client.put_item(TableName=names["approvals"], Item=serialize(approval)) + transact = client.transact_write_items + + def race_then_transact(**kwargs): + if race == "approve": + approval["status"] = "APPROVED" + client.put_item(TableName=names["approvals"], Item=serialize(approval)) + else: + task["status"] = "CANCELLED" + client.put_item(TableName=names["tasks"], Item=serialize(task)) + return transact(**kwargs) + + monkeypatch.setattr(client, "transact_write_items", race_then_transact) + if race == "cancel": + with pytest.raises(ClientError) as error: + checkpoint.refresh_and_reconcile_after_resume(park) + assert error.value.response["Error"]["Code"] == "TransactionCanceledException" + else: + checkpoint.refresh_and_reconcile_after_resume(park) + assert client.scan(TableName=names["events"], ConsistentRead=True)["Count"] == 0 + + @pytest.mark.parametrize("wake", ["approved", "expired_pending", "cancelled", "changed_intent"]) + def test_http_hooks_use_real_transactions_and_preserve_original_gate(self, local_tables, wake): + import httpx + + import server + from microvm_lifecycle import register_task, unregister_task + + client, names = local_tables + original = make_park() + task, approval = make_rows(original) + client.put_item(TableName=names["tasks"], Item=serialize(task)) + client.put_item(TableName=names["approvals"], Item=serialize(approval)) + context = register_task(original.task_id, original.microvm_id) + + async def exercise(): + await context.tool_started(original.tool_use_id) + park = context.park_approval( + original.request_id, original.tool_use_id, original.deadline, record=original.record + ) + assert park is not None + async with httpx.AsyncClient( + transport=httpx.ASGITransport(app=server.app), base_url="http://test" + ) as http: + prefix = server.MICROVM_HOOK_PREFIX + for _ in range(2): + assert (await http.post(prefix + "/suspend", json={})).status_code == 200 + assert client.scan(TableName=names["events"], ConsistentRead=True)["Count"] == 1 + task["microvm_lifecycle"].update(action="resume", generation="wake") + if wake == "approved": + approval["status"] = "APPROVED" + elif wake == "expired_pending": + assert isinstance(original.deadline, Deadline) + original.deadline.remaining = 0 + elif wake == "cancelled": + task["status"] = "CANCELLED" + else: + task["microvm_lifecycle"]["request_id"] = "another-gate" + client.put_item(TableName=names["tasks"], Item=serialize(task)) + client.put_item(TableName=names["approvals"], Item=serialize(approval)) + + response = await http.post(prefix + "/resume", json={"microvmId": "vm"}) + if wake in {"cancelled", "changed_intent"}: + assert response.status_code == 409 + with pytest.raises(LifecycleUnavailable): + await context.wait_until_open() + else: + assert response.status_code == 200 + await asyncio.wait_for(context.leave_approval(park), timeout=1) + assert park.deadline is original.deadline + if wake == "expired_pending": + assert park.deadline.remaining_s() == 0 + assert (await http.post(prefix + "/resume", json={})).status_code == 200 + for kind, expected in [("tasks", task), ("approvals", approval)]: + rows = client.scan(TableName=names[kind], ConsistentRead=True)["Items"] + assert rows == [serialize(expected)] + assert client.scan(TableName=names["events"], ConsistentRead=True)["Count"] == 1 + + try: + asyncio.run(exercise()) + finally: + unregister_task(context) diff --git a/agent/tests/test_microvm_credentials.py b/agent/tests/test_microvm_credentials.py new file mode 100644 index 000000000..769d9de35 --- /dev/null +++ b/agent/tests/test_microvm_credentials.py @@ -0,0 +1,335 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Real botocore objects and loopback HTTP exercise the credential boundary.""" + +from __future__ import annotations + +import asyncio +import json +import threading +from datetime import UTC, datetime, timedelta, timezone +from unittest.mock import MagicMock +from urllib.error import HTTPError +from urllib.request import ProxyHandler, Request, build_opener + +import pytest +from botocore.config import Config +from botocore.credentials import Credentials, DeferredRefreshableCredentials + +import aws_session +from microvm_credentials import ScopedCredentialBroker +from microvm_lifecycle import LifecycleUnavailable, MicrovmLifecycle + + +@pytest.fixture(autouse=True) +def _reset(): + aws_session.reset_session_cache() + yield + aws_session.reset_session_cache() + + +@pytest.fixture +def anyio_backend(): + return "asyncio" + + +def _metadata(key: str) -> dict[str, str]: + return { + "access_key": key, + "secret_key": "synthetic-secret", + "token": "synthetic-token", + "expiry_time": (datetime.now(UTC) + timedelta(hours=1)).isoformat(), + } + + +def _envelope() -> dict[str, str]: + return { + "AccessKeyId": "SYNTHETIC", + "SecretAccessKey": "synthetic-secret", + "Token": "synthetic-token", + "Expiration": (datetime.now(UTC) + timedelta(hours=1)).strftime("%Y-%m-%dT%H:%M:%SZ"), + } + + +def _fetch(broker: ScopedCredentialBroker, *, token: str | None = None, path: str = ""): + url = broker.environment["AWS_CONTAINER_CREDENTIALS_FULL_URI"] + path + assert url.startswith("http://127.0.0.1:") + request = Request(url) # noqa: S310 — endpoint comes from this test's loopback broker + if token is not None: + request.add_header("Authorization", token) + with build_opener(ProxyHandler({})).open(request, timeout=3) as response: + assert response.headers["Cache-Control"] == "no-store" + return json.load(response) + + +class TestRetainedCredentials: + def _configure(self, monkeypatch): + monkeypatch.setenv( + aws_session.SESSION_ROLE_ARN_ENV, "arn:aws:iam::123456789012:role/session" + ) + monkeypatch.setenv("AWS_REGION", "us-west-2") + aws_session.configure_session("user", "owner/repo", "task") + order = [] + + def ambient_refresh(): + order.append("ambient") + return _metadata(f"AMBIENT_{len(order)}") + + ambient = DeferredRefreshableCredentials( + method="container-role", refresh_using=ambient_refresh + ) + sts = MagicMock() + sts._request_signer._credentials = ambient + + def assume(**_kwargs): + ambient.get_frozen_credentials() + order.append("tenant") + return { + "Credentials": { + "AccessKeyId": f"TENANT_{sts.assume_role.call_count}", + "SecretAccessKey": "synthetic-secret", + "SessionToken": "synthetic-token", + "Expiration": datetime.now(UTC) + timedelta(hours=1), + } + } + + sts.assume_role.side_effect = assume + monkeypatch.setattr("boto3.client", lambda *_a, **_kw: sts) + return sts, ambient, order + + def test_resume_refreshes_retained_clients_in_order_with_identical_tags(self, monkeypatch): + sts, ambient, order = self._configure(monkeypatch) + client = aws_session.tenant_client("s3", config=Config(signature_version="s3v4")) + resource = aws_session.tenant_resource("dynamodb") + original = aws_session.get_session().get_credentials() + before = aws_session.export_microvm_credentials("task") + tags = sts.assume_role.call_args.kwargs["Tags"] + assert order == ["ambient", "tenant"] + order.clear() + + aws_session.refresh_microvm_credentials("task") + + assert order == ["ambient", "tenant"] + assert client._request_signer._credentials is original + assert resource.meta.client._request_signer._credentials is original + assert sts._request_signer._credentials is ambient + assert aws_session.get_session().get_credentials() is original + assert ( + sts.assume_role.call_args.kwargs["Tags"] + == tags + == [ + {"Key": "user_id", "Value": "user"}, + {"Key": "repo", "Value": "owner/repo"}, + {"Key": "task_id", "Value": "task"}, + ] + ) + after = aws_session.export_microvm_credentials("task") + assert before["AccessKeyId"] != after["AccessKeyId"] + signed = client.generate_presigned_url( + "get_object", Params={"Bucket": "synthetic", "Key": "synthetic"} + ) + assert after["AccessKeyId"] in signed + assert before["AccessKeyId"] not in signed + + @pytest.mark.parametrize("microvm", [False, True]) + def test_short_sts_network_budget_applies_only_to_microvm(self, monkeypatch, microvm): + import microvm_lifecycle + + sts, _, _ = self._configure(monkeypatch) + make_client = MagicMock(return_value=sts) + monkeypatch.setattr("boto3.client", make_client) + context = microvm_lifecycle.register_task("task", "vm") if microvm else None + try: + aws_session.export_microvm_credentials("task") + config = make_client.call_args.kwargs["config"] + assert config.connect_timeout == (2 if microvm else 60) + assert config.read_timeout == (2 if microvm else 60) + assert config.retries == ({"total_max_attempts": 1} if microvm else None) + finally: + if context is not None: + microvm_lifecycle.unregister_task(context) + + def test_failed_ambient_refresh_never_attempts_tenant_renewal(self, monkeypatch): + sts, ambient, _ = self._configure(monkeypatch) + aws_session.export_microvm_credentials("task") + count = sts.assume_role.call_count + ambient._refresh_using = MagicMock(side_effect=RuntimeError("synthetic outage")) + with pytest.raises(RuntimeError, match="synthetic outage"): + aws_session.refresh_microvm_credentials("task") + assert sts.assume_role.call_count == count + + def test_failed_tenant_refresh_is_mandatory_even_before_expiry(self, monkeypatch): + sts, _, _ = self._configure(monkeypatch) + aws_session.export_microvm_credentials("task") + sts.assume_role.side_effect = RuntimeError("synthetic denial") + with pytest.raises(RuntimeError, match="synthetic denial"): + aws_session.refresh_microvm_credentials("task") + + def test_identity_cannot_change_after_clients_exist(self, monkeypatch): + self._configure(monkeypatch) + aws_session.export_microvm_credentials("task") + aws_session.configure_session("user", "owner/repo", "task") + with pytest.raises(aws_session.SessionScopingError, match="Cannot change identity"): + aws_session.configure_session("another", "owner/other", "another-task") + with pytest.raises(aws_session.SessionScopingError, match="original scoped task"): + aws_session.export_microvm_credentials("another-task") + assert aws_session._tags["user_id"] == "user" + + def test_static_runtime_credentials_cannot_claim_wake_renewal(self, monkeypatch): + sts, _, _ = self._configure(monkeypatch) + sts._request_signer._credentials = Credentials("STATIC", "synthetic") + aws_session.export_microvm_credentials("task") + with pytest.raises(aws_session.SessionScopingError, match="refreshable"): + aws_session.refresh_microvm_credentials("task") + + def test_absent_scoping_or_ambient_provider_is_rejected(self, monkeypatch): + self._configure(monkeypatch) + aws_session.export_microvm_credentials("task") + aws_session._ambient_credentials.clear() + with pytest.raises(aws_session.SessionScopingError, match="No runtime"): + aws_session.refresh_microvm_credentials("task") + monkeypatch.setattr(aws_session, "_scoped", False) + with pytest.raises(aws_session.SessionScopingError, match="original scoped task"): + aws_session.export_microvm_credentials("task") + + def test_expired_replacement_is_rejected(self): + metadata = _metadata("EXPIRED") + metadata["expiry_time"] = (datetime.now(UTC) - timedelta(seconds=1)).isoformat() + credentials = DeferredRefreshableCredentials(method="test", refresh_using=lambda: metadata) + with pytest.raises(RuntimeError, match="still expired"): + aws_session._locked_refresh(credentials, force=True) + + def test_expiration_is_utc_and_keys_are_a_coherent_pair(self): + metadata = _metadata("PAIR") + metadata["expiry_time"] = ( + datetime.now(UTC).astimezone(timezone(timedelta(hours=5))) + timedelta(hours=1) + ).isoformat() + credentials = DeferredRefreshableCredentials(method="test", refresh_using=lambda: metadata) + result = aws_session._locked_refresh(credentials, force=False) + expiry = datetime.fromisoformat(result["Expiration"]) + assert 3598 < (expiry - datetime.now(UTC)).total_seconds() <= 3600 + assert result["AccessKeyId"] == "PAIR" + + @pytest.mark.parametrize( + ("remaining_seconds", "force", "must_fail"), + [(12 * 60, False, False), (5 * 60, False, True), (12 * 60, True, True)], + ) + def test_refresh_outage_preserves_only_advisory_cached_credentials( + self, remaining_seconds, force, must_fail + ): + metadata = _metadata("CACHED") + credentials = DeferredRefreshableCredentials(method="test", refresh_using=lambda: metadata) + aws_session._locked_refresh(credentials, force=False) + credentials._expiry_time = datetime.now(UTC) + timedelta(seconds=remaining_seconds) + credentials._refresh_using = MagicMock(side_effect=RuntimeError("synthetic STS outage")) + if must_fail: + with pytest.raises(RuntimeError, match="synthetic STS outage"): + aws_session._locked_refresh(credentials, force=force) + else: + result = aws_session._locked_refresh(credentials, force=force) + assert result["AccessKeyId"] == "CACHED" + assert result["Token"] == metadata["token"] + credentials._refresh_using.assert_called_once() + + +class TestScopedBroker: + def test_requires_auth_and_scrubs_only_child_environment(self, monkeypatch): + monkeypatch.setenv("AWS_ACCESS_KEY_ID", "PARENT") + monkeypatch.setenv("AWS_PROFILE", "operator") + monkeypatch.setenv("AWS_WEB_IDENTITY_TOKEN_FILE", "/synthetic/token") + monkeypatch.setenv("AWS_CONTAINER_AUTHORIZATION_TOKEN_FILE", "/synthetic/ambient-token") + monkeypatch.setenv("AWS_FUTURE_CREDENTIAL_SOURCE", "synthetic") + monkeypatch.setenv("AWS_REGION", "us-west-2") + provider = MagicMock(side_effect=_envelope) + broker = ScopedCredentialBroker(MicrovmLifecycle("task", "vm"), provider=provider) + try: + import os + + assert os.environ["AWS_ACCESS_KEY_ID"] == "PARENT" + assert os.environ["AWS_PROFILE"] == "operator" + for key in ( + "AWS_ACCESS_KEY_ID", + "AWS_PROFILE", + "AWS_WEB_IDENTITY_TOKEN_FILE", + "AWS_CONTAINER_AUTHORIZATION_TOKEN_FILE", + "AWS_FUTURE_CREDENTIAL_SOURCE", + ): + assert broker.environment[key] == "" + assert "HOME" not in broker.environment + assert "AWS_REGION" not in broker.environment + for token, path, status in [(None, "", 403), ("wrong", "", 403), ("wrong", "/x", 404)]: + with pytest.raises(HTTPError) as error: + _fetch(broker, token=token, path=path) + assert error.value.code == status + provider.assert_not_called() + result = _fetch(broker, token=broker.environment["AWS_CONTAINER_AUTHORIZATION_TOKEN"]) + assert result["AccessKeyId"] == "SYNTHETIC" + finally: + broker.close() + broker.close() + assert not broker._thread.is_alive() + + def test_provider_failure_has_no_secret_details_or_fallback(self): + provider = MagicMock(side_effect=RuntimeError("do-not-leak-this")) + broker = ScopedCredentialBroker(MicrovmLifecycle("task", "vm"), provider=provider) + try: + with pytest.raises(HTTPError) as error: + _fetch(broker, token=broker.environment["AWS_CONTAINER_AUTHORIZATION_TOKEN"]) + assert error.value.code == 503 + assert b"do-not-leak-this" not in error.value.read() + provider.assert_called_once() + finally: + broker.close() + + @pytest.mark.anyio + async def test_suspend_drains_broker_and_resume_failure_keeps_it_closed(self): + class Deadline: + def remaining_s(self) -> float: + return 60 + + lifecycle = MicrovmLifecycle("task", "vm") + await lifecycle.tool_started("tool") + lifecycle.park_approval("gate", "tool", Deadline()) + entered, release = threading.Event(), threading.Event() + + def provider(): + entered.set() + assert release.wait(2) + return _envelope() + + broker = ScopedCredentialBroker(lifecycle, provider=provider) + try: + fetch = asyncio.create_task( + asyncio.to_thread( + _fetch, broker, token=broker.environment["AWS_CONTAINER_AUTHORIZATION_TOKEN"] + ) + ) + assert await asyncio.to_thread(entered.wait, 1) + checkpoint = MagicMock() + suspend = asyncio.create_task(lifecycle.suspend(checkpoint, budget_s=2)) + await asyncio.sleep(0.03) + checkpoint.assert_not_called() + release.set() + await fetch + await suspend + checkpoint.assert_called_once() + with pytest.raises(HTTPError) as error: + await asyncio.to_thread( + _fetch, broker, token=broker.environment["AWS_CONTAINER_AUTHORIZATION_TOKEN"] + ) + assert error.value.code == 503 + with pytest.raises(RuntimeError, match="renewal failed"): + await lifecycle.resume( + MagicMock(side_effect=RuntimeError("renewal failed")), budget_s=1 + ) + with pytest.raises(LifecycleUnavailable): + await lifecycle.wait_until_open() + with pytest.raises(HTTPError) as error: + await asyncio.to_thread( + _fetch, broker, token=broker.environment["AWS_CONTAINER_AUTHORIZATION_TOKEN"] + ) + assert error.value.code == 503 + finally: + release.set() + broker.close() diff --git a/agent/tests/test_microvm_http.py b/agent/tests/test_microvm_http.py new file mode 100644 index 000000000..e0c77b9e2 --- /dev/null +++ b/agent/tests/test_microvm_http.py @@ -0,0 +1,498 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Exercise production HTTP routes with real controller state and bounded work.""" + +from __future__ import annotations + +import asyncio +import json +import threading +import time +from dataclasses import dataclass +from unittest.mock import MagicMock + +import httpx +import pytest +from botocore.exceptions import ClientError +from fastapi import Request + +import microvm_http +import microvm_lifecycle as lifecycle +import server +from microvm_diagnostics import lifecycle_stage + +PREFIX = server.MICROVM_HOOK_PREFIX + + +def diagnostic_records(capsys): + return [ + json.loads(line) + for line in capsys.readouterr().out.splitlines() + if line.startswith("{") and '"event": "microvm_hook_' in line + ] + + +@dataclass +class Deadline: + remaining: float = 60 + + def remaining_s(self) -> float: + return self.remaining + + +@pytest.fixture +def anyio_backend(): + return "asyncio" + + +@pytest.fixture +def context(): + context = lifecycle.register_task("http-task", "http-vm") + yield context + lifecycle.unregister_task(context) + + +@pytest.fixture +def callbacks(monkeypatch): + suspend, resume = MagicMock(), MagicMock() + monkeypatch.setattr(microvm_http, "checkpoint_before_suspend", suspend) + monkeypatch.setattr(microvm_http, "refresh_and_reconcile_after_resume", resume) + return suspend, resume + + +@pytest.fixture +async def client(): + async with httpx.AsyncClient( + transport=httpx.ASGITransport(app=server.app), base_url="http://test" + ) as client: + yield client + + +async def park(context): + await context.tool_started("tool") + deadline = Deadline() + parked = context.park_approval("gate", "tool", deadline) + assert parked is not None + return parked, deadline + + +@pytest.mark.anyio +class TestLifecycleHttp: + @pytest.mark.parametrize("action", ["suspend", "resume"]) + @pytest.mark.parametrize( + ("outcome", "expected_status"), + [("acknowledged", 200), ("unavailable", 409), ("invalid", 400), ("failed", 503)], + ) + async def test_lifecycle_responses_close_even_when_client_requests_keep_alive( + self, context, callbacks, client, action, outcome, expected_status + ): + await park(context) + if action == "resume": + assert (await client.post(PREFIX + "/suspend", json={})).status_code == 200 + if outcome == "unavailable": + lifecycle.unregister_task(context) + elif outcome == "failed": + callback = callbacks[0] if action == "suspend" else callbacks[1] + callback.side_effect = RuntimeError("callback failed") + response = await client.post( + PREFIX + "/" + action, + content=b"{" if outcome == "invalid" else b"{}", + headers={"Connection": "keep-alive"}, + ) + assert response.status_code == expected_status + assert response.headers["connection"] == "close" + + async def test_hooks_have_distinct_correlated_timelines( + self, context, callbacks, client, capsys + ): + await park(context) + for action in ["suspend", "resume"]: + response = await client.post(PREFIX + "/" + action, json={}) + assert response.status_code == 200 + records = diagnostic_records(capsys) + starts = [row for row in records if row["event"] == "microvm_hook_started"] + ends = [row for row in records if row["event"] == "microvm_hook_finished"] + assert len(starts) == len(ends) == 2 + assert starts[0]["hook_id"] != starts[1]["hook_id"] + for start, end in zip(starts, ends, strict=True): + assert start["hook_id"] == end["hook_id"] + assert end["task_id"] == "http-task" + assert end["microvm_id"] == "http-vm" + assert end["request_id"] == "gate" + assert end["pid"] > 0 + assert end["elapsed_ms"] >= 0 + assert end["http_status"] == 200 + assert end["code"] == "acknowledged" + assert end["late"] is False + assert ends[0]["phase"] == "suspend-ready" + assert ends[1]["phase"] == "parked" + + async def test_failed_refresh_logs_stage_and_aws_identity_without_secrets( + self, context, callbacks, client, capsys + ): + await park(context) + assert (await client.post(PREFIX + "/suspend", json={})).status_code == 200 + capsys.readouterr() + + def fail_refresh(_park): + with lifecycle_stage("credential-refresh"): + raise ClientError( + { + "Error": {"Code": "AccessDenied", "Message": "secret-credential"}, + "ResponseMetadata": { + "RequestId": "aws-request-123", + "HTTPHeaders": {"Authorization": "secret-header"}, + }, + }, + "AssumeRole", + ) + + callbacks[1].side_effect = fail_refresh + response = await client.post(PREFIX + "/resume", json={"ignored": "secret-body"}) + records = diagnostic_records(capsys) + serialized = json.dumps(records) + response.text + assert "secret-" not in serialized + failure = next(row for row in records if row["event"] == "microvm_hook_stage_failed") + assert failure["stage"] == "credential-refresh" + assert failure["error_type"] == "ClientError" + assert failure["aws_error_code"] == "AccessDenied" + assert failure["aws_request_id"] == "aws-request-123" + end = records[-1] + assert end["stage"] == "credential-refresh" + assert end["phase"] == "failed" + assert end["http_status"] == response.status_code == 503 + assert all(row["hook_id"] == end["hook_id"] for row in records) + with pytest.raises(lifecycle.LifecycleUnavailable): + await context.wait_until_open() + + async def test_logging_failure_does_not_change_hook_result( + self, context, callbacks, client, monkeypatch + ): + await park(context) + monkeypatch.setattr( + "microvm_diagnostics.print", MagicMock(side_effect=OSError("closed")), raising=False + ) + assert (await client.post(PREFIX + "/suspend", json={})).status_code == 200 + assert (await client.post(PREFIX + "/resume", json={})).status_code == 200 + await context.wait_until_open() + + async def test_timeout_reports_blocked_stage_and_marks_late_thread( + self, context, callbacks, client, monkeypatch, capsys + ): + await park(context) + assert (await client.post(PREFIX + "/suspend", json={})).status_code == 200 + capsys.readouterr() + release, finished = threading.Event(), threading.Event() + + def slow_refresh(_park): + try: + with lifecycle_stage("credential-refresh"): + assert release.wait(2) + finally: + finished.set() + + callbacks[1].side_effect = slow_refresh + monkeypatch.setattr(microvm_http, "LIFECYCLE_HANDLER_BUDGET_S", 0.05) + try: + response = await client.post(PREFIX + "/resume", json={}) + assert response.status_code == 503 + before = diagnostic_records(capsys) + end = before[-1] + assert end["stage"] == "credential-refresh" + assert end["code"] == "MICROVM_LIFECYCLE_TIMEOUT" + assert end["phase"] == "failed" + finally: + release.set() + assert await asyncio.to_thread(finished.wait, 2) + after = diagnostic_records(capsys) + assert after + assert all(row["late"] and row["hook_id"] == end["hook_id"] for row in after) + assert all(row["event"] != "microvm_hook_finished" for row in after) + with pytest.raises(lifecycle.LifecycleUnavailable): + await context.wait_until_open() + + async def test_duplicate_hooks_acknowledge_without_repeating_work( + self, context, callbacks, client + ): + suspend, resume = callbacks + parked, deadline = await park(context) + for body in [{"microvmId": "http-vm"}, {"microvmId": ""}]: + response = await client.post(PREFIX + "/suspend", json=body) + assert response.status_code == 200 + assert response.json()["request_id"] == "gate" + suspend.assert_called_once_with(parked) + leave = asyncio.create_task(context.leave_approval(parked)) + await asyncio.sleep(0.03) + assert not leave.done() + deadline.remaining = 0 + response = await client.post(PREFIX + "/resume", content=b"") + assert response.status_code == 200 + await leave + response = await client.post(PREFIX + "/resume", json={"microvmId": "http-vm"}) + assert response.status_code == 200 + resume.assert_called_once_with(parked) + assert parked.deadline is deadline + assert deadline.remaining_s() == 0 + assert (await client.post(PREFIX + "/suspend", json={})).status_code == 409 + + async def test_new_gate_cannot_reuse_an_old_wake_acknowledgment( + self, context, callbacks, client + ): + parked, _ = await park(context) + assert (await client.post(PREFIX + "/suspend", json={})).status_code == 200 + assert (await client.post(PREFIX + "/resume", json={})).status_code == 200 + await context.leave_approval(parked) + context.tool_finished("tool") + await context.tool_started("next-tool") + next_park = context.park_approval("next-gate", "next-tool", Deadline()) + assert next_park is not None + assert (await client.post(PREFIX + "/resume", json={})).status_code == 409 + assert (await client.post(PREFIX + "/suspend", json={})).status_code == 200 + response = await client.post(PREFIX + "/resume", json={}) + assert response.status_code == 200 + assert response.json()["request_id"] == "next-gate" + callbacks[1].assert_called_with(next_park) + assert callbacks[1].call_count == 2 + + @pytest.mark.parametrize( + ("content", "status"), + [ + (b"{", 400), + (b"[]", 400), + (b'{"microvmId":1}', 400), + (b'{"microvmId":" http-vm"}', 400), + (b"x" * 4097, 413), + ], + ) + async def test_invalid_body_never_reaches_lifecycle_work( + self, context, callbacks, client, content, status + ): + await park(context) + assert (await client.post(PREFIX + "/suspend", content=content)).status_code == status + assert (await client.post(PREFIX + "/resume", content=content)).status_code == status + for callback in callbacks: + callback.assert_not_called() + + async def test_wrong_vm_or_no_registered_task_rejects(self, context, callbacks, client): + await park(context) + assert ( + await client.post(PREFIX + "/suspend", json={"microvmId": "another-vm"}) + ).status_code == 409 + lifecycle.unregister_task(context) + assert (await client.post(PREFIX + "/suspend", json={})).status_code == 409 + assert (await client.post(PREFIX + "/resume", json={})).status_code == 409 + for callback in callbacks: + callback.assert_not_called() + + async def test_unparked_resume_or_parallel_tools_never_acknowledges( + self, context, callbacks, client + ): + assert (await client.post(PREFIX + "/suspend", json={})).status_code == 409 + await park(context) + assert (await client.post(PREFIX + "/resume", json={})).status_code == 409 + await context.tool_started("parallel") + assert (await client.post(PREFIX + "/suspend", json={})).status_code == 409 + for callback in callbacks: + callback.assert_not_called() + + async def test_checkpoint_failure_keeps_guest_awake_and_disables_another_suspend( + self, context, callbacks, client + ): + suspend, _ = callbacks + parked, _ = await park(context) + suspend.side_effect = RuntimeError("do-not-leak-secret") + response = await client.post(PREFIX + "/suspend", json={}) + assert response.status_code == 503 + assert "do-not-leak-secret" not in response.text + assert (await client.post(PREFIX + "/suspend", json={})).status_code == 409 + await context.leave_approval(parked) + + async def test_failed_refresh_closes_barrier_without_retrying(self, context, callbacks, client): + _, resume = callbacks + await park(context) + assert (await client.post(PREFIX + "/suspend", json={})).status_code == 200 + resume.side_effect = RuntimeError("do-not-leak-secret") + response = await client.post(PREFIX + "/resume", json={}) + assert response.status_code == 503 + assert "do-not-leak-secret" not in response.text + with pytest.raises(lifecycle.LifecycleUnavailable): + await context.wait_until_open() + assert (await client.post(PREFIX + "/resume", json={})).status_code == 409 + resume.assert_called_once() + + @pytest.mark.parametrize("action", ["suspend", "resume"]) + async def test_timed_out_callback_cannot_acknowledge_late( + self, context, callbacks, client, monkeypatch, action + ): + await park(context) + if action == "resume": + assert (await client.post(PREFIX + "/suspend", json={})).status_code == 200 + entered, release, finished = threading.Event(), threading.Event(), threading.Event() + + def block(_park): + entered.set() + try: + assert release.wait(2) + finally: + finished.set() + + callbacks[action == "resume"].side_effect = block + monkeypatch.setattr(microvm_http, "LIFECYCLE_HANDLER_BUDGET_S", 0.05) + started = time.monotonic() + try: + response = await client.post(PREFIX + "/" + action, json={}) + assert response.status_code == 503 + assert response.json()["code"] == "MICROVM_LIFECYCLE_TIMEOUT" + assert time.monotonic() - started < 0.5 + assert entered.is_set() + finally: + release.set() + assert await asyncio.to_thread(finished.wait, 1) + assert (await client.post(PREFIX + "/" + action, json={})).status_code == 409 + if action == "resume": + with pytest.raises(lifecycle.LifecycleUnavailable): + await context.wait_until_open() + else: + await context.wait_until_open() + + async def test_terminate_invalidates_an_inflight_refresh( + self, context, callbacks, client, monkeypatch + ): + await park(context) + assert (await client.post(PREFIX + "/suspend", json={})).status_code == 200 + entered, release = threading.Event(), threading.Event() + + def refresh(_park): + entered.set() + assert release.wait(2) + + callbacks[1].side_effect = refresh + monkeypatch.setattr(server, "_debug_cw", MagicMock()) + request = asyncio.create_task(client.post(PREFIX + "/resume", json={})) + try: + assert await asyncio.to_thread(entered.wait, 1) + assert ( + await client.post(PREFIX + "/terminate", content=b"bad body") + ).status_code == 200 + with pytest.raises(lifecycle.LifecycleUnavailable): + await context.wait_until_open() + finally: + release.set() + assert (await request).status_code == 409 + with pytest.raises(lifecycle.LifecycleUnavailable): + await context.wait_until_open() + + async def test_concurrent_resume_reports_busy_while_first_finishes( + self, context, callbacks, client, capsys + ): + await park(context) + assert (await client.post(PREFIX + "/suspend", json={})).status_code == 200 + entered, release = threading.Event(), threading.Event() + + def refresh(_park): + entered.set() + assert release.wait(2) + + callbacks[1].side_effect = refresh + request = asyncio.create_task(client.post(PREFIX + "/resume", json={})) + try: + assert await asyncio.to_thread(entered.wait, 1) + assert (await client.post(PREFIX + "/resume", json={})).status_code == 409 + finally: + release.set() + assert (await request).status_code == 200 + callbacks[1].assert_called_once() + records = [row for row in diagnostic_records(capsys) if row["action"] == "resume"] + ends = [row for row in records if row["event"] == "microvm_hook_finished"] + assert [row["http_status"] for row in ends] == [409, 200] + assert ends[0]["hook_id"] != ends[1]["hook_id"] + for end in ends: + matching = [row for row in records if row["hook_id"] == end["hook_id"]] + assert matching[0]["event"] == "microvm_hook_started" + assert matching[-1] == end + + async def test_request_body_read_shares_the_total_budget( + self, context, callbacks, monkeypatch, capsys + ): + await park(context) + monkeypatch.setattr(microvm_http, "LIFECYCLE_HANDLER_BUDGET_S", 0.02) + + async def slow_receive(): + await asyncio.sleep(1) + return {"type": "http.request", "body": b"{}", "more_body": False} + + request = Request({"type": "http", "method": "POST", "headers": []}, slow_receive) + response = await microvm_http.microvm_suspend(request) + assert response.status_code == 503 + for callback in callbacks: + callback.assert_not_called() + end = diagnostic_records(capsys)[-1] + assert end["stage"] == "body-read" + assert end["code"] == "MICROVM_LIFECYCLE_TIMEOUT" + + async def test_cancelled_handler_logs_cancellation_and_still_propagates_it( + self, context, callbacks, capsys + ): + await park(context) + entered = asyncio.Event() + + async def receive(): + entered.set() + await asyncio.Event().wait() + + request = Request({"type": "http", "method": "POST", "headers": []}, receive) + pending = asyncio.create_task(microvm_http.microvm_suspend(request)) + await entered.wait() + pending.cancel() + with pytest.raises(asyncio.CancelledError): + await pending + end = diagnostic_records(capsys)[-1] + assert end["code"] == "MICROVM_LIFECYCLE_CANCELLED" + assert end["http_status"] is None + assert end["stage"] == "body-read" + for callback in callbacks: + callback.assert_not_called() + + async def test_terminate_closes_barrier_and_answers_even_if_body_stalls( + self, context, monkeypatch + ): + await park(context) + monkeypatch.setattr(server, "_TERMINATE_BODY_BUDGET_SECONDS", 0.02) + monkeypatch.setattr(server, "_debug_cw", MagicMock()) + + async def slow_receive(): + with pytest.raises(lifecycle.LifecycleUnavailable): + await context.wait_until_open() + await asyncio.sleep(1) + return {"type": "http.request", "body": b"{}", "more_body": False} + + request = Request({"type": "http", "method": "POST", "headers": []}, slow_receive) + started = time.monotonic() + response = await server.microvm_terminate(request) + assert response["status"] == "acknowledged" + assert time.monotonic() - started < 0.5 + + async def test_expired_or_failed_progress_cannot_repeat_suspend_ack( + self, context, callbacks, client + ): + _, deadline = await park(context) + assert (await client.post(PREFIX + "/suspend", json={})).status_code == 200 + deadline.remaining = 0 + assert (await client.post(PREFIX + "/suspend", json={})).status_code == 409 + deadline.remaining = 60 + context.progress_write_failed() + assert (await client.post(PREFIX + "/suspend", json={})).status_code == 409 + callbacks[0].assert_called_once() + + async def test_validate_checks_new_routes_without_creating_aws_clients( + self, monkeypatch, client + ): + forbidden = MagicMock(side_effect=AssertionError("build hook initialized AWS")) + monkeypatch.setattr("aws_session.tenant_client", forbidden) + monkeypatch.setattr("aws_session.platform_client", forbidden) + response = await client.post(PREFIX + "/validate") + assert response.status_code == 200 + assert response.json()["checks"]["hook_routes_registered"] is True + forbidden.assert_not_called() + assert 0 < microvm_http.LIFECYCLE_HANDLER_BUDGET_S < microvm_http.LIFECYCLE_HOOK_TIMEOUT_S diff --git a/agent/tests/test_microvm_lifecycle.py b/agent/tests/test_microvm_lifecycle.py new file mode 100644 index 000000000..7f913c9e1 --- /dev/null +++ b/agent/tests/test_microvm_lifecycle.py @@ -0,0 +1,568 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 + +"""Concurrency regressions for the guest approval/sleep boundary.""" + +import asyncio +import threading +from unittest.mock import Mock + +import pytest + +from hooks import _ApprovalDeadline +from microvm_lifecycle import ( + LifecycleUnavailable, + MicrovmLifecycle, + get_context, + register_task, + unregister_task, +) + + +@pytest.fixture +def anyio_backend(): + return "asyncio" + + +@pytest.fixture(autouse=True) +def reset_progress_state(): + from progress_writer import _reset_circuit_breakers + + _reset_circuit_breakers() + yield + _reset_circuit_breakers() + + +async def parked(context=None): + context = context or MicrovmLifecycle("task", "microvm") + deadline = Mock(remaining_s=Mock(return_value=300)) + await context.tool_started("tool") + park = context.park_approval("request", "tool", deadline) + assert park is not None + return context, park + + +@pytest.mark.anyio +async def test_suspend_blocks_decision_and_new_tool_until_refresh_finishes(): + context, park = await parked() + checkpoint = Mock() + await context.suspend(checkpoint, budget_s=1) + checkpoint.assert_called_once_with(park) + decision = asyncio.create_task(context.leave_approval(park)) + new_tool = asyncio.create_task(context.tool_started("parallel")) + await asyncio.sleep(0.04) + assert not decision.done() + assert not new_tool.done() + refreshing = threading.Event() + release = threading.Event() + + def refresh(same_park): + assert same_park is park + refreshing.set() + assert release.wait(2) + + wake = asyncio.create_task(context.resume(refresh, budget_s=1)) + try: + assert await asyncio.to_thread(refreshing.wait, 1) + assert not decision.done() + assert not new_tool.done() + finally: + release.set() + await wake + await asyncio.wait_for(asyncio.gather(decision, new_tool), 1) + + +@pytest.mark.anyio +async def test_parallel_tool_prevents_suspend_until_its_post_hook(): + context, park = await parked() + await context.tool_started("other") + checkpoint = Mock() + with pytest.raises(LifecycleUnavailable): + await context.suspend(checkpoint, budget_s=1) + checkpoint.assert_not_called() + context.tool_finished("other") + assert await context.suspend(checkpoint, budget_s=1) is park + + +@pytest.mark.anyio +@pytest.mark.parametrize( + "reason", ["unknown-tool", "duplicate-tool", "detached-tool", "dropped-event"] +) +async def test_unaccounted_work_or_progress_disables_sleep_without_blocking_tools(reason): + context, park = await parked() + if reason == "unknown-tool": + await context.tool_started(None) + elif reason == "duplicate-tool": + await context.tool_started("tool") + elif reason == "detached-tool": + context.disable_suspend() + else: + context.progress_write_failed() + with context.activity(): + pass # A later success does not recover the lost event. + with pytest.raises(LifecycleUnavailable): + await context.suspend(Mock(), budget_s=1) + await context.leave_approval(park) + await context.tool_started("next") + + +@pytest.mark.anyio +async def test_suspend_drains_inflight_progress_before_checkpoint(): + context, park = await parked() + entered = threading.Event() + release = threading.Event() + + def write(): + with context.activity(): + entered.set() + assert release.wait(2) + + writer = asyncio.create_task(asyncio.to_thread(write)) + assert await asyncio.to_thread(entered.wait, 1) + checkpoint = Mock() + suspend = asyncio.create_task(context.suspend(checkpoint, budget_s=1)) + try: + await asyncio.sleep(0.04) + checkpoint.assert_not_called() + finally: + release.set() + await writer + assert await suspend is park + checkpoint.assert_called_once() + + +@pytest.mark.anyio +async def test_failed_inflight_progress_rejects_suspend(): + context, park = await parked() + entered = threading.Event() + release = threading.Event() + + def write(): + with context.activity(): + entered.set() + assert release.wait(2) + context.progress_write_failed() + + writer = asyncio.create_task(asyncio.to_thread(write)) + assert await asyncio.to_thread(entered.wait, 1) + checkpoint = Mock() + suspend = asyncio.create_task(context.suspend(checkpoint, budget_s=1)) + await asyncio.sleep(0.04) + release.set() + await writer + with pytest.raises(LifecycleUnavailable): + await suspend + checkpoint.assert_not_called() + await context.leave_approval(park) + + +@pytest.mark.anyio +async def test_checkpoint_failure_releases_unsuspended_wait_and_disables_retry(): + context, park = await parked() + with pytest.raises(OSError, match="durability"): + await context.suspend(Mock(side_effect=OSError("durability")), budget_s=1) + with pytest.raises(LifecycleUnavailable): + await context.suspend(Mock(), budget_s=1) + await context.leave_approval(park) + + +@pytest.mark.anyio +async def test_expiry_during_checkpoint_refuses_freeze(): + context, park = await parked() + + def checkpoint(_): + park.deadline.remaining_s.return_value = 0 + + with pytest.raises(LifecycleUnavailable): + await context.suspend(checkpoint, budget_s=1) + await context.leave_approval(park) + + +@pytest.mark.anyio +async def test_late_refresh_after_timeout_cannot_release_coding(): + context, park = await parked() + await context.suspend(Mock(), budget_s=1) + entered = threading.Event() + release = threading.Event() + completed = threading.Event() + + def refresh(_): + entered.set() + assert release.wait(2) + completed.set() + + wake = asyncio.create_task(context.resume(refresh, budget_s=0.1)) + try: + assert await asyncio.to_thread(entered.wait, 1) + with pytest.raises(TimeoutError): + await wake + with pytest.raises(LifecycleUnavailable): + await context.leave_approval(park) + finally: + release.set() + assert await asyncio.to_thread(completed.wait, 1) + with pytest.raises(LifecycleUnavailable): + await context.tool_started("next") + with pytest.raises(LifecycleUnavailable): + await context.resume(Mock(), budget_s=1) + + +@pytest.mark.anyio +async def test_close_during_refresh_invalidates_completion(): + context, park = await parked() + await context.suspend(Mock(), budget_s=1) + with pytest.raises(LifecycleUnavailable, match="superseded"): + await context.resume(lambda _: context.close(), budget_s=1) + with pytest.raises(LifecycleUnavailable): + await context.leave_approval(park) + + +@pytest.mark.anyio +async def test_resume_failure_and_cancellation_stay_closed(): + for failure in (OSError("STS unavailable"), asyncio.CancelledError()): + context, park = await parked() + await context.suspend(Mock(), budget_s=1) + with pytest.raises(type(failure)): + await context.resume(Mock(side_effect=failure), budget_s=1) + with pytest.raises(LifecycleUnavailable): + await context.leave_approval(park) + + +@pytest.mark.anyio +async def test_concurrent_lifecycle_calls_do_not_share_transition_ownership(): + context, park = await parked() + entered = threading.Event() + release = threading.Event() + + def checkpoint(_): + entered.set() + assert release.wait(2) + + suspend = asyncio.create_task(context.suspend(checkpoint, budget_s=1)) + try: + assert await asyncio.to_thread(entered.wait, 1) + with pytest.raises(LifecycleUnavailable): + await context.suspend(Mock(), budget_s=1) + with pytest.raises(LifecycleUnavailable): + await context.resume(Mock(), budget_s=1) + finally: + release.set() + await suspend + await context.resume(Mock(), budget_s=1) + # The same gate cannot sleep again while the released waiter catches up. + with pytest.raises(LifecycleUnavailable): + await context.suspend(Mock(), budget_s=1) + await context.leave_approval(park) + context.tool_finished("tool") + await context.tool_started("next") + next_park = context.park_approval("next-gate", "next", park.deadline) + assert await context.suspend(Mock(), budget_s=1) is next_park + + +@pytest.mark.anyio +async def test_original_deadline_survives_freeze_and_expired_resume(monkeypatch): + clock = {"wall": 1000.0, "mono": 50.0} + # Patch only the module method through a proxy; replacing global monotonic + # would also freeze asyncio's own timeout scheduler. + monkeypatch.setattr( + "hooks.time", + Mock(time=lambda: clock["wall"], monotonic=lambda: clock["mono"]), + ) + deadline = _ApprovalDeadline.from_recorded("1970-01-01T00:16:40Z", 300) + context = MicrovmLifecycle("task", "microvm") + await context.tool_started("tool") + park = context.park_approval("request", "tool", deadline) + await context.suspend(Mock(), budget_s=1) + clock["wall"] += 400 # Guest monotonic clock did not advance while frozen. + refresh = Mock() + assert await context.resume(refresh, budget_s=1) is park + assert park.deadline is deadline + assert park.deadline.remaining_s() == 0 + refresh.assert_called_once_with(park) + await context.leave_approval(park) + + +@pytest.mark.anyio +async def test_leaving_approval_before_suspend_wins_without_checkpoint(): + context, park = await parked() + await context.leave_approval(park) + checkpoint = Mock() + with pytest.raises(LifecycleUnavailable): + await context.suspend(checkpoint, budget_s=1) + checkpoint.assert_not_called() + + +def test_registry_rejects_second_pipeline_and_unregister_uses_identity(): + context = register_task("task", "microvm") + try: + assert get_context("task") is context + assert get_context("other") is None + with pytest.raises(LifecycleUnavailable): + register_task("other", "other-vm") + unregister_task(MicrovmLifecycle("task", "other-vm")) + assert get_context("task") is context + finally: + unregister_task(context) + assert get_context("task") is None + + +@pytest.mark.anyio +async def test_progress_writer_acknowledgments_are_shared_and_missing_table_refuses_sleep( + monkeypatch, +): + from progress_writer import _ProgressWriter + + context = register_task("progress-task", "microvm") + try: + await context.tool_started("tool") + context.park_approval("gate", "tool", Mock(remaining_s=Mock(return_value=300))) + monkeypatch.setenv("TASK_EVENTS_TABLE_NAME", "events") + writer = _ProgressWriter("progress-task") + writer._table = Mock() + writer._put_event("agent_milestone", {"message": "saved"}) + writer._table.put_item.assert_called_once() + + # A second writer's dropped event must invalidate the same task's + # barrier, even after the first writer successfully writes again. + missing = _ProgressWriter("progress-task") + missing._table_name = None + missing._put_event("agent_milestone", {"message": "lost"}) + writer._put_event("agent_milestone", {"message": "later success"}) + with pytest.raises(LifecycleUnavailable): + await context.suspend(Mock(), budget_s=1) + finally: + unregister_task(context) + + +@pytest.mark.anyio +async def test_real_progress_writer_failure_cannot_acknowledge_suspend(monkeypatch): + from progress_writer import _ProgressWriter + + context = register_task("write-error-task", "microvm") + try: + await context.tool_started("tool") + context.park_approval("gate", "tool", Mock(remaining_s=Mock(return_value=300))) + monkeypatch.setenv("TASK_EVENTS_TABLE_NAME", "events") + writer = _ProgressWriter("write-error-task") + writer._table = Mock() + writer._table.put_item.side_effect = OSError("write reply lost") + writer._put_event("agent_milestone", {"message": "uncertain"}) + with pytest.raises(LifecycleUnavailable): + await context.suspend(Mock(), budget_s=1) + finally: + unregister_task(context) + + +@pytest.mark.anyio +async def test_progress_during_continuation_capture_is_acknowledged_before_suspend(monkeypatch): + from progress_writer import _ProgressWriter + + context = register_task("capture-progress-task", "microvm") + try: + _, park = await parked(context) + monkeypatch.setenv("TASK_EVENTS_TABLE_NAME", "events") + writer = _ProgressWriter(context.task_id) + writer._table = Mock() + async with context.continuation_checkpoint(park): + # SDK progress can arrive while the workspace upload awaits a thread. + writer._put_event("agent_turn", {"message": "waiting for approval"}) + writer._table.put_item.assert_called_once() + assert not context.diagnostic_snapshot()["progress_failed"] + with pytest.raises(LifecycleUnavailable), context.activity(): + pytest.fail("General activity must remain paused during capture") + await context.suspend(Mock(), budget_s=1) + # The progress exception is limited to capture, never actual suspension. + writer._put_event("agent_turn", {"message": "cannot write while suspended"}) + writer._table.put_item.assert_called_once() + assert context.diagnostic_snapshot()["progress_failed"] + finally: + unregister_task(context) + + +@pytest.mark.anyio +async def test_capture_waits_for_progress_started_during_upload(monkeypatch): + from progress_writer import _ProgressWriter + + context = register_task("capture-progress-drain-task", "microvm") + entered = threading.Event() + release = threading.Event() + write_task = None + capture_task = None + try: + _, park = await parked(context) + monkeypatch.setenv("TASK_EVENTS_TABLE_NAME", "events") + writer = _ProgressWriter(context.task_id) + writer._table = Mock() + + def slow_write(**_): + entered.set() + assert release.wait(2) + + writer._table.put_item.side_effect = slow_write + + async def capture(): + nonlocal write_task + async with context.continuation_checkpoint(park, drain_budget_s=1): + write_task = asyncio.create_task( + asyncio.to_thread(writer._put_event, "agent_turn", {"message": "in flight"}) + ) + assert await asyncio.to_thread(entered.wait, 1) + + capture_task = asyncio.create_task(capture()) + assert await asyncio.to_thread(entered.wait, 1) + await asyncio.sleep(0.04) + assert not capture_task.done() + assert context.diagnostic_snapshot()["phase"] == "checkpointing" + release.set() + await asyncio.wait_for(capture_task, 1) + assert write_task is not None + await write_task + assert context.diagnostic_snapshot()["phase"] == "checkpoint-ready" + await context.suspend(Mock(), budget_s=1) + finally: + release.set() + if capture_task: + await asyncio.gather(capture_task, return_exceptions=True) + if write_task: + await asyncio.gather(write_task, return_exceptions=True) + unregister_task(context) + + +@pytest.mark.anyio +async def test_progress_failure_during_capture_prevents_checkpoint_publication(monkeypatch): + from progress_writer import _ProgressWriter + + context = register_task("capture-progress-error-task", "microvm") + try: + _, park = await parked(context) + monkeypatch.setenv("TASK_EVENTS_TABLE_NAME", "events") + writer = _ProgressWriter(context.task_id) + writer._table = Mock() + writer._table.put_item.side_effect = OSError("write reply lost") + with pytest.raises(LifecycleUnavailable, match="Progress was not acknowledged"): + async with context.continuation_checkpoint(park): + writer._put_event("agent_turn", {"message": "uncertain"}) + assert context.diagnostic_snapshot()["phase"] == "parked" + with pytest.raises(LifecycleUnavailable): + await context.suspend(Mock(), budget_s=1) + finally: + unregister_task(context) + + +@pytest.mark.anyio +async def test_resume_reseeds_from_fresh_os_entropy_after_refresh(monkeypatch): + from microvm_lifecycle import reseed_random + + entropy = Mock(side_effect=[b"r" * 32, b"w" * 32]) + seed = Mock() + monkeypatch.setattr("microvm_lifecycle.os.urandom", entropy) + monkeypatch.setattr("microvm_lifecycle.random.seed", seed) + reseed_random() # Same entry point used by /run. + context, _ = await parked() + await context.suspend(Mock(), budget_s=1) + + def refresh(_): + seed.assert_called_once_with(b"r" * 32) + + await context.resume(refresh, budget_s=1) + assert [call.args for call in entropy.call_args_list] == [(32,), (32,)] + assert [call.args for call in seed.call_args_list] == [(b"r" * 32,), (b"w" * 32,)] + + +@pytest.mark.anyio +async def test_heartbeat_does_not_write_until_resume_finishes(monkeypatch): + import server + + context = register_task("heartbeat-task", "microvm") + write = Mock() + monkeypatch.setattr(server.task_state, "write_heartbeat", write) + try: + await context.tool_started("tool") + context.park_approval("gate", "tool", Mock(remaining_s=Mock(return_value=300))) + await context.suspend(Mock(), budget_s=1) + server._heartbeat_worker(context.task_id, Mock(wait=Mock(side_effect=[False, True]))) + write.assert_not_called() + await context.resume(Mock(), budget_s=1) + server._heartbeat_worker(context.task_id, Mock(wait=Mock(side_effect=[False, True]))) + write.assert_called_once_with(context.task_id) + finally: + unregister_task(context) + + +@pytest.mark.anyio +@pytest.mark.parametrize("background", [False, True, "serialized"]) +async def test_sdk_failure_hook_releases_tool_and_detached_work_disables_sleep( + monkeypatch, background +): + import hooks + + async def allow(*args, **kwargs): + return {"hookSpecificOutput": {"permissionDecision": "allow"}} + + monkeypatch.setattr(hooks, "pre_tool_use_hook", allow) + context = register_task("sdk-tools-task", "microvm") + try: + matchers = hooks.build_hook_matchers(engine=Mock(), task_id=context.task_id) + pre = matchers["PreToolUse"][0].hooks[0] + tool_input = ( + '{"run_in_background": true}' + if background == "serialized" + else {"run_in_background": background} + ) + await pre({"tool_name": "Bash", "tool_input": tool_input}, "first", {}) + await matchers["PostToolUseFailure"][0].hooks[0]({}, "first", {}) + await pre({"tool_name": "Bash", "tool_input": {}}, "approval", {}) + context.park_approval("gate", "approval", Mock(remaining_s=Mock(return_value=300))) + if background: + with pytest.raises(LifecycleUnavailable): + await context.suspend(Mock(), budget_s=1) + else: + await context.suspend(Mock(), budget_s=1) + finally: + unregister_task(context) + + +@pytest.mark.anyio +@pytest.mark.parametrize("already_started", [False, True]) +async def test_late_sdk_callback_keeps_closed_context_after_registry_removal( + monkeypatch, already_started +): + import hooks + + entered = asyncio.Event() + release = asyncio.Event() + + async def allow(*args, **kwargs): + entered.set() + await release.wait() + return {"hookSpecificOutput": {"permissionDecision": "allow"}} + + monkeypatch.setattr(hooks, "pre_tool_use_hook", allow) + monkeypatch.setattr(hooks, "log_error_cw", Mock()) + context = register_task("late-sdk-task", "microvm") + matchers = hooks.build_hook_matchers(engine=Mock(), task_id=context.task_id) + pre = matchers["PreToolUse"][0].hooks[0] + pending = None + try: + if already_started: + pending = asyncio.create_task(pre({"tool_input": {}}, "tool", {})) + await asyncio.wait_for(entered.wait(), 1) + unregister_task(context) + release.set() + result = await pending if pending else await pre({"tool_input": {}}, "tool", {}) + assert result["hookSpecificOutput"]["permissionDecision"] == "deny" + assert entered.is_set() is already_started + finally: + release.set() + unregister_task(context) + if pending: + await asyncio.gather(pending, return_exceptions=True) + + +@pytest.mark.anyio +@pytest.mark.parametrize("budget", [0, -1, float("nan"), float("inf")]) +async def test_bad_budget_does_not_close_approval_gate(budget): + context, park = await parked() + with pytest.raises(ValueError): + await context.suspend(Mock(), budget_s=budget) + await context.leave_approval(park) diff --git a/agent/tests/test_payload_bootstrap.py b/agent/tests/test_payload_bootstrap.py new file mode 100644 index 000000000..12f9b465c --- /dev/null +++ b/agent/tests/test_payload_bootstrap.py @@ -0,0 +1,367 @@ +"""Real bootstrap parsing/stream handling with AWS and HTTPS service boundaries stubbed.""" + +import hashlib +import json +import traceback +from email.message import Message +from io import BytesIO +from unittest.mock import MagicMock +from urllib.error import HTTPError +from urllib.parse import urlencode + +import pytest +from botocore.response import StreamingBody +from fastapi.testclient import TestClient + +import aws_session +import payload_bootstrap as bootstrap +import server + + +class HttpBody(BytesIO): + def __init__(self, raw: bytes, length: int): + super().__init__(raw) + self.headers = {"Content-Length": str(length)} + + +def signed_url(task_id="task-1", bucket="payload-bucket", region="us-east-1"): + suffix = "amazonaws.com.cn" if region.startswith("cn-") else "amazonaws.com" + query = urlencode( + { + "X-Amz-Algorithm": "AWS4-HMAC-SHA256", + "X-Amz-Credential": f"EXAMPLE/20260913/{region}/s3/aws4_request", + "X-Amz-Date": "20260913T120000Z", + "X-Amz-Expires": "900", + "X-Amz-SignedHeaders": "host", + "X-Amz-Signature": "a" * 64, + "X-Amz-Security-Token": "BEARER-SECRET", + } + ) + return f"https://{bucket}.s3.{region}.{suffix}/{task_id}/payload.json?{query}" + + +@pytest.fixture +def transport(monkeypatch): + config = { + "task_table_name": "Tasks", + "approval_requests_api_url": "https://approval.execute-api.us-east-1.amazonaws.com/v1", + "task_events_table_name": "Events", + "agent_session_role_arn": "arn:aws:iam::123456789012:role/AgentSession", + "github_token_secret_arn": ( + "arn:aws:secretsmanager:us-west-2:123456789012:secret:github-token-123456" + ), + } + manifest = {"version": 2, "backend": "lambda-microvm", "platform_config": config} + raw_manifest = json.dumps(manifest).encode() + key = "bootstrap/" + hashlib.sha256(raw_manifest).hexdigest() + ".json" + reference = { + "version": 2, + "task_id": "task-1", + "bootstrap_s3_uri": f"s3://payload-bucket/{key}", + "payload_url": signed_url(), + "expires_at": 1800000000000, + } + payload = { + "task_id": "task-1", + "repo_url": "org/repo", + "prompt": "do it", + "github_token": "fake", + } + document = { + "version": 2, + "task_id": "task-1", + "agent_payload": payload, + "platform_config": config, + } + streams = [] + s3 = MagicMock() + + def get_object(**kwargs): + assert kwargs == {"Bucket": "payload-bucket", "Key": key} + body = StreamingBody(BytesIO(raw_manifest), len(raw_manifest)) + streams.append(body) + return {"Body": body, "ContentLength": len(raw_manifest)} + + s3.get_object.side_effect = get_object + platform = MagicMock(return_value=s3) + monkeypatch.setattr(aws_session, "platform_client", platform) + opener = MagicMock() + + def response(_url, **kwargs): + assert kwargs == {"timeout": 10} + raw = json.dumps(document).encode() + return HttpBody(raw, len(raw)) + + opener.open.side_effect = response + build = MagicMock(return_value=opener) + monkeypatch.setattr(bootstrap, "build_opener", build) + return { + "reference": reference, + "config": config, + "payload": payload, + "document": document, + "manifest": manifest, + "s3": s3, + "platform": platform, + "opener": opener, + "build": build, + "streams": streams, + } + + +def test_authenticates_manifest_before_downloading_and_preserves_cross_region_secret(transport): + payload, config = bootstrap.resolve_payload_reference(transport["reference"], "lambda-microvm") + assert payload == transport["payload"] + assert config == transport["config"] + assert ":us-west-2:" in config["github_token_secret_arn"] + transport["platform"].assert_called_once() + assert transport["platform"].call_args.args == ("s3",) + assert len(transport["streams"]) == 1 + assert transport["streams"][0]._raw_stream.closed + assert transport["opener"].open.call_args.args == (transport["reference"]["payload_url"],) + assert isinstance(transport["build"].call_args.args[1], bootstrap._NoRedirect) + assert transport["build"].call_args.args[0].proxies == {} + + +@pytest.mark.parametrize("mutation", ["none", "url", "payload", "traversal"]) +def test_replacement_reference_binds_task_attempt_and_exact_object(transport, mutation): + reference = transport["reference"] + reference["attempt_id"] = "replacement-2" + reference["payload_url"] = signed_url("task-1/replacement-2") + transport["payload"]["attempt_id"] = "replacement-2" + if mutation == "url": + reference["payload_url"] = signed_url("task-1/replacement-other") + elif mutation == "payload": + transport["payload"]["attempt_id"] = "replacement-other" + elif mutation == "traversal": + reference["attempt_id"] = "../replacement-other" + if mutation == "none": + payload, _ = bootstrap.resolve_payload_reference(reference, "lambda-microvm") + assert payload["attempt_id"] == "replacement-2" + else: + with pytest.raises(ValueError): + bootstrap.resolve_payload_reference(reference, "lambda-microvm") + + +def test_large_registry_bundle_reaches_run_mapper_and_mcp_loader(transport, monkeypatch, tmp_path): + """Resolve real v2 bytes and carry the bundle from /run into the real local loader.""" + from registry.loader import apply_resolved_assets + + runtime = { + "transport": "http", + "url": "https://mcp.example.com/tools", + "headers": {"X-Registry-Context": "x" * 6000}, + } + assets = [ + { + "kind": "mcp_server", + "namespace": "acme", + "name": "large", + "version": "1.0.0", + "runtime": runtime, + } + ] + transport["payload"]["resolved_assets"] = assets + # Isolate real config installation; no task thread or CloudWatch client runs. + monkeypatch.setattr(server.os, "environ", dict(server.os.environ)) + monkeypatch.setattr(server, "_debug_cw", lambda *args, **kwargs: None) + spawn = MagicMock() + monkeypatch.setattr(server, "_spawn_background", spawn) + with TestClient(server.app) as client: + response = client.post( + server.MICROVM_HOOK_PREFIX + "/run", + json={ + "microvmId": "mvm-registry", + "runHookPayload": json.dumps(transport["reference"]), + }, + ) + assert response.status_code == 200 + spawn.assert_called_once() + received = spawn.call_args.args[0]["resolved_assets"] + assert received == assets + assert apply_resolved_assets(str(tmp_path), received) == ["acme__large"] + saved = json.loads((tmp_path / ".mcp.json").read_text()) + assert saved["mcpServers"]["acme__large"] == { + "type": "http", + "url": runtime["url"], + "headers": runtime["headers"], + } + + +@pytest.mark.parametrize("reference", [{}, {"agent_payload": {}}, {"version": 1}, [], "raw"]) +def test_legacy_and_malformed_envelopes_never_read_or_start(reference, transport): + with pytest.raises(ValueError, match="v2 is required"): + bootstrap.resolve_payload_reference(reference, "lambda-microvm") + transport["platform"].assert_not_called() + + +@pytest.mark.parametrize( + "uri", + [ + "http://169.254.169.254/metadata", + "s3://payload-bucket/task-2/payload.json", + "s3://payload-bucket/bootstrap/../task-2/payload.json", + "s3://payload-bucket/bootstrap/not-a-digest.json", + "s3://payload-bucket/bootstrap/" + "a" * 64 + ".json?versionId=evil", + ], +) +def test_manifest_shape_rejects_other_objects_before_aws(transport, uri): + transport["reference"]["bootstrap_s3_uri"] = uri + with pytest.raises(ValueError): + bootstrap.resolve_payload_reference(transport["reference"], "lambda-microvm") + transport["platform"].assert_not_called() + + +def test_foreign_manifest_denial_prevents_download_and_config_install( + transport, monkeypatch, capfd +): + transport["s3"].get_object.side_effect = PermissionError("AccessDenied") + install = MagicMock() + spawn = MagicMock() + monkeypatch.setattr(server, "_install_platform_config", install) + monkeypatch.setattr(server, "_spawn_background", spawn) + with TestClient(server.app) as client: + response = client.post( + server.MICROVM_HOOK_PREFIX + "/run", + json={ + "microvmId": "mvm", + "runHookPayload": json.dumps(transport["reference"]), + }, + ) + assert response.status_code == 500 + assert response.json()["code"] == "MICROVM_RUN_PAYLOAD_UNREADABLE" + install.assert_not_called() + spawn.assert_not_called() + transport["opener"].open.assert_not_called() + assert "BEARER-SECRET" not in capfd.readouterr().out + + +def test_run_hook_never_logs_the_chained_http_error_url(transport, monkeypatch, capfd): + url = transport["reference"]["payload_url"] + transport["opener"].open.side_effect = HTTPError(url, 403, url, Message(), None) + install = MagicMock() + monkeypatch.setattr(server, "_install_platform_config", install) + with TestClient(server.app) as client: + response = client.post( + server.MICROVM_HOOK_PREFIX + "/run", + json={ + "microvmId": "mvm", + "runHookPayload": json.dumps(transport["reference"]), + }, + ) + assert response.status_code == 500 + install.assert_not_called() + assert "BEARER-SECRET" not in response.text + assert "BEARER-SECRET" not in capfd.readouterr().out + + +@pytest.mark.parametrize( + "mutate", + [ + lambda d: d["platform_config"].update( + github_token_secret_arn="arn:aws:secretsmanager:us-west-2:123456789012:secret:bgagent-linear-oauth-victim" + ), + lambda d: d.update(task_id="task-2"), + lambda d: d["agent_payload"].update(task_id="task-2"), + lambda d: d.update(platform_config=None), + ], +) +def test_same_account_workspace_redirect_or_task_substitution_is_rejected(transport, mutate): + # Detach the document from the fixture's authenticated manifest values. + transport["document"]["platform_config"] = dict(transport["config"]) + mutate(transport["document"]) + with pytest.raises(ValueError): + bootstrap.resolve_payload_reference(transport["reference"], "lambda-microvm") + + +@pytest.mark.parametrize( + "url", + [ + "http://169.254.169.254/latest/meta-data", + signed_url(task_id="task-2"), + signed_url(bucket="another-bucket"), + signed_url().replace("amazonaws.com/", "amazonaws.com.evil.example/"), + signed_url().replace("https://", "https://user@"), + signed_url().replace("amazonaws.com/", "amazonaws.com:443/"), + signed_url() + "#fragment", + signed_url() + "&X-Amz-Signature=" + "b" * 64, + signed_url().replace("X-Amz-Expires=900", "X-Amz-Expires=999999"), + ], +) +def test_download_reference_rejects_wrong_task_hosts_redirect_coordinates_and_ambiguity( + transport, url +): + transport["reference"]["payload_url"] = url + with pytest.raises(ValueError, match="exact S3 object"): + bootstrap.resolve_payload_reference(transport["reference"], "lambda-microvm") + transport["opener"].open.assert_not_called() + + +@pytest.mark.parametrize( + "raw,length", + [ + (b'{"truncated":', 13), + (b"\xff", 1), + (b"[]", 2), + (b"{}", 40), + (b"{}", 0), + (b"{}", bootstrap.CONTRACT["max_payload_bytes"] + 1), + ], +) +def test_bad_payload_bytes_are_unreadable_not_bad_envelopes(transport, raw, length): + body = HttpBody(raw, length) + transport["opener"].open.side_effect = None + transport["opener"].open.return_value = body + with pytest.raises(bootstrap.PayloadFetchError): + bootstrap.resolve_payload_reference(transport["reference"], "lambda-microvm") + assert body.closed + + +@pytest.mark.parametrize("mode", ["digest", "short", "closed", "too_large", "array"]) +def test_bad_manifest_streams_fail_before_download_and_are_closed(transport, mode): + raw = b"[]" if mode == "array" else b"{}" + length = 100 if mode == "short" else len(raw) + if mode == "too_large": + length = bootstrap.CONTRACT["max_manifest_bytes"] + 1 + body = StreamingBody(BytesIO(raw), length) + if mode == "closed": + body.close() + transport["s3"].get_object.side_effect = None + transport["s3"].get_object.return_value = {"Body": body, "ContentLength": length} + with pytest.raises(bootstrap.PayloadFetchError): + bootstrap.resolve_payload_reference(transport["reference"], "lambda-microvm") + assert body._raw_stream.closed + transport["opener"].open.assert_not_called() + + +@pytest.mark.parametrize("code", [301, 307, 403, 500]) +def test_redirect_expiry_and_service_errors_do_not_echo_capabilities(transport, code): + url = transport["reference"]["payload_url"] + transport["opener"].open.side_effect = HTTPError(url, code, url, Message(), None) + with pytest.raises(bootstrap.PayloadFetchError, match=f"HTTP {code}") as error: + bootstrap.resolve_payload_reference(transport["reference"], "lambda-microvm") + assert "BEARER-SECRET" not in str(error.value) + # ECS's Python boot command can print an uncaught exception. Its formatted + # traceback must be safe too, not just the MicroVM route's JSON response. + assert "BEARER-SECRET" not in "".join(traceback.format_exception(error.value)) + assert bootstrap._NoRedirect().redirect_request(None, None, code, "", {}, url) is None + + +def test_ecs_consumes_reference_before_running_code(transport, monkeypatch): + monkeypatch.setenv("TASK_ID", "task-1") + monkeypatch.setenv("AGENT_PAYLOAD_REF", json.dumps(transport["reference"])) + resolve = MagicMock(return_value=(transport["payload"], {})) + monkeypatch.setattr(bootstrap, "resolve_payload_reference", resolve) + assert bootstrap.load_ecs_payload() == transport["payload"] + import os + + assert "AGENT_PAYLOAD_REF" not in os.environ + assert resolve.call_args.args[1] == "ecs" + + +def test_ecs_rejects_another_tasks_reference(transport, monkeypatch): + monkeypatch.setenv("TASK_ID", "task-2") + monkeypatch.setenv("AGENT_PAYLOAD_REF", json.dumps(transport["reference"])) + with pytest.raises(ValueError, match="ECS task identity"): + bootstrap.load_ecs_payload() + transport["platform"].assert_not_called() diff --git a/agent/tests/test_pipeline.py b/agent/tests/test_pipeline.py index e6b26ea6d..44f94f597 100644 --- a/agent/tests/test_pipeline.py +++ b/agent/tests/test_pipeline.py @@ -1,6 +1,7 @@ """Unit tests for pipeline.py — cedar_policies injection and pure helpers.""" import os +from contextlib import ExitStack from unittest.mock import MagicMock, patch import pytest @@ -2340,3 +2341,139 @@ def test_setup_failure_swaps_eyes_to_cross( _args, kwargs = m_finished.call_args assert kwargs.get("success") is False assert kwargs.get("started_reaction_id") == "reaction-42" + + +def test_finished_pipeline_cannot_report_success_without_committing_result(monkeypatch): + import task_state + from pipeline import _persist_finished_task + + monkeypatch.setattr( + task_state, + "write_terminal", + MagicMock(return_value=task_state.TerminalWriteOutcome.FAILED), + ) + with pytest.raises(task_state.TerminalWriteError, match="Task result was not committed"): + _persist_finished_task("task", "COMPLETED", {"status": "success"}) + + +@pytest.mark.parametrize("outcome", ["written", "disabled", "superseded"]) +def test_finished_pipeline_accepts_persistence_local_run_or_supersession(monkeypatch, outcome): + import task_state + from pipeline import _persist_finished_task + + monkeypatch.setattr( + task_state, + "write_terminal", + MagicMock(return_value=task_state.TerminalWriteOutcome(outcome)), + ) + _persist_finished_task("task", "COMPLETED", {"status": "success"}) + + +@pytest.mark.parametrize("repo_url", ["owner/repo", ""]) +@pytest.mark.parametrize("outcome", ["failed", "superseded"]) +def test_terminal_persistence_through_task_entry_point(monkeypatch, repo_url, outcome): + """Both actual pipeline paths must distinguish storage failure from cancellation.""" + import task_state + from pipeline import run_task + + monkeypatch.setenv("AWS_REGION", "us-east-1") + monkeypatch.setenv("ARTIFACTS_BUCKET_NAME", "artifacts-bkt") + + async def agent_result(*args, **kwargs): + return AgentResult( + status="success", turns=1, num_turns=1, result_text="The README describes the app." + ) + + finished = MagicMock() + terminal = MagicMock(return_value=task_state.TerminalWriteOutcome(outcome)) + with ExitStack() as stack: + for context in ( + patch("runner.run_agent", side_effect=agent_result), + patch( + "repo.setup_repo", return_value=RepoSetup(repo_dir="/workspace/repo", branch="test") + ), + patch("pipeline.task_span"), + patch("pipeline.discover_project_config"), + patch("pipeline.build_system_prompt"), + patch("pipeline.resolve_linear_api_token"), + patch("pipeline.configure_channel_mcp"), + patch("pipeline.react_task_started", return_value="started-reaction"), + patch("pipeline.comment_task_started"), + patch("pipeline.transition_task_started"), + patch("pipeline.react_task_finished", finished), + patch("pipeline.ensure_committed", return_value=False), + patch("pipeline.verify_build", return_value=VerifyOutcome(passed=True)), + patch("pipeline.verify_lint", return_value=VerifyOutcome(passed=True)), + patch("pipeline.ensure_pr", return_value="https://github.com/owner/repo/pull/1"), + patch("pipeline.get_disk_usage", return_value=0), + patch("pipeline.print_metrics"), + patch("pipeline._maybe_upload_trace", return_value=None), + patch("aws_session.tenant_client", return_value=MagicMock()), + patch.object(task_state, "write_running"), + patch.object(task_state, "write_terminal", terminal), + ): + stack.enter_context(context) + + def execute(): + return run_task( + repo_url=repo_url, + task_description="Read the README", + github_token="ghp_test", + aws_region="us-east-1", + task_id="terminal-race", + channel_source="linear", + channel_metadata=_LINEAR_META, + resolved_workflow=None + if repo_url + else {"id": "default/agent-v1", "version": "1.0.0"}, + ) + + if outcome == "failed": + with pytest.raises( + task_state.TerminalWriteError, match="Task result was not committed" + ): + execute() + else: + result = execute() + assert result["status"] == "success" + terminal.assert_called_once() + assert not any(call.kwargs.get("success") is False for call in finished.call_args_list) + assert terminal.call_args_list[0].args[1] == "COMPLETED" + + +def test_normal_terminal_writes_cannot_bypass_outcome_handling(): + """Direct best-effort writes are reserved for the existing crash handler.""" + import ast + from pathlib import Path + + import pipeline + + class TerminalCalls(ast.NodeVisitor): + function = "" + in_exception = False + + def visit_FunctionDef(self, node): + previous = self.function + self.function = node.name + self.generic_visit(node) + self.function = previous + + def visit_ExceptHandler(self, node): + previous = self.in_exception + self.in_exception = True + self.generic_visit(node) + self.in_exception = previous + + def visit_Call(self, node): + if ( + isinstance(node.func, ast.Attribute) + and isinstance(node.func.value, ast.Name) + and node.func.value.id == "task_state" + and node.func.attr == "write_terminal" + ): + assert self.function == "_persist_finished_task" or ( + self.function == "run_task" and self.in_exception + ), f"Unchecked terminal outcome at pipeline.py:{node.lineno}" + self.generic_visit(node) + + TerminalCalls().visit(ast.parse(Path(pipeline.__file__).read_text())) diff --git a/agent/tests/test_policy_three_outcome.py b/agent/tests/test_policy_three_outcome.py index 244977d4e..63b5748ef 100644 --- a/agent/tests/test_policy_three_outcome.py +++ b/agent/tests/test_policy_three_outcome.py @@ -443,19 +443,18 @@ def test_soft_deny_force_push_returns_require_approval(self): d = engine.evaluate_tool_use("Bash", {"command": "git push --force origin feature"}) assert d.outcome == Outcome.REQUIRE_APPROVAL assert "force_push_any" in d.matching_rule_ids - assert d.timeout_s == 300 + assert d.timeout_s == 0 assert d.severity == "medium" def test_soft_deny_multi_match_merges_annotations(self): - # force_push_any (300s, medium) + force_push_main (600s, high) both - # match "git push --force origin main". Merge picks min(300, 600)=300s - # and max(medium, high)=high. §6.3. + # Both built-in rules have no automatic expiry. Their merged severity + # is max(medium, high)=high. engine = PolicyEngine(task_type="new_task", repo="owner/repo") d = engine.evaluate_tool_use("Bash", {"command": "git push --force origin main"}) assert d.outcome == Outcome.REQUIRE_APPROVAL assert "force_push_any" in d.matching_rule_ids assert "force_push_main" in d.matching_rule_ids - assert d.timeout_s == 300 # min across rules + task default + assert d.timeout_s == 0 # neither rule supplies a positive deadline assert d.severity == "high" # max across rules def test_default_allow_on_no_match(self): @@ -543,7 +542,7 @@ def test_denied_then_retry_hits_cache(self): engine.recent_decisions.record("Bash", sha, "DENIED", "user said force-push is too risky") d = engine.evaluate_tool_use("Bash", tool_input) assert d.outcome == Outcome.DENY - assert "Recent DENIED" in d.reason + assert "Recorded DENIED" in d.reason def test_cache_does_not_shadow_hard_deny(self): engine = PolicyEngine(task_type="new_task", repo="owner/repo") @@ -605,7 +604,7 @@ def test_semantic_retry_hits_rule_cache(self): "Bash", {"command": "git push --force origin some-other-branch"} ) assert d.outcome == Outcome.DENY - assert "Recent DENIED on rule 'force_push_any'" in d.reason + assert "Recorded DENIED on rule 'force_push_any'" in d.reason assert d.cache_hit_metadata is not None assert d.cache_hit_metadata["matched_rule_id"] == "force_push_any" assert d.cache_hit_metadata["cached_decision"] == "DENIED" diff --git a/agent/tests/test_poll_for_decision.py b/agent/tests/test_poll_for_decision.py index 67739b708..8845d7564 100644 --- a/agent/tests/test_poll_for_decision.py +++ b/agent/tests/test_poll_for_decision.py @@ -28,8 +28,7 @@ def test_n_consecutive_failures_returns_timed_out_with_reason(self, monkeypatch) without further polling. """ # Tiny intervals so the loop iterates fast but doesn't immediately - # bail via the ``sleep_for <= 0`` early-return on line 723 of - # hooks.py. + # bail via the ``sleep_for <= 0`` early return. monkeypatch.setattr(hooks, "POLL_FAST_INTERVAL_S", 0.001) monkeypatch.setattr(hooks, "POLL_FAST_DURATION_S", 0.001) monkeypatch.setattr(hooks, "POLL_SLOW_INTERVAL_S", 0.001) @@ -43,7 +42,7 @@ def test_n_consecutive_failures_returns_timed_out_with_reason(self, monkeypatch) hooks._poll_for_decision( task_id="01KTASK", request_id="01KREQ", - timeout_s=300, # large enough that the deadline doesn't fire first + deadline=hooks._ApprovalDeadline.from_recorded(hooks._iso_now(), 300), progress=progress, ts=ts, ) @@ -59,8 +58,7 @@ def test_degraded_emitted_once_at_threshold(self, monkeypatch): on every subsequent poll — IMPL-22 / §13.2. """ # Tiny intervals so the loop iterates fast but doesn't immediately - # bail via the ``sleep_for <= 0`` early-return on line 723 of - # hooks.py. + # bail via the ``sleep_for <= 0`` early return. monkeypatch.setattr(hooks, "POLL_FAST_INTERVAL_S", 0.001) monkeypatch.setattr(hooks, "POLL_FAST_DURATION_S", 0.001) monkeypatch.setattr(hooks, "POLL_SLOW_INTERVAL_S", 0.001) @@ -73,7 +71,7 @@ def test_degraded_emitted_once_at_threshold(self, monkeypatch): hooks._poll_for_decision( task_id="01KTASK", request_id="01KREQ", - timeout_s=300, + deadline=hooks._ApprovalDeadline.from_recorded(hooks._iso_now(), 300), progress=progress, ts=ts, ) @@ -100,8 +98,7 @@ def test_recovery_resets_failure_counter(self, monkeypatch): outage. """ # Tiny intervals so the loop iterates fast but doesn't immediately - # bail via the ``sleep_for <= 0`` early-return on line 723 of - # hooks.py. + # bail via the ``sleep_for <= 0`` early return. monkeypatch.setattr(hooks, "POLL_FAST_INTERVAL_S", 0.001) monkeypatch.setattr(hooks, "POLL_FAST_DURATION_S", 0.001) monkeypatch.setattr(hooks, "POLL_SLOW_INTERVAL_S", 0.001) @@ -126,7 +123,7 @@ def test_recovery_resets_failure_counter(self, monkeypatch): hooks._poll_for_decision( task_id="01KTASK", request_id="01KREQ", - timeout_s=300, + deadline=hooks._ApprovalDeadline.from_recorded(hooks._iso_now(), 300), progress=progress, ts=ts, ) @@ -143,8 +140,7 @@ def test_deadline_beats_failures_when_timeout_short(self, monkeypatch): IMPL-24 design. """ # Tiny intervals so the loop iterates fast but doesn't immediately - # bail via the ``sleep_for <= 0`` early-return on line 723 of - # hooks.py. + # bail via the ``sleep_for <= 0`` early return. monkeypatch.setattr(hooks, "POLL_FAST_INTERVAL_S", 0.001) monkeypatch.setattr(hooks, "POLL_FAST_DURATION_S", 0.001) monkeypatch.setattr(hooks, "POLL_SLOW_INTERVAL_S", 0.001) @@ -153,12 +149,12 @@ def test_deadline_beats_failures_when_timeout_short(self, monkeypatch): ts.get_approval_row.return_value = {"status": "PENDING"} progress = MagicMock() - # 0 timeout → loop returns immediately at the deadline check. + # An already elapsed explicit deadline returns immediately. outcome = _run( hooks._poll_for_decision( task_id="01KTASK", request_id="01KREQ", - timeout_s=0, + deadline=hooks._ApprovalDeadline(0, 0), progress=progress, ts=ts, ) diff --git a/agent/tests/test_runner.py b/agent/tests/test_runner.py index ffeeb87c4..0a9872eba 100644 --- a/agent/tests/test_runner.py +++ b/agent/tests/test_runner.py @@ -1,7 +1,7 @@ """Unit tests for runner.py helpers. -The full ``run_agent`` path is integration-tested via test_pipeline.py -with a mocked ``pipeline.run_agent``. This module covers the narrower +Pipeline tests mock ``pipeline.run_agent``; they do not exercise its SDK loop. +This module covers client/broker ownership and the narrower ``_initialize_policy_engine_and_hooks`` helper extracted in Chunk 7 so the policy-engine bootstrap + ``pre_approvals_loaded`` emission can be verified without spinning up the Claude Agent SDK client. @@ -12,7 +12,7 @@ import asyncio import subprocess from typing import Any -from unittest.mock import MagicMock, patch +from unittest.mock import AsyncMock, MagicMock, patch import pytest @@ -48,6 +48,165 @@ def _config(**overrides: Any) -> TaskConfig: return TaskConfig(**base) +class TestClaudeSessionOwnership: + @pytest.mark.parametrize("microvm", [False, True]) + @pytest.mark.parametrize( + "failure", [None, "connect", "query", "receive", "cancel", "hook-denied"] + ) + def test_broker_selection_and_cleanup_on_every_session_exit( + self, monkeypatch, microvm, failure + ): + import claude_agent_sdk + + import microvm_credentials + import microvm_lifecycle + + config = _config() + context = microvm_lifecycle.register_task(config.task_id, "vm") if microvm else None + client = MagicMock() + client.connect = AsyncMock() + client.query = AsyncMock() + client.disconnect = AsyncMock() + if failure in {"connect", "query"}: + getattr(client, failure).side_effect = RuntimeError("synthetic failure") + + async def messages(): + if failure == "receive": + raise RuntimeError("synthetic receive failure") + if failure == "cancel": + raise asyncio.CancelledError + if failure == "hook-denied": + if context is not None: + await context.tool_started("denied-call") + await context.tool_started("other-active-call") + yield claude_agent_sdk.UserMessage( + content=[ + claude_agent_sdk.ToolResultBlock( + tool_use_id="denied-call", + content="Project hook denied", + is_error=True, + ), + ] + ) + if context is not None: + assert context.diagnostic_snapshot()["active_tools"] == 1 + assert "other-active-call" in context._tools + yield claude_agent_sdk.ResultMessage( + subtype="success", + duration_ms=1, + duration_api_ms=1, + is_error=False, + num_turns=0, + session_id="synthetic", + total_cost_usd=0, + usage={}, + ) + + client.receive_response = messages + make_client = MagicMock(return_value=client) + monkeypatch.setattr(claude_agent_sdk, "ClaudeSDKClient", make_client) + broker = MagicMock() + broker.environment = {"ABCA_MICROVM_CREDENTIAL_BROKER": "1"} + make_broker = MagicMock(return_value=broker) + monkeypatch.setattr(microvm_credentials, "ScopedCredentialBroker", make_broker) + monkeypatch.setattr(runner, "_setup_agent_env", lambda _config: None) + monkeypatch.setattr(runner, "_log_claude_cli_version", lambda: None) + monkeypatch.setattr(runner, "_initialize_policy_engine_and_hooks", lambda **_kw: (None, {})) + monkeypatch.setattr(runner, "_register_gateway_server", lambda _servers: None) + monkeypatch.setattr(runner, "build_clarification_server", lambda: None) + monkeypatch.setattr(runner, "_ProgressWriter", MagicMock()) + monkeypatch.setattr(runner, "log_error_cw", MagicMock()) + try: + if failure == "cancel": + with pytest.raises(asyncio.CancelledError): + asyncio.run(runner.run_agent("probe", "probe", config, trajectory=MagicMock())) + else: + result = asyncio.run( + runner.run_agent("probe", "probe", config, trajectory=MagicMock()) + ) + expected = "error" if failure not in {None, "hook-denied"} else "success" + assert result.status == expected + options = make_client.call_args.kwargs["options"] + if microvm: + make_broker.assert_called_once_with(context) + assert options.env == broker.environment + broker.close.assert_called_once() + else: + make_broker.assert_not_called() + assert options.env == {} + client.disconnect.assert_awaited_once() + finally: + if context is not None: + microvm_lifecycle.unregister_task(context) + + +@pytest.mark.parametrize("exhausted", [None, "dollars", "turns"]) +def test_replacement_runner_preserves_total_limits_and_reports_cumulative_usage( + monkeypatch, exhausted +): + from types import SimpleNamespace + + import claude_agent_sdk + + from continuation_runtime import bind_runtime + + config = _config(max_turns=10, max_budget_usd=1.0) + runtime: Any = SimpleNamespace( + store=MagicMock(), + restored={"session_id": "saved-session"}, + context=SimpleNamespace(turns_used=10 if exhausted == "turns" else 4), + prior_cost_usd=1.0 if exhausted == "dollars" else 0.25, + prior_token_usage={"input_tokens": 100, "output_tokens": 10}, + client=None, + ) + client = MagicMock() + client.connect = AsyncMock() + client.query = AsyncMock() + client.disconnect = AsyncMock() + + async def messages(): + yield claude_agent_sdk.ResultMessage( + subtype="success", + duration_ms=1, + duration_api_ms=1, + is_error=False, + num_turns=2, + session_id="saved-session", + total_cost_usd=0.1, + usage={"input_tokens": 50, "output_tokens": 5}, + ) + + client.receive_response = messages + make_client = MagicMock(return_value=client) + monkeypatch.setattr(claude_agent_sdk, "ClaudeSDKClient", make_client) + monkeypatch.setattr(runner, "_setup_agent_env", lambda _: None) + monkeypatch.setattr(runner, "_log_claude_cli_version", lambda: None) + monkeypatch.setattr(runner, "_initialize_policy_engine_and_hooks", lambda **_: (None, {})) + monkeypatch.setattr(runner, "_register_gateway_server", lambda _: None) + monkeypatch.setattr(runner, "build_clarification_server", lambda: None) + monkeypatch.setattr(runner, "_ProgressWriter", MagicMock()) + with bind_runtime(runtime): + if exhausted: + with pytest.raises(RuntimeError, match="before the saved continuation"): + asyncio.run( + runner.run_agent("continue", "saved system", config, trajectory=MagicMock()) + ) + make_client.assert_not_called() + return + result = asyncio.run( + runner.run_agent("continue", "saved system", config, trajectory=MagicMock()) + ) + options = make_client.call_args.kwargs["options"] + assert options.resume == "saved-session" + assert options.max_turns == 6 + assert options.max_budget_usd == pytest.approx(0.75) + assert result.cost_usd == pytest.approx(0.35) + assert result.num_turns == 6 + assert result.usage is not None + assert result.usage.input_tokens == 150 + assert result.usage.output_tokens == 15 + + class TestInitializePolicyEngineAndHooks: """Bootstrap the per-task PolicyEngine + hooks without the SDK loop. diff --git a/agent/tests/test_server.py b/agent/tests/test_server.py index 067e01c6a..9350ef1cc 100644 --- a/agent/tests/test_server.py +++ b/agent/tests/test_server.py @@ -9,9 +9,8 @@ import threading import time from pathlib import Path -from types import SimpleNamespace from typing import Any -from unittest.mock import MagicMock +from unittest.mock import MagicMock, patch import pytest from fastapi.testclient import TestClient @@ -19,35 +18,48 @@ import server -@pytest.fixture(autouse=True) -def reset_server_state(): - """Reset the pipeline registry, joining any thread still running on the way out. - - `/run` and `/invocations` answer while the pipeline thread is only just starting, - so that thread usually looks up `server.run_task` AFTER the test body has - returned. Clearing the registry without joining orphans it: the stubs are then - undone, and the thread goes on to run a REAL pipeline — task-state writes, a - heartbeat, a clone — inside whichever test happens to be running next. That is a - cross-test AWS call arriving from a thread nothing is waiting on, and it stays - invisible until some later test asserts that no AWS seam was touched. Joining - here is what keeps a pipeline thread from outliving the test that spawned it. - - The live threads are read from `threading.enumerate()` rather than from - `_active_threads`, because a test may substitute that registry with one that - refuses to be read; it stays clearable, which is all this fixture asks of it. - Joining the pipeline thread also retires its heartbeat, which the pipeline stops - on its way out. - """ - server._background_pipeline_failed = False +def _join_server_threads(timeout: float = 5.0) -> None: + """Reap tracked work before restoring test state; retain leaked handles on failure.""" + deadline = time.monotonic() + timeout + with server._threads_lock: + threads = list(server._active_threads) + # The pipeline may need _threads_lock to finish; never join while holding it. + for thread in threads: + if thread.is_alive(): + thread.join(timeout=max(0.0, deadline - time.monotonic())) with server._threads_lock: + leaked = [thread.name for thread in server._active_threads if thread.is_alive()] + if leaked: + pytest.fail(f"Server pipeline threads did not exit within {timeout}s: {leaked}") server._active_threads.clear() - yield - for thread in threading.enumerate(): - if thread.name.startswith("pipeline-") and thread.is_alive(): - thread.join(timeout=10) + + +@pytest.fixture(autouse=True) +def reset_server_state(monkeypatch, env_guard): + # Dependencies force this teardown to precede mock/environment restoration. + # A still-starting pipeline resolves run_task from the module at call time. + _join_server_threads() server._background_pipeline_failed = False + try: + yield + finally: + _join_server_threads() + server._background_pipeline_failed = False + + +def test_server_thread_cleanup_keeps_leaked_handles_and_reports_their_names(): + release = threading.Event() + thread = threading.Thread(target=release.wait, name="deliberately-blocked-pipeline") + thread.start() with server._threads_lock: - server._active_threads.clear() + server._active_threads.append(thread) + try: + with pytest.raises(pytest.fail.Exception, match="deliberately-blocked-pipeline"): + _join_server_threads(timeout=0) + assert thread in server._active_threads + finally: + release.set() + thread.join(timeout=5) @pytest.fixture @@ -61,11 +73,12 @@ def test_ping_healthy_by_default(client): assert r.json() == {"status": "healthy"} -def test_background_thread_failure_503_and_backup_terminal_write(client, monkeypatch): +@pytest.mark.parametrize("outcome", ["written", "failed", "superseded"]) +def test_background_thread_failure_503_and_backup_terminal_write(client, monkeypatch, outcome): def boom(**_kwargs): raise RuntimeError("simulated pipeline crash") - mock_write = MagicMock() + mock_write = MagicMock(return_value=server.task_state.TerminalWriteOutcome(outcome)) monkeypatch.setattr(server, "run_task", boom) monkeypatch.setattr(server.task_state, "write_terminal", mock_write) @@ -103,13 +116,8 @@ def boom(**_kwargs): assert body["status"] == "unhealthy" assert body["reason"] == "background_pipeline_failed" - # Race: /ping flips to 503 as soon as ``_background_pipeline_failed = True`` - # is set in the except block, but ``task_state.write_terminal(...)`` happens - # a few lines later (after ``print()`` + ``traceback.print_exc()``). Wait - # for the mock to actually be invoked before asserting. - deadline2 = time.time() + 5.0 - while time.time() < deadline2 and not mock_write.called: - time.sleep(0.05) + # The thread has exited, including the backup write. Even a failed backup + # must leave /ping unhealthy so the coordinator can recover the task. mock_write.assert_called() call_kw = mock_write.call_args assert call_kw[0][0] == "task-crash-1" @@ -480,36 +488,47 @@ def test_debug_cw_exc_appends_the_traceback(monkeypatch, capfd): assert "Traceback" in out -def test_debug_cw_write_blocking_bumps_failure_counter_on_boto_error(monkeypatch): - """On boto errors the failure counter increments so operators can alarm. - - AgentCore doesn't forward container stdout to APPLICATION_LOGS, so a - broken ``_debug_cw`` is invisible except for this counter. If the - counter ever stops bumping on error the blind-debug alarm breaks - silently. - """ - # Seed the counter to a known value so we can assert the delta without - # being sensitive to other tests. - with server._debug_cw_failures_lock: - server._debug_cw_failures = 0 - - # Stub ``boto3.client`` to raise so the except branch (which bumps - # the counter) runs. - class _BrokenBoto3: - @staticmethod - def client(*args, **kwargs): - raise RuntimeError("simulated boto failure") +@pytest.mark.parametrize("writer", ["debug", "warn"]) +@pytest.mark.parametrize("stage", ["client", "stream", "events"]) +def test_cloudwatch_failures_emit_structured_stdout_without_recursion( + writer, stage, monkeypatch, capfd +): + """The fallback survives a broken writer without exposing the failed log text.""" + import aws_session + + class StreamExists(Exception): + pass + + logs = MagicMock() + logs.exceptions.ResourceAlreadyExistsException = StreamExists + failure = RuntimeError("BEARER-SECRET in SDK error") + factory = MagicMock(return_value=logs) + if stage == "client": + factory.side_effect = failure + elif stage == "stream": + logs.create_log_stream.side_effect = failure + else: + logs.put_log_events.side_effect = failure + monkeypatch.setattr(aws_session, "platform_client", factory) - monkeypatch.setitem(__import__("sys").modules, "boto3", _BrokenBoto3) + def forbidden(*args, **kwargs): + pytest.fail("the fallback must not call a CloudWatch writer") - server._debug_cw_write_blocking( - log_group="/some/log-group", - task_id="t-1", - stamped="2026-01-01T00:00:00Z hello", + monkeypatch.setattr(server, "_debug_cw", forbidden) + monkeypatch.setattr(server, "_warn_cw", forbidden) + getattr(server, f"_{writer}_cw_write_blocking")( + log_group="/test/logs", task_id="task-log-failure", stamped="PRIVATE-TASK-PROMPT" ) - - with server._debug_cw_failures_lock: - assert server._debug_cw_failures == 1 + factory.assert_called_once() + output = capfd.readouterr().out + assert json.loads(output) == { + "event": "cloudwatch_write_failed", + "writer": writer, + "task_id": "task-log-failure", + "error_type": "RuntimeError", + } + assert "BEARER-SECRET" not in output + assert "PRIVATE-TASK-PROMPT" not in output # Chunk 7c — _warn_cw parallels _debug_cw so warn-level invocation-payload @@ -566,33 +585,6 @@ def start(self) -> None: ) -def test_warn_cw_write_blocking_bumps_failure_counter_on_boto_error(monkeypatch): - """Warn-path boto errors bump the same failure counter as debug. - - A single alarm surface is intentional (§server.py comment on - ``_debug_cw_failures``). If the counter ever stops bumping on a - warn write failure the blind-warn alarm breaks silently. - """ - with server._debug_cw_failures_lock: - server._debug_cw_failures = 0 - - class _BrokenBoto3: - @staticmethod - def client(*args, **kwargs): - raise RuntimeError("simulated boto failure") - - monkeypatch.setitem(__import__("sys").modules, "boto3", _BrokenBoto3) - - server._warn_cw_write_blocking( - log_group="/some/log-group", - task_id="t-1", - stamped="[server/warn] malformed payload", - ) - - with server._debug_cw_failures_lock: - assert server._debug_cw_failures == 1 - - def test_warn_cw_write_blocking_uses_server_warn_stream(monkeypatch): """Warn writes land in ``server_warn/``, not the debug stream. @@ -612,12 +604,9 @@ def create_log_stream(self, *, logGroupName, logStreamName): def put_log_events(self, *, logGroupName, logStreamName, logEvents): captured_streams.append(logStreamName) - class _FakeBoto3: - @staticmethod - def client(*args, **kwargs): - return _FakeLogs() - - monkeypatch.setitem(__import__("sys").modules, "boto3", _FakeBoto3) + # This test owns stream routing. Credential/signing behavior belongs to the + # aws_session tests, so stub the attributed client factory at its boundary. + monkeypatch.setattr("aws_session.platform_client", lambda *_args, **_kwargs: _FakeLogs()) server._warn_cw_write_blocking( log_group="/some/log-group", @@ -926,13 +915,45 @@ def _platform_config(**overrides) -> dict: return config -def _run_hook_body(envelope: dict, microvm_id: str = "microvm-abc") -> dict: - """Wrap an ABCA payload envelope in the service's ``/run`` request body. +# Route tests isolate the authenticated transport boundary. The real manifest +# IAM/stream/URL consumer is exercised in test_payload_bootstrap.py. +_run_documents: dict[str, tuple[dict, Any]] = {} - The service passes ``runHookPayload`` through as an opaque STRING (it never - parses it), so the double encoding here is the real wire shape, not a test - artifact. - """ + +@pytest.fixture(autouse=True) +def _authenticated_test_transport(monkeypatch): + import payload_bootstrap + + _run_documents.clear() + monkeypatch.setattr( + payload_bootstrap, + "_manifest", + lambda uri, backend: ("payload-bucket", _run_documents[uri][1]), + ) + monkeypatch.setattr(payload_bootstrap, "_download", lambda url: _run_documents[url][0]) + yield + _run_documents.clear() + + +def _run_hook_body(envelope: dict, microvm_id: str = "microvm-abc") -> dict: + """Wrap pipeline/config fixtures in the current v2 service envelope.""" + from tests.test_payload_bootstrap import signed_url + + if "agent_payload" in envelope and isinstance(envelope["agent_payload"], dict): + payload = envelope["agent_payload"] + task_id = payload.get("task_id", "invalid") + config = envelope.get("platform_config", _platform_config()) + uri = f"s3://payload-bucket/bootstrap/{'a' * 64}.json" + url = signed_url(task_id=task_id) + document = { + "version": 2, + "task_id": task_id, + "agent_payload": payload, + "platform_config": config, + } + _run_documents[uri] = (document, config) + _run_documents[url] = (document, config) + envelope = {"version": 2, "task_id": task_id, "bootstrap_s3_uri": uri, "payload_url": url} return {"microvmId": microvm_id, "runHookPayload": json.dumps(envelope)} @@ -987,14 +1008,13 @@ def test_ready_does_not_start_a_pipeline(self, client, monkeypatch, warm_ready): with server._threads_lock: assert server._active_threads == [] - def test_suspend_and_resume_are_NOT_served(self, client): - # Declaring a hook nothing answers fails the corresponding build or - # lifecycle transition, so the construct declares exactly the hooks the - # agent serves. /validate + /terminate joined that set in P2; /suspend + - # /resume need the ComputeStrategy interface widening (P3), so they must - # still 404 — the assertion that keeps the construct honest. + def test_suspend_and_resume_require_a_registered_task(self, client): + # Runtime routes exist before the image declares the P3 capability. + # Snapshot warm-up has no task and cannot authorize a lifecycle change. for hook in ("suspend", "resume"): - assert client.post(f"{server.MICROVM_HOOK_PREFIX}/{hook}").status_code == 404 + response = client.post(f"{server.MICROVM_HOOK_PREFIX}/{hook}") + assert response.status_code == 409 + assert response.json()["code"] == "MICROVM_LIFECYCLE_UNAVAILABLE" class TestMicrovmReadyHookWarmUp: @@ -1119,10 +1139,9 @@ def exploding(argv, **kwargs): def test_the_warm_up_makes_no_aws_call_even_with_a_log_group_baked( self, client, monkeypatch, capfd, warm_ready ): - # /ready runs under the BUILD role: a Logs write can only fail (and each - # failure pollutes the shared _debug_cw_failures alarm), and any boto3 - # client built here freezes the build role's credential chain and the build - # region into the snapshot. Adding a subprocess must not have changed that. + # The BUILD role cannot write application logs outside its MicroVM log + # namespace. Initializing AWS clients here can preserve build credentials + # in the snapshot. Warm-up subprocesses must not change the stdout-only rule. monkeypatch.setenv("LOG_GROUP_NAME", "/abca/agent") def forbidden(*_args, **_kwargs): @@ -1320,21 +1339,44 @@ def fake_run(argv, **kwargs): assert "skipping best-effort warm-up of 'opt'" in capfd.readouterr().out -class TestMicrovmRunHookInlinePayload: - """Inline envelope: ``{"agent_payload": {...}}``. +class TestMicrovmRunHookVerifiedPayload: + """Map an authenticated v2 payload to the asynchronous task pipeline. - The exception rather than the rule — the service caps ``runHookPayload`` at - 4 096 bytes and a hydrated payload is larger — but it is the branch that - proves the payload→pipeline mapping without any S3 involvement. + The module fixture supplies manifest/download bytes; the shared consumer's + authentication and transport checks have their own regression suite. """ + @pytest.mark.parametrize("attempt", ["replacement-2", "../other", "", 123]) + def test_attempt_identity_is_validated_before_starting( + self, client, monkeypatch, cached_github_token, attempt + ): + spawn = MagicMock() + monkeypatch.setattr(server, "_spawn_background", spawn) + response = client.post( + RUN_HOOK, + json=_run_hook_body( + { + "agent_payload": { + "task_id": "task-attempt", + "repo_url": "org/repo", + "prompt": "Continue", + "github_token": "ghp_x", + "attempt_id": attempt, + }, + } + ), + ) + if attempt == "replacement-2": + assert response.status_code == 200 + assert spawn.call_args.args[0]["attempt_id"] == "replacement-2" + else: + assert response.status_code == 400 + assert response.json()["code"] == "MICROVM_ATTEMPT_ID_INVALID" + spawn.assert_not_called() + def test_accepts_the_payload_and_starts_the_pipeline_asynchronously( - self, client, monkeypatch, baked_platform_env + self, client, monkeypatch, cached_github_token ): - # `baked_platform_env`: this envelope carries no `platform_config`, which - # since review N2 requires the effective env to supply the required values - # (a legacy image that bakes its own). This class is about the - # payload->pipeline mapping, not about config delivery. started = threading.Event() seen: dict = {} @@ -1357,7 +1399,7 @@ def fake_run_task(**kwargs): "aws_region": "us-east-1", } }, - microvm_id="microvm-inline", + microvm_id="microvm-verified", ), ) @@ -1366,7 +1408,7 @@ def fake_run_task(**kwargs): assert body["status"] == "accepted" assert body["task_id"] == "t-microvm-1" # Echoed so a MicroVM log line can be joined to the control-plane id. - assert body["microvm_id"] == "microvm-inline" + assert body["microvm_id"] == "microvm-verified" assert started.wait(timeout=5.0), "pipeline thread did not start" # Same mapping the /invocations path performs: prompt→task_description, @@ -1375,11 +1417,7 @@ def fake_run_task(**kwargs): assert seen["repo_url"] == "org/repo" assert seen["task_description"] == "Fix the bug" - def test_returns_before_the_pipeline_finishes(self, client, monkeypatch, baked_platform_env): - # `baked_platform_env`: this envelope carries no `platform_config`, which - # since review N2 requires the effective env to supply the required values - # (a legacy image that bakes its own). This class is about the - # payload->pipeline mapping, not about config delivery. + def test_returns_before_the_pipeline_finishes(self, client, monkeypatch, cached_github_token): release = threading.Event() entered = threading.Event() @@ -1407,12 +1445,8 @@ def slow_run_task(**_kwargs): release.set() def test_uses_the_same_model_id_and_prompt_aliases_as_invocations( - self, client, monkeypatch, baked_platform_env + self, client, monkeypatch, cached_github_token ): - # `baked_platform_env`: this envelope carries no `platform_config`, which - # since review N2 requires the effective env to supply the required values - # (a legacy image that bakes its own). This class is about the - # payload->pipeline mapping, not about config delivery. seen: dict = {} started = threading.Event() @@ -1444,118 +1478,6 @@ def fake_run_task(**kwargs): assert seen["channel_source"] == "linear" -class TestMicrovmRunHookS3Payload: - """S3-pointer envelope: ``{"agent_payload_s3_uri": "s3://bucket/key"}``. - - The DOMINANT path on this backend: with a 4 096-byte ``runHookPayload`` cap, - any hydrated payload is offloaded to the platform payload bucket and only the - pointer travels in the hook body. - """ - - def test_fetches_the_payload_from_s3_and_starts_the_pipeline( - self, client, monkeypatch, baked_platform_env - ): - # `baked_platform_env`: this envelope carries no `platform_config`, which - # since review N2 requires the effective env to supply the required values - # (a legacy image that bakes its own). This class is about the - # payload->pipeline mapping, not about config delivery. - seen: dict = {} - started = threading.Event() - - def fake_run_task(**kwargs): - seen.update(kwargs) - started.set() - - fetched: dict = {} - - def fake_fetch(uri): - fetched["uri"] = uri - return {"task_id": "t-s3", "repo_url": "org/repo", "prompt": "from s3"} - - monkeypatch.setattr(server, "_fetch_microvm_payload_from_s3", fake_fetch) - monkeypatch.setattr(server, "run_task", fake_run_task) - monkeypatch.setattr(server.task_state, "write_terminal", MagicMock()) - - r = client.post( - RUN_HOOK, - json=_run_hook_body({"agent_payload_s3_uri": "s3://payload-bucket/t-s3/payload.json"}), - ) - - assert r.status_code == 200 - assert r.json()["task_id"] == "t-s3" - assert fetched["uri"] == "s3://payload-bucket/t-s3/payload.json" - assert started.wait(timeout=5.0) - assert seen["task_description"] == "from s3" - - def test_parses_bucket_and_key_out_of_the_uri(self, monkeypatch): - captured: dict = {} - - class _Body: - @staticmethod - def read(): - return b'{"task_id": "t-1", "repo_url": "o/r"}' - - class _S3: - @staticmethod - def get_object(**kwargs): - captured.update(kwargs) - return {"Body": _Body} - - import boto3 - - monkeypatch.setattr(boto3, "client", lambda *_a, **_k: _S3) - - payload = server._fetch_microvm_payload_from_s3("s3://my-bucket/prefix/t-1/payload.json") - - # Key keeps every slash after the bucket — a naive split would truncate it. - assert captured == {"Bucket": "my-bucket", "Key": "prefix/t-1/payload.json"} - assert payload == {"task_id": "t-1", "repo_url": "o/r"} - - def test_rejects_a_uri_with_no_key(self, monkeypatch): - with pytest.raises(ValueError, match="not a bucket/key URI"): - server._fetch_microvm_payload_from_s3("s3://bucket-only") - - def test_rejects_a_non_object_s3_body(self, monkeypatch): - class _Body: - @staticmethod - def read(): - return b"[1, 2, 3]" - - class _S3: - @staticmethod - def get_object(**_kwargs): - return {"Body": _Body} - - import boto3 - - monkeypatch.setattr(boto3, "client", lambda *_a, **_k: _S3) - - # `_PayloadFetchError`, NOT `ValueError` (review N1): the object is bad, not - # the orchestrator's envelope, so the handler must route it to its RETRYABLE - # 500 rather than the "retrying cannot help" 400. Asserting the type is the - # point — `ValueError` here would silently restore the misclassification. - with pytest.raises(server._PayloadFetchError, match="expected an object"): - server._fetch_microvm_payload_from_s3("s3://b/k") - assert not issubclass(server._PayloadFetchError, ValueError) - - def test_s3_failure_returns_500_and_starts_nothing(self, client, monkeypatch): - def boom(_uri): - raise RuntimeError("AccessDenied") - - monkeypatch.setattr(server, "_fetch_microvm_payload_from_s3", boom) - monkeypatch.setattr(server, "run_task", MagicMock()) - - r = client.post(RUN_HOOK, json=_run_hook_body({"agent_payload_s3_uri": "s3://bucket/key"})) - - # 500, not 400: the body was well-formed, the fetch was not. Retrying an - # identical body CAN help here, unlike a malformed envelope. - assert r.status_code == 500 - assert r.json()["code"] == "MICROVM_RUN_PAYLOAD_UNREADABLE" - assert "AccessDenied" in r.json()["message"] - with server._threads_lock: - assert server._active_threads == [] - - class TestMicrovmRunHookRejections: """Every shape the agent cannot act on must fail LOUDLY, before spawning. @@ -1566,15 +1488,15 @@ class TestMicrovmRunHookRejections: @pytest.mark.parametrize( "run_hook_payload,expected_fragment", [ - ("", "runHookPayload is empty"), - (" ", "runHookPayload is empty"), + ("", "not valid JSON"), + (" ", "not valid JSON"), ("not json at all", "not valid JSON"), - ('"a string"', "must be a JSON object"), - ("[1,2,3]", "must be a JSON object"), - ('{"agent_payload": "not-an-object"}', "agent_payload must be an object"), - ('{"agent_payload_s3_uri": "https://example.com/x"}', "must be an s3:// URI"), - ('{"agent_payload_s3_uri": 42}', "must be an s3:// URI"), - ('{"something_else": 1}', "neither agent_payload nor agent_payload_s3_uri"), + ('"a string"', "v2 is required"), + ("[1,2,3]", "v2 is required"), + ('{"agent_payload": "not-an-object"}', "v2 is required"), + ('{"agent_payload_s3_uri": "https://example.com/x"}', "v2 is required"), + ('{"agent_payload_s3_uri": 42}', "v2 is required"), + ('{"something_else": 1}', "v2 is required"), ], ) def test_returns_400_with_a_named_code( @@ -1603,9 +1525,9 @@ def test_a_missing_body_field_is_a_400_not_a_422(self, client, monkeypatch): assert r.json()["code"] == "MICROVM_RUN_PAYLOAD_INVALID" def test_incomplete_task_record_reuses_the_invocations_rejection_shape( - self, client, monkeypatch, baked_platform_env + self, client, monkeypatch, cached_github_token ): - # `baked_platform_env` so the run gets PAST config delivery and reaches the + # Cache the GitHub token so this test reaches the # task-record check this test is actually about. monkeypatch.setattr(server, "run_task", MagicMock()) @@ -1633,12 +1555,8 @@ class TestMicrovmRunHookHeaderPosture: """ def test_session_id_and_workload_token_resolve_empty( - self, client, monkeypatch, baked_platform_env + self, client, monkeypatch, cached_github_token ): - # `baked_platform_env`: this envelope carries no `platform_config`, which - # since review N2 requires the effective env to supply the required values - # (a legacy image that bakes its own). This class is about the - # payload->pipeline mapping, not about config delivery. seen: dict = {} started = threading.Event() @@ -1744,31 +1662,14 @@ def test_merge_branches_non_string_entries_filtered(self): @pytest.fixture -def baked_platform_env(env_guard): - """Simulate a legacy/hand-built image that BAKES its own required config. - - Needed by every ``/run`` test whose envelope carries no ``platform_config``. - Since review N2 the no-config branch re-runs the required-key check against the - EFFECTIVE environment and rejects when it is unsatisfied — because - ``aws_session`` silently drops tenant scoping when ``AGENT_SESSION_ROLE_ARN`` - is unset, and running a task unscoped is worse than refusing it. That check is - what this fixture satisfies, and satisfying it is exactly what a real legacy - image does: the compatibility path is "the snapshot supplies the values", not - "nobody supplies them". - - Depends on ``env_guard`` so the writes are reverted with everything else. - """ - for key in server.MICROVM_PLATFORM_CONFIG_REQUIRED_KEYS: - os.environ[server.MICROVM_PLATFORM_CONFIG_ENV_BY_KEY[key]] = _platform_config_value(key) - # A baked GITHUB_TOKEN_SECRET_ARN makes `resolve_github_token` reach for real - # Secrets Manager; pre-seeding its cache keeps these tests offline. Realistic - # for the image this fixture models, and it changes nothing they assert. - os.environ["GITHUB_TOKEN"] = "ghp_baked_for_tests" # noqa: S105 -- test placeholder, not a secret +def cached_github_token(env_guard): + """Keep pipeline mapping tests offline after authenticated config installation.""" + os.environ["GITHUB_TOKEN"] = "ghp_cached_for_tests" # noqa: S105 -- test placeholder yield @pytest.fixture -def env_guard(): +def env_guard(monkeypatch): """Snapshot/restore ``os.environ`` around a test that installs into it. ``_install_platform_config`` writes to the REAL process environment (that is @@ -1776,6 +1677,10 @@ def env_guard(): without this, one platform_config test would leak table names and a bogus ``AGENT_SESSION_ROLE_ARN`` into every test that runs after it (the conftest ``_clean_env`` fixture only strips the subset it knows about). + + All server tests use this through reset_server_state. Depending on monkeypatch + keeps its original-value restoration last, after joining work and restoring + direct environment writes. """ before = dict(os.environ) yield @@ -1796,8 +1701,68 @@ def test_allowlist_is_sourced_from_the_shared_contract(self): from shared_constants import SHARED_CONSTANTS contract = SHARED_CONSTANTS["microvm_platform_config"] + assert set(contract) == {"env_by_key", "required", "arn_keys", "account_anchor_key"} assert contract["env_by_key"] == server.MICROVM_PLATFORM_CONFIG_ENV_BY_KEY assert frozenset(contract["required"]) == server.MICROVM_PLATFORM_CONFIG_REQUIRED_KEYS + assert frozenset(contract["arn_keys"]) == server.MICROVM_PLATFORM_CONFIG_ARN_KEYS + assert contract["account_anchor_key"] == server.MICROVM_PLATFORM_CONFIG_ACCOUNT_ANCHOR_KEY + + def test_arn_fields_and_account_anchor_are_explicitly_reviewed(self): + assert { + "github_token_secret_arn", + "linear_oauth_secret_arn", + "jira_oauth_secret_arn", + "agent_session_role_arn", + } == server.MICROVM_PLATFORM_CONFIG_ARN_KEYS + assert server.MICROVM_PLATFORM_CONFIG_ACCOUNT_ANCHOR_KEY == "agent_session_role_arn" + + @pytest.mark.parametrize( + "key,env_name", + [("new_resource_arn", "NEW_RESOURCE"), ("new_resource", "NEW_RESOURCE_ARN")], + ) + def test_new_arn_fields_cannot_omit_arn_validation(self, monkeypatch, key, env_name): + monkeypatch.setattr( + server, + "MICROVM_PLATFORM_CONFIG_ENV_BY_KEY", + {**server.MICROVM_PLATFORM_CONFIG_ENV_BY_KEY, key: env_name}, + ) + with pytest.raises(ValueError, match=r"ARN-shaped key.*missing from arn_keys"): + server._validate_platform_config_contract() + + monkeypatch.setattr( + server, + "MICROVM_PLATFORM_CONFIG_ARN_KEYS", + server.MICROVM_PLATFORM_CONFIG_ARN_KEYS | {key}, + ) + server._validate_platform_config_contract() + + @pytest.mark.parametrize( + "constant,value,error", + [ + ("MICROVM_PLATFORM_CONFIG_ARN_KEYS", frozenset(), "arn_keys must not be empty"), + ( + "MICROVM_PLATFORM_CONFIG_ARN_KEYS", + frozenset({"unknown_arn"}), + "arn_keys names key.*absent from env_by_key", + ), + ( + "MICROVM_PLATFORM_CONFIG_ACCOUNT_ANCHOR_KEY", + "task_table_name", + "must be one of arn_keys", + ), + ( + "MICROVM_PLATFORM_CONFIG_ACCOUNT_ANCHOR_KEY", + "linear_oauth_secret_arn", + "must also be listed in .required", + ), + ], + ) + def test_invalid_arn_contract_fails_before_serving_tasks( + self, monkeypatch, constant, value, error + ): + monkeypatch.setattr(server, constant, value) + with pytest.raises(ValueError, match=error): + server._validate_platform_config_contract() def test_wire_contract_is_exactly_the_documented_key_set(self): # Spelled out on purpose: this is the wire contract Stage B's producer is @@ -1807,12 +1772,16 @@ def test_wire_contract_is_exactly_the_documented_key_set(self): "task_table_name": "TASK_TABLE_NAME", "task_events_table_name": "TASK_EVENTS_TABLE_NAME", "task_approvals_table_name": "TASK_APPROVALS_TABLE_NAME", + "approval_requests_api_url": "APPROVAL_REQUESTS_API_URL", "nudges_table_name": "NUDGES_TABLE_NAME", "log_group_name": "LOG_GROUP_NAME", "artifacts_bucket_name": "ARTIFACTS_BUCKET_NAME", "trace_artifacts_bucket_name": "TRACE_ARTIFACTS_BUCKET_NAME", + "continuation_bucket_name": "CONTINUATION_BUCKET_NAME", "github_token_secret_arn": "GITHUB_TOKEN_SECRET_ARN", "linear_oauth_secret_arn": "LINEAR_OAUTH_SECRET_ARN", + "linear_vault_enabled": "LINEAR_VAULT_ENABLED", + "linear_workload_identity_name": "LINEAR_WORKLOAD_IDENTITY_NAME", "jira_oauth_secret_arn": "JIRA_OAUTH_SECRET_ARN", "agent_session_role_arn": "AGENT_SESSION_ROLE_ARN", "aws_sdk_ua_app_id": "AWS_SDK_UA_APP_ID", @@ -1826,10 +1795,11 @@ def test_wire_contract_is_exactly_the_documented_key_set(self): "anthropic_model": "ANTHROPIC_MODEL", } - def test_required_subset_is_exactly_the_four_run_blocking_keys(self): + def test_required_subset_includes_the_trusted_approval_service(self): assert ( frozenset( { + "approval_requests_api_url", "task_table_name", "task_events_table_name", "github_token_secret_arn", @@ -2009,6 +1979,44 @@ def test_every_allowlisted_key_is_installable(self, env_guard): for key, env_name in server.MICROVM_PLATFORM_CONFIG_ENV_BY_KEY.items(): assert os.environ[env_name] == _platform_config_value(key) + def test_installed_vault_config_resolves_token_without_fallback(self, env_guard): + from config import resolve_linear_api_token + + os.environ.pop("LINEAR_API_TOKEN", None) + os.environ.pop("LINEAR_VAULT_ENABLED", None) + os.environ.pop("LINEAR_WORKLOAD_IDENTITY_NAME", None) + os.environ["AWS_REGION"] = "us-east-1" + server._install_platform_config( + _platform_config( + linear_vault_enabled="true", + linear_workload_identity_name="abca_linear_oauth", + ) + ) + client = MagicMock() + client.get_workload_access_token_for_user_id.return_value = { + "workloadAccessToken": "workload-token", + } + client.get_resource_oauth2_token.return_value = {"accessToken": "linear-vault-token"} + with patch("aws_session.platform_client", return_value=client) as make_client: + assert ( + resolve_linear_api_token( + { + "linear_provider_name": "bgagent-linear-oauth-acme", + "linear_workspace_id": "workspace-id", + "linear_vault_user_id": "linear-ws-acme", + "linear_oauth_secret_arn": _platform_config_value( + "linear_oauth_secret_arn" + ), + } + ) + == "linear-vault-token" + ) + make_client.assert_called_once_with("bedrock-agentcore", region_name="us-east-1") + client.get_workload_access_token_for_user_id.assert_called_once_with( + workloadName="abca_linear_oauth", userId="linear-ws-acme" + ) + client.get_secret_value.assert_not_called() + def test_payload_wins_over_a_pre_existing_image_env_value(self, env_guard): # The load-bearing precedence rule: image env is frozen at snapshot time, # the payload describes the live deployment. @@ -2201,20 +2209,10 @@ def test_an_arn_in_a_foreign_account_is_rejected(self, env_guard): assert "GITHUB_TOKEN_SECRET_ARN" not in os.environ def test_an_in_account_redirect_is_NOT_rejected(self, env_guard): - # The KNOWN LIMITATION, pinned as a test so it cannot be quietly mistaken for - # coverage. This check compares partition + account only, so a block naming - # another workspace's channel-OAuth secret in the SAME account is accepted — - # and the `bgagent-*-oauth-*` grants are prefix grants, so IAM would allow - # that read too. - # - # It is left open because it is currently unreachable from the guest: - # `platform_config` is produced by the orchestrator Lambda, and the MicroVM - # execution role holds `grantRead` ONLY on the payload bucket, so a running - # MicroVM can read another task's payload but cannot write one. - # - # If this test ever needs to flip to `pytest.raises`, the escalation is a - # name-shape check tying `github_token_secret_arn` to the task's own - # channel/workspace — see `MICROVM_PLATFORM_CONFIG_ARN_KEYS`. + # Unit scope: the installer checks internal ARN agreement, not origin. + # The v2 resolver authenticates the manifest and compares the downloaded + # config BEFORE calling this installer. Its own route/transport tests + # reject same-account workspace substitutions. installed = server._install_platform_config( _platform_config( github_token_secret_arn=( @@ -2226,10 +2224,8 @@ def test_an_in_account_redirect_is_NOT_rejected(self, env_guard): assert "GITHUB_TOKEN_SECRET_ARN" in installed def test_a_wholesale_partition_swap_is_NOT_rejected(self, env_guard): - # The other half of the same limitation, and the reason the docstrings say - # "internally consistent" rather than "pinned to this deployment": the anchor - # travels in the block it validates, so a payload that moves EVERY ARN to - # another partition agrees with itself and passes. IAM is what refuses it. + # A self-consistent block passes this internal-consistency helper. + # Deployment provenance is enforced earlier by the v2 bootstrap resolver. installed = server._install_platform_config( _platform_config( agent_session_role_arn=f"arn:aws-cn:iam::{_TEST_ACCOUNT}:role/r", @@ -2386,16 +2382,20 @@ def test_the_two_codes_are_distinct(self, env_guard): class TestMicrovmRunHookPlatformConfig: """``platform_config`` arrives on the ``/run`` hook as a SIBLING of ``agent_payload``.""" + @pytest.fixture(autouse=True) + def _task_identity(self, request): + self.task_id = f"t-pc-{request.node.name}" + def _payload(self, **extra) -> dict: return { - "task_id": "t-pc", + "task_id": self.task_id, "repo_url": "org/repo", "prompt": "do it", "github_token": "ghp_x", **extra, } - def test_inline_envelope_installs_the_config_and_accepts_the_task( + def test_verified_payload_installs_the_config_and_accepts_the_task( self, client, monkeypatch, env_guard ): monkeypatch.setattr(server, "run_task", MagicMock()) @@ -2506,194 +2506,6 @@ def test_a_non_object_block_returns_400(self, client, monkeypatch, env_guard): assert r.status_code == 400 assert r.json()["code"] == "MICROVM_RUN_PLATFORM_CONFIG_INVALID" - def test_an_envelope_without_platform_config_is_accepted_when_the_image_bakes_it( - self, client, monkeypatch, baked_platform_env, capfd - ): - # P1 compatibility, PRECISELY scoped (review N2): image snapshot and - # orchestrator Lambda deploy on independent cadences, so a new image must not - # require a Stage-B orchestrator — PROVIDED the values come from somewhere. - # Here the image bakes them, which is what the compatibility path is for. - monkeypatch.setattr(server, "run_task", MagicMock()) - monkeypatch.setattr(server.task_state, "write_terminal", MagicMock()) - - r = client.post(RUN_HOOK, json=_run_hook_body({"agent_payload": self._payload()})) - - assert r.status_code == 200 - # Still pre-install (nothing was installed), so the warning is stdout-only: - # `_warn_cw` here would spawn the CloudWatch thread off the snapshot's own - # baked env — the very thing it is warning about. See `TestMicrovmRunHook - # PreInstallAwsSilence`. - assert "[server/run-pre-config] /run hook received no platform_config" in ( - capfd.readouterr().out - ) - - def test_no_platform_config_and_no_baked_env_is_refused_not_run_unscoped( - self, client, monkeypatch, env_guard, capfd - ): - # Review N2, the version-skew case: a pre-Stage-B orchestrator launching a P2 - # image. The image bakes NOTHING by design (`imageEnvironmentVariables` - # defaults to `{}`), so nothing supplies `AGENT_SESSION_ROLE_ARN` — and - # `aws_session.get_session` silently falls back to the ambient compute role - # with tenant scoping OFF when it is unset. Refusing is the only correct - # answer; the previous behaviour was a 200 and a stdout breadcrumb. - run_task = MagicMock() - monkeypatch.setattr(server, "run_task", run_task) - monkeypatch.setattr(server.task_state, "write_terminal", MagicMock()) - for key in server.MICROVM_PLATFORM_CONFIG_REQUIRED_KEYS: - os.environ.pop(server.MICROVM_PLATFORM_CONFIG_ENV_BY_KEY[key], None) - - r = client.post(RUN_HOOK, json=_run_hook_body({"agent_payload": self._payload()})) - - assert r.status_code == 400 - body = r.json() - assert body["code"] == "MICROVM_RUN_PLATFORM_CONFIG_INCOMPLETE" - # The response NAMES the unset variables — the whole point is that a skewed - # deployment is diagnosable rather than mysterious. - assert body["missing_env"] == sorted( - server.MICROVM_PLATFORM_CONFIG_ENV_BY_KEY[key] - for key in server.MICROVM_PLATFORM_CONFIG_REQUIRED_KEYS - ) - assert "AGENT_SESSION_ROLE_ARN" in body["missing_env"] - # Nothing was started. - run_task.assert_not_called() - # And the rejection is attributable in the log, not just in the response. - assert "/run hook REJECTED: no platform_config" in capfd.readouterr().out - - def test_a_partially_baked_env_names_only_what_is_actually_missing( - self, client, monkeypatch, baked_platform_env - ): - # The realistic skew: an image that bakes SOME config. The rejection must - # name only the genuinely-unset variables, or an operator chases the wrong - # one. - monkeypatch.setattr(server, "run_task", MagicMock()) - monkeypatch.setattr(server.task_state, "write_terminal", MagicMock()) - os.environ.pop("AGENT_SESSION_ROLE_ARN", None) - - r = client.post(RUN_HOOK, json=_run_hook_body({"agent_payload": self._payload()})) - - assert r.status_code == 400 - assert r.json()["missing_env"] == ["AGENT_SESSION_ROLE_ARN"] - - def test_s3_pointer_takes_the_config_from_the_outer_envelope( - self, client, monkeypatch, env_guard - ): - # The producer's pointer form: the bare task payload lands in S3 and the - # config rides beside the pointer, inside the 4 KB hook body. - monkeypatch.setattr( - server, - "_fetch_microvm_payload_from_s3", - lambda _uri: self._payload(task_id="t-outer"), - ) - monkeypatch.setattr(server, "run_task", MagicMock()) - monkeypatch.setattr(server.task_state, "write_terminal", MagicMock()) - - r = client.post( - RUN_HOOK, - json=_run_hook_body( - { - "agent_payload_s3_uri": "s3://bucket/t-outer/payload.json", - "platform_config": _platform_config(task_table_name="outer-table"), - } - ), - ) - - assert r.status_code == 200 - assert r.json()["task_id"] == "t-outer" - assert os.environ["TASK_TABLE_NAME"] == "outer-table" - - def test_s3_pointer_takes_the_config_merged_into_the_fetched_object( - self, client, monkeypatch, env_guard - ): - # The producer ALSO merges the config into the S3 object, so the agent - # gets it whichever end of the fetch it reads. A stray platform_config key - # left in the bare payload is inert — the extractor reads named fields. - fetched = self._payload(task_id="t-inner") - fetched["platform_config"] = _platform_config(task_table_name="inner-table") - monkeypatch.setattr(server, "_fetch_microvm_payload_from_s3", lambda _uri: fetched) - monkeypatch.setattr(server, "run_task", MagicMock()) - monkeypatch.setattr(server.task_state, "write_terminal", MagicMock()) - - r = client.post( - RUN_HOOK, - json=_run_hook_body({"agent_payload_s3_uri": "s3://bucket/t-inner/payload.json"}), - ) - - assert r.status_code == 200 - assert r.json()["task_id"] == "t-inner" - assert os.environ["TASK_TABLE_NAME"] == "inner-table" - - def test_s3_object_may_itself_be_the_full_envelope(self, client, monkeypatch, env_guard): - monkeypatch.setattr( - server, - "_fetch_microvm_payload_from_s3", - lambda _uri: { - "agent_payload": self._payload(task_id="t-nested"), - "platform_config": _platform_config(task_table_name="nested-table"), - }, - ) - monkeypatch.setattr(server, "run_task", MagicMock()) - monkeypatch.setattr(server.task_state, "write_terminal", MagicMock()) - - r = client.post( - RUN_HOOK, - json=_run_hook_body({"agent_payload_s3_uri": "s3://bucket/t-nested/payload.json"}), - ) - - assert r.status_code == 200 - assert r.json()["task_id"] == "t-nested" - assert os.environ["TASK_TABLE_NAME"] == "nested-table" - - def test_the_fetched_object_wins_over_the_outer_envelope(self, client, monkeypatch, env_guard): - fetched = self._payload(task_id="t-prec") - fetched["platform_config"] = _platform_config(task_table_name="inner-wins") - monkeypatch.setattr(server, "_fetch_microvm_payload_from_s3", lambda _uri: fetched) - monkeypatch.setattr(server, "run_task", MagicMock()) - monkeypatch.setattr(server.task_state, "write_terminal", MagicMock()) - - r = client.post( - RUN_HOOK, - json=_run_hook_body( - { - "agent_payload_s3_uri": "s3://bucket/t-prec/payload.json", - "platform_config": _platform_config(task_table_name="outer-loses"), - } - ), - ) - - assert r.status_code == 200 - assert os.environ["TASK_TABLE_NAME"] == "inner-wins" - - def test_a_nested_agent_payload_of_the_wrong_type_is_a_retryable_500(self, client, monkeypatch): - # Review N1: this is a problem with the FETCHED OBJECT, not with the envelope - # the orchestrator built, so it belongs on the retryable 500 branch. The old - # 400 told the operator "the orchestrator built a bad envelope; retrying - # cannot help" — both halves wrong for a racing or half-written S3 object. - monkeypatch.setattr( - server, - "_fetch_microvm_payload_from_s3", - lambda _uri: {"agent_payload": "not-an-object"}, - ) - monkeypatch.setattr(server, "run_task", MagicMock()) - - r = client.post(RUN_HOOK, json=_run_hook_body({"agent_payload_s3_uri": "s3://b/k"})) - - assert r.status_code == 500 - assert r.json()["code"] == "MICROVM_RUN_PAYLOAD_UNREADABLE" - assert "agent_payload in the S3 payload must be an object" in r.json()["message"] - - def test_resolve_returns_the_config_alongside_the_payload(self): - payload, config = server._resolve_microvm_run_payload( - json.dumps({"agent_payload": {"task_id": "t"}, "platform_config": {"a": "b"}}) - ) - assert payload == {"task_id": "t"} - assert config == {"a": "b"} - - def test_resolve_returns_none_for_an_envelope_without_a_config(self): - _payload, config = server._resolve_microvm_run_payload( - json.dumps({"agent_payload": {"task_id": "t"}}) - ) - assert config is None - # -------------------------------------------------------------------------- # /validate + /terminate (ADR-021 P2) @@ -2720,10 +2532,33 @@ def test_returns_200_with_the_individual_checks(self, client): "hook_routes_registered": True, "python_version_supported": True, "platform_config_contract_loaded": True, + "image_lifecycle_protocol_supported": True, } assert body["hook_prefix"] == server.MICROVM_HOOK_PREFIX assert body["platform_config_keys"] == len(server.MICROVM_PLATFORM_CONFIG_ENV_BY_KEY) + @pytest.mark.parametrize( + "marker", + [None, str(server.SHARED_CONSTANTS["microvm_lifecycle"]["protocol_version"]), "", "999"], + ) + def test_build_rejects_an_image_marker_the_source_does_not_support( + self, client, monkeypatch, marker + ): + contract = server.SHARED_CONSTANTS["microvm_lifecycle"] + key = contract["image_protocol_env"] + if marker is None: + monkeypatch.delenv(key, raising=False) + else: + monkeypatch.setenv(key, marker) + forbidden = MagicMock(side_effect=AssertionError("image validation contacted AWS")) + monkeypatch.setattr("aws_session.platform_client", forbidden) + monkeypatch.setattr("aws_session.tenant_client", forbidden) + response = client.post(VALIDATE_HOOK) + expected = marker in (None, str(contract["protocol_version"])) + assert response.status_code == (200 if expected else 503) + assert response.json()["checks"]["image_lifecycle_protocol_supported"] is expected + forbidden.assert_not_called() + def test_makes_zero_aws_calls_even_with_a_log_group_configured( self, client, monkeypatch, capfd ): @@ -2755,10 +2590,9 @@ def forbidden(*_args, **_kwargs): def test_ready_is_also_aws_silent_with_a_log_group_configured( self, client, monkeypatch, capfd, warm_ready ): - # /ready runs under the same build role, so the same rule applies. It used - # to route through _debug_cw, whose write can only FAIL under a role with - # no Logs grant — and each failure bumps the shared _debug_cw_failures - # counter, poisoning the "debug path is blind" signal on every build. + # /ready shares the build role's namespace-scoped Logs permissions. + # Build diagnostics stay on stdout so no runtime logging client or + # build-role credential state is initialized before the snapshot. monkeypatch.setenv("LOG_GROUP_NAME", "/abca/agent") def forbidden(*_args, **_kwargs): @@ -2808,7 +2642,9 @@ def test_reports_a_missing_hook_route(self, client, monkeypatch): assert "hook_routes_registered" in r.json()["failed_checks"] assert r.json()["missing_routes"] == [ "/typo/prefix/ready", + "/typo/prefix/resume", "/typo/prefix/run", + "/typo/prefix/suspend", "/typo/prefix/terminate", "/typo/prefix/validate", ] @@ -2912,12 +2748,8 @@ def test_never_writes_terminal_task_status(self, client, monkeypatch): write_heartbeat.assert_not_called() def test_returns_200_without_joining_a_running_pipeline( - self, client, monkeypatch, baked_platform_env + self, client, monkeypatch, cached_github_token ): - # `baked_platform_env`: this envelope carries no `platform_config`, which - # since review N2 requires the effective env to supply the required values - # (a legacy image that bakes its own). This class is about the - # payload->pipeline mapping, not about config delivery. # A drain can take minutes (that is lifespan's job on graceful shutdown); # the hook budget is 1-60 s, so /terminate must observe and return. release = threading.Event() @@ -2977,14 +2809,14 @@ def test_an_unreadable_thread_count_reports_None_not_a_confident_zero( # COUNT is the only way to reach it, which is why patching `_debug_cw` (the # existing test above) cannot — by then `active` is already an int. class _UnreadableThreadList(list): - """Raises when COUNTED, but still clearable by the reset fixture.""" + """Simulate an unreadable registry only during the hook call.""" def __iter__(self): raise RuntimeError("thread registry read exploded") - monkeypatch.setattr(server, "_active_threads", _UnreadableThreadList()) - - r = client.post(TERMINATE_HOOK, json={"microvmId": "m-unknown"}) + with monkeypatch.context() as patch: + patch.setattr(server, "_active_threads", _UnreadableThreadList()) + r = client.post(TERMINATE_HOOK, json={"microvmId": "m-unknown"}) assert r.status_code == 200 body = r.json() @@ -2995,56 +2827,6 @@ def __iter__(self): assert "best-effort step failed" in capfd.readouterr().out -class TestMicrovmPayloadFetchAttribution: - """Every outbound AWS call carries ABCA's solution attribution (#319).""" - - def test_the_s3_payload_fetch_goes_through_the_attributed_factory(self, monkeypatch): - captured: dict = {} - - class _Body: - @staticmethod - def read(): - return b'{"task_id": "t-1"}' - - def fake_platform_client(service_name, **kwargs): - captured["service"] = service_name - captured["kwargs"] = kwargs - return SimpleNamespace(get_object=lambda **_kw: {"Body": _Body}) - - import aws_session - - monkeypatch.setattr(aws_session, "platform_client", fake_platform_client) - - assert server._fetch_microvm_payload_from_s3("s3://b/k") == {"task_id": "t-1"} - assert captured["service"] == "s3" - - def test_the_fetch_client_carries_the_md_user_agent_segment(self, monkeypatch): - # A naked boto3.client('s3') would silently drop the md/ segment. Assert on - # the OUTCOME (the UA on the config) rather than on which helper was used. - captured: dict = {} - - class _Body: - @staticmethod - def read(): - return b'{"task_id": "t-1"}' - - def fake_boto3_client(service_name, **kwargs): - captured["service"] = service_name - captured["config"] = kwargs.get("config") - return SimpleNamespace(get_object=lambda **_kw: {"Body": _Body}) - - import boto3 - - monkeypatch.setattr(boto3, "client", fake_boto3_client) - - server._fetch_microvm_payload_from_s3("s3://b/k") - - import ua - - assert captured["service"] == "s3" - assert ua.static_user_agent_extra() in captured["config"].user_agent_extra - - class TestSnapshotCredentialHygiene: """Nothing on the server's import or /ready path may cache an SDK session. @@ -3130,316 +2912,6 @@ def test_import_and_build_hooks_create_no_boto3_session(self): } -class TestMicrovmRunHookPreInstallAwsSilence: - """Before ``platform_config`` is installed, ``/run`` may touch exactly ONE AWS seam. - - Same defect class the build hooks avoid, one phase later: until the install has - run, ``LOG_GROUP_NAME`` is whatever the snapshot happens to carry, so a - ``_debug_cw`` on this path would resolve credentials and pin - ``boto3.DEFAULT_SESSION`` *before* the orchestrator's own region / - ``AWS_SDK_UA_APP_ID`` / session role are in the environment. The sole permitted - pre-install call is the S3 payload fetch, because the config is inside the - object being fetched. - - Every test here runs with a **baked ``LOG_GROUP_NAME``** — the hostile case the - fix exists for. Without it, ``_debug_cw`` degrades to stdout on its own and the - assertions would pass vacuously. - """ - - def _payload(self, **extra) -> dict: - return { - "task_id": "t-silent", - "repo_url": "org/repo", - "prompt": "do it", - "github_token": "ghp_x", - **extra, - } - - @pytest.fixture(autouse=True) - def _disable_background_pipeline(self, monkeypatch): - """Keep handler-only assertions isolated from asynchronous pipeline work. - - Mocking ``run_task`` is insufficient because ``_spawn_background`` returns - before its thread necessarily dereferences that global. Pytest can restore - the mock between parametrized cases while the prior thread is still - starting, letting it run the real pipeline under the next case's AWS seam - guard. Stub the spawn boundary instead: these tests specify only the - pre-install handler phase and none need a pipeline thread. - """ - monkeypatch.setattr(server, "_spawn_background", MagicMock()) - - @pytest.fixture - def seam_guard(self, monkeypatch): - """Arm every AWS/credential seam to raise until the pre-install phase is OVER. - - Two things end that phase, and only two: - - * ``_install_platform_config`` returning a **non-empty** env list — a real - install. Flipping on *any* return would be a hole big enough to drive B2 - through: the ``raw is None`` early return installs nothing and returns - ``[]``, so treating it as "installed" disarms the guard for the entire - legacy no-``platform_config`` path — which is exactly where a ``_warn_cw`` - was spawning the CloudWatch thread off the snapshot's baked env. - * ``_extract_invocation_params`` being entered. Past that point the legacy - path is *allowed* to talk to AWS: running on the snapshot's own env is the - documented P1-compatibility behaviour, so the accepted-line ``_debug_cw`` - and the pipeline below it are legitimate. Everything the handler does - *before* it — including the "no platform_config" warning — is not. - - A rejection path reaches neither, so the seams stay armed for the whole - request: a rejected run installed nothing and has no more right to an AWS - call than it had before. - """ - state: dict[str, Any] = { - "install_phase_done": False, - "installed_env": None, - "violations": [], - } - real_install = server._install_platform_config - real_extract = server._extract_invocation_params - - def spy_install(raw): - result = real_install(raw) - state["installed_env"] = result - if result: - state["install_phase_done"] = True - return result - - def spy_extract(*args, **kwargs): - state["install_phase_done"] = True - return real_extract(*args, **kwargs) - - monkeypatch.setattr(server, "_install_platform_config", spy_install) - monkeypatch.setattr(server, "_extract_invocation_params", spy_extract) - - def guard(name): - def _seam(*_args, **_kwargs): - if not state["install_phase_done"]: - state["violations"].append(name) - raise AssertionError(f"{name} touched before platform_config was installed") - return MagicMock() - - return _seam - - import boto3 - - import aws_session - - # Kept so a test can re-enable exactly the ONE permitted pre-install seam - # (the S3 payload fetch) and assert on it positively. - state["real_platform_client"] = aws_session.platform_client - - for module, attr in ( - (boto3, "client"), - (boto3, "Session"), - (aws_session, "platform_client"), - (aws_session, "tenant_client"), - (aws_session, "tenant_resource"), - (aws_session, "get_session"), - (server, "_debug_cw"), - (server, "_warn_cw"), - (server, "_debug_cw_exc"), - ): - monkeypatch.setattr(module, attr, guard(f"{module.__name__}.{attr}")) - - monkeypatch.setenv("LOG_GROUP_NAME", "/abca/agent") - return state - - @pytest.mark.parametrize("with_config", [True, False], ids=["with-config", "no-config"]) - def test_no_cloudwatch_or_credential_seam_is_touched_before_the_install( - self, client, monkeypatch, baked_platform_env, seam_guard, capfd, with_config - ): - # The ``no-config`` arm is the legacy P1 envelope, and it is the harder case: - # nothing is ever installed, so EVERY line up to param extraction — including - # the "running on the snapshot's frozen env" warning itself — is still - # pre-install. A ``_warn_cw`` there would spawn the CloudWatch writer thread - # and pin ``boto3.DEFAULT_SESSION`` off the baked ``LOG_GROUP_NAME`` this - # fixture sets, which is precisely the defect the warning is reporting. - # - # ``baked_platform_env`` (rather than ``env_guard``) so that arm reaches the - # ACCEPT path: since review N2 a no-config run with an unsatisfied effective - # env is refused, and a rejected run installs nothing and so proves nothing - # about the seams staying silent all the way to param extraction. Baking the - # env is also the only shape in which the no-config path is legitimate. - monkeypatch.setattr(server, "run_task", MagicMock()) - monkeypatch.setattr(server.task_state, "write_terminal", MagicMock()) - - envelope: dict[str, Any] = {"agent_payload": self._payload()} - if with_config: - envelope["platform_config"] = _platform_config() - - r = client.post(RUN_HOOK, json=_run_hook_body(envelope)) - - assert r.status_code == 200 - assert seam_guard["violations"] == [] - assert seam_guard["install_phase_done"] is True - if with_config: - assert seam_guard["installed_env"] - else: - # Vacuously "installed": the early return the flag must NOT trust. - assert seam_guard["installed_env"] == [] - # The warning still reaches an operator — stdout, via the pre-install sink. - assert ( - "[server/run-pre-config] /run hook received no platform_config" - in capfd.readouterr().out - ) - - def test_the_no_config_REJECTION_also_touches_no_seam( - self, client, monkeypatch, env_guard, seam_guard, capfd - ): - # Review N2's rejection is itself a pre-install path, and it emits a NEW log - # line — so it needs the same guarantee as every other rejection here: the - # refusal must not be the thing that pins `boto3.DEFAULT_SESSION` off the - # snapshot's baked `LOG_GROUP_NAME` (which `seam_guard` sets). - monkeypatch.setattr(server, "run_task", MagicMock()) - for key in server.MICROVM_PLATFORM_CONFIG_REQUIRED_KEYS: - os.environ.pop(server.MICROVM_PLATFORM_CONFIG_ENV_BY_KEY[key], None) - - r = client.post(RUN_HOOK, json=_run_hook_body({"agent_payload": self._payload()})) - - assert r.status_code == 400 - assert r.json()["code"] == "MICROVM_RUN_PLATFORM_CONFIG_INCOMPLETE" - assert seam_guard["violations"] == [] - assert "/run hook REJECTED: no platform_config" in capfd.readouterr().out - - def test_the_received_line_is_stdout_only( - self, client, monkeypatch, env_guard, seam_guard, capfd - ): - monkeypatch.setattr(server, "run_task", MagicMock()) - monkeypatch.setattr(server.task_state, "write_terminal", MagicMock()) - - client.post( - RUN_HOOK, - json=_run_hook_body( - {"agent_payload": self._payload(), "platform_config": _platform_config()}, - microvm_id="microvm-quiet", - ), - ) - - out = capfd.readouterr().out - assert "[server/run-pre-config] /run hook received:" in out - assert "microvm-quiet" in out - - def test_the_s3_payload_fetch_is_the_only_pre_install_aws_call( - self, client, monkeypatch, env_guard, seam_guard - ): - # The permitted exception, asserted positively: exactly one client, for s3, - # while the CloudWatch/credential seams stay armed. - services: list[str] = [] - - class _Body: - @staticmethod - def read(): - return json.dumps(self._payload(task_id="t-from-s3")).encode() - - def recording_client(service_name, **_kwargs): - services.append(service_name) - return SimpleNamespace(get_object=lambda **_kw: {"Body": _Body}) - - import boto3 - - import aws_session - - # Re-enable the one permitted seam, and only it: the fetch must still go - # through the attributed factory (#319), which delegates to boto3.client. - monkeypatch.setattr(aws_session, "platform_client", seam_guard["real_platform_client"]) - monkeypatch.setattr(boto3, "client", recording_client) - monkeypatch.setattr(server, "run_task", MagicMock()) - monkeypatch.setattr(server.task_state, "write_terminal", MagicMock()) - - r = client.post( - RUN_HOOK, - json=_run_hook_body( - { - "agent_payload_s3_uri": "s3://payload-bucket/t-from-s3/payload.json", - "platform_config": _platform_config(), - } - ), - ) - - assert r.status_code == 200 - assert r.json()["task_id"] == "t-from-s3" - assert services == ["s3"] - assert seam_guard["violations"] == [] - - def test_a_malformed_envelope_is_rejected_without_touching_a_seam( - self, client, monkeypatch, seam_guard, capfd - ): - monkeypatch.setattr(server, "run_task", MagicMock()) - - r = client.post(RUN_HOOK, json={"microvmId": "m", "runHookPayload": "not json"}) - - assert r.status_code == 400 - assert r.json()["code"] == "MICROVM_RUN_PAYLOAD_INVALID" - assert seam_guard["violations"] == [] - assert seam_guard["install_phase_done"] is False - # The reason still reaches an operator: stdout here, and the response body - # (which the MicroVM service surfaces) in every case. - assert "[server/run-pre-config] /run hook rejected:" in capfd.readouterr().out - - def test_a_failed_payload_fetch_is_reported_without_touching_a_seam( - self, client, monkeypatch, seam_guard, capfd - ): - def boom(_uri): - raise RuntimeError("AccessDenied") - - monkeypatch.setattr(server, "_fetch_microvm_payload_from_s3", boom) - monkeypatch.setattr(server, "run_task", MagicMock()) - - r = client.post(RUN_HOOK, json=_run_hook_body({"agent_payload_s3_uri": "s3://b/k"})) - - assert r.status_code == 500 - assert r.json()["code"] == "MICROVM_RUN_PAYLOAD_UNREADABLE" - assert seam_guard["violations"] == [] - out = capfd.readouterr().out - assert "[server/run-pre-config] /run hook payload fetch FAILED" in out - # The traceback is preserved on the stdout line (it is the only diagnostic - # the response body does not carry). - assert "Traceback" in out - - @pytest.mark.parametrize( - "config,expected_code", - [ - ({"ld_preload": "/tmp/evil.so"}, "MICROVM_RUN_PLATFORM_CONFIG_INVALID"), - ({"log_group_name": "lg"}, "MICROVM_RUN_PLATFORM_CONFIG_INCOMPLETE"), - ], - ) - def test_a_rejected_platform_config_touches_no_seam( - self, client, monkeypatch, env_guard, seam_guard, config, expected_code - ): - monkeypatch.setattr(server, "run_task", MagicMock()) - - r = client.post( - RUN_HOOK, - json=_run_hook_body({"agent_payload": self._payload(), "platform_config": config}), - ) - - assert r.status_code == 400 - assert r.json()["code"] == expected_code - # Nothing was installed, so nothing earned the right to an AWS call. - assert seam_guard["violations"] == [] - assert seam_guard["install_phase_done"] is False - - def test_the_accepted_line_correlates_task_and_microvm_ids( - self, client, monkeypatch, env_guard, capfd - ): - # The pre-install "received" line is stdout-only now, so the first line that - # reaches the task's log group has to join both ids by itself. - monkeypatch.setattr(server, "run_task", MagicMock()) - monkeypatch.setattr(server.task_state, "write_terminal", MagicMock()) - - client.post( - RUN_HOOK, - json=_run_hook_body( - {"agent_payload": self._payload(), "platform_config": _platform_config()}, - microvm_id="microvm-joined", - ), - ) - - out = capfd.readouterr().out - assert "/run hook accepted task_id='t-silent' microvm_id='microvm-joined'" in out - - class TestTerminateHookBodyTolerance: """``/terminate`` must answer 200 for ANY body — that is why it takes the raw request. @@ -3556,3 +3028,45 @@ def test_every_unusable_body_yields_empty(self, raw): def test_ignores_unrelated_fields(self): assert server._parse_terminate_microvm_id(b'{"reason": "idle", "x": 1}') == "" + + +@pytest.mark.parametrize("crash", [False, True]) +@pytest.mark.parametrize("attempt", ["", "replacement-2"]) +def test_microvm_pipeline_registers_identity_and_always_removes_lifecycle( + monkeypatch, crash, attempt +): + from microvm_lifecycle import get_context + + monkeypatch.delenv("UV_LINK_MODE", raising=False) + observed = [] + + def run_task(**kwargs): + context = get_context("lifecycle-server-task") + assert context is not None + assert context.microvm_id == "microvm-server" + assert context.attempt_id == (attempt or "lifecycle-server-task") + assert "microvm_id" not in kwargs + assert os.environ["UV_LINK_MODE"] == "copy" + observed.append(context) + if crash: + raise RuntimeError("pipeline crashed") + + monkeypatch.setattr(server, "run_task", run_task) + monkeypatch.setattr(server, "_heartbeat_worker", lambda *_: None) + monkeypatch.setattr(server.task_state, "write_terminal", MagicMock()) + server._run_task_background( + repo_url="owner/repo", + task_description="test", + issue_number="", + github_token="test-token", + anthropic_model="model", + max_turns=1, + max_budget_usd=None, + aws_region="us-west-2", + task_id="lifecycle-server-task", + microvm_id="microvm-server", + attempt_id=attempt, + ) + assert len(observed) == 1 + assert get_context("lifecycle-server-task") is None + assert server._background_pipeline_failed is crash diff --git a/agent/tests/test_server_isolation.py b/agent/tests/test_server_isolation.py new file mode 100644 index 000000000..0939d4aab --- /dev/null +++ b/agent/tests/test_server_isolation.py @@ -0,0 +1,98 @@ +"""Exercise server fixture teardown through a real, isolated pytest run.""" + +from __future__ import annotations + +import os +import subprocess +import sys +import textwrap +from pathlib import Path + + +def test_pipeline_finishes_before_test_mocks_and_environment_are_restored(tmp_path: Path): + """Force a pipeline to resolve run_task only when fixture teardown joins it.""" + test_file = tmp_path / "test_fixture_order.py" + test_file.write_text( + textwrap.dedent( + """ + import os + import threading + from unittest.mock import MagicMock + + import server + from test_server import env_guard, reset_server_state + + release = threading.Event() + entered = threading.Event() + first_run = MagicMock() + observed_env = [] + pipeline = None + original_join = None + + def test_first(monkeypatch, env_guard): + global pipeline, original_join + os.environ["ABCA_TEST_THREAD_ORIGIN"] = "first-test" + monkeypatch.setattr(server, "_debug_cw", lambda *a, **kw: None) + monkeypatch.setattr(server, "run_task", first_run) + + def delayed_lookup(**kwargs): + entered.set() + if release.wait(timeout=10): + observed_env.append(os.environ.get("ABCA_TEST_THREAD_ORIGIN")) + server.run_task(**kwargs) + + monkeypatch.setattr(server, "_run_task_background", delayed_lookup) + pipeline = server._spawn_background({"task_id": "fixture-first"}) + assert entered.wait(timeout=5) + original_join = pipeline.join + + def join_and_release(timeout=None): + release.set() + original_join(timeout=timeout) + + # Teardown must use this test's join and run_task before monkeypatch + # restores either. No timing-dependent sleep is needed. + monkeypatch.setattr(pipeline, "join", join_and_release) + + def test_second(monkeypatch): + next_run = MagicMock() + monkeypatch.setattr(server, "run_task", next_run) + try: + assert not pipeline.is_alive(), "prior test leaked a live pipeline" + first_run.assert_called_once_with(task_id="fixture-first") + assert observed_env == ["first-test"] + assert "ABCA_TEST_THREAD_ORIGIN" not in os.environ + assert server._active_threads == [] + finally: + # Also reap the deliberately blocked worker when testing the + # broken fixture, without ever invoking the real pipeline. + release.set() + original_join(timeout=5) + next_run.assert_not_called() + """ + ) + ) + agent_dir = Path(__file__).resolve().parents[1] + env: dict[str, str] = dict(os.environ) + env["PYTHONPATH"] = os.pathsep.join((str(agent_dir / "src"), str(agent_dir / "tests"))) + env.pop("ABCA_TEST_THREAD_ORIGIN", None) + result = subprocess.run( + [ + sys.executable, + "-m", + "pytest", + "-q", + "--no-cov", + "-c", + str(agent_dir / "pyproject.toml"), + str(test_file), + ], + cwd=agent_dir, + env=env, + text=True, + capture_output=True, + timeout=20, + check=False, + ) + assert result.returncode == 0, result.stdout + result.stderr + assert "2 passed" in result.stdout diff --git a/agent/tests/test_task_state.py b/agent/tests/test_task_state.py index b0eef03d9..0ce93921b 100644 --- a/agent/tests/test_task_state.py +++ b/agent/tests/test_task_state.py @@ -1,5 +1,9 @@ -"""Unit tests for pure functions in task_state.py.""" +"""Task persistence behavior and the agent's IAM write contract.""" +import ast +import json +import re +from pathlib import Path from unittest.mock import MagicMock import pytest @@ -8,6 +12,340 @@ from task_state import TaskFetchError, _build_logs_url, _now_iso +@pytest.mark.parametrize("writer", [task_state.write_running, task_state.write_heartbeat]) +@pytest.mark.parametrize( + ("reasons", "expected_warning"), + [ + ([{"Code": "ConditionalCheckFailed"}, {"Code": "None"}], False), + ([{"Code": "None"}, {"Code": "ConditionalCheckFailed"}], True), + ([{"Code": "ConditionalCheckFailed"}, {"Code": "ConditionalCheckFailed"}], True), + ([], True), + ], +) +def test_status_race_logging_preserves_worker_fence_failures( + monkeypatch, writer, reasons, expected_warning +): + from botocore.exceptions import ClientError + + monkeypatch.setattr(task_state, "_get_table", MagicMock()) + monkeypatch.setattr( + task_state, + "_update_task", + MagicMock( + side_effect=ClientError( + { + "Error": { + "Code": "TransactionCanceledException", + "Message": "transaction failed", + }, + "CancellationReasons": reasons, + }, + "TransactWriteItems", + ) + ), + ) + log = MagicMock() + monkeypatch.setattr(task_state, "log", log) + writer("owned-task") + assert any(call.args[0] == "WARN" for call in log.call_args_list) is expected_warning + + +@pytest.mark.parametrize( + ("codes", "can_heal"), + [ + (["ConditionalCheckFailed", "None"], True), + (["None", "ConditionalCheckFailed"], False), + (["ConditionalCheckFailed", "ConditionalCheckFailed"], False), + (["TransactionConflict", "None"], False), + ([], False), + ], +) +def test_terminal_transaction_heals_trace_only_when_worker_lease_passed( + monkeypatch, codes, can_heal +): + from botocore.exceptions import ClientError + + monkeypatch.setattr(task_state, "_get_table", MagicMock()) + error = ClientError( + { + "Error": {"Code": "TransactionCanceledException", "Message": "cancelled"}, + "CancellationReasons": [{"Code": code} for code in codes], + }, + "TransactWriteItems", + ) + monkeypatch.setattr(task_state, "_update_task", MagicMock(side_effect=error)) + heal = MagicMock(return_value=True) + report = MagicMock() + monkeypatch.setattr(task_state, "write_trace_uri_conditional", heal) + monkeypatch.setattr(task_state, "log_error_cw", report) + task_state.write_terminal("task", "COMPLETED", {"trace_s3_uri": "s3://bucket/trace"}) + assert heal.called is can_heal + if can_heal: + heal.assert_called_once_with("task", "s3://bucket/trace") + else: + assert "TransactionCanceledException" in report.call_args.args[0] + assert "CancellationReasons" in report.call_args.args[0] + + +@pytest.mark.parametrize( + ("codes", "benign"), + [ + (["ConditionalCheckFailed", "None"], True), + (["None", "ConditionalCheckFailed"], False), + (["ConditionalCheckFailed", "ConditionalCheckFailed"], False), + ], +) +def test_trace_transaction_classifies_status_race_separately_from_lease_loss( + monkeypatch, codes, benign +): + from botocore.exceptions import ClientError + + error = ClientError( + { + "Error": {"Code": "TransactionCanceledException"}, + "CancellationReasons": [{"Code": code} for code in codes], + }, + "TransactWriteItems", + ) + monkeypatch.setattr(task_state, "_get_table", MagicMock()) + monkeypatch.setattr(task_state, "_update_task", MagicMock(side_effect=error)) + report = MagicMock() + monkeypatch.setattr(task_state, "log", report) + assert not task_state.write_trace_uri_conditional("task", "s3://bucket/trace") + assert report.call_args.args[0] == ("INFO" if benign else "WARN") + if not benign: + assert "CancellationReasons=" in report.call_args.args[1] + + +@pytest.mark.parametrize( + "codes", [["TransactionConflict", "None"], ["None", "TransactionConflict"]] +) +def test_task_transaction_retries_conflict_with_identical_ownership_fence(monkeypatch, codes): + from botocore.exceptions import ClientError + + error = ClientError( + { + "Error": {"Code": "TransactionCanceledException"}, + "CancellationReasons": [{"Code": code} for code in codes], + }, + "TransactWriteItems", + ) + client = MagicMock() + client.transact_write_items.side_effect = [error, {"committed": True}] + monkeypatch.setattr(task_state.time, "sleep", MagicMock()) + lease = { + "ConditionCheck": { + "ConditionExpression": "lease_attempt_id = :attempt", + "ExpressionAttributeValues": {":attempt": "original-worker"}, + } + } + monkeypatch.setattr(task_state, "_lease_check", lambda *args, **kwargs: lease) + operation = {"TableName": "tasks", "Key": {"task_id": {"S": "task"}}} + assert task_state._update_task(client, "task", low_level=True, **operation) == { + "committed": True + } + assert client.transact_write_items.call_count == 2 + for call in client.transact_write_items.call_args_list: + assert call.kwargs["TransactItems"] == [{"Update": operation}, lease] + + +@pytest.mark.parametrize( + ("codes", "attempts"), + [ + (["TransactionConflict", "None"], 3), + (["TransactionConflict", "ConditionalCheckFailed"], 1), + (["None", "ConditionalCheckFailed"], 1), + ([], 1), + ], +) +def test_transaction_retry_is_bounded_and_never_retries_ownership_denial( + monkeypatch, codes, attempts +): + from botocore.exceptions import ClientError + + error = ClientError( + { + "Error": {"Code": "TransactionCanceledException"}, + "CancellationReasons": [{"Code": code} for code in codes], + }, + "TransactWriteItems", + ) + client = MagicMock() + client.transact_write_items.side_effect = error + monkeypatch.setattr(task_state.time, "sleep", MagicMock()) + with pytest.raises(ClientError): + task_state._transact_with_conflict_retry(client, []) + assert client.transact_write_items.call_count == attempts + + +@pytest.mark.parametrize( + ("failure", "expected"), + [ + (None, task_state.TerminalWriteOutcome.WRITTEN), + ("ConditionalCheckFailedException", task_state.TerminalWriteOutcome.SUPERSEDED), + ("AccessDeniedException", task_state.TerminalWriteOutcome.FAILED), + ], +) +def test_terminal_returns_persistence_outcome(monkeypatch, failure, expected): + from botocore.exceptions import ClientError + + monkeypatch.setattr(task_state, "_get_table", MagicMock()) + monkeypatch.setattr( + task_state, + "_update_task", + MagicMock( + side_effect=( + ClientError({"Error": {"Code": failure}}, "UpdateItem") if failure else None + ) + ), + ) + assert task_state.write_terminal("task", "COMPLETED") is expected + + +class TestAgentWriteContract: + def test_current_task_writers_fit_the_deployed_attribute_allowlist(self, monkeypatch): + """Exercise real writers; detect a new field before IAM rejects it live. + + This checks request/contract compatibility, not AWS IAM enforcement. + """ + table = MagicMock() + client = MagicMock() + monkeypatch.setattr(task_state, "_get_table", lambda: table) + monkeypatch.setenv("TASK_TABLE_NAME", "Tasks") + monkeypatch.setenv("TASK_APPROVALS_TABLE_NAME", "Approvals") + monkeypatch.setenv("AWS_REGION", "us-east-1") + monkeypatch.setenv("LOG_GROUP_NAME", "/test") + + task_state.write_running("t1") + task_state.write_heartbeat("t1") + task_state.write_terminal( + "t1", + "COMPLETED", + { + "pr_url": "https://example.com/pr/1", + "error": "example", + "cost_usd": 1, + "duration_s": 10, + "turns": 3, + "turns_attempted": 3, + "turns_completed": 2, + "prompt_version": "v1", + "memory_written": True, + "build_passed": True, + "lint_passed": True, + "code_changed": True, + "head_sha": "abc", + "answer_text": "done", + "otel_trace_id": "trace", + "trace_s3_uri": "s3://b/trace", + "artifact_uri": "s3://b/artifact", + }, + ) + assert task_state.write_trace_uri_conditional("t1", "s3://b/trace") + task_state.transact_write_approval_request( + "t1", + "r1", + { + "task_id": "t1", + "request_id": "r1", + "status": "PENDING", + "tool_name": "Bash", + "tool_input_preview": "example", + "tool_input_sha256": "a" * 64, + "reason": "approval required", + "severity": "high", + "matching_rule_ids": ["rule1"], + "created_at": "2026-09-13T00:00:00Z", + "timeout_s": 300, + "ttl": 1800000000, + "user_id": "u1", + "repo": "owner/repo", + }, + client=client, + ) + task_state.transact_resume_from_approval("t1", "r1", client=client) + assert task_state.increment_approval_gate_count_in_ddb("t1", client=client) + + requests = [c.kwargs for c in table.update_item.call_args_list] + requests += [c.kwargs for c in client.update_item.call_args_list] + for call in client.transact_write_items.call_args_list: + for item in call.kwargs["TransactItems"]: + action, request = next(iter(item.items())) + if request["TableName"] == "Tasks": + assert action == "Update" # Never Put/Delete the task row. + requests.append(request) + assert len(requests) == 7 + table.put_item.assert_not_called() + table.delete_item.assert_not_called() + + contract = Path(__file__).resolve().parents[2] / ( + "cdk/src/constructs/agent-task-write-attributes.json" + ) + allowed = set(json.loads(contract.read_text())) + seen: set[str] = set() + for request in requests: + # Current writers use flat attributes. Resolve aliases and ignore + # value placeholders/operators/functions in their DDB expressions. + expression = request["UpdateExpression"] + " " + request.get("ConditionExpression", "") + expression = re.sub(r":[A-Za-z0-9_]+", "", expression) + for alias, name in request.get("ExpressionAttributeNames", {}).items(): + expression = expression.replace(alias, name) + names = set(re.findall(r"[A-Za-z_][A-Za-z0-9_]*", expression)) + names -= {"SET", "REMOVE", "ADD", "IN", "AND", "attribute_not_exists"} + names |= set(request["Key"]) + assert names <= allowed, ( + f"Task writer needs a reviewed IAM contract update: {names - allowed}" + ) + seen |= names + # No stale writable attribute may linger after its writer is removed. + assert seen == allowed + + def test_write_inventory_requires_review_when_a_new_writer_is_added(self): + tree = ast.parse(Path(task_state.__file__).read_text()) + writers = { + node.name + for node in tree.body + if isinstance(node, ast.FunctionDef) + and any( + isinstance(child, ast.Call) + and ( + ( + isinstance(child.func, ast.Attribute) + and child.func.attr + in { + "update_item", + "put_item", + "delete_item", + "transact_write_items", + "batch_writer", + } + ) + or ( + isinstance(child.func, ast.Name) + and child.func.id + in {"_update_task", "_transact_task", "_transact_with_conflict_retry"} + ) + ) + for child in ast.walk(node) + ) + } + assert writers == { + "_transact_with_conflict_retry", + "_update_task", + "_transact_task", + "write_running", + "write_heartbeat", + "write_terminal", + "write_trace_uri_conditional", + "transact_write_approval_request", + "transact_resume_from_approval", + "increment_approval_gate_count_in_ddb", + "publish_continuation_checkpoint", + "consume_restored_continuation", + "best_effort_update_approval_status", # Writes only the supporting approvals table. + } + + class TestNowIso: def test_format(self): result = _now_iso() @@ -97,82 +435,6 @@ def get_item(self, Key): assert "ProvisionedThroughputExceededException" in str(exc_info.value) -class TestWriteSessionInfo: - """Rev-5 OBS-4: interactive path writes session_id + agent_runtime_arn.""" - - def test_writes_session_id_and_arn(self, monkeypatch): - calls: list[dict] = [] - - class _FakeTable: - def update_item(self, **kwargs): - calls.append(kwargs) - - monkeypatch.setattr(task_state, "_get_table", lambda: _FakeTable()) - - task_state.write_session_info( - "t-interactive", - "sess-abc123", - "arn:aws:bedrock-agentcore:us-east-1:123:runtime/jwt-xyz", - ) - - assert len(calls) == 1 - call = calls[0] - assert call["Key"] == {"task_id": "t-interactive"} - assert "session_id = :sid" in call["UpdateExpression"] - assert "agent_runtime_arn = :arn" in call["UpdateExpression"] - assert "compute_type = :ct" in call["UpdateExpression"] - assert "compute_metadata = :cm" in call["UpdateExpression"] - values = call["ExpressionAttributeValues"] - assert values[":sid"] == "sess-abc123" - assert values[":arn"] == "arn:aws:bedrock-agentcore:us-east-1:123:runtime/jwt-xyz" - assert values[":ct"] == "agentcore" - assert values[":cm"] == { - "runtimeArn": "arn:aws:bedrock-agentcore:us-east-1:123:runtime/jwt-xyz" - } - - def test_noop_when_both_empty(self, monkeypatch): - calls: list[dict] = [] - - class _FakeTable: - def update_item(self, **kwargs): - calls.append(kwargs) - - monkeypatch.setattr(task_state, "_get_table", lambda: _FakeTable()) - - task_state.write_session_info("t-empty", "", "") - assert calls == [] - - def test_skips_silently_when_task_already_advanced(self, monkeypatch): - from botocore.exceptions import ClientError - - class _FakeTable: - def update_item(self, **kwargs): - raise ClientError( - {"Error": {"Code": "ConditionalCheckFailedException"}}, - "UpdateItem", - ) - - monkeypatch.setattr(task_state, "_get_table", lambda: _FakeTable()) - - # Must NOT raise — the conditional failure is expected when the - # task has already transitioned past SUBMITTED/HYDRATING. - task_state.write_session_info("t-raced", "sess-x", "arn:x") - - def test_writes_only_session_when_arn_missing(self, monkeypatch): - calls: list[dict] = [] - - class _FakeTable: - def update_item(self, **kwargs): - calls.append(kwargs) - - monkeypatch.setattr(task_state, "_get_table", lambda: _FakeTable()) - - task_state.write_session_info("t-partial", "sess-only", "") - assert len(calls) == 1 - assert "session_id = :sid" in calls[0]["UpdateExpression"] - assert "agent_runtime_arn" not in calls[0]["UpdateExpression"] - - class TestWriteRunningMaintainsStatusCreatedAt: """Regression guard: ``write_running`` must rewrite ``status_created_at`` so the ``UserStatusIndex`` GSI sort key reflects the current status. @@ -807,6 +1069,28 @@ def test_condition_failed_reason_includes_both_branches( class TestTransactResumeFromApproval: + def test_resume_refreshes_heartbeat_in_the_same_conditional_write( + self, approval_tables_env, monkeypatch + ): + """A poll immediately after a long approval wait must see a fresh heartbeat.""" + monkeypatch.setattr(task_state, "_now_iso", lambda: "2026-09-13T12:05:00Z") + client = MagicMock() + + task_state.transact_resume_from_approval("01KTASK", "01KREQ", client=client) + + client.transact_write_items.assert_called_once() + updates = client.transact_write_items.call_args.kwargs["TransactItems"] + assert len(updates) == 1 + update = updates[0]["Update"] + assert "agent_heartbeat_at = :heartbeat" in update["UpdateExpression"] + assert "#s = :running" in update["UpdateExpression"] + assert update["ExpressionAttributeValues"][":heartbeat"] == {"S": "2026-09-13T12:05:00Z"} + assert update["ConditionExpression"] == ( + "#s = :awaiting AND awaiting_approval_request_id = :rid" + ) + # An independent write would leave the old heartbeat visible after RUNNING. + client.update_item.assert_not_called() + def test_env_missing_raises(self, monkeypatch): monkeypatch.delenv("TASK_TABLE_NAME", raising=False) monkeypatch.delenv("TASK_APPROVALS_TABLE_NAME", raising=False) @@ -884,12 +1168,12 @@ def test_reason_optional_attached(self, approval_tables_env): client.update_item.return_value = {} task_state.best_effort_update_approval_status( - "01KTASK", "01KREQ", "DENIED", reason="no prod pushes", client=client + "01KTASK", "01KREQ", "TIMED_OUT", reason="polling failed", client=client ) call = client.update_item.call_args assert "deny_reason = :reason" in call.kwargs["UpdateExpression"] - assert call.kwargs["ExpressionAttributeValues"][":reason"] == {"S": "no prod pushes"} + assert call.kwargs["ExpressionAttributeValues"][":reason"] == {"S": "polling failed"} def test_conditional_check_failed_returns_false(self, approval_tables_env): """IMPL-24 — this is the VM-throttle race signal the hook re-reads on.""" @@ -1045,3 +1329,52 @@ def test_extract_cancellation_reasons(self): def test_extract_cancellation_reasons_none_on_plain_exception(self): assert task_state._extract_cancellation_reasons(RuntimeError()) == [] + + +class TestContinuationClaimReadback: + @pytest.mark.parametrize( + "changed", ["none", "cancelled", "worker", "record", "request", "lease"] + ) + def test_lost_response_requires_exact_consumed_assignment( + self, approval_tables_env, monkeypatch, changed + ): + from boto3.dynamodb.types import TypeSerializer + + record = { + "version": 1, + "state": "RESTORING", + "worker_id": "microvm-new", + "identity": { + "task_id": "task", + "attempt_id": "microvm-old", + "request_id": "request", + "user_id": "user", + "repo": "owner/repo", + }, + "manifest": {"key": "exact-saved-key"}, + } + consumed = {**record, "state": "CONSUMED"} + task = {"status": "RUNNING", "session_id": "microvm-new", "continuation": consumed} + if changed == "cancelled": + task["status"] = "CANCELLED" + elif changed == "worker": + task["session_id"] = "microvm-other" + elif changed == "record": + task["continuation"] = {**consumed, "manifest": {"key": "different"}} + elif changed == "request": + task["awaiting_approval_request_id"] = "another-request" + client = MagicMock() + client.transact_write_items.side_effect = TimeoutError("response lost after commit") + serialize = TypeSerializer().serialize + client.get_item.return_value = {"Item": {k: serialize(v) for k, v in task.items()}} + lease = MagicMock(side_effect=RuntimeError("lease lost") if changed == "lease" else None) + monkeypatch.setattr(task_state, "verify_worker_lease", lease) + if changed == "none": + task_state.consume_restored_continuation("task", "microvm-new", record, client=client) + lease.assert_called_once_with("task", client=client) + else: + with pytest.raises(RuntimeError): + task_state.consume_restored_continuation( + "task", "microvm-new", record, client=client + ) + assert client.get_item.call_args.kwargs["ConsistentRead"] is True diff --git a/agent/tests/test_workflow_runner.py b/agent/tests/test_workflow_runner.py index d71494dde..116e55946 100644 --- a/agent/tests/test_workflow_runner.py +++ b/agent/tests/test_workflow_runner.py @@ -487,6 +487,73 @@ def _real_ctx(workflow: Workflow, **kw): return StepContext(workflow=workflow, config=config, **kw) +@pytest.mark.parametrize( + "prepared_prompt", + [ + "Recorded APPROVED. Saved proposal: " + '{"command": "cat marker", "description": "Read marker"}', + "Recorded DENIED. Do not work around this decision.", + "Recorded TIMED_OUT. Continue from the saved decision.", + "Original task with local attachment references.", + "", + ], +) +def test_hydration_preserves_the_prepared_prompt_sent_to_the_agent(monkeypatch, prepared_prompt): + from models import AgentResult, HydratedContext + from workflow.runner import _handle_hydrate_context, _handle_run_agent + + wf = _workflow([{"kind": "hydrate_context"}, {"kind": "run_agent"}]) + ctx = _real_ctx( + wf, + hydrated=HydratedContext(user_prompt="Original task before the human decision"), + user_prompt=prepared_prompt, + system_prompt="Saved system prompt", + ) + received = [] + + async def capture_prompt(prompt, system_prompt, *_args, **_kwargs): + received.append((prompt, system_prompt)) + return AgentResult(status="success") + + monkeypatch.setattr("runner.run_agent", capture_prompt) + _handle_hydrate_context(wf.steps[0], ctx) + _handle_run_agent(wf.steps[1], ctx) + + assert received == [ + (prepared_prompt or "Original task before the human decision", "Saved system prompt") + ] + + +def test_closed_microvm_cannot_run_delivery_steps_after_its_agent_unwinds(monkeypatch): + from microvm_lifecycle import register_task, unregister_task + from models import AgentResult + from workflow.runner import _handle_run_agent + + wf = _workflow([{"kind": "run_agent"}, {"kind": "ensure_pr", "strategy": "create"}]) + ctx = _real_ctx(wf, user_prompt="test", system_prompt="test") + lifecycle = register_task(ctx.config.task_id, "microvm-test") + + async def finish(*_args, **_kwargs): + lifecycle.close() + return AgentResult(status="success") + + def must_not_deliver(*_args): + raise AssertionError("closed worker reached its delivery step") + + monkeypatch.setattr("runner.run_agent", finish) + try: + result = run_workflow( + wf, ctx, handlers={"run_agent": _handle_run_agent, "ensure_pr": must_not_deliver} + ) + assert not result.succeeded + assert result.failed_step is not None + assert result.failed_step.error is not None + assert "Worker execution is closed" in result.failed_step.error + assert len(result.outcomes) == 1 + finally: + unregister_task(lifecycle) + + class TestVerifyHandlers: def test_verify_build_regression_only_passes_when_broken_before(self, monkeypatch): from models import RepoSetup diff --git a/agent/workflows/coding/pr-review-v1.yaml b/agent/workflows/coding/pr-review-v1.yaml index 30d6033fd..bf43ba60f 100644 --- a/agent/workflows/coding/pr-review-v1.yaml +++ b/agent/workflows/coding/pr-review-v1.yaml @@ -2,9 +2,10 @@ # mutating it. Selected via workflow_ref (pr_review → coding/pr-review-v1). # # read_only:true with Write/Edit dropped from allowed_tools (validator rule 6 — -# read-only tier may not grant mutating tools). Read-only is enforced by Cedar -# `context.read_only == true` (hard-deny read_only_forbid_write/edit) AND by the -# trimmed allowed_tools — defense in depth. It resolves the existing PR URL +# read-only tier may not list Write/Edit). Cedar's `context.read_only == true` +# hard-denies Write/Edit. The SDK allowed_tools list controls auto-approval; +# omitting a tool from it does not remove that tool from the available surface. +# Bash commands still depend on their own policy rules. It resolves the PR URL # without pushing via ensure_pr(strategy: resolve); the post_review step is not # used (its registered handler is intentionally unimplemented). The Cedar # principal is the id-derived audit identity "pr_review" — enforcement keys off diff --git a/agent/workflows/default/agent-v1.yaml b/agent/workflows/default/agent-v1.yaml index f0910879b..1061f5b3b 100644 --- a/agent/workflows/default/agent-v1.yaml +++ b/agent/workflows/default/agent-v1.yaml @@ -6,7 +6,7 @@ # hatch — so its behavior is auditable and overridable per-repo via the Blueprint. # # Deliberately conservative because it runs when *nothing* was specified: -# requires_repo:false (no clone), a read-leaning tool set (no Bash/Write/Edit), +# requires_repo:false (repository optional), a read-leaning auto-approval list, # and delivery to S3 + a comment milestone. A caller who wants coding selects # (or maps to) a coding workflow. soft_deny is still mandatory (read_only:false) # so any future tool addition stays HITL-gated. @@ -23,7 +23,9 @@ domain: hybrid description: >- Run the user's request through the agent and deliver the result. Minimal default — no repo, build, or PR assumptions. -# No clone; if a repo is supplied it is hydrated as context, not scaffolded. +# Without a repo, pipeline.py skips cloning and delivers an S3 artifact. When a +# repo is supplied, the current shared pipeline clones it and runs its repo-bound +# build/PR post-hooks. This flag alone does not disable that delivery path. requires_repo: false read_only: false prompt: @@ -33,8 +35,8 @@ hydration: sources: [task_description, attachments, memory] agent_config: tier: standard - # Conservative: no Bash/Write/Edit by default. The default must not silently - # mutate a filesystem or push code on a submission that never asked for it. + # No Bash/Write/Edit in SDK auto-approval. This list does not hide those tools; + # applicable Cedar rules govern approval/denial when they are requested. allowed_tools: [Read, Glob, Grep, WebFetch] cedar_policy_modules: [builtin/hard_deny, builtin/soft_deny] repo_config: diff --git a/cdk/AGENTS.md b/cdk/AGENTS.md index e328711e1..35f6c75d5 100644 --- a/cdk/AGENTS.md +++ b/cdk/AGENTS.md @@ -27,8 +27,11 @@ mise //cdk:destroy # destroy stack | Code | Test location | |------|---------------| | Shared handler logic | `cdk/test/handlers/shared/*.test.ts` | +| Trusted approval writer | `src/handlers/request-approval.ts`, `src/constructs/approval-request-service.ts`; handler/unit, DynamoDB Local and construct tests | | Handler entrypoints | `cdk/test/handlers/orchestrate-task.test.ts`, `create-task.test.ts`, `webhook-create-task.test.ts` | | Constructs | `cdk/test/constructs/task-orchestrator.test.ts`, `task-api.test.ts` | +| MicroVM lifecycle/capacity transactions | `test/handlers/shared/*-local.test.ts` and `test/handlers/request-approval-local.test.ts`; use a loopback DynamoDB Local endpoint in `ABCA_DDB_LOCAL_ENDPOINT` (mandatory in CI) | +| Live verification harnesses | `test/live/`; explicitly invoked against an owned fixture, excluded from normal Jest collection | Construct tests: synthesize each distinct stack config once in `beforeAll`, assert against cached `Template` — do not re-synth per test. Bundling is disabled globally via `test/setup/disable-bundling.ts` (see Common mistakes). @@ -40,6 +43,7 @@ Construct tests: synthesize each distinct stack config once in `beforeAll`, asse | `cdk/src/stacks/` | WRITE | Stack definitions | | `cdk/src/constructs/` | WRITE | Reusable constructs | | `cdk/src/handlers/shared/types.ts` | WRITE | API types (mirror to `cli/src/types.ts`) | +| `cdk/src/handlers/shared/canonical-json.ts` | WRITE | Runtime receipt encoding; preserve stored bytes and keep migration encoding separate | | `cdk/test/` | WRITE | Unit / snapshot tests | | `cli/src/types.ts` | WRITE (sync) | Must match shared types | diff --git a/cdk/bootstrap/BOOTSTRAP_HASH b/cdk/bootstrap/BOOTSTRAP_HASH index 3ca756db5..61b09e65c 100644 --- a/cdk/bootstrap/BOOTSTRAP_HASH +++ b/cdk/bootstrap/BOOTSTRAP_HASH @@ -1 +1 @@ -d30eb8e63c8f6fd5551de03a72acb342231175aaf417bd98811d6a07b0be3afc +dc6301b65558c8fb51200e7b712a754e1b85adda7a512b14ec4f647dc954f541 diff --git a/cdk/bootstrap/BOOTSTRAP_VERSION b/cdk/bootstrap/BOOTSTRAP_VERSION index dc1e644a1..f8e233b27 100644 --- a/cdk/bootstrap/BOOTSTRAP_VERSION +++ b/cdk/bootstrap/BOOTSTRAP_VERSION @@ -1 +1 @@ -1.6.0 +1.9.0 diff --git a/cdk/bootstrap/bootstrap-template.yaml b/cdk/bootstrap/bootstrap-template.yaml index b4568679f..934df554c 100644 --- a/cdk/bootstrap/bootstrap-template.yaml +++ b/cdk/bootstrap/bootstrap-template.yaml @@ -1,12 +1,13 @@ # GENERATED FILE - DO NOT EDIT DIRECTLY # This template is generated by: npx tsx scripts/generate-bootstrap-template.ts -# ABCA Bootstrap Policy Version: 1.6.0 -# ABCA Bootstrap Policy Hash: d30eb8e63c8f6fd5551de03a72acb342231175aaf417bd98811d6a07b0be3afc +# ABCA Bootstrap Policy Version: 1.9.0 +# ABCA Bootstrap Policy Hash: dc6301b65558c8fb51200e7b712a754e1b85adda7a512b14ec4f647dc954f541 # # Based on the default CDK bootstrap template with the following modifications: # - BootstrapVariant set to "ABCA: Least-Privilege Bootstrap" # - ComputeTypes parameter added for compute-variant selection # - IncludeComputeEcs / IncludeComputeLambdaMicrovms conditions added +# - Execution role may pass only itself to CloudFormation for nested stacks # - 6 AWS::IAM::ManagedPolicy resources replace AdministratorAccess; each # PolicyDocument is a minified JSON string so the template stays under the # 51,200-char CloudFormation inline-template limit (#864) @@ -754,6 +755,20 @@ Resources: - PermissionsBoundarySet - Fn::Sub: arn:${AWS::Partition}:iam::${AWS::AccountId}:policy/${InputPermissionsBoundary} - Ref: AWS::NoValue + Policies: + - PolicyName: PassExecutionRoleToCloudFormation + PolicyDocument: + Version: '2012-10-17' + Statement: + - Sid: PassSelfToCloudFormation + Effect: Allow + Action: iam:PassRole + Resource: + Fn::Sub: >- + arn:${AWS::Partition}:iam::${AWS::AccountId}:role/cdk-${Qualifier}-cfn-exec-role-${AWS::AccountId}-${AWS::Region} + Condition: + StringEquals: + iam:PassedToService: cloudformation.amazonaws.com CdkBoostrapPermissionsBoundaryPolicy: Condition: ShouldCreatePermissionsBoundary Type: AWS::IAM::ManagedPolicy @@ -853,7 +868,7 @@ Resources: ManagedPolicyName: Fn::Sub: cdk-${Qualifier}-IaCRole-ABCA-Compute-LambdaMicrovms-${AWS::AccountId}-${AWS::Region} PolicyDocument: >- - {"Statement":[{"Action":["lambda:CreateMicrovmImage","lambda:GetMicrovmImage","lambda:UpdateMicrovmImage","lambda:DeleteMicrovmImage","lambda:ListMicrovmImages","lambda:GetMicrovmImageVersion","lambda:UpdateMicrovmImageVersion","lambda:DeleteMicrovmImageVersion","lambda:ListMicrovmImageVersions","lambda:GetMicrovmImageBuild","lambda:ListMicrovmImageBuilds","lambda:ListManagedMicrovmImages","lambda:ListManagedMicrovmImageVersions","lambda:CreateNetworkConnector","lambda:GetNetworkConnector","lambda:UpdateNetworkConnector","lambda:DeleteNetworkConnector","lambda:ListNetworkConnectors","lambda:PassNetworkConnector"],"Effect":"Allow","Resource":"*","Sid":"LambdaMicrovms"},{"Action":"iam:PassRole","Effect":"Allow","Resource":["arn:aws:iam::*:role/backgroundagent-dev-LambdaMicrovmComputeBuild*","arn:aws:iam::*:role/backgroundagent-dev-LambdaMicrovmComputeConnector*"],"Sid":"MicrovmPassRoles"}],"Version":"2012-10-17"} + {"Statement":[{"Action":["lambda:CreateMicrovmImage","lambda:GetMicrovmImage","lambda:UpdateMicrovmImage","lambda:DeleteMicrovmImage","lambda:ListMicrovmImages","lambda:GetMicrovmImageVersion","lambda:UpdateMicrovmImageVersion","lambda:DeleteMicrovmImageVersion","lambda:ListMicrovmImageVersions","lambda:GetMicrovmImageBuild","lambda:ListMicrovmImageBuilds","lambda:ListManagedMicrovmImages","lambda:ListManagedMicrovmImageVersions","lambda:CreateNetworkConnector","lambda:GetNetworkConnector","lambda:UpdateNetworkConnector","lambda:DeleteNetworkConnector","lambda:ListNetworkConnectors","lambda:PassNetworkConnector"],"Effect":"Allow","Resource":"*","Sid":"LambdaMicrovms"},{"Action":"iam:PassRole","Effect":"Allow","Resource":["arn:aws:iam::*:role/backgroundagent-dev-LambdaMicrovmComputeBuild*","arn:aws:iam::*:role/backgroundagent-dev-LambdaMicrovmComputeConnector*","arn:aws:iam::*:role/backgroundagent-dev-MicrovmBuildRole","arn:aws:iam::*:role/backgroundagent-dev-MicrovmConnectorRole"],"Sid":"MicrovmPassRoles"},{"Action":["ssm:GetParameters","ssm:PutParameter","ssm:DeleteParameter","ssm:AddTagsToResource","ssm:RemoveTagsFromResource","ssm:ListTagsForResource"],"Effect":"Allow","Resource":"arn:aws:ssm:*:*:parameter/backgroundagent-*/microvm-approval-suspend-enabled","Sid":"MicrovmSuspendConfiguration"}],"Version":"2012-10-17"} Description: 'ABCA Bootstrap: IaCRole-ABCA-Compute-LambdaMicrovms permissions for CloudFormation execution role' Condition: IncludeComputeLambdaMicrovms Outputs: @@ -884,10 +899,10 @@ Outputs: Value: '32' BootstrapPolicyVersion: Description: The version of the ABCA bootstrap policy bundle - Value: 1.6.0 + Value: 1.9.0 BootstrapPolicyHash: Description: SHA-256 hash of the ABCA bootstrap policy bundle for drift detection - Value: d30eb8e63c8f6fd5551de03a72acb342231175aaf417bd98811d6a07b0be3afc + Value: dc6301b65558c8fb51200e7b712a754e1b85adda7a512b14ec4f647dc954f541 BootstrapPolicySet: Description: Comma-separated list of active ABCA bootstrap policy names Value: diff --git a/cdk/bootstrap/policies/compute-lambda-microvm.json b/cdk/bootstrap/policies/compute-lambda-microvm.json index dd7656ad6..227dd2cda 100644 --- a/cdk/bootstrap/policies/compute-lambda-microvm.json +++ b/cdk/bootstrap/policies/compute-lambda-microvm.json @@ -31,9 +31,24 @@ "Effect": "Allow", "Resource": [ "arn:aws:iam::*:role/backgroundagent-dev-LambdaMicrovmComputeBuild*", - "arn:aws:iam::*:role/backgroundagent-dev-LambdaMicrovmComputeConnector*" + "arn:aws:iam::*:role/backgroundagent-dev-LambdaMicrovmComputeConnector*", + "arn:aws:iam::*:role/backgroundagent-dev-MicrovmBuildRole", + "arn:aws:iam::*:role/backgroundagent-dev-MicrovmConnectorRole" ], "Sid": "MicrovmPassRoles" + }, + { + "Action": [ + "ssm:GetParameters", + "ssm:PutParameter", + "ssm:DeleteParameter", + "ssm:AddTagsToResource", + "ssm:RemoveTagsFromResource", + "ssm:ListTagsForResource" + ], + "Effect": "Allow", + "Resource": "arn:aws:ssm:*:*:parameter/backgroundagent-*/microvm-approval-suspend-enabled", + "Sid": "MicrovmSuspendConfiguration" } ], "Version": "2012-10-17" diff --git a/cdk/package.json b/cdk/package.json index bccf3df3e..bf838e8ab 100644 --- a/cdk/package.json +++ b/cdk/package.json @@ -4,7 +4,7 @@ "private": true, "description": "ABCA CDK application", "license": "MIT-0", - "_cedar_parity_pin": "DO NOT BUMP @cedar-policy/cedar-wasm IN ISOLATION. The Python cedarpy binding (agent/pyproject.toml) and this WASM binding share a Rust core and must move together. Drift between bindings can produce divergent (decision, matching_rule_ids) on the same (policy, input). See docs/design/CEDAR_HITL_GATES.md §15.6 (decision #23). The contracts/cedar-parity/ golden fixtures are how CI catches divergence; bumping either side requires bumping the other AND refreshing the fixtures in the same commit.", + "_cedar_parity_pin": "DO NOT BUMP @cedar-policy/cedar-wasm IN ISOLATION. The Python cedarpy binding (agent/pyproject.toml) and this WASM binding share a Rust core and must move together. Drift between bindings can produce divergent (decision, matching_rule_ids) on the same (policy, input). See docs/design/CEDAR_HITL_GATES.md \u00a715.6 (decision #23). The contracts/cedar-parity/ golden fixtures are how CI catches divergence; bumping either side requires bumping the other AND refreshing the fixtures in the same commit.", "scripts": { "compile": "tsc --build tsconfig.json", "watch": "tsc --build -w tsconfig.json", @@ -29,6 +29,7 @@ "@aws-sdk/client-s3": "^3.1078.0", "@aws-sdk/client-secrets-manager": "^3.1078.0", "@aws-sdk/client-sns": "^3.1078.0", + "@aws-sdk/client-ssm": "^3.1078.0", "@aws-sdk/client-sts": "^3.1078.0", "@aws-sdk/credential-provider-node": "^3.972.61", "@aws-sdk/lib-dynamodb": "^3.1078.0", @@ -137,6 +138,9 @@ "tsconfig": "tsconfig.jest.json" } ] + }, + "moduleNameMapper": { + "^(\\.{1,2}/.*)\\.js$": "$1" } } } diff --git a/cdk/scripts/README.md b/cdk/scripts/README.md index 548533c34..42688a709 100644 --- a/cdk/scripts/README.md +++ b/cdk/scripts/README.md @@ -7,5 +7,10 @@ Bundling for Lambda assets is handled at synth time; the **`bundle`** task in ** | `generate-bootstrap-artifacts.ts` | Regenerates `cdk/bootstrap/policies/*.json`, `BOOTSTRAP_VERSION`, `BOOTSTRAP_HASH` from the typed policies in `src/bootstrap/policies/` | `mise //cdk:bootstrap:generate` | | `generate-bootstrap-template.ts` | Regenerates `cdk/bootstrap/bootstrap-template.yaml` (least-privilege CDK bootstrap, `ComputeTypes`-gated compute policies) | `mise //cdk:bootstrap:generate` | | `package-microvm-artifact.sh` | Packages `agent/` + `contracts/` + `Dockerfile` into the zip artifact an `AWS::Lambda::MicrovmImage` builds from, and uploads it to the CDK-created artifact bucket (ADR-021) | run directly — see the script header for the full bootstrap sequence | +| `build-microvm-artifact.py` | Builds the deterministic zip and digest consumed by the packaging script | Called by `package-microvm-artifact.sh`; `python3 cdk/scripts/build-microvm-artifact.py --help` for local packaging | -`package-microvm-artifact.sh` exists because CloudFormation cannot produce its own MicroVM `codeArtifact`: the image resource consumes a zip that must already be in S3, and there is no CDK asset type for "zip + Dockerfile a MicroVM image builds from". Everything else on that backend (buckets, roles, network connector, log group, the image resource itself) is CDK-managed in `src/constructs/lambda-microvm-compute.ts`. +`package-microvm-artifact.sh` exists because CloudFormation cannot produce its own MicroVM `codeArtifact`: the image resource consumes a zip that must already be in S3, and there is no CDK asset type for "zip + Dockerfile a MicroVM image builds from". Everything else on that backend (buckets, roles, network connectors, log group, the image resource itself) is CDK-managed by `src/constructs/lambda-microvm-compute.ts`, normally inside `lambda-microvm-stack.ts`. + +New installations must explicitly select `microvm_nested_stack=true`; the nested layout requires bootstrap bundle 1.9.0. Before upgrading an +existing flat deployment, set and retain `microvm_nested_stack=false` until completing the +[resource migration](../../docs/verification/645-p3-nested-stack.md). diff --git a/cdk/scripts/build-microvm-artifact.py b/cdk/scripts/build-microvm-artifact.py new file mode 100644 index 000000000..3ff0a9cb8 --- /dev/null +++ b/cdk/scripts/build-microvm-artifact.py @@ -0,0 +1,107 @@ +#!/usr/bin/env python3 +# MIT No Attribution +# +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# +# Permission is hereby granted, free of charge, to any person obtaining a copy of +# the Software without restriction, including without limitation the rights to +# use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of +# the Software, and to permit persons to whom the Software is furnished to do so. +# +# THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR +# IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, +# FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE +# AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER +# LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, +# OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE +# SOFTWARE. + +"""Package the Dockerfile's local inputs into a reproducible MicroVM build ZIP.""" + +import argparse +import base64 +import hashlib +import json +from pathlib import Path +import shutil +import stat +import zipfile + + +# Keep these in step with the local COPY sources in agent/Dockerfile. The +# packaging regression checks both directions; COPY --from uses remote stages. +INPUTS = ( + "agent/pyproject.toml", + "agent/uv.lock", + "agent/src", + "agent/policies", + "agent/workflows", + "agent/prepare-commit-msg.sh", + "agent/managed-settings.json", + "contracts", +) +IGNORED_DIRS = {"__pycache__", ".pytest_cache", ".ruff_cache", ".mypy_cache", "node_modules", ".git"} + + +def source_files(root: Path): + """Yield only regular build inputs; never follow a link outside the tree.""" + def walk(path: Path): + if path.is_symlink(): + raise ValueError(f"MicroVM build inputs must not contain symlinks: {path}") + if path.is_dir(): + for child in sorted(path.iterdir()): + if child.name in IGNORED_DIRS or child.name == ".DS_Store": + continue + if child.suffix in {".pyc", ".pyo"}: + continue + yield from walk(child) + elif path.is_file(): + yield path + else: + raise ValueError(f"Missing or unsupported MicroVM build input: {path}") + + dockerfile = root / "agent/Dockerfile" + if dockerfile.is_symlink() or not dockerfile.is_file(): + raise ValueError("agent/Dockerfile must be a regular file") + yield "Dockerfile", dockerfile + for name in INPUTS: + for path in walk(root / name): + yield path.relative_to(root).as_posix(), path + + +def build(root: Path, output: Path): + files = sorted(source_files(root)) + output.parent.mkdir(parents=True, exist_ok=True) + with zipfile.ZipFile(output, "w", compression=zipfile.ZIP_DEFLATED) as archive: + for name, path in files: + # File dates, checkout locations, uid/gid and umask must not cause + # needless image versions. Preserve just the executable permission. + info = zipfile.ZipInfo(name, date_time=(1980, 1, 1, 0, 0, 0)) + info.create_system = 3 + mode = 0o755 if path.stat().st_mode & 0o111 else 0o644 + info.external_attr = (stat.S_IFREG | mode) << 16 + info.compress_type = zipfile.ZIP_DEFLATED + with path.open("rb") as source, archive.open(info, "w") as target: + shutil.copyfileobj(source, target) + digest = hashlib.sha256() + with output.open("rb") as source: + for block in iter(lambda: source.read(1024 * 1024), b""): + digest.update(block) + return { + "sha256": digest.hexdigest(), + "checksum_sha256": base64.b64encode(digest.digest()).decode("ascii"), + "file_count": len(files), + "size_bytes": output.stat().st_size, + } + + +def main(): + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--repo-root", required=True, type=Path) + parser.add_argument("--output", required=True, type=Path) + args = parser.parse_args() + print(json.dumps(build(args.repo_root.resolve(), args.output.resolve()))) + + +if __name__ == "__main__": + main() diff --git a/cdk/scripts/generate-bootstrap-template.ts b/cdk/scripts/generate-bootstrap-template.ts index a74e07053..f2d945d51 100644 --- a/cdk/scripts/generate-bootstrap-template.ts +++ b/cdk/scripts/generate-bootstrap-template.ts @@ -31,6 +31,7 @@ import { join } from 'node:path'; import * as yaml from 'js-yaml'; +import { nestedStackExecutionPolicy } from '../src/bootstrap/nested-stack-policy'; import { applicationPolicy, computeAgentcorePolicy, @@ -193,7 +194,7 @@ export function buildTemplate(): any { } // --- Step 5: Modify CloudFormationExecutionRole ManagedPolicyArns --- - // Replace the conditional that falls back to AdministratorAccess with our inline policies. + // Replace the conditional that falls back to AdministratorAccess with our managed policies. // Keep the CloudFormationExecutionPolicies parameter override for flexibility. const coreRefs = [ { Ref: 'IaCRoleABCAInfrastructure' }, @@ -218,6 +219,17 @@ export function buildTemplate(): any { ], }; + // Nested stacks inherit the parent execution role. CloudFormation checks that + // the caller can pass it, including during change-set validation (#645). + const executionRole = template.Resources.CloudFormationExecutionRole.Properties; + executionRole.Policies = [ + ...(executionRole.Policies ?? []), + { + PolicyName: 'PassExecutionRoleToCloudFormation', + PolicyDocument: nestedStackExecutionPolicy(), + }, + ]; + // --- Step 6: Add outputs --- template.Outputs.BootstrapPolicyVersion = { Description: 'The version of the ABCA bootstrap policy bundle', @@ -283,6 +295,7 @@ export function renderTemplate(): string { '# - BootstrapVariant set to "ABCA: Least-Privilege Bootstrap"', '# - ComputeTypes parameter added for compute-variant selection', '# - IncludeComputeEcs / IncludeComputeLambdaMicrovms conditions added', + '# - Execution role may pass only itself to CloudFormation for nested stacks', '# - 6 AWS::IAM::ManagedPolicy resources replace AdministratorAccess; each', '# PolicyDocument is a minified JSON string so the template stays under the', '# 51,200-char CloudFormation inline-template limit (#864)', diff --git a/cdk/scripts/package-microvm-artifact.sh b/cdk/scripts/package-microvm-artifact.sh index f825063ca..c0dc13153 100755 --- a/cdk/scripts/package-microvm-artifact.sh +++ b/cdk/scripts/package-microvm-artifact.sh @@ -18,15 +18,23 @@ # --------------------------------------------------------------------------- # BOOTSTRAP SEQUENCE (first time) # --------------------------------------------------------------------------- +# Choose microvm_nested_stack=true for new/already-nested installations, +# or false for existing flat installations, in the commands below. +# +# 0. Use bootstrap policy bundle >= 1.9.0 for the nested layout. +# Before upgrading an existing flat deployment, set and retain +# microvm_nested_stack=false (bundle >= 1.8.0) on every deploy below until +# its resources are migrated; see docs/verification/645-p3-nested-stack.md. +# # 1. Deploy the MicroVM substrate WITHOUT an image. Synth warns that no image # is configured; that is expected — the artifact bucket must exist before # you can upload to it. # -# MISE_EXPERIMENTAL=1 mise //cdk:deploy -- --context compute_type=lambda-microvm +# MISE_EXPERIMENTAL=1 mise //cdk:deploy -- --context compute_type=lambda-microvm --context microvm_nested_stack=REPLACE_WITH_TRUE_OR_FALSE # # 2. Package + upload the artifact (this script). It reads the bucket name and -# object key straight from the stack outputs, so there is nothing to copy -# by hand: +# base object key from stack outputs, uploads an immutable hash-suffixed +# artifact, and prints the digest required by the subsequent deployment: # # cdk/scripts/package-microvm-artifact.sh --stack-name backgroundagent-dev # @@ -43,13 +51,15 @@ # image resource and injects MICROVM_IMAGE_IDENTIFIER into the orchestrator: # # MISE_EXPERIMENTAL=1 mise //cdk:deploy -- \ -# --context compute_type=lambda-microvm \ +# --context compute_type=lambda-microvm --context microvm_nested_stack=REPLACE_WITH_TRUE_OR_FALSE \ # --context microvm_base_image_arn= \ -# --context microvm_base_image_version= +# --context microvm_base_image_version= \ +# --context microvm_artifact_sha256= # -# On subsequent agent changes only step 2 is needed, followed by a CloudFormation -# update of the image resource (the service builds a NEW image version from the -# refreshed artifact). +# On subsequent agent changes repeat step 2 and deploy with the NEW digest. +# Changing the digest changes CodeArtifact.Uri, causing a new image version. +# Retain the digest in the deployment's context for later unrelated deploys; +# overwriting a fixed key alone never signals a CloudFormation image update. # # --------------------------------------------------------------------------- # THE OUT-OF-BAND ALTERNATIVE (--create-image) @@ -62,7 +72,7 @@ # — `run-microvm` rejects a bare name): # # MISE_EXPERIMENTAL=1 mise //cdk:deploy -- \ -# --context compute_type=lambda-microvm \ +# --context compute_type=lambda-microvm --context microvm_nested_stack=REPLACE_WITH_TRUE_OR_FALSE \ # --context microvm_image_identifier= # # --------------------------------------------------------------------------- @@ -74,51 +84,25 @@ # is therefore: # # Dockerfile <- verbatim copy of agent/Dockerfile (MicroVM needs it at the root) -# agent/ <- minus .venv/, __pycache__, test caches +# agent/ <- only local Dockerfile COPY inputs, without bytecode/caches # contracts/ <- cross-language constants the agent reads at runtime # # --------------------------------------------------------------------------- -# !! A P2 IMAGE IS FULLY WIRED, BUT NOT SMOKE-VERIFIED !! +# VERIFICATION # --------------------------------------------------------------------------- -# This script packages and uploads a real artifact, and the image the service -# builds from it will reach ACTIVE, accept a `runHookPayload`, and launch. What -# it does NOT have is any smoke-parity guarantee. -# -# ADR-021 sub-decision 3's hook-phasing table (corrected after the live P1 -# verification run, then completed in P2) is now: +# Test the selected image/coordinator together before enabling automatic sleep. +# See docs/verification/README.md for acceptance criteria and remaining PR checks. # -# /ready, /run declared by the CDK construct AND served by -# the agent in P1 (agent/src/server.py). -# /ready is MANDATORY: create-microvm-image -# refuses any lifecycle hook without it, and an -# image with no hooks at all cannot receive a -# runHookPayload — so "declare /run in P1, -# serve it in P2" was never a reachable state. -# /validate, /terminate declared AND served in P2. /validate is a -# build-time self-check that makes ZERO AWS -# calls (it runs under the build role, which -# holds no Bedrock/Secrets/DynamoDB grants); -# /terminate is a best-effort in-guest -# breadcrumb that must not write terminal task -# status — the orchestrator finalizes the task -# and THEN calls TerminateMicrovm. -# /suspend, /resume P3. A hook the service calls but nothing -# answers fails its lifecycle transition, so -# each is enabled only once it is served. +# Managed images declare all six hooks: +# /ready, /validate Local build-time checks and warm-up; no AWS calls. +# /run Authenticated task bootstrap and asynchronous execution. +# /terminate Close the coding barrier and acknowledge teardown without +# finalizing the task; the worker may be retiring mid-task. +# /suspend, /resume Checkpoint and credential/gate reconciliation. The +# coordinator verifies the launched image's protocol marker +# before allowing automatic sleep. # -# A P2 smoke run HAS now completed clone → change → PR on this substrate -# (2026-08-07: two tasks COMPLETED with pull requests, live progress events, and -# the 45 s agent heartbeat observed). What is still missing is a run with NO manual -# intervention: that smoke needed a live IAM workaround, and the two defects behind -# it (ADR-021 P2r2-F9 / P2r2-F10 — the `iam:PassedToService` condition on both -# `iam:PassRole` paths) are fixed in source but not yet re-exercised live. Keep -# production repos on compute_type=agentcore or ecs until a clean run is on record, -# and note that the CDK-managed image path additionally needs bootstrap policy -# bundle >= 1.6.0 (see the banner after upload). The Dockerfile is the P2 tuned base -# and is copied unmodified — further customization (e.g., Alpine adoption) would be -# P2.5 work. -# -# Requires: awscli v2, zip, rsync, python3 (none of which are installed by this script). +# Requires: awscli v2 with conditional PutObject/checksum support, python3. set -euo pipefail @@ -196,7 +180,7 @@ case " ${SUPPORTED_MEMORY_MIB} " in ;; esac -for tool in aws zip python3 rsync; do +for tool in aws python3; do command -v "$tool" >/dev/null 2>&1 || { echo "error: '$tool' is required but not on PATH" >&2; exit 1; } done @@ -224,6 +208,9 @@ for output in stacks[0].get("Outputs", []): ARTIFACT_BUCKET="$(stack_output MicrovmArtifactBucketName)" ARTIFACT_KEY="$(stack_output MicrovmArtifactObjectKey)" +ARTIFACT_BASE_KEY="$(stack_output MicrovmArtifactBaseObjectKey)" +# Before hashed-artifact support, ObjectKey was always the unsuffixed base. +ARTIFACT_BASE_KEY="${ARTIFACT_BASE_KEY:-${ARTIFACT_KEY}}" BUILD_ROLE_ARN="$(stack_output MicrovmBuildRoleArn)" EGRESS_CONNECTORS="$(stack_output MicrovmEgressConnectorArns)" # BUILD-time connectors (TCP 443 + 80). `agent/Dockerfile` runs `apt-get`, which @@ -247,7 +234,7 @@ error: stack '${STACK_NAME}' has no MicrovmArtifactBucketName/MicrovmArtifactObj That means this stack was not deployed with the lambda-microvm compute backend. Deploy it first: - MISE_EXPERIMENTAL=1 mise //cdk:deploy -- --context compute_type=lambda-microvm + MISE_EXPERIMENTAL=1 mise //cdk:deploy -- --context compute_type=lambda-microvm --context microvm_nested_stack=REPLACE_WITH_TRUE_OR_FALSE EOF exit 1 fi @@ -265,18 +252,14 @@ echo " log group : ${LOG_GROUP}" print_p1_reminder() { cat <<'EOF' -!! REMINDER (ADR-021 P2): smoke-verified ONCE, and only WITH a manual workaround. - The image is creatable and launchable, the agent serves all four declared hooks - (/ready + /validate on the build path, /run + /terminate at runtime), the - execution role holds its full runtime permission set, and a 2026-08-07 run took - two tasks clone -> change -> PR to COMPLETED with a live 45 s heartbeat. - NOT verified: an UNATTENDED run. That smoke needed a live IAM workaround, and the - two defects behind it (ADR-021 P2r2-F9 / P2r2-F10) are fixed in source but not - re-exercised. The CDK-managed image path also needs bootstrap bundle >= 1.6.0. - Keep production repos on compute_type=agentcore or ecs until a clean run is on - record. /suspend and /resume stay disabled until P3. CDK synth emits the same - warning (abca:microvm-image-p1-smoke-unverified) on every deploy that configures - an image. +REMINDER (ADR-021): verify the deployed image and coordinator together before + enabling automatic suspension. Managed images declare all six hooks; + compatible coordinator, IAM and image configuration are required. + Nested deployments require bootstrap bundle 1.9.0 or later. Existing flat + deployments need the staged migration in docs/verification/645-p3-nested-stack.md. + Acceptance criteria and remaining PR checks: docs/verification/README.md. + Preserve a compatible published coordinator and explicit image pin for rollback. + CDK retains warning ID abca:microvm-image-p1-smoke-unverified for compatibility. EOF } @@ -291,69 +274,86 @@ cleanup() { } trap cleanup EXIT -echo "==> Staging agent tree in ${STAGE_DIR}" -# The MicroVM build context is the zip root, and agent/Dockerfile COPYs -# repo-root-relative paths — so the staged tree mirrors the repo, with the -# Dockerfile additionally promoted to the root where the service looks for it. -cp "${REPO_ROOT}/agent/Dockerfile" "${STAGE_DIR}/Dockerfile" - -# Excludes match the build inputs the Dockerfile never COPYs but which dominate -# the zip size: the local virtualenv, Python/pytest caches, and node_modules. -rsync -a \ - --exclude '.venv/' \ - --exclude '__pycache__/' \ - --exclude '.pytest_cache/' \ - --exclude '.ruff_cache/' \ - --exclude '.mypy_cache/' \ - --exclude 'node_modules/' \ - --exclude '*.pyc' \ - "${REPO_ROOT}/agent" "${STAGE_DIR}/" -rsync -a --exclude 'node_modules/' "${REPO_ROOT}/contracts" "${STAGE_DIR}/" - -ARTIFACT_ZIP="${STAGE_DIR}.zip" -rm -f "${ARTIFACT_ZIP}" -echo "==> Zipping to ${ARTIFACT_ZIP}" -( cd "${STAGE_DIR}" && zip -q -r "${ARTIFACT_ZIP}" . ) -echo " $(du -h "${ARTIFACT_ZIP}" | cut -f1) artifact" +ARTIFACT_ZIP="${STAGE_DIR}/agent-artifact.zip" +echo "==> Packaging reproducible build inputs to ${ARTIFACT_ZIP}" +ARTIFACT_INFO="$(python3 "${REPO_ROOT}/cdk/scripts/build-microvm-artifact.py" \ + --repo-root "${REPO_ROOT}" --output "${ARTIFACT_ZIP}")" +ARTIFACT_SHA256="$(python3 -c 'import json,sys; print(json.load(sys.stdin)["sha256"])' <<<"${ARTIFACT_INFO}")" +ARTIFACT_CHECKSUM="$(python3 -c 'import json,sys; print(json.load(sys.stdin)["checksum_sha256"])' <<<"${ARTIFACT_INFO}")" +echo " ${ARTIFACT_INFO}" # --- Upload ------------------------------------------------------------------- +if [[ "${CREATE_IMAGE}" -eq 0 ]]; then + ARTIFACT_KEY="${ARTIFACT_BASE_KEY%.zip}-${ARTIFACT_SHA256}.zip" +else + # The manual create API starts a build explicitly; retain its fixed-key path. + ARTIFACT_KEY="${ARTIFACT_BASE_KEY}" +fi echo "==> Uploading to s3://${ARTIFACT_BUCKET}/${ARTIFACT_KEY}" -aws s3 cp "${ARTIFACT_ZIP}" "s3://${ARTIFACT_BUCKET}/${ARTIFACT_KEY}" -rm -f "${ARTIFACT_ZIP}" +if [[ "${CREATE_IMAGE}" -eq 0 ]]; then + # S3 verifies the bytes against the checksum. Never overwrite a managed build's + # inputs, including a repeated invocation or concurrent publisher. + if aws s3api put-object --bucket "${ARTIFACT_BUCKET}" --key "${ARTIFACT_KEY}" \ + --body "${ARTIFACT_ZIP}" --if-none-match '*' \ + --checksum-algorithm SHA256 --checksum-sha256 "${ARTIFACT_CHECKSUM}" \ + >"${STAGE_DIR}/upload.json" 2>"${STAGE_DIR}/upload.err"; then + echo " Created immutable artifact" + else + UPLOAD_STATUS=$? + case "$(cat "${STAGE_DIR}/upload.err")" in + *"(PreconditionFailed)"*) + EXISTING_CHECKSUM="$(aws s3api head-object --bucket "${ARTIFACT_BUCKET}" \ + --key "${ARTIFACT_KEY}" --checksum-mode ENABLED \ + --query ChecksumSHA256 --output text)" + if [[ "${EXISTING_CHECKSUM}" != "${ARTIFACT_CHECKSUM}" ]]; then + echo "error: existing artifact checksum does not match; refusing to overwrite ${ARTIFACT_KEY}" >&2 + exit 1 + fi + echo " Reusing checksum-verified existing artifact" + ;; + *) + cat "${STAGE_DIR}/upload.err" >&2 + exit "${UPLOAD_STATUS}" + ;; + esac + fi +else + aws s3api put-object --bucket "${ARTIFACT_BUCKET}" --key "${ARTIFACT_KEY}" \ + --body "${ARTIFACT_ZIP}" \ + --checksum-algorithm SHA256 --checksum-sha256 "${ARTIFACT_CHECKSUM}" +fi if [[ "${CREATE_IMAGE}" -eq 0 ]]; then cat < Artifact uploaded. Next: create (or update) the image. - CDK-managed (recommended) — redeploy with the base image pinned. + CDK-managed (recommended) — redeploy with the base image and artifact pinned. + Keep this digest with the deployment's context; it identifies these exact ZIP bytes. - !! RE-BOOTSTRAP REQUIRED (bootstrap policy bundle >= 1.6.0) !! - This path took two live-verified fixes to work. The first (ADR-021 P2-F2: the L1 - sent hook paths and \`arm64\` where CloudFormation wants ENABLED / ARM_64) is - DISCHARGED — change-set early validation now passes. The second (ADR-021 - P2r2-F9) is a BOOTSTRAP change: CloudFormation could not pass the MicroVM build - role, because the deploy role's \`iam:PassRole\` carried an - \`iam:PassedToService\` condition the Lambda MicroVMs service does not satisfy. - The fix is the \`MicrovmPassRoles\` statement in the conditional - IaCRole-ABCA-Compute-LambdaMicrovms policy, which only reaches your account when - you re-bootstrap: + The nested layout requires bootstrap policy bundle >= 1.9.0. + Flat P3 deployments require >= 1.8.0. Check the installed bundle: aws cloudformation describe-stacks --stack-name CDKToolkit \\ --query 'Stacks[0].Outputs[?OutputKey==\`BootstrapPolicyVersion\`].OutputValue' --output text - # if that is below 1.6.0: MISE_EXPERIMENTAL=1 mise //cdk:bootstrap # ComputeTypes must include lambda-microvm - Without it the image resource fails with - "is not authorized to perform: iam:PassRole on resource: - ...LambdaMicrovmComputeBuildRole... (Service: LambdaMicrovms, Status Code: 403)". - Then: + An older bundle may deny iam:PassRole for the image's build role. Updating the + source bundle alone does not update the account's installed policies. + + Existing flat stacks: set --context microvm_nested_stack=false before upgrading + and retain it until the resource migration is complete. Changing the layout directly + can replace resources or delete bucket contents. See: + docs/verification/645-p3-nested-stack.md + + For a new nested deployment (or a completed migration): aws lambda-microvms list-managed-microvm-images MISE_EXPERIMENTAL=1 mise //cdk:deploy -- \\ - --context compute_type=lambda-microvm \\ + --context compute_type=lambda-microvm --context microvm_nested_stack=REPLACE_WITH_TRUE_OR_FALSE \\ --context microvm_base_image_arn= \\ - --context microvm_base_image_version= + --context microvm_base_image_version= \\ + --context microvm_artifact_sha256=${ARTIFACT_SHA256} Out of band — re-run this script with: @@ -404,12 +404,11 @@ echo "==> Creating MicroVM image '${IMAGE_NAME}' (${MEMORY_MIB} MiB baseline)" # * `/ready` is MANDATORY whenever any lifecycle hook is enabled: # "The ready (/ready) MicroVM image hook must be enabled when any MicroVM # lifecycle hook (run, resume, suspend, or terminate) is enabled." -# * all four hooks the agent serves are enabled: `/ready` + `/validate` (build) -# and `/run` + `/terminate` (runtime). `/suspend` and `/resume` stay DISABLED -# until P3 implements them — a hook the service calls but nothing answers -# fails the corresponding build or lifecycle transition. +# * all six served hooks are enabled: `/ready` + `/validate` (build), `/run`, +# `/terminate`, `/suspend` and `/resume` (runtime). The non-secret lifecycle +# protocol marker describes the source's checkpoint/credential/gate protocol. # * the timeouts mirror the construct's constants -# (`RUN_/READY_/VALIDATE_/TERMINATE_HOOK_TIMEOUT_SECONDS` in +# (`RUN_/READY_/VALIDATE_/TERMINATE_/LIFECYCLE_HOOK_TIMEOUT_SECONDS` in # `cdk/src/constructs/lambda-microvm-compute.ts`), which carry the rationale # for each value. A bash helper cannot import them, and "keep the two in step" # as prose already FAILED once — `readyTimeoutInSeconds` stayed at 60 here when @@ -436,7 +435,8 @@ CREATE_RESPONSE="$(aws lambda-microvms create-microvm-image \ --resources "[{\"minimumMemoryInMiB\":${MEMORY_MIB}}]" \ --egress-network-connectors "${BUILD_EGRESS_CONNECTORS}" \ --logging "{\"cloudWatch\":{\"logGroup\":\"${LOG_GROUP}\"}}" \ - --hooks '{"port":8080,"microvmHooks":{"run":"ENABLED","runTimeoutInSeconds":60,"terminate":"ENABLED","terminateTimeoutInSeconds":15},"microvmImageHooks":{"ready":"ENABLED","readyTimeoutInSeconds":300,"validate":"ENABLED","validateTimeoutInSeconds":60}}' \ + --hooks '{"port":8080,"microvmHooks":{"run":"ENABLED","runTimeoutInSeconds":60,"terminate":"ENABLED","terminateTimeoutInSeconds":15,"suspend":"ENABLED","suspendTimeoutInSeconds":30,"resume":"ENABLED","resumeTimeoutInSeconds":30},"microvmImageHooks":{"ready":"ENABLED","readyTimeoutInSeconds":300,"validate":"ENABLED","validateTimeoutInSeconds":60}}' \ + --environment-variables '{"ABCA_MICROVM_LIFECYCLE_PROTOCOL":"1"}' \ --tags "abca:compute-backend=lambda-microvm" \ --output json)" @@ -480,7 +480,7 @@ cat < = { + 'AWS::SSM::Parameter': ['ssm:PutParameter', 'ssm:GetParameters', 'ssm:DeleteParameter'], 'AWS::ApiGateway::Authorizer': ['apigateway:POST'], 'AWS::ApiGateway::Method': ['apigateway:POST'], 'AWS::ApiGateway::RequestValidator': ['apigateway:POST'], diff --git a/cdk/src/bootstrap/version.ts b/cdk/src/bootstrap/version.ts index 2bdb6caac..a2634e16e 100644 --- a/cdk/src/bootstrap/version.ts +++ b/cdk/src/bootstrap/version.ts @@ -19,6 +19,7 @@ import { createHash } from 'node:crypto'; +import { nestedStackExecutionPolicy } from './nested-stack-policy'; import { allPolicies } from './policies'; /** @@ -44,25 +45,45 @@ import { allPolicies } from './policies'; * `policies/compute-lambda-microvm.ts`), which is exactly the class of breakage a * patch-level bump would under-advertise. * - * It is 1.6.0 and not 1.4.0 because #629 and #246 landed on `main` first and took + * The 1.6.0 release was not 1.4.0 because #629 and #246 landed on `main` first and took * 1.4.0 and 1.5.0 in the interim. The number is an operator-visible contract (the * `CDKToolkit` stack's `BootstrapPolicyVersion` output), so re-using a published * version would make two different bundles indistinguishable to the `>=` check * operators are told to run. + * + * 1.6.0 → 1.7.0 adds an exact-self, CloudFormation-only PassRole inline policy + * for nested stacks (#645 clean deployment). The policy hash now includes that + * inline policy and every nested JSON field; the old root-key replacer omitted + * Action/Resource/Condition changes from its input. + * + * 1.7.0 → 1.8.0 adds scoped SSM parameter lifecycle/tag permissions for the P3 + * MicroVM suspension switch. Re-bootstrap before deploying the live parameter. + * + * 1.8.0 → 1.9.0 admits the exact nested MicroVM build/operator role names to + * the backend-specific PassRole statement. Re-bootstrap before the nested split. */ -export const BOOTSTRAP_VERSION = '1.6.0'; +export const BOOTSTRAP_VERSION = '1.9.0'; + +function canonicalize(value: unknown): unknown { + if (Array.isArray(value)) return value.map(canonicalize); + if (value !== null && typeof value === 'object') { + const entries = value as Record; + return Object.fromEntries( + Object.keys(entries).sort().map((key) => [key, canonicalize(entries[key])]), + ); + } + return value; +} /** - * Computes a SHA-256 hash over all bootstrap policies. - * The hash is deterministic: policies are serialized with sorted keys - * so that object property ordering does not affect the digest. + * Hashes all ABCA managed and inline execution policies. Sort object keys + * recursively without dropping nested fields; preserve array ordering. */ export function computeBootstrapHash(): string { - const policies = allPolicies(); - const normalized = policies.map((p) => { - const json = p.toJSON(); - return JSON.stringify(json, Object.keys(json).sort()); - }); - const payload = JSON.stringify(normalized); + const policies = [ + ...allPolicies().map((policy) => policy.toJSON()), + nestedStackExecutionPolicy(), + ]; + const payload = JSON.stringify(canonicalize(policies)); return createHash('sha256').update(payload).digest('hex'); } diff --git a/cdk/src/constructs/agent-session-role.ts b/cdk/src/constructs/agent-session-role.ts index 1602b734d..1c9453d58 100644 --- a/cdk/src/constructs/agent-session-role.ts +++ b/cdk/src/constructs/agent-session-role.ts @@ -24,6 +24,77 @@ import * as iam from 'aws-cdk-lib/aws-iam'; import * as s3 from 'aws-cdk-lib/aws-s3'; import { NagSuppressions } from 'cdk-nag'; import { Construct } from 'constructs'; +import agentTaskWriteAttributes from './agent-task-write-attributes.json'; +import constants from '../../../contracts/constants.json'; + +/** + * Task reporting may update only the attributes written by task_state.py. + * Keep whole-row replacement/deletion and coordinator metadata out of this + * grant. DynamoDB evaluates each transaction item using its item action, so + * approval-related TaskTable updates receive the same restriction. + * + * This protects writes, not reads: agents may read their complete task record. + * The JSON list is also checked against actual Python writer requests in tests. + * Used for both scoped sessions and the legacy ECS direct-grant fallback. + */ +export function grantAgentTaskTableAccess( + table: dynamodb.ITable, + grantee: iam.IGrantable, + taskScoped: boolean, +): void { + const leadingKeys = taskScoped + ? { 'dynamodb:LeadingKeys': ['${aws:PrincipalTag/task_id}'] } + : {}; + const readLeadingKeys = taskScoped + ? { + 'dynamodb:LeadingKeys': [ + '${aws:PrincipalTag/task_id}', `${constants.microvm_continuation.lease_key_prefix}\${aws:PrincipalTag/task_id}`, + ], + } + : {}; + grantee.grantPrincipal.addToPrincipalPolicy(new iam.PolicyStatement({ + actions: ['dynamodb:GetItem', 'dynamodb:BatchGetItem', 'dynamodb:Query', 'dynamodb:ConditionCheckItem'], + resources: [table.tableArn], + ...(taskScoped ? { + conditions: { + 'ForAllValues:StringEquals': readLeadingKeys, + 'Null': { 'dynamodb:LeadingKeys': 'false' }, + }, + } : {}), + })); + grantee.grantPrincipal.addToPrincipalPolicy(new iam.PolicyStatement({ + actions: ['dynamodb:UpdateItem'], + resources: [table.tableArn], + conditions: { + 'ForAllValues:StringEquals': { + ...leadingKeys, + 'dynamodb:Attributes': agentTaskWriteAttributes, + }, + // ForAllValues alone also matches an absent context key. Require the + // attribute list (and session key when scoped) to be present. + 'Null': { + 'dynamodb:Attributes': 'false', + ...(taskScoped ? { 'dynamodb:LeadingKeys': 'false' } : {}), + }, + }, + })); +} + +/** Workers observe decisions; only the control plane writes approval records. */ +export function grantAgentApprovalReadAccess( + table: dynamodb.ITable, grantee: iam.IGrantable, taskScoped: boolean, +): void { + grantee.grantPrincipal.addToPrincipalPolicy(new iam.PolicyStatement({ + actions: ['dynamodb:GetItem', 'dynamodb:BatchGetItem', 'dynamodb:Query', 'dynamodb:ConditionCheckItem'], + resources: [table.tableArn], + ...(taskScoped ? { + conditions: { + 'ForAllValues:StringEquals': { 'dynamodb:LeadingKeys': ['${aws:PrincipalTag/task_id}'] }, + 'Null': { 'dynamodb:LeadingKeys': 'false' }, + }, + } : {}), + })); +} /** S3 key prefixes the agent writes/reads, scoped per tenant. */ const TRACE_KEY_PREFIX = 'traces'; @@ -36,17 +107,26 @@ const ARTIFACT_KEY_PREFIX = 'artifacts'; */ export interface AgentSessionRoleProps { /** - * Compute roles (AgentCore Runtime ExecutionRole and/or ECS Fargate task - * role) permitted to assume this SessionRole and pass session tags. These - * are the only principals trusted to mint scoped credentials, so they bound - * the trust surface. Both run the same trusted agent code, which sources the + * Compute roles (AgentCore Runtime, ECS Fargate or Lambda MicroVM) permitted + * to assume this SessionRole and pass session tags. These principals mint + * scoped credentials. The agent code sources the * `{user_id, repo, task_id}` tag values from the resolved TaskConfig. */ readonly assumingRoles: iam.IRole[]; /** - * The four task-scoped DynamoDB tables, all partitioned by `task_id`. The - * SessionRole receives item-level access constrained by a + * The main task table: own-task reads and attribute-scoped reporting updates. + * Task creation/deletion, owner identity, compute handles, start receipts and + * capacity reservations belong to the coordinator. + */ + readonly taskTable: dynamodb.ITable; + /** Decisions and notification markers are control-plane-owned. */ + readonly approvalsTable: dynamodb.ITable; + + /** + * Supporting task-scoped tables (events, nudges), all partitioned + * by `task_id`. Do not include taskTable or approvalsTable here: that would + * bypass their write restrictions. The SessionRole receives access constrained by a * `dynamodb:LeadingKeys` condition on `aws:PrincipalTag/task_id`, so a * session can only touch its own task's rows. Order is irrelevant. */ @@ -76,9 +156,10 @@ export interface AgentSessionRoleProps { * subprocess is then attributed per `{user_id, repo}` in CUR 2.0 / Cost * Explorer via the session tags this role already carries. * - * The compute role keeps its own Bedrock grant: attribution is a billing - * control that fails open (the credential helper falls back to compute-role - * creds if the assume-role fails), so model invocation never depends on this. + * The compute role keeps its own Bedrock grant. On AgentCore/ECS, the export + * helper can fall back to compute-role credentials if attribution fails. + * MicroVM workers instead use the retained task-scoped credential provider; + * failed renewal blocks wake rather than falling back to ambient credentials. * Omit (e.g. isolated construct tests) to skip the Bedrock grant. */ readonly invokableModels?: bedrock.IBedrockInvokable[]; @@ -98,17 +179,20 @@ export interface AgentSessionRoleProps { * - S3 trace writes and attachment reads are scoped to the * `/${aws:PrincipalTag/user_id}/` object prefix. * - * The result: a compromised agent session can reach only its own task's data, - * not other tenants' — enforced at the IAM layer rather than in application - * code. Backend-agnostic: the same role serves agents booted under either the - * AgentCore Runtime execution role or the ECS Fargate task role. + * Existing session credentials are limited to their tagged task. Compute roles + * choose these tags when assuming this role; the trust policy does not bind + * those choices to a particular task. This is not an isolation claim for a + * compromised worker that can obtain ambient compute credentials. + * TaskTable writes additionally exclude coordinator-owned attributes and + * whole-row replacement/deletion. All three compute backends share this role. * - * CloudWatch Logs remains on the compute role (shared, non-tenant access). The + * CloudWatch Logs remains on the compute role (shared access). The * compute role *also* keeps `InvokeModel`; this role adds a parallel, session- * tagged Bedrock grant (#215) used by the Claude Code subprocess for cost - * attribution. Long-task safety on the 1-hour-capped chained session is handled - * by Claude Code's `awsCredentialExport` refresh, and the helper falls back to - * the compute role if assume fails — so model invocation never breaks. + * attribution. AgentCore/ECS use `awsCredentialExport` with expiry and an ambient + * fallback for attribution failures. MicroVM uses the parent's retained scoped + * provider with synchronous renewal; the export helper returns no credentials. + * The background export refresh is not a MicroVM wake barrier. */ export class AgentSessionRole extends Construct { /** Actions sufficient for the agent's DynamoDB access. Excludes Scan. */ @@ -136,6 +220,10 @@ export class AgentSessionRole extends Construct { 'AgentSessionRole requires at least one assuming role (the compute role[s] that mint scoped credentials)', ); } + if (props.taskScopedTables.some((table) => + [props.taskTable.tableArn, props.approvalsTable.tableArn].includes(table.tableArn))) { + throw new Error('taskTable and approvalsTable must not appear in taskScopedTables; they require restricted writes'); + } const [firstAssumingRole] = props.assumingRoles; @@ -153,7 +241,10 @@ export class AgentSessionRole extends Construct { maxSessionDuration: Duration.hours(1), }); - // --- DynamoDB: item access gated by task_id leading-key --- + grantAgentTaskTableAccess(props.taskTable, this.role, true); + grantAgentApprovalReadAccess(props.approvalsTable, this.role, true); + + // --- Supporting tables: item access gated by task_id leading-key --- // One statement per table keeps the resource ARNs explicit. The condition // requires the request's partition key (task_id) to equal the session's // task_id tag. ForAllValues is required by DynamoDB for LeadingKeys. @@ -166,6 +257,7 @@ export class AgentSessionRole extends Construct { 'ForAllValues:StringEquals': { 'dynamodb:LeadingKeys': ['${aws:PrincipalTag/task_id}'], }, + 'Null': { 'dynamodb:LeadingKeys': 'false' }, }, }), ); @@ -216,8 +308,8 @@ export class AgentSessionRole extends Construct { // Reuse grantInvoke so this role's Bedrock permissions exactly mirror the // compute role's (cross-region profiles fan out to the foundation model in // every routed region — replicating that by hand would risk an AccessDenied - // on a cross-region route). Claude Code assumes this role (via its - // awsCredentialExport helper) so InvokeModel rides the session's + // on a cross-region route). Claude Code uses this role through the export + // helper or MicroVM's scoped provider, so InvokeModel rides the session's // {user_id, repo, task_id} tags, surfacing per-user/repo Bedrock spend in // CUR 2.0 / Cost Explorer. No PrincipalTag condition: the tags are for // billing attribution, not access scoping, so a condition would add no @@ -238,6 +330,7 @@ export class AgentSessionRole extends Construct { 'Resource wildcards are the per-object suffix under a tenant-scoped ' + 'prefix (traces/${aws:PrincipalTag/user_id}/*, ' + 'attachments/${aws:PrincipalTag/user_id}/*, ' + + 'continuations/${aws:PrincipalTag/task_id}/*, ' + 'artifacts/${aws:PrincipalTag/task_id}/*) and the DynamoDB item ' + 'set gated by a dynamodb:LeadingKeys = ${aws:PrincipalTag/task_id} ' + 'condition — narrower than the compute role this replaces. Bedrock ' diff --git a/cdk/src/constructs/agent-task-write-attributes.json b/cdk/src/constructs/agent-task-write-attributes.json new file mode 100644 index 000000000..49e4c9ff6 --- /dev/null +++ b/cdk/src/constructs/agent-task-write-attributes.json @@ -0,0 +1,29 @@ +[ + "task_id", + "status", + "status_created_at", + "started_at", + "completed_at", + "logs_url", + "agent_heartbeat_at", + "awaiting_approval_request_id", + "approval_gate_count", + "continuation", + "pr_url", + "error_message", + "cost_usd", + "duration_s", + "turns", + "turns_attempted", + "turns_completed", + "prompt_version", + "memory_written", + "build_passed", + "lint_passed", + "code_changed", + "head_sha", + "answer_text", + "otel_trace_id", + "trace_s3_uri", + "artifact_uri" +] diff --git a/cdk/src/constructs/approval-request-service.ts b/cdk/src/constructs/approval-request-service.ts new file mode 100644 index 000000000..40784c302 --- /dev/null +++ b/cdk/src/constructs/approval-request-service.ts @@ -0,0 +1,103 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import * as path from 'node:path'; +import { Duration, NestedStack } from 'aws-cdk-lib'; +import * as apigw from 'aws-cdk-lib/aws-apigateway'; +import type * as dynamodb from 'aws-cdk-lib/aws-dynamodb'; +import * as iam from 'aws-cdk-lib/aws-iam'; +import { Architecture, Runtime } from 'aws-cdk-lib/aws-lambda'; +import { NodejsFunction } from 'aws-cdk-lib/aws-lambda-nodejs'; +import * as logs from 'aws-cdk-lib/aws-logs'; +import { NagSuppressions } from 'cdk-nag'; +import type { Construct } from 'constructs'; + +const REQUEST_TIMEOUT_SECONDS = 15; +/** + * Trusted approval writer. A separate API keeps routes inside this child stack. + * IAM binds a signed POST path to the session's task tag; direct Lambda invoke + * would lose that binding because Lambda does not expose caller session tags. + */ +export class ApprovalRequestService extends NestedStack { + public readonly api: apigw.RestApi; + public readonly fn: NodejsFunction; + + constructor(scope: Construct, id: string, props: { + taskTable: dynamodb.ITable; + approvalsTable: dynamodb.ITable; + }) { + super(scope, id); + this.fn = new NodejsFunction(this, 'RequestFn', { + entry: path.join(__dirname, '..', 'handlers', 'request-approval.ts'), + handler: 'handler', + runtime: Runtime.NODEJS_24_X, + architecture: Architecture.ARM_64, + timeout: Duration.seconds(REQUEST_TIMEOUT_SECONDS), + memorySize: 256, + environment: { + ABCA_COMPONENT: 'approval', + TASK_TABLE_NAME: props.taskTable.tableName, + TASK_APPROVALS_TABLE_NAME: props.approvalsTable.tableName, + }, + }); + this.fn.addToRolePolicy(new iam.PolicyStatement({ + actions: ['dynamodb:GetItem', 'dynamodb:UpdateItem', 'dynamodb:ConditionCheckItem'], + resources: [props.taskTable.tableArn], + })); + this.fn.addToRolePolicy(new iam.PolicyStatement({ + actions: ['dynamodb:PutItem', 'dynamodb:UpdateItem'], + resources: [props.approvalsTable.tableArn], + })); + const accessLogs = new logs.LogGroup(this, 'AccessLogs', { retention: logs.RetentionDays.ONE_MONTH }); + this.api = new apigw.RestApi(this, 'Api', { + description: 'IAM-authenticated worker approval requests; no human decision endpoint.', + // The parent TaskApi configures the regional API Gateway logging role. + cloudWatchRole: false, + deployOptions: { + stageName: 'v1', + accessLogDestination: new apigw.LogGroupLogDestination(accessLogs), + accessLogFormat: apigw.AccessLogFormat.jsonWithStandardFields(), + loggingLevel: apigw.MethodLoggingLevel.INFO, + throttlingRateLimit: 60, + throttlingBurstLimit: 100, + }, + }); + this.api.root.addResource('tasks').addResource('{task_id}').addMethod( + 'POST', new apigw.LambdaIntegration(this.fn), + { authorizationType: apigw.AuthorizationType.IAM }, + ); + NagSuppressions.addResourceSuppressions(this.fn, [{ + id: 'AwsSolutions-IAM4', reason: 'AWSLambdaBasicExecutionRole provides Lambda runtime logging.', + }], true); + NagSuppressions.addResourceSuppressions(this.api, [ + { id: 'AwsSolutions-APIG2', reason: 'The handler validates request shape and allowed fields. Action descriptions and policy metadata remain worker assertions, not independently verified policy results.' }, + { id: 'AwsSolutions-APIG3', reason: 'Machine-only IAM-signed API; session policy restricts POST to its tagged task path, with stage throttling.' }, + { id: 'AwsSolutions-COG4', reason: 'Workers authenticate with task-scoped AWS credentials, not human Cognito credentials.' }, + ], true); + } + + /** No wildcard task path on production session credentials. */ + public grantRequests(grantee: iam.IGrantable): void { + grantee.grantPrincipal.addToPrincipalPolicy(new iam.PolicyStatement({ + actions: ['execute-api:Invoke'], + resources: [this.api.arnForExecuteApi('POST', '/tasks/${aws:PrincipalTag/task_id}', 'v1')], + conditions: { Null: { 'aws:PrincipalTag/task_id': 'false' } }, + })); + } +} diff --git a/cdk/src/constructs/concurrency-reconciler.ts b/cdk/src/constructs/concurrency-reconciler.ts index c0fed93dd..b28a9e1b6 100644 --- a/cdk/src/constructs/concurrency-reconciler.ts +++ b/cdk/src/constructs/concurrency-reconciler.ts @@ -41,7 +41,7 @@ const RECONCILER_MEMORY_MB = 256; */ export interface ConcurrencyReconcilerProps { /** - * The DynamoDB task table (has UserStatusIndex GSI). + * The DynamoDB task table (reservations are scanned from the base table). */ readonly taskTable: dynamodb.ITable; @@ -59,8 +59,8 @@ export interface ConcurrencyReconcilerProps { /** * Scheduled Lambda that reconciles user concurrency counters by comparing - * the active_count in the concurrency table against actual active tasks - * in the task table. Corrects drift caused by orchestrator crashes. + * active_count against saved task reservations, including approval waits. + * Releases terminal reservations left behind by interrupted cleanup. */ export class ConcurrencyReconciler extends Construct { public readonly fn: lambda.NodejsFunction; @@ -89,6 +89,7 @@ export class ConcurrencyReconciler extends Construct { }); props.taskTable.grantReadData(this.fn); + props.taskTable.grant(this.fn, 'dynamodb:UpdateItem'); props.userConcurrencyTable.grantReadWriteData(this.fn); const schedule = props.schedule ?? Duration.minutes(DEFAULT_SCHEDULE_MINUTES); diff --git a/cdk/src/constructs/continuation-bucket.ts b/cdk/src/constructs/continuation-bucket.ts new file mode 100644 index 000000000..777852d9a --- /dev/null +++ b/cdk/src/constructs/continuation-bucket.ts @@ -0,0 +1,73 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { Duration, RemovalPolicy } from 'aws-cdk-lib'; +import * as iam from 'aws-cdk-lib/aws-iam'; +import * as s3 from 'aws-cdk-lib/aws-s3'; +import { NagSuppressions } from 'cdk-nag'; +import { Construct } from 'constructs'; +import constants from '../../../contracts/constants.json'; + +/** Durable task recovery data. Active checkpoints must not expire with debug traces. */ +export class ContinuationBucket extends Construct { + public readonly bucket: s3.Bucket; + + constructor(scope: Construct, id: string) { + super(scope, id); + this.bucket = new s3.Bucket(this, 'Bucket', { + versioned: true, + blockPublicAccess: s3.BlockPublicAccess.BLOCK_ALL, + encryption: s3.BucketEncryption.S3_MANAGED, + enforceSSL: true, + removalPolicy: RemovalPolicy.RETAIN, + lifecycleRules: [{ + id: 'abandoned-multipart-uploads', + abortIncompleteMultipartUploadAfter: Duration.days(1), + }], + }); + NagSuppressions.addResourceSuppressions(this.bucket, [{ + id: 'AwsSolutions-S1', + reason: 'Recovery data is private and task-scoped through IAM; coordinator cleanup runs after task closure. No public or anonymous access path. Server access logging is not enabled for task storage.', + }]); + } + + /** A worker can save/read its own task versions, never list or remove data. */ + public grantWorker(grantee: iam.IGrantable): void { + grantee.grantPrincipal.addToPrincipalPolicy(new iam.PolicyStatement({ + actions: ['s3:GetObject', 's3:GetObjectVersion', 's3:PutObject'], + resources: [this.bucket.arnForObjects( + `${constants.microvm_continuation.object_key_prefix}\${aws:PrincipalTag/task_id}/*`, + )], + })); + } + + /** Coordinator owns launch inputs and cleanup after a task closes. */ + public grantCoordinator(grantee: iam.IGrantable): void { + const prefix = constants.microvm_continuation.object_key_prefix; + grantee.grantPrincipal.addToPrincipalPolicy(new iam.PolicyStatement({ + actions: ['s3:GetObject', 's3:GetObjectVersion', 's3:PutObject', 's3:DeleteObject', 's3:DeleteObjectVersion'], + resources: [this.bucket.arnForObjects(`${prefix}*`)], + })); + grantee.grantPrincipal.addToPrincipalPolicy(new iam.PolicyStatement({ + actions: ['s3:ListBucketVersions'], + resources: [this.bucket.bucketArn], + conditions: { StringLike: { 's3:prefix': `${prefix}*` } }, + })); + } +} diff --git a/cdk/src/constructs/ecs-agent-cluster.ts b/cdk/src/constructs/ecs-agent-cluster.ts index 7e487f2bc..1cdbf78cb 100644 --- a/cdk/src/constructs/ecs-agent-cluster.ts +++ b/cdk/src/constructs/ecs-agent-cluster.ts @@ -29,7 +29,7 @@ import * as secretsmanager from 'aws-cdk-lib/aws-secretsmanager'; import { NagSuppressions } from 'cdk-nag'; import { Construct, type Node } from 'constructs'; import { AgentMemory } from './agent-memory'; -import { AgentSessionRole } from './agent-session-role'; +import { AgentSessionRole, grantAgentTaskTableAccess } from './agent-session-role'; import { PLATFORM_DEFAULT_AUX_MODEL_ID, PLATFORM_DEFAULT_MODEL_ID, @@ -38,6 +38,7 @@ import { resolveBedrockModelIds, } from './bedrock-models'; import { LinearIdentityVault } from './linear-identity-vault'; +import { grantWorkerBootstrap } from './payload-bootstrap-permissions'; import { buildAppId } from './solution-ua-aspect'; import { ToolGateway } from './tool-gateway'; @@ -46,6 +47,8 @@ export interface EcsAgentClusterProps { readonly agentImageAsset: ecr_assets.DockerImageAsset; readonly taskTable: dynamodb.ITable; readonly taskEventsTable: dynamodb.ITable; + /** Approval storage. Required for human approval gates; optional in isolated tests. */ + readonly taskApprovalsTable?: dynamodb.ITable; readonly userConcurrencyTable: dynamodb.ITable; readonly githubTokenSecret: secretsmanager.ISecret; readonly memoryId?: string; @@ -58,14 +61,12 @@ export interface EcsAgentClusterProps { readonly taskSizing?: EcsTaskSizing; /** - * S3 bucket holding per-task ECS payloads. The orchestrator writes the - * payload (incl. the large hydrated_context, which can't fit in the 8 KB - * RunTask containerOverrides limit) here and passes only an - * `AGENT_PAYLOAD_S3_URI` pointer; the container fetches it on boot. The task - * role gets **read-only** on this bucket — the container runs untrusted repo - * code, so it must not be able to delete payloads (the trusted orchestrator - * owns write + delete). When omitted (isolated construct tests / deployments - * that still pass the payload inline), no grant or env var is added. + * S3 storage for deployment manifests, task payloads and private launch + * references. The v2 coordinator sends AGENT_PAYLOAD_REF, containing a + * single-object signed download URL. The task role can read only bootstrap/*; + * object reads elsewhere and bucket listing are explicitly denied. The + * coordinator owns writes and cleanup. Optional for isolated construct tests; + * production ECS launches require this bucket and a matching v2 image. */ readonly payloadBucket?: s3.IBucket; @@ -98,6 +99,7 @@ export interface EcsAgentClusterProps { * retains the direct grants. */ readonly agentSessionRole?: AgentSessionRole; + readonly approvalRequestsApiUrl?: string; /** * AgentCore Memory for cross-task learning. When provided, the ECS task role @@ -151,8 +153,8 @@ const HTTPS_PORT = 443; * ~3.1 GB of the 16 GB, because ``MISE_JOBS=1`` serialises the packages so peak * is max-single-package rather than sum-of-all. Nearly 5x headroom. * - * Disk is the tighter constraint and is sized less aggressively for that - * reason: the same build peaked at ~14.7 GiB, so Fargate's 21 GiB floor leaves + * Disk is the tighter constraint: the same build peaked at ~14.7 GiB, so + * Fargate's 20 GiB default leaves * only ~1.4x — a heavier dependency cache or a second build sharing the task * would run it out of space and surface as a spurious build failure. 50 GiB * restores real margin, and ephemeral storage is a small fraction of the @@ -170,16 +172,9 @@ const HTTPS_PORT = 443; * execution role, because splitting them is how a grant silently lands on one * def and not the other. Do not read the name as a privilege boundary. */ -// A MODEST default: 4 vCPU / 16 GB, and Fargate's own 20 GiB disk. -// -// Deliberately not the Fargate ceiling. A default is what an adopter who changes -// nothing gets, and at 16 vCPU / 120 GB that is roughly 5x the per-build cost of -// this size in us-east-1 on-demand. Under-provisioning surfaces as a slow or -// OOM-ing build, which is diagnosable and fixable with one prop; over-provisioning -// surfaces as a bill, which is not. A large TypeScript + Python monorepo genuinely -// needs more — raise it through {@link EcsTaskSizing}, up to Fargate's 16 vCPU / -// 120 GB maximum, and raise ephemeral storage with it if concurrent builds run the -// disk out of space. +// Build defaults: 4 vCPU / 16 GiB RAM / 50 GiB disk. +// Planning defaults: 2 vCPU / 8 GiB RAM / Fargate's 20 GiB default disk. +// EcsTaskSizing overrides these values for the target repository's workload. const DEFAULT_BUILD_TASK_CPU = 4096; const DEFAULT_BUILD_TASK_MEMORY_MIB = 16384; const DEFAULT_BUILD_TASK_EPHEMERAL_STORAGE_GIB = 50; @@ -189,7 +184,7 @@ const DEFAULT_PLANNING_TASK_MEMORY_MIB = 8192; /** * Per-task Fargate sizing overrides. Every field is optional; anything left * unset uses the default above. A consumer with a lighter repo should shrink the - * build task (for example 4 vCPU / 16 GB) to cut cost; a heavy monorepo can keep + * build task to cut cost; a heavy monorepo can keep * or raise it up to the Fargate ceiling of 16 vCPU / 120 GB. Values are passed * straight to the Fargate task definition, so they must be a valid Fargate * cpu/memory combination (see the AWS Fargate docs) — an invalid pair fails at @@ -245,8 +240,10 @@ export interface EcsTaskSizing { * override key should fail synth, not disable an isolation control. */ const RESERVED_BUILD_ENV_KEYS = new Set([ + 'APPROVAL_REQUESTS_API_URL', 'TASK_TABLE_NAME', 'TASK_EVENTS_TABLE_NAME', + 'TASK_APPROVALS_TABLE_NAME', 'USER_CONCURRENCY_TABLE_NAME', 'LOG_GROUP_NAME', 'GITHUB_TOKEN_SECRET_ARN', @@ -311,12 +308,10 @@ export class EcsAgentCluster extends Construct { public readonly taskDefinition: ecs.FargateTaskDefinition; /** * The smaller read-only PLANNING task def (8 GB / 2 vCPU) — for any read-only - * workflow that clones + reads + emits an artifact but never builds. Same - * image/role/env/grants as the build def (shared task+execution role + a shared - * container spec, so a grant present on one def but missing on the other can't - * silently diverge); the ONLY difference is cpu/mem. The orchestrator selects - * this for read-only workflows on an ECS repo, so planning doesn't - * over-allocate the large build task. + * workflow that clones + reads + emits an artifact but never builds. Both + * definitions share their image, roles, grants and platform environment. + * Sizing, disk and build-tool settings differ. The orchestrator selects this + * definition for read-only workflows on an ECS repo. */ public readonly planningTaskDefinition: ecs.FargateTaskDefinition; public readonly securityGroup: ec2.SecurityGroup; @@ -327,6 +322,10 @@ export class EcsAgentCluster extends Construct { constructor(scope: Construct, id: string, props: EcsAgentClusterProps) { super(scope, id); + if ((props.taskApprovalsTable || props.approvalRequestsApiUrl) + && (!props.agentSessionRole || !props.approvalRequestsApiUrl)) { + throw new Error('ECS approvals require agentSessionRole and approvalRequestsApiUrl with a task-scoped invocation grant'); + } this.containerName = 'AgentContainer'; // ECS Cluster with Fargate capacity provider and container insights @@ -403,13 +402,16 @@ export class EcsAgentCluster extends Construct { inferenceProfileId(bedrockGeoRegion, PLATFORM_DEFAULT_AUX_MODEL_ID), TASK_TABLE_NAME: props.taskTable.tableName, TASK_EVENTS_TABLE_NAME: props.taskEventsTable.tableName, + ...(props.taskApprovalsTable && { + TASK_APPROVALS_TABLE_NAME: props.taskApprovalsTable.tableName, + }), + ...(props.approvalRequestsApiUrl && { APPROVAL_REQUESTS_API_URL: props.approvalRequestsApiUrl }), USER_CONCURRENCY_TABLE_NAME: props.userConcurrencyTable.tableName, LOG_GROUP_NAME: logGroup.logGroupName, GITHUB_TOKEN_SECRET_ARN: props.githubTokenSecret.secretArn, ...(props.memoryId && { MEMORY_ID: props.memoryId }), - // The payload bucket name so the orchestrator-issued AGENT_PAYLOAD_S3_URI - // can be fetched. (The orchestrator sets the URI per-task via container - // override; this is set here for parity with the runtime env.) + // Deployment metadata; the per-task AGENT_PAYLOAD_REF supplies the + // manifest URI and signed payload URL. IAM authenticates the manifest. ...(props.payloadBucket && { ECS_PAYLOAD_BUCKET: props.payloadBucket.bucketName }), // Artifact workflows (planning/analysis) deliver their document to // this bucket. The AgentCore runtime has ARTIFACTS_BUCKET_NAME; the ECS task @@ -478,51 +480,19 @@ export class EcsAgentCluster extends Construct { this.taskDefinition = makeTaskDef('TaskDef', buildCpu, buildMemory, { // Heavy CI-parity builds legitimately run longer than the 1800s default. BUILD_VERIFY_TIMEOUT_S: '3600', - // Pin the jest test fleet to an ABSOLUTE worker count on ECS. jest's - // `maxWorkers: 25%` is CORE-relative → 4 workers on this 16-vCPU box. - // Measured, the test suite at 4 workers peaks at only ~2.2 GB (whole process - // tree) — not tens of GB. Container OOMs were not driven by this test - // suite's worker count; they were driven by TOTAL concurrency — a - // full-parallel `mise run build` running every package's test/build legs - // plus the resident coding agent all at once. So the real memory driver is - // cross-package build parallelism, not jest's internal workers. 4 is - // comfortably safe on the 120 GB box even alongside the other packages + - // agent. Kept as an explicit env (not core-relative) so a future bigger box - // can't silently over-spawn. The test script reads JEST_MAX_WORKERS (default - // 25%), so this only pins the shared ECS box — CI (2–4 cores) and dev - // machines keep 25%, unaffected. + // Repositories that honor JEST_MAX_WORKERS use four Jest workers on build + // tasks. An absolute value keeps that limit stable when task CPU changes; + // it does not set worker counts for other test runners. JEST_MAX_WORKERS: '4', - // Serialize the mise task graph so the build steps' peak memory doesn't sum - // and OOM the task. `mise run build` fans out its `depends` (the per-package - // build/quality legs) up to MISE_JOBS in parallel (default 4); each package - // then spawns its OWN worker fleet (jest, pytest, esbuild, cdk synth). The - // measured memory driver of the OOMs was this CROSS-PACKAGE storm summing on - // top of the resident coding agent — not any single package. At 120 GB - // (Fargate's max at 16 vCPU) there is no more RAM to add, so the remedy is to - // cut peak parallelism. MISE_JOBS=1 runs the packages SEQUENTIALLY → peak ≈ - // max(single package) instead of sum(all packages), while still building - // every package and keeping BOTH gates (baseline + post-agent). - // Within-package parallelism (JEST_MAX_WORKERS=4, pytest) is untouched, so a - // single package still uses the box's cores. Cost is wall-clock (~serial - // sum, still minutes) — trivial against BUILD_VERIFY_TIMEOUT_S=3600. Without - // this, the post-agent build OOM'd (exit 137) stacking on the still-resident - // agent; a gate that OOMs verified NOTHING — serializing lets it actually - // COMPLETE and gate. Only affects `mise run ` (the build legs); the - // agent's direct `uv run pytest` calls are unaffected. + // Run one mise task at a time to limit overlap between package builds. + // Individual tools can still run their own workers, and the coding agent + // remains resident. This trades build duration for lower peak memory; + // it does not cap direct pytest/Jest calls or guarantee a workload fits. MISE_JOBS: '1', - // Skip the target repo's pre-push TEST hook inside the agent container. - // `mise run install` installs prek git hooks, incl. a pre-push hook that - // re-runs the FULL cdk+cli+agent test suite on every `git push`. In this - // container that suite already ran TWICE (baseline + post-agent build gate) - // and GitHub CI runs it again — so the pre-push run is pure redundancy, AND - // it runs UNcapped (no JEST_MAX_WORKERS), stacking on the resident agent → - // OOM. The agent's only escape was `git push --no-verify`, which silently - // bypassed ALL hooks (incl. the security scan) and trained a - // skip-verification habit. SKIP is the pre-commit/prek standard env var - // (comma-separated hook ids); scoping it to the tests hook lets the push - // succeed WITHOUT --no-verify while KEEPING the pre-push security scan. - // Propagates to both the platform push (post_hooks.py) and the agent's own - // git-tool pushes via shell.py::_clean_env (blacklist — passes SKIP through). + // For repositories using this pre-commit/prek hook ID, skip that named + // pre-push test hook. Other hook IDs remain enabled. This does not establish + // that the repository's tests already ran; verification depends on its + // configured workflow. shell.py::_clean_env passes SKIP to git subprocesses. SKIP: 'monorepo-tests-pre-push', // Caller overrides win: the values above are tuned for one monorepo's // toolchain, so a deployment with a different build shape replaces them @@ -545,24 +515,20 @@ export class EcsAgentCluster extends Construct { if (props.agentSessionRole) { props.agentSessionRole.admitComputeRole(taskRole); } else { - props.taskTable.grantReadWriteData(taskRole); + grantAgentTaskTableAccess(props.taskTable, taskRole, false); props.taskEventsTable.grantReadWriteData(taskRole); } - // UserConcurrencyTable is user-scoped (not task_id leading-key-able) and is - // touched by the reconciler/orchestrator path; keep it on the task role. - props.userConcurrencyTable.grantReadWriteData(taskRole); + // Capacity counters are coordinator-owned. The agent never accesses them. // Secrets Manager read for GitHub token (read once at startup, before the // agent assumes the SessionRole — stays on the task role). props.githubTokenSecret.grantRead(taskRole); - // Read-only on the ECS payload bucket so the container can fetch its payload - // (AGENT_PAYLOAD_S3_URI) at boot. READ only — the container runs untrusted - // repo code, so it must not be able to write or delete payloads (the trusted - // orchestrator owns write + delete). Stays on the task role (read once at - // startup, before the agent assumes any SessionRole). + // Only deployment manifests use worker credentials. The boot helper reads + // its exact task object through a signed URL and removes that capability + // from the environment before starting repository code. if (props.payloadBucket) { - props.payloadBucket.grantRead(taskRole); + grantWorkerBootstrap(props.payloadBucket, taskRole); } // Artifact workflows (planning/analysis) deliver their document to the @@ -700,7 +666,7 @@ export class EcsAgentCluster extends Construct { NagSuppressions.addResourceSuppressions(taskRole, [ { id: 'AwsSolutions-IAM5', - reason: 'DynamoDB index/* wildcards from CDK grantReadWriteData (UserConcurrencyTable, and task tables only when no SessionRole is wired); Secrets Manager wildcards from CDK grantRead (GitHub token) and the bgagent-linear-oauth-*/bgagent-jira-oauth-* prefix grant (ABCA-488 — per-workspace channel OAuth tokens are created by the CLI at setup, name unknown at synth, GetSecretValue only); CloudWatch Logs wildcards from CDK grantWrite; S3 object/* wildcard from CDK grantRead on the ECS payload bucket (read-only, scoped to that bucket — #502). Bedrock InvokeModel is scoped to explicit model/inference-profile ARNs (no wildcard resource). ec2:DescribeAvailabilityZones requires Resource:* (EC2 describe actions have no resource-level scoping) — read-only, no mutation/data access; needed so a CDK target repo\'s `cdk synth` build gate can resolve AZ context on a fresh clone (ECS-parity, no cdk.context.json cache in the container).', + reason: 'DynamoDB index/* wildcards from the legacy TaskEventsTable grant when no SessionRole is wired (TaskTable allows only reporting updates; the worker has no UserConcurrency access); Secrets Manager wildcards from CDK grantRead (GitHub token) and the bgagent-linear-oauth-*/bgagent-jira-oauth-* prefix grant (ABCA-488 — per-workspace channel OAuth tokens are created by the CLI at setup, name unknown at synth, GetSecretValue only); CloudWatch Logs wildcards from CDK grantWrite; Worker S3 GetObject is restricted to bootstrap/*; other object reads and payload-bucket listing are explicitly denied (#700). Bedrock InvokeModel is scoped to explicit model/inference-profile ARNs (no wildcard resource). ec2:DescribeAvailabilityZones requires Resource:* (EC2 describe actions have no resource-level scoping) — read-only, no mutation/data access; needed so a CDK target repo\'s `cdk synth` build gate can resolve AZ context on a fresh clone (ECS-parity, no cdk.context.json cache in the container).', }, { id: 'AwsSolutions-ECS2', diff --git a/cdk/src/constructs/ecs-payload-bucket.ts b/cdk/src/constructs/ecs-payload-bucket.ts index 87dc47553..4c1b07402 100644 --- a/cdk/src/constructs/ecs-payload-bucket.ts +++ b/cdk/src/constructs/ecs-payload-bucket.ts @@ -55,24 +55,17 @@ export interface EcsPayloadBucketProps { } /** - * S3 bucket for ECS task payloads (#502). + * Storage for ECS v2 bootstrap manifests, task payloads and private references. * - * The ECS compute strategy cannot pass the orchestrator payload (repo URL, - * prompt, and the large ``hydrated_context``) inline: a Fargate ``RunTask`` - * caps the entire ``containerOverrides`` blob at 8192 bytes, and the hydrated - * context routinely exceeds that, so the call is rejected with - * ``InvalidParameterException``. (AgentCore is unaffected — it passes the - * payload in the ``InvokeAgentRuntime`` request body, which has no comparable - * limit.) Instead, the orchestrator writes the payload to - * ``s3:////payload.json`` and passes only a small - * ``AGENT_PAYLOAD_S3_URI`` pointer in the override; the container fetches and - * parses it on boot. + * RunTask caps the entire overrides object at 8192 bytes. The coordinator + * writes task instructions to /payload.json and passes AGENT_PAYLOAD_REF, + * containing an authenticated-manifest URI and a signed one-object URL. + * EcsAgentCluster grants the worker only bootstrap/* reads and explicitly denies + * other object reads/listing. The coordinator owns writes, signing and cleanup. * - * Dedicated (not co-tenant with attachments/traces) so the boundary is - * structural: the ECS task role gets S3 **read** here and nowhere else, the - * attachments feature can never collide with payload keys, and the tight - * 1-day TTL is whole-bucket rather than a prefix-scoped rule grafted onto a - * shared bucket. + * A dedicated bucket separates boot data from attachments/traces and gives it + * a one-day lifecycle backstop. Finalize deletes payload.json and launch.json; + * shared deployment manifests expire through the lifecycle rule. * * Security / hygiene (parity with TraceArtifactsBucket): * - ``blockPublicAccess: BLOCK_ALL`` + ``enforceSSL: true`` — no public read, diff --git a/cdk/src/constructs/fanout-consumer.ts b/cdk/src/constructs/fanout-consumer.ts index f232897e3..5ab025043 100644 --- a/cdk/src/constructs/fanout-consumer.ts +++ b/cdk/src/constructs/fanout-consumer.ts @@ -63,6 +63,9 @@ export interface FanOutConsumerProps { */ readonly taskTable?: dynamodb.ITable; + /** Approval state and successful per-request notification receipts. */ + readonly taskApprovalsTable?: dynamodb.ITable; + /** * RepoTable — GitHub dispatcher reads per-repo * `github_token_secret_arn` overrides. Optional: if omitted, falls @@ -214,6 +217,10 @@ export class FanOutConsumer extends Construct { props.taskTable.grantReadWriteData(this.fn); this.fn.addEnvironment('TASK_TABLE_NAME', props.taskTable.tableName); } + if (props.taskApprovalsTable) { + props.taskApprovalsTable.grant(this.fn, 'dynamodb:GetItem', 'dynamodb:UpdateItem'); + this.fn.addEnvironment('TASK_APPROVALS_TABLE_NAME', props.taskApprovalsTable.tableName); + } if (props.repoTable) { props.repoTable.grantReadData(this.fn); this.fn.addEnvironment('REPO_TABLE_NAME', props.repoTable.tableName); diff --git a/cdk/src/constructs/lambda-microvm-compute.ts b/cdk/src/constructs/lambda-microvm-compute.ts index e429b0728..6d7303df7 100644 --- a/cdk/src/constructs/lambda-microvm-compute.ts +++ b/cdk/src/constructs/lambda-microvm-compute.ts @@ -26,43 +26,22 @@ import * as s3 from 'aws-cdk-lib/aws-s3'; import * as secretsmanager from 'aws-cdk-lib/aws-secretsmanager'; import { NagSuppressions } from 'cdk-nag'; import { Construct } from 'constructs'; -// Cross-language contract (S9): `microvm_hook_budgets` couples THIS construct's -// `/ready` hook timeout to the agent's own warm-up ceiling in -// `agent/src/server.py`. Imported (not copied) so `tsc` fails on a renamed field, -// and `scripts/check-constants-sync.ts` enforces the `warmup_total < ready_hook` -// invariant plus the no-literal-redeclaration rule on both sides. See -// `contracts/constants.md`. -// Single source of truth for the supported-Region list. ADR-021's -// `microvm-regions.ts` header is explicit that the list "is the ONLY place the -// list is declared — do not copy it", so the synth-time gate IMPORTS it rather -// than duplicating it. The module is a dependency-free pair of pure constants -// (no AWS SDK, no Lambda-runtime code), so pulling it into the CDK app tree -// costs nothing and cannot drift. +// Shared hook budgets and Region constants prevent cross-package drift. import { AgentMemory } from './agent-memory'; import { AgentSessionRole } from './agent-session-role'; import { resolveBedrockModelIds } from './bedrock-models'; +import { grantWorkerBootstrap } from './payload-bootstrap-permissions'; import sharedConstants from '../../../contracts/constants.json'; import { LAMBDA_MICROVM_SUPPORTED_REGIONS, isLambdaMicrovmRegionSupported } from '../handlers/shared/microvm-regions'; /** - * Lifecycle expiry for MicroVM `/run` hook payloads, in days. - * - * Mirrors {@link ECS_PAYLOAD_TTL_DAYS} in shape but NOT in role: on ECS the - * orchestrator deletes the payload at finalize and the rule is only a crash - * backstop, whereas the MicroVM strategy never deletes (ADR-021 sub-decision 3 - * / `microvmPayloadKey`) — so on this backend the lifecycle rule is the ONLY - * reaper. Payloads carry the hydrated prompt context, so the TTL stays as tight - * as the read-once-at-`/run` access pattern allows. + * Fallback expiry for /run payloads and private launch references. Finalization + * deletes task objects; S3 lifecycle also cleans abandoned objects asynchronously. */ export const MICROVM_PAYLOAD_TTL_DAYS = 1; /** - * Cost-allocation tag applied to every resource this construct creates - * (ADR-021 sub-decision 4, "Cost attribution"). - * - * The stack-level `compute_type` tag in `main.ts` carries a single value and is - * already imprecise with two backends; these per-resource tags make MicroVM - * spend attributable regardless of what the stack-level tag says. + * Per-resource cost tag identifying MicroVM infrastructure in mixed-backend stacks. */ export const MICROVM_BACKEND_TAG_KEY = 'abca:compute-backend'; @@ -76,10 +55,9 @@ export const MICROVM_BACKEND_TAG_VALUE = 'lambda-microvm'; export const MICROVM_LOG_GROUP_PREFIX = '/aws/lambda-microvms'; /** - * Default S3 key the packaging helper (`cdk/scripts/package-microvm-artifact.sh`) - * uploads the zip + Dockerfile artifact to, inside the artifact bucket this - * construct creates. Kept in one place so the script, the `CfnMicrovmImage` - * `codeArtifact.uri`, and the build role's `s3:GetObject` scope cannot drift. + * Base S3 key for image artifacts. Managed builds insert the ZIP's SHA-256 + * before `.zip`, so changing code changes CodeArtifact.Uri. The unsuffixed key + * remains available to the explicit out-of-band image builder. */ export const MICROVM_ARTIFACT_OBJECT_KEY = 'microvm-images/agent-artifact.zip'; @@ -88,31 +66,12 @@ export const MICROVM_ARTIFACT_OBJECT_KEY = 'microvm-images/agent-artifact.zip'; * (`agent/Dockerfile` → `EXPOSE 8080`), and therefore the port the MicroVM * lifecycle-hook listener is configured for. */ -const AGENT_HOOK_PORT = 8080; +const AGENT_HOOK_PORT = sharedConstants.microvm_lifecycle.hook_port; +const LIFECYCLE_HOOK_TIMEOUT_SECONDS = sharedConstants.microvm_hook_budgets.lifecycle_hook_timeout_seconds; /** - * Value every hook field on `AWS::Lambda::MicrovmImage` takes to turn a hook ON. - * - * **The field is an ENUM, not a path** — `[DISABLED, ENABLED]`. The CDK L1 types - * `hooks.microvmHooks.run` and its three siblings as plain `string` and documents - * no allowed values, which is why this construct originally sent the agent's - * route there. CloudFormation **rejected all four at change-set early validation** - * (live 2026-08-06, ADR-021 P2-F2 — the stack was never touched, so there was no - * rollback to read): - * - * > /aws/lambda-microvms/runtime/v1/run is not a valid enum value. Supported - * > values: [DISABLED, ENABLED] (at - * > /Resources/…/Properties/Hooks/MicrovmHooks/Run) - * - * So the CloudFormation surface is IDENTICAL to the `CreateMicrovmImage` API - * surface (`--hooks '{"microvmHooks":{"run":"ENABLED",…}}'`), not different from - * it as the previous comment here claimed. There is no hook-path field on either - * surface: the service calls fixed well-known routes, which - * {@link MICROVM_AGENT_HOOK_ROUTES} records and the live build/run logs confirm. - * - * `DISABLED` is never emitted: a hook the agent does not serve is OMITTED rather - * than disabled, so the exactness test can assert the declared set in both - * directions (see `/suspend` + `/resume`, P3). + * Hook fields accept ENABLED/DISABLED, not paths. Managed images enable all six + * served hooks; the coordinator separately controls automatic suspension. */ const HOOK_ENABLED = 'ENABLED'; @@ -123,166 +82,48 @@ const HOOK_ENABLED = 'ENABLED'; const MICROVM_HOOK_ROUTE_PREFIX = '/aws/lambda-microvms/runtime/v1'; /** - * The service's fixed hook ROUTES, keyed by hook name — the paths - * `agent/src/server.py` must serve. - * - * ## ⚠️ These are AGENT ROUTE CONSTANTS ONLY. Never send them to an AWS API. - * - * They were previously passed as the `hooks.*` property VALUES on the L1, on the - * reasoning that the generated CloudFormation type accepts strings and documents - * no allowed-value constraint. CloudFormation refused every one of them - * (P2-F2 — see {@link HOOK_ENABLED} for the verbatim early-validation output): - * the hook fields are `ENABLED`/`DISABLED` enums on BOTH the CloudFormation and - * the API surface, and neither surface accepts a path at all. - * - * The routes are still worth declaring here because they are a CROSS-PACKAGE - * CONTRACT that nothing else in the CDK tree records: the service POSTs to these - * exact paths (live 2026-08-06 — `"POST /aws/lambda-microvms/runtime/v1/ready - * HTTP/1.1" 200 OK`, and the same for `/validate`, `/run` and `/terminate`), so a - * prefix drift between this map and the agent's `MICROVM_HOOK_PREFIX` surfaces as - * a failed image build (`/ready`, `/validate`) or a failed lifecycle transition on - * a real task (`/run`, `/terminate`). The construct test compares THIS map — not - * the rendered template, which no longer contains a path — against the routes the - * agent serves. + * Fixed service routes served by agent/src/server.py. Contract tests compare + * these paths with the guest routes; AWS hook properties take {@link HOOK_ENABLED}. */ export const MICROVM_AGENT_HOOK_ROUTES = { ready: `${MICROVM_HOOK_ROUTE_PREFIX}/ready`, validate: `${MICROVM_HOOK_ROUTE_PREFIX}/validate`, run: `${MICROVM_HOOK_ROUTE_PREFIX}/run`, terminate: `${MICROVM_HOOK_ROUTE_PREFIX}/terminate`, + suspend: `${MICROVM_HOOK_ROUTE_PREFIX}/suspend`, + resume: `${MICROVM_HOOK_ROUTE_PREFIX}/resume`, } as const; /** - * `/run` runtime-hook budget (seconds). - * - * `/run` is how the task payload reaches the agent (`runHookPayload`, ADR-021 - * sub-decision 3) — there is no other orchestrator→agent channel on this - * backend. The hook only validates + starts the pipeline asynchronously, so it - * stays well inside the service's 1–60 s runtime-hook window. + * /run validates the launch payload and starts the pipeline asynchronously. + * It must return within the service runtime-hook window. */ const RUN_HOOK_TIMEOUT_SECONDS = 60; /** - * `/ready` build-hook budget (seconds). - * - * `/ready` is **mandatory, not optional**: `CreateMicrovmImage` rejects an image - * that enables ANY lifecycle hook without it (live 2026-07-31): - * - * > The ready (/ready) MicroVM image hook must be enabled when any MicroVM - * > lifecycle hook (run, resume, suspend, or terminate) is enabled. The ready - * > hook signals when the application has finished initializing so the snapshot - * > is taken in a ready state. - * - * So ADR-021's original "declare `/run` in P1, serve it in P2" plan was not a - * reachable service state: `/ready` + `/run` land together in P1. - * - * **The budget is 300 s, not 60 s, because as of the P2-F5 fix `/ready` does real - * work.** It no longer just reports that uvicorn is bound: it warms the 225 MiB - * `claude` binary (`agent/src/server.py` → `_warm_snapshot_binaries`) so the - * binary's pages are resident when the snapshot is taken, instead of being faulted - * in lazily on the first task and blowing a timeout there (the defect that failed - * every P2 smoke task at turn 0). A cold 225 MiB `exec` is the one thing on this - * path that can plausibly take tens of seconds, and the build hook window allows - * up to 3600 s, so a 60 s budget would trade the runtime failure for a build - * failure. It stays far below the service ceiling: the cost of a too-generous - * budget is only how long the service waits before calling a permanently-wedged - * snapshot broken. - * - * The number is chosen against the agent's own warm-up ceiling, not guessed — and - * it is not declared here either. Both this budget and the agent's - * `_READY_WARMUP_TOTAL_BUDGET_SECONDS` come from `contracts/constants.json` → - * `microvm_hook_budgets`, because the relationship between them is the invariant - * that matters and a relationship cannot be enforced from one side. The agent's - * ceiling bounds the WHOLE warm-up (required command + every best-effort one, - * which share the remainder) at 240 s, leaving ~60 s here for uvicorn scheduling - * and the request itself. Per-command timeouts deliberately do NOT compose on the - * agent side — three commands at 120 s each would be 360 s and would blow this - * budget, turning a fix for a runtime failure into a build failure. - * `scripts/check-constants-sync.ts` fails the build if the contract ever stops - * satisfying `warmup_total < ready_hook`, so the two numbers cannot drift apart in - * a single-sided edit. + * /ready warms the required agent binary before the image snapshot is taken. + * The shared contract keeps the total guest warm-up budget below this hook timeout; + * scripts/check-constants-sync.ts enforces that relationship. */ const READY_HOOK_TIMEOUT_SECONDS = sharedConstants.microvm_hook_budgets.ready_hook_timeout_seconds; /** - * `/validate` build-hook budget (seconds). - * - * `/validate` is declared as of P2, when the agent started serving it. What it - * asserts is narrower than ADR-021 first sketched, and the narrowing is a - * consequence of THIS construct's IAM: it runs during the image build under - * {@link LambdaMicrovmCompute.buildRole}, which holds only `s3:GetObject` on the - * artifact plus log writes. So the "deeper warm-up assertions" (Bedrock - * reachability, Memory access, tool availability) are not implementable here — - * each would `AccessDenied` and fail every build. The agent's hook is therefore an - * in-process self-check (server alive, every declared hook route registered, - * interpreter floor, cross-package `platform_config` contract loaded), which is - * exactly the class of failure a build hook CAN catch: a typo'd hook prefix would - * otherwise surface as a failed lifecycle transition on the first real task - * instead of as a failed build. - * - * The checks themselves are sub-millisecond (no AWS calls, no I/O beyond a stdout - * line), so the budget is not sized for the work: it is sized for the - * still-initialising path, where the agent answers **503** until module import - * completes. 60 s covers that with orders of magnitude to spare. This used to be - * an alias for {@link READY_HOOK_TIMEOUT_SECONDS} on the argument that one number - * should cover both build hooks; the two DECOUPLED when `/ready` gained the - * binary warm-up (P2-F5) and `/validate` did not, so sharing a number would now - * mean sizing `/validate` for work it does not do. Set explicitly rather than - * relying on the service's 30 s default: a permanently failing check SHOULD fail - * the image build, and the budget is what decides how long the service waits - * before calling it that. + * /validate checks local readiness, hook registration and configuration contracts. + * It makes no AWS calls: the build role lacks runtime data and model permissions. */ const VALIDATE_HOOK_TIMEOUT_SECONDS = 60; /** - * `/terminate` runtime-hook budget (seconds). - * - * `/terminate` is declared as of P2. It is a log-and-acknowledge breadcrumb, NOT - * a shutdown mechanism: the orchestrator finalizes the task and *then* calls - * `TerminateMicrovm`, so the hook must not write terminal task status (it would - * race the finalization it follows) and must not join the pipeline thread. Its - * value is the last structured line in the task's log group from inside the guest. - * - * This is the one hook where a GENEROUS budget buys nothing and costs something. - * There is nothing to drain — `_ProgressWriter` does a synchronous `put_item` per - * event, so every progress write is already durable when this hook is called — - * and the handler never joins the pipeline thread, so it completes in - * milliseconds by construction. Meanwhile the budget bounds how long teardown - * waits on a guest that is WEDGED, and a MicroVM that has not finished - * terminating is still holding the account memory quota that gates admission for - * everyone else. - * - * So this is set near the bottom of the service's 1–60 s window rather than at - * it: 15 s is ~three orders of magnitude above the measured work, which absorbs - * a scheduling delay on a guest still saturated by a build (the realistic reason - * a fast handler answers slowly), while keeping teardown prompt. Exceeding it - * costs only a reported hook failure — the task is already finalized and - * `TerminateMicrovm` removes the VM regardless — which is why erring tight is - * the safe direction here and erring generous is not. + * /terminate logs and acknowledges teardown without joining the pipeline or + * writing task status. Termination can interrupt a task; coordinator recovery owns + * its resulting state. Keep the hook budget short so a stuck guest cannot delay + * teardown for the full runtime-hook window. */ const TERMINATE_HOOK_TIMEOUT_SECONDS = 15; /** - * BASELINE memory sizes (MiB) the service accepts for a MicroVM image. - * - * NOT a range and NOT a per-VM ceiling: the service enumerates the allowed - * baseline values per base image and rejects anything else. Live 2026-07-31 - * against `arn:aws:lambda:us-east-1:aws:microvm-image:al2023-1`: - * - * > The requested memory size of 32768 MiB is not supported by base MicroVM - * > image …al2023-1. Supported memory sizes in MiB are: - * > [512, 1024, 2048, 4096, 8192]. - * - * The service then scales a running MicroVM VERTICALLY on demand, up to a - * 32 GiB / 16 vCPU peak (developer-guide sizing table) — so 8 GiB is the top of - * the *configurable baseline*, not the amount of memory a task can use. The - * boundary probe above establishes what the FIELD accepts; only the guide - * establishes what the field MEANS (ADR-021's source hierarchy). - * - * Kept as an exported constant so {@link LambdaMicrovmComputeProps.minimumMemoryInMiB} - * can fail at synth with the real list instead of at image-create time. The - * individually named sizes exist only to keep the list out of `no-magic-numbers` - * territory — the list itself is the contract. + * Baseline values accepted by the al2023-1 image in live validation. This list + * does not establish guest-visible launch memory or capacity-change timing. */ const MEMORY_512_MIB = 512; const MEMORY_1_GIB_IN_MIB = 1024; @@ -298,15 +139,8 @@ export const MICROVM_SUPPORTED_MEMORY_MIB: readonly number[] = [ ]; /** - * Baseline memory the image declares, in MiB — the largest baseline the service - * accepts (see {@link MICROVM_SUPPORTED_MEMORY_MIB}). - * - * The ABCA agent is a build-heavy workload, so P1 asks for the top of the - * accepted baseline list rather than a smaller one it would spend the whole task - * scaling up from. Automatic vertical scaling then supplies burst capacity to a - * 32 GiB peak; nothing here requests that, and nothing can. - * - * This was `32768` until live verification refuted it as a *baseline* value. + * Baseline used by the verified ABCA workloads. Lower values require workload + * measurements; the accepted-value probe did not measure a performance advantage. */ export const DEFAULT_MINIMUM_MEMORY_MIB = MEMORY_8_GIB_IN_MIB; @@ -317,55 +151,26 @@ const LOG_RETENTION = logs.RetentionDays.THREE_MONTHS; const HTTPS_PORT = 443; /** - * HTTP port. Allowed on the **build-time** connector only. - * - * `agent/Dockerfile` installs Debian packages, and `apt-get` fetches over plain - * HTTP. With a 443-only egress path every snapshot build failed (live - * 2026-07-31): `Could not connect to deb.debian.org:80 … E: Unable to locate - * package curl` → `exit code: 100`. DNS resolved fine — the port was the sole - * cause. Opening 80 for the build path (and only the build path) made the build - * succeed immediately; see {@link LambdaMicrovmCompute.buildSecurityGroup}. + * Build-only HTTP egress for apt-get; runtime egress remains HTTPS-only. */ const HTTP_PORT = 80; /** - * Graviton/ARM64: the agent image is ARM64 on every backend. - * - * The value is the service's **enum member spelling**, `ARM_64` — not the - * lowercase `arm64` Docker/CDK use elsewhere. The CDK L1 types - * `cpuConfigurations[].architecture` as a plain `string` and documents no allowed - * values, and `arm64` was rejected at change-set early validation (live - * 2026-08-06, ADR-021 P2-F2): *"arm64 is not a valid enum value. Supported - * values: [ARM_64]"*. Matches `--cpu-configurations '[{"architecture":"ARM_64"}]'` - * in `cdk/scripts/package-microvm-artifact.sh`, which had it right all along. + * The MicroVM API spells its ARM64 enum ARM_64; Docker uses arm64. */ const CPU_ARCHITECTURE = 'ARM_64'; +const MAX_MANAGED_IMAGE_VERSION_LENGTH = 64; /** - * Resource-name half of the Lambda-managed **`NO_INGRESS`** network connector - * ARN, i.e. everything after `…:aws:network-connector:`. - * - * Why this exists at all: `RunMicrovm` does **not** default to "no ingress". A - * launch that omits `ingressNetworkConnectors` entirely came back (live - * 2026-07-31) with a service-attached PUBLIC connector and a public endpoint: - * - * ``` - * "ingressNetworkConnectors": ["arn:aws:lambda:us-east-1:aws:network-connector:aws-network-connector:HTTP_INGRESS"] - * ``` - * - * ADR-021's "no inbound exposure" posture therefore has to be an **explicit - * control**, not an omission: the strategy passes the `NO_INGRESS` connector on - * every `RunMicrovm`. The ARN shape is taken verbatim from that observed - * `HTTP_INGRESS` ARN, with the connector name swapped. + * Explicit no-ingress connector. Omitting ingress connectors let the service + * attach HTTP_INGRESS in live validation. NO_INGRESS can still return an endpoint + * URL; an unauthenticated 403 tests authentication, not valid-token reachability. */ export const MICROVM_NO_INGRESS_CONNECTOR_RESOURCE = 'aws-network-connector:NO_INGRESS'; /** - * ARN of the Lambda-managed `NO_INGRESS` connector in {@link scope}'s - * partition/Region (see {@link MICROVM_NO_INGRESS_CONNECTOR_RESOURCE}). - * - * Account segment is the literal `aws` — these connectors are service-owned, not - * account-owned, exactly like the AWS-managed policy ARNs. + * ARN of the service-owned NO_INGRESS connector in this partition and Region. + * Its account segment is the literal aws. */ export function microvmNoIngressConnectorArn(scope: Construct): string { return Stack.of(scope).formatArn({ @@ -378,36 +183,14 @@ export function microvmNoIngressConnectorArn(scope: Construct): string { } /** - * CDK context flag that bypasses the synth-time Region gate. - * - * Exists because {@link LAMBDA_MICROVM_SUPPORTED_REGIONS} rots by design: when - * AWS launches Lambda MicroVMs in a new Region, an operator there must not have - * to wait for an ABCA release. The live probes (CLI onboarding, `platform - * doctor`) already accept the new Region, so this flag only unblocks synth. + * Escape hatch for Regions launched after the static support list was updated. + * This bypasses only synth validation; live service checks still apply. */ export const MICROVM_REGION_OVERRIDE_CONTEXT = 'microvm_region_override'; /** - * Fail synth when the `lambda-microvm` backend is enabled in a Region that is - * not in the statically documented support list (ADR-021 sub-decision 4, - * "Regional availability enforcement" — the synth/deploy row). - * - * Three behaviours worth knowing: - * - * - **Unresolved (token) Region → check SKIPPED.** A region-agnostic app - * (`new Stack(app, 'X')` with no `env`, or `env.region` left to the CLI) - * resolves `Stack.region` to the `AWS::Region` pseudo-parameter, whose value - * is unknowable at synth. Comparing a token against a Region list would - * reject every region-agnostic synth — including `cdk synth` on a developer - * box with no `CDK_DEFAULT_REGION` — so the static layer stands down and the - * live probes (onboarding / doctor / orchestration classification) carry the - * enforcement. This is a deliberate hole in the *static* layer only. - * - **Escape hatch.** `--context microvm_region_override=true` skips the check; - * the error message names the flag so an operator in a just-launched Region - * is never blocked on a code change. - * - **Failure is a synth-time throw, not a warning.** ADR-021 requires "synth - * fails when ComputeTypes includes lambda-microvm in an unlisted Region"; a - * warning would let a broken deploy through to a runtime AccessDenied. + * Reject a concrete unsupported Region at synth. Unresolved Regions defer to + * runtime checks; microvm_region_override permits newly supported Regions. */ export function assertLambdaMicrovmRegionSupported(scope: Construct): void { const region = Stack.of(scope).region; @@ -443,32 +226,21 @@ export function assertLambdaMicrovmRegionSupported(scope: Construct): void { } /** - * The four operator-supplied image inputs, read from CDK context by the stack. - * - * Extracted into a type so the stack can resolve them ONCE, before `TaskApi` is - * constructed, and hand the same object to this construct — see - * {@link isLambdaMicrovmImageConfigured} for why that ordering matters. + * Shared image inputs resolved before TaskApi so lifecycle IAM and this construct + * use the same image-availability decision. */ export interface LambdaMicrovmImageInputs { readonly baseImageArn?: string; readonly baseImageVersion?: string; + readonly artifactSha256?: string; + readonly managedImageVersion?: string; readonly externalImageIdentifier?: string; readonly externalImageVersion?: string; } /** - * True when {@link inputs} selects one of the two image-provisioning states - * (managed-base-image build, or an out-of-band image), i.e. the deployment will - * have a MicroVM image and therefore an `imageArn` to scope IAM against. - * - * Exists so the three-state decision documented on {@link LambdaMicrovmCompute} - * is made in exactly ONE place. `TaskApi` is constructed before this construct - * (the cancel Lambda's ARN is needed earlier, hence the `Lazy.string` holders in - * `stacks/agent.ts`), so the stack must know whether an image will exist *before* - * the construct that creates it runs. Without a shared predicate the stack and - * the construct would each re-derive that answer and could drift — and a drift - * here means either a missing cancel grant or a grant scoped to an image that - * does not exist. + * Whether managed or external image inputs select an image. The stack uses this + * before constructing TaskApi to decide whether to grant lifecycle permissions. */ export function isLambdaMicrovmImageConfigured(inputs: LambdaMicrovmImageInputs): boolean { return Boolean((inputs.baseImageArn && inputs.baseImageVersion) || inputs.externalImageIdentifier); @@ -478,112 +250,77 @@ export function isLambdaMicrovmImageConfigured(inputs: LambdaMicrovmImageInputs) * Properties for {@link LambdaMicrovmCompute}. */ export interface LambdaMicrovmComputeProps extends LambdaMicrovmImageInputs { + /** Stable parent deployment name for image/connector names when nested. */ + readonly deploymentName?: string; + + /** + * Parent-owned execution role. Keeping this beside AgentSessionRole avoids + * a parent/child cycle through the session role's trust and AssumeRole grant. + * Omitted by standalone constructs, which create their own execution role. + */ + readonly executionRole?: iam.Role; + + /** Explicit names used by nested deployments to preserve bootstrap PassRole scope. */ + readonly buildRoleName?: string; + readonly connectorOperatorRoleName?: string; + /** - * Platform VPC. Egress leaves the MicroVM through a `AWS::Lambda::NetworkConnector` - * bound to this VPC's private-with-egress subnets, so the DNS Firewall / - * security-group / flow-log stack applies to MicroVM traffic unchanged - * (ADR-021 security table: "Egress ... None" delta). + * Platform VPC for build and runtime connectors, using private-with-egress + * subnets and the existing DNS Firewall, NAT and flow logs. */ readonly vpc: ec2.IVpc; /** - * Per-task SessionRole (#209). When provided, the MicroVM **execution role** - * is admitted to the SessionRole's trust via - * {@link AgentSessionRole.admitComputeRole} — the mechanism was designed for - * exactly this (ADR-021 sub-decision 4) — so tenant-data access stays on the - * tag-scoped SessionRole instead of the execution role. Mirrors how - * `EcsAgentCluster` delegates the Fargate task role. Omitted in isolated - * construct tests, in which case NO tenant-data access is granted at all - * (this construct never grants DynamoDB directly: unlike the ECS backend - * there is no legacy direct-grant path to preserve). + * Per-task role providing tag-scoped tenant access. The execution role can assume + * it when supplied; omitting it grants no direct DynamoDB or artifact access. */ readonly agentSessionRole?: AgentSessionRole; /** - * GitHub PAT secret. When provided, the MicroVM **execution role** gets - * `grantRead` on it. - * - * This grant stays on the execution role rather than moving to the SessionRole - * because of WHEN it is used: the agent resolves the token at startup, before - * it has assumed the SessionRole — the same ordering that keeps the grant on the - * ECS task role and the AgentCore runtime role. Without it the MicroVM cannot - * clone, push, or open a PR, which is the whole task. - * - * Omitted in isolated construct tests → no grant. + * GitHub credential read by the execution role before the task assumes its + * SessionRole. Omit only when that startup credential is unnecessary. */ readonly githubTokenSecret?: secretsmanager.ISecret; /** - * AgentCore Memory for cross-task learning. When provided, the execution role - * gets read+write so the agent's `write_task_episode` / `write_repo_learnings` - * (`bedrock-agentcore:CreateEvent`) succeed on this substrate. - * - * Exactly the prop `EcsAgentCluster` takes, for exactly the same reason: the - * `MEMORY_ID` the agent receives (in `agent_payload`, unchanged by ADR-021 P2) - * makes it ATTEMPT the write, and without the grant that attempt fails closed on - * AccessDenied. `memory.py` treats that as an infra failure — logged, - * non-fatal — so learning would silently never persist on a MicroVM-only - * deployment. Omitted in isolated construct tests / memory-less deployments. + * Platform Memory used for cross-task learning. Grants execution-role read/write + * access; omit for deployments without Memory. */ readonly agentMemory?: AgentMemory; /** - * The platform's APPLICATION_LOGS group — the same log group whose NAME travels - * to the guest as `platform_config.log_group_name` (`stacks/agent.ts` → - * `TaskOrchestrator.agentPlatformConfig` → `LOG_GROUP_NAME`). When provided, the - * MicroVM **execution role** gets `logs:CreateLogStream` + `logs:PutLogEvents` - * on it. - * - * Not optional in spirit — omitted only in isolated construct tests. P2 wired - * the name into `platform_config`, which makes the agent ATTEMPT the write, and - * shipped without the matching grant, so every structured per-task log line was - * denied (live 2026-08-07, ADR-021 P2-F4): - * - * > User: …:assumed-role/…LambdaMicrovmComputeExecutionRo…/Lambda-microvmsExecutor-… - * > is not authorized to perform: logs:CreateLogStream on resource: - * > …:log-group:/aws/vendedlogs/bedrock-agentcore/runtime/APPLICATION_LOGS/… - * - * The role's OTHER logs grant ({@link LambdaMicrovmCompute.grantMicrovmLogWrites}) - * is scoped to the service's own `/aws/lambda-microvms/*` namespace and cannot - * cover this group — the two namespaces are unrelated. Non-fatal (the agent - * degrades to stdout, which the MicroVM log group captures) but it empties the - * platform's canonical per-task observability streams, `METRICS_REPORT` - * included, on this backend only. Exactly the omission class the P2 Bedrock / - * Secrets Manager / Memory grants exist to close. + * Platform APPLICATION_LOGS group named in platform_config. Its write grant is + * separate from the service-owned /aws/lambda-microvms log namespace. */ readonly applicationLogGroup?: logs.ILogGroup; /** - * ARN of the Lambda-managed base MicroVM image to build on - * (`aws lambda-microvms list-managed-microvm-images`), e.g. - * `arn:aws:lambda:us-east-1:aws:microvm-image:al2023-1`. - * - * Supplying this **and** {@link baseImageVersion} switches the construct into - * its primary mode: it synthesizes an `AWS::Lambda::MicrovmImage` (L1) whose - * `codeArtifact.uri` points at {@link artifactObjectKey} in the artifact - * bucket this construct creates. There is no default — base-image ARNs are - * account/Region-scoped service data that is only discoverable through a live - * API call, so hardcoding one would be a guess that rots. + * Service-managed base image ARN, discovered with list-managed-microvm-images. + * With baseImageVersion and artifactSha256, creates a CloudFormation-managed image. */ readonly baseImageArn?: string; /** - * Version of {@link baseImageArn} - * (`aws lambda-microvms list-managed-microvm-image-versions`). Required - * alongside `baseImageArn` because CloudFormation marks it required on - * `AWS::Lambda::MicrovmImage` even though the API treats it as optional. + * Base image version, required alongside baseImageArn by CloudFormation. */ readonly baseImageVersion?: string; /** - * Identifier (name or ARN) of a MicroVM image built **out of band** — i.e. by - * running `cdk/scripts/package-microvm-artifact.sh` and then - * `aws lambda-microvms create-microvm-image` by hand. - * - * Only consulted when {@link baseImageArn} is absent. It exists so an - * operator can iterate on the snapshot (which takes minutes and often several - * attempts) without a stack update per attempt, then hand the finished image - * to the orchestrator. + * SHA-256 of the uploaded ZIP, printed by package-microvm-artifact.sh. + * Required for managed images: a mutable fixed key does not trigger updates. + */ + readonly artifactSha256?: string; + + /** + * Optional version to RUN from the managed image (for example `7.0`). + * This does not alter the image build or transfer its CloudFormation ownership. + * Omit to run the latest active version; pin a verified version for rollback. + */ + readonly managedImageVersion?: string; + + /** + * Name or ARN of an image built outside this construct. Used only when managed + * base-image inputs are absent; enables image iteration without a stack build. */ readonly externalImageIdentifier?: string; @@ -602,137 +339,48 @@ export interface LambdaMicrovmComputeProps extends LambdaMicrovmImageInputs { readonly imageName?: string; /** - * S3 key of the zip + Dockerfile artifact inside the artifact bucket. + * Base S3 key of the zip + Dockerfile artifact. Managed builds append the + * artifact digest before `.zip`; the manual builder uses this base key. * @default MICROVM_ARTIFACT_OBJECT_KEY */ readonly artifactObjectKey?: string; /** - * **Baseline** memory the image declares, in MiB. - * - * This is the size the MicroVM STARTS at, not a cap on what it may use: the - * service scales a running MicroVM vertically on demand, up to a 32 GiB / - * 16 vCPU peak, with no field to request that. So the practical effect of this - * prop is where the VM begins (and what it is baseline-priced at), not whether - * a build will fit. - * - * Must be one of {@link MICROVM_SUPPORTED_MEMORY_MIB} — the construct throws at - * synth otherwise, because the service rejects any other baseline at - * image-create time with a `ValidationException` an operator would only see - * minutes into a build. - * - * @default 8192 — the largest accepted baseline (see DEFAULT_MINIMUM_MEMORY_MIB) + * Service baseline in MiB; must be one of {@link MICROVM_SUPPORTED_MEMORY_MIB}. + * This setting does not guarantee workload fit or capacity-change timing. + * @default 8192 — the baseline used in live ABCA verification */ readonly minimumMemoryInMiB?: number; /** - * Non-secret environment variables baked into the snapshot at build time. - * - * Deliberately empty by default, and expected to STAY empty. ADR-021 - * sub-decision 3 forbids secrets, tokens, and per-task identity in the snapshot - * — and P2 resolved the remaining question (where the agent's non-secret - * configuration parity with the ECS container comes from) in favour of the - * `/run` payload's `platform_config` block, NOT this prop. A snapshot is shared - * across every task and every deployment that reuses it, so a table or bucket - * name baked in here would be a deploy-time value frozen at image-build time — - * stale the moment the stack is redeployed. Reach for this only for genuinely - * image-invariant settings (a locale, a toolchain path). - * @default {} — no baked configuration + * Image-invariant, non-secret settings only. The construct also adds its lifecycle + * protocol marker. Deployment configuration, credentials and task identity arrive + * through /run and must not be baked into a shared snapshot. + * @default {} — no caller-supplied settings */ readonly imageEnvironmentVariables?: Record; } /** - * AWS Lambda MicroVMs compute backend — infrastructure half (ADR-021, - * sub-decision 4; P1 of the phased rollout). - * - * Provisions, in dependency order: - * - * 1. **Egress network connectors** (`AWS::Lambda::NetworkConnector`) on the - * platform VPC's private-with-egress subnets — TWO of them, sharing one - * operator role: - * - the **runtime** connector with a 443-only security group. This is what - * keeps the ADR's "Egress: no delta vs AgentCore/ECS" claim true — MicroVM - * traffic traverses the same NAT / DNS Firewall / flow logs as the other - * two backends. - * - the **build-time** connector with a 443 **and 80** security group, - * referenced only by the image resource. `agent/Dockerfile` runs - * `apt-get`, which is plain HTTP; a 443-only build path fails every - * snapshot build (see {@link HTTP_PORT}). Runtime egress stays 443-only. - * 2. **Artifact bucket** for the zip + Dockerfile the service builds the - * snapshot from, and a **payload bucket** for `/run` payloads that exceed - * the 4 KB `runHookPayload` cap (which is nearly all of them). - * 3. **Build role** — assumed by Lambda during image creation: `s3:GetObject` - * on the artifact object and CloudWatch Logs writes. Without it Lambda - * cannot emit build logs, which makes a failed snapshot build undebuggable. - * 4. **Execution role** — assumed by the running MicroVM: CloudWatch Logs (both - * the service's own `/aws/lambda-microvms/*` namespace and the platform - * APPLICATION_LOGS group whose name `platform_config` delivers), read-only on - * the payload bucket, the P2 runtime-parity grants (GitHub PAT + - * channel-OAuth secret reads, scoped Bedrock invocation, AgentCore Memory, - * `ec2:DescribeAvailabilityZones` for a CDK repo's synth gate), and — when a - * SessionRole is wired — admission to the per-task SessionRole, which is the - * ONLY path to tenant data. - * 5. **MicroVM image** (`AWS::Lambda::MicrovmImage`) — see {@link baseImageArn} - * for why this is conditional. - * - * ## Image provisioning: three states, one construct - * - * | Props supplied | What happens | When to use it | - * |---|---|---| - * | `baseImageArn` + `baseImageVersion` | `AWS::Lambda::MicrovmImage` L1 is synthesized from `s3:///`; {@link imageIdentifier} is its ARN | steady state | - * | `externalImageIdentifier` | no image resource; the supplied identifier is resolved to its exact ARN and handed to the orchestrator | iterating on the snapshot out of band | - * | neither | roles + buckets + connectors only; a synth-time **warning**, no image, and no `MICROVM_IMAGE_IDENTIFIER` for the orchestrator | first deploy — you cannot upload the artifact before the bucket that holds it exists | - * - * That third state is not an oversight: the artifact bucket is created by this - * stack, so the very first `--context compute_type=lambda-microvm` deploy has - * nowhere to have put the zip yet. It is a **warning rather than a throw** - * precisely so the bootstrap sequence (deploy → run the packaging script - * against the now-existing bucket → redeploy with `microvm_base_image_arn`) is - * possible at all. A `lambda-microvm` task submitted in that interim window - * fails fast with the strategy's own "stack deployed without the MicroVM - * substrate" error, which names the remedy. - * - * ## ⚠️ A P2 substrate is fully wired, but still NOT smoke-verified - * - * Reaching state 1 or 2 provisions a complete substrate, a buildable image, and - * a payload-deliverable `/run` path: P1 declares AND the agent serves `/ready` - * and `/run` (`agent/src/server.py`), because live verification proved the - * original "declare in P1, serve in P2" split was not a reachable service state - * — `CreateMicrovmImage` refuses any lifecycle hook without `/ready` (see - * {@link READY_HOOK_TIMEOUT_SECONDS}), and an image with no hooks at all cannot - * receive a `runHookPayload`. + * Lambda MicroVM infrastructure: build/runtime VPC connectors, image artifacts, + * bootstrap payloads, logs, roles and an optional managed image (ADR-021). * - * P2 adds the two halves an agent needs to actually finish a task: the runtime - * IAM parity on the execution role (see item 4 above) and non-secret - * configuration delivery through the `/run` payload's `platform_config` block - * (`handlers/shared/strategies/lambda-microvm-strategy.ts`) — the substitute for - * the env block the other two backends get at deploy time, since the snapshot must - * not bake it in. It also declares the two hooks the agent gained in the same - * phase: `/validate` (build-time self-check — see - * {@link VALIDATE_HOOK_TIMEOUT_SECONDS} for why the build role's permissions bound - * what it can assert) and `/terminate` (in-guest teardown breadcrumb — - * {@link TERMINATE_HOOK_TIMEOUT_SECONDS}). + * Runtime egress permits HTTPS; the separate build connector also permits HTTP + * for apt-get. Every launch explicitly selects NO_INGRESS. The execution role + * reads bootstrap manifests and startup credentials, invokes models, writes logs + * and uses Memory. Tenant data remains on the per-task SessionRole; lifecycle + * control remains on coordinator and decision-handler roles. * - * What is still unverified is the thing no amount of wiring can assert: an - * end-to-end clone → change → PR run on this substrate, plus egress specifics and - * heartbeat/progress behaviour from a live MicroVM. That is what the - * `abca:microvm-image-p1-smoke-unverified` warning below says, and it is repeated - * in `cdk/scripts/package-microvm-artifact.sh`. Only `/suspend` and `/resume` - * remain undeclared, until P3 implements them: a hook the service calls but - * nothing answers fails the corresponding lifecycle transition. + * Image provisioning has three states: + * - baseImageArn + baseImageVersion + artifactSha256: build a managed image. + * - externalImageIdentifier: use an image built outside this construct. + * - neither: create infrastructure for the initial artifact upload, with a warning; + * tasks cannot run until an image is configured. * - * ## Deliberately NOT here - * - * The execution role gets no artifacts-bucket grant and no DynamoDB grant: an - * artifact delivery write goes through the SessionRole's - * `artifacts/${task_id}/*` statement (the AgentCore runtime role has no direct - * grant either) and every table the agent touches is `task_id`-partitioned - * SessionRole territory. It also has no UserConcurrencyTable grant — that counter - * is orchestrator/reconciler-owned and the agent path never writes it. - * `lambda:SuspendMicrovm` / `lambda:ResumeMicrovm` are absent (P3), and - * `lambda:CreateMicrovmAuthToken` is granted to no role in any phase — no JWE - * consumer exists (sub-decision 3). + * Managed images enable all six served hooks. Automatic approval sleep requires + * a compatible image and coordinator plus the deployment enable switch. P3 live + * acceptance is recorded in docs/verification/README.md; new + * installations must verify their own configuration before enabling sleep. */ export class LambdaMicrovmCompute extends Construct { /** S3 bucket holding the zip + Dockerfile the snapshot is built from. */ @@ -740,8 +388,10 @@ export class LambdaMicrovmCompute extends Construct { /** Key of the artifact object inside {@link artifactBucket}. */ public readonly artifactObjectKey: string; + /** Unsuffixed key used by the manual builder and packaging helper. */ + public readonly artifactBaseObjectKey: string; - /** S3 bucket holding oversized `/run` payloads (S3-pointer delivery). */ + /** S3 bucket holding bootstrap manifests, task payloads and private launch references. */ public readonly payloadBucket: s3.Bucket; /** Role Lambda assumes while building the snapshot image. */ @@ -751,10 +401,7 @@ export class LambdaMicrovmCompute extends Construct { public readonly executionRole: iam.Role; /** - * Role Lambda assumes to manage the connectors' ENIs in the platform VPC. - * - * REQUIRED for `VPC_EGRESS` connectors, despite the generated L1 typing - * `operatorRole` as optional — see the comment at the connector below. + * Service role managing connector ENIs; required for VPC_EGRESS. */ public readonly connectorOperatorRole: iam.Role; @@ -774,11 +421,7 @@ export class LambdaMicrovmCompute extends Construct { public readonly buildEgressConnectorArns: string[]; /** - * Ingress connectors passed on every `RunMicrovm` - * (`MICROVM_INGRESS_CONNECTOR_ARNS`). In P1–P3 this is exactly the - * Lambda-managed `NO_INGRESS` connector: the service's default is a PUBLIC - * `HTTP_INGRESS`, so "no inbound" has to be requested explicitly (see - * {@link MICROVM_NO_INGRESS_CONNECTOR_RESOURCE}). + * Explicit NO_INGRESS connector passed on every RunMicrovm request. */ public readonly ingressConnectorArns: string[]; @@ -797,16 +440,14 @@ export class LambdaMicrovmCompute extends Construct { /** Image name used for the image resource and the log group. */ public readonly imageName: string; - /** BASELINE memory (MiB) declared on the image; the service bursts above it. */ + /** + * Configured image memory baseline in MiB. + */ public readonly minimumMemoryInMiB: number; /** - * Value for `MICROVM_IMAGE_IDENTIFIER`. **Always a full image ARN**, never a - * bare name — `RunMicrovm` rejects bare names outright - * (`ValidationException: Malformed ARN - doesn't start with 'arn:'`, live - * 2026-07-31), and so does `list-microvm-image-builds`. `undefined` in the - * neither-input-supplied bootstrap state, in which case the stack must not - * inject the MicroVM env block at all. + * Full image ARN for RunMicrovm. Undefined during the no-image bootstrap phase; + * bare external names are resolved to ARNs for both launch and IAM. */ public readonly imageIdentifier?: string; @@ -814,13 +455,8 @@ export class LambdaMicrovmCompute extends Construct { public readonly imageVersion?: string; /** - * The image's IAM resource ARN. Always set whenever {@link imageIdentifier} - * is — a bare image name is resolved to its full - * `…:microvm-image:` ARN — so the orchestrator's `lambda:RunMicrovm` / - * `GetMicrovm` / `TerminateMicrovm` grant is *always* scoped to this one - * platform-created image and never widens to an account-level wildcard. - * `undefined` only in the no-image bootstrap state, where no grant is issued - * at all. + * Exact image ARN used to scope lifecycle IAM; undefined only before an image + * is configured. Image versions are separate request fields, not ARN suffixes. */ public readonly imageArn?: string; @@ -833,8 +469,23 @@ export class LambdaMicrovmCompute extends Construct { assertLambdaMicrovmRegionSupported(this); const stack = Stack.of(this); - this.artifactObjectKey = props.artifactObjectKey ?? MICROVM_ARTIFACT_OBJECT_KEY; - this.imageName = props.imageName ?? sanitizeImageName(`${stack.stackName}-abca-agent`); + const deploymentName = props.deploymentName ?? stack.stackName; + if (Token.isUnresolved(deploymentName)) { + throw new Error('Nested MicroVM resources require a concrete deploymentName from the parent stack'); + } + const managedImage = Boolean(props.baseImageArn && props.baseImageVersion); + if ((managedImage || props.artifactSha256 !== undefined) + && !/^[a-f0-9]{64}$/.test(props.artifactSha256 ?? '')) { + throw new Error( + 'Managed MicroVM images require microvm_artifact_sha256 (64 lowercase hex characters). ' + + 'Run cdk/scripts/package-microvm-artifact.sh and deploy with its printed artifact digest.', + ); + } + this.artifactBaseObjectKey = props.artifactObjectKey ?? MICROVM_ARTIFACT_OBJECT_KEY; + this.artifactObjectKey = props.artifactSha256 + ? `${this.artifactBaseObjectKey.replace(/\.zip$/, '')}-${props.artifactSha256}.zip` + : this.artifactBaseObjectKey; + this.imageName = props.imageName ?? sanitizeImageName(`${deploymentName}-abca-agent`); // Fail at SYNTH on an unsupported memory size. The service enumerates the // sizes a base image accepts and rejects anything else at create time, which @@ -844,10 +495,8 @@ export class LambdaMicrovmCompute extends Construct { throw new Error( `minimumMemoryInMiB=${this.minimumMemoryInMiB} is not a BASELINE memory size AWS Lambda ` + `MicroVMs accepts. Supported baselines (MiB): ${MICROVM_SUPPORTED_MEMORY_MIB.join(', ')}. ` - + `The default is ${DEFAULT_MINIMUM_MEMORY_MIB}, the largest accepted baseline. Note this is ` - + 'a BASELINE, not a cap: the service scales a running MicroVM vertically to a 32 GiB / ' - + '16 vCPU peak on its own, so asking for more here is neither possible nor necessary. A ' - + 'repo needing more than that SUSTAINED belongs on compute_type=ecs (16 vCPU / 120 GB).', + + `The default is ${DEFAULT_MINIMUM_MEMORY_MIB}. Choose a supported baseline and verify ` + + 'that the workload fits; this setting alone does not establish available peak capacity.', ); } @@ -891,51 +540,13 @@ export class LambdaMicrovmCompute extends Construct { 'Allow HTTP egress for apt-get during the snapshot build (build path only)', ); - // OPERATOR ROLE — required, despite the generated L1 typing it optional. - // - // `CfnNetworkConnector.operatorRole?: string` reads like an opt-in, and this - // construct originally left it unset "so Lambda manages the ENIs with its own - // service-linked role". Live deploy 2026-07-31 refuted that outright — the - // connector never reaches CREATE_COMPLETE without it: - // - // NetworkConnectorOperatorRole is required for VPC_EGRESS connector type - // (Service: Lambda, Status Code: 400) HandlerErrorCode: InvalidRequest - // - // The permission set below is the minimal recipe validated standalone in that - // run: `AWSLambdaVPCAccessExecutionRole` plus the ENI/tag/private-IP actions - // the managed policy omits. - // - // ONE role for BOTH connectors: they differ only in security group, both are - // created and owned by this construct in the same VPC, and a second identical - // role would double the IAM surface a reviewer has to check for no isolation - // gain (the role manages ENIs, not traffic). - // - // --- TRUST POLICY: bare service principal, NO source conditions (P2-F1/F3) --- - // - // Shared by all three MicroVM-facing roles (this one, `buildRole`, - // `executionRole`) because the decision is one decision. `lambda.amazonaws.com` - // is the correct principal — there is no `microvms.lambda.amazonaws.com`, and - // using one is rejected at role-creation time with MalformedPolicyDocument. - // - // ⚠️ **`aws:SourceAccount`/`aws:SourceArn` are deliberately absent. Adding one - // back re-breaks the deploy.** The Lambda MicroVMs service populates NO source - // condition key when it assumes these roles, so a trust policy carrying one is - // unassumable: both network connectors CREATE_FAILED deterministically, and - // `RunMicrovm` surfaced the same root cause as a misleading caller-side - // `iam:PassRole` denial on the orchestrator. Removing the conditions fixed both - // within seconds. This looks like a regression to anyone applying the standard - // confused-deputy pattern — and it already WAS one in the other direction: P1's - // working probe had no conditions, and the P1 F2 fix added them "to mirror the - // build/execution roles". `sts:TagSession` stays; it was never implicated. - // - // ADR-021 §4 is authoritative for the rest: the live evidence (P2-F1 / P2-F3, - // verbatim failures, the two-arm PassRole experiment, the contaminated control - // that produced run 1's false negative), the per-role compensating-controls - // table, and the conditions under which the condition could be restored. The - // trust shape here is asserted by `test/constructs/lambda-microvm-compute.test.ts` - // ("NO source-key condition on any MicroVM-facing role trust"). + // VPC_EGRESS requires an operator role shared by the two connectors. + // Live P2 checks rejected source-conditioned trust on all three roles and + // conditioned PassRole on the build/execution paths. Keep lambda.amazonaws.com + // without source conditions; ADR-021 §4 records evidence and compensating scopes. const microvmAssumedBy = new iam.ServicePrincipal('lambda.amazonaws.com'); this.connectorOperatorRole = new iam.Role(this, 'ConnectorOperatorRole', { + roleName: props.connectorOperatorRoleName, assumedBy: microvmAssumedBy, description: 'ABCA Lambda MicroVMs network-connector operator role: lets Lambda manage the connector ' @@ -958,26 +569,18 @@ export class LambdaMicrovmCompute extends Construct { 'ec2:UnassignPrivateIpAddresses', 'ec2:CreateTags', ], - // EC2 network-interface APIs are largely non-resource-scopable (the ENI - // does not exist when CreateNetworkInterface is authorized, and the - // Describe* calls take no resource), which is exactly why the AWS-managed - // VPC-access policy uses `*` too. See the cdk-nag suppression below. + // Describe actions need wildcard resources. ENI mutations share this tested + // wildcard statement; their scope is a separate IAM-hardening consideration. resources: ['*'], })); - // The connector, not the MicroVM, owns the ENIs — which is why - // `lambda:PassNetworkConnector` is required on the orchestrator even for - // AWS-managed connectors (ADR-021 sub-decision 4). - // - // `associatedComputeResourceTypes: ['MicroVm']` is the only value the - // service accepts today (CloudFormation: "Currently, only MicroVm is - // supported"). + // Connectors own the VPC interfaces. The service accepts MicroVm here. const vpcEgressSubnetIds = props.vpc.selectSubnets({ subnetType: ec2.SubnetType.PRIVATE_WITH_EGRESS, }).subnetIds; this.egressConnector = new lambda.CfnNetworkConnector(this, 'EgressConnector', { - name: sanitizeImageName(`${stack.stackName}-microvm-egress`), + name: sanitizeImageName(`${deploymentName}-microvm-egress`), operatorRole: this.connectorOperatorRole.roleArn, configuration: { vpcEgressConfiguration: { @@ -994,7 +597,7 @@ export class LambdaMicrovmCompute extends Construct { // (443 + 80). Referenced by the image resource / packaging script, never by // `RunMicrovm`, so a running agent still gets the 443-only posture. this.buildEgressConnector = new lambda.CfnNetworkConnector(this, 'BuildEgressConnector', { - name: sanitizeImageName(`${stack.stackName}-microvm-build-egress`), + name: sanitizeImageName(`${deploymentName}-microvm-build-egress`), operatorRole: this.connectorOperatorRole.roleArn, configuration: { vpcEgressConfiguration: { @@ -1036,8 +639,8 @@ export class LambdaMicrovmCompute extends Construct { { id: 'microvm-artifact-mpu-abort', enabled: true, - // The artifact is tens/hundreds of MB, so uploads are multipart; a - // failed `aws s3 cp` otherwise leaves billable parts forever. + // Abort abandoned multipart uploads from alternative publishers. + // The packaging helper currently uses a single PutObject. abortIncompleteMultipartUploadAfter: Duration.days(1), }, ], @@ -1064,19 +667,11 @@ export class LambdaMicrovmCompute extends Construct { autoDeleteObjects: true, }); - // --- Roles --- - // - // TRUST POLICY: all three MicroVM-facing roles share `microvmAssumedBy` — - // the BARE `lambda.amazonaws.com` service principal, with NO source - // conditions. The warning and the pointer live at that constant's - // declaration, above the connector operator role that needs it first; ADR-021 - // §4 carries the evidence (P2-F1 / P2-F3) and the per-role compensating - // controls. The - // build and execution roles additionally need `sts:TagSession` alongside - // `sts:AssumeRole` (developer guide, "Trust policies"), which - // {@link grantTagSession} adds. + // Build and execution roles also require sts:TagSession. Trust constraints + // are documented beside microvmAssumedBy above. this.buildRole = new iam.Role(this, 'BuildRole', { + roleName: props.buildRoleName, assumedBy: microvmAssumedBy, description: 'ABCA Lambda MicroVMs image-build role: reads the zip+Dockerfile artifact from S3 ' @@ -1084,91 +679,44 @@ export class LambdaMicrovmCompute extends Construct { }); grantTagSession(this.buildRole, microvmAssumedBy); - // s3:GetObject only, scoped to the single artifact key — not the bucket. - // The build role runs the `/ready` and `/validate` build hooks, i.e. code - // from the repo under build, so it gets the narrowest possible read. + // Exact selected artifact plus the legacy manual-build key, never the + // bucket or every hash. The managed image only reads its immutable key. this.buildRole.addToPrincipalPolicy(new iam.PolicyStatement({ actions: ['s3:GetObject'], - resources: [this.artifactBucket.arnForObjects(this.artifactObjectKey)], + resources: [...new Set([this.artifactObjectKey, this.artifactBaseObjectKey])] + .map(key => this.artifactBucket.arnForObjects(key)), })); // Build role: keeps `logs:CreateLogGroup` — see `grantMicrovmLogWrites`. this.grantMicrovmLogWrites(this.buildRole, { allowCreateLogGroup: true }); - this.executionRole = new iam.Role(this, 'ExecutionRole', { - assumedBy: microvmAssumedBy, - description: - 'ABCA Lambda MicroVMs execution role: assumed by the running MicroVM and its runtime ' - + 'lifecycle hooks; writes logs and reads out-of-band /run payloads.', - }); - grantTagSession(this.executionRole, microvmAssumedBy); + this.executionRole = props.executionRole ?? createMicrovmExecutionRole(this, 'ExecutionRole'); // Execution role: NO `logs:CreateLogGroup`. It runs untrusted repo code and // live evidence shows it only ever writes into the pre-created group — see // `grantMicrovmLogWrites` for the runbook citations and the re-verify note. this.grantMicrovmLogWrites(this.executionRole, { allowCreateLogGroup: false }); - // The APPLICATION_LOGS group the agent is TOLD to write to (P2-F4). Separate - // from `grantMicrovmLogWrites` above and not reachable from it: that grant - // covers the service-owned `/aws/lambda-microvms/*` namespace, while - // `platform_config.log_group_name` points at the platform's vended - // `/aws/vendedlogs/bedrock-agentcore/runtime/APPLICATION_LOGS/` group — - // the one the dashboard and every per-task log query read. Scoped to that one - // group (CDK `grantWrite` → `logs:CreateLogStream` + `logs:PutLogEvents` on the - // group's ARN, stream wildcard only), so this adds no cross-log-group reach. - // See `applicationLogGroup` for the denial this fixes. + // Platform task logs use a separate namespace from MicroVM service logs. props.applicationLogGroup?.grantWrite(this.executionRole); - // READ-ONLY on the payload bucket (ADR-021: "The MicroVM execution role - // shall hold read-only access to the payload bucket, scoped to that - // bucket"). Read-only is not a nicety: the MicroVM runs untrusted repo - // code, so it must not be able to clobber another task's payload. Write + - // lifecycle stay with the trusted orchestrator. - this.payloadBucket.grantRead(this.executionRole); - - // Tenant-data access is delegated to the per-task SessionRole, exactly as - // EcsAgentCluster does for the Fargate task role. NOTE the asymmetry with - // that construct: there is no `else` branch granting DynamoDB directly — - // this backend has no legacy deployments to keep working, so a missing - // SessionRole means no tenant-data access rather than broad access. - // - // This is ALSO why nothing below grants DynamoDB: every table the agent - // touches is `task_id`-partitioned and reachable only through the - // SessionRole's `dynamodb:LeadingKeys` condition. `admitComputeRole` wires - // both halves that needs (trust on the SessionRole + `sts:AssumeRole` / - // `sts:TagSession` here), so P2 adds nothing to this seam — asserted by a - // unit test, because a "convenience" direct grant is exactly how per-tenant - // isolation gets lost. + // Authenticate only this deployment's manifests. Payload and launch reads + // using worker credentials are explicitly denied; one-object signed URLs + // carry the coordinator's authorization instead. + grantWorkerBootstrap(this.payloadBucket, this.executionRole); + + // Tenant access requires the task-scoped role; there is no direct-grant fallback. if (props.agentSessionRole) { props.agentSessionRole.admitComputeRole(this.executionRole); } - // --- P2 runtime parity on the EXECUTION role (ADR-021 "smoke parity") --- - // - // Feature-derived, not copied from `ecs-agent-cluster`: each grant below - // exists because a specific agent code path fails without it on THIS - // substrate. The ECS task role's remaining grants are deliberately absent — - // the UserConcurrencyTable (orchestrator/reconciler-owned; the agent path - // never writes it) and any artifacts-bucket access (delivery writes go - // through the SessionRole's `artifacts/${task_id}/*` statement, so the - // AgentCore runtime role has no direct grant either and neither does this). - + // Runtime startup grants complement task-scoped tenant permissions. // Secrets Manager, part 1: the GitHub PAT, read at startup before the agent // assumes the SessionRole. if (props.githubTokenSecret) { props.githubTokenSecret.grantRead(this.executionRole); } - // Secrets Manager, part 2: per-workspace Linear/Jira OAuth tokens. Same shape - // and same reason as `ecs-agent-cluster`'s grant (ABCA-488): the CLI creates - // `bgagent-linear-oauth-` / `bgagent-jira-oauth-` at setup, so - // the name is unknown at synth and a PREFIX grant is the only expressible - // scope. For a Linear/Jira-channel task the agent resolves that token at - // startup (`config.resolve_linear_api_token` / - // `resolve_jira_oauth_token`) to fire the 👀→✅ reaction and drive the channel - // MCP; without the grant the fetch hits AccessDenied and both silently no-op - // (logged by the token resolver, but invisible to the user in the channel). - // - // `GetSecretValue` ONLY — the agent reads; the orchestrator owns refresh / - // PutSecretValue. + // Workspace OAuth secrets are created after synth. Keep reads prefix-scoped; + // refresh and secret writes belong to control-plane resolvers. this.executionRole.addToPrincipalPolicy(new iam.PolicyStatement({ actions: ['secretsmanager:GetSecretValue'], resources: [ @@ -1187,17 +735,8 @@ export class LambdaMicrovmCompute extends Construct { ], })); - // Bedrock model invocation — scoped to explicit foundation-model and - // cross-Region inference-profile ARNs (parity with the AgentCore runtime and - // the ECS task role), NEVER `Resource: '*'`. The model set comes from the - // shared, context-overridable list (`constructs/bedrock-models.ts`) so no - // backend can drift from the others. - // - // Required on the COMPUTE role even though the SessionRole carries a - // session-tagged Bedrock grant for cost attribution (#215): that attribution - // is designed to FAIL OPEN — Claude Code's credential helper falls back to - // ambient compute-role credentials when the assume-role fails — so without - // this grant the fallback path AccessDenies and the task dies at turn 0. + // Use the shared model allowlist. The compute role supports the credential + // helper fallback when task-tagged model attribution cannot assume its role. const bedrockResources: string[] = []; for (const modelId of resolveBedrockModelIds(this.node)) { bedrockResources.push( @@ -1225,37 +764,37 @@ export class LambdaMicrovmCompute extends Construct { resources: bedrockResources, })); - // AgentCore Memory read+write, so cross-task learning actually persists on - // this substrate. `MEMORY_ID` already reaches the agent inside - // `agent_payload` (unchanged by P2 — it is task data, not platform config), - // which means the agent ATTEMPTS the write regardless; the grant is what - // decides whether it lands or fails closed. + // Memory is a standalone service shared by all compute backends. if (props.agentMemory) { props.agentMemory.grantReadWrite(this.executionRole); } - // A CDK-based target repo's build gate runs `cdk synth`, and a stack wired to - // a concrete env ({account, region}) does a synth-time availability-zone - // context lookup. On a developer box the gitignored cdk.context.json caches - // the answer; the agent clones fresh, so there is no cache and synth fires the - // live lookup. Without this grant the role hits AccessDenied → "Synthesis - // finished with errors" → a FALSE build-gate failure on code that builds fine - // everywhere else (the exact regression the ECS task role hit). Read-only - // describe with no resource-level scoping in IAM, so `Resource: '*'` is - // mandatory (suppressed below); it grants no mutation and no data access. + // Fresh CDK clones can need the read-only availability-zone context lookup. this.executionRole.addToPrincipalPolicy(new iam.PolicyStatement({ actions: ['ec2:DescribeAvailabilityZones'], resources: ['*'], })); // --- Image --- + if (props.managedImageVersion !== undefined + && (!props.baseImageArn || !props.baseImageVersion + || typeof props.managedImageVersion !== 'string' + || Token.isUnresolved(props.managedImageVersion) + || props.managedImageVersion.length > MAX_MANAGED_IMAGE_VERSION_LENGTH + || !/^[1-9]\d*(?:\.\d+)?$/.test(props.managedImageVersion))) { + throw new Error('microvm_managed_image_version requires managed base-image inputs and a literal positive image version'); + } // Branch through the shared predicate's two components rather than an // ad-hoc condition, so this construct and the stack's pre-TaskApi decision // (isLambdaMicrovmImageConfigured) can never disagree. if (props.baseImageArn && props.baseImageVersion) { + const requestedProtocol = props.imageEnvironmentVariables?.[sharedConstants.microvm_lifecycle.image_protocol_env]; + if (requestedProtocol !== undefined && requestedProtocol !== String(sharedConstants.microvm_lifecycle.protocol_version)) { + throw new Error('The managed MicroVM lifecycle protocol marker is owned by the image source'); + } this.image = new lambda.CfnMicrovmImage(this, 'Image', { name: this.imageName, - description: `ABCA agent snapshot for ${stack.stackName} (ADR-021 lambda-microvm backend)`, + description: `ABCA agent snapshot for ${deploymentName} (ADR-021 lambda-microvm backend)`, baseImageArn: props.baseImageArn, baseImageVersion: props.baseImageVersion, buildRoleArn: this.buildRole.roleArn, @@ -1274,43 +813,31 @@ export class LambdaMicrovmCompute extends Construct { logging: { cloudWatch: { logGroup: this.logGroup.logGroupName } }, // No extra OS capabilities: the agent runs ordinary user-space tooling. additionalOsCapabilities: [], - // Nothing baked in — see `imageEnvironmentVariables`. - environmentVariables: Object.entries(props.imageEnvironmentVariables ?? {}) + // The non-secret protocol marker belongs to this immutable image version. + // Task/deployment identity still arrives only through /run. + environmentVariables: Object.entries({ + ...props.imageEnvironmentVariables, + [sharedConstants.microvm_lifecycle.image_protocol_env]: + String(sharedConstants.microvm_lifecycle.protocol_version), + }) .map(([key, value]) => ({ key, value })), hooks: { port: AGENT_HOOK_PORT, microvmHooks: { - // Values are the `ENABLED` enum, NOT the hook route: CloudFormation - // rejects a path on every one of these four fields (P2-F2 — see - // {@link HOOK_ENABLED}). The routes are fixed and service-owned; - // {@link MICROVM_AGENT_HOOK_ROUTES} records them for the agent's - // benefit and must never be sent here again. - // - // `/run` is the payload-delivery channel (and, since P2, the - // platform-configuration channel); `/terminate` is the in-guest - // teardown breadcrumb. The agent serves BOTH — enabling a runtime - // hook nothing answers fails the corresponding lifecycle transition, - // so each is enabled only once it is served. - // - // `/suspend` and `/resume` stay OMITTED (not `DISABLED`) until P3, - // where the suspend/resume interface widening lands across all three - // strategies. Note that termination does NOT depend on this hook: - // `TerminateMicrovm` removes the VM with or without in-guest - // cooperation, which is what makes a best-effort `/terminate` safe to - // declare. + // Fixed service routes use enum switches. Automatic sleep remains + // coordinator-gated even though the image serves both lifecycle hooks. run: HOOK_ENABLED, runTimeoutInSeconds: RUN_HOOK_TIMEOUT_SECONDS, terminate: HOOK_ENABLED, terminateTimeoutInSeconds: TERMINATE_HOOK_TIMEOUT_SECONDS, + suspend: HOOK_ENABLED, + suspendTimeoutInSeconds: LIFECYCLE_HOOK_TIMEOUT_SECONDS, + resume: HOOK_ENABLED, + resumeTimeoutInSeconds: LIFECYCLE_HOOK_TIMEOUT_SECONDS, }, microvmImageHooks: { - // `/ready` is MANDATORY whenever any lifecycle hook is enabled — the - // service refuses the create otherwise (see - // READY_HOOK_TIMEOUT_SECONDS), which is why it moved from P2 to P1. - // `/validate` joins it in P2, now that the agent serves a real - // (AWS-call-free — see VALIDATE_HOOK_TIMEOUT_SECONDS) self-check: a - // hook that 404s or reports failure fails every image build, so it - // could not be enabled before there was something behind it. + // Ready is required when runtime hooks are enabled; validate checks + // the local image contract without runtime AWS credentials. ready: HOOK_ENABLED, readyTimeoutInSeconds: READY_HOOK_TIMEOUT_SECONDS, validate: HOOK_ENABLED, @@ -1324,33 +851,15 @@ export class LambdaMicrovmCompute extends Construct { this.imageIdentifier = this.image.attrImageArn; this.imageArn = this.image.attrImageArn; - // Version intentionally unpinned: the service resolves the latest ACTIVE - // version, which is what a redeploy-after-rebuild flow wants. Pinning to + // Without an explicit runtime pin, the service resolves the latest ACTIVE + // version for the redeploy-after-rebuild flow. Pinning automatically to // `attrLatestActiveImageVersion` would be empty on the very first create - // (the build has not finished) and would force a stack update per rebuild. - this.imageVersion = undefined; + // (the build has not finished). An operator can instead select a known + // version without changing the managed image resource itself. + this.imageVersion = props.managedImageVersion; } else if (props.externalImageIdentifier) { this.imageVersion = props.externalImageVersion; - // An operator may pass a bare image NAME (that is what - // `create-microvm-image --name` takes), but a bare name is useless to BOTH - // consumers: it is not an IAM resource, and — refuting this construct's - // former comment — `RunMicrovm` rejects it outright - // (`ValidationException: Malformed ARN - doesn't start with 'arn:'`, live - // 2026-07-31), as does `list-microvm-image-builds`. So the name is resolved - // to its exact ARN ONCE here and that single value feeds both - // `MICROVM_IMAGE_IDENTIFIER` and the lifecycle IAM scope. - // - // The Service Authorization Reference gives the `microvmImage` resource an - // unambiguous shape — - // `arn:${Partition}:lambda:${Region}:${Account}:microvm-image:${MicrovmImageName}` - // — matching the ARN the live run observed, so the ARN is derivable from - // the name plus this stack's partition/Region/account. `formatArn` emits - // the Aws.PARTITION / Aws.REGION / Aws.ACCOUNT_ID pseudo-parameters, so it - // is correct in a region-agnostic app too. - // - // Note the version is NOT part of the resource ARN: the SAR pattern ends - // at the image name, and `RunMicrovm` carries `imageVersion` as a separate - // request field, so pinning a version must never change the IAM scope. + // Launch and IAM require the full image ARN; version is a separate field. this.imageArn = props.externalImageIdentifier.startsWith('arn:') ? props.externalImageIdentifier : stack.formatArn({ @@ -1368,49 +877,35 @@ export class LambdaMicrovmCompute extends Construct { + 'deploy (the artifact bucket must exist before the artifact can be uploaded). Next: run ' + 'cdk/scripts/package-microvm-artifact.sh to upload the zip+Dockerfile, then redeploy with ' + '--context microvm_base_image_arn= --context microvm_base_image_version= ' + + '--context microvm_artifact_sha256= ' + '(or point at an image you built by hand with --context microvm_image_identifier=).', ); } if (this.imageIdentifier) { - // Emitted on EVERY deploy that configures an image, in both image states. - // Not a throw and not suppressible: the substrate now looks like a working - // backend in every observable way — the image builds, launches, receives a - // payload, and the execution role holds the full runtime permission set — - // while nothing has exercised clone → change → PR on it. The warning's job is - // to keep "deploy succeeded" from reading as "backend works". - // - // The id is deliberately UNCHANGED across P1→P2 (operators grep for it, and a - // rename would read as "the old warning is gone, so it must be fine"). + // Keep operational guidance independent of one installation's image + // numbers and rollout dates. Retain the ID for existing operator filters. Annotations.of(this).addWarningV2( 'abca:microvm-image-p1-smoke-unverified', - 'A MicroVM image is configured. A P2 smoke run HAS now completed clone -> change -> PR on ' - + 'this substrate (2026-08-07: two tasks COMPLETED with pull requests, progress streaming to ' - + 'bgagent watch, and the 45s agent heartbeat observed live), the agent serves the /ready, ' - + '/validate, /run and /terminate hooks, and the execution role holds its full runtime ' - + 'permission set. What is still MISSING is a run with no manual intervention: that smoke ' - + 'needed a live IAM workaround, and the two defects behind it (ADR-021 P2r2-F9 / P2r2-F10 — ' - + 'the iam:PassedToService condition on both PassRole paths) are fixed in source but NOT yet ' - + 're-exercised live. ALSO REQUIRED: re-bootstrap to policy bundle 1.6.0 or the CDK-managed ' - + 'image path fails with iam:PassRole AccessDenied on the build role. So the backend still ' - + 'carries no smoke-parity guarantee for an unattended deployment - keep production repos on ' - + 'compute_type=agentcore or ecs until a clean run is on record. Only the /suspend and ' - + '/resume runtime hooks remain undeclared, until P3 implements them: a hook the service ' - + 'calls but nothing answers fails the corresponding lifecycle transition. ' - + "(This warning's id still reads p1- by design: it is frozen across phases so operator " - + 'greps and suppression lists keep matching — read the text, not the id, for the phase.)', + 'A MicroVM image is configured. Before enabling automatic suspension, verify the configured ' + + 'image and coordinator together using the P3 acceptance procedure. The coordinator checks the ' + + 'actual launched image version before allowing sleep. The agent serves /ready, /validate, ' + + '/run, /terminate, /suspend and /resume; managed images declare all six. ' + + 'Nested deployments require bundle 1.9.0 and a reviewed migration from existing flat stacks. ' + + 'Preserve a compatible coordinator and explicit image version for rollback. Follow ' + + 'docs/verification/README.md and docs/verification/645-p3-nested-stack.md.', ); } NagSuppressions.addResourceSuppressions([this.artifactBucket, this.payloadBucket], [ { id: 'AwsSolutions-S1', - reason: 'Artifact bucket holds a single build input (the agent zip+Dockerfile) read only by ' + reason: 'Artifact bucket holds versioned agent zip+Dockerfile build inputs read only by ' + 'the Lambda MicroVMs build role; the payload bucket holds ephemeral per-task /run payloads ' - + `with a ${MICROVM_PAYLOAD_TTL_DAYS}-day TTL, written only by the orchestrator (grantPut) and ` - + 'read only by the MicroVM execution role, both scoped to the bucket. Object-level access ' - + 'logging (a second log bucket + CloudTrail data events) is not justified for a single ' - + 'build input or for transient boot payloads.', + + `with a ${MICROVM_PAYLOAD_TTL_DAYS}-day TTL, written only by the orchestrator and ` + + 'read through single-object signed URLs; the worker reads only bootstrap manifests. Object-level access ' + + 'logging (a second log bucket + CloudTrail data events) is not justified for these ' + + 'build inputs or for transient boot payloads.', }, ], true); @@ -1419,9 +914,9 @@ export class LambdaMicrovmCompute extends Construct { id: 'AwsSolutions-IAM5', reason: 'CloudWatch Logs wildcard is the service-owned ' + `${MICROVM_LOG_GROUP_PREFIX}/* namespace (log stream names are minted per MicroVM, so no ` - + 'synth-time ARN exists); S3 object/* wildcard comes from CDK grantRead on the dedicated ' - + 'payload bucket (read-only, scoped to that bucket — ADR-021 sub-decision 3). The build ' - + 'role\'s s3:GetObject is scoped to a single object key, not a wildcard. On the execution ' + + 'synth-time ARN exists); worker S3 GetObject is limited to bootstrap/* in its payload bucket, ' + + 'with explicit denial outside that prefix and for bucket listing. The build ' + + 'role\'s s3:GetObject names the selected artifact and manual-build key. On the execution ' + 'role (ADR-021 P2 runtime parity, mirroring the ECS task role): the second Logs grant is ' + 'CDK grantWrite (CreateLogStream + PutLogEvents only) on the SINGLE platform ' + 'APPLICATION_LOGS group whose name platform_config delivers to the guest, whose ARN ends ' @@ -1457,63 +952,19 @@ export class LambdaMicrovmCompute extends Construct { }, { id: 'AwsSolutions-IAM5', - reason: 'EC2 network-interface APIs are not meaningfully resource-scopable here: ' - + 'CreateNetworkInterface is authorized before the ENI exists, and the Describe* calls ' - + 'take no resource at all. The role is assumable ONLY by lambda.amazonaws.com, holds no ' - + 'data-plane permission, and is used solely to attach the two platform-owned connectors ' - + 'to the platform VPC. An aws:SourceAccount confused-deputy condition is NOT available ' - + 'on this trust: the Lambda MicroVMs service presents no source key when it assumes the ' - + 'role, and adding one makes the connector un-creatable (live-verified, ADR-021 P2-F1 — ' - + 'see ADR-021 section 4 for the per-role compensating controls). The ' - + 'AWS-managed VPC-access policy uses the same wildcard for the same reason.', + reason: 'EC2 Describe actions require wildcard resources. ENI mutations retain the ' + + 'wildcard scope validated with the MicroVM service; that does not establish that ' + + 'narrower mutation permissions are impossible. The role manages the two platform ' + + 'connectors, trusts lambda.amazonaws.com and has no tenant-data grant. Source-conditioned ' + + 'trust failed the recorded live checks; ADR-021 section 4 documents the limitation.', }, ], true); } /** - * CloudWatch Logs writes scoped to the MicroVM log namespace. - * - * `CreateLogStream` + `PutLogEvents` go to both MicroVM-facing roles; the - * `logs:CreateLogGroup` half is **build-role only**, and that asymmetry is - * evidence-based rather than tidiness: - * - * - The service documents `CreateLogGroup` for the BUILD role, and losing build - * logs costs you the one artifact you need when a snapshot build fails — a - * failure mode we have actually hit (ADR-021 P1 4.3: the 443-only SG made the - * image unbuildable, and the root cause was only readable from this group). - * So the build role keeps it. - * - The EXECUTION role does not get it. Across all three live runs (P1, P2 run - * 1, P2 run 2) the only group ever *named* under this prefix in any log or - * inventory was `/aws/lambda-microvms/`, the one CloudFormation - * pre-creates below, and both build-time and guest-runtime lines landed in it - * (P1 runbook line 1833 records that single group being deleted with the - * stack). A create right the runtime never exercises does not belong on the - * role that runs untrusted repo code. - * - * EVIDENCE STRENGTH, stated honestly, because it is a security narrowing: - * the corroborating inventory is an ABSENCE measured AFTER teardown, not a - * during-run enumeration. `645-p2-smoke-runbook.md` **§8.6** ("Billing - * confirmed stopped") records `/aws/lambda-microvms/*` log groups: **none** - * once the stack was deleted, and **§8.8** item 4 (a separate section — the - * deliberately-retained list) names the only service-vended groups created - * outside CloudFormation as `/aws/bedrock-agentcore/runtimes/…` and - * `/aws/lambda/backgroundagent-dev-…`. That combination is load-bearing - * because a service-created group is NOT a CloudFormation resource and so - * would have survived the stack delete and appeared in §8.6 — but it is - * inference from an absence, not a positive observation that no sub-group was - * ever created mid-run. Treat it as strong-but-indirect. - * - * ⚠️ RE-VERIFY on the pending clean re-run (ADR-021 P2 "the row is not yet - * fully closed"), and make it a DURING-RUN enumeration this time — an - * `aws logs describe-log-groups --log-group-name-prefix /aws/lambda-microvms/` - * taken while a task is `RUNNING` is the positive observation the post-teardown - * absence above only implies. If guest logging ever goes silent on this backend, - * this narrowing is the first thing to re-widen — the symptom would be an - * `AccessDeniedException` naming `logs:CreateLogGroup` in the guest's stdout - * fallback, which the MicroVM group still captures. - * - * @param role - the role to grant. - * @param options - `allowCreateLogGroup` gates the build-role-only half. + * Scope log writes to the MicroVM namespace. Only the build role can create log + * groups; runtime logs use the pre-created image group. Investigate a specific + * runtime denial before widening the grant. */ private grantMicrovmLogWrites( role: iam.IRole, @@ -1543,17 +994,23 @@ export class LambdaMicrovmCompute extends Construct { } } +/** Create the runtime role in its owning stack, independently of image resources. */ +export function createMicrovmExecutionRole(scope: Construct, id: string): iam.Role { + const principal = new iam.ServicePrincipal('lambda.amazonaws.com'); + const role = new iam.Role(scope, id, { + assumedBy: principal, + description: + 'ABCA Lambda MicroVMs execution role: assumed by the running MicroVM and its runtime ' + + 'lifecycle hooks; writes logs and reads deployment bootstrap manifests.', + }); + grantTagSession(role, principal); + Tags.of(role).add(MICROVM_BACKEND_TAG_KEY, MICROVM_BACKEND_TAG_VALUE); + return role; +} + /** - * Add `sts:TagSession` alongside the `sts:AssumeRole` CDK's `assumedBy` emits. - * - * The MicroVM service needs BOTH actions (developer guide, "Trust policies"), - * but `iam.Role`'s `assumedBy` only renders `sts:AssumeRole`. Passing a second - * statement through `assumeRolePolicy` keeps the two halves identical — which - * since P2-F1/F3 means "identical and unconditioned": `principal.policyFragment. - * conditions` is now empty, and the pass-through is kept deliberately rather - * than hardcoding `{}` so that if a source-condition key ever becomes usable on - * this path (see the trust-policy block in the constructor), adding it to the - * principal fixes BOTH actions instead of half-closing the hole. + * Add the service-required sts:TagSession action with the same principal and + * conditions as sts:AssumeRole. */ function grantTagSession(role: iam.Role, principal: iam.ServicePrincipal): void { role.assumeRolePolicy?.addStatements(new iam.PolicyStatement({ @@ -1564,11 +1021,7 @@ function grantTagSession(role: iam.Role, principal: iam.ServicePrincipal): void } /** - * Reduce a candidate name to the character set MicroVM image / network - * connector names accept (alphanumerics, `-`, `_`) and cap its length. - * - * Stack names can contain characters these APIs reject, and an unresolved - * (token) stack name would otherwise produce a name containing `${Token[...]}`. + * Restrict image/connector names to the service character set and length limit. */ function sanitizeImageName(candidate: string): string { const MAX_NAME_LENGTH = 64; diff --git a/cdk/src/constructs/lambda-microvm-stack.ts b/cdk/src/constructs/lambda-microvm-stack.ts new file mode 100644 index 000000000..a53855632 --- /dev/null +++ b/cdk/src/constructs/lambda-microvm-stack.ts @@ -0,0 +1,78 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { NestedStack, Stack, Tags, Token } from 'aws-cdk-lib'; +import type * as iam from 'aws-cdk-lib/aws-iam'; +import { Construct } from 'constructs'; +import { + LambdaMicrovmCompute, + type LambdaMicrovmComputeProps, + MICROVM_BACKEND_TAG_KEY, + MICROVM_BACKEND_TAG_VALUE, +} from './lambda-microvm-compute'; + +export interface LambdaMicrovmStackProps extends Omit< + LambdaMicrovmComputeProps, 'buildRoleName' | 'connectorOperatorRoleName' +> { + readonly deploymentName: string; + readonly executionRole: iam.Role; + /** Distinct image/connector/log names for an overlapping flat-to-nested migration. */ + readonly resourceNamePrefix?: string; +} + +/** Child deployment for MicroVM resources; shared runtime trust stays in the parent. */ +export class LambdaMicrovmStack extends NestedStack { + public readonly compute: LambdaMicrovmCompute; + + constructor(scope: Construct, id: string, props: LambdaMicrovmStackProps) { + super(scope, id, { + description: 'ABCA Lambda MicroVM image, storage, build roles and network connectors', + }); + if (Stack.of(props.executionRole) !== this.nestedStackParent) { + throw new Error('LambdaMicrovmStack executionRole must be owned by its parent stack'); + } + if (props.resourceNamePrefix !== undefined + && (Token.isUnresolved(props.resourceNamePrefix) + || !/^[A-Za-z0-9][A-Za-z0-9-]{0,39}$/.test(props.resourceNamePrefix))) { + throw new Error('microvm_resource_name_prefix must be 1–40 letters, digits or hyphens and start with a letter or digit'); + } + // CDK creates bucket-cleanup providers at stack scope, outside Compute. + // Those providers are also MicroVM-specific in this child. + Tags.of(this).add(MICROVM_BACKEND_TAG_KEY, MICROVM_BACKEND_TAG_VALUE); + // Automatic IAM names include the generated nested-stack name and can lose + // their discriminating role suffix to truncation. Explicit parent names keep + // the bootstrap's unconditioned PassRole grant limited to these two roles. + const roleName = (suffix: string): string => { + const name = `${props.deploymentName}-${suffix}`; + if (Token.isUnresolved(name) || !/^[A-Za-z0-9+=,.@_-]{1,64}$/.test(name)) { + throw new Error(`Nested MicroVM role name must be concrete and at most 64 characters: ${suffix}`); + } + return name; + }; + this.compute = new LambdaMicrovmCompute(this, 'Compute', { + ...props, + // Keep bootstrap-authorized IAM role names tied to the parent deployment. + // Only service names change; the old and new image may coexist during a + // reviewed migration without moving the shared execution role. + deploymentName: props.resourceNamePrefix ?? props.deploymentName, + buildRoleName: roleName('MicrovmBuildRole'), + connectorOperatorRoleName: roleName('MicrovmConnectorRole'), + }); + } +} diff --git a/cdk/src/constructs/linear-identity-vault.ts b/cdk/src/constructs/linear-identity-vault.ts index a176945fe..765f837cb 100644 --- a/cdk/src/constructs/linear-identity-vault.ts +++ b/cdk/src/constructs/linear-identity-vault.ts @@ -70,7 +70,7 @@ export interface LinearIdentityVaultProps { /** * The Linear identity vault's workload identity. Grant helpers wire the token * data-plane permissions onto whichever principal resolves Linear tokens - * (webhook processor, orchestrator, agent session role). + * (webhook processor, orchestrator, compute execution role). */ export class LinearIdentityVault extends Construct { /** The provisioned workload identity name (stable natural id). */ diff --git a/cdk/src/constructs/linear-integration.ts b/cdk/src/constructs/linear-integration.ts index 92ccca407..0afb843f3 100644 --- a/cdk/src/constructs/linear-integration.ts +++ b/cdk/src/constructs/linear-integration.ts @@ -18,7 +18,7 @@ */ import * as path from 'path'; -import { ArnFormat, Aspects, Duration, RemovalPolicy, Stack } from 'aws-cdk-lib'; +import { ArnFormat, Aspects, Duration, Fn, RemovalPolicy, Stack } from 'aws-cdk-lib'; import * as apigw from 'aws-cdk-lib/aws-apigateway'; import * as cognito from 'aws-cdk-lib/aws-cognito'; import * as dynamodb from 'aws-cdk-lib/aws-dynamodb'; @@ -78,6 +78,11 @@ export interface LinearIntegrationProps { /** The DynamoDB task events table. */ readonly taskEventsTable: dynamodb.ITable; + /** Enables task-owner decisions from replies to Linear approval comments. */ + readonly taskApprovalsTable?: dynamodb.ITable; + readonly lambdaMicrovmImageArn?: string; + readonly continuationBucketName?: string; + /** Monthly user/team budget configuration and spend table. */ readonly budgetTable?: dynamodb.ITable; @@ -159,7 +164,7 @@ export interface LinearIntegrationProps { * provider name; Phase 2.0b OAuth migration). Webhook processor and * orchestrator use this to look up which credential provider holds the * workspace's OAuth token. - * - LinearWebhookDedupTable (60s TTL dedup for webhook retries) + * - LinearWebhookDedupTable (8-hour TTL dedup for webhook retries) * - Lambda handlers for the webhook receiver, async processor, and account linking * - API Gateway routes under /linear/* * - Two Secrets Manager secrets (webhook signing secret + personal API token) @@ -178,7 +183,7 @@ export class LinearIntegration extends Construct { */ public readonly workspaceRegistryTable: dynamodb.Table; - /** Webhook dedup table — (issue_id, action) keys with 60s TTL. */ + /** Webhook dedup table — data.id/action/webhookTimestamp keys with 8-hour TTL. */ public readonly webhookDedupTable: dynamodb.Table; /** Linear webhook signing secret (placeholder — populated by `bgagent linear setup`). */ @@ -206,8 +211,7 @@ export class LinearIntegration extends Construct { this.userMappingTable = userMapping.table; this.workspaceRegistryTable = workspaceRegistry.table; - // Dedup table: linear webhook retries collapse to a single processor invoke - // within the 60s TTL window. Keyed on `{issue_id}#{action}`. + // The receiver deduplicates data.id/action/webhookTimestamp for 8 hours. this.webhookDedupTable = new dynamodb.Table(this, 'WebhookDedupTable', { partitionKey: { name: 'dedup_key', type: dynamodb.AttributeType.STRING }, billingMode: dynamodb.BillingMode.PAY_PER_REQUEST, @@ -236,6 +240,18 @@ export class LinearIntegration extends Construct { // the task-orchestrator. Used by the webhook processor's PDF attachment path. const attachmentScreeningBundling: lambda.BundlingOptions = { ...commonBundling, + // Approval replies need the current MicroVM client and durable Invoke + // fields, which cannot depend on the SDK version supplied by Lambda. + ...(props.taskApprovalsTable && { + externalModules: [ + '@aws-sdk/client-dynamodb', + '@aws-sdk/client-ecs', + '@aws-sdk/client-bedrock-runtime', + '@aws-sdk/client-secrets-manager', + '@aws-sdk/lib-dynamodb', + '@aws-sdk/util-dynamodb', + ], + }), nodeModules: ['pdf-parse'], }; @@ -245,6 +261,12 @@ export class LinearIntegration extends Construct { TASK_EVENTS_TABLE_NAME: props.taskEventsTable.tableName, TASK_RETENTION_DAYS: String(props.taskRetentionDays ?? DEFAULT_TASK_RETENTION_DAYS), }; + if (props.taskApprovalsTable) { + createTaskEnv.TASK_APPROVALS_TABLE_NAME = props.taskApprovalsTable.tableName; + } + if (props.continuationBucketName) { + createTaskEnv.CONTINUATION_BUCKET_NAME = props.continuationBucketName; + } if (props.repoTable) { createTaskEnv.REPO_TABLE_NAME = props.repoTable.tableName; } @@ -304,7 +326,7 @@ export class LinearIntegration extends Construct { }), // Throttle the seed-time root release to the free concurrency // budget (see prop doc). Only wired when both tables are present. - ...(props.orchestrationTable && props.userConcurrencyTable && { + ...((props.orchestrationTable || props.continuationBucketName) && props.userConcurrencyTable && { USER_CONCURRENCY_TABLE_NAME: props.userConcurrencyTable.tableName, MAX_CONCURRENT_TASKS_PER_USER: String(props.maxConcurrentTasksPerUser ?? 10), }), @@ -389,6 +411,20 @@ export class LinearIntegration extends Construct { } props.taskTable.grantReadWriteData(webhookProcessorFn); props.taskEventsTable.grantReadWriteData(webhookProcessorFn); + props.taskApprovalsTable?.grantReadWriteData(webhookProcessorFn); + if (props.taskApprovalsTable && props.lambdaMicrovmImageArn) { + webhookProcessorFn.addToRolePolicy(new iam.PolicyStatement({ + actions: ['lambda:GetMicrovm', 'lambda:ResumeMicrovm'], + resources: [props.lambdaMicrovmImageArn, `${props.lambdaMicrovmImageArn}:*`], + })); + } + if (props.continuationBucketName && props.orchestratorFunctionArn) { + props.userConcurrencyTable?.grantReadWriteData(webhookProcessorFn); + const coordinatorArn = Fn.join(':', Array.from({ length: 7 }, (_, i) => Fn.select(i, Fn.split(':', props.orchestratorFunctionArn!)))); + webhookProcessorFn.addToRolePolicy(new iam.PolicyStatement({ + actions: ['lambda:InvokeFunction'], resources: [`${coordinatorArn}:*`], + })); + } if (props.repoTable) { props.repoTable.grantReadData(webhookProcessorFn); } diff --git a/cdk/src/constructs/microvm-continuation-manager.ts b/cdk/src/constructs/microvm-continuation-manager.ts new file mode 100644 index 000000000..62594c7e6 --- /dev/null +++ b/cdk/src/constructs/microvm-continuation-manager.ts @@ -0,0 +1,99 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +const RECONCILER_TIMEOUT_MINUTES = 5; +const RECONCILER_SCHEDULE_MINUTES = 5; +import * as path from 'path'; +import { Duration } from 'aws-cdk-lib'; +import type * as dynamodb from 'aws-cdk-lib/aws-dynamodb'; +import * as events from 'aws-cdk-lib/aws-events'; +import * as targets from 'aws-cdk-lib/aws-events-targets'; +import * as iam from 'aws-cdk-lib/aws-iam'; +import { Architecture, Runtime } from 'aws-cdk-lib/aws-lambda'; +import * as lambda from 'aws-cdk-lib/aws-lambda-nodejs'; +import { NagSuppressions } from 'cdk-nag'; +import { Construct } from 'constructs'; +import type { ContinuationBucket } from './continuation-bucket'; + +export interface MicrovmContinuationManagerProps { + readonly taskTable: dynamodb.ITable; + readonly approvalsTable: dynamodb.ITable; + readonly userConcurrencyTable: dynamodb.ITable; + readonly continuationBucket: ContinuationBucket; + /** Unqualified function ARN; saved tasks select their original published version. */ + readonly orchestratorFunctionArn: string; + readonly imageArn: string; + readonly maxConcurrentTasksPerUser?: number; +} + +/** Recover lost continuation signals and clean saved objects after confirmed shutdown. */ +export class MicrovmContinuationManager extends Construct { + public readonly fn: lambda.NodejsFunction; + + constructor(scope: Construct, id: string, props: MicrovmContinuationManagerProps) { + super(scope, id); + this.fn = new lambda.NodejsFunction(this, 'ReconcilerFn', { + entry: path.join(__dirname, '..', 'handlers', 'reconcile-microvm-continuations.ts'), + handler: 'handler', + runtime: Runtime.NODEJS_24_X, + architecture: Architecture.ARM_64, + timeout: Duration.minutes(RECONCILER_TIMEOUT_MINUTES), + memorySize: 256, + environment: { + ABCA_COMPONENT: 'orchestr', + TASK_TABLE_NAME: props.taskTable.tableName, + TASK_APPROVALS_TABLE_NAME: props.approvalsTable.tableName, + USER_CONCURRENCY_TABLE_NAME: props.userConcurrencyTable.tableName, + CONTINUATION_BUCKET_NAME: props.continuationBucket.bucket.bucketName, + ORCHESTRATOR_FUNCTION_ARN: props.orchestratorFunctionArn, + MAX_CONCURRENT_TASKS_PER_USER: String(props.maxConcurrentTasksPerUser ?? 10), + }, + // Bundle the pinned Lambda serializer: DurableExecutionName is required + // for deduplication and may be absent from the runtime's older SDK. + bundling: { + externalModules: ['@aws-sdk/client-dynamodb', '@aws-sdk/lib-dynamodb'], + // Shared supervisor imports reach attachment screening; pdf-parse needs + // its packaged worker assets, just as in the main orchestrator bundle. + nodeModules: ['pdf-parse'], + }, + }); + props.taskTable.grantReadWriteData(this.fn); + props.approvalsTable.grantReadWriteData(this.fn); + props.userConcurrencyTable.grantReadWriteData(this.fn); + props.continuationBucket.grantCoordinator(this.fn); + this.fn.addToRolePolicy(new iam.PolicyStatement({ + actions: ['lambda:GetMicrovm', 'lambda:TerminateMicrovm'], + resources: [props.imageArn, `${props.imageArn}:*`], + })); + this.fn.addToRolePolicy(new iam.PolicyStatement({ + actions: ['lambda:InvokeFunction'], resources: [`${props.orchestratorFunctionArn}:*`], + })); + new events.Rule(this, 'Schedule', { + schedule: events.Schedule.rate(Duration.minutes(RECONCILER_SCHEDULE_MINUTES)), + targets: [new targets.LambdaFunction(this.fn)], + }); + NagSuppressions.addResourceSuppressions(this.fn, [ + { id: 'AwsSolutions-IAM4', reason: 'AWSLambdaBasicExecutionRole supplies CloudWatch runtime logs.' }, + { + id: 'AwsSolutions-IAM5', + reason: 'DynamoDB index/* grants accompany the task tables; S3 is restricted to the continuation prefix; MicroVM state/termination uses one image and its versions; invoke uses only retained versions of the one coordinator.', + }, + ], true); + } +} diff --git a/cdk/src/constructs/payload-bootstrap-permissions.ts b/cdk/src/constructs/payload-bootstrap-permissions.ts new file mode 100644 index 000000000..c6f215c0e --- /dev/null +++ b/cdk/src/constructs/payload-bootstrap-permissions.ts @@ -0,0 +1,67 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import * as iam from 'aws-cdk-lib/aws-iam'; +import * as s3 from 'aws-cdk-lib/aws-s3'; +import constants from '../../../contracts/constants.json'; + +/** Authenticate deployment settings with IAM; task payloads require a one-object capability. */ +export function grantWorkerBootstrap(bucket: s3.IBucket, worker: iam.IGrantable): void { + const manifests = bucket.arnForObjects(`${constants.payload_bootstrap.manifest_prefix}*`); + worker.grantPrincipal.addToPrincipalPolicy(new iam.PolicyStatement({ + actions: ['s3:GetObject'], + resources: [manifests], + })); + // An Allow alone is not an authentication boundary: another bucket could + // publicly grant access to an attacker's fake deployment manifest. + worker.grantPrincipal.addToPrincipalPolicy(new iam.PolicyStatement({ + effect: iam.Effect.DENY, + actions: ['s3:GetObject*'], + notResources: [manifests], + })); + worker.grantPrincipal.addToPrincipalPolicy(new iam.PolicyStatement({ + effect: iam.Effect.DENY, + actions: ['s3:List*'], + resources: [bucket.bucketArn], + })); +} + +export function grantCoordinatorPayloads(bucket: s3.IBucket, coordinator: iam.IGrantable): void { + // S3 returns AccessDenied rather than NoSuchKey for an absent launch record + // without ListBucket. Only the trusted coordinator needs this permission. + coordinator.grantPrincipal.addToPrincipalPolicy(new iam.PolicyStatement({ + actions: ['s3:ListBucket'], + resources: [bucket.bucketArn], + })); + const taskObjects = [ + bucket.arnForObjects('*/payload.json'), + bucket.arnForObjects(`*/${constants.payload_bootstrap.launch_filename}`), + ]; + coordinator.grantPrincipal.addToPrincipalPolicy(new iam.PolicyStatement({ + actions: ['s3:PutObject'], + resources: [ + ...taskObjects, + bucket.arnForObjects(`${constants.payload_bootstrap.manifest_prefix}*`), + ], + })); + coordinator.grantPrincipal.addToPrincipalPolicy(new iam.PolicyStatement({ + actions: ['s3:GetObject', 's3:DeleteObject'], + resources: taskObjects, + })); +} diff --git a/cdk/src/constructs/registry.ts b/cdk/src/constructs/registry.ts index 0332f4669..23d9e1c3a 100644 --- a/cdk/src/constructs/registry.ts +++ b/cdk/src/constructs/registry.ts @@ -23,7 +23,7 @@ // this wraps the CDK Provider framework: an `onEvent` Lambda starts the mutation // and an `isComplete` Lambda is polled until the registry reaches a stable state. import * as path from 'path'; -import { CustomResource, Duration, NestedStack, type NestedStackProps, Stack } from 'aws-cdk-lib'; +import { CfnResource, CustomResource, Duration, Names, NestedStack, type NestedStackProps, Stack } from 'aws-cdk-lib'; import * as iam from 'aws-cdk-lib/aws-iam'; import { Architecture, Runtime } from 'aws-cdk-lib/aws-lambda'; import * as lambda from 'aws-cdk-lib/aws-lambda-nodejs'; @@ -147,6 +147,23 @@ export class AgentRegistry extends Construct { totalTimeout: TOTAL_TIMEOUT, }); + // CloudFormation's generated waiter name omits the parent stack prefix, + // falling outside the bootstrap policy's backgroundagent-dev-* namespace. + // Provider has no public naming option, so use its L1 escape hatch. Include + // the outer stack name and a path hash, even inside a nested stack. + const waiters = provider.node.findAll().filter( + (node): node is CfnResource => CfnResource.isCfnResource(node) + && node.cfnResourceType === 'AWS::StepFunctions::StateMachine', + ); + if (waiters.length !== 1) { + throw new Error(`Expected one Agent Registry provider waiter, found ${waiters.length}`); + } + waiters[0].addPropertyOverride('StateMachineName', Names.uniqueResourceName(provider, { + maxLength: 80, + separator: '-', + allowedSpecialCharacters: '-', + })); + const resource = new CustomResource(this, 'Resource', { serviceToken: provider.serviceToken, // Changing the resource type forces replacement from the retired preview diff --git a/cdk/src/constructs/slack-integration.ts b/cdk/src/constructs/slack-integration.ts index fef7a117a..4140a91fe 100644 --- a/cdk/src/constructs/slack-integration.ts +++ b/cdk/src/constructs/slack-integration.ts @@ -64,6 +64,9 @@ export interface SlackIntegrationProps { /** The DynamoDB task events table (must have DynamoDB Streams enabled). */ readonly taskEventsTable: dynamodb.ITable; + /** Approval records closed atomically when their owner cancels a task. */ + readonly taskApprovalsTable?: dynamodb.ITable; + /** Monthly user/team budget configuration and spend table. */ readonly budgetTable?: dynamodb.ITable; @@ -351,6 +354,9 @@ export class SlackIntegration extends Construct { environment: { SLACK_SIGNING_SECRET_ARN: this.signingSecret.secretArn, TASK_TABLE_NAME: props.taskTable.tableName, + TASK_EVENTS_TABLE_NAME: props.taskEventsTable.tableName, + TASK_RETENTION_DAYS: String(props.taskRetentionDays ?? DEFAULT_TASK_RETENTION_DAYS), + ...(props.taskApprovalsTable && { TASK_APPROVALS_TABLE_NAME: props.taskApprovalsTable.tableName }), SLACK_USER_MAPPING_TABLE_NAME: this.userMappingTable.tableName, }, bundling: commonBundling, @@ -358,6 +364,8 @@ export class SlackIntegration extends Construct { this.signingSecret.grantRead(slackInteractionsFn); slackInteractionsFn.addToRolePolicy(readSlackSecretsPolicy); props.taskTable.grantReadWriteData(slackInteractionsFn); + props.taskEventsTable.grant(slackInteractionsFn, 'dynamodb:PutItem'); + props.taskApprovalsTable?.grant(slackInteractionsFn, 'dynamodb:GetItem', 'dynamodb:UpdateItem'); this.userMappingTable.grantReadData(slackInteractionsFn); // --- Slash Command Acknowledger --- diff --git a/cdk/src/constructs/stranded-task-reconciler.ts b/cdk/src/constructs/stranded-task-reconciler.ts index cb3dbdbc6..1cfd54ef8 100644 --- a/cdk/src/constructs/stranded-task-reconciler.ts +++ b/cdk/src/constructs/stranded-task-reconciler.ts @@ -30,8 +30,8 @@ import { Construct } from 'constructs'; /** Default stranded-timeout (seconds; 20 minutes). */ const DEFAULT_STRANDED_TIMEOUT_SECONDS = 1200; -/** Default approval-stranded timeout (seconds; 2 hours). */ -const DEFAULT_APPROVAL_STRANDED_TIMEOUT_SECONDS = 7200; +/** Backstop for approval waits without a checkpoint (eight hours plus 30 minutes). */ +const DEFAULT_APPROVAL_STRANDED_TIMEOUT_SECONDS = 30600; /** Default task-record retention used for event TTL (days). */ const DEFAULT_TASK_RETENTION_DAYS = 90; @@ -54,8 +54,10 @@ export interface StrandedTaskReconcilerProps { /** TaskEventsTable (handler writes task_stranded + task_failed events). */ readonly taskEventsTable: dynamodb.ITable; + /** Close unanswered requests when the owning task is declared stranded. */ + readonly taskApprovalsTable?: dynamodb.ITable; - /** UserConcurrencyTable (handler decrements active_count on fail). */ + /** UserConcurrencyTable (handler atomically releases held task reservations). */ readonly userConcurrencyTable: dynamodb.ITable; /** @@ -77,13 +79,12 @@ export interface StrandedTaskReconcilerProps { readonly strandedTimeoutSeconds?: number; /** - * Cedar HITL approval-stranded timeout (seconds). Tasks in - * AWAITING_APPROVAL older than this are transitioned to FAILED. - * Longer than the stranded-timeout because approvals legitimately - * sit for up to an hour (§7.3). Set via + * Backstop for AWAITING_APPROVAL tasks without a saved continuation. + * Allows the worker's full service lifetime and coordinator cleanup. + * Saved MicroVM requests are handled by the continuation manager. Set via * ``APPROVAL_STRANDED_TIMEOUT_SECONDS``. * - * @default 7200 (2 hours — double §7.3's 1-hour ceiling + an hour grace) + * @default 30600 (8.5 hours) */ readonly approvalStrandedTimeoutSeconds?: number; @@ -126,6 +127,7 @@ export class StrandedTaskReconciler extends Construct { // Solution-attribution component label (#319): orchestration plane. ABCA_COMPONENT: 'orchestr', TASK_TABLE_NAME: props.taskTable.tableName, + ...(props.taskApprovalsTable && { TASK_APPROVALS_TABLE_NAME: props.taskApprovalsTable.tableName }), TASK_EVENTS_TABLE_NAME: props.taskEventsTable.tableName, USER_CONCURRENCY_TABLE_NAME: props.userConcurrencyTable.tableName, STRANDED_TIMEOUT_SECONDS: String(strandedTimeout), @@ -140,9 +142,10 @@ export class StrandedTaskReconciler extends Construct { // TaskTable: read (query by StatusIndex) + conditional UpdateItem to // transition stranded rows to FAILED. props.taskTable.grantReadWriteData(this.fn); + props.taskApprovalsTable?.grantReadWriteData(this.fn); // TaskEvents: write task_stranded + task_failed events. props.taskEventsTable.grantWriteData(this.fn); - // Concurrency: decrement active_count on fail. + // Concurrency: read/release a task-owned reservation transactionally. props.userConcurrencyTable.grantReadWriteData(this.fn); const schedule = props.schedule ?? Duration.minutes(DEFAULT_SCHEDULE_MINUTES); diff --git a/cdk/src/constructs/task-api.ts b/cdk/src/constructs/task-api.ts index be735cad2..6e9da19c2 100644 --- a/cdk/src/constructs/task-api.ts +++ b/cdk/src/constructs/task-api.ts @@ -200,10 +200,8 @@ export interface TaskApiProps { /** * IAM resource ARN of the MicroVM image this deployment provisioned (ADR-021 - * sub-decision 4). When provided, the cancel Lambda gets - * `lambda:TerminateMicrovm` **scoped to that one image**, so a cancelled - * MicroVM-backed task actually stops billing — mirroring the conditional - * AgentCore `RUNTIME_ARN` / `ecsClusterArn` wiring above. + * sub-decision 4). When provided, cancel gets TerminateMicrovm and the approval + * decision handlers get GetMicrovm/ResumeMicrovm, scoped to that one image. * * An ARN rather than an on/off boolean so the grant is exactly scoped. `TaskApi` * is constructed before `LambdaMicrovmCompute` (the cancel Lambda's ARN is @@ -228,7 +226,7 @@ export interface TaskApiProps { readonly attachmentsBucket?: s3.IBucket; /** - * User concurrency table for admission control during confirm-uploads. + * User concurrency table for the advisory confirm-uploads pre-check. * Required when attachmentsBucket is provided. */ readonly userConcurrencyTable?: dynamodb.ITable; @@ -261,6 +259,8 @@ export interface TaskApiProps { * - DELETE /api-keys/{key_id} → deleteApiKey (Cognito) */ export class TaskApi extends Construct { + private readonly approvalDecisionFunctions: lambda.NodejsFunction[] = []; + /** * The API Gateway REST API. */ @@ -654,6 +654,9 @@ export class TaskApi extends Construct { }); const cancelTaskEnv: Record = { ...commonEnv }; + if (props.taskApprovalsTable) { + cancelTaskEnv.TASK_APPROVALS_TABLE_NAME = props.taskApprovalsTable.tableName; + } const stopSessionArn = props.agentCoreStopSessionRuntimeArn; if (stopSessionArn) { cancelTaskEnv.RUNTIME_ARN = stopSessionArn; @@ -669,10 +672,8 @@ export class TaskApi extends Construct { architecture: Architecture.ARM_64, environment: cancelTaskEnv, bundling: commonBundling, - // Cancel performs: DDB GetItem + DDB UpdateItem + ECS StopTask or - // AgentCore StopRuntimeSession + DDB PutItem. The default 3s timeout - // is not enough once cold-start TLS handshakes for bedrock-agentcore - // are added. 15s gives comfortable headroom. + // Cancel reads state, atomically closes an unanswered approval, stops + // compute and writes an event. The state operation has its own 5s bound. timeout: Duration.seconds(API_HANDLER_TIMEOUT_SECONDS), memorySize: API_HANDLER_MEMORY_MB, }); @@ -711,6 +712,7 @@ export class TaskApi extends Construct { props.taskEventsTable.grantReadWriteData(createTaskFn); props.taskTable.grantReadWriteData(cancelTaskFn); props.taskEventsTable.grantReadWriteData(cancelTaskFn); + props.taskApprovalsTable?.grant(cancelTaskFn, 'dynamodb:GetItem', 'dynamodb:UpdateItem'); if (stopSessionArn) { cancelTaskFn.addToRolePolicy(new iam.PolicyStatement({ @@ -733,15 +735,14 @@ export class TaskApi extends Construct { // ADR-021: cancelling a `lambda-microvm` task must actively terminate the // MicroVM — leaving it to the 8-hour `maximumDurationInSeconds` cap would - // keep billing 16 vCPU and would hold account memory quota that gates - // admission of new tasks. Conditional for the same reason the AgentCore and - // ECS grants above are: a deployment without the backend gets no grant. + // keep running compute or retained snapshots billable. Conditional for the + // same reason the AgentCore and ECS grants above are: a deployment without + // the backend gets no grant. // // ONLY `lambda:TerminateMicrovm`. `cancel-task.ts` sends // `TerminateMicrovmCommand` and nothing else — it does not read MicroVM // state first — so `lambda:GetMicrovm` would be a permission with no caller. - // (The approve/deny Lambdas get `ResumeMicrovm` + `GetMicrovm` in P3, where - // a state read is genuinely needed for the resume reconciliation.) + // The approve/deny Lambdas separately get ResumeMicrovm + GetMicrovm below. // // Resource is the MicroVM *image*, not the running instance: every MicroVM // lifecycle action authorizes against `microvm-image:` (Service @@ -836,7 +837,7 @@ export class TaskApi extends Construct { props.taskEventsTable.grantReadWriteData(confirmUploadsFn); props.attachmentsBucket.grantReadWrite(confirmUploadsFn); props.attachmentsBucket.grantDelete(confirmUploadsFn); - props.userConcurrencyTable.grantReadWriteData(confirmUploadsFn); + props.userConcurrencyTable.grantReadData(confirmUploadsFn); if (props.orchestratorFunctionArn) { confirmUploadsFn.addToRolePolicy(new iam.PolicyStatement({ @@ -999,6 +1000,10 @@ export class TaskApi extends Construct { ...commonEnv, TASK_APPROVALS_TABLE_NAME: props.taskApprovalsTable.tableName, }; + const decisionBundling = { + ...commonBundling, + externalModules: commonBundling.externalModules?.filter(name => name !== '@aws-sdk/client-lambda'), + }; // ApproveTaskFn — POST /tasks/{task_id}/approve const approveTaskFn = new lambda.NodejsFunction(this, 'ApproveTaskFn', { @@ -1007,7 +1012,7 @@ export class TaskApi extends Construct { runtime: Runtime.NODEJS_24_X, architecture: Architecture.ARM_64, environment: approvalEnv, - bundling: commonBundling, + bundling: decisionBundling, timeout: Duration.seconds(API_HANDLER_TIMEOUT_SECONDS), memorySize: API_HANDLER_MEMORY_MB, }); @@ -1022,13 +1027,22 @@ export class TaskApi extends Construct { runtime: Runtime.NODEJS_24_X, architecture: Architecture.ARM_64, environment: approvalEnv, - bundling: commonBundling, + bundling: decisionBundling, timeout: Duration.seconds(API_HANDLER_TIMEOUT_SECONDS), memorySize: API_HANDLER_MEMORY_MB, }); props.taskTable.grantReadWriteData(denyTaskFn); props.taskApprovalsTable.grantReadWriteData(denyTaskFn); props.taskEventsTable.grantReadWriteData(denyTaskFn); + this.approvalDecisionFunctions.push(approveTaskFn, denyTaskFn); + if (props.lambdaMicrovmImageArn) { + for (const decisionFn of [approveTaskFn, denyTaskFn]) { + decisionFn.addToRolePolicy(new iam.PolicyStatement({ + actions: ['lambda:GetMicrovm', 'lambda:ResumeMicrovm'], + resources: [props.lambdaMicrovmImageArn, `${props.lambdaMicrovmImageArn}:*`], + })); + } + } // GetPendingFn — GET /pending const getPendingFn = new lambda.NodejsFunction(this, 'GetPendingFn', { @@ -1041,6 +1055,7 @@ export class TaskApi extends Construct { timeout: Duration.seconds(10), memorySize: API_HANDLER_MEMORY_MB, }); + props.taskTable.grant(getPendingFn, 'dynamodb:BatchGetItem'); // Least-privilege: GetPendingFn only reads (Query on // user_id-status-index for the user's pending rows) and writes // a synthetic ``RATE##PENDING`` rate-limit row @@ -1428,7 +1443,7 @@ export class TaskApi extends Construct { }, { id: 'AwsSolutions-IAM5', - reason: 'DynamoDB index/* wildcards generated by CDK grantReadWriteData/grantReadData for GSI access; ecs:StopTask is conditioned on the cluster ARN; lambda:TerminateMicrovm is scoped to the single platform MicroVM image ARN plus a :* version-suffix sibling (delivered as a Lazy.string because TaskApi is built before the MicroVM construct) — ADR-021', + reason: 'DynamoDB index/* wildcards generated by CDK grantReadWriteData/grantReadData for GSI access; ecs:StopTask is conditioned on the cluster ARN; Lambda MicroVM cancel/wake actions (TerminateMicrovm/GetMicrovm/ResumeMicrovm) are scoped to the single platform MicroVM image ARN plus a :* version-suffix sibling (delivered as a Lazy.string because TaskApi is built before the MicroVM construct) — ADR-021', }, ], true); } @@ -1440,4 +1455,21 @@ export class TaskApi extends Construct { }, ], true); } + + /** Wire after the coordinator and bucket exist, avoiding a construction-order cycle. */ + public enableMicrovmContinuations( + bucketName: string, coordinatorArn: string, concurrencyTable: dynamodb.ITable, + maxConcurrentTasksPerUser: number, + ): void { + for (const fn of this.approvalDecisionFunctions) { + fn.addEnvironment('CONTINUATION_BUCKET_NAME', bucketName); + fn.addEnvironment('ORCHESTRATOR_FUNCTION_ARN', coordinatorArn); + fn.addEnvironment('USER_CONCURRENCY_TABLE_NAME', concurrencyTable.tableName); + fn.addEnvironment('MAX_CONCURRENT_TASKS_PER_USER', String(maxConcurrentTasksPerUser)); + concurrencyTable.grantReadWriteData(fn); + fn.addToRolePolicy(new iam.PolicyStatement({ + actions: ['lambda:InvokeFunction'], resources: [`${coordinatorArn}:*`], + })); + } + } } diff --git a/cdk/src/constructs/task-approvals-table.ts b/cdk/src/constructs/task-approvals-table.ts index 20bb4e6ef..159bae293 100644 --- a/cdk/src/constructs/task-approvals-table.ts +++ b/cdk/src/constructs/task-approvals-table.ts @@ -60,9 +60,9 @@ export interface TaskApprovalsTableProps { * * Schema: `task_id` (PK, ULID matching TaskTable) + `request_id` (SK, * ULID minted by the agent). Each row represents one human-in-the-loop - * approval gate; the agent writes PENDING, the ApproveTaskFn / - * DenyTaskFn Lambdas (Chunk 5) update to APPROVED / DENIED, and the - * reconciler sweeps STRANDED rows. + * approval gate. The trusted approval request service creates PENDING rows; + * owner-authenticated decision handlers record APPROVED / DENIED. Task closure + * cancels unanswered requests. Workers have no direct approval-row write grant. * * A GSI (`user_id-status-index`) supports the `bgagent pending` access * pattern — `user_id = :caller AND status = :pending` — without @@ -75,8 +75,9 @@ export interface TaskApprovalsTableProps { * streams wired into the fan-out Lambda. Enabling streams here would * create duplicate fan-out paths. * - * TTL is sized by the agent as `created_at_epoch + timeout_s + 120s` - * so rows never expire during the decision window (§10.1). + * Pending rows have no TTL, including requests with an explicit deadline. + * Task closure applies retention cleanup; capacity delays must never erase + * the recorded decision before a replacement worker consumes it. */ export class TaskApprovalsTable extends Construct { /** @@ -116,10 +117,9 @@ export class TaskApprovalsTable extends Construct { // GSI for GET /v1/pending — user_id PK + status SK (§10.1). // - // Projection is INCLUDE with exactly the non-key attributes the - // pending-list endpoint needs: keeps per-write cost small while - // keeping the list response small enough to render in the CLI - // without additional GetItem round-trips. + // Preserve the existing INCLUDE projection. The pending-list endpoint uses + // this GSI to discover candidates, then strongly reads approval/task rows + // before returning them; GSI results alone may contain stale decisions. this.table.addGlobalSecondaryIndex({ indexName: USER_STATUS_INDEX_NAME, partitionKey: { @@ -140,19 +140,8 @@ export class TaskApprovalsTable extends Construct { 'reason', 'created_at', 'timeout_s', - // Cedar HITL: surface which rule(s) fired on the gate in the - // pending-list response so `bgagent pending` can show _why_ - // without a second read against the base table. Projected - // because the handler reads rows through this GSI. - // - // ARCHITECTURAL NOTE: DynamoDB rejects in-place updates to - // ``nonKeyAttributes`` on an existing GSI. Any future field - // that needs to appear on the pending view must be decided - // here at design time — adding one post-hoc requires a - // destructive migration (delete + recreate the table, or - // create a parallel GSI under a new name with shadow - // backfill). Chunks that extend TaskApprovalsTable should - // audit this list before shipping. + // DynamoDB rejects in-place projection changes. New response fields + // can come from the existing base-table read without changing this GSI. 'matching_rule_ids', ], }); diff --git a/cdk/src/constructs/task-orchestrator.ts b/cdk/src/constructs/task-orchestrator.ts index ba3398883..94795b79e 100644 --- a/cdk/src/constructs/task-orchestrator.ts +++ b/cdk/src/constructs/task-orchestrator.ts @@ -18,7 +18,7 @@ */ import * as path from 'path'; -import { ArnFormat, Duration, Stack } from 'aws-cdk-lib'; +import { ArnFormat, Duration, RemovalPolicy, Stack } from 'aws-cdk-lib'; import * as cloudwatch from 'aws-cdk-lib/aws-cloudwatch'; import * as dynamodb from 'aws-cdk-lib/aws-dynamodb'; import * as iam from 'aws-cdk-lib/aws-iam'; @@ -26,6 +26,7 @@ import { Runtime, Architecture } from 'aws-cdk-lib/aws-lambda'; import * as lambda from 'aws-cdk-lib/aws-lambda-nodejs'; import * as s3 from 'aws-cdk-lib/aws-s3'; import * as secretsmanager from 'aws-cdk-lib/aws-secretsmanager'; +import * as ssm from 'aws-cdk-lib/aws-ssm'; import { NagSuppressions } from 'cdk-nag'; import { Construct } from 'constructs'; @@ -51,10 +52,17 @@ const ORCHESTRATOR_TIMEOUT_SECONDS = 60; /** Orchestrator Lambda memory (MB). */ const ORCHESTRATOR_MEMORY_MB = 1024; +import type { ContinuationBucket } from './continuation-bucket'; +import { grantCoordinatorPayloads } from './payload-bootstrap-permissions'; + /** * Properties for TaskOrchestrator construct. */ export interface TaskOrchestratorProps { + /** Close outstanding requests and apply retention after task completion. */ + readonly taskApprovalsTable?: dynamodb.ITable; + /** Versioned durable checkpoints and launch inputs for MicroVM worker replacement. */ + readonly continuationBucket?: ContinuationBucket; /** * The DynamoDB task table. */ @@ -175,12 +183,11 @@ export interface TaskOrchestratorProps { }; /** - * S3 bucket for per-task ECS payloads. When provided (alongside - * ``ecsConfig``), the orchestrator writes the payload here and passes only an - * ``AGENT_PAYLOAD_S3_URI`` pointer in the RunTask override (the full payload - * exceeds the 8 KB containerOverrides limit), then deletes the object in the - * finalize step. The orchestrator gets write + delete; the ECS task role gets - * read-only (granted on the bucket by ``EcsAgentCluster``). + * S3 storage for ECS v2 bootstrap manifests, task payloads and private launch + * references. The coordinator publishes/signs/replays the reference delivered + * in AGENT_PAYLOAD_REF, then deletes both task objects at finalization. + * EcsAgentCluster separately grants worker bootstrap-only reads and denies + * other object reads and payload-bucket listing. */ readonly ecsPayloadBucket?: s3.IBucket; @@ -209,11 +216,9 @@ export interface TaskOrchestratorProps { * ## Names, ARNs — and NO grants * * Every field is an identifier, never a secret value, and NONE of them adds an - * IAM grant to the orchestrator role: it forwards these strings and never calls - * the resources they name (the agent does, through its own execution role / - * SessionRole). The approvals and nudges tables in particular stay ungranted to - * the orchestrator, which is asserted by a unit test — a "while I'm here" grant - * would hand the orchestration plane tenant-data access it has never needed. + * IAM grant to the orchestrator role. P3's separate `microvmConfig` grants + * explicit approval reads/condition checks for lifecycle supervision. Forwarding + * these names alone still grants no access to approvals, nudges or tenant roles. * * ## All-or-nothing, and wired unconditionally * @@ -224,8 +229,8 @@ export interface TaskOrchestratorProps { * only ever fire for a hand-edited Lambda environment — never because a * deploy-time gate and a per-repo `compute_type` disagreed. * - * Optional as a prop only so isolated construct tests can omit it. Four of the - * thirteen `platform_config` keys come from env vars the orchestrator already + * Optional as a prop only so isolated construct tests can omit it. Some + * `platform_config` keys come from env vars the orchestrator already * carries for its own work (`TASK_TABLE_NAME`, `TASK_EVENTS_TABLE_NAME`, * `GITHUB_TOKEN_SECRET_ARN`) or from the stack-wide `SolutionUaAspect` * (`AWS_SDK_UA_APP_ID`), so they are deliberately NOT repeated here. @@ -237,6 +242,7 @@ export interface TaskOrchestratorProps { * fails closed with `approval_write_failed`. */ readonly taskApprovalsTableName: string; + readonly approvalRequestsApiUrl?: string; /** Nudges table (`NUDGES_TABLE_NAME`) the agent polls for mid-task nudges. */ readonly nudgesTableName: string; /** Application log group (`LOG_GROUP_NAME`) the agent writes progress logs to. */ @@ -294,8 +300,8 @@ export interface TaskOrchestratorProps { * * `ingressConnectorArns` is required for a different reason — it is a security * control whose absence has a *wider* meaning than "off" (see the field). Only - * `imageVersion` is genuinely optional, and its absent state ("let the service - * resolve the latest ACTIVE version") is a real, intended configuration. + * `imageVersion` may be omitted to resolve the latest ACTIVE version. + * Approval suspension defaults off while wake/cleanup stay available. */ readonly microvmConfig?: { /** @@ -326,6 +332,10 @@ export interface TaskOrchestratorProps { * active version, which is what a rebuild-in-place flow wants. */ readonly imageVersion?: string; + /** Coordinator reads and condition-checks the current gate before sleeping. */ + readonly approvalsTable: dynamodb.ITable; + /** Static opt-in and live Parameter Store value for new suspends; default false. */ + readonly approvalSuspendEnabled?: boolean; /** Role the MicroVM assumes at runtime; passed on `RunMicrovm`. */ readonly executionRoleArn: string; /** Egress network connectors; comma-joined into the env var. */ @@ -350,9 +360,9 @@ export interface TaskOrchestratorProps { /** * Bucket for `/run` payloads that exceed the 4 KB `runHookPayload` cap — * i.e. nearly all of them, since a hydrated payload is bigger than that. - * The orchestrator gets **write only**: unlike the ECS payload bucket - * there is no finalize-time delete on this backend (the bucket's lifecycle - * rule is the reaper), so `grantDelete` would be an unused permission. + * The orchestrator uploads payloads and deletes `/payload.json` at + * finalization. It does not read them; the execution role is the reader. + * Lifecycle expiry is the fallback if finalization or deletion fails. */ readonly payloadBucket: s3.IBucket; }; @@ -388,6 +398,9 @@ export class TaskOrchestrator extends Construct { constructor(scope: Construct, id: string, props: TaskOrchestratorProps) { super(scope, id); + if (props.agentPlatformConfig && !props.agentPlatformConfig.approvalRequestsApiUrl) { + throw new Error('agentPlatformConfig requires approvalRequestsApiUrl; deploy the matching approval service'); + } if (props.guardrailId && !props.guardrailVersion) { throw new Error('guardrailVersion is required when guardrailId is provided'); } @@ -397,6 +410,12 @@ export class TaskOrchestrator extends Construct { const handlersDir = path.join(__dirname, '..', 'handlers'); const maxConcurrent = props.maxConcurrentTasksPerUser ?? 10; + const suspendParameter = props.microvmConfig ? new ssm.StringParameter(this, 'MicrovmApprovalSuspendEnabled', { + parameterName: `/${Stack.of(this).stackName}/microvm-approval-suspend-enabled`, + stringValue: String(props.microvmConfig.approvalSuspendEnabled ?? false), + description: 'Allow new approval suspensions; existing durable executions reread before suspending.', + allowedPattern: '^(true|false)$', + }) : undefined; // Hydration pulls in bedrock-agentcore (bundled), durable-execution, and // attachment screening (URL resolution). pdf-parse is needed for PDF text @@ -419,7 +438,6 @@ export class TaskOrchestrator extends Construct { '@aws-sdk/client-ecs', '@aws-sdk/client-lambda', '@aws-sdk/client-bedrock-runtime', - '@aws-sdk/client-s3', '@aws-sdk/client-secrets-manager', '@aws-sdk/lib-dynamodb', '@aws-sdk/util-dynamodb', @@ -440,10 +458,14 @@ export class TaskOrchestrator extends Construct { executionTimeout: Duration.hours(DURABLE_EXECUTION_TIMEOUT_HOURS), retentionPeriod: Duration.days(DURABLE_RETENTION_DAYS), }, + // Durable executions replay their original code and environment after a + // deployment. Keep published versions until no execution can resume them. + currentVersionOptions: { removalPolicy: RemovalPolicy.RETAIN }, environment: { // Solution-attribution component label (#319): orchestration plane. ABCA_COMPONENT: 'orchestr', TASK_TABLE_NAME: props.taskTable.tableName, + ...(props.taskApprovalsTable && { TASK_APPROVALS_TABLE_NAME: props.taskApprovalsTable.tableName }), TASK_EVENTS_TABLE_NAME: props.taskEventsTable.tableName, USER_CONCURRENCY_TABLE_NAME: props.userConcurrencyTable.tableName, RUNTIME_ARN: props.runtimeArn, @@ -488,6 +510,12 @@ export class TaskOrchestrator extends Construct { // unconditional; there is no "no ingress configured" state to express. MICROVM_INGRESS_CONNECTOR_ARNS: props.microvmConfig.ingressConnectorArns.join(','), MICROVM_PAYLOAD_BUCKET: props.microvmConfig.payloadBucket.bucketName, + ...(props.continuationBucket && { + CONTINUATION_BUCKET_NAME: props.continuationBucket.bucket.bucketName, + }), + TASK_APPROVALS_TABLE_NAME: props.microvmConfig.approvalsTable.tableName, + MICROVM_APPROVAL_SUSPEND_ENABLED: String(props.microvmConfig.approvalSuspendEnabled ?? false), + MICROVM_APPROVAL_SUSPEND_PARAMETER_NAME: suspendParameter!.parameterName, ...(props.microvmConfig.imageVersion && { MICROVM_IMAGE_VERSION: props.microvmConfig.imageVersion, }), @@ -501,7 +529,10 @@ export class TaskOrchestrator extends Construct { // PLATFORM_CONFIG_ENV_VARS map verbatim — one stack value, one name, three // backends. NO IAM grant accompanies any of these (see the prop docs). ...(props.agentPlatformConfig && { - TASK_APPROVALS_TABLE_NAME: props.agentPlatformConfig.taskApprovalsTableName, + TASK_APPROVALS_TABLE_NAME: props.microvmConfig?.approvalsTable.tableName ?? props.agentPlatformConfig.taskApprovalsTableName, + ...(props.agentPlatformConfig.approvalRequestsApiUrl && { + APPROVAL_REQUESTS_API_URL: props.agentPlatformConfig.approvalRequestsApiUrl, + }), NUDGES_TABLE_NAME: props.agentPlatformConfig.nudgesTableName, LOG_GROUP_NAME: props.agentPlatformConfig.logGroupName, ARTIFACTS_BUCKET_NAME: props.agentPlatformConfig.artifactsBucketName, @@ -524,6 +555,7 @@ export class TaskOrchestrator extends Construct { // DynamoDB grants props.taskTable.grantReadWriteData(this.fn); + props.taskApprovalsTable?.grantReadWriteData(this.fn); props.taskEventsTable.grantReadWriteData(this.fn); props.userConcurrencyTable.grantReadWriteData(this.fn); if (props.repoTable) { @@ -535,24 +567,11 @@ export class TaskOrchestrator extends Construct { props.attachmentsBucket.grantReadWrite(this.fn); } - // ECS payload bucket — the orchestrator writes the payload before - // RunTask and deletes it at finalize. Write + delete only (it never reads - // its own payload back; the ECS container is the reader, with its own - // read-only grant from EcsAgentCluster). - if (props.ecsPayloadBucket) { - props.ecsPayloadBucket.grantPut(this.fn); - props.ecsPayloadBucket.grantDelete(this.fn); - } - - // ADR-021: MicroVM payload bucket. WRITE only — the strategy uploads an - // oversized /run payload and never reads it back (the MicroVM execution - // role is the reader, with its own read-only grant), and it never deletes - // (the bucket's lifecycle rule reaps). No grantDelete, deliberately: the ECS - // path has one because the orchestrator deletes at finalize; this one does - // not, so the grant would be dead permission. - if (props.microvmConfig) { - props.microvmConfig.payloadBucket.grantPut(this.fn); - } + // Publish manifests/payloads, sign one-object reads and persist private + // launch references for replay. Workers cannot read the launch records. + if (props.ecsPayloadBucket) grantCoordinatorPayloads(props.ecsPayloadBucket, this.fn); + if (props.microvmConfig) grantCoordinatorPayloads(props.microvmConfig.payloadBucket, this.fn); + props.continuationBucket?.grantCoordinator(this.fn); // Durable execution managed policy this.fn.role!.addManagedPolicy( @@ -645,20 +664,17 @@ export class TaskOrchestrator extends Construct { // Lambda MicroVMs compute strategy permissions (only when configured). // - // EXACTLY the four control-plane actions the P1 strategy calls, per + // Control-plane actions used by the strategy, per // ADR-021's "only the MicroVM lifecycle actions it calls" requirement: // RunMicrovm — startSession // GetMicrovm — pollSession + // GetMicrovmImageVersion — attest the actual launched snapshot's lifecycle hooks // TerminateMicrovm — stopSession / finalize (the active cleanup path) + // SuspendMicrovm / ResumeMicrovm — durable approval-wait supervision // PassNetworkConnector — required to attach egress connectors, even the // AWS-managed ones // // NOT granted, deliberately: - // - lambda:SuspendMicrovm / lambda:ResumeMicrovm — the ADR's grant list - // names them, but P1 has no suspend/resume code path. They land with the - // P3 interface widening (mandatory suspendSession/resumeSession across - // all three strategies) together with the approve/deny Lambdas' - // conditional ResumeMicrovm + GetMicrovm. // - lambda:CreateMicrovmAuthToken — granted to no role in any phase; no // JWE consumer exists (ADR-021 sub-decision 3). if (props.microvmConfig) { @@ -691,19 +707,26 @@ export class TaskOrchestrator extends Construct { actions: [ 'lambda:RunMicrovm', 'lambda:GetMicrovm', + 'lambda:GetMicrovmImageVersion', 'lambda:TerminateMicrovm', + 'lambda:SuspendMicrovm', + 'lambda:ResumeMicrovm', ], resources: microvmImageResources, })); + this.fn.addToRolePolicy(new iam.PolicyStatement({ + sid: 'MicrovmApprovalObservation', + actions: ['dynamodb:GetItem', 'dynamodb:ConditionCheckItem'], + resources: [props.microvmConfig.approvalsTable.tableArn], + })); + this.fn.addToRolePolicy(new iam.PolicyStatement({ + sid: 'MicrovmSuspendConfiguration', + actions: ['ssm:GetParameter'], + resources: [suspendParameter!.parameterArn], + })); - // `lambda:PassNetworkConnector` supports NO resource-level permissions - // (the Service Authorization Reference lists no resource type for it), so - // `Resource: '*'` is mandatory — a narrowed ARN would simply never match - // and RunMicrovm would fail with AccessDenied. It is also why the ADR - // notes the action is needed "even for the default connectors": the - // AWS-managed connectors live in the `aws` account, outside any ARN we - // could enumerate. The action only permits *passing* a connector to a - // service, not creating or reading one. + // PassNetworkConnector has no resource-level authorization support. Its + // wildcard permits attaching connectors, not creating or inspecting them. this.fn.addToRolePolicy(new iam.PolicyStatement({ sid: 'MicrovmPassNetworkConnector', actions: ['lambda:PassNetworkConnector'], @@ -821,7 +844,7 @@ export class TaskOrchestrator extends Construct { }, { id: 'AwsSolutions-IAM5', - reason: 'DynamoDB index/* wildcards generated by CDK grantReadWriteData; AgentCore runtime/* required for sub-resource invocation; Secrets Manager wildcards generated by CDK grantRead; AgentCore Memory wildcards generated by CDK grantRead/grantWrite; ECS RunTask/DescribeTasks/StopTask conditioned on cluster ARN; iam:PassRole scoped to ECS task/execution roles and conditioned on ecs-tasks.amazonaws.com; S3 object/* wildcard from CDK grantPut on the dedicated MicroVM payload bucket; MicroVM lifecycle actions (RunMicrovm/GetMicrovm/TerminateMicrovm) are scoped to the single platform MicroVM image ARN plus a :* version-suffix sibling (every one of them authorizes against the image resource, not the per-session instance; no account-wide wildcard is used); lambda:PassNetworkConnector requires Resource:* because the action supports no resource-level permissions and the AWS-managed connectors live outside this account; iam:PassRole is scoped to the MicroVM execution role and conditioned on lambda.amazonaws.com; Agent Registry read scoped to the wired registry ARN, with a record/* suffix wildcard because record ids are server-assigned and unknown at synth (#246)', + reason: 'DynamoDB index/* wildcards generated by CDK grantReadWriteData; AgentCore runtime/* required for sub-resource invocation; Secrets Manager wildcards generated by CDK grantRead; AgentCore Memory wildcards generated by CDK grantRead/grantWrite; ECS RunTask/DescribeTasks/StopTask conditioned on cluster ARN; iam:PassRole scoped to ECS task/execution roles and conditioned on ecs-tasks.amazonaws.com; S3 writes restricted to bootstrap manifests and task payload/launch objects; GetObject and DeleteObject restricted to */payload.json and */launch.json for signing, replay and cleanup; ListBucket is scoped to each payload bucket so absent launch records return NoSuchKey; MicroVM launch/state/sleep/wake/cleanup and image-capability actions (RunMicrovm/GetMicrovm/SuspendMicrovm/ResumeMicrovm/TerminateMicrovm/GetMicrovmImageVersion) are scoped to the single platform MicroVM image ARN plus a :* version-suffix sibling (every one of them authorizes against the image resource, not the per-session instance; no account-wide wildcard is used); lambda:PassNetworkConnector requires Resource:* because the action supports no resource-level permissions and the AWS-managed connectors live outside this account; iam:PassRole is scoped to the exact MicroVM execution role without iam:PassedToService (ADR-021 P2r2-F10); Agent Registry read scoped to the wired registry ARN, with a record/* suffix wildcard because record ids are server-assigned and unknown at synth (#246)', }, ], true); } diff --git a/cdk/src/constructs/task-status.ts b/cdk/src/constructs/task-status.ts index 3451155c1..960083691 100644 --- a/cdk/src/constructs/task-status.ts +++ b/cdk/src/constructs/task-status.ts @@ -112,7 +112,7 @@ export const VALID_TRANSITIONS: Readonly:*`], + }], true); } /** diff --git a/cdk/src/constructs/user-concurrency-table.ts b/cdk/src/constructs/user-concurrency-table.ts index 223fd3213..200f6a890 100644 --- a/cdk/src/constructs/user-concurrency-table.ts +++ b/cdk/src/constructs/user-concurrency-table.ts @@ -49,7 +49,8 @@ export interface UserConcurrencyTableProps { * * Schema: user_id (PK). Each item holds an atomic counter (active_count) * representing the number of currently running tasks for the user. - * The application layer uses conditional updates for increment/decrement. + * The application layer changes the counter and its reservation_version in + * transactions with per-task reservation markers. */ export class UserConcurrencyTable extends Construct { /** diff --git a/cdk/src/handlers/approve-task.ts b/cdk/src/handlers/approve-task.ts index 5c9a701c4..6331541c1 100644 --- a/cdk/src/handlers/approve-task.ts +++ b/cdk/src/handlers/approve-task.ts @@ -18,9 +18,10 @@ */ import { TransactionCanceledException } from '@aws-sdk/client-dynamodb'; -import { PutCommand, TransactWriteCommand, UpdateCommand } from '@aws-sdk/lib-dynamodb'; -import type { APIGatewayProxyEvent, APIGatewayProxyResult } from 'aws-lambda'; +import { TransactWriteCommand, UpdateCommand } from '@aws-sdk/lib-dynamodb'; +import type { APIGatewayProxyEvent, APIGatewayProxyResult, Context } from 'aws-lambda'; import { ulid } from 'ulid'; +import { recordDecisionPostCommit } from './shared/approval-decision'; import { VALID_APPROVAL_SCOPE_PREFIXES, parseApprovalScope } from './shared/approval-scope'; import { extractUserId } from './shared/gateway'; import { logger } from './shared/logger'; @@ -64,25 +65,38 @@ const AUDIT_EVENT_RETENTION_DAYS = Number(process.env.TASK_RETENTION_DAYS ?? '90 * @param event - API Gateway proxy event. * @returns API Gateway proxy result. */ -export async function handler(event: APIGatewayProxyEvent): Promise { +export async function handler( + event: APIGatewayProxyEvent, context?: Pick, +): Promise { + return recordApprovalForUser({ + userId: extractUserId(event), taskId: event.pathParameters?.task_id, body: event.body, + }, context); +} + +/** Shared decision path. Callers must authenticate and map the platform user first. */ +export async function recordApprovalForUser( + input: { userId: string | null; taskId?: string; body?: string | null; decisionSource?: string }, + context?: Pick, +): Promise { + const invocationStartedMs = Date.now(); const requestId = ulid(); try { // 1. Auth - const callerUserId = extractUserId(event); + const callerUserId = input.userId; if (!callerUserId) { return errorResponse(401, ErrorCode.UNAUTHORIZED, 'Missing or invalid authentication.', requestId); } // 2. Path + body - const taskId = event.pathParameters?.task_id; + const taskId = input.taskId; if (!taskId) { return errorResponse(400, ErrorCode.VALIDATION_ERROR, 'Missing task_id path parameter.', requestId); } let parsed: ApprovalRequest | null = null; try { - parsed = event.body ? JSON.parse(event.body) as ApprovalRequest : null; + parsed = input.body ? JSON.parse(input.body) as ApprovalRequest : null; } catch { return errorResponse(400, ErrorCode.VALIDATION_ERROR, 'Request body must be valid JSON.', requestId); } @@ -122,10 +136,10 @@ export async function handler(event: APIGatewayProxyEvent): Promise#MINUTE#` - // so the existing grantReadWriteData wiring carries forward; TTL - // reaps the counter after ~120s. + // 3. Per-user per-minute rate limit, shared with deny. The approvals-table + // partition key is RATE##APPROVE; the sort key is MINUTE#. + // TTL makes old counters eligible for eventual cleanup, not deletion at + // an exact time. Each minute uses its own counter regardless of that delay. const minuteBucket = formatMinuteBucket(new Date()); try { await ddb.send(new UpdateCommand({ @@ -168,9 +182,11 @@ export async function handler(event: APIGatewayProxyEvent): Promise :epoch)', ExpressionAttributeNames: { '#status': 'status', '#scope': 'scope', @@ -181,6 +197,8 @@ export async function handler(event: APIGatewayProxyEvent): Promise { /** * Non-mutating read to check if the user is at their concurrency limit. - * Used as a fast pre-check before expensive screening; the actual atomic - * increment happens in checkConcurrency() during transitionToSubmitted. + * Used as an advisory pre-check before expensive screening. The orchestrator + * reserves capacity atomically after submission, queuing if capacity fills. */ async function preCheckConcurrency(userId: string): Promise { try { @@ -745,7 +735,7 @@ async function preCheckConcurrency(userId: string): Promise { return activeCount < MAX_CONCURRENT; } catch (err: any) { // Only swallow DDB throttling errors — these are transient and the atomic - // check in transitionToSubmitted is the authoritative gate. + // orchestrator admission transaction is the authoritative gate. const throttleErrors = ['ProvisionedThroughputExceededException', 'RequestLimitExceeded', 'ThrottlingException']; if (throttleErrors.includes(err?.name)) { logger.warn('Pre-check concurrency throttled — allowing request to proceed', { @@ -765,66 +755,3 @@ async function preCheckConcurrency(userId: string): Promise { throw err; } } - -async function checkConcurrency(userId: string): Promise { - try { - await ddb.send(new UpdateCommand({ - TableName: CONCURRENCY_TABLE_NAME, - Key: { user_id: userId }, - UpdateExpression: 'SET active_count = if_not_exists(active_count, :zero) + :one, updated_at = :now', - ConditionExpression: 'attribute_not_exists(active_count) OR active_count < :max', - ExpressionAttributeValues: { - ':zero': 0, - ':one': 1, - ':max': MAX_CONCURRENT, - ':now': new Date().toISOString(), - }, - })); - return true; - } catch (err: any) { - if (err.name === 'ConditionalCheckFailedException') { - return false; - } - throw err; - } -} - -async function decrementConcurrency(userId: string): Promise { - const maxAttempts = 3; - for (let attempt = 0; attempt < maxAttempts; attempt++) { - try { - await ddb.send(new UpdateCommand({ - TableName: CONCURRENCY_TABLE_NAME, - Key: { user_id: userId }, - UpdateExpression: 'SET active_count = active_count - :one, updated_at = :now', - ConditionExpression: 'attribute_exists(active_count) AND active_count > :zero', - ExpressionAttributeValues: { - ':one': 1, - ':zero': 0, - ':now': new Date().toISOString(), - }, - })); - return; - } catch (err: any) { - if (err.name === 'ConditionalCheckFailedException') { - // Counter already at 0 or doesn't exist — nothing to roll back - return; - } - if (attempt < maxAttempts - 1) { - // Retry transient DDB errors (throttling, network) with backoff - await new Promise(resolve => setTimeout(resolve, 100 * Math.pow(2, attempt))); - continue; - } - logger.error('Failed to decrement concurrency counter after retries (leak possible)', { - user_id: userId, - attempts: maxAttempts, - error: err instanceof Error ? err.message : String(err), - metric_type: 'concurrency_counter_leak', - }); - throw new Error( - `Concurrency counter decrement failed for user ${userId} after ${maxAttempts} attempts. ` + - 'Manual intervention may be required to reset the counter.', - ); - } - } -} diff --git a/cdk/src/handlers/deny-task.ts b/cdk/src/handlers/deny-task.ts index 6758d9886..dcf248352 100644 --- a/cdk/src/handlers/deny-task.ts +++ b/cdk/src/handlers/deny-task.ts @@ -18,9 +18,10 @@ */ import { TransactionCanceledException } from '@aws-sdk/client-dynamodb'; -import { PutCommand, TransactWriteCommand, UpdateCommand } from '@aws-sdk/lib-dynamodb'; -import type { APIGatewayProxyEvent, APIGatewayProxyResult } from 'aws-lambda'; +import { TransactWriteCommand, UpdateCommand } from '@aws-sdk/lib-dynamodb'; +import type { APIGatewayProxyEvent, APIGatewayProxyResult, Context } from 'aws-lambda'; import { ulid } from 'ulid'; +import { recordDecisionPostCommit } from './shared/approval-decision'; import { scanDenyReason } from './shared/deny-reason-scanner'; import { extractUserId } from './shared/gateway'; import { logger } from './shared/logger'; @@ -62,25 +63,38 @@ const AUDIT_EVENT_RETENTION_DAYS = Number(process.env.TASK_RETENTION_DAYS ?? '90 * @param event - API Gateway proxy event. * @returns API Gateway proxy result. */ -export async function handler(event: APIGatewayProxyEvent): Promise { +export async function handler( + event: APIGatewayProxyEvent, context?: Pick, +): Promise { + return recordDenialForUser({ + userId: extractUserId(event), taskId: event.pathParameters?.task_id, body: event.body, + }, context); +} + +/** Shared decision path. Callers must authenticate and map the platform user first. */ +export async function recordDenialForUser( + input: { userId: string | null; taskId?: string; body?: string | null; decisionSource?: string }, + context?: Pick, +): Promise { + const invocationStartedMs = Date.now(); const requestId = ulid(); try { // 1. Auth - const callerUserId = extractUserId(event); + const callerUserId = input.userId; if (!callerUserId) { return errorResponse(401, ErrorCode.UNAUTHORIZED, 'Missing or invalid authentication.', requestId); } // 2. Path + body - const taskId = event.pathParameters?.task_id; + const taskId = input.taskId; if (!taskId) { return errorResponse(400, ErrorCode.VALIDATION_ERROR, 'Missing task_id path parameter.', requestId); } let parsed: DenyRequest | null = null; try { - parsed = event.body ? JSON.parse(event.body) as DenyRequest : null; + parsed = input.body ? JSON.parse(input.body) as DenyRequest : null; } catch { return errorResponse(400, ErrorCode.VALIDATION_ERROR, 'Request body must be valid JSON.', requestId); } @@ -149,9 +163,11 @@ export async function handler(event: APIGatewayProxyEvent): Promise :epoch)', ExpressionAttributeNames: { '#status': 'status' }, ExpressionAttributeValues: { ':denied': 'DENIED', @@ -159,6 +175,8 @@ export async function handler(event: APIGatewayProxyEvent): Promise> ...TERMINAL_EVENT_TYPES, 'pr_created', ]), - // Linear posts deterministic status comments on the platform tier - // (ADR-016: Linear is fully deterministic — the agent has no Linear MCP - // and posts nothing itself). Two events: + // Linear posts approval requests/outcomes and deterministic status comments + // on the platform tier (ADR-016: the agent has no Linear MCP). // * ``pr_created`` — the first-run "🔗 PR opened" courtesy comment (or, // for a comment-iteration, matures the threaded reply to "🔄 Working"). // This replaces the agent's old step-2 MCP save_comment. @@ -185,12 +162,12 @@ export const CHANNEL_DEFAULTS: Record> // OOM) before any PR, which is the motivating case: without a // platform-side comment the requester gets no completion signal at all. // - // Linear's `save_comment` doesn't support edit, so each is post-once (no - // live updates a la GitHub edit-in-place), idempotent across partial-batch - // retries via per-event markers. The start "🤖 Starting" comment is posted + // Approval messages and first-run status comments use delivery receipts; + // iteration replies are edited in place. The start "🤖 Starting" comment is posted // even earlier, at task-admission in the webhook processor (ADR-016 P4.5). linear: new Set([ ...TERMINAL_EVENT_TYPES, + ...APPROVAL_NOTIFICATION_EVENTS, 'pr_created', // Include task_timed_out so a Linear standalone iteration // that TIMES OUT still settles (its 👀→✅/❌ + terminal reply come through the @@ -354,7 +331,7 @@ export function parseStreamRecord(record: DynamoDBRecord): FanOutEvent | null { * ``trajectory_uploaded``, ``trace_truncated``. Only ``pr_created`` * is currently in any channel's default filter (§6.2 Slack + GitHub). */ -const ROUTABLE_MILESTONES: ReadonlySet = new Set(['pr_created']); +const ROUTABLE_MILESTONES: ReadonlySet = new Set(['pr_created', ...APPROVAL_NOTIFICATION_EVENTS]); /** * Unwrap ``agent_milestone`` events to their milestone name for @@ -481,6 +458,7 @@ async function loadTaskForComment(taskId: string): Promise { const result = await ddb.send(new GetCommand({ TableName: tableName, Key: { task_id: taskId }, + ConsistentRead: true, })); return (result.Item as TaskRecord | undefined) ?? null; } @@ -1177,6 +1155,45 @@ async function dispatchToLinear(event: FanOutEvent): Promise { return; } + const effectiveType = effectiveEventType(event); + if (isApprovalNotification(effectiveType)) { + const notification = await loadApprovalNotification(ddb, task, effectiveType, event.metadata ?? {}, 'linear'); + if (!notification) return; + const thread = { + workspaceId, + issueId, + taskId: task.task_id, + requestId: notification.requestId, + userId: notification.userId, + }; + const ctx = { linearWorkspaceId: workspaceId, registryTableName }; + const body = approvalNotificationMarkdown(notification); + if (effectiveType !== 'approval_requested') { + // Retention follows the saved closure even if Linear cannot receive its notice. + await closeLinearApprovalThread(ddb, process.env.TASK_APPROVALS_TABLE_NAME!, thread); + } + const result = effectiveType === 'approval_requested' + ? await postIdentifiedComment(ctx, { + id: await saveLinearApprovalThread(ddb, process.env.TASK_APPROVALS_TABLE_NAME!, thread), issueId, body, + }) + : await postIssueComment(ctx, issueId, body); + if (!result.ok) { + logger.warn('Linear approval notification failed', { + event: 'fanout.linear.approval_post_failed', + task_id: task.task_id, + request_id: notification.requestId, + retryable: result.retryable, + }); + if (result.retryable) throw new Error('Retryable Linear approval notification failure'); + return; + } + await markApprovalNotificationDelivered(ddb, notification); + logger.info('Linear approval notification delivered', { + event: 'fanout.linear.approval_dispatched', task_id: task.task_id, request_id: notification.requestId, + }); + return; // An approval message must never settle the task's final reply. + } + // Iteration-UX: this task is a comment-iteration when it carries a maturing // reply id (set at trigger time). For those, the progress + terminal status // lives in that ONE edited reply, NOT in fresh top-level comments. @@ -1264,8 +1281,8 @@ async function dispatchToLinear(event: FanOutEvent): Promise { return; // milestones never post the terminal status comment } - // Idempotency across partial-batch retries: Linear has no comment - // edit API, so a re-run of this dispatcher (e.g. a sibling channel's + // Idempotency across partial-batch retries: a re-run of this dispatcher + // (e.g. a sibling channel's // infra rejection pushed the whole stream record into // ``batchItemFailures``) would post a duplicate final-status comment. // The marker is persisted after the first successful post below. @@ -1943,8 +1960,7 @@ export async function routeEvent( * ``constructs/fanout-consumer.ts``) can honor partial-batch semantics. * Without a structured return, a single poisonous record would cause * Lambda to retry the **entire batch** from the stream checkpoint, - * replaying every sibling event and defeating the per-task ordering - * guarantee promised by ``ParallelizationFactor: 1`` upstream. + * replaying successful earlier events unnecessarily. * * Partial-failure surface (per-record try/catch below): * - ``routeEvent`` wraps each dispatcher in ``Promise.allSettled``, so @@ -1960,10 +1976,10 @@ export async function routeEvent( * future refactor (e.g. a stricter ``parseStreamRecord``) from * crashing the whole batch. * - * On any caught throw we push ``{ itemIdentifier: record.eventID }`` so - * Lambda retries ONLY that record, isolating the poison pill per - * design §6 + §8.9 expectations. Successful records are NOT in - * ``batchItemFailures`` and advance the stream checkpoint normally. + * Failures return the DynamoDB ``SequenceNumber``, not the opaque ``eventID``. + * Lambda checkpoints at the lowest failed sequence and retries that record and + * the following records. Successful later records can therefore be delivered + * again; channel delivery receipts still matter. * * Two review findings shaped this shape: the fanout handler used to return * ``void`` despite ``reportBatchItemFailures: true``, and a ``routeEvent`` @@ -1986,6 +2002,15 @@ export const handler = async ( let processed = 0; let dispatched = 0; let dropped = 0; + const retryRecord = (record: DynamoDBRecord): void => { + const sequence = record.dynamodb?.SequenceNumber; + if (!sequence) { + // A malformed failed record must not acknowledge the batch or return an + // invalid retry cursor. Reject the invocation so Lambda retries the batch. + throw new Error('Failed DynamoDB record is missing its sequence number'); + } + batchItemFailures.push({ itemIdentifier: sequence }); + }; // v1: no per-task override; every event uses the channel defaults. // Chunk K wires a DDB read here to load ``TaskRecord.notifications``. @@ -2028,21 +2053,13 @@ export const handler = async ( // attempt has a chance to succeed. Without this push, a transient // failure would be silently dropped — the regression that // motivated this fix. - if (outcome.infraRejections.length > 0 && record.eventID !== undefined) { - batchItemFailures.push({ itemIdentifier: record.eventID }); - } + if (outcome.infraRejections.length > 0) retryRecord(record); } catch (err) { // Poison-pill isolation: one record's unhandled throw must not // crash the batch. See the handler doc block for the full list of // paths that can reach here (notably AccessDeniedException from // ``resolveTokenSecretArn``). // - // ``eventID`` is the stream-record identifier Lambda uses for the - // retry cursor; on Kinesis-style event-source-mappings with - // ``reportBatchItemFailures: true`` the service retries all - // records at-or-after the lowest-sequence failure. Returning even - // one failed itemIdentifier is enough to preserve ordering across - // the whole batch for that task. const eventID = record.eventID; logger.warn('[fanout] record threw — flagging for partial-batch retry', { event: 'fanout.record.failed', @@ -2050,9 +2067,7 @@ export const handler = async ( error: err instanceof Error ? err.message : String(err), error_name: err instanceof Error ? err.name : undefined, }); - if (eventID !== undefined) { - batchItemFailures.push({ itemIdentifier: eventID }); - } + retryRecord(record); } } diff --git a/cdk/src/handlers/get-pending.ts b/cdk/src/handlers/get-pending.ts index 713a231dd..909c235f5 100644 --- a/cdk/src/handlers/get-pending.ts +++ b/cdk/src/handlers/get-pending.ts @@ -17,7 +17,7 @@ * SOFTWARE. */ -import { QueryCommand, UpdateCommand } from '@aws-sdk/lib-dynamodb'; +import { BatchGetCommand, QueryCommand, type QueryCommandInput, UpdateCommand } from '@aws-sdk/lib-dynamodb'; import type { APIGatewayProxyEvent, APIGatewayProxyResult } from 'aws-lambda'; import { ulid } from 'ulid'; import { extractUserId } from './shared/gateway'; @@ -29,12 +29,15 @@ import { makeDocClient } from './shared/ua'; const ddb = makeDocClient(); const TASK_APPROVALS_TABLE_NAME = process.env.TASK_APPROVALS_TABLE_NAME; -if (!TASK_APPROVALS_TABLE_NAME) { - throw new Error('get-pending handler requires TASK_APPROVALS_TABLE_NAME env var'); +const TASK_TABLE_NAME = process.env.TASK_TABLE_NAME ?? ''; +if (!TASK_APPROVALS_TABLE_NAME || !TASK_TABLE_NAME) { + throw new Error('get-pending handler requires TASK_APPROVALS_TABLE_NAME and TASK_TABLE_NAME env vars'); } const USER_STATUS_INDEX_NAME = process.env.USER_STATUS_INDEX_NAME ?? 'user_id-status-index'; const PENDING_RATE_LIMIT_PER_MINUTE = Number(process.env.PENDING_RATE_LIMIT_PER_MINUTE ?? '10'); const PENDING_LIST_LIMIT = 100; +const PENDING_READ_TIMEOUT_MS = 5_000; +const TASK_READ_ATTEMPTS = 3; /** * GET /v1/pending — List pending approvals owned by the caller (§7.7). @@ -100,20 +103,16 @@ export async function handler(event: APIGatewayProxyEvent): Promise>; - const pending: PendingApprovalSummary[] = items.map((row) => { + const { liveItems, omitted } = await readLivePendingRows(userId); + if (omitted > 0) { + logger.info('Omitted approvals whose tasks are no longer waiting for them', { + event: 'pending_approvals_inactive', + user_id: userId, + omitted, + request_id: requestId, + }); + } + const pending: PendingApprovalSummary[] = liveItems.map((row) => { const created_at = String(row.created_at ?? ''); const timeout_s = Number(row.timeout_s ?? 0); const expires_at = computeExpiresAt(created_at, timeout_s); @@ -148,6 +147,71 @@ export async function handler(event: APIGatewayProxyEvent): Promise>; + omitted: number; +}> { + const liveItems: Array> = []; + let omitted = 0; + let lastKey: QueryCommandInput['ExclusiveStartKey']; + const abortSignal = AbortSignal.timeout(PENDING_READ_TIMEOUT_MS); + do { + // The limit applies to displayed requests, not stale rows examined. Older + // cancellations can otherwise fill page one and hide every current request. + abortSignal.throwIfAborted(); + const result = await ddb.send(new QueryCommand({ + TableName: TASK_APPROVALS_TABLE_NAME, + IndexName: USER_STATUS_INDEX_NAME, + KeyConditionExpression: 'user_id = :user AND #status = :pending', + ExpressionAttributeNames: { '#status': 'status' }, + ExpressionAttributeValues: { ':user': userId, ':pending': 'PENDING' }, + Limit: PENDING_LIST_LIMIT - liveItems.length, + ...(lastKey && { ExclusiveStartKey: lastKey }), + }), { abortSignal }); + const items = (result.Items ?? []) as ReadonlyArray>; + const tasks = await readWaitingTasks(items, abortSignal); + for (const row of items) { + const task = tasks.get(String(row.task_id)); + // Bind the eventually consistent GSI row to explicit task/request state. + // This does not assess whether the requested action remains relevant. + if (task?.user_id === userId && task.status === 'AWAITING_APPROVAL' + && task.awaiting_approval_request_id === row.request_id) { + liveItems.push(row); + } else { + omitted++; + } + } + lastKey = result.LastEvaluatedKey; + } while (lastKey && liveItems.length < PENDING_LIST_LIMIT); + return { liveItems, omitted }; +} + +async function readWaitingTasks( + items: ReadonlyArray>, + abortSignal: AbortSignal, +): Promise>> { + const tasks = new Map>(); + let keys = [...new Set(items.map(row => String(row.task_id)))].map(task_id => ({ task_id })); + for (let attempt = 0; keys.length > 0 && attempt < TASK_READ_ATTEMPTS; attempt++) { + const response = await ddb.send(new BatchGetCommand({ + RequestItems: { + [TASK_TABLE_NAME]: { + Keys: keys, + ConsistentRead: true, + ProjectionExpression: 'task_id, user_id, #status, awaiting_approval_request_id', + ExpressionAttributeNames: { '#status': 'status' }, + }, + }, + }), { abortSignal }); + for (const task of response.Responses?.[TASK_TABLE_NAME] ?? []) { + tasks.set(String(task.task_id), task); + } + keys = (response.UnprocessedKeys?.[TASK_TABLE_NAME]?.Keys ?? []) as typeof keys; + } + if (keys.length > 0) throw new Error('Pending approval task reads remained unprocessed'); + return tasks; +} + function coerceSeverity(value: unknown): Severity { if (value === 'low' || value === 'medium' || value === 'high') { return value; @@ -160,7 +224,8 @@ function coerceStringList(value: unknown): readonly string[] { return value.filter((v): v is string => typeof v === 'string'); } -function computeExpiresAt(createdAt: string, timeoutS: number): string { +function computeExpiresAt(createdAt: string, timeoutS: number): string | null { + if (timeoutS === 0) return null; if (!createdAt || !Number.isFinite(timeoutS) || timeoutS <= 0) { return createdAt; } diff --git a/cdk/src/handlers/linear-webhook-processor.ts b/cdk/src/handlers/linear-webhook-processor.ts index 1f3eb36e1..0b1b179d2 100644 --- a/cdk/src/handlers/linear-webhook-processor.ts +++ b/cdk/src/handlers/linear-webhook-processor.ts @@ -26,6 +26,7 @@ import type { ScreeningConfig } from './shared/attachment-screening'; import { buildClarifyResumeDescription, isClarifyHold } from './shared/clarify-resume'; import { createTaskCore } from './shared/create-task-core'; import { renderMaturingReply } from './shared/iteration-reply'; +import { handleLinearApprovalReply } from './shared/linear-approval-reply'; import { cleanupPreScreenedAttachments, downloadScreenAndStoreLinearAttachments, LinearAttachmentError } from './shared/linear-attachments'; import { deleteComment, @@ -480,24 +481,6 @@ function patchChildOwnAttachments( }; } -/** - * Post a Linear comment + ❌ reaction without ever propagating an error. - * - * Phase 2.0b-O2: feedback is workspace-scoped — the resolver looks up - * the per-workspace OAuth token via `LinearWorkspaceRegistryTable` and - * issues a Bearer token. If the workspace isn't registered (drop-on-the-floor - * for unmapped orgs) the feedback path no-ops cleanly. - * - * Two failure modes handled here: - * - `LINEAR_WORKSPACE_REGISTRY_TABLE_NAME` env var unset (deploy misconfig) — - * skip with a clear diagnostic instead of letting the resolver fail - * per-call. - * - `reportIssueFailure` throws synchronously (today impossible thanks to the - * helper's internal `Promise.allSettled`, but a future refactor could - * break that contract). Catching here means a synchronous throw can't - * bubble up and fail the Lambda — which would trigger SQS retries on a - * poison message. - */ /** * Iteration-UX: post the IMMEDIATE threaded "👀 On it" reply under the trigger * comment, synchronously at trigger time. This is what kills the multi-minute @@ -535,6 +518,7 @@ async function postIterationAck( } } +/** Report workspace-scoped failure feedback best-effort; skip missing OAuth routing. */ async function safeReportIssueFailure( issueId: string, linearWorkspaceId: string | undefined, @@ -689,9 +673,8 @@ export async function handler(event: ProcessorEvent): Promise { return; } - // A Comment with an @bgagent mention on an orchestrated sub-issue - // re-iterates that sub-issue's PR (the reconciler then cascades the - // re-stack). Handled on a separate path from Issue → task creation. + // Comments route approval replies first, then @bgagent task/iteration + // requests. Issue events use the separate task-creation path below. if (payload.type === 'Comment') { await handleCommentTrigger(payload as LinearCommentEvent); return; @@ -1834,6 +1817,15 @@ async function handleNearMissMention(payload: LinearCommentEvent): Promise * a clean no-op (no failure comment — comments are conversational). */ async function handleCommentTrigger(payload: LinearCommentEvent): Promise { + if (process.env.TASK_APPROVALS_TABLE_NAME && WORKSPACE_REGISTRY_TABLE + && await handleLinearApprovalReply(payload, { + ddb, + approvalsTable: process.env.TASK_APPROVALS_TABLE_NAME, + taskTable: process.env.TASK_TABLE_NAME!, + registryTable: WORKSPACE_REGISTRY_TABLE, + lookupUser: lookupPlatformUser, + })) return; + // Orchestration must be enabled + a workspace token resolvable. if (!ORCHESTRATION_TABLE || !WORKSPACE_REGISTRY_TABLE) { return; @@ -3288,6 +3280,7 @@ async function lookupPlatformUser(workspaceId: string, userId: string): Promise< const result = await ddb.send(new GetCommand({ TableName: USER_MAPPING_TABLE, Key: { linear_identity: key }, + ConsistentRead: true, })); if (!result.Item || result.Item.status === 'pending') return null; return (result.Item.platform_user_id as string) ?? null; diff --git a/cdk/src/handlers/linear-webhook.ts b/cdk/src/handlers/linear-webhook.ts index 0cba78c70..a9878edf1 100644 --- a/cdk/src/handlers/linear-webhook.ts +++ b/cdk/src/handlers/linear-webhook.ts @@ -74,8 +74,8 @@ interface LinearWebhookEnvelope { * * Verifies the `Linear-Signature` HMAC over the raw body, rejects stale * `webhookTimestamp` values (replay protection), dedups on - * `(issue_id, action)` with a 60s TTL, and async-invokes the processor - * Lambda so we can ack within Linear's 5s timeout. + * `(data.id, action, webhookTimestamp)` with an eight-hour TTL, and invokes + * the processor asynchronously to acknowledge the delivery promptly. */ export async function handler(event: APIGatewayProxyEvent): Promise { try { diff --git a/cdk/src/handlers/orchestrate-task.ts b/cdk/src/handlers/orchestrate-task.ts index 3f448c133..2d8ed913f 100644 --- a/cdk/src/handlers/orchestrate-task.ts +++ b/cdk/src/handlers/orchestrate-task.ts @@ -18,11 +18,17 @@ */ import { withDurableExecution, type DurableExecutionHandler } from '@aws/durable-execution-sdk-js'; +import { CONTINUATION_RETRY_POLL_SECONDS, CONTINUATION_TRANSITION_POLL_SECONDS } from './shared/microvm-continuation-timing'; import { TaskStatus, TERMINAL_STATUSES } from '../constructs/task-status'; import { resolveComputeStrategy } from './shared/compute-strategy'; +import { MicrovmStartUncertainError } from './shared/error-classifier'; import { reportIssueFailure as reportJiraIssueFailure } from './shared/jira-feedback'; import { reportIssueFailure } from './shared/linear-feedback'; import { logger } from './shared/logger'; +import { runMicrovmContinuation } from './shared/microvm-continuation-runner'; +import { saveContinuationLaunch } from './shared/microvm-continuation-storage'; +import { stopMicrovmWithDiagnostics } from './shared/microvm-supervisor'; +import { pollMicrovmTask } from './shared/microvm-task-poll'; import { admissionControl, buildComputeMetadata, @@ -35,7 +41,6 @@ import { loadTask, pollTaskStatus, queueTask, - reconcileMicrovmSubstrateState, transitionTask, type PollState, } from './shared/orchestrator'; @@ -43,6 +48,7 @@ import { runPreflightChecks } from './shared/preflight'; import { isAutoRetried, startSessionWithRetry } from './shared/session-start-retry'; import { deleteEcsPayload } from './shared/strategies/ecs-strategy'; import { deleteMicrovmPayload } from './shared/strategies/lambda-microvm-strategy'; +import { releaseTaskSlot } from './shared/task-concurrency'; import type { TaskRecord } from './shared/types'; import { workflowIsReadOnly, workflowRequiresRepo } from './shared/workflows'; @@ -56,6 +62,8 @@ interface OrchestrateTaskEvent { * durable-execution idempotency) can mistake it for a replay. */ readonly queue_pickup_id?: string; + readonly continuation_request_id?: string; + readonly continuation_attempt_id?: string; } const MAX_POLL_ATTEMPTS = 1020; // ~8.5h at 30s intervals @@ -66,6 +74,13 @@ const MAX_CONSECUTIVE_ECS_COMPLETED_POLLS = 5; const DEFAULT_POLL_INTERVAL_SECONDS = 30; const durableHandler: DurableExecutionHandler = async (event, context) => { + if (event.continuation_request_id !== undefined || event.continuation_attempt_id !== undefined) { + return runMicrovmContinuation({ + task_id: event.task_id, + continuation_request_id: event.continuation_request_id ?? '', + continuation_attempt_id: event.continuation_attempt_id ?? '', + }, context); + } const { task_id: taskId } = event; // Step 1: Load task record @@ -95,7 +110,7 @@ const durableHandler: DurableExecutionHandler = asyn // up, flips QUEUED -> SUBMITTED, and re-invokes this orchestrator. const admitted = await context.step('admission-control', async () => { // Re-read status to detect external cancellation between steps - const current = await loadTask(taskId); + const current = await loadTask(taskId, true); if (TERMINAL_STATUSES.includes(current.status)) { return false; } @@ -131,13 +146,14 @@ const durableHandler: DurableExecutionHandler = asyn }); if (!admitted) { + await context.step('release-before-admission', () => releaseTaskSlot(taskId, task.user_id)); return; } // Step 2b: Pre-flight checks — verify external dependencies before consuming AgentCore runtime const preflightPassed = await context.step('pre-flight', async () => { try { - const current = await loadTask(taskId); + const current = await loadTask(taskId, true); if (TERMINAL_STATUSES.includes(current.status)) { return false; } @@ -166,6 +182,7 @@ const durableHandler: DurableExecutionHandler = asyn }); if (!preflightPassed) { + await context.step('release-before-work', () => releaseTaskSlot(taskId, task.user_id)); return; } @@ -184,11 +201,15 @@ const durableHandler: DurableExecutionHandler = asyn // Returns the full SessionHandle (serializable) so ECS polling can use it in step 5. const sessionHandle = await context.step('start-session', async () => { let autoRetried = false; + let failureStatus: TaskRecord['status'] = TaskStatus.HYDRATING; // Hoisted out of the `try` so the catch can reap a MicroVM that STARTED but - // whose handle never made it into DynamoDB — see the catch block. + // whose registration did not complete — see the catch block. let strategy: ReturnType | undefined; let startedHandle: Awaited>['handle'] | undefined; try { + if (blueprintConfig.compute_type === 'lambda-microvm') { + await saveContinuationLaunch(taskId, task.user_id, payload, blueprintConfig); + } strategy = resolveComputeStrategy(blueprintConfig); const startInput = { taskId, @@ -223,6 +244,15 @@ const durableHandler: DurableExecutionHandler = asyn // can load the handle they resume from — ADR-021 sub-decision 2). const computeMetadata = buildComputeMetadata(handle); + const current = handle.strategyType === 'lambda-microvm' ? await loadTask(taskId, true) : undefined; + if (current && TERMINAL_STATUSES.includes(current.status)) { + await strategy.stopSession(handle); + return null; + } + const alreadyRegistered = current?.session_id === handle.sessionId + && (current.status === TaskStatus.RUNNING || current.status === TaskStatus.AWAITING_APPROVAL); + if (alreadyRegistered) return handle; + await transitionTask(taskId, TaskStatus.HYDRATING, TaskStatus.RUNNING, { session_id: handle.sessionId, started_at: new Date().toISOString(), @@ -230,10 +260,17 @@ const durableHandler: DurableExecutionHandler = asyn compute_metadata: computeMetadata, ...(handle.strategyType === 'agentcore' && { agent_runtime_arn: handle.runtimeArn }), }); - await emitTaskEvent(taskId, 'session_started', { - session_id: handle.sessionId, - strategy_type: handle.strategyType, - }, correlation); + try { + await emitTaskEvent(taskId, 'session_started', { + session_id: handle.sessionId, + strategy_type: handle.strategyType, + }, correlation); + } catch (emitErr) { + if (handle.strategyType !== 'lambda-microvm') throw emitErr; + log.warn('session_started event failed after the MicroVM was registered', { + session_id: handle.sessionId, error: String(emitErr), + }); + } log.info('Session started', { session_id: handle.sessionId, @@ -242,15 +279,39 @@ const durableHandler: DurableExecutionHandler = asyn return handle; } catch (err) { - // ORPHAN REAP (ADR-021). `RunMicrovm` may have already succeeded and left a - // MicroVM RUNNING — the throw could have come from `buildComputeMetadata`, - // the `transitionTask` write, or the `session_started` emit. Nothing - // self-terminates on this substrate (live-verified: a MicroVM with no - // working hook reached RUNNING in 12 s and stayed RUNNING with no - // stateReason), and the handle only ever existed in this Lambda's memory — - // once we throw, no poll and no finalize step will ever see it. So the VM - // would bill until `maximumDurationInSeconds` (8 h) expired while also - // holding account memory quota that gates admission for everyone else. + if (blueprintConfig.compute_type === 'lambda-microvm') { + try { + const current = await loadTask(taskId, true); + // A lost DynamoDB update response does not undo its committed result. + if (startedHandle && current.session_id === startedHandle.sessionId + && (current.status === TaskStatus.RUNNING || current.status === TaskStatus.AWAITING_APPROVAL)) { + return startedHandle; + } + if (TERMINAL_STATUSES.includes(current.status)) { + if (startedHandle && strategy) await strategy.stopSession(startedHandle); + if (!startedHandle && err instanceof MicrovmStartUncertainError) { + log.error('Task became terminal while its MicroVM start outcome is unknown', { + task_id: taskId, client_token: taskId, task_status: current.status, + }); + try { + await emitTaskEvent(taskId, 'microvm_start_outcome_unknown', { + client_token: taskId, task_status: current.status, + }, correlation); + } catch (eventErr) { + log.warn('Could not record the unknown MicroVM start event', { error: String(eventErr) }); + } + } + return null; + } + failureStatus = current.status; + } catch (readErr) { + log.warn('Could not reconcile MicroVM registration after start failure', { error: String(readErr) }); + } + } + // Registration failed and a strong read could not establish a committed + // RUNNING/approval-wait task. Reap the known computer. The start receipt + // retains its ID for recovery if cleanup itself fails; the service's + // eight-hour maximum duration remains the final lifetime bound. // // Best-effort in the strongest sense: `stopSession` is internally // non-throwing for this backend, and the extra try/catch guarantees that @@ -290,10 +351,39 @@ const durableHandler: DurableExecutionHandler = asyn // Without this, a double-transient failure was told "reply to retry" instead // of "I already retried" — the exact confusion the marker exists to prevent. const retriedNote = (autoRetried || isAutoRetried(err)) ? ' [auto-retried]' : ''; - await failTask(taskId, TaskStatus.HYDRATING, `Session start failed: ${String(err)}${retriedNote}`, task.user_id, true, task.repo); + const detail = err instanceof MicrovmStartUncertainError + ? `MICROVM_START_OUTCOME_UNKNOWN: ${String(err)}` + : String(err); + const errorMessage = `Session start failed: ${detail}${retriedNote}`; + if (blueprintConfig.compute_type === 'lambda-microvm') { + // The durable finalization step below owns the terminal event and slot + // release. Replaying this step after FAILED must not release it twice. + try { + await transitionTask(taskId, failureStatus, TaskStatus.FAILED, { + completed_at: new Date().toISOString(), error_message: errorMessage, + }); + } catch (transitionErr) { + const current = await loadTask(taskId, true); + if (!TERMINAL_STATUSES.includes(current.status)) throw transitionErr; + } + return null; + } + await failTask(taskId, failureStatus, errorMessage, task.user_id, true, task.repo); throw err; } - }); + }, blueprintConfig.compute_type === 'lambda-microvm' + ? { retryStrategy: () => ({ shouldRetry: false }) } + : undefined); + + if (!sessionHandle) { + // The task ended before registration, including a persisted start failure. + // Reach normal finalization without starting another VM. + await context.step('finalize-before-session', async () => { + await finalizeTask(taskId, { attempts: 0 }, task.user_id); + await deleteMicrovmPayload(taskId); + }); + return; + } // Resolve the compute strategy once and reuse it across poll iterations // instead of constructing a new instance on every cycle. @@ -301,10 +391,7 @@ const durableHandler: DurableExecutionHandler = asyn ? resolveComputeStrategy(blueprintConfig) : undefined; - // Kept as a SEPARATE local rather than widening `computeStrategy`'s condition: - // the ECS cross-check below is gated on `computeStrategy` truthiness, so - // reusing that local for lambda-microvm would route MicroVM polls through the - // ECS exit-code/patience logic. Two locals keep the ECS path byte-identical. + // ECS crash checks and MicroVM lifecycle supervision have separate policies. const microvmStrategy = blueprintConfig.compute_type === 'lambda-microvm' ? resolveComputeStrategy(blueprintConfig) : undefined; @@ -314,18 +401,26 @@ const durableHandler: DurableExecutionHandler = asyn // While RUNNING, the runtime updates `agent_heartbeat_at`; if that timestamp // goes stale, `pollTaskStatus` sets `sessionUnhealthy` so we fail fast instead // of waiting the full MAX_POLL_ATTEMPTS window (~8.5h) after a silent crash. - // HYDRATING without transition to RUNNING is still bounded by MAX_NON_RUNNING_POLLS (~5min). + // ECS/AgentCore retain their poll-count startup/total bounds. MicroVM uses + // persisted wall-clock deadlines so faster transition polls cannot shorten a session. const finalPollState = await context.waitForCondition( 'await-agent-completion', async (state) => { + if (microvmStrategy && sessionHandle.strategyType === 'lambda-microvm') { + return pollMicrovmTask({ + taskId, + userId: task.user_id, + handle: sessionHandle, + strategy: microvmStrategy, + pollIntervalMs: blueprintConfig.poll_interval_ms ?? DEFAULT_POLL_INTERVAL_SECONDS * 1000, + suspendEnabled: process.env.MICROVM_APPROVAL_SUSPEND_ENABLED === 'true', + emitEvent: (type, metadata, options) => emitTaskEvent(taskId, type, metadata, correlation, options), + }, state); + } const ddbState = await pollTaskStatus(taskId, state, blueprintConfig.compute_type); let consecutiveEcsPollFailures = 0; let consecutiveEcsCompletedPolls = 0; - // Carried forward by default: an unrelated poll (or a MicroVM poll that - // threw) must not silently re-arm the once-per-episode anomaly event. - let microvmSuspendAnomalyReported = state.microvmSuspendAnomalyReported ?? false; - // ECS compute-level crash detection: if DDB is not terminal, check ECS task status if ( ddbState.lastStatus && @@ -374,62 +469,31 @@ const durableHandler: DurableExecutionHandler = asyn } } - // Lambda MicroVMs substrate cross-check (ADR-021 sub-decision 1). Same - // division of labour as the ECS block above — the strategy reports raw - // substrate state and `reconcileMicrovmSubstrateState` interprets it - // against the DDB status — but the rules differ: a `suspended` VM is - // healthy during an approval wait, an anomaly (not a failure) otherwise, - // and a terminal VM with a non-terminal task row is a substrate failure. - if ( - ddbState.lastStatus - && !TERMINAL_STATUSES.includes(ddbState.lastStatus) - && microvmStrategy - && sessionHandle.strategyType === 'lambda-microvm' - ) { - try { - const substrateStatus = await microvmStrategy.pollSession(sessionHandle); - const { taskFailed, suspendAnomalyReported } = await reconcileMicrovmSubstrateState({ - taskId, - ddbStatus: ddbState.lastStatus, - substrate: substrateStatus, - microvmId: sessionHandle.microvmId, - userId: task.user_id, - correlation, - log, - repo: task.repo, - // Threaded so `microvm_suspend_anomaly` is emitted once per anomaly - // EPISODE rather than on every ~30 s poll; a non-anomalous observation - // re-arms it (see reconcileMicrovmSubstrateState). - suspendAnomalyReported: microvmSuspendAnomalyReported, - }); - microvmSuspendAnomalyReported = suspendAnomalyReported; - if (taskFailed) { - return { attempts: ddbState.attempts, lastStatus: TaskStatus.FAILED }; - } - } catch (err) { - // Non-fatal: a GetMicrovm hiccup must not abort the durable poll step. - // The task stays bounded by MAX_POLL_ATTEMPTS (~8.5 h) and, once the - // agent is RUNNING, by its own terminal write. A repeated-failure - // escalation counter (the ECS `MAX_CONSECUTIVE_ECS_POLL_FAILURES` - // analogue) is deliberately deferred — it would add PollState fields - // that P3's suspend policy will need to reshape anyway. - log.warn('MicroVM pollSession check failed (non-fatal)', { - error: err instanceof Error ? err.message : String(err), - }); - } - } - - return { ...ddbState, consecutiveEcsPollFailures, consecutiveEcsCompletedPolls, microvmSuspendAnomalyReported }; + return { ...ddbState, consecutiveEcsPollFailures, consecutiveEcsCompletedPolls }; }, { initialState: { attempts: 0 }, waitStrategy: (state: PollState) => { + if (state.microvmParked || state.microvmOwnershipLost) return { shouldContinue: false }; + if (state.microvmRetiring) { + return { + shouldContinue: true, + delay: { seconds: state.microvmRetirementError ? CONTINUATION_RETRY_POLL_SECONDS : CONTINUATION_TRANSITION_POLL_SECONDS }, + }; + } if (state.lastStatus && TERMINAL_STATUSES.includes(state.lastStatus)) { return { shouldContinue: false }; } if (state.sessionUnhealthy) { return { shouldContinue: false }; } + if (state.microvmSupervisor) { + if (state.microvmFailureReason || state.microvmOwnershipLost) return { shouldContinue: false }; + return { + shouldContinue: true, + delay: { seconds: Math.max(1, Math.ceil(state.microvmSupervisor.nextPollInMs / 1000)) }, + }; + } if (state.attempts >= MAX_POLL_ATTEMPTS) { return { shouldContinue: false }; } @@ -449,39 +513,39 @@ const durableHandler: DurableExecutionHandler = asyn // Step 6: Finalize — update terminal status, emit events, release concurrency await context.step('finalize', async () => { - await finalizeTask(taskId, finalPollState, task.user_id); - // The task is terminal — the substrate has long since read its payload, so - // delete the ephemeral S3 payload object now. Best-effort (both deleters - // swallow errors) and a no-op for AgentCore tasks / deployments without a - // payload bucket; the bucket's 1-day lifecycle rule is the backstop if this - // delete or the whole step never runs. - // - // Both payload-carrying backends get this, and the MicroVM one is NOT - // optional polish: its execution role holds `grantRead` on the WHOLE payload - // bucket (the guest must read its object before any tenant identity exists), - // keys are `/payload.json`, and the guest runs untrusted repo code — - // so a TTL-only reaper left every finished task's hydrated prompt readable by - // any concurrently running MicroVM for up to ~24 h. See - // `deleteMicrovmPayload`. + if (finalPollState.microvmParked) { + await emitTaskEvent(taskId, 'continuation_parked', { + microvm_id: sessionHandle.sessionId, + detail: 'Your approval request is still available. The saved task will continue on another worker after your answer.', + }, correlation); + await deleteMicrovmPayload(taskId); + return; + } + let finalized = false; + try { + if (!finalPollState.microvmOwnershipLost) { + finalized = (await finalizeTask(taskId, finalPollState, task.user_id)) !== false; + } + } finally { + // Even a database finalization failure must not lose cleanup of this handle. + // A replacement worker, if any, is never followed or terminated here. + if (microvmStrategy && sessionHandle.strategyType === 'lambda-microvm') { + await stopMicrovmWithDiagnostics({ + taskId, + handle: sessionHandle, + strategy: microvmStrategy, + emitEvent: (type, metadata, options) => emitTaskEvent(taskId, type, metadata, correlation, options), + }); + } + } + if (!finalized) return; + // Delete task instructions and their saved signed download capability after + // finalization. Shared manifests remain; bucket lifecycle is a backstop. if (blueprintConfig.compute_type === 'ecs') { await deleteEcsPayload(taskId); } else if (blueprintConfig.compute_type === 'lambda-microvm') { await deleteMicrovmPayload(taskId); } - // ADR-021: "When the orchestrator finalizes a `lambda-microvm` task, the - // orchestrator shall call terminate-microvm (termination shall not rely on - // any substrate timeout)." Without this the VM lingers until - // `maximumDurationInSeconds` (8 h) expires — with `idlePolicy` omitted there - // is no tighter substrate bound — so we would keep paying for a full 8-hour - // reservation after every task, and every SUSPENDED/RUNNING VM keeps counting - // against the account memory quota that gates admission. - // - // `stopSession` is internally best-effort (it swallows and level-differentiates - // every failure), so this cannot fail the finalize step or strand the task in - // a non-terminal state. - if (microvmStrategy && sessionHandle.strategyType === 'lambda-microvm') { - await microvmStrategy.stopSession(sessionHandle); - } }); }; diff --git a/cdk/src/handlers/reconcile-admission-queue.ts b/cdk/src/handlers/reconcile-admission-queue.ts index 0b08c48bf..849ca9785 100644 --- a/cdk/src/handlers/reconcile-admission-queue.ts +++ b/cdk/src/handlers/reconcile-admission-queue.ts @@ -213,7 +213,9 @@ async function requeueAfterInvokeFailure(taskId: string): Promise { TableName: TASK_TABLE, Key: { task_id: { S: taskId } }, UpdateExpression: 'SET #s = :queued, updated_at = :now, status_created_at = :sca', - ConditionExpression: '#s = :submitted', + // A lost invoke response may hide a successful admission. Never put a + // task that now holds capacity back into the no-reservation queue. + ConditionExpression: '#s = :submitted AND attribute_not_exists(concurrency_slot)', ExpressionAttributeNames: { '#s': 'status' }, ExpressionAttributeValues: { ':queued': { S: 'QUEUED' }, diff --git a/cdk/src/handlers/reconcile-concurrency.ts b/cdk/src/handlers/reconcile-concurrency.ts index 008f0e8a2..a8c97daef 100644 --- a/cdk/src/handlers/reconcile-concurrency.ts +++ b/cdk/src/handlers/reconcile-concurrency.ts @@ -17,108 +17,151 @@ * SOFTWARE. */ -import { DynamoDBClient, ScanCommand, QueryCommand, UpdateItemCommand } from '@aws-sdk/client-dynamodb'; +import { randomUUID } from 'node:crypto'; +import { GetCommand, ScanCommand, type ScanCommandOutput, UpdateCommand } from '@aws-sdk/lib-dynamodb'; +import { ACTIVE_STATUSES, TERMINAL_STATUSES } from '../constructs/task-status'; import { logger } from './shared/logger'; -import { makeClient } from './shared/ua'; +import { workerLeaseKey } from './shared/microvm-continuation-types'; +import { releaseTaskSlot, type ReservationTask } from './shared/task-concurrency'; +import { makeDocClient } from './shared/ua'; -const ddb = makeClient(DynamoDBClient); +const ddb = makeDocClient(); const TASK_TABLE = process.env.TASK_TABLE_NAME!; -const CONCURRENCY_TABLE = process.env.USER_CONCURRENCY_TABLE_NAME!; +const COUNTER_TABLE = process.env.USER_CONCURRENCY_TABLE_NAME!; -/** - * Count actual active tasks for a user by querying the UserStatusIndex GSI. - */ -async function countActiveTasks(userId: string): Promise { - let count = 0; - let lastKey: Record | undefined; - - do { - const resp = await ddb.send(new QueryCommand({ - TableName: TASK_TABLE, - IndexName: 'UserStatusIndex', - KeyConditionExpression: 'user_id = :uid', - FilterExpression: '#s IN (:s1, :s2, :s3, :s4)', - ExpressionAttributeNames: { '#s': 'status' }, - ExpressionAttributeValues: { - ':uid': { S: userId }, - ':s1': { S: 'SUBMITTED' }, - ':s2': { S: 'HYDRATING' }, - ':s3': { S: 'RUNNING' }, - ':s4': { S: 'FINALIZING' }, - }, - Select: 'COUNT', - ExclusiveStartKey: lastKey, - })); - count += resp.Count ?? 0; - lastKey = resp.LastEvaluatedKey; - } while (lastKey); +interface CounterSnapshot { + readonly user_id: string; + readonly active_count?: number; + readonly reservation_version?: string; +} - return count; +interface Reservations { + held: number; + ambiguous: boolean; + terminal: string[]; } /** - * Scheduled handler: scan the concurrency table and reconcile each user's - * active_count against actual active tasks in the task table. + * Read counters BEFORE scanning reservations. Every reservation mutation also + * changes its counter revision; a conditional repair then detects changes + * anywhere during the scan, including an increment/decrement with no net change. + * Strong base-table reads avoid the GSI lag that can hide a newly held slot. */ export async function handler(): Promise { logger.info('Concurrency reconciler started'); - let corrected = 0; - let scanned = 0; - let errors = 0; + const counters = new Map(); let lastKey: Record | undefined; - do { - const scanResp = await ddb.send(new ScanCommand({ - TableName: CONCURRENCY_TABLE, + const page: ScanCommandOutput = await ddb.send(new ScanCommand({ + TableName: COUNTER_TABLE, + ConsistentRead: true, + ProjectionExpression: 'user_id, active_count, reservation_version', ExclusiveStartKey: lastKey, })); + for (const row of page.Items ?? []) { + if (typeof row.user_id === 'string') counters.set(row.user_id, row as CounterSnapshot); + } + lastKey = page.LastEvaluatedKey; + } while (lastKey); - for (const rawItem of scanResp.Items ?? []) { - const userId = rawItem.user_id?.S; - const storedCount = Number(rawItem.active_count?.N ?? '0'); - if (!userId) continue; - scanned++; - - try { - const actualCount = await countActiveTasks(userId); + const reservations = new Map(); + do { + const page: ScanCommandOutput = await ddb.send(new ScanCommand({ + TableName: TASK_TABLE, + ConsistentRead: true, + ProjectionExpression: 'task_id, user_id, #status, concurrency_slot, continuation, microvm_start, session_id, repo', + ExpressionAttributeNames: { '#status': 'status' }, + ExclusiveStartKey: lastKey, + })); + for (const row of page.Items ?? []) { + const task = row as ReservationTask; + if (!task.task_id || !task.user_id) continue; + const active = ACTIVE_STATUSES.some(status => status === task.status); + if (!task.concurrency_slot && !active) continue; + const owned = reservations.get(task.user_id) ?? { held: 0, ambiguous: false, terminal: [] }; + let parked = false; + if (task.status === 'AWAITING_APPROVAL' && task.concurrency_slot?.state === 'released' + && row.continuation?.state === 'PARKED') { + const saved = await ddb.send(new GetCommand({ + TableName: TASK_TABLE, Key: workerLeaseKey(task.task_id), ConsistentRead: true, + })); + parked = saved.Item?.lease_state === 'PARKED' + && saved.Item.lease_attempt_id === row.microvm_start?.clientToken + && saved.Item.lease_user_id === task.user_id && saved.Item.lease_repo === (row.repo ?? '') + && saved.Item.lease_microvm_id === row.session_id; + } + if (task.concurrency_slot?.state === 'held') { + // Terminal tasks still own their seat until release commits. Repair + // that total first, then release terminal reservations through the + // same transaction used by the orchestrator and stranded-task cleaner. + owned.held++; + if (TERMINAL_STATUSES.some(status => status === task.status)) owned.terminal.push(task.task_id); + } else if ((active && !parked) || (task.concurrency_slot && task.concurrency_slot.state !== 'released')) { + // Older active tasks and not-yet-admitted SUBMITTED tasks cannot be + // distinguished by status. Wait for them to settle; never guess a count. + owned.ambiguous = true; + } + reservations.set(task.user_id, owned); + } + lastKey = page.LastEvaluatedKey; + } while (lastKey); - if (storedCount !== actualCount) { - logger.info('Drift detected', { userId, storedCount, actualCount }); - try { - await ddb.send(new UpdateItemCommand({ - TableName: CONCURRENCY_TABLE, - Key: { user_id: { S: userId } }, - UpdateExpression: 'SET active_count = :count, updated_at = :now', - ConditionExpression: 'active_count = :stored', - ExpressionAttributeValues: { - ':count': { N: String(actualCount) }, - ':now': { S: new Date().toISOString() }, - ':stored': { N: String(storedCount) }, - }, - })); - corrected++; - } catch (updateErr: unknown) { - if (updateErr && typeof updateErr === 'object' && 'name' in updateErr && updateErr.name === 'ConditionalCheckFailedException') { - logger.info('Concurrent update detected, skipping', { userId }); - } else { - throw updateErr; - } + let corrected = 0; + let errors = 0; + const users = new Set([...counters.keys(), ...reservations.keys()]); + for (const userId of users) { + const snapshot = counters.get(userId); + const owned = reservations.get(userId) ?? { held: 0, ambiguous: false, terminal: [] }; + try { + const stored = snapshot?.active_count ?? 0; + if (!Number.isSafeInteger(stored)) throw new Error('Concurrency counter is not an integer'); + if (owned.ambiguous) { + logger.warn('Skipping capacity repair while task reservation ownership is ambiguous', { + user_id: userId, error_id: 'CONCURRENCY_RESERVATION_UNKNOWN', + }); + } else if (stored !== owned.held) { + const conditions: string[] = []; + const values: Record = { + ':count': owned.held, ':now': new Date().toISOString(), ':version': randomUUID(), + }; + if (!snapshot) { + conditions.push('attribute_not_exists(user_id)'); + } else { + conditions.push('attribute_exists(user_id)'); + if (snapshot.active_count === undefined) {conditions.push('attribute_not_exists(active_count)');} else { + conditions.push('active_count = :stored'); + values[':stored'] = stored; + } + if (snapshot.reservation_version === undefined) {conditions.push('attribute_not_exists(reservation_version)');} else { + conditions.push('reservation_version = :observed'); + values[':observed'] = snapshot.reservation_version; } } - } catch (err: unknown) { - errors++; - logger.warn('Per-user reconciliation failed, continuing', { - userId, - error: err instanceof Error ? err.message : String(err), - }); + try { + await ddb.send(new UpdateCommand({ + TableName: COUNTER_TABLE, + Key: { user_id: userId }, + UpdateExpression: 'SET active_count = :count, updated_at = :now, reservation_version = :version', + ConditionExpression: conditions.join(' AND '), + ExpressionAttributeValues: values, + })); + corrected++; + logger.info('Corrected capacity counter from saved reservations', { + user_id: userId, stored_count: stored, reservation_count: owned.held, + }); + } catch (error) { + if ((error as { name?: string })?.name !== 'ConditionalCheckFailedException') throw error; + logger.info('Capacity changed during reconciliation; skipped stale repair', { user_id: userId }); + } } + for (const taskId of owned.terminal) await releaseTaskSlot(taskId, userId); + } catch (error) { + errors++; + logger.warn('Per-user capacity reconciliation failed, continuing', { user_id: userId, error: String(error) }); } - - lastKey = scanResp.LastEvaluatedKey; - } while (lastKey); - - if (errors === scanned && scanned > 0) { - logger.error('All users failed reconciliation — possible systemic issue', { scanned, errors }); } - logger.info('Concurrency reconciler finished', { scanned, corrected, errors }); + if (errors === users.size && users.size > 0) { + logger.error('All users failed reconciliation — possible systemic issue', { scanned: users.size, errors }); + } + logger.info('Concurrency reconciler finished', { scanned: users.size, corrected, errors }); } diff --git a/cdk/src/handlers/reconcile-microvm-continuations.ts b/cdk/src/handlers/reconcile-microvm-continuations.ts new file mode 100644 index 000000000..7939eba7b --- /dev/null +++ b/cdk/src/handlers/reconcile-microvm-continuations.ts @@ -0,0 +1,208 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { GetCommand, ScanCommand, UpdateCommand, type ScanCommandOutput } from '@aws-sdk/lib-dynamodb'; +import type { Context } from 'aws-lambda'; +import { TaskStatus, TERMINAL_STATUSES } from '../constructs/task-status'; +import { closeTaskApprovals } from './shared/close-task-approvals'; +import { logger } from './shared/logger'; +import { dispatchMicrovmContinuation } from './shared/microvm-continuation-dispatch'; +import { retireCheckpointedMicrovm, type ContinuableTask } from './shared/microvm-continuation-retirement'; +import { deleteClosedTaskContinuations } from './shared/microvm-continuation-storage'; +import { CONTINUATION_IO_TIMEOUT_MS, CONTINUATION_RETIREMENT_TIMEOUT_MS } from './shared/microvm-continuation-timing'; +import { workerLeaseKey, type MicrovmHandle } from './shared/microvm-continuation-types'; +import { microvmErrorIdentity } from './shared/microvm-control'; +import { LambdaMicrovmComputeStrategy, MICROVM_MAX_DURATION_SECONDS } from './shared/strategies/lambda-microvm-strategy'; +import { releaseTaskSlot } from './shared/task-concurrency'; +import { makeDocClient } from './shared/ua'; + +const ddb = makeDocClient(); +const strategy = new LambdaMicrovmComputeStrategy(); +const TABLE = process.env.TASK_TABLE_NAME!; +const CURSOR_KEY = { task_id: 'continuation-manager#cursor' }; +const CURSOR_WRITE_TIMEOUT_MS = 5000; +const MIN_REMAINING_MS = 45_000; +const BATCH_SIZE = 4; + +/** Every mutation below rechecks the current owner/attempt; scan rows are only hints. */ +export async function reconcileMicrovmContinuation(task: ContinuableTask): Promise { + const options = { abortSignal: AbortSignal.timeout(CONTINUATION_RETIREMENT_TIMEOUT_MS) }; + // The scan can predate a replacement or cancellation. Resolve physical identity + // again before issuing a stop against a terminal task. + if (TERMINAL_STATUSES.includes(task.status)) { + const latest = await ddb.send(new GetCommand({ + TableName: TABLE, Key: { task_id: task.task_id }, ConsistentRead: true, + }), options); + if (latest.Item?.user_id !== task.user_id || !TERMINAL_STATUSES.includes(latest.Item.status)) return; + task = latest.Item as ContinuableTask; + const handle = task.microvm_start?.handle as MicrovmHandle | undefined; + await closeTaskApprovals(task.task_id, task.user_id, options); + if (handle) { + await strategy.stopSession(handle, options); + const observed = await strategy.pollSession(handle, options); + if (!['TERMINATED', 'NOT_FOUND'].includes(observed.microvmState ?? '')) return; + } else { + // An unanswered RunMicrovm response has no handle to inspect. Its fixed + // service lifetime is the only safe bound on a potentially created worker. + const started = Date.parse(task.microvm_start?.createdAt ?? ''); + if (task.microvm_start && !Number.isFinite(started)) { + throw new Error('MICROVM_CONTINUATION_START_TIME_INVALID: cannot confirm the unknown worker lifetime'); + } + if (Number.isFinite(started) && Date.now() < started + MICROVM_MAX_DURATION_SECONDS * 1000) return; + } + let attempt = task.microvm_start?.clientToken; + const withoutStart = !task.microvm_start; + let leaseState: string | undefined; + if (withoutStart) { + // Replacement admission removes the old start receipt. Its new launch + // token lives on the coordinator-owned slot and lease, not the task id. + const lease = (await ddb.send(new GetCommand({ + TableName: TABLE, Key: workerLeaseKey(task.task_id), ConsistentRead: true, + }), options)).Item; + if (lease) { + // claimMicrovmStart persists its receipt before calling AWS and only + // while the task is active. This terminal task has no receipt, so a + // matching admitted attempt cannot still launch. Finalization may have + // already released its slot; that does not erase its attempt identity. + const notLaunched = ['ACTIVE', 'FENCED'].includes(lease.lease_state) && !lease.lease_microvm_id + && ['held', 'released'].includes(task.concurrency_slot?.state ?? '') + && task.concurrency_slot?.attempt_id === lease.lease_attempt_id + && task.continuation?.attempt_id === lease.lease_attempt_id; + if (lease.lease_user_id !== task.user_id + || (!['PARKED', 'CLOSED'].includes(lease.lease_state) && !notLaunched) + || typeof lease.lease_attempt_id !== 'string' || !lease.lease_attempt_id) { + throw new Error('MICROVM_CONTINUATION_LEASE_INVALID: missing start does not prove worker shutdown'); + } + attempt = lease.lease_attempt_id; + leaseState = lease.lease_state; + } + } + await ddb.send(new UpdateCommand({ + TableName: TABLE, + Key: workerLeaseKey(task.task_id), + UpdateExpression: 'SET #ttl = :ttl, lease_state = :closed, lease_user_id = :user, lease_attempt_id = :attempt', + ConditionExpression: 'attribute_not_exists(task_id) OR (lease_user_id = :user AND lease_attempt_id = :attempt' + + (withoutStart ? ' AND lease_state = :observedState' : '') + + (['ACTIVE', 'FENCED'].includes(leaseState ?? '') ? ' AND attribute_not_exists(lease_microvm_id)' : '') + ')', + ExpressionAttributeNames: { '#ttl': 'ttl' }, + ExpressionAttributeValues: { + ':user': task.user_id, + ':closed': 'CLOSED', + ':attempt': attempt ?? task.task_id, + ...(withoutStart ? { ':observedState': leaseState ?? 'CLOSED' } : {}), + ':ttl': Math.floor(Date.now() / 1000) + Number(process.env.TASK_RETENTION_DAYS ?? '90') * 86400, + }, + }), options); + await releaseTaskSlot(task.task_id, task.user_id); + await deleteClosedTaskContinuations(task.task_id, task.user_id, options); + return; + } + if (task.status !== TaskStatus.AWAITING_APPROVAL || !task.awaiting_approval_request_id) return; + const handle = task.microvm_start?.handle as MicrovmHandle | undefined; + if (task.continuation?.state === 'READY' || task.continuation?.state === 'FENCED') { + if (!handle) throw new Error('MICROVM_CONTINUATION_HANDLE_MISSING: cannot confirm source shutdown'); + const observed = await strategy.pollSession(handle, options); + const terminal = ['TERMINATED', 'NOT_FOUND'].includes(observed.microvmState ?? ''); + const sourceStarted = observed.microvmStartedAtMs ?? Date.parse(task.microvm_start?.createdAt ?? ''); + await retireCheckpointedMicrovm({ + taskId: task.task_id, + userId: task.user_id, + handle, + strategy, + force: terminal, + abortSignal: options.abortSignal, + sessionDeadlineMs: Number.isFinite(sourceStarted) + ? sourceStarted + (observed.microvmMaximumDurationSeconds ?? MICROVM_MAX_DURATION_SECONDS) * 1000 + : Infinity, + }); + } + await dispatchMicrovmContinuation(task.task_id, task.user_id, task.awaiting_approval_request_id, options); +} + +/** + * Backstop for lost approval invokes, unfinished retirement, and closed-task + * object cleanup. The task table is scanned with the same bounded-page pattern + * as the concurrency reconciler; reserved lease items have no launch receipt. + */ +export async function handler(_event: unknown, context: Pick): Promise { + const saved = await ddb.send(new GetCommand({ + TableName: TABLE, Key: CURSOR_KEY, ConsistentRead: true, + }), { abortSignal: AbortSignal.timeout(CONTINUATION_IO_TIMEOUT_MS) }); + const initialCursor: Record | undefined = saved.Item?.cursor; + let lastKey = initialCursor; + const saveCursor = async (cursor?: Record) => { + const values = { + ...(cursor && { ':cursor': cursor }), + ...(initialCursor && { ':previous': initialCursor }), + }; + try { + await ddb.send(new UpdateCommand({ + TableName: TABLE, + Key: CURSOR_KEY, + UpdateExpression: cursor ? 'SET #cursor = :cursor' : 'REMOVE #cursor', + ConditionExpression: initialCursor ? '#cursor = :previous' : 'attribute_not_exists(#cursor)', + ExpressionAttributeNames: { '#cursor': 'cursor' }, + ...(Object.keys(values).length > 0 && { ExpressionAttributeValues: values }), + }), { abortSignal: AbortSignal.timeout(CURSOR_WRITE_TIMEOUT_MS) }); + } catch (error) { + if (!(error instanceof Error) || error.name !== 'ConditionalCheckFailedException') throw error; + logger.info('Another continuation sweep advanced the cursor; preserving its progress'); + } + }; + let processed = 0; + let failures = 0; + do { + if (context.getRemainingTimeInMillis() < MIN_REMAINING_MS) { + await saveCursor(lastKey); + return; + } + const page: ScanCommandOutput = await ddb.send(new ScanCommand({ + TableName: TABLE, + ConsistentRead: true, + Limit: 100, + FilterExpression: 'attribute_exists(continuation_launch)', + ExclusiveStartKey: lastKey, + }), { abortSignal: AbortSignal.timeout(CONTINUATION_IO_TIMEOUT_MS) }); + // Small batches bound both downstream concurrency and the invocation budget. + const tasks = page.Items ?? []; + let lastProcessedKey = lastKey; + for (let index = 0; index < tasks.length; index += BATCH_SIZE) { + if (context.getRemainingTimeInMillis() < MIN_REMAINING_MS) { + await saveCursor(lastProcessedKey); + logger.warn('Continuation reconciliation reached its invocation budget', { processed, failures }); + return; + } + // The slice bounds work to BATCH_SIZE. + // eslint-disable-next-line @cdklabs/promiseall-no-unbounded-parallelism + await Promise.all(tasks.slice(index, index + BATCH_SIZE).map(async row => { + try { + await reconcileMicrovmContinuation(row as ContinuableTask); + processed++; + } catch (error) { + failures++; + logger.warn('Continuation reconciliation will retry', { task_id: row.task_id, ...microvmErrorIdentity(error) }); + } + })); + lastProcessedKey = { task_id: tasks[Math.min(index + BATCH_SIZE - 1, tasks.length - 1)].task_id }; + } + lastKey = page.LastEvaluatedKey; + } while (lastKey); + await saveCursor(); + logger.info('Continuation reconciliation finished', { processed, failures }); +} diff --git a/cdk/src/handlers/reconcile-stranded-tasks.ts b/cdk/src/handlers/reconcile-stranded-tasks.ts index 4d5fa988f..8a1098294 100644 --- a/cdk/src/handlers/reconcile-stranded-tasks.ts +++ b/cdk/src/handlers/reconcile-stranded-tasks.ts @@ -30,14 +30,10 @@ * in `orchestrator.ts` via the `agent_heartbeat_at` timeout path — this * reconciler targets `SUBMITTED`, `HYDRATING`, and `AWAITING_APPROVAL`. * - * AWAITING_APPROVAL reconciliation (§13.6): - * If the agent container evicts mid-approval, neither the poll loop - * nor the resume transaction ever fires; the task sits in - * AWAITING_APPROVAL indefinitely while its approval row's TTL - * eventually reaps the row. This reconciler sweeps any - * AWAITING_APPROVAL task whose age exceeds the stranded timeout and - * transitions it to FAILED with a specific reason so the user sees - * a clear failure rather than a silent hang. + * AWAITING_APPROVAL tasks with a saved MicroVM checkpoint belong to the + * continuation manager and can remain open across workers. For other tasks, + * the backstop allows the full eight-hour worker lifetime plus cleanup grace; + * it detects a lost worker/coordinator, not an unanswered human deadline. */ import { @@ -47,13 +43,14 @@ import { PutItemCommand, } from '@aws-sdk/client-dynamodb'; import { ulid } from 'ulid'; +import { closeTaskApprovals } from './shared/close-task-approvals'; import { logger } from './shared/logger'; +import { releaseTaskSlot } from './shared/task-concurrency'; import { makeClient } from './shared/ua'; const ddb = makeClient(DynamoDBClient); const TASK_TABLE = process.env.TASK_TABLE_NAME!; const EVENTS_TABLE = process.env.TASK_EVENTS_TABLE_NAME!; -const CONCURRENCY_TABLE = process.env.USER_CONCURRENCY_TABLE_NAME!; /** Stranded-task timeout. The orchestrator Lambda is async-invoked and * the agent runtime has a cold-start path; 1200 s covers Lambda retries @@ -65,17 +62,12 @@ const STRANDED_TIMEOUT_SECONDS = Number( const TASK_RETENTION_DAYS = Number(process.env.TASK_RETENTION_DAYS ?? '90'); /** - * Separate (longer) timeout for AWAITING_APPROVAL tasks — approvals - * legitimately sit for an hour (the §7.3 ceiling for per-task approval - * timeout). A task's approval row carries its own per-row TTL; this - * sweep is the backstop for the case where the row gets reaped by DDB - * TTL but the TaskTable row never gets unstuck. - * - * Default: 7200s (2 hours) — double the §7.3 1-hour ceiling + an hour - * grace so this reconciler never races the happy-path timer. + * Backstop for approval waits without a recoverable checkpoint: the worker's + * eight-hour lifetime plus 30 minutes for its coordinator to close the task. + * Approval records do not expire merely because their worker stopped. */ const APPROVAL_STRANDED_TIMEOUT_SECONDS = Number( - process.env.APPROVAL_STRANDED_TIMEOUT_SECONDS ?? '7200', + process.env.APPROVAL_STRANDED_TIMEOUT_SECONDS ?? '30600', ); interface StrandedCandidate { @@ -123,6 +115,9 @@ async function findStrandedCandidates( const userId = item.user_id?.S; const createdAt = item.created_at?.S; if (!taskId || !userId || !createdAt) continue; + const continuationState = item.continuation?.M?.state?.S; + if (status === 'AWAITING_APPROVAL' + && ['READY', 'FENCED', 'PARKED', 'STARTING', 'RESTORING'].includes(continuationState ?? '')) continue; // Age by time-in-CURRENT-status, not creation time (#441). A task // that waited in the admission queue longer than the stranded @@ -180,14 +175,23 @@ async function failStrandedTask(task: StrandedCandidate): Promise { UpdateExpression: 'SET #s = :failed, updated_at = :now, completed_at = :now, ' + 'error_message = :err, status_created_at = :sca', - ConditionExpression: '#s = :expected', - ExpressionAttributeNames: { '#s': 'status' }, + ConditionExpression: '#s = :expected' + (task.status === 'AWAITING_APPROVAL' + ? ' AND (attribute_not_exists(continuation.#state) OR NOT (continuation.#state IN (:ready, :fenced, :parked, :starting, :restoring)))' + : ''), + ExpressionAttributeNames: { '#s': 'status', ...(task.status === 'AWAITING_APPROVAL' && { '#state': 'state' }) }, ExpressionAttributeValues: { ':failed': { S: 'FAILED' }, ':expected': { S: task.status }, ':now': { S: now }, ':err': { S: errorMessage }, ':sca': { S: `FAILED#${now}` }, + ...(task.status === 'AWAITING_APPROVAL' && { + ':ready': { S: 'READY' }, + ':fenced': { S: 'FENCED' }, + ':parked': { S: 'PARKED' }, + ':starting': { S: 'STARTING' }, + ':restoring': { S: 'RESTORING' }, + }), }, })); } catch (err: unknown) { @@ -286,32 +290,11 @@ async function failStrandedTask(task: StrandedCandidate): Promise { } } - // 3. Release the concurrency slot. Best-effort; drift is later corrected - // by the concurrency reconciler. - try { - await ddb.send(new UpdateItemCommand({ - TableName: CONCURRENCY_TABLE, - Key: { user_id: { S: task.user_id } }, - UpdateExpression: 'SET active_count = active_count - :one, updated_at = :now', - ConditionExpression: 'active_count > :zero', - ExpressionAttributeValues: { - ':one': { N: '1' }, - ':zero': { N: '0' }, - ':now': { S: now }, - }, - })); - } catch (decrErr: unknown) { - if (decrErr && typeof decrErr === 'object' && 'name' in decrErr - && decrErr.name !== 'ConditionalCheckFailedException') { - logger.warn('Failed to decrement concurrency for stranded task', { - task_id: task.task_id, - user_id: task.user_id, - error: decrErr instanceof Error ? decrErr.message : String(decrErr), - }); - } - // ConditionalCheckFailedException means the counter is already 0 — - // drift the concurrency reconciler will eventually catch. - } + // 3. Cooperate with normal finalization through the task-owned marker. + // If this fails after the terminal write, the capacity reconciler retries + // release for terminal held reservations on its next sweep. + await releaseTaskSlot(task.task_id, task.user_id); + await closeTaskApprovals(task.task_id, task.user_id); return true; } diff --git a/cdk/src/handlers/request-approval.ts b/cdk/src/handlers/request-approval.ts new file mode 100644 index 000000000..31a534d5a --- /dev/null +++ b/cdk/src/handlers/request-approval.ts @@ -0,0 +1,205 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { GetCommand, TransactWriteCommand } from '@aws-sdk/lib-dynamodb'; +import type { APIGatewayProxyEvent, APIGatewayProxyResult } from 'aws-lambda'; +import { logger } from './shared/logger'; +import { workerLeaseKey } from './shared/microvm-continuation-types'; +import { errorResponse, successResponse } from './shared/response'; +import { makeDocClient } from './shared/ua'; +import constants from '../../../contracts/constants.json'; + +const ddb = makeDocClient(); +const TASKS = process.env.TASK_TABLE_NAME!; +const APPROVALS = process.env.TASK_APPROVALS_TABLE_NAME!; +const MAX_ID_LENGTH = 128; +const MAX_TEXT_LENGTH = 8192; +const REQUEST_FIELDS = new Set([ + 'task_id', 'request_id', 'tool_name', 'tool_input_preview', 'tool_input_sha256', + 'reason', 'severity', 'matching_rule_ids', 'status', 'created_at', 'timeout_s', + 'deadline_epoch', 'user_id', 'repo', +]); + +interface RequestInput { + readonly operation: 'create' | 'timeout'; + readonly task_id: string; + readonly request_id: string; + readonly worker_attempt_id?: string; + readonly approval?: Record; + readonly reason?: string; +} + +function validShortText(value: unknown): value is string { + return typeof value === 'string' && value.length > 0 && value.length <= MAX_ID_LENGTH; +} + +function validId(value: unknown): value is string { + return validShortText(value) && /^[A-Za-z0-9][A-Za-z0-9_-]*$/.test(value); +} + +/** Never copy decision, notification, retention or arbitrary caller fields. */ +function validateRequest(input: RequestInput, task: Record): Record { + const row = input.approval; + if (!row || Object.keys(row).some(key => !REQUEST_FIELDS.has(key)) + || row.task_id !== input.task_id || row.request_id !== input.request_id + || row.user_id !== task.user_id || row.repo !== (task.repo ?? '') + || row.status !== 'PENDING' + || !validShortText(row.tool_name) || typeof row.tool_input_preview !== 'string' || row.tool_input_preview.length > MAX_TEXT_LENGTH + || typeof row.tool_input_sha256 !== 'string' || !/^[a-f0-9]{64}$/.test(row.tool_input_sha256) + || typeof row.reason !== 'string' || row.reason.length > MAX_TEXT_LENGTH + || !['low', 'medium', 'high'].includes(row.severity as string) + || !Array.isArray(row.matching_rule_ids) || row.matching_rule_ids.length > 500 + || !row.matching_rule_ids.every(validShortText) + || typeof row.created_at !== 'string' || !Number.isFinite(Date.parse(row.created_at)) + || typeof row.timeout_s !== 'number' || !Number.isInteger(row.timeout_s) + || row.timeout_s < 0 || row.timeout_s > constants.approval_timeout_s.max + || (row.timeout_s === 0 ? row.deadline_epoch !== undefined + : row.deadline_epoch !== Math.floor(Date.parse(row.created_at) / 1000) + row.timeout_s)) { + throw new Error('APPROVAL_REQUEST_INVALID'); + } + return row; +} + +/** + * Worker-callable writer: creates pending requests or records a non-human timeout. + * Human decisions remain exclusively in approve/deny handlers. Workers have no + * direct approvals-table write permission, including whole-row replacement. + */ +export async function recordWorkerRequest(input: RequestInput): Promise> { + if (!input || !validId(input.task_id) || !validId(input.request_id) + || (input.worker_attempt_id !== undefined && !validId(input.worker_attempt_id)) + || !['create', 'timeout'].includes(input.operation)) { + return { ok: false, code: 'APPROVAL_REQUEST_INVALID' }; + } + try { + const task = (await ddb.send(new GetCommand({ + TableName: TASKS, Key: { task_id: input.task_id }, ConsistentRead: true, + }))).Item; + if (!task || typeof task.user_id !== 'string') return { ok: false, code: 'APPROVAL_TASK_MISSING' }; + const lease = task.compute_type === 'lambda-microvm' ? [{ + ConditionCheck: { + TableName: TASKS, + Key: workerLeaseKey(input.task_id), + ConditionExpression: 'lease_state = :active AND lease_attempt_id = :attempt AND lease_user_id = :user', + ExpressionAttributeValues: { + ':active': 'ACTIVE', ':attempt': input.worker_attempt_id ?? '', ':user': task.user_id, + }, + }, + }] : []; + if (input.operation === 'create') { + const row = validateRequest(input, task); + await ddb.send(new TransactWriteCommand({ + TransactItems: [{ + Put: { + TableName: APPROVALS, + Item: row, + ConditionExpression: 'attribute_not_exists(request_id)', + }, + }, { + Update: { + TableName: TASKS, + Key: { task_id: input.task_id }, + UpdateExpression: 'SET #status = :awaiting, awaiting_approval_request_id = :request', + ConditionExpression: '#status = :running AND user_id = :user', + ExpressionAttributeNames: { '#status': 'status' }, + ExpressionAttributeValues: { + ':awaiting': 'AWAITING_APPROVAL', + ':running': 'RUNNING', + ':request': input.request_id, + ':user': task.user_id, + }, + }, + }, ...lease], + })); + } else { + // A worker may fail closed after a deadline or polling failure; it cannot + // turn either into human DENIED/APPROVED or replace an existing decision. + if (input.approval !== undefined || (input.reason !== undefined + && (typeof input.reason !== 'string' || input.reason.length > MAX_TEXT_LENGTH))) { + return { ok: false, code: 'APPROVAL_REQUEST_INVALID' }; + } + await ddb.send(new TransactWriteCommand({ + TransactItems: [{ + Update: { + TableName: APPROVALS, + Key: { task_id: input.task_id, request_id: input.request_id }, + UpdateExpression: 'SET #status = :timeout, decided_at = :now' + + (input.reason !== undefined ? ', deny_reason = :reason' : ''), + ConditionExpression: '#status = :pending AND user_id = :user', + ExpressionAttributeNames: { '#status': 'status' }, + ExpressionAttributeValues: { + ':timeout': 'TIMED_OUT', + ':pending': 'PENDING', + ':user': task.user_id, + ':now': new Date().toISOString(), + ...(input.reason !== undefined ? { ':reason': input.reason } : {}), + }, + }, + }, { + ConditionCheck: { + TableName: TASKS, + Key: { task_id: input.task_id }, + ConditionExpression: '#status = :awaiting AND awaiting_approval_request_id = :request AND user_id = :user', + ExpressionAttributeNames: { '#status': 'status' }, + ExpressionAttributeValues: { ':awaiting': 'AWAITING_APPROVAL', ':request': input.request_id, ':user': task.user_id }, + }, + }, ...lease], + })); + } + return { ok: true }; + } catch (error) { + const failure = error as { name?: string; message?: string; CancellationReasons?: { Code?: string }[] }; + if (failure.message === 'APPROVAL_REQUEST_INVALID') return { ok: false, code: failure.message }; + const reasons = failure.CancellationReasons?.map(reason => ({ Code: reason.Code })); + logger.warn('Worker approval request was not recorded', { + task_id: input.task_id, + request_id: input.request_id, + operation: input.operation, + error_type: failure.name, + cancellation_reasons: reasons, + }); + return { ok: false, code: failure.name ?? 'APPROVAL_WRITE_FAILED', cancellation_reasons: reasons }; + } +} + +/** IAM authorizes the signed task path before invoking this Lambda. */ +export async function handler(event: APIGatewayProxyEvent): Promise { + const requestId = event.requestContext?.requestId ?? 'unknown'; + if (!event.requestContext?.identity?.userArn || !validId(event.pathParameters?.task_id)) { + return errorResponse(403, 'APPROVAL_CALLER_UNAUTHENTICATED', 'IAM authentication required', requestId); + } + let input: RequestInput; + try { + input = JSON.parse(event.isBase64Encoded + ? Buffer.from(event.body ?? '', 'base64').toString('utf8') : event.body ?? ''); + } catch { + return errorResponse(400, 'APPROVAL_REQUEST_INVALID', 'Invalid JSON request', requestId); + } + if (!input || input.task_id !== event.pathParameters!.task_id) { + return errorResponse(400, 'APPROVAL_TASK_MISMATCH', 'Task must match the signed path', requestId); + } + const result = await recordWorkerRequest(input); + if (result.ok) return successResponse(200, result, requestId); + const code = String(result.code); + const statusCode = code === 'TransactionCanceledException' ? 409 + : code === 'APPROVAL_TASK_MISSING' ? 404 : code === 'APPROVAL_REQUEST_INVALID' ? 400 : 503; + return errorResponse(statusCode, code, 'Approval write was not acknowledged', requestId, { + cancellation_reasons: result.cancellation_reasons ?? [], + }); +} diff --git a/cdk/src/handlers/shared/agent-heartbeat.ts b/cdk/src/handlers/shared/agent-heartbeat.ts new file mode 100644 index 000000000..7e8f4cbae --- /dev/null +++ b/cdk/src/handlers/shared/agent-heartbeat.ts @@ -0,0 +1,36 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +// SPDX-License-Identifier: MIT-0 + +/** Shared by the two runtimes that run server.py's periodic heartbeat worker. */ +export const AGENT_HEARTBEAT_GRACE_SEC = 120; +export const AGENT_HEARTBEAT_STALE_SEC = 240; + +export function evaluateAgentHeartbeat( + startedAtMs: number | undefined, heartbeatAtMs: number | undefined, nowMs: number, +): 'stale' | 'missing' | undefined { + if (startedAtMs === undefined || !Number.isFinite(startedAtMs)) return undefined; + const age = (nowMs - startedAtMs) / 1000; + if (heartbeatAtMs !== undefined) { + return age > AGENT_HEARTBEAT_GRACE_SEC + && (nowMs - heartbeatAtMs) / 1000 > AGENT_HEARTBEAT_STALE_SEC ? 'stale' : undefined; + } + return age > AGENT_HEARTBEAT_GRACE_SEC + AGENT_HEARTBEAT_STALE_SEC ? 'missing' : undefined; +} diff --git a/cdk/src/handlers/shared/approval-decision.ts b/cdk/src/handlers/shared/approval-decision.ts new file mode 100644 index 000000000..c80ace88c --- /dev/null +++ b/cdk/src/handlers/shared/approval-decision.ts @@ -0,0 +1,111 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +// SPDX-License-Identifier: MIT-0 + +import { PutCommand, type DynamoDBDocumentClient } from '@aws-sdk/lib-dynamodb'; +import type { Context } from 'aws-lambda'; +import { ulid } from 'ulid'; +import { logger } from './logger'; +import { APPROVAL_AUDIT_TIMEOUT_MS, approvalPostCommitOptions, wakeMicrovmAfterApproval } from './microvm-approval-wake'; +import { microvmErrorIdentity } from './microvm-control'; + +export interface DecisionPostCommitInput { + readonly ddb: DynamoDBDocumentClient; + readonly eventsTableName: string; + readonly taskId: string; + readonly callerUserId: string; + readonly requestId: string; + readonly decision: 'APPROVED' | 'DENIED'; + /** Decision-specific audit fields (approve: `scope`; deny: `reason`). */ + readonly auditMetadata: Record; + readonly decidedAt: string; + readonly nowEpoch: number; + readonly retentionDays: number; + readonly invocationStartedMs: number; + readonly context?: Pick; +} + +/** + * Shared approve/deny tail after the decision transaction commits: write the + * `approval_decision_recorded` audit event, then attempt the optional MicroVM + * wake. Neither step may fail the request — the human decision is already + * committed on TaskApprovalsTable, and the durable supervisor recovers a + * sleeping worker if the wake cannot finish here. + */ +export async function recordDecisionPostCommit(input: DecisionPostCommitInput): Promise { + const postCommit = approvalPostCommitOptions(input.invocationStartedMs, input.context); + const ttl = input.nowEpoch + input.retentionDays * 86400; + try { + const abortSignal = AbortSignal.any([postCommit.abortSignal!, AbortSignal.timeout(APPROVAL_AUDIT_TIMEOUT_MS)]); + abortSignal.throwIfAborted(); + await input.ddb.send(new PutCommand({ + TableName: input.eventsTableName, + Item: { + task_id: input.taskId, + event_id: ulid(), + event_type: 'approval_decision_recorded', + timestamp: input.decidedAt, + ttl, + metadata: { + request_id: input.requestId, + status: input.decision, + ...input.auditMetadata, + decided_at: input.decidedAt, + caller_user_id: input.callerUserId, + }, + }, + }), { abortSignal }); + } catch (auditErr) { + logger.warn('approval_decision_recorded audit write failed (decision already committed)', { + task_id: input.taskId, + request_id: input.requestId, + ...microvmErrorIdentity(auditErr), + }); + } + + try { + await wakeMicrovmAfterApproval({ + taskId: input.taskId, + userId: input.callerUserId, + requestId: input.requestId, + decision: input.decision, + options: postCommit, + emitEvent: async (eventType, metadata, options) => { + options.abortSignal?.throwIfAborted(); + await input.ddb.send(new PutCommand({ + TableName: input.eventsTableName, + Item: { + task_id: input.taskId, + user_id: input.callerUserId, + event_id: ulid(), + event_type: eventType, + timestamp: new Date().toISOString(), + ttl, + metadata, + }, + }), options); + }, + }); + } catch (wakeError) { + logger.warn('MicroVM wake helper failed after decision commit', { + task_id: input.taskId, request_id: input.requestId, ...microvmErrorIdentity(wakeError), + }); + } +} diff --git a/cdk/src/handlers/shared/approval-notifications.ts b/cdk/src/handlers/shared/approval-notifications.ts new file mode 100644 index 000000000..4382d7a0d --- /dev/null +++ b/cdk/src/handlers/shared/approval-notifications.ts @@ -0,0 +1,233 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { GetCommand, UpdateCommand, type DynamoDBDocumentClient } from '@aws-sdk/lib-dynamodb'; +import { scanDenyReason } from './deny-reason-scanner'; +import { logger } from './logger'; +import type { TaskRecord } from './types'; +import { TaskStatus, TERMINAL_STATUSES } from '../../constructs/task-status'; + +const REQUEST_ID_MAX_LENGTH = 128; +const SEVERITY_MAX_LENGTH = 20; +const SCOPE_MAX_LENGTH = 150; + +/** Decisions are acknowledged at commit time, even when the worker cannot wake. */ +export const APPROVAL_NOTIFICATION_EVENTS = [ + 'approval_requested', 'approval_decision_recorded', 'approval_timed_out', + 'approval_cancelled', 'approval_stranded', +] as const; + +export function isApprovalNotification(eventType: string): boolean { + return (APPROVAL_NOTIFICATION_EVENTS as readonly string[]).includes(eventType); +} + +export interface ApprovalNotification { + readonly title: string; + readonly text: string; + readonly taskId: string; + readonly requestId: string; + readonly userId: string; + readonly marker: string; +} + +function text(value: unknown, max = 500): string { + if (typeof value !== 'string') return ''; + // Redact before truncation so cutting a token cannot hide its recognizable shape. + const clean = scanDenyReason(value) + .replace(/\u001b\[[0-?]*[ -/]*[@-~]/g, '') + .replace(/[\u0000-\u0008\u000b-\u001f\u007f\u200e\u200f\u202a-\u202e\u2066-\u2069]/g, ''); + return clean.length > max ? `${clean.slice(0, max)}… [shortened]` : clean; +} + +/** Describe only recorded arguments; tool names cannot establish intent or safety. */ +function describeAction(row: Record, task: TaskRecord): string[] { + const tool = text(row.tool_name, 100); + const preview = typeof row.tool_input_preview === 'string' ? row.tool_input_preview : ''; + let input: Record | undefined; + try { + const parsed: unknown = JSON.parse(preview); + if (parsed && typeof parsed === 'object' && !Array.isArray(parsed)) input = parsed as Record; + } catch { + // The guest stores a bounded preview, which may end partway through JSON. + } + const quote = (value: unknown) => JSON.stringify(text(value)); + const path = typeof input?.file_path === 'string' ? input.file_path : ''; + const workspace = `/workspace/${task.task_id}/`; + const target = quote(path.startsWith(workspace) ? path.slice(workspace.length) : path); + let action: string; + if (tool === 'Read' && path) { + action = `The agent wants to read the file ${target}.`; + } else if (tool === 'Write' && path) { + action = `The agent wants to write to ${target}, creating the file or replacing its contents.`; + } else if (tool === 'Edit' && path) { + action = `The agent wants to change text in ${target}.`; + } else if (tool === 'Bash') { + action = 'The agent wants to run a shell command.'; + } else if (tool === 'WebFetch' && typeof input?.url === 'string') { + action = `The agent wants to fetch content from ${quote(input.url)}.`; + } else if (tool === 'Glob' || tool === 'Grep') { + action = `The agent wants to search ${tool === 'Glob' ? 'file names' : 'file contents'}.`; + } else { + action = `The agent wants to call the tool ${quote(tool)}.`; + } + const lines = [action]; + if (task.repo) lines.push(`Repository: ${text(task.repo)}`); + if (typeof input?.description === 'string' && input.description.trim()) { + lines.push(`Agent's explanation: ${text(input.description)}`); + } + if (tool === 'Bash' && typeof input?.command === 'string') { + lines.push(`Command: ${quote(input.command)}`); + } + // Always retain the preview: summaries must not hide flags, ranges or edits. + lines.push(`Saved arguments: ${text(preview) || '(not available)'}`); + if (!input || preview.endsWith('...') || preview.length > 500) { + lines.push('The saved arguments are incomplete or could not be interpreted. They may not show the full action.'); + } + return lines; +} + +/** Read saved state instead of showing an old, delayed "please approve" event. */ +export async function loadApprovalNotification( + ddb: DynamoDBDocumentClient, + task: TaskRecord, + eventType: string, + metadata: Record, + channel: 'slack' | 'linear', +): Promise { + if (!isApprovalNotification(eventType)) return null; + const tableName = process.env.TASK_APPROVALS_TABLE_NAME; + if (!tableName) throw new Error('Approval notifications require TASK_APPROVALS_TABLE_NAME'); + // The deployed stranded-task reconciler closes the task, leaves its approval + // row PENDING, and emits a milestone without request_id. Recover that identity + // only from the consistently read, failed owning task and its saved cause. + const legacyStranded = eventType === 'approval_stranded' + && task.status === TaskStatus.FAILED + && metadata.reason === 'STRANDED_NO_HEARTBEAT' + && task.error_message?.startsWith('Approval stranded:'); + const requestId = metadata.request_id ?? (legacyStranded ? task.awaiting_approval_request_id : undefined); + if (!requestId && eventType === 'approval_stranded') { + logger.warn('Stranded approval notification has no saved request identity', { + event: 'approval_notification_request_missing', task_id: task.task_id, + }); + return null; + } + if (typeof requestId !== 'string' || !requestId || requestId.length > REQUEST_ID_MAX_LENGTH) { + throw new Error('Approval notification is missing a valid request_id'); + } + const response = await ddb.send(new GetCommand({ + TableName: tableName, Key: { task_id: task.task_id, request_id: requestId }, ConsistentRead: true, + })); + const row = response.Item; + const marker = `notified_${channel}_${eventType}`; + if (!row || row.user_id !== task.user_id || row[marker]) return null; + const status = legacyStranded && row.status === 'PENDING' + && task.awaiting_approval_request_id === requestId ? 'STRANDED' : row.status; + const expected = eventType === 'approval_requested' ? ['PENDING'] + : eventType === 'approval_decision_recorded' ? ['APPROVED', 'DENIED'] + : eventType === 'approval_timed_out' ? ['TIMED_OUT'] + : eventType === 'approval_cancelled' ? ['CANCELLED'] : ['STRANDED']; + if (!expected.includes(status)) return null; + if (status === 'PENDING' + && (task.status !== 'AWAITING_APPROVAL' || task.awaiting_approval_request_id !== requestId)) return null; + + const title = status === 'PENDING' ? 'Approval needed' + : status === 'APPROVED' ? 'Approval recorded' + : status === 'DENIED' ? 'Denial recorded' + : status === 'TIMED_OUT' ? 'Approval request timed out' + : status === 'CANCELLED' ? 'Approval request cancelled' : 'Approval wait could not continue'; + const lines = [title]; + if (status === 'PENDING') { + lines.push('', ...describeAction(row, task), ''); + const reason = text(row.reason); + lines.push(reason.startsWith('Soft-deny:') + ? 'Why approval is required: your configured policy requires a human decision for this action.' + : `Why approval is required: ${reason || 'No explanation was saved with this request.'}`); + lines.push('Approve: allow this action once.', + 'Deny: block this action and return the decision to the agent.'); + const deadline = Date.parse(row.created_at) + Number(row.timeout_s) * 1000; + if (Number.isFinite(deadline) && Number(row.timeout_s) > 0) { + lines.push(`Decision deadline: ${new Date(deadline).toISOString()}`); + } + if (channel === 'linear') { + lines.push('', 'Reply approve or deny to this comment while signed in as the task owner.'); + } + lines.push('', 'Technical details / CLI alternative', + `Task: ${task.task_id}`, `Request: ${requestId}`, + `Tool: ${text(row.tool_name, 100)}`, `Policy severity: ${text(row.severity, SEVERITY_MAX_LENGTH)}`, + `Policy detail: ${reason}`); + // Never interpolate untrusted text into suggested shell commands. + if ([task.task_id, requestId].every(id => /^[A-Za-z0-9_-]{1,128}$/.test(id))) { + lines.push('Respond using the CLI while signed in as the task owner:', + `bgagent approve ${task.task_id} ${requestId} --scope this_call`, + `bgagent deny ${task.task_id} ${requestId}`); + } + lines.push('Run bgagent pending to see currently open requests.'); + } else if (status === 'APPROVED') { + lines.push(`Task: ${task.task_id}`, `Request: ${requestId}`); + lines.push(`Scope: ${text(row.scope, SCOPE_MAX_LENGTH)}`); + lines.push(TERMINAL_STATUSES.includes(task.status) + ? `The decision is saved. Task status: ${task.status}. The task has already ended.` + : 'The decision is saved. The agent will continue when its worker is ready.'); + } else if (status === 'DENIED') { + lines.push(`Task: ${task.task_id}`, `Request: ${requestId}`); + lines.push(`Reason: ${text(row.deny_reason) || 'No reason supplied.'}`); + lines.push(TERMINAL_STATUSES.includes(task.status) + ? `The decision is saved. Task status: ${task.status}. The task has already ended.` + : 'The decision is saved. The agent will receive the denial when its worker is ready.'); + } else { + lines.push(`Task: ${task.task_id}`, `Request: ${requestId}`); + // reason is the original policy explanation, not the cause of closure. + // The guest stores a polling failure in deny_reason on its timeout path. + const reason = status === 'TIMED_OUT' + ? text(row.deny_reason) || 'The configured decision deadline was reached.' + : status === 'CANCELLED' + ? text(row.cancellation_reason) || 'Task cancelled by its owner.' + : text(task.error_message) || 'The agent could not continue this approval wait.'; + lines.push(`Reason: ${reason}`); + } + return { title, text: lines.join('\n'), taskId: task.task_id, requestId, userId: task.user_id, marker }; +} + +/** Persist only after the external service accepted the message, so failures retry. */ +export async function markApprovalNotificationDelivered( + ddb: DynamoDBDocumentClient, + notification: ApprovalNotification, +): Promise { + try { + await ddb.send(new UpdateCommand({ + TableName: process.env.TASK_APPROVALS_TABLE_NAME!, + Key: { task_id: notification.taskId, request_id: notification.requestId }, + UpdateExpression: 'SET #marker = :now', + ConditionExpression: 'attribute_exists(task_id) AND user_id = :user', + ExpressionAttributeNames: { '#marker': notification.marker }, + ExpressionAttributeValues: { ':now': new Date().toISOString(), ':user': notification.userId }, + })); + } catch (error) { + if ((error as { name?: string }).name !== 'ConditionalCheckFailedException') throw error; + // TTL removal after a successful post must not recreate the approval row. + logger.info('Approval disappeared before its notification receipt was saved', { + event: 'approval_notification_receipt_missing', task_id: notification.taskId, request_id: notification.requestId, + }); + } +} + +/** Escape preview Markdown, including backticks, so repository text cannot break out of its code fence. */ +export function approvalNotificationMarkdown(notification: ApprovalNotification): string { + return `\`\`\`text\n${notification.text.replace(/`/g, 'ˋ')}\n\`\`\``; +} diff --git a/cdk/src/handlers/shared/builtin-policies.ts b/cdk/src/handlers/shared/builtin-policies.ts index e4c983f3f..e2dd5d9fc 100644 --- a/cdk/src/handlers/shared/builtin-policies.ts +++ b/cdk/src/handlers/shared/builtin-policies.ts @@ -126,20 +126,20 @@ permit (principal, action, resource); // Every rule in this file MUST carry: // @tier("soft") // @rule_id("...") — stable ID for --pre-approve rule:X -// @approval_timeout_s — integer seconds >= 30 (<120 emits WARN per IMPL-25) // @severity — "low" | "medium" | "high" // @category — optional free-form UX grouping +// An optional @approval_timeout_s sets a positive explicit deadline (>= 30s). +// Omitting it uses the task setting, whose default is no automatic expiry. // // Blueprints may OPT OUT of specific rules here via // \`security.cedarPolicies.disable: [rule_id]\`. They may NOT disable any // rule in hard_deny.cedar (blueprint loader rejects those at task start). -// Gate any git --force / -f push. 300s default approval window, medium severity. +// Gate any git --force / -f push, with medium severity. // Covers both long-form (--force) and short-form (-f) variants, including // the bare \`git push -f\` invocation with no branch argument. @tier("soft") @rule_id("force_push_any") -@approval_timeout_s("300") @severity("medium") @category("destructive") forbid (principal, action == Agent::Action::"execute_bash", resource) @@ -147,12 +147,11 @@ when { context.command like "*git push --force*" || context.command like "*git push -f *" || context.command like "*git push -f" }; -// Force-push to main/prod specifically — longer window, higher severity. +// Force-push to main/prod specifically — higher severity. // Multi-match with force_push_any is expected: the engine's annotation -// merging picks min(300, 600)=300s and max(medium, high)=high. +// merging picks max(medium, high)=high. There is no implicit approval expiry. @tier("soft") @rule_id("force_push_main") -@approval_timeout_s("600") @severity("high") @category("destructive") forbid (principal, action == Agent::Action::"execute_bash", resource) @@ -165,7 +164,6 @@ when { context.command like "*git push --force origin main*" // agent bypasses PR workflow by pushing directly. @tier("soft") @rule_id("push_to_protected_branch") -@approval_timeout_s("300") @severity("medium") @category("destructive") forbid (principal, action == Agent::Action::"execute_bash", resource) @@ -174,20 +172,18 @@ when { context.command like "*git push origin main*" || context.command like "*git push origin prod*" || context.command like "*git push origin release/*" }; -// Writes to \`.env\` files typically contain secrets. 600s window, high severity. +// Writes to \`.env\` files typically contain secrets. High severity. @tier("soft") @rule_id("write_env_files") -@approval_timeout_s("600") @severity("high") @category("filesystem") forbid (principal, action == Agent::Action::"write_file", resource) when { context.file_path like "*.env" }; // Writes to any path containing "credentials" — SSH keys, AWS creds, -// service-account JSON, etc. 300s window, high severity. +// service-account JSON, etc. High severity. @tier("soft") @rule_id("write_credentials") -@approval_timeout_s("300") @severity("high") @category("auth") forbid (principal, action == Agent::Action::"write_file", resource) diff --git a/cdk/src/handlers/shared/canonical-json.ts b/cdk/src/handlers/shared/canonical-json.ts new file mode 100644 index 000000000..f89bab658 --- /dev/null +++ b/cdk/src/handlers/shared/canonical-json.ts @@ -0,0 +1,26 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +/** Stable JSON bytes for ABCA receipts; preserve this ordering for stored hashes. */ +export function canonicalJson(value: unknown): string { + return JSON.stringify(value, (_key, item: unknown) => + item && typeof item === 'object' && !Array.isArray(item) + ? Object.fromEntries(Object.entries(item).sort(([a], [b]) => a.localeCompare(b))) + : item); +} diff --git a/cdk/src/handlers/shared/close-task-approvals.ts b/cdk/src/handlers/shared/close-task-approvals.ts new file mode 100644 index 000000000..96a092208 --- /dev/null +++ b/cdk/src/handlers/shared/close-task-approvals.ts @@ -0,0 +1,96 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { GetCommand, QueryCommand, UpdateCommand, type QueryCommandOutput } from '@aws-sdk/lib-dynamodb'; +import type { SessionControlOptions } from './compute-strategy'; +import { makeDocClient } from './ua'; +import { TERMINAL_STATUSES } from '../../constructs/task-status'; + +const ddb = makeDocClient(); +const CLEANUP_TIMEOUT_MS = 10_000; +const UPDATE_BATCH_SIZE = 4; + +/** Task closure closes its unanswered requests; no action-content recheck is involved. */ +export async function closeTaskApprovals( + taskId: string, userId: string, options: SessionControlOptions = { abortSignal: AbortSignal.timeout(CLEANUP_TIMEOUT_MS) }, +): Promise { + if (!process.env.TASK_APPROVALS_TABLE_NAME) return; + const current = await ddb.send(new GetCommand({ + TableName: process.env.TASK_TABLE_NAME!, Key: { task_id: taskId }, ConsistentRead: true, + }), options); + if (current.Item?.user_id !== userId || !TERMINAL_STATUSES.includes(current.Item.status)) return; + const terminalStatus: string = current.Item.status; + let key: Record | undefined; + const now = new Date().toISOString(); + const ttl = Math.floor(Date.now() / 1000) + Number(process.env.TASK_RETENTION_DAYS ?? '90') * 86400; + do { + const page: QueryCommandOutput = await ddb.send(new QueryCommand({ + TableName: process.env.TASK_APPROVALS_TABLE_NAME, + ConsistentRead: true, + KeyConditionExpression: 'task_id = :task', + FilterExpression: 'attribute_not_exists(#ttl) OR #status = :pending', + ExpressionAttributeNames: { '#ttl': 'ttl', '#status': 'status' }, + ExpressionAttributeValues: { ':task': taskId, ':pending': 'PENDING' }, + ExclusiveStartKey: key, + }), options); + const rows = (page.Items ?? []).filter(row => + row.user_id === userId && typeof row.request_id === 'string' && typeof row.status === 'string'); + for (let index = 0; index < rows.length; index += UPDATE_BATCH_SIZE) { + // Each slice bounds concurrent writes; completed rows are filtered on retry. + // eslint-disable-next-line @cdklabs/promiseall-no-unbounded-parallelism + await Promise.all(rows.slice(index, index + UPDATE_BATCH_SIZE).map(async row => { + const pending = row.status === 'PENDING'; + try { + await ddb.send(new UpdateCommand({ + TableName: process.env.TASK_APPROVALS_TABLE_NAME, + Key: { task_id: taskId, request_id: row.request_id }, + UpdateExpression: 'SET #ttl = if_not_exists(#ttl, :ttl)' + (pending + ? ', #status = :cancelled, decided_at = :now, cancellation_reason = :reason' : ''), + ConditionExpression: 'user_id = :user AND #status = :observed', + ExpressionAttributeNames: { '#ttl': 'ttl', '#status': 'status' }, + ExpressionAttributeValues: { + ':ttl': ttl, + ':user': userId, + ':observed': row.status, + ...(pending && { + ':cancelled': 'CANCELLED', + ':now': now, + ':reason': `Owning task is ${terminalStatus.toLowerCase()}.`, + }), + }, + }), options); + } catch (error) { + if ((error as { name?: string }).name !== 'ConditionalCheckFailedException') throw error; + // An answer can win after the query. Preserve it and add only retention. + await ddb.send(new UpdateCommand({ + TableName: process.env.TASK_APPROVALS_TABLE_NAME, + Key: { task_id: taskId, request_id: row.request_id }, + UpdateExpression: 'SET #ttl = if_not_exists(#ttl, :ttl)', + ConditionExpression: 'user_id = :user AND #status <> :pending', + ExpressionAttributeNames: { '#ttl': 'ttl', '#status': 'status' }, + ExpressionAttributeValues: { ':ttl': ttl, ':user': userId, ':pending': 'PENDING' }, + }), options).catch(race => { + if ((race as { name?: string }).name !== 'ConditionalCheckFailedException') throw race; + }); + } + })); + } + key = page.LastEvaluatedKey; + } while (key); +} diff --git a/cdk/src/handlers/shared/compute-strategy.ts b/cdk/src/handlers/shared/compute-strategy.ts index aec753f3a..f4070d37c 100644 --- a/cdk/src/handlers/shared/compute-strategy.ts +++ b/cdk/src/handlers/shared/compute-strategy.ts @@ -17,6 +17,8 @@ * SOFTWARE. */ +import type { MicrovmState } from '@aws-sdk/client-lambda-microvms'; +import type { MicrovmImageMetadata } from './microvm-image-capability'; import type { BlueprintConfig, ComputeType } from './repo-config'; import { AgentCoreComputeStrategy } from './strategies/agentcore-strategy'; import { EcsComputeStrategy } from './strategies/ecs-strategy'; @@ -36,16 +38,15 @@ import { LambdaMicrovmComputeStrategy } from './strategies/lambda-microvm-strate * ADR-021 sub-decision 1: the MicroVM variant carries ``microvmId`` (every * lifecycle API — suspend/resume/terminate/get — takes only that identifier) * and ``endpoint`` (minted per session by ``RunMicrovm``, required for any - * future orchestrator→agent HTTP interaction). The image ARN/version is - * deliberately NOT in the handle: like the ECS task-definition ARN it is - * deployment-time configuration consumed by ``startSession`` from the - * construct-injected environment and recorded in the session-start log entry - * for diagnostics, not per-session lifecycle state. + * future orchestrator→agent HTTP interaction). P3 additionally retains the actual + * image ARN/version and verified lifecycle protocol. These describe the snapshot + * that launched this worker; current deployment settings cannot substitute for it. + * Legacy handles remain usable for cleanup, with new suspension disabled. */ export type SessionHandle = | { readonly sessionId: string; readonly strategyType: 'agentcore'; readonly runtimeArn: string } | { readonly sessionId: string; readonly strategyType: 'ecs'; readonly clusterArn: string; readonly taskArn: string } - | { readonly sessionId: string; readonly strategyType: 'lambda-microvm'; readonly microvmId: string; readonly endpoint: string }; + | ({ readonly sessionId: string; readonly strategyType: 'lambda-microvm'; readonly microvmId: string; readonly endpoint: string } & MicrovmImageMetadata); /** * Substrate-observed session state. Deliberately mechanical: the strategy @@ -71,26 +72,52 @@ export type SessionHandle = * ~12 s — reached the operator as the bare, and therefore fabricated, * ``"substrate state completed"``. * - * It is OPTIONAL and OPAQUE: no control flow may branch on its content (that would - * put substrate interpretation back in the strategy), and it is for the reconcile - * ``detail`` string and logs only. + * It is OPTIONAL and OPAQUE to the strategy. The orchestrator retains it as + * diagnostic detail and recognizes known run-rejection and resume-hook failure shapes + * to choose a stable failure code. Consumers classify that code, so arbitrary + * words in the reason cannot change the category or user-facing retry advice. * - * Declared on all four variants for UNIFORMITY, though only ``completed`` and - * ``failed`` are read today (``reconcileMicrovmSubstrateState`` returns early for - * the other two). The wide union is deliberate rather than dead weight: - * ``suspended.reason`` has a named future consumer — P3's suspend/resume policy - * (ADR-021 sub-decision 2) has to distinguish an orchestrator-intended suspend - * during an approval wait from a substrate-side one, and ``stateReason`` is the - * only evidence the substrate offers for that. Narrowing the union now would mean - * widening it again there, and a per-variant union would invite call sites to - * branch on which variant carries a reason — the opposite of the opacity rule - * above. + * Declared on all four variants for uniform diagnostics. Terminal failure + * formatting consumes ``completed`` and ``failed``; ``suspended.reason`` can + * explain an observation without deciding whether it is healthy. The policy + * must distinguish intended suspension through durable orchestrator intent and + * task/approval state, not by parsing this service-provided text. Keeping the + * field on every variant preserves the same diagnostic shape. */ -export type SessionStatus = +/** UNKNOWN/NOT_FOUND are local observations, not AWS service states. */ +export type MicrovmObservedState = MicrovmState | 'UNKNOWN' | 'NOT_FOUND'; + +export type SessionStatus = ( | { readonly status: 'running'; readonly reason?: string } | { readonly status: 'suspended'; readonly reason?: string } | { readonly status: 'completed'; readonly reason?: string } - | { readonly status: 'failed'; readonly error: string; readonly reason?: string }; + | { readonly status: 'failed'; readonly error: string; readonly reason?: string } +) & { + /** Explicit MicroVM observation; coarse `running` also covers pending/unknown. */ + readonly microvmState?: MicrovmObservedState; + /** Service observations used to retain the original lifetime across durable replay. */ + readonly microvmStartedAtMs?: number; + readonly microvmMaximumDurationSeconds?: number; +}; + +/** A caller may impose a shorter total budget across several control operations. */ +export interface SessionControlOptions { + readonly abortSignal?: AbortSignal; +} + +/** + * `supported: true` means the lifecycle command was acknowledged, not that the + * target state has been reached. Callers must observe/reconcile the session. + * Failures throw; they must never be disguised as an unsupported capability. + */ +export type SessionLifecycleResult = + | { readonly supported: false } + | { readonly supported: true }; + +/** Optional evidence from best-effort cleanup. A request is not confirmed teardown. */ +export type SessionStopResult = + | { readonly outcome: 'requested' | 'not-found' | 'terminated' } + | { readonly outcome: 'unconfirmed'; readonly error_type: string; readonly aws_request_id?: string }; export interface ComputeStrategy { readonly type: ComputeType; @@ -120,9 +147,13 @@ export interface ComputeStrategy { * the build def (never worse than today). */ readOnly?: boolean; + /** Coordinator-owned checkpoint recovery pins the original worker image. */ + microvmImage?: { readonly imageArn: string; readonly imageVersion: string }; }): Promise; - pollSession(handle: SessionHandle): Promise; - stopSession(handle: SessionHandle): Promise; + pollSession(handle: SessionHandle, options?: SessionControlOptions): Promise; + stopSession(handle: SessionHandle, options?: SessionControlOptions): Promise; + suspendSession(handle: SessionHandle, options?: SessionControlOptions): Promise; + resumeSession(handle: SessionHandle, options?: SessionControlOptions): Promise; } export function resolveComputeStrategy(blueprintConfig: BlueprintConfig): ComputeStrategy { diff --git a/cdk/src/handlers/shared/create-task-core.ts b/cdk/src/handlers/shared/create-task-core.ts index de9a92135..e08e23ab6 100644 --- a/cdk/src/handlers/shared/create-task-core.ts +++ b/cdk/src/handlers/shared/create-task-core.ts @@ -48,6 +48,9 @@ import { type CreateTaskRequest, createAttachmentRecord, INITIAL_APPROVALS_MAX_ENTRIES, + MICROVM_SLEEP_AFTER_S_DEFAULT, + MICROVM_SLEEP_AFTER_S_MAX, + MICROVM_SLEEP_AFTER_S_MIN, type InlineAttachment, type PresignedAttachment, type TaskRecord, @@ -309,9 +312,8 @@ export async function createTaskCore( } // Cedar HITL — validate approval_timeout_s if supplied (§7.3 step 5). - // maxLifetime-based ceiling clip is applied at orchestrator - // invocation time; at submit time we only enforce the `[floor, cap]` - // envelope. + // Zero retains unanswered requests. Positive explicit deadlines use the + // supported range; the worker's resource lifetime is managed separately. let approvalTimeoutS: number | undefined; if (body.approval_timeout_s !== undefined) { if (typeof body.approval_timeout_s !== 'number' @@ -323,12 +325,12 @@ export async function createTaskCore( requestId, ); } - if (body.approval_timeout_s < APPROVAL_TIMEOUT_S_MIN + if ((body.approval_timeout_s !== 0 && body.approval_timeout_s < APPROVAL_TIMEOUT_S_MIN) || body.approval_timeout_s > APPROVAL_TIMEOUT_S_MAX) { return errorResponse( 400, ErrorCode.VALIDATION_ERROR, - `Invalid approval_timeout_s. Must be between ${APPROVAL_TIMEOUT_S_MIN}s ` + `Invalid approval_timeout_s. Must be 0 (no automatic expiry), or between ${APPROVAL_TIMEOUT_S_MIN}s ` + `and ${APPROVAL_TIMEOUT_S_MAX}s.`, requestId, ); @@ -336,6 +338,15 @@ export async function createTaskCore( approvalTimeoutS = body.approval_timeout_s; } + const microvmSleepAfterS = body.microvm_sleep_after_s === undefined + ? MICROVM_SLEEP_AFTER_S_DEFAULT : body.microvm_sleep_after_s; + if (typeof microvmSleepAfterS !== 'number' || !Number.isInteger(microvmSleepAfterS) + || microvmSleepAfterS < MICROVM_SLEEP_AFTER_S_MIN || microvmSleepAfterS > MICROVM_SLEEP_AFTER_S_MAX) { + return errorResponse(400, ErrorCode.VALIDATION_ERROR, + `Invalid microvm_sleep_after_s. Must be an integer between ${MICROVM_SLEEP_AFTER_S_MIN} ` + + `and ${MICROVM_SLEEP_AFTER_S_MAX} seconds (0 disables sleep).`, requestId); + } + // Cedar HITL — validate initial_approvals if supplied (§7.3 step 4). let initialApprovals: string[] | undefined; if (body.initial_approvals !== undefined) { @@ -811,6 +822,8 @@ export async function createTaskCore( // payload supplied them; ``approval_timeout_s`` defaults to the // engine default at agent runtime when absent here. ...(approvalTimeoutS !== undefined && { approval_timeout_s: approvalTimeoutS }), + // Capture the default so future deployments cannot change this task's preference. + microvm_sleep_after_s: microvmSleepAfterS, ...(initialApprovals !== undefined && { initial_approvals: initialApprovals }), // Persisted counter the stranded-approval reconciler + agent // counter both read (§13.6). Seeded to 0 at task-create time. diff --git a/cdk/src/handlers/shared/deny-reason-scanner.ts b/cdk/src/handlers/shared/deny-reason-scanner.ts index e38af68ed..b5f455038 100644 --- a/cdk/src/handlers/shared/deny-reason-scanner.ts +++ b/cdk/src/handlers/shared/deny-reason-scanner.ts @@ -33,8 +33,8 @@ * * Keep the pattern set in sync with `agent/src/output_scanner.py`; the * agent-side scanner is the canonical source for PostToolUse output - * redaction, and this module is the REST-side port for the deny-reason - * path only. If the agent-side patterns change, update here too. + * redaction, and this module also protects approval notification previews. + * If the agent-side patterns change, update here too. */ interface SecretPattern { diff --git a/cdk/src/handlers/shared/error-classifier.ts b/cdk/src/handlers/shared/error-classifier.ts index 8d6b29c14..dfe06bd5a 100644 --- a/cdk/src/handlers/shared/error-classifier.ts +++ b/cdk/src/handlers/shared/error-classifier.ts @@ -41,6 +41,11 @@ export const ErrorCategory = { export type ErrorCategoryType = (typeof ErrorCategory)[keyof typeof ErrorCategory]; +/** A start was sent, but its response does not establish whether AWS created a VM. */ +export class MicrovmStartUncertainError extends Error { + override name = 'MicrovmStartUncertainError'; +} + /** * WHO should act, and whether retrying the SAME request can help — the axis a * channel reader needs to answer "just retry, or tell my admin?". Distinct from @@ -90,7 +95,114 @@ interface ErrorPattern { readonly classification: ErrorClassification; } +/** Stable codes written by the orchestrator; diagnostic text cannot override them. */ +const MICROVM_TERMINAL_CLASSIFICATIONS: Readonly> = { + MICROVM_RESUME_HOOK_FAILED: { + category: ErrorCategory.COMPUTE, + title: 'The MicroVM could not wake after being paused', + description: + 'AWS could not complete the worker’s wake hook. An accepted wake request does not mean the worker resumed. The task may already have made changes before it paused.', + remedy: + 'An ABCA admin should inspect the task ID and MicroVM ID in the coordinator and approval API logs, including the AWS request ID, ' + + 'then check microvm_hook_started, microvm_hook_stage_failed and microvm_hook_finished in /aws/lambda-microvms/. ' + + 'A transport error can occur before any guest hook log. The service’s "connection was refused" wording alone does not establish that the listener was closed. ' + + 'Check saved task progress and cleanup before starting a replacement task; retrying alone is not a verified fix.', + retryable: false, + errorClass: ErrorClass.SERVICE, + }, + MICROVM_RUN_HOOK_REJECTED: { + category: ErrorCategory.CONFIG, + title: 'The MicroVM rejected its own run payload', + description: + 'The agent\'s /run hook rejected the request with an HTTP 4xx response before reporting a task result. Common causes are a malformed payload, invalid platform configuration, or mismatched orchestrator and image versions.', + remedy: + 'Retrying as-is will not help — the guest will reject the identical payload again. ' + + 'Read the agent\'s structured response in the MicroVM log group (/aws/lambda-microvms/) for the MICROVM_RUN_* code and any missing environment variables. ' + + 'MICROVM_RUN_PLATFORM_CONFIG_INCOMPLETE means the orchestrator predates the platform_config contract — an admin should redeploy the stack so the orchestrator and image versions match.', + retryable: false, + errorClass: ErrorClass.SERVICE, + }, + MICROVM_SUBSTRATE_TERMINATED: { + category: ErrorCategory.COMPUTE, + title: 'The MicroVM stopped before the agent reported a result', + description: + 'The MicroVM ended before the agent reported a task result. Any explanation supplied by AWS is preserved in the original error message for diagnosis.', + remedy: + 'This is a compute-substrate fault, not a problem with your request — reply here to try again. ' + + 'If the message names a lifecycle-hook HTTP status, check that named hook in the MicroVM log group for its structured diagnostic code. ' + + 'Otherwise check the MicroVM logs for the session and whether the task is exceeding the 8-hour session cap; ' + + 'a long-running repo may belong on --compute-type ecs.', + retryable: true, + errorClass: ErrorClass.TRANSIENT, + }, +}; +const MICROVM_TERMINAL_PREFIX = 'MicroVM substrate terminated before the agent wrote a terminal status: '; +const MICROVM_RUN_HOOK_4XX = /^Run lifecycle hook returned HTTP status 4\d{2}(?:\.|$)/i; +const MICROVM_RESUME_HOOK_FAILURE = /^Resume lifecycle hook (?:failed|timed out|connection was refused|returned HTTP status [45]\d{2})(?:\.|$)/i; + +function microvmTerminalCode(stateReason?: string): string { + if (stateReason && MICROVM_RUN_HOOK_4XX.test(stateReason)) return 'MICROVM_RUN_HOOK_REJECTED'; + if (stateReason && MICROVM_RESUME_HOOK_FAILURE.test(stateReason)) return 'MICROVM_RESUME_HOOK_FAILED'; + return 'MICROVM_SUBSTRATE_TERMINATED'; +} + +/** + * Preserve the service reason for diagnosis, independently of the persisted code. + * GetMicrovm exposes hook status only in stateReason. Recognize observed service + * message shapes at this boundary; unknown wording stays a generic terminal failure. + * The strategy itself continues to report state without applying health policy. + */ +export function formatMicrovmTerminalFailure(detail: string, stateReason?: string): string { + const code = microvmTerminalCode(stateReason); + const reason = stateReason ? ` (${stateReason})` : ''; + return `${code}: ${MICROVM_TERMINAL_PREFIX}${detail}${reason}`; +} + +/** Classify persisted MicroVM terminal failures, including records without a code. */ +export function classifyMicrovmTerminalFailure(errorMessage?: string | null): ErrorClassification | null { + if (!errorMessage) return null; + const code = /^(MICROVM_RUN_HOOK_REJECTED|MICROVM_RESUME_HOOK_FAILED|MICROVM_SUBSTRATE_TERMINATED): /.exec(errorMessage)?.[1]; + if (code) return MICROVM_TERMINAL_CLASSIFICATIONS[code]; + if (MICROVM_RUN_HOOK_4XX.test(errorMessage)) return MICROVM_TERMINAL_CLASSIFICATIONS.MICROVM_RUN_HOOK_REJECTED; + if (MICROVM_RESUME_HOOK_FAILURE.test(errorMessage)) return MICROVM_TERMINAL_CLASSIFICATIONS.MICROVM_RESUME_HOOK_FAILED; + + // Old records have only the descriptive prefix. Do not let other words in + // their service reason override the known terminal failure. + if (errorMessage.startsWith(MICROVM_TERMINAL_PREFIX)) { + const reason = errorMessage.slice(errorMessage.indexOf(' (') + 2).replace(/\)$/, ''); + return MICROVM_TERMINAL_CLASSIFICATIONS[microvmTerminalCode(reason)]; + } + return null; +} + const PATTERNS: readonly ErrorPattern[] = [ + { + // The supervisor persists this reason for a permanent read failure or after + // exhausting its bounded retry count. Neither outcome establishes whether + // the retained worker was healthy or whether cleanup has completed. + pattern: /^MicroVM supervisor: substrate-read-failed(?:-repeatedly)?$/, + classification: { + category: ErrorCategory.COMPUTE, + title: 'The MicroVM status could not be checked', + description: 'ABCA ended the task after its required worker-state checks failed.', + remedy: + 'An ABCA admin should find microvm_supervisor_request_failed in the coordinator logs using the task ID and MicroVM ID. ' + + 'Check stage, error_type and any AWS request ID. Confirm worker termination and review saved task progress before starting a replacement task.', + retryable: false, + errorClass: ErrorClass.SERVICE, + }, + }, + { + pattern: /MICROVM_START_(?:OUTCOME_UNKNOWN|INPUT_CHANGED|STATE_INVALID|TASK_CLOSED|RECEIPT_SAVE_FAILED):/, + classification: { + category: ErrorCategory.COMPUTE, + title: 'The MicroVM start requires reconciliation', + description: 'The platform cannot safely issue another start for this task.', + remedy: 'An admin should inspect the task\'s saved MicroVM start receipt and any recorded computer ID before submitting another task. An unknown start may still have created a MicroVM; do not assume that a missing response means it failed.', + retryable: false, + errorClass: ErrorClass.SERVICE, + }, + }, // --- Auth --- { pattern: /INSUFFICIENT_GITHUB_REPO_PERMISSIONS/i, @@ -205,7 +317,7 @@ const PATTERNS: readonly ErrorPattern[] = [ // anchored on a `lambda`/`microvm` marker precisely so an AgentCore or ECS // endpoint failure cannot be hijacked into MicroVM copy — their endpoint // hosts are `bedrock-agentcore.*` / `ecs.*`. - pattern: /(?:UnknownEndpoint|Inaccessible host|Could not resolve endpoint).{0,120}lambda|(?:lambda[- ]?microvms?|microvms?).{0,60}(?:not available|not supported|unavailable)|(?:not available|not supported|unavailable).{0,60}(?:lambda[- ]?microvms?)/i, + pattern: /(?:UnknownEndpoint|Inaccessible host|Could not resolve endpoint).{0,120}lambda|(?:lambda[- ]?microvms?|microvms?)\s+(?:service\s+)?(?:(?:is|are)\s+)?(?:not available|not supported|unavailable)\s+in\s+[^.\n]{0,50}\bregion\b/i, classification: { category: ErrorCategory.CONFIG, title: 'Lambda MicroVMs is not available in this Region', @@ -224,7 +336,7 @@ const PATTERNS: readonly ErrorPattern[] = [ }, // --- Lambda MicroVMs (ADR-021) --- // - // SCOPING: the three SDK-exception entries below are anchored on the + // SCOPING: the SDK-exception entries below are anchored on the // ``MicroVM failed`` marker that // ``lambda-microvm-strategy.wrapMicrovmError`` puts on every error it lets // escape (see ``MICROVM_ERROR_MARKER``). The anchor is mandatory, not @@ -237,9 +349,8 @@ const PATTERNS: readonly ErrorPattern[] = [ // classified them before this section existed (a precise earlier pattern, or // UNKNOWN). // - // Both orders are accepted in each pattern so a future wrapper that puts the - // exception name ahead of the marker still matches; ``[\s\S]`` rather than - // ``.`` because SDK messages can span lines. + // ``[\s\S]`` rather than ``.`` because SDK messages can span lines. The older + // quota/throttle/not-found patterns also accept the reverse wrapper order. // // ORDERING: this section sits immediately ABOVE the generic // `Session start failed` catch-all. That is only safe BECAUSE of the marker: @@ -249,66 +360,6 @@ const PATTERNS: readonly ErrorPattern[] = [ // the MICROVM_* env vars to check) instead of "Check AgentCore Runtime or ECS // cluster health" — advice that names the wrong substrate entirely. If the // marker anchor is ever dropped, these MUST move back below the catch-all. - { - // A LIFECYCLE-HOOK 4xx, which the substrate reports in `stateReason` and - // `reconcileMicrovmSubstrateState` appends to the persisted message. - // - // ORDERING: this MUST stay above the generic `MicroVM substrate terminated` - // entry below, which matches the same message (both strings travel in one - // `error_message`) and would otherwise win with `retryable: true`. - // - // WHY it is a distinct, NON-retryable entry: every 4xx the guest can answer is - // a config/envelope fault that an identical retry cannot fix — - // `MICROVM_RUN_PAYLOAD_INVALID` (bad envelope), - // `MICROVM_RUN_PLATFORM_CONFIG_INVALID` (bad `platform_config` value) and - // `MICROVM_RUN_PLATFORM_CONFIG_INCOMPLETE` (missing required key / version - // skew), all in `agent/src/server.py`'s `microvm_run`. The one genuinely - // retryable guest answer, `MICROVM_RUN_PAYLOAD_UNREADABLE`, is deliberately a - // **500**, which is why this pattern is scoped to `4\d\d` and a 5xx still falls - // through to the retryable entry below. Without this, a permanently-skewed - // deployment rendered as TRANSIENT with the remedy "reply here to try again", - // i.e. an invitation to loop forever. - // - // WHAT THE MESSAGE DOES *NOT* CARRY, measured: the guest's structured response - // BODY does not reach `stateReason`. Live evidence - // (`docs/verification/645-p2-smoke-runbook.md` §6.1) shows the service supplies - // exactly `"Run lifecycle hook returned HTTP status 400. Please check your hook - // endpoint and application logs for more details."` — the status code and - // nothing else. So this entry anchors on the STATUS, which is the only - // discriminator that actually travels; the `code` is available to the operator - // only in the guest log group, which is what the remedy sends them to. If a - // future service release ever enriches `stateReason` with the body, a - // code-anchored entry becomes possible and would be strictly better. - pattern: /Run lifecycle hook returned HTTP status 4\d\d/i, - classification: { - category: ErrorCategory.CONFIG, - title: 'The MicroVM rejected its own run payload', - description: - 'The Lambda MicroVMs service delivered this task to the in-guest agent\'s /run lifecycle hook and the agent answered 4xx, so the service reaped the MicroVM within ~12 s without the task ever starting. The agent refuses a run for exactly three reasons, all of them wiring faults rather than transient ones: the orchestrator built a malformed envelope, a platform_config value was invalid, or a required platform_config key was missing (a version-skewed orchestrator paired with a current image). The agent writes a structured reason to its log group before answering.', - remedy: - 'Retrying as-is will not help — the guest will reject the identical payload again. ' - + 'Read the agent\'s own structured response body in the MicroVM log group (/aws/lambda-microvms/): it carries a MICROVM_RUN_* code that names which of the three faults occurred, and for a missing-config fault it lists the exact environment variables. ' - + 'MICROVM_RUN_PLATFORM_CONFIG_INCOMPLETE means the orchestrator predates the platform_config contract — an admin should redeploy the stack so the orchestrator and image versions match.', - retryable: false, - errorClass: ErrorClass.SERVICE, - }, - }, - { - pattern: /MicroVM substrate terminated before the agent wrote a terminal status/i, - classification: { - category: ErrorCategory.COMPUTE, - title: 'The MicroVM stopped before the agent reported a result', - description: - 'The Lambda MicroVM running this task reached a terminal state while the task was still mid-flight, so no result was ever written. When the substrate supplied a reason (GetMicrovm\'s stateReason) the orchestrator appends it in parentheses on the message above — read that first, because it distinguishes the common causes (a /run lifecycle-hook 4xx, which the service reaps in ~12 s) from the rarer ones (the session duration cap, a host fault, or an external terminate).', - remedy: - 'This is a compute-substrate fault, not a problem with your request — reply here to try again. ' - + 'If the message names a lifecycle-hook HTTP status, the guest rejected the run: check the MicroVM log group for the agent\'s structured response body. ' - + 'Otherwise check the MicroVM logs for the session and whether the task is exceeding the 8-hour session cap; ' - + 'a long-running repo may belong on --compute-type ecs.', - retryable: true, - errorClass: ErrorClass.TRANSIENT, - }, - }, { // A `platform_config` block the orchestrator could not even assemble: a // REQUIRED key is absent from the orchestrator Lambda's own environment, so @@ -335,18 +386,14 @@ const PATTERNS: readonly ErrorPattern[] = [ }, }, { - // Account-level MicroVM memory quota. Deliberately TRANSIENT rather than - // SERVICE: this quota is capacity-shaped, not configuration-shaped — it - // frees as running/suspended MicroVMs terminate (AWS counts SUSPENDED VMs - // toward the quota), which is the same "wait and retry" character as the - // existing per-user `concurrency limit` entry above. The remedy still names - // the quota-increase path for the case where the ceiling is genuinely too - // low, so a persistently failing deployment is not left guessing. + // Account-level capacity may become available as other sessions terminate. + // Suspended-session quota accounting still needs live verification; do not + // promise that suspending a session frees capacity. pattern: /MicroVM [\w ]+failed[\s\S]*ServiceQuotaExceededException|ServiceQuotaExceededException[\s\S]*MicroVM [\w ]+failed/i, classification: { category: ErrorCategory.COMPUTE, title: 'Couldn\'t start — the MicroVM compute quota is currently exhausted', - description: 'Starting the MicroVM was rejected because the account\'s Lambda MicroVMs quota (memory across running and suspended MicroVMs) is fully consumed.', + description: 'Starting the MicroVM was rejected because the account\'s Lambda MicroVMs compute quota is fully consumed.', remedy: 'Wait for in-flight tasks to finish and retry — the quota frees as MicroVMs terminate. ' + 'If the platform hits this routinely, request a Lambda MicroVMs quota increase in Service Quotas, ' @@ -387,6 +434,50 @@ const PATTERNS: readonly ErrorPattern[] = [ }, }, + { + pattern: /MicroVM [\w ]+failed[\s\S]*(?:AccessDeniedException|UnauthorizedException)/i, + classification: { + category: ErrorCategory.AUTH, + title: 'The platform is not authorized to run this MicroVM', + description: 'AWS denied the MicroVM lifecycle request with the current permissions.', + remedy: 'An admin should check the orchestrator role, iam:PassRole, and the MicroVM execution role against the deployed bootstrap bundle and stack. Correct the permissions before retrying.', + retryable: false, + errorClass: ErrorClass.SERVICE, + }, + }, + { + pattern: /MicroVM [\w ]+failed[\s\S]*(?:ValidationException|InvalidParameterValueException)/i, + classification: { + category: ErrorCategory.CONFIG, + title: 'The MicroVM request contains invalid configuration', + description: 'AWS rejected a value in the MicroVM request.', + remedy: 'An admin should check the image identifier/version, connector ARNs, execution role and hook payload against the service contract, then redeploy the corrected configuration.', + retryable: false, + errorClass: ErrorClass.SERVICE, + }, + }, + { + pattern: /MicroVM[\s\S]*(?:host|capacity)[\s\S]*(?:unavailable|not available)|MicroVM[\s\S]*InsufficientCapacityException/i, + classification: { + category: ErrorCategory.COMPUTE, + title: 'The MicroVM host or capacity is temporarily unavailable', + description: 'AWS could not provide the host or capacity needed for this MicroVM.', + remedy: 'Retry the task. If the failure persists, an admin should check AWS service health and the available MicroVM capacity.', + retryable: true, + errorClass: ErrorClass.TRANSIENT, + }, + }, + { + pattern: /MicroVM [\w ]+failed[\s\S]*(?:TimeoutError|RequestTimeout|ECONNRESET|ETIMEDOUT|InternalServerException|ServiceException)/i, + classification: { + category: ErrorCategory.COMPUTE, + title: 'The MicroVM service response was interrupted', + description: 'The request timed out, the connection failed, or AWS returned a server error. The request may already have succeeded.', + remedy: 'The platform retries the same saved start request within its recovery window. If recovery fails, an admin should inspect the saved start receipt before creating a new task.', + retryable: true, + errorClass: ErrorClass.TRANSIENT, + }, + }, { pattern: /Session start failed/i, classification: { @@ -867,6 +958,10 @@ export function classifyError(errorMessage: string | undefined | null): ErrorCla return null; } + // Read the orchestrator's code before scanning any appended service text. + const microvmFailure = classifyMicrovmTerminalFailure(errorMessage); + if (microvmFailure) return microvmFailure; + // Environmental blockers carry a canonical ``BLOCKED[]`` prefix // and an extractable resource — check them first so the remedy can name the // exact secret / host rather than falling through to a generic pattern. diff --git a/cdk/src/handlers/shared/failure-reply.ts b/cdk/src/handlers/shared/failure-reply.ts index ff65ad963..3fe70d513 100644 --- a/cdk/src/handlers/shared/failure-reply.ts +++ b/cdk/src/handlers/shared/failure-reply.ts @@ -41,7 +41,7 @@ * iteration on the same PR. Pure and deterministic; no I/O. */ -import { classifyError, retryGuidance } from './error-classifier'; +import { classifyError, classifyMicrovmTerminalFailure, retryGuidance } from './error-classifier'; import type { TaskStatusType } from '../../constructs/task-status'; /** Max chars of the raw agent error surfaced inline (the rest is in CloudWatch). */ @@ -96,6 +96,7 @@ const BUILD_GATE_TIMEOUT_RE = /agent_status=['"]?(success|end_turn)['"]?.*build_ * that surfaces build_passed directly). */ function isBuildFailure(input: Pick): boolean { + if (classifyMicrovmTerminalFailure(input.errorMessage)) return false; if (input.errorMessage && BUILD_GATE_FAILED_RE.test(input.errorMessage)) { return true; } @@ -104,13 +105,14 @@ function isBuildFailure(input: Pick): boolean { + if (classifyMicrovmTerminalFailure(input.errorMessage)) return false; return !!input.errorMessage && BUILD_GATE_TIMEOUT_RE.test(input.errorMessage); } -/** Collapse whitespace + clip to EXCERPT_MAX chars with an ellipsis. Strips the - * internal `[auto-retried]` marker (it drives the guidance, not user-facing text). */ +/** Collapse whitespace and clip the raw detail; hide internal classification and retry markers. */ function excerpt(raw: string): string { - const oneLine = raw.replace(/\s*\[auto-retried\]\s*/gi, ' ').replace(/\s+/g, ' ').trim(); + const oneLine = raw.replace(/^MICROVM_(?:RUN_HOOK_REJECTED|SUBSTRATE_TERMINATED): /, '') + .replace(/\s*\[auto-retried\]\s*/gi, ' ').replace(/\s+/g, ' ').trim(); return oneLine.length > EXCERPT_MAX ? `${oneLine.slice(0, EXCERPT_MAX)}…` : oneLine; } @@ -165,6 +167,9 @@ export function renderFailureReply(input: FailureReplyInput): string { /** True when the orchestrator marked this failure as already auto-retried once. */ function wasAutoRetried(errorMessage?: string | null): boolean { + // A terminal-state observation is not a session-start retry. Any matching + // text in its AWS diagnostic reason must not claim that ABCA retried it. + if (classifyMicrovmTerminalFailure(errorMessage)) return false; return !!errorMessage && /\[auto-retried\]/i.test(errorMessage); } diff --git a/cdk/src/handlers/shared/linear-approval-reply.ts b/cdk/src/handlers/shared/linear-approval-reply.ts new file mode 100644 index 000000000..11274639a --- /dev/null +++ b/cdk/src/handlers/shared/linear-approval-reply.ts @@ -0,0 +1,129 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { GetCommand, type DynamoDBDocumentClient } from '@aws-sdk/lib-dynamodb'; +import { + closeLinearApprovalThread, linearApprovalCommentId, parseLinearApprovalReply, readLinearApprovalThread, +} from './linear-approval-thread'; +import { postIdentifiedComment, readLinearApprovalComment } from './linear-feedback'; +import { logger } from './logger'; + +interface ApprovalReplyEvent { + action: string; + organizationId?: string; + actor?: { id?: string }; + data: { id: string; body?: string; parentId?: string; issueId?: string; issue?: { id?: string } }; +} + +interface ApprovalReplyDependencies { + ddb: DynamoDBDocumentClient; + approvalsTable: string; + taskTable: string; + registryTable: string; + lookupUser: (workspaceId: string, actorId: string) => Promise; +} + +/** A verified webhook may decide only the exact gate bound to its thread. */ +export async function handleLinearApprovalReply( + event: ApprovalReplyEvent, deps: ApprovalReplyDependencies, +): Promise { + const decision = parseLinearApprovalReply(event.data.body); + const workspaceId = event.organizationId; + const parentId = event.data.parentId; + if (event.action !== 'create' || !decision || !workspaceId || !parentId || !event.data.id) return false; + const thread = await readLinearApprovalThread(deps.ddb, deps.approvalsTable, workspaceId, parentId); + if (!thread) return false; + const issueId = event.data.issueId ?? event.data.issue?.id; + if (issueId !== thread.issueId) return true; + const ctx = { linearWorkspaceId: workspaceId, registryTableName: deps.registryTable }; + const comment = await readLinearApprovalComment(ctx, event.data.id); + if (!comment || comment.id !== event.data.id || comment.botActor || !comment.user?.id + || comment.issue?.id !== issueId || comment.parent?.id !== parentId + || parseLinearApprovalReply(comment.body) !== decision) { + logger.warn('Linear approval webhook does not match a human comment', { + task_id: thread.taskId, request_id: thread.requestId, comment_id: event.data.id, + }); + return true; + } + const respond = async (body: string): Promise => { + const posted = await postIdentifiedComment(ctx, { + id: linearApprovalCommentId({ ...thread, requestId: `${thread.requestId}#reply#${event.data.id}` }), + issueId, + parentId, + body, + }); + if (!posted.ok) { + logger.warn('Linear approval acknowledgement failed', { + task_id: thread.taskId, request_id: thread.requestId, comment_id: event.data.id, retryable: posted.retryable, + }); + if (posted.retryable) throw new Error('Retryable Linear approval acknowledgement failure'); + } + }; + const userId = await deps.lookupUser(workspaceId, comment.user.id); + if (!userId || userId !== thread.userId) { + await respond('Only the task owner can decide this request. Link your Linear account to ABCA, then reply again.'); + return true; + } + const task = (await deps.ddb.send(new GetCommand({ + TableName: deps.taskTable, Key: { task_id: thread.taskId }, ConsistentRead: true, + }))).Item; + if (!task || task.user_id !== userId || task.channel_source !== 'linear' + || task.channel_metadata?.linear_workspace_id !== workspaceId + || task.channel_metadata?.linear_issue_id !== issueId) { + await respond('This approval is no longer available for this task. No decision was recorded.'); + return true; + } + const source = JSON.stringify(['linear', workspaceId, event.data.id]); + const expectedStatus = decision === 'approve' ? 'APPROVED' : 'DENIED'; + const readApproval = async () => (await deps.ddb.send(new GetCommand({ + TableName: deps.approvalsTable, Key: { task_id: thread.taskId, request_id: thread.requestId }, ConsistentRead: true, + }))).Item; + const alreadyRecorded = (row: Awaited>) => + row?.user_id === userId && row.status === expectedStatus && row.decision_source === source; + let recorded = alreadyRecorded(await readApproval()); + if (!recorded) { + // Loaded only for configured approval integrations; ordinary comment paths + // do not require the decision handlers' environment variables. + const record = decision === 'approve' + ? (await import('../approve-task.js')).recordApprovalForUser + : (await import('../deny-task.js')).recordDenialForUser; + const result = await record({ + userId, + taskId: thread.taskId, + decisionSource: source, + body: JSON.stringify({ request_id: thread.requestId, decision }), + }); + recorded = result.statusCode === 202 || alreadyRecorded(await readApproval()); + if (!recorded) { + if (result.statusCode >= 500) throw new Error(`Linear approval decision failed: HTTP ${result.statusCode}`); + await respond(result.statusCode === 429 + ? 'Too many approval decisions. Wait a minute, then send a new reply.' + : 'This request is already closed, expired, or no longer waiting. No new decision was recorded.'); + return true; + } + } + await closeLinearApprovalThread(deps.ddb, deps.approvalsTable, thread); + await respond(decision === 'approve' + ? 'Approved for this action once. The decision is saved; the agent will continue when its worker is ready.' + : 'Denied. The decision is saved; the agent will receive your denial when its worker is ready.'); + logger.info('Linear approval reply recorded', { + task_id: thread.taskId, request_id: thread.requestId, comment_id: event.data.id, decision, + }); + return true; +} diff --git a/cdk/src/handlers/shared/linear-approval-thread.ts b/cdk/src/handlers/shared/linear-approval-thread.ts new file mode 100644 index 000000000..d84006626 --- /dev/null +++ b/cdk/src/handlers/shared/linear-approval-thread.ts @@ -0,0 +1,100 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { createHash } from 'node:crypto'; +import { GetCommand, UpdateCommand, type DynamoDBDocumentClient } from '@aws-sdk/lib-dynamodb'; + +const RETENTION_DAYS = 90; +const RETENTION_SECONDS = RETENTION_DAYS * 24 * 60 * 60; + +export interface LinearApprovalThread { + readonly workspaceId: string; + readonly issueId: string; + readonly taskId: string; + readonly requestId: string; + readonly userId: string; +} + +/** Linear requires UUIDv4; stable hash-derived bits make comment retries address the same ID. */ +export function linearApprovalCommentId(thread: LinearApprovalThread): string { + const hash = createHash('sha256').update(JSON.stringify([ + 'abca-linear-approval-v1', thread.workspaceId, thread.issueId, thread.taskId, thread.requestId, + ])).digest('hex'); + const groups = hash.match(/^(.{8})(.{4}).(.{3}).(.{3})(.{12})/)!; + return `${groups[1]}-${groups[2]}-4${groups[3]}-a${groups[4]}-${groups[5]}`; +} + +function threadKey(workspaceId: string, commentId: string): { task_id: string; request_id: string } { + return { task_id: `LINEAR_COMMENT#${workspaceId}#${commentId}`, request_id: 'APPROVAL' }; +} + +/** Mapping is coordinator-owned: task-scoped worker credentials cannot write this key. */ +export async function saveLinearApprovalThread( + ddb: DynamoDBDocumentClient, tableName: string, thread: LinearApprovalThread, +): Promise { + const commentId = linearApprovalCommentId(thread); + await ddb.send(new UpdateCommand({ + TableName: tableName, + Key: threadKey(thread.workspaceId, commentId), + UpdateExpression: 'SET #kind = :kind, #thread = :thread', + ConditionExpression: 'attribute_not_exists(task_id) OR (#kind = :kind AND #thread = :thread)', + ExpressionAttributeNames: { '#kind': 'kind', '#thread': 'thread' }, + ExpressionAttributeValues: { ':kind': 'linear_approval_thread', ':thread': thread }, + })); + return commentId; +} + +export async function readLinearApprovalThread( + ddb: DynamoDBDocumentClient, tableName: string, workspaceId: string, commentId: string, +): Promise { + const result = await ddb.send(new GetCommand({ + TableName: tableName, Key: threadKey(workspaceId, commentId), ConsistentRead: true, + })); + const thread = result.Item?.thread as LinearApprovalThread | undefined; + if (result.Item?.kind !== 'linear_approval_thread' || !thread + || ![thread.workspaceId, thread.issueId, thread.taskId, thread.requestId, thread.userId] + .every(value => typeof value === 'string' && value.length > 0) + || thread.workspaceId !== workspaceId || linearApprovalCommentId(thread) !== commentId) return null; + return thread; +} + +/** Pending mappings have no TTL, like pending approvals; closure starts retention. */ +export async function closeLinearApprovalThread( + ddb: DynamoDBDocumentClient, tableName: string, thread: LinearApprovalThread, +): Promise { + try { + await ddb.send(new UpdateCommand({ + TableName: tableName, + Key: threadKey(thread.workspaceId, linearApprovalCommentId(thread)), + UpdateExpression: 'SET #ttl = if_not_exists(#ttl, :ttl)', + ConditionExpression: 'attribute_exists(task_id)', + ExpressionAttributeNames: { '#ttl': 'ttl' }, + ExpressionAttributeValues: { ':ttl': Math.floor(Date.now() / 1000) + RETENTION_SECONDS }, + })); + } catch (error) { + if ((error as { name?: string }).name !== 'ConditionalCheckFailedException') throw error; + } +} + +/** Only an explicit answer in an approval thread is a decision, never prose or an edit. */ +export function parseLinearApprovalReply(body: unknown): 'approve' | 'deny' | null { + if (typeof body !== 'string') return null; + const match = /^(approve|deny)[.!]?$/i.exec(body.trim()); + return match ? match[1]!.toLowerCase() as 'approve' | 'deny' : null; +} diff --git a/cdk/src/handlers/shared/linear-feedback.ts b/cdk/src/handlers/shared/linear-feedback.ts index 83de57b4d..6b88cc06c 100644 --- a/cdk/src/handlers/shared/linear-feedback.ts +++ b/cdk/src/handlers/shared/linear-feedback.ts @@ -408,6 +408,75 @@ export async function postIssueComment( return graphqlRequest(token, COMMENT_CREATE_MUTATION, { issueId, body }); } +/** + * Read consent from Linear itself. Webhook authentication alone is insufficient: + * legacy worker credentials can read the OAuth bundle containing the HMAC key. + * Lookup failures must retry, never become permission to use webhook fields. + */ +export async function readLinearApprovalComment( + ctx: LinearFeedbackContext, + id: string, +): Promise<{ + id: string; + body: string; + user?: { id: string } | null; + botActor?: { id: string } | null; + issue: { id: string }; + parent?: { id: string } | null; +} | null> { + const token = await resolveToken(ctx); + if (!token) throw new Error('Linear approval verification token unavailable'); + const result = await graphqlData(token, ` + query VerifyApprovalComment($id: String!) { + organization { id } + viewer { id } + comment(id: $id) { id body user { id } botActor { id } issue { id } parent { id } } + }`, { id }); + if (!result.ok) throw new Error('Linear approval comment verification unavailable'); + if ((result.value.organization as { id?: string } | undefined)?.id !== ctx.linearWorkspaceId) { + throw new Error('Linear approval verification workspace mismatch'); + } + const viewerId = (result.value.viewer as { id?: string } | undefined)?.id; + if (!viewerId) throw new Error('Linear approval verification identity unavailable'); + const comment = result.value.comment as Awaited> ?? null; + // Diagnostic user-mode OAuth tokens can post genuine human comments. Workers + // holding that token must not manufacture consent from its own identity. + // Read the identity from Linear, including for installations predating this check. + if (comment?.user?.id === viewerId) return null; + return comment; +} + +/** Retry-safe posting for approval prompts and acknowledgements. */ +export async function postIdentifiedComment( + ctx: LinearFeedbackContext, + input: { id: string; issueId: string; body: string; parentId?: string }, +): Promise { + const token = await resolveToken(ctx); + if (!token) return { ok: false, retryable: false }; + const created = await graphqlData(token, ` + mutation ApprovalComment($input: CommentCreateInput!) { + commentCreate(input: $input) { success comment { id } } + }`, { input }); + if (created.ok && (created.value.commentCreate as { success?: boolean } | undefined)?.success) { + return { ok: true }; + } + // A successful write can lose its response. A duplicate ID is acceptable only + // when destination and content match, ignoring transport line endings only. + const existing = await graphqlData(token, ` + query ApprovalComment($id: String!) { + comment(id: $id) { body issue { id } parent { id } } + }`, { id: input.id }); + if (!existing.ok) return { ok: false, retryable: true }; + const comment = existing.value.comment as { + body?: string; issue?: { id?: string }; parent?: { id?: string }; + } | undefined; + const normalizeBody = (body: string) => body.replace(/\r\n?/g, '\n').replace(/\n+$/, ''); + return typeof comment?.body === 'string' && normalizeBody(comment.body) === normalizeBody(input.body) + && comment.issue?.id === input.issueId + && comment.parent?.id === input.parentId + ? { ok: true } : { ok: false, retryable: false }; +} + /** * Upsert the orchestration's live status block — ONE comment on the parent epic * that is rewritten as the run progresses, rather than a new comment per diff --git a/cdk/src/handlers/shared/microvm-approval-wake.ts b/cdk/src/handlers/shared/microvm-approval-wake.ts new file mode 100644 index 000000000..1f5350e3d --- /dev/null +++ b/cdk/src/handlers/shared/microvm-approval-wake.ts @@ -0,0 +1,178 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +// SPDX-License-Identifier: MIT-0 + +import { GetMicrovmCommand, LambdaMicrovmsClient, ResumeMicrovmCommand } from '@aws-sdk/client-lambda-microvms'; +import type { Context } from 'aws-lambda'; +import type { SessionControlOptions } from './compute-strategy'; +import { logger } from './logger'; +import { dispatchMicrovmContinuation } from './microvm-continuation-dispatch'; +import { microvmErrorIdentity, microvmRequestIdentity } from './microvm-control'; +import { readMicrovmLifecycleSnapshot, saveMicrovmLifecycleIntent, type MicrovmLifecycleSnapshot } from './microvm-lifecycle'; +import { makeClient } from './ua'; +import { TaskStatus } from '../../constructs/task-status'; + +const API_TIMEOUT_MS = 15_000; +const RESPONSE_RESERVE_MS = 1_000; +export const APPROVAL_POST_COMMIT_TIMEOUT_MS = 8_000; +export const APPROVAL_AUDIT_TIMEOUT_MS = 2_000; +let client: LambdaMicrovmsClient | undefined; + +/** Reserve time to send 202 after a committed decision, including slow prior work. */ +export function approvalPostCommitOptions( + invocationStartedMs: number, context?: Pick, +): SessionControlOptions { + const remaining = context?.getRemainingTimeInMillis() ?? API_TIMEOUT_MS - (Date.now() - invocationStartedMs); + const budget = Math.floor(Math.min(APPROVAL_POST_COMMIT_TIMEOUT_MS, remaining - RESPONSE_RESERVE_MS)); + return { + abortSignal: Number.isSafeInteger(budget) && budget > 0 + ? AbortSignal.timeout(budget) : AbortSignal.abort(new Error('Approval post-commit budget exhausted')), + }; +} + +interface ApprovalWakeInput { + readonly taskId: string; + readonly userId: string; + readonly requestId: string; + readonly decision: 'APPROVED' | 'DENIED'; + readonly options: SessionControlOptions; + readonly emitEvent: (eventType: string, metadata: Record, options: SessionControlOptions) => Promise; +} + +function relevant(snapshot: MicrovmLifecycleSnapshot, input: ApprovalWakeInput): boolean { + if (snapshot.status === TaskStatus.AWAITING_APPROVAL) { + return snapshot.requestId === input.requestId + && (snapshot.approval.kind !== 'present' || snapshot.approval.status === input.decision); + } + // The guest may already have consumed this decision while an earlier Suspend + // is still in flight. Fence only this decision's old intent, never a new gate. + return snapshot.status === TaskStatus.RUNNING && snapshot.requestId === null + && snapshot.intent?.request_id === input.requestId; +} + +/** + * Optional latency improvement after the approval transaction commits. The + * durable supervisor remains responsible for retries and observed RUNNING. + * A retired checkpoint is dispatched to its original coordinator version. + * This path cannot suspend, terminate or rewrite the human decision. + */ +export async function wakeMicrovmAfterApproval(input: ApprovalWakeInput): Promise { + const options = { abortSignal: input.options.abortSignal ?? AbortSignal.timeout(APPROVAL_POST_COMMIT_TIMEOUT_MS) }; + let stage = 'task-read'; + let microvmId: string | undefined; + const report = async (reason: string, error?: unknown) => { + const metadata = { + task_id: input.taskId, + request_id: input.requestId, + ...(microvmId && { microvm_id: microvmId }), + stage, + reason, + ...(error !== undefined && microvmErrorIdentity(error)), + }; + logger.warn('MicroVM approval wake deferred to durable supervisor', metadata); + try { + const abortSignal = AbortSignal.any([options.abortSignal, AbortSignal.timeout(APPROVAL_AUDIT_TIMEOUT_MS)]); + abortSignal.throwIfAborted(); + await input.emitEvent('microvm_resume_orphan', metadata, { abortSignal }); + } catch (auditError) { + logger.warn('MicroVM resume audit failed after decision commit', { ...metadata, ...microvmErrorIdentity(auditError) }); + } + }; + try { + options.abortSignal?.throwIfAborted(); + if (await dispatchMicrovmContinuation(input.taskId, input.userId, input.requestId, options)) return; + const snapshot = await readMicrovmLifecycleSnapshot(input.taskId, input.userId, options); + options.abortSignal.throwIfAborted(); + if (!snapshot) return; + microvmId = snapshot.handle.microvmId; + if (!relevant(snapshot, input)) { + await report('task-or-gate-changed'); + return; + } + stage = 'wake-intent'; + const saved = await saveMicrovmLifecycleIntent(snapshot, 'resume', Date.now(), options); + if (saved.status !== 'saved') { + await report(`intent-${saved.status}`); + return; + } + + // Save even if AWS still reports RUNNING/SUSPENDING. A delayed Suspend must + // not strand the worker after the decision handler has returned. + stage = 'substrate-read'; + options.abortSignal?.throwIfAborted(); + client ??= makeClient(LambdaMicrovmsClient); + const observed = await client.send(new GetMicrovmCommand({ microvmIdentifier: microvmId }), options); + options.abortSignal?.throwIfAborted(); + logger.info('MicroVM observed after approval decision', { + task_id: input.taskId, + request_id: input.requestId, + microvm_id: microvmId, + observed_state: observed.state, + generation: saved.intent.generation, + intent_requested_at_ms: saved.intent.requested_at_ms, + ...microvmRequestIdentity(observed), + }); + if (observed.state !== 'SUSPENDED') { + if (observed.state !== 'RUNNING' && observed.state !== 'SUSPENDING') await report('state-not-resumable'); + return; + } + + stage = 'pre-resume-read'; + const current = await readMicrovmLifecycleSnapshot(input.taskId, input.userId, options); + options.abortSignal?.throwIfAborted(); + if (!current || current.handle.microvmId !== microvmId || current.handle.endpoint !== snapshot.handle.endpoint + || current.handle.sessionId !== snapshot.handle.sessionId + || current.requestId !== snapshot.requestId || current.status !== snapshot.status + || current.intent?.generation !== saved.intent.generation + || (current.status === TaskStatus.AWAITING_APPROVAL && !relevant(current, input))) { + await report('worker-or-gate-changed-before-resume'); + return; + } + + stage = 'resume-request'; + const startedAt = Date.now(); + const diagnostic = { + task_id: input.taskId, + request_id: input.requestId, + microvm_id: microvmId, + generation: saved.intent.generation, + intent_requested_at_ms: saved.intent.requested_at_ms, + image_arn: snapshot.handle.imageArn, + image_version: snapshot.handle.imageVersion, + }; + logger.info('MicroVM wake request started after approval decision', diagnostic); + try { + const response = await client.send(new ResumeMicrovmCommand({ microvmIdentifier: microvmId }), options); + logger.info('MicroVM wake requested after approval decision', { + ...diagnostic, elapsed_ms: Date.now() - startedAt, ...microvmRequestIdentity(response), + }); + } catch (error) { + await report('resume-request-failed', error); + } finally { + // Check after both acknowledged and uncertain outcomes. No next action + // follows a cancellation, changed gate or ownership loss in this handler. + stage = 'post-resume-read'; + options.abortSignal?.throwIfAborted(); + await readMicrovmLifecycleSnapshot(input.taskId, input.userId, options); + } + } catch (error) { + await report('wake-reconciliation-failed', error); + } +} diff --git a/cdk/src/handlers/shared/microvm-continuation-dispatch.ts b/cdk/src/handlers/shared/microvm-continuation-dispatch.ts new file mode 100644 index 000000000..689a845fa --- /dev/null +++ b/cdk/src/handlers/shared/microvm-continuation-dispatch.ts @@ -0,0 +1,110 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { createHash } from 'node:crypto'; +import { InvokeCommand, LambdaClient } from '@aws-sdk/client-lambda'; +import { GetCommand } from '@aws-sdk/lib-dynamodb'; +import type { SessionControlOptions } from './compute-strategy'; +import { logger } from './logger'; +import type { ContinuableTask } from './microvm-continuation-retirement'; +import type { MicrovmContinuationEvent } from './microvm-continuation-runner'; +import { admitContinuation } from './microvm-continuation-start'; +import { CONTINUATION_IO_TIMEOUT_MS } from './microvm-continuation-timing'; +import { validAttemptId } from './microvm-continuation-types'; +import { microvmErrorIdentity } from './microvm-control'; +import { continuationEnabled } from './microvm-worker-lease'; +import { makeClient, makeDocClient } from './ua'; +import { TaskStatus } from '../../constructs/task-status'; + +const ddb = makeDocClient(); +const MAX_VALIDATION_DETAIL_LENGTH = 512; +let lambda: LambdaClient | undefined; + +/** + * True means this is a retired/restoring worker; callers must not /resume its + * source handle. Admission and invocation can be retried by the periodic scan. + */ +export async function dispatchMicrovmContinuation( + taskId: string, userId: string, requestId: string, + options: SessionControlOptions = { abortSignal: AbortSignal.timeout(CONTINUATION_IO_TIMEOUT_MS) }, +): Promise { + if (!continuationEnabled()) return false; + const response = await ddb.send(new GetCommand({ + TableName: process.env.TASK_TABLE_NAME!, Key: { task_id: taskId }, ConsistentRead: true, + }), options); + const task = response.Item as ContinuableTask | undefined; + if (!task || task.user_id !== userId || task.compute_type !== 'lambda-microvm' + || task.status !== TaskStatus.AWAITING_APPROVAL || task.awaiting_approval_request_id !== requestId) return false; + if (!['FENCED', 'PARKED', 'STARTING', 'RESTORING'].includes(task.continuation?.state ?? '')) return false; + // Retirement still owns termination and capacity release. + if (task.continuation?.state === 'FENCED') return true; + const version = task.continuation_launch?.orchestrator_version; + const configuredArn = process.env.ORCHESTRATOR_FUNCTION_ARN ?? ''; + const match = /^(arn:[^:]+:lambda:[^:]+:\d+:function:[^:]+)(?::[^:]+)?$/.exec(configuredArn); + if (!version || !/^\d+$/.test(version) || !match) { + throw new Error('MICROVM_CONTINUATION_COORDINATOR_INVALID: original published function is unavailable'); + } + const admission = await admitContinuation( + taskId, userId, requestId, Number(process.env.MAX_CONCURRENT_TASKS_PER_USER ?? '10'), options, + ); + if (admission.kind !== 'ready') return true; + const attemptId = admission.task.continuation?.attempt_id; + if (!validAttemptId(attemptId)) throw new Error('MICROVM_CONTINUATION_ASSIGNMENT_INVALID'); + const event: MicrovmContinuationEvent = { + task_id: taskId, continuation_request_id: requestId, continuation_attempt_id: attemptId, + }; + const payload = JSON.stringify(event); + // Lambda accepts at most 64 characters. Keep the complete 256-bit identity; + // a descriptive prefix would exceed the service limit and strand admission. + const name = createHash('sha256').update(payload).digest('hex'); + lambda ??= makeClient(LambdaClient); + try { + const result = await lambda.send(new InvokeCommand({ + FunctionName: `${match[1]}:${version}`, + InvocationType: 'Event', + DurableExecutionName: name, + Payload: Buffer.from(payload), + }), options); + if (result.StatusCode !== 202) throw new Error('MICROVM_CONTINUATION_DISPATCH_FAILED: invocation was not accepted'); + logger.info('Saved task continuation dispatched', { + task_id: taskId, + request_id: requestId, + attempt_id: attemptId, + coordinator_version: version, + durable_execution_name: name, + }); + } catch (error) { + logger.warn('Saved task continuation dispatch needs reconciliation', { + task_id: taskId, + request_id: requestId, + attempt_id: attemptId, + operation: 'Invoke', + coordinator_version: version, + durable_execution_name_length: name.length, + // This Invoke carries only task/request/attempt identifiers, never a + // prompt, credential or signed URL. Its validation message is safe and + // necessary to diagnose errors that a class name alone cannot explain. + ...(error instanceof Error && error.name === 'ValidationException' + ? { validation_detail: error.message.slice(0, MAX_VALIDATION_DETAIL_LENGTH) } : {}), + ...microvmErrorIdentity(error), + }); + throw error; + } + return true; +} diff --git a/cdk/src/handlers/shared/microvm-continuation-retirement.ts b/cdk/src/handlers/shared/microvm-continuation-retirement.ts new file mode 100644 index 000000000..c87a9af05 --- /dev/null +++ b/cdk/src/handlers/shared/microvm-continuation-retirement.ts @@ -0,0 +1,300 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { randomUUID } from 'node:crypto'; +import { GetCommand, TransactWriteCommand } from '@aws-sdk/lib-dynamodb'; +import { canonicalJson } from './canonical-json'; +import type { ComputeStrategy, SessionControlOptions } from './compute-strategy'; +import { logger } from './logger'; +import { verifyContinuationCheckpoint } from './microvm-continuation-storage'; +import { CONTINUATION_RETIREMENT_TIMEOUT_MS, CONTINUATION_STOP_TIMEOUT_MS } from './microvm-continuation-timing'; +import { CONTINUATION, type ContinuationRecord, type MicrovmHandle, type WorkerLease, validateContinuation, workerLeaseKey } from './microvm-continuation-types'; +import { continuationEnabled } from './microvm-worker-lease'; +import { MICROVM_SLEEP_AFTER_S_DEFAULT, type TaskRecord } from './types'; +import { makeDocClient } from './ua'; +import { TaskStatus } from '../../constructs/task-status'; + +const ddb = makeDocClient(); +const TABLE = process.env.TASK_TABLE_NAME!; +const APPROVALS = process.env.TASK_APPROVALS_TABLE_NAME!; +const COUNTERS = process.env.USER_CONCURRENCY_TABLE_NAME!; + +interface HeldSlot { + readonly state: 'held' | 'released'; + readonly acquired_at: string; + readonly released_at?: string; + readonly attempt_id?: string; +} + +export type ContinuableTask = TaskRecord & { + readonly concurrency_slot?: HeldSlot; + readonly microvm_start?: { readonly clientToken: string; readonly handle?: MicrovmHandle; readonly createdAt?: string }; +}; + +export type RetirementResult = 'not-due' | 'stopping' | 'parked' | 'ownership-lost'; + +async function readTask(taskId: string, options: SessionControlOptions): Promise { + const result = await ddb.send(new GetCommand({ + TableName: TABLE, Key: { task_id: taskId }, ConsistentRead: true, + }), options); + return result.Item as ContinuableTask | undefined; +} + +function sameRecord(left: unknown, right: unknown): boolean { + return canonicalJson(left) === canonicalJson(right); +} + +async function leaseMatches( + task: ContinuableTask, handle: MicrovmHandle, state: 'FENCED' | 'PARKED', options: SessionControlOptions, +): Promise { + if (!task.microvm_start?.clientToken) return false; + const result = await ddb.send(new GetCommand({ + TableName: TABLE, Key: workerLeaseKey(task.task_id), ConsistentRead: true, + }), options); + const lease = result.Item as WorkerLease | undefined; + return lease?.lease_state === state && lease.lease_attempt_id === task.microvm_start.clientToken + && lease.lease_microvm_id === handle.microvmId && lease.lease_user_id === task.user_id + && lease.lease_repo === (task.repo ?? ''); +} + +async function fence(task: ContinuableTask, handle: MicrovmHandle, options: SessionControlOptions): Promise { + const record = task.continuation!; + const fenced: ContinuationRecord = { ...record, state: 'FENCED', source_handle: handle }; + const attempt = task.microvm_start?.clientToken; + if (!attempt || record.identity.attempt_id !== handle.microvmId) { + throw new Error('MICROVM_CONTINUATION_INVALID: source worker has no matching launch receipt'); + } + await verifyContinuationCheckpoint(record, options); + try { + await ddb.send(new TransactWriteCommand({ + TransactItems: [ + { + Update: { + TableName: TABLE, + Key: { task_id: task.task_id }, + UpdateExpression: 'SET continuation = :fenced', + ConditionExpression: 'user_id = :user AND #status = :awaiting AND session_id = :vm ' + + 'AND awaiting_approval_request_id = :request AND continuation = :record', + ExpressionAttributeNames: { '#status': 'status' }, + ExpressionAttributeValues: { + ':user': task.user_id, + ':awaiting': TaskStatus.AWAITING_APPROVAL, + ':vm': handle.microvmId, + ':request': record.identity.request_id, + ':record': record, + ':fenced': fenced, + }, + }, + }, + { + Update: { + TableName: TABLE, + Key: workerLeaseKey(task.task_id), + UpdateExpression: 'SET lease_state = :fenced', + ConditionExpression: 'lease_state = :active AND lease_attempt_id = :attempt ' + + 'AND lease_microvm_id = :vm AND lease_user_id = :user', + ExpressionAttributeValues: { + ':fenced': 'FENCED', + ':active': 'ACTIVE', + ':attempt': attempt, + ':vm': handle.microvmId, + ':user': task.user_id, + }, + }, + }, + { + ConditionCheck: { + TableName: APPROVALS, + Key: { task_id: task.task_id, request_id: record.identity.request_id }, + ConditionExpression: 'user_id = :user AND #status IN (:pending, :approved, :denied, :timedout)', + ExpressionAttributeNames: { '#status': 'status' }, + ExpressionAttributeValues: { + ':user': task.user_id, + ':pending': 'PENDING', + ':approved': 'APPROVED', + ':denied': 'DENIED', + ':timedout': 'TIMED_OUT', + }, + }, + }, + ], + }), options); + return fenced; + } catch (error) { + const latest = await readTask(task.task_id, options); + if (latest?.user_id === task.user_id && latest.status === TaskStatus.AWAITING_APPROVAL + && sameRecord(latest.continuation, fenced) + && await leaseMatches(latest, handle, 'FENCED', options)) return fenced; + // Original worker resumed or cancellation won before the fence. + if (!latest || latest.status !== TaskStatus.AWAITING_APPROVAL + || !sameRecord(latest.continuation, record)) return undefined; + throw error; + } +} + +/** Close the old reservation only after GetMicrovm confirmed terminal/not-found. */ +async function parkAfterTermination(task: ContinuableTask, record: ContinuationRecord, options: SessionControlOptions): Promise { + if (task.concurrency_slot?.state !== 'held' || !task.microvm_start?.clientToken) { + throw new Error('MICROVM_CONTINUATION_RESERVATION_INVALID: fenced worker has no held capacity'); + } + let emptyCounter = false; + for (let attempt = 0; attempt < 2; attempt++) { + const now = new Date().toISOString(); + const revision = randomUUID(); + const parked: ContinuationRecord = { ...record, state: 'PARKED', parked_at: now }; + try { + await ddb.send(new TransactWriteCommand({ + ClientRequestToken: revision, + TransactItems: [ + { + Update: { + TableName: TABLE, + Key: { task_id: task.task_id }, + UpdateExpression: 'SET continuation = :parked, concurrency_slot.#state = :released, ' + + 'concurrency_slot.released_at = :now REMOVE #ttl', + ConditionExpression: 'user_id = :user AND #status = :awaiting AND continuation = :record ' + + 'AND concurrency_slot = :slot', + ExpressionAttributeNames: { '#status': 'status', '#state': 'state', '#ttl': 'ttl' }, + ExpressionAttributeValues: { + ':user': task.user_id, + ':awaiting': TaskStatus.AWAITING_APPROVAL, + ':record': record, + ':slot': task.concurrency_slot, + ':parked': parked, + ':released': 'released', + ':now': now, + }, + }, + }, + { + Update: { + TableName: TABLE, + Key: workerLeaseKey(task.task_id), + UpdateExpression: 'SET lease_state = :parked', + ConditionExpression: 'lease_state = :fenced AND lease_attempt_id = :attempt AND lease_microvm_id = :vm', + ExpressionAttributeValues: { + ':parked': 'PARKED', + ':fenced': 'FENCED', + ':attempt': task.microvm_start.clientToken, + ':vm': record.source_handle!.microvmId, + }, + }, + }, + { + Update: { + TableName: COUNTERS, + Key: { user_id: task.user_id }, + UpdateExpression: emptyCounter + ? 'SET active_count = if_not_exists(active_count, :zero), updated_at = :now, reservation_version = :revision' + : 'SET active_count = active_count - :one, updated_at = :now, reservation_version = :revision', + ConditionExpression: emptyCounter + ? 'attribute_not_exists(active_count) OR active_count >= :zero' + : 'active_count > :zero', + ExpressionAttributeValues: { + ':zero': 0, ...(!emptyCounter && { ':one': 1 }), ':now': now, ':revision': revision, + }, + }, + }, + ], + }), options); + if (emptyCounter) { + logger.warn('Parked continuation whose capacity counter was already empty', { + task_id: task.task_id, error_id: 'CONCURRENCY_EMPTY_COUNTER', + }); + } + return true; + } catch (error) { + const latest = await readTask(task.task_id, options); + if (latest?.user_id === task.user_id && latest.continuation?.state === 'PARKED' + && latest.concurrency_slot?.state === 'released' + && sameRecord(latest.continuation.identity, record.identity) + && await leaseMatches(latest, record.source_handle!, 'PARKED', options)) return true; + if (!latest || latest.status !== TaskStatus.AWAITING_APPROVAL + || !sameRecord(latest.continuation, record)) return false; + const failure = error as { name?: string; CancellationReasons?: { Code?: string }[] }; + if (!emptyCounter && failure.name === 'TransactionCanceledException' + && failure.CancellationReasons?.[2]?.Code === 'ConditionalCheckFailed') { + emptyCounter = true; + continue; + } + throw error; + } + } + throw new Error('MICROVM_CONTINUATION_RESERVATION_INVALID: capacity release was not acknowledged'); +} + +/** + * One bounded retirement cycle. A worker is fenced before termination; its + * reservation remains held until the service confirms it cannot execute. + */ +export async function retireCheckpointedMicrovm(input: { + taskId: string; + userId: string; + handle: MicrovmHandle; + strategy: ComputeStrategy; + sessionDeadlineMs: number; + force?: boolean; + abortSignal?: AbortSignal; +}): Promise { + if (!continuationEnabled()) return 'not-due'; + const timeout = AbortSignal.timeout(CONTINUATION_RETIREMENT_TIMEOUT_MS); + const options = { abortSignal: input.abortSignal ? AbortSignal.any([input.abortSignal, timeout]) : timeout }; + const task = await readTask(input.taskId, options); + if (!task || task.user_id !== input.userId || task.session_id !== input.handle.microvmId) { + return 'ownership-lost'; + } + if (task.status !== TaskStatus.AWAITING_APPROVAL || !task.continuation) return 'not-due'; + const record = task.continuation; + validateContinuation(record, task); + if (record.state === 'PARKED' || record.state === 'FENCED') { + // Workers can publish the continuation attribute, but only the coordinator + // can advance the separate lease. Never treat a worker-written state label + // as proof that shutdown or capacity release actually happened. + if (!await leaseMatches(task, input.handle, record.state, options) + || (record.state === 'PARKED' && task.concurrency_slot?.state !== 'released')) { + throw new Error('MICROVM_CONTINUATION_LEASE_INVALID: retirement state has no coordinator authority'); + } + if (record.state === 'PARKED') return 'parked'; + } + if (record.state !== 'READY' && record.state !== 'FENCED') return 'not-due'; + if (record.state === 'READY') { + const approval = await ddb.send(new GetCommand({ + TableName: APPROVALS, + Key: { task_id: task.task_id, request_id: record.identity.request_id }, + ConsistentRead: true, + }), options); + const now = Date.now(); + const created = Date.parse(approval.Item?.created_at ?? ''); + const sleep = task.microvm_sleep_after_s ?? MICROVM_SLEEP_AFTER_S_DEFAULT; + const ageDue = sleep > 0 && Number.isFinite(created) && now - created >= CONTINUATION.park_after_seconds * 1000; + const lifetimeDue = Number.isFinite(input.sessionDeadlineMs) + && now >= input.sessionDeadlineMs - CONTINUATION.retirement_margin_seconds * 1000; + if (!input.force && !ageDue && !lifetimeDue) return 'not-due'; + } + const fenced = record.state === 'FENCED' ? record : await fence(task, input.handle, options); + if (!fenced) return 'not-due'; + if (fenced.source_handle?.microvmId !== input.handle.microvmId) { + throw new Error('MICROVM_CONTINUATION_INVALID: retirement handle does not match the fence'); + } + const signal = AbortSignal.any([options.abortSignal, AbortSignal.timeout(CONTINUATION_STOP_TIMEOUT_MS)]); + await input.strategy.stopSession(input.handle, { abortSignal: signal }); + const state = await input.strategy.pollSession(input.handle, { abortSignal: signal }); + if (state.microvmState !== 'TERMINATED' && state.microvmState !== 'NOT_FOUND') return 'stopping'; + return await parkAfterTermination(task, fenced, options) ? 'parked' : 'not-due'; +} diff --git a/cdk/src/handlers/shared/microvm-continuation-runner.ts b/cdk/src/handlers/shared/microvm-continuation-runner.ts new file mode 100644 index 000000000..8e53c1382 --- /dev/null +++ b/cdk/src/handlers/shared/microvm-continuation-runner.ts @@ -0,0 +1,283 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import type { DurableContext, WaitForConditionDecision } from '@aws/durable-execution-sdk-js'; +import { TransactWriteCommand } from '@aws-sdk/lib-dynamodb'; +import { resolveComputeStrategy, type ComputeStrategy } from './compute-strategy'; +import type { ContinuableTask } from './microvm-continuation-retirement'; +import { admitContinuation } from './microvm-continuation-start'; +import { loadContinuationLaunch } from './microvm-continuation-storage'; +import { CONTINUATION_IO_TIMEOUT_MS, CONTINUATION_POLL_INTERVAL_MS, CONTINUATION_RETRY_POLL_SECONDS, CONTINUATION_START_ATTEMPTS, CONTINUATION_TRANSITION_POLL_SECONDS } from './microvm-continuation-timing'; +import { type MicrovmHandle, validAttemptId, workerLeaseKey } from './microvm-continuation-types'; +import { microvmErrorIdentity } from './microvm-control'; +import { MICROVM_MAX_POLL_FAILURES, stopMicrovmWithDiagnostics } from './microvm-supervisor'; +import { pollMicrovmTask } from './microvm-task-poll'; +import { emitTaskEvent, envelopeFor, finalizeTask, loadTask, type PollState } from './orchestrator'; +import { deleteMicrovmPayload } from './strategies/lambda-microvm-strategy'; +import { makeDocClient } from './ua'; +import { TaskStatus, TERMINAL_STATUSES } from '../../constructs/task-status'; + +const ddb = makeDocClient(); +const TABLE = process.env.TASK_TABLE_NAME!; +const RESTORE_TIMEOUT_MS = 900_000; + +export interface MicrovmContinuationEvent { + readonly task_id: string; + readonly continuation_request_id: string; + readonly continuation_attempt_id: string; +} + +function sameAttempt(task: ContinuableTask, event: MicrovmContinuationEvent): boolean { + return task.microvm_start?.clientToken === event.continuation_attempt_id + || task.continuation?.attempt_id === event.continuation_attempt_id; +} + +/** + * Fence the assigned process while closing a failed restoration. A lost reply + * is resolved by a strong read; a different attempt is never changed. + */ +export async function failContinuationAttempt( + event: MicrovmContinuationEvent, userId: string, detail: string, +): Promise { + const task = await loadTask(event.task_id, true) as ContinuableTask; + if (task.user_id !== userId || !sameAttempt(task, event) || TERMINAL_STATUSES.includes(task.status)) return; + try { + await ddb.send(new TransactWriteCommand({ + TransactItems: [ + { + Update: { + TableName: TABLE, + Key: { task_id: event.task_id }, + UpdateExpression: 'SET #status = :failed, error_message = :detail, completed_at = :now, ' + + 'updated_at = :now, status_created_at = :statusTime', + ConditionExpression: 'user_id = :user AND #status = :status ' + + 'AND (continuation.attempt_id = :attempt OR microvm_start.clientToken = :attempt)', + ExpressionAttributeNames: { '#status': 'status' }, + ExpressionAttributeValues: { + ':failed': TaskStatus.FAILED, + ':detail': detail, + ':now': new Date().toISOString(), + ':statusTime': `FAILED#${new Date().toISOString()}`, + ':user': userId, + ':status': task.status, + ':attempt': event.continuation_attempt_id, + }, + }, + }, + { + Update: { + TableName: TABLE, + Key: workerLeaseKey(event.task_id), + UpdateExpression: 'SET lease_state = :fenced', + ConditionExpression: 'lease_user_id = :user AND lease_attempt_id = :attempt AND lease_state = :active', + ExpressionAttributeValues: { + ':user': userId, ':attempt': event.continuation_attempt_id, ':active': 'ACTIVE', ':fenced': 'FENCED', + }, + }, + }, + ], + })); + } catch (error) { + const latest = await loadTask(event.task_id, true) as ContinuableTask; + if (latest.user_id === userId && sameAttempt(latest, event) && !TERMINAL_STATUSES.includes(latest.status)) throw error; + } +} + +export interface RestoreState { + readonly deadlineMs: number; + readonly ready?: boolean; + readonly closed?: boolean; + readonly ownershipLost?: boolean; + readonly failure?: string; + readonly consecutivePollFailures?: number; +} + +/** Restoration has its own bounded startup window; it is not a /resume hook. */ +export async function pollContinuationRestore( + event: MicrovmContinuationEvent, userId: string, handle: MicrovmHandle, + strategy: ComputeStrategy, previous: RestoreState, +): Promise { + const task = await loadTask(event.task_id, true) as ContinuableTask; + if (task.user_id !== userId || task.session_id !== handle.microvmId + || task.compute_metadata?.microvmId !== handle.microvmId) return { ...previous, ownershipLost: true }; + if (TERMINAL_STATUSES.includes(task.status)) return { ...previous, closed: true }; + if (task.status === TaskStatus.RUNNING || task.status === TaskStatus.FINALIZING + || (task.status === TaskStatus.AWAITING_APPROVAL + && task.awaiting_approval_request_id !== event.continuation_request_id)) { + return { ...previous, ready: true }; + } + if (!sameAttempt(task, event) || task.continuation?.state !== 'RESTORING') { + return { ...previous, failure: 'MICROVM_CONTINUATION_ASSIGNMENT_CHANGED: restoration no longer owns its saved request' }; + } + if (Date.now() >= previous.deadlineMs) { + return { ...previous, failure: 'MICROVM_CONTINUATION_RESTORE_TIMEOUT: saved workspace and conversation were not restored within 15 minutes' }; + } + try { + const observed = await strategy.pollSession(handle, { abortSignal: AbortSignal.timeout(CONTINUATION_IO_TIMEOUT_MS) }); + if (observed.status === 'completed' || observed.status === 'failed') { + return { ...previous, failure: `MICROVM_CONTINUATION_WORKER_STOPPED: ${observed.reason ?? ('error' in observed ? observed.error : observed.status)}` }; + } + return { deadlineMs: previous.deadlineMs, consecutivePollFailures: 0 }; + } catch (error) { + const failures = (previous.consecutivePollFailures ?? 0) + 1; + return { + deadlineMs: previous.deadlineMs, + consecutivePollFailures: failures, + ...(failures >= MICROVM_MAX_POLL_FAILURES && { + failure: `MICROVM_CONTINUATION_POLL_FAILED: ${microvmErrorIdentity(error).error_type}`, + }), + }; + } +} + +export function continuationWaitStrategy(state: PollState): WaitForConditionDecision { + if (state.microvmParked || state.microvmOwnershipLost + || (state.lastStatus && TERMINAL_STATUSES.includes(state.lastStatus))) return { shouldContinue: false }; + if (state.microvmRetiring) { + return { + shouldContinue: true, delay: { seconds: state.microvmRetirementError ? CONTINUATION_RETRY_POLL_SECONDS : CONTINUATION_TRANSITION_POLL_SECONDS }, + }; + } + if (state.microvmFailureReason || state.sessionUnhealthy) return { shouldContinue: false }; + return { + shouldContinue: true, + delay: { seconds: Math.max(1, Math.ceil((state.microvmSupervisor?.nextPollInMs ?? CONTINUATION_POLL_INTERVAL_MS) / 1000)) }, + }; +} + +/** One durable execution for one assigned replacement; dispatch supplies a stable execution name. */ +export async function runMicrovmContinuation(event: MicrovmContinuationEvent, context: DurableContext): Promise { + if (![event.task_id, event.continuation_request_id, event.continuation_attempt_id].every(validAttemptId)) { + throw new Error('MICROVM_CONTINUATION_EVENT_INVALID: invalid task, request or worker attempt'); + } + const task = await context.step('continuation-task', async () => loadTask(event.task_id, true) as Promise); + if (!sameAttempt(task, event) || TERMINAL_STATUSES.includes(task.status)) return; + const { correlation, log } = envelopeFor(task); + const emit = (type: string, metadata: Record, options?: { abortSignal?: AbortSignal }) => + emitTaskEvent(task.task_id, type, metadata, correlation, options); + let handle: MicrovmHandle | undefined; + let strategy: ComputeStrategy | undefined; + try { + const launch = await context.step('continuation-launch-inputs', async () => { + if (!task.continuation_launch) throw new Error('MICROVM_CONTINUATION_INPUT_INVALID: saved launch is missing'); + const saved = await loadContinuationLaunch(task.task_id, task.user_id, task.continuation_launch); + if (saved.orchestrator_version !== process.env.AWS_LAMBDA_FUNCTION_VERSION) { + throw new Error('MICROVM_CONTINUATION_VERSION_CHANGED: recovery requires the original published coordinator'); + } + return saved; + }); + strategy = resolveComputeStrategy(launch.blueprint); + const assigned = await context.step('continuation-assignment', async () => { + // Dispatch already admitted this attempt. This call verifies its readonly + // lease and held slot; it cannot re-admit a different request. + const current = await loadTask(task.task_id, true) as ContinuableTask; + if (!sameAttempt(current, event)) return null; + const admission = await admitContinuation(task.task_id, task.user_id, event.continuation_request_id, 1); + return admission.kind === 'ready' && sameAttempt(admission.task, event) ? admission.task : null; + }); + if (!assigned) return; + const source = assigned.continuation?.source_handle; + if (!source?.imageArn || !source.imageVersion || !assigned.continuation?.started_at) { + throw new Error('MICROVM_CONTINUATION_IMAGE_INVALID: saved worker image or assignment time is missing'); + } + const startInput = { + taskId: task.task_id, + userId: task.user_id, + blueprintConfig: launch.blueprint, + payload: { + ...launch.payload, + attempt_id: event.continuation_attempt_id, + task_started_at: assigned.continuation.started_at, + }, + microvmImage: { imageArn: source.imageArn, imageVersion: source.imageVersion }, + }; + const started = await context.step('continuation-start', () => strategy!.startSession(startInput), { + // All retries use the immutable start receipt/token, including lost replies. + retryStrategy: (_error, attempt) => ({ shouldRetry: attempt < CONTINUATION_START_ATTEMPTS, delay: { seconds: 10 } }), + }); + if (started.strategyType !== 'lambda-microvm') throw new Error('MICROVM_CONTINUATION_BACKEND_CHANGED'); + handle = started; + const restoring = await context.waitForCondition('continuation-restore', state => + pollContinuationRestore(event, task.user_id, started, strategy!, state), { + initialState: { deadlineMs: Date.parse(assigned.continuation.started_at) + RESTORE_TIMEOUT_MS }, + waitStrategy: state => ({ + shouldContinue: !state.ready && !state.closed && !state.ownershipLost && !state.failure, + delay: { seconds: 5 }, + }), + }); + if (restoring.failure) throw new Error(restoring.failure); + if (restoring.ownershipLost) { + await context.step('continuation-stop-old-worker', () => stopMicrovmWithDiagnostics({ + taskId: task.task_id, handle: started, strategy: strategy!, emitEvent: emit, + })); + return; + } + const final = restoring.closed ? { attempts: 0 } : await context.waitForCondition( + 'continuation-agent-completion', state => pollMicrovmTask({ + taskId: task.task_id, + userId: task.user_id, + handle: started, + strategy: strategy!, + pollIntervalMs: launch.blueprint.poll_interval_ms ?? CONTINUATION_POLL_INTERVAL_MS, + suspendEnabled: process.env.MICROVM_APPROVAL_SUSPEND_ENABLED === 'true', + emitEvent: emit, + }, state), { initialState: { attempts: 0 }, waitStrategy: continuationWaitStrategy }, + ); + await context.step('continuation-finalize', async () => { + if (final.microvmParked) { + await emit('continuation_parked', { + microvm_id: started.microvmId, + detail: 'Your approval request is still available. The saved task will continue on another worker after your answer.', + }); + } else { + try { + const current = await loadTask(task.task_id, true); + if (!final.microvmOwnershipLost && current.session_id === started.microvmId) { + await finalizeTask(task.task_id, final, task.user_id); + } + } finally { + await stopMicrovmWithDiagnostics({ + taskId: task.task_id, handle: started, strategy: strategy!, emitEvent: emit, + }); + } + } + await deleteMicrovmPayload(task.task_id, event.continuation_attempt_id); + }); + } catch (error) { + // Recover a handle saved before the start step's reply was lost. + await context.step('continuation-failed', async () => { + const current = await loadTask(task.task_id, true) as ContinuableTask; + if (!sameAttempt(current, event)) return; + const savedHandle = current.microvm_start?.handle as MicrovmHandle | undefined; + const ownedHandle = handle ?? savedHandle; + await failContinuationAttempt(event, task.user_id, `Saved-task continuation failed: ${String(error)}`); + if (ownedHandle) { + const cleanup = strategy ?? resolveComputeStrategy({ compute_type: 'lambda-microvm' } as Parameters[0]); + await stopMicrovmWithDiagnostics({ taskId: task.task_id, handle: ownedHandle, strategy: cleanup, emitEvent: emit }); + } + await finalizeTask(task.task_id, { attempts: 0 }, task.user_id); + await deleteMicrovmPayload(task.task_id, event.continuation_attempt_id); + log.error('Saved-task continuation failed', { + request_id: event.continuation_request_id, + attempt_id: event.continuation_attempt_id, + ...microvmErrorIdentity(error), + }); + }); + } +} diff --git a/cdk/src/handlers/shared/microvm-continuation-start.ts b/cdk/src/handlers/shared/microvm-continuation-start.ts new file mode 100644 index 000000000..f889f90a5 --- /dev/null +++ b/cdk/src/handlers/shared/microvm-continuation-start.ts @@ -0,0 +1,207 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { randomUUID } from 'node:crypto'; +import { GetCommand, TransactWriteCommand } from '@aws-sdk/lib-dynamodb'; +import type { SessionControlOptions } from './compute-strategy'; +import type { ContinuableTask } from './microvm-continuation-retirement'; +import { CONTINUATION_IO_TIMEOUT_MS } from './microvm-continuation-timing'; +import { type ContinuationRecord, type WorkerLease, validateContinuation, workerLeaseKey } from './microvm-continuation-types'; +import { makeDocClient } from './ua'; +import { TaskStatus } from '../../constructs/task-status'; + +const ddb = makeDocClient(); +const TABLE = process.env.TASK_TABLE_NAME!; +const APPROVALS = process.env.TASK_APPROVALS_TABLE_NAME!; +const COUNTERS = process.env.USER_CONCURRENCY_TABLE_NAME!; + +export type ContinuationAdmission = + | { readonly kind: 'ready'; readonly task: ContinuableTask } + | { readonly kind: 'waiting' | 'closed' | 'capacity' }; + +async function readTask(taskId: string, options: SessionControlOptions): Promise { + const result = await ddb.send(new GetCommand({ + TableName: TABLE, Key: { task_id: taskId }, ConsistentRead: true, + }), options); + return result.Item as ContinuableTask | undefined; +} + +/** A replay can use only the same active assignment, never revive a retired lease. */ +async function readActiveAssignment(task: ContinuableTask, options: SessionControlOptions): Promise { + const record = task.continuation; + if (!record?.attempt_id || task.concurrency_slot?.state !== 'held' + || task.concurrency_slot.attempt_id !== record.attempt_id) return false; + const result = await ddb.send(new GetCommand({ + TableName: TABLE, Key: workerLeaseKey(task.task_id), ConsistentRead: true, + }), options); + const lease = result.Item as WorkerLease | undefined; + if (record.state === 'RESTORING' && (!task.session_id || record.worker_id !== task.session_id + || lease?.lease_microvm_id !== task.session_id)) return false; + return lease?.lease_state === 'ACTIVE' && lease.lease_attempt_id === record.attempt_id + && lease.lease_user_id === task.user_id && lease.lease_repo === (task.repo ?? ''); +} + +/** + * Claim capacity and assign a fresh worker token only for a resolved PARKED + * request. The old worker was already confirmed stopped before PARKED. + */ +export async function admitContinuation( + taskId: string, userId: string, requestId: string, limit: number, + options: SessionControlOptions = { abortSignal: AbortSignal.timeout(CONTINUATION_IO_TIMEOUT_MS) }, +): Promise { + const task = await readTask(taskId, options); + if (!task || task.user_id !== userId || task.status !== TaskStatus.AWAITING_APPROVAL + || task.awaiting_approval_request_id !== requestId || !task.continuation) return { kind: 'closed' }; + const record = task.continuation; + validateContinuation(record, task); + if (record.state === 'STARTING' || record.state === 'RESTORING') { + if (!await readActiveAssignment(task, options)) { + throw new Error('MICROVM_CONTINUATION_LEASE_INVALID: replacement assignment is not active'); + } + return { kind: 'ready', task }; + } + if (record.state !== 'PARKED') return { kind: 'waiting' }; + const approval = await ddb.send(new GetCommand({ + TableName: APPROVALS, Key: { task_id: taskId, request_id: requestId }, ConsistentRead: true, + }), options); + if (approval.Item?.user_id !== userId || approval.Item.status === 'CANCELLED') return { kind: 'closed' }; + const finiteDeadline = Date.parse(approval.Item?.created_at ?? '') + Number(approval.Item?.timeout_s) * 1000; + const expired = approval.Item?.status === 'PENDING' && Number(approval.Item.timeout_s) > 0 + && Number.isFinite(finiteDeadline) && Date.now() >= finiteDeadline; + if (!expired && !['APPROVED', 'DENIED', 'TIMED_OUT'].includes(approval.Item?.status)) return { kind: 'waiting' }; + if (task.concurrency_slot?.state !== 'released' || !task.microvm_start?.clientToken + || !record.source_handle || record.source_handle.microvmId !== task.session_id + || !Number.isSafeInteger(limit) || limit <= 0) { + throw new Error('MICROVM_CONTINUATION_RESERVATION_INVALID: parked task has no released source reservation'); + } + const attempt = randomUUID(); + const now = new Date().toISOString(); + const starting: ContinuationRecord = { + ...record, state: 'STARTING', attempt_id: attempt, started_at: now, + }; + const slot = { state: 'held' as const, acquired_at: now, attempt_id: attempt }; + const lease: WorkerLease = { + ...workerLeaseKey(taskId), + lease_state: 'ACTIVE', + lease_attempt_id: attempt, + lease_user_id: userId, + lease_repo: record.identity.repo, + }; + try { + await ddb.send(new TransactWriteCommand({ + ClientRequestToken: attempt, + TransactItems: [ + { + Update: { + TableName: TABLE, + Key: { task_id: taskId }, + UpdateExpression: 'SET continuation = :starting, concurrency_slot = :slot ' + + 'REMOVE session_id, compute_metadata, microvm_start, microvm_lifecycle, agent_heartbeat_at', + ConditionExpression: 'user_id = :user AND #status = :awaiting AND awaiting_approval_request_id = :request ' + + 'AND continuation = :record AND concurrency_slot.#state = :released', + ExpressionAttributeNames: { '#status': 'status', '#state': 'state' }, + ExpressionAttributeValues: { + ':user': userId, + ':awaiting': TaskStatus.AWAITING_APPROVAL, + ':request': requestId, + ':record': record, + ':released': 'released', + ':starting': starting, + ':slot': slot, + }, + }, + }, + { + Update: { + TableName: COUNTERS, + Key: { user_id: userId }, + UpdateExpression: 'SET active_count = if_not_exists(active_count, :zero) + :one, ' + + 'updated_at = :now, reservation_version = :revision', + ConditionExpression: 'attribute_not_exists(active_count) OR active_count < :limit', + ExpressionAttributeValues: { + ':zero': 0, ':one': 1, ':now': now, ':revision': attempt, ':limit': limit, + }, + }, + }, + { + Put: { + TableName: TABLE, + Item: lease, + ConditionExpression: 'lease_state = :parked AND lease_attempt_id = :source ' + + 'AND lease_microvm_id = :vm AND lease_user_id = :user', + ExpressionAttributeValues: { + ':parked': 'PARKED', + ':source': task.microvm_start.clientToken, + ':vm': record.source_handle.microvmId, + ':user': userId, + }, + }, + }, + expired ? { + Update: { + TableName: APPROVALS, + Key: { task_id: taskId, request_id: requestId }, + UpdateExpression: 'SET #status = :timedout, decided_at = :now', + ConditionExpression: 'user_id = :user AND #status = :pending AND created_at = :created AND timeout_s = :timeout', + ExpressionAttributeNames: { '#status': 'status' }, + ExpressionAttributeValues: { + ':user': userId, + ':pending': 'PENDING', + ':timedout': 'TIMED_OUT', + ':now': now, + ':created': approval.Item!.created_at, + ':timeout': approval.Item!.timeout_s, + }, + }, + } : { + ConditionCheck: { + TableName: APPROVALS, + Key: { task_id: taskId, request_id: requestId }, + ConditionExpression: 'user_id = :user AND #status IN (:approved, :denied, :timedout)', + ExpressionAttributeNames: { '#status': 'status' }, + ExpressionAttributeValues: { + ':user': userId, ':approved': 'APPROVED', ':denied': 'DENIED', ':timedout': 'TIMED_OUT', + }, + }, + }, + ], + }), options); + const updated: ContinuableTask = { + ...task, + continuation: starting, + concurrency_slot: slot, + session_id: undefined, + compute_metadata: undefined, + microvm_start: undefined, + agent_heartbeat_at: undefined, + }; + return { kind: 'ready', task: updated }; + } catch (error) { + const latest = await readTask(taskId, options); + if (!latest || latest.user_id !== userId || latest.status !== TaskStatus.AWAITING_APPROVAL + || latest.awaiting_approval_request_id !== requestId) return { kind: 'closed' }; + if (['STARTING', 'RESTORING'].includes(latest.continuation?.state ?? '') + && await readActiveAssignment(latest, options)) return { kind: 'ready', task: latest }; + const failure = error as { name?: string; CancellationReasons?: { Code?: string }[] }; + if (failure.name === 'TransactionCanceledException' + && failure.CancellationReasons?.[1]?.Code === 'ConditionalCheckFailed' + && latest.continuation?.state === 'PARKED') return { kind: 'capacity' }; + throw error; + } +} diff --git a/cdk/src/handlers/shared/microvm-continuation-storage.ts b/cdk/src/handlers/shared/microvm-continuation-storage.ts new file mode 100644 index 000000000..c3c7473a7 --- /dev/null +++ b/cdk/src/handlers/shared/microvm-continuation-storage.ts @@ -0,0 +1,285 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { createHash } from 'node:crypto'; +import { Readable } from 'node:stream'; +import { DeleteObjectsCommand, GetObjectCommand, HeadObjectCommand, ListObjectVersionsCommand, PutObjectCommand, S3Client } from '@aws-sdk/client-s3'; +import { GetCommand, UpdateCommand } from '@aws-sdk/lib-dynamodb'; +import { canonicalJson } from './canonical-json'; +import type { SessionControlOptions } from './compute-strategy'; +import { CONTINUATION_IO_TIMEOUT_MS } from './microvm-continuation-timing'; +import { + CONTINUATION, type ContinuationLaunchReceipt, type ContinuationRecord, validAttemptId, +} from './microvm-continuation-types'; +import { continuationEnabled } from './microvm-worker-lease'; +import type { BlueprintConfig } from './repo-config'; +import { makeClient, makeDocClient } from './ua'; +import constants from '../../../../contracts/constants.json'; +import { TERMINAL_STATUSES } from '../../constructs/task-status'; + +const ddb = makeDocClient(); +let s3: S3Client | undefined; +function storage(): S3Client { + return s3 ??= makeClient(S3Client); +} +const TABLE = process.env.TASK_TABLE_NAME!; +const MAX_BYTES = constants.payload_bootstrap.max_payload_bytes; + +interface LaunchInputs { + readonly version: number; + readonly task_id: string; + readonly user_id: string; + readonly payload: Record; + readonly blueprint: BlueprintConfig; + readonly orchestrator_version: string; +} + +function canonical(value: unknown): Buffer { + return Buffer.from(canonicalJson(value)); +} + +function hash(value: Buffer): string { + return createHash('sha256').update(value).digest('hex'); +} + +async function readObject(key: string, versionId?: string, options?: SessionControlOptions) { + const timeout = AbortSignal.timeout(CONTINUATION_IO_TIMEOUT_MS); + const signal = options?.abortSignal ? AbortSignal.any([options.abortSignal, timeout]) : timeout; + const response = await storage().send(new GetObjectCommand({ + Bucket: process.env.CONTINUATION_BUCKET_NAME!, + Key: key, + ...(versionId && { VersionId: versionId }), + }), { abortSignal: signal }); + if (!(response.Body instanceof Readable) || !response.ContentLength || response.ContentLength > MAX_BYTES + || !response.VersionId || response.VersionId === 'null' + || (versionId && response.VersionId !== versionId)) { + if (response.Body instanceof Readable) response.Body.destroy(); + throw new Error('MICROVM_CONTINUATION_STORAGE_INVALID: saved object is incomplete or unversioned'); + } + const stream = response.Body; + const abort = () => { stream.destroy(new Error('MICROVM_CONTINUATION_STORAGE_TIMEOUT: download did not finish')); }; + signal.addEventListener('abort', abort, { once: true }); + const chunks: Buffer[] = []; + let size = 0; + try { + signal.throwIfAborted(); + for await (const chunk of stream) { + signal.throwIfAborted(); + const bytes = Buffer.from(chunk); + size += bytes.length; + if (size > response.ContentLength || size > MAX_BYTES) { + throw new Error('MICROVM_CONTINUATION_STORAGE_INVALID: download exceeded its declared size'); + } + chunks.push(bytes); + } + } finally { + signal.removeEventListener('abort', abort); + stream.destroy(); + } + const bytes = Buffer.concat(chunks, size); + if (bytes.length !== response.ContentLength) { + throw new Error('MICROVM_CONTINUATION_STORAGE_INVALID: launch object length changed'); + } + return { bytes, versionId: response.VersionId }; +} + +/** Save exact hydrated inputs independently of the short-lived bootstrap URL. */ +export async function saveContinuationLaunch( + taskId: string, userId: string, payload: Record, blueprint: BlueprintConfig, +): Promise { + if (!continuationEnabled()) return; + const version = process.env.AWS_LAMBDA_FUNCTION_VERSION ?? ''; + if (!validAttemptId(taskId) || payload.task_id !== taskId || payload.user_id !== userId || !/^\d+$/.test(version)) { + throw new Error('MICROVM_CONTINUATION_INPUT_INVALID: task identity or published coordinator version is missing'); + } + const inputs: LaunchInputs = { + version: CONTINUATION.version, + task_id: taskId, + user_id: userId, + payload, + blueprint, + orchestrator_version: version, + }; + const bytes = canonical(inputs); + if (!bytes.length || bytes.length > MAX_BYTES) { + throw new Error('MICROVM_CONTINUATION_INPUT_INVALID: launch inputs exceed the storage bound'); + } + const sha256 = hash(bytes); + const key = `${CONTINUATION.object_key_prefix}${taskId}/launch/${sha256}.json`; + let putError: unknown; + try { + await storage().send(new PutObjectCommand({ + Bucket: process.env.CONTINUATION_BUCKET_NAME!, + Key: key, + Body: bytes, + ContentType: 'application/json', + ServerSideEncryption: 'AES256', + ChecksumSHA256: createHash('sha256').update(bytes).digest('base64'), + IfNoneMatch: '*', + }), { abortSignal: AbortSignal.timeout(CONTINUATION_IO_TIMEOUT_MS) }); + } catch (error) { + putError = error; + } + // Read back even after a successful Put. This also recovers a lost reply or + // an identical publication by a replay, without changing the version pointer. + let stored: Awaited>; + try { + stored = await readObject(key); + } catch (error) { + throw putError ?? error; + } + if (!stored.bytes.equals(bytes)) { + throw new Error('MICROVM_CONTINUATION_STORAGE_INVALID: launch readback does not match published bytes'); + } + const receipt: ContinuationLaunchReceipt = { + version: CONTINUATION.version, + key, + version_id: stored.versionId, + sha256, + size_bytes: bytes.length, + orchestrator_version: version, + }; + try { + await ddb.send(new UpdateCommand({ + TableName: TABLE, + Key: { task_id: taskId }, + UpdateExpression: 'SET continuation_launch = :receipt REMOVE #ttl', + ConditionExpression: 'user_id = :user AND #status = :hydrating ' + + 'AND (attribute_not_exists(continuation_launch) OR continuation_launch = :receipt)', + ExpressionAttributeNames: { '#status': 'status', '#ttl': 'ttl' }, + ExpressionAttributeValues: { ':receipt': receipt, ':user': userId, ':hydrating': 'HYDRATING' }, + })); + } catch (error) { + const current = await ddb.send(new GetCommand({ + TableName: TABLE, Key: { task_id: taskId }, ConsistentRead: true, + })); + if (current.Item?.user_id !== userId + || !canonical(current.Item.continuation_launch ?? null).equals(canonical(receipt))) throw error; + } +} + +/** The task record pins the version; workers cannot choose or overwrite this pointer. */ +export async function loadContinuationLaunch( + taskId: string, userId: string, receipt: ContinuationLaunchReceipt, +): Promise { + if (!validAttemptId(taskId) || receipt?.version !== CONTINUATION.version + || !/^[a-f0-9]{64}$/.test(receipt.sha256) + || receipt.key !== `${CONTINUATION.object_key_prefix}${taskId}/launch/${receipt.sha256}.json` + || !receipt.version_id || receipt.version_id === 'null' + || !Number.isSafeInteger(receipt.size_bytes) || receipt.size_bytes <= 0 || receipt.size_bytes > MAX_BYTES + || !/^\d+$/.test(receipt.orchestrator_version)) { + throw new Error('MICROVM_CONTINUATION_INPUT_INVALID: saved launch receipt is invalid'); + } + const stored = await readObject(receipt.key, receipt.version_id); + if (stored.bytes.length !== receipt.size_bytes || hash(stored.bytes) !== receipt.sha256) { + throw new Error('MICROVM_CONTINUATION_STORAGE_INVALID: saved launch checksum does not match'); + } + const inputs = JSON.parse(stored.bytes.toString('utf8')) as LaunchInputs; + if (!inputs || inputs.version !== CONTINUATION.version || inputs.task_id !== taskId || inputs.user_id !== userId + || inputs.payload?.task_id !== taskId || inputs.payload.user_id !== userId + || inputs.blueprint?.compute_type !== 'lambda-microvm' + || inputs.orchestrator_version !== receipt.orchestrator_version) { + throw new Error('MICROVM_CONTINUATION_INPUT_INVALID: saved launch belongs to a different task'); + } + return inputs; +} + +/** Confirm all acknowledged object versions before retiring the only worker. */ +export async function verifyContinuationCheckpoint(record: ContinuationRecord, options?: SessionControlOptions): Promise { + const stored = await readObject(record.manifest.key, record.manifest.version_id, options); + if (stored.bytes.length !== record.manifest.size_bytes || hash(stored.bytes) !== record.manifest.sha256) { + throw new Error('MICROVM_CONTINUATION_STORAGE_INVALID: checkpoint manifest checksum does not match'); + } + const manifest = JSON.parse(stored.bytes.toString('utf8')); + if (!manifest || manifest.version !== CONTINUATION.version + || !canonical(manifest.identity ?? null).equals(canonical(record.identity))) { + throw new Error('MICROVM_CONTINUATION_STORAGE_INVALID: checkpoint manifest identity does not match'); + } + const identity = record.identity; + const prefix = `${CONTINUATION.object_key_prefix}${identity.task_id}/${identity.attempt_id}/${identity.request_id}/`; + for (const object of [ + { receipt: manifest.conversation, prefix, suffix: '.json', limit: CONTINUATION.max_conversation_bytes }, + { receipt: manifest.workspace, prefix: `${prefix}workspace/`, suffix: '.tar', limit: CONTINUATION.max_workspace_bytes }, + ]) { + const receipt = object.receipt; + if (!receipt || !/^[a-f0-9]{64}$/.test(receipt.sha256) + || receipt.key !== `${object.prefix}${receipt.sha256}${object.suffix}` + || typeof receipt.version_id !== 'string' || !receipt.version_id || receipt.version_id === 'null' + || !Number.isSafeInteger(receipt.size_bytes) || receipt.size_bytes <= 0 || receipt.size_bytes > object.limit) { + throw new Error('MICROVM_CONTINUATION_STORAGE_INVALID: checkpoint contains an invalid object receipt'); + } + const response = await storage().send(new HeadObjectCommand({ + Bucket: process.env.CONTINUATION_BUCKET_NAME!, + Key: receipt.key, + VersionId: receipt.version_id, + ChecksumMode: 'ENABLED', + }), { + abortSignal: options?.abortSignal + ? AbortSignal.any([options.abortSignal, AbortSignal.timeout(CONTINUATION_IO_TIMEOUT_MS)]) + : AbortSignal.timeout(CONTINUATION_IO_TIMEOUT_MS), + }); + if (response.VersionId !== receipt.version_id || response.ContentLength !== receipt.size_bytes + || response.ChecksumSHA256 !== Buffer.from(receipt.sha256, 'hex').toString('base64')) { + throw new Error('MICROVM_CONTINUATION_STORAGE_INVALID: checkpoint object version, length or checksum does not match'); + } + } +} + +/** + * Called only after the last worker is confirmed stopped. Remove every version, + * including superseded checkpoints, before clearing the task's cleanup marker. + * Active records have no object expiry: a pending answer can outlive a worker. + */ +export async function deleteClosedTaskContinuations( + taskId: string, userId: string, options: SessionControlOptions, +): Promise { + if (!validAttemptId(taskId)) throw new Error('MICROVM_CONTINUATION_CLEANUP_INVALID'); + const current = await ddb.send(new GetCommand({ + TableName: TABLE, Key: { task_id: taskId }, ConsistentRead: true, + }), options); + if (current.Item?.user_id !== userId || !TERMINAL_STATUSES.includes(current.Item.status)) return; + const prefix = `${CONTINUATION.object_key_prefix}${taskId}/`; + // Always list from the start after deleting a page. No cursor can skip a + // version whose neighbour was removed in the preceding batch. + while (true) { + options.abortSignal?.throwIfAborted(); + const page = await storage().send(new ListObjectVersionsCommand({ + Bucket: process.env.CONTINUATION_BUCKET_NAME!, Prefix: prefix, MaxKeys: 1000, + }), options); + const objects = [...(page.Versions ?? []), ...(page.DeleteMarkers ?? [])].map(object => { + if (!object.Key?.startsWith(prefix) || !object.VersionId) { + throw new Error('MICROVM_CONTINUATION_CLEANUP_INVALID: unexpected object identity'); + } + return { Key: object.Key, VersionId: object.VersionId }; + }); + if (!objects.length) break; + const removed = await storage().send(new DeleteObjectsCommand({ + Bucket: process.env.CONTINUATION_BUCKET_NAME!, Delete: { Objects: objects, Quiet: true }, + }), options); + if (removed.Errors?.length) throw new Error('MICROVM_CONTINUATION_CLEANUP_FAILED: object versions remain'); + } + await ddb.send(new UpdateCommand({ + TableName: TABLE, + Key: { task_id: taskId }, + UpdateExpression: 'SET continuation_cleanup_at = :now REMOVE continuation, continuation_launch', + ConditionExpression: 'user_id = :user AND #status = :status', + ExpressionAttributeNames: { '#status': 'status' }, + ExpressionAttributeValues: { ':user': userId, ':status': current.Item.status, ':now': new Date().toISOString() }, + }), options); +} diff --git a/cdk/src/handlers/shared/microvm-continuation-timing.ts b/cdk/src/handlers/shared/microvm-continuation-timing.ts new file mode 100644 index 000000000..e075fd41e --- /dev/null +++ b/cdk/src/handlers/shared/microvm-continuation-timing.ts @@ -0,0 +1,27 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +/** Transport limits fit inside a durable Lambda's 60-second invocation budget. */ +export const CONTINUATION_IO_TIMEOUT_MS = 10_000; +export const CONTINUATION_RETIREMENT_TIMEOUT_MS = 40_000; +export const CONTINUATION_STOP_TIMEOUT_MS = 25_000; +export const CONTINUATION_POLL_INTERVAL_MS = 30_000; +export const CONTINUATION_TRANSITION_POLL_SECONDS = 5; +export const CONTINUATION_RETRY_POLL_SECONDS = 30; +export const CONTINUATION_START_ATTEMPTS = 4; diff --git a/cdk/src/handlers/shared/microvm-continuation-types.ts b/cdk/src/handlers/shared/microvm-continuation-types.ts new file mode 100644 index 000000000..07da23233 --- /dev/null +++ b/cdk/src/handlers/shared/microvm-continuation-types.ts @@ -0,0 +1,99 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import type { SessionHandle } from './compute-strategy'; +import constants from '../../../../contracts/constants.json'; + +export const CONTINUATION = constants.microvm_continuation; +export type MicrovmHandle = Extract; + +export interface ContinuationIdentity { + readonly task_id: string; + /** Physical worker that captured the source checkpoint. */ + readonly attempt_id: string; + readonly request_id: string; + readonly user_id: string; + readonly repo: string; +} + +export interface ContinuationReceipt { + readonly kind: 'manifest'; + readonly key: string; + readonly version_id: string; + readonly sha256: string; + readonly size_bytes: number; +} + +export interface ContinuationRecord { + readonly version: number; + /** PARKED means the source worker has retired; the guest's local parked phase only means a safe approval wait. */ + readonly state: 'READY' | 'FENCED' | 'PARKED' | 'STARTING' | 'RESTORING' | 'CONSUMED'; + readonly identity: ContinuationIdentity; + readonly manifest: ContinuationReceipt; + readonly source_handle?: MicrovmHandle; + readonly parked_at?: string; + /** Logical launch token, assigned before calling RunMicrovm. */ + readonly attempt_id?: string; + readonly worker_id?: string; + readonly started_at?: string; +} + +export interface ContinuationLaunchReceipt { + readonly version: number; + readonly key: string; + readonly version_id: string; + readonly sha256: string; + readonly size_bytes: number; + readonly orchestrator_version: string; +} + +export interface WorkerLease { + readonly task_id: string; + readonly lease_attempt_id: string; + readonly lease_state: 'ACTIVE' | 'FENCED' | 'PARKED' | 'CLOSED'; + readonly lease_user_id: string; + readonly lease_repo: string; + readonly lease_microvm_id?: string; +} + +export function workerLeaseKey(taskId: string): { task_id: string } { + return { task_id: CONTINUATION.lease_key_prefix + taskId }; +} + +export function validAttemptId(value: unknown): value is string { + return typeof value === 'string' && /^[A-Za-z0-9][A-Za-z0-9_-]{0,127}$/.test(value); +} + +/** A worker may publish a pointer only within its own complete checkpoint prefix. */ +export function validateContinuation(record: ContinuationRecord, task: { + task_id: string; user_id: string; repo?: string; awaiting_approval_request_id?: string; +}): void { + const identity = record?.identity; + const receipt = record?.manifest; + if (record?.version !== CONTINUATION.version || !identity || !receipt + || identity.task_id !== task.task_id || identity.user_id !== task.user_id || identity.repo !== (task.repo ?? '') + || identity.request_id !== task.awaiting_approval_request_id + || !validAttemptId(identity.task_id) || !validAttemptId(identity.attempt_id) || !validAttemptId(identity.request_id) + || receipt.kind !== 'manifest' || !/^[a-f0-9]{64}$/.test(receipt.sha256) + || receipt.key !== `${CONTINUATION.object_key_prefix}${identity.task_id}/${identity.attempt_id}/${identity.request_id}/manifest/${receipt.sha256}.json` + || typeof receipt.version_id !== 'string' || !receipt.version_id || receipt.version_id === 'null' + || !Number.isSafeInteger(receipt.size_bytes) || receipt.size_bytes <= 0 || receipt.size_bytes > CONTINUATION.max_manifest_bytes) { + throw new Error('MICROVM_CONTINUATION_INVALID: checkpoint identity or receipt does not match the pending request'); + } +} diff --git a/cdk/src/handlers/shared/microvm-control.ts b/cdk/src/handlers/shared/microvm-control.ts new file mode 100644 index 000000000..3c55b02ca --- /dev/null +++ b/cdk/src/handlers/shared/microvm-control.ts @@ -0,0 +1,38 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +// SPDX-License-Identifier: MIT-0 + +/** Keep service correlation identifiers from successful replies, never the reply body. */ +export function microvmRequestIdentity(response: unknown): { aws_request_id?: string } { + const requestId = (response as { $metadata?: { requestId?: unknown } } | undefined)?.$metadata?.requestId; + return typeof requestId === 'string' && /^[A-Za-z0-9-]{1,128}$/.test(requestId) + ? { aws_request_id: requestId } : {}; +} + +/** Control-plane diagnostics must not copy SDK messages, payloads or credentials. */ +export function microvmErrorIdentity(error: unknown): { error_type: string; aws_request_id?: string } { + const outer = error as { name?: unknown; cause?: unknown; $metadata?: { requestId?: unknown } } | undefined; + const cause = outer?.cause as typeof outer; + const name = cause?.name ?? outer?.name; + return { + error_type: typeof name === 'string' && /^[A-Za-z0-9_]{1,100}$/.test(name) ? name : 'Error', + ...microvmRequestIdentity({ $metadata: { requestId: cause?.$metadata?.requestId ?? outer?.$metadata?.requestId } }), + }; +} diff --git a/cdk/src/handlers/shared/microvm-image-capability.ts b/cdk/src/handlers/shared/microvm-image-capability.ts new file mode 100644 index 000000000..9b338b327 --- /dev/null +++ b/cdk/src/handlers/shared/microvm-image-capability.ts @@ -0,0 +1,82 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +// SPDX-License-Identifier: MIT-0 + +import type { GetMicrovmImageVersionOutput } from '@aws-sdk/client-lambda-microvms'; +import sharedConstants from '../../../../contracts/constants.json'; + +export const MICROVM_LIFECYCLE_PROTOCOL = String(sharedConstants.microvm_lifecycle.protocol_version); +export const MICROVM_IMAGE_PROTOCOL_ENV = sharedConstants.microvm_lifecycle.image_protocol_env; +// Optional discovery/enrichment must not hold up an already registered worker. +export const MICROVM_IMAGE_CAPABILITY_REQUEST_TIMEOUT_MS = 3_000; +const MAX_IMAGE_IDENTITY_LENGTH = 2048; + +/** Coordinator-owned evidence for the image version that actually launched a VM. */ +export interface MicrovmImageMetadata { + readonly imageArn?: string; + readonly imageVersion?: string; + readonly lifecycleProtocol?: string; +} + +function nonblank(value: unknown): value is string { + return typeof value === 'string' && value.length > 0 && value.length <= MAX_IMAGE_IDENTITY_LENGTH + && value === value.trim() && !/[\u0000-\u001f\u007f]/.test(value); +} + +/** Legacy or malformed identity never acquires capability from deployment settings. */ +export function readMicrovmImageMetadata(value: unknown): Record { + if (!value || typeof value !== 'object' || Array.isArray(value)) return {}; + const metadata = value as MicrovmImageMetadata; + if (!nonblank(metadata.imageArn) || !nonblank(metadata.imageVersion)) return {}; + return { + imageArn: metadata.imageArn, + imageVersion: metadata.imageVersion, + ...(metadata.lifecycleProtocol === MICROVM_LIFECYCLE_PROTOCOL + && { lifecycleProtocol: MICROVM_LIFECYCLE_PROTOCOL }), + }; +} + +export function supportsMicrovmLifecycle(value: unknown): boolean { + return readMicrovmImageMetadata(value).lifecycleProtocol === MICROVM_LIFECYCLE_PROTOCOL; +} + +/** Check the exact immutable version returned by Run, never a latest-version alias. */ +export function verifyMicrovmImageLifecycle( + identity: Required>, + version: Pick, +): boolean { + const hooks = version.hooks; + const runtime = hooks?.microvmHooks; + const timeout = sharedConstants.microvm_hook_budgets.lifecycle_hook_timeout_seconds; + return version.imageArn === identity.imageArn + && version.imageVersion === identity.imageVersion + && version.environmentVariables?.[MICROVM_IMAGE_PROTOCOL_ENV] === MICROVM_LIFECYCLE_PROTOCOL + && hooks?.port === sharedConstants.microvm_lifecycle.hook_port + && hooks.microvmImageHooks?.ready === 'ENABLED' + && hooks.microvmImageHooks?.validate === 'ENABLED' + && runtime?.run === 'ENABLED' + && runtime.terminate === 'ENABLED' + && runtime.suspend === 'ENABLED' + && runtime.resume === 'ENABLED' + && Number.isInteger(runtime.suspendTimeoutInSeconds) + && runtime.suspendTimeoutInSeconds! >= timeout + && Number.isInteger(runtime.resumeTimeoutInSeconds) + && runtime.resumeTimeoutInSeconds! >= timeout; +} diff --git a/cdk/src/handlers/shared/microvm-lifecycle-policy.ts b/cdk/src/handlers/shared/microvm-lifecycle-policy.ts new file mode 100644 index 000000000..1a9a25785 --- /dev/null +++ b/cdk/src/handlers/shared/microvm-lifecycle-policy.ts @@ -0,0 +1,127 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import type { SessionStatus } from './compute-strategy'; +import { supportsMicrovmLifecycle } from './microvm-image-capability'; +import { intentMatchesGate, type MicrovmLifecycleSnapshot } from './microvm-lifecycle'; +import { MICROVM_SLEEP_AFTER_S_DEFAULT, MICROVM_SLEEP_AFTER_S_MAX } from './types'; +import { TaskStatus, TERMINAL_STATUSES } from '../../constructs/task-status'; + +// Initial policy values, not service limits. Live timings must validate them. +export const MICROVM_WAKE_MARGIN_MS = 60_000; +export const MICROVM_MIN_USEFUL_SLEEP_MS = 30_000; +export const MICROVM_TRANSITION_POLL_MS = 5_000; + +export interface MicrovmLifecyclePolicyInput { + readonly snapshot: MicrovmLifecycleSnapshot; + readonly substrate: SessionStatus; + readonly nowMs: number; + readonly sessionDeadlineMs: number; + readonly pollIntervalMs: number; + /** Stops new suspends; already-sleeping VMs can still wake or terminate. */ + readonly suspendEnabled: boolean; +} + +export type MicrovmLifecycleDecision = ( + | { readonly action: 'wait' | 'terminate' | 'reconcile-terminal' } + | { readonly action: 'suspend' | 'resume'; readonly requestReady: boolean } +) & { readonly reason: string; readonly nextPollInMs: number }; + +/** Pure policy: no API calls, status changes, or approval decisions. */ +export function decideMicrovmLifecycle(input: MicrovmLifecyclePolicyInput): MicrovmLifecycleDecision { + const { snapshot, substrate, nowMs, sessionDeadlineMs, pollIntervalMs } = input; + if (![nowMs, sessionDeadlineMs, pollIntervalMs].every(value => Number.isSafeInteger(value) && value >= 0) || pollIntervalMs === 0) { + throw new Error('MicroVM lifecycle policy requires valid millisecond times and a positive poll interval'); + } + const state = substrate.microvmState ?? 'UNKNOWN'; + const terminal = state === 'TERMINATING' || state === 'TERMINATED' || state === 'NOT_FOUND'; + const nextPollInMs = Math.max(1, Math.min(pollIntervalMs, sessionDeadlineMs - nowMs)); + const transitionPoll = Math.min(nextPollInMs, MICROVM_TRANSITION_POLL_MS); + if (TERMINAL_STATUSES.some(status => status === snapshot.status) || snapshot.status === TaskStatus.FINALIZING) { + return { action: terminal ? 'wait' : 'terminate', reason: 'task-closed', nextPollInMs }; + } + if (terminal) return { action: 'reconcile-terminal', reason: 'substrate-terminal', nextPollInMs }; + if (nowMs >= sessionDeadlineMs) return { action: 'terminate', reason: 'session-deadline', nextPollInMs }; + if (state !== 'RUNNING' && state !== 'SUSPENDING' && state !== 'SUSPENDED') { + return { action: 'wait', reason: state === 'PENDING' ? 'starting' : 'unconfirmed-state', nextPollInMs: transitionPoll }; + } + + const sleeping = state === 'SUSPENDING' || state === 'SUSPENDED'; + const wake = (reason: string): MicrovmLifecycleDecision => ({ + action: 'resume', requestReady: state === 'SUSPENDED', reason, nextPollInMs: transitionPoll, + }); + const sameGate = intentMatchesGate(snapshot); + if (snapshot.status !== TaskStatus.AWAITING_APPROVAL) { + if (snapshot.status !== TaskStatus.RUNNING && snapshot.status !== TaskStatus.HYDRATING) { + return { action: 'wait', reason: 'task-not-active', nextPollInMs }; + } + if (sleeping || (snapshot.intent?.action === 'suspend')) return wake('suspended-outside-gate'); + return { action: 'wait', reason: 'working', nextPollInMs }; + } + + const approval = snapshot.approval; + if (approval.kind !== 'present') { + // A failed/missing read is not permission to sleep. Preserve wake intent + // even while RUNNING if an earlier suspend might still be in flight. + if (sleeping || snapshot.intent) return wake(`approval-${approval.kind}`); + return { action: 'wait', reason: `approval-${approval.kind}`, nextPollInMs: transitionPoll }; + } + if (approval.status !== 'PENDING') return wake('approval-terminal'); + if (sameGate && snapshot.intent?.deadline_ms !== approval.deadlineMs) return wake('approval-deadline-changed'); + if (sameGate && snapshot.intent?.action === 'resume') { + return sleeping ? wake('wake-intent') : { action: 'wait', reason: 'wake-intent', nextPollInMs: transitionPoll }; + } + if (approval.createdAtMs > nowMs) return wake('approval-time-invalid'); + const wakeAt = Math.min(approval.deadlineMs ?? sessionDeadlineMs, sessionDeadlineMs) - MICROVM_WAKE_MARGIN_MS; + if (nowMs >= wakeAt) return wake('wake-deadline'); + + const sleepAfterSeconds = snapshot.sleepAfterSeconds === undefined + ? MICROVM_SLEEP_AFTER_S_DEFAULT : snapshot.sleepAfterSeconds; + // A malformed stored preference loses savings, never wake or cleanup. + if (!Number.isInteger(sleepAfterSeconds) || sleepAfterSeconds <= 0 || sleepAfterSeconds > MICROVM_SLEEP_AFTER_S_MAX) { + return sleeping || snapshot.intent ? wake('task-sleep-disabled') + : { action: 'wait', reason: 'task-sleep-disabled', nextPollInMs }; + } + + if (sleeping) { + if (!sameGate || snapshot.intent?.action !== 'suspend' || snapshot.intent.deadline_ms !== approval.deadlineMs) { + return wake('unintended-suspension'); + } + return { action: 'wait', reason: 'intentionally-suspended', nextPollInMs: Math.min(nextPollInMs, wakeAt - nowMs) }; + } + if (!input.suspendEnabled || !supportsMicrovmLifecycle(snapshot.handle)) { + return { action: 'wait', reason: 'suspend-disabled', nextPollInMs }; + } + // A prior gate's in-flight suspend must be resolved conservatively. Persist a + // wake for this gate rather than attributing that old sleep request to it. + if (snapshot.intent?.action === 'suspend' && !sameGate) return wake('previous-gate-suspend'); + const graceEndsAt = approval.createdAtMs + sleepAfterSeconds * 1000; + if (nowMs < graceEndsAt) { + return { action: 'wait', reason: 'suspend-grace', nextPollInMs: Math.min(nextPollInMs, graceEndsAt - nowMs, wakeAt - nowMs) }; + } + if (wakeAt - nowMs < MICROVM_MIN_USEFUL_SLEEP_MS) { + return { action: 'wait', reason: 'short-window', nextPollInMs: Math.min(nextPollInMs, wakeAt - nowMs) }; + } + return { + action: 'suspend', + requestReady: true, + reason: 'pending-long-gate', + nextPollInMs: Math.min(transitionPoll, wakeAt - nowMs), + }; +} diff --git a/cdk/src/handlers/shared/microvm-lifecycle.ts b/cdk/src/handlers/shared/microvm-lifecycle.ts new file mode 100644 index 000000000..e8e73d1bf --- /dev/null +++ b/cdk/src/handlers/shared/microvm-lifecycle.ts @@ -0,0 +1,309 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { randomUUID } from 'node:crypto'; +import { GetCommand, TransactWriteCommand } from '@aws-sdk/lib-dynamodb'; +import type { SessionControlOptions, SessionHandle } from './compute-strategy'; +import { readMicrovmImageMetadata, supportsMicrovmLifecycle } from './microvm-image-capability'; +import type { ApprovalStatus } from './types'; +import { makeDocClient } from './ua'; +import { TaskStatus, type TaskStatusType } from '../../constructs/task-status'; + +type MicrovmHandle = Extract; +export type LifecycleAction = 'suspend' | 'resume'; + +/** Internal coordinator data. Retain it until task expiry; never erase a wake intent. */ +export interface MicrovmLifecycleIntent { + readonly version: 1; + readonly generation: string; + readonly microvm_id: string; + readonly request_id: string | null; + readonly action: LifecycleAction; + readonly requested_at_ms: number; + readonly deadline_ms: number | null; +} + +export type LifecycleApproval = + | { + readonly kind: 'present'; + readonly status: ApprovalStatus; + readonly created_at: string; + readonly timeout_s: number; + readonly createdAtMs: number; + readonly deadlineMs: number | null; + } + | { readonly kind: 'none' | 'missing' | 'invalid' } + | { readonly kind: 'unavailable'; readonly errorType: string }; + +/** A read is only an observation; saving intent rechecks identity and generation atomically. */ +export interface MicrovmLifecycleSnapshot { + readonly taskId: string; + readonly userId: string; + readonly status: TaskStatusType; + readonly handle: MicrovmHandle; + readonly requestId: string | null; + readonly intent?: MicrovmLifecycleIntent; + readonly approval: LifecycleApproval; + /** Latest persisted guest heartbeat; used only to bound recovery liveness grace. */ + readonly heartbeatAtMs?: number; + readonly taskStartedAtMs?: number; + /** Persisted approval-wait preference; absent legacy values use the current default. */ + readonly sleepAfterSeconds?: number; +} + +export type SaveLifecycleResult = + | { readonly status: 'saved'; readonly intent: MicrovmLifecycleIntent } + | { readonly status: 'stale' | 'ineligible' }; + +// Per request/read sequence. A write plus lost-reply recovery can take two such +// budgets; the caller must also bound its whole reconciliation cycle. +export const MICROVM_LIFECYCLE_STORE_TIMEOUT_MS = 5_000; +const ddb = makeDocClient(); +const TASK_TABLE = process.env.TASK_TABLE_NAME!; +const APPROVALS_TABLE = process.env.TASK_APPROVALS_TABLE_NAME!; +const APPROVAL_STATUSES: readonly ApprovalStatus[] = ['PENDING', 'APPROVED', 'DENIED', 'CANCELLED', 'TIMED_OUT', 'STRANDED']; +const LIVE_TASK_STATUSES: readonly TaskStatusType[] = [TaskStatus.HYDRATING, TaskStatus.RUNNING, TaskStatus.AWAITING_APPROVAL]; + +function storeSignal(options?: SessionControlOptions): AbortSignal { + const limit = AbortSignal.timeout(MICROVM_LIFECYCLE_STORE_TIMEOUT_MS); + const signal = options?.abortSignal ? AbortSignal.any([limit, options.abortSignal]) : limit; + signal.throwIfAborted(); + return signal; +} + +function nonblank(value: unknown): value is string { + return typeof value === 'string' && value.trim().length > 0; +} +function timestamp(value: unknown): value is number { + return typeof value === 'number' && Number.isSafeInteger(value) && value >= 0; +} +function validIntent(value: unknown): value is MicrovmLifecycleIntent { + if (!value || typeof value !== 'object') return false; + const item = value as MicrovmLifecycleIntent; + return item.version === 1 && nonblank(item.generation) && nonblank(item.microvm_id) + && (item.request_id === null || nonblank(item.request_id)) + && (item.action === 'suspend' || item.action === 'resume') + && timestamp(item.requested_at_ms) && (item.deadline_ms === null || timestamp(item.deadline_ms)) + && (item.action !== 'suspend' || item.request_id !== null); +} + +function parseApproval(row: Record | undefined, taskId: string, userId: string, requestId: string): LifecycleApproval { + if (!row) return { kind: 'missing' }; + if (row.task_id !== taskId || row.user_id !== userId || row.request_id !== requestId + || !APPROVAL_STATUSES.includes(row.status as ApprovalStatus) + || typeof row.created_at !== 'string' || !/^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}(\.\d{3})?Z$/.test(row.created_at) + || !timestamp(row.timeout_s)) return { kind: 'invalid' }; + const createdAtMs = Date.parse(row.created_at); + const canonical = row.created_at.includes('.') ? row.created_at : row.created_at.replace('Z', '.000Z'); + const deadlineMs = row.timeout_s === 0 ? null : createdAtMs + row.timeout_s * 1000; + if (!timestamp(createdAtMs) || new Date(createdAtMs).toISOString() !== canonical + || (deadlineMs !== null && !timestamp(deadlineMs))) return { kind: 'invalid' }; + return { + kind: 'present', + status: row.status as ApprovalStatus, + created_at: row.created_at, + timeout_s: row.timeout_s, + createdAtMs, + deadlineMs, + }; +} + +/** Missing/non-MicroVM tasks are inapplicable. Invalid identity/state fails visibly. */ +export async function readMicrovmLifecycleSnapshot( + taskId: string, userId: string, options?: SessionControlOptions, +): Promise { + const abortSignal = storeSignal(options); + const result = await ddb.send(new GetCommand({ TableName: TASK_TABLE, Key: { task_id: taskId }, ConsistentRead: true }), { abortSignal }); + abortSignal.throwIfAborted(); + const task = result.Item; + if (!task) return undefined; + if (task.user_id !== userId || task.task_id !== taskId) throw new Error('MicroVM lifecycle task identity mismatch'); + if (task.compute_type !== 'lambda-microvm') return undefined; + const metadata = task.compute_metadata; + if (!nonblank(task.session_id) || metadata?.microvmId !== task.session_id || !nonblank(metadata?.endpoint) + || !Object.values(TaskStatus).includes(task.status)) throw new Error('MicroVM lifecycle task has invalid handle or status'); + const requestId = task.awaiting_approval_request_id ?? null; + if ((requestId !== null && !nonblank(requestId)) + || (task.status === TaskStatus.AWAITING_APPROVAL && requestId === null) + || ((task.status === TaskStatus.RUNNING || task.status === TaskStatus.HYDRATING) && requestId !== null)) { + throw new Error('MicroVM lifecycle task has inconsistent approval identity'); + } + if (task.microvm_lifecycle !== undefined && !validIntent(task.microvm_lifecycle)) { + throw new Error('MicroVM lifecycle intent is malformed or has an unsupported version'); + } + let approval: LifecycleApproval = { kind: 'none' }; + // Cancellation/terminal writers can retain the old gate pointer. They need + // cleanup, not another approval read or a wake, so preserve that identity only. + if (task.status === TaskStatus.AWAITING_APPROVAL && requestId !== null) { + try { + const response = await ddb.send(new GetCommand({ + TableName: APPROVALS_TABLE, + Key: { task_id: taskId, request_id: requestId }, + ConsistentRead: true, + }), { abortSignal }); + abortSignal.throwIfAborted(); + approval = parseApproval(response.Item, taskId, userId, requestId); + } catch (error) { + // This explicit observation forbids suspend and permits conservative wake. + // The caller must count/report the failure; no raw SDK message is retained. + const name = (error as { name?: unknown })?.name; + approval = { kind: 'unavailable', errorType: typeof name === 'string' && /^[A-Za-z0-9_]{1,100}$/.test(name) ? name : 'Error' }; + } + } + return { + taskId, + userId, + status: task.status, + requestId, + approval, + intent: task.microvm_lifecycle, + ...(task.microvm_sleep_after_s !== undefined && { sleepAfterSeconds: task.microvm_sleep_after_s }), + ...(typeof task.agent_heartbeat_at === 'string' && Number.isSafeInteger(Date.parse(task.agent_heartbeat_at)) + && { heartbeatAtMs: Date.parse(task.agent_heartbeat_at) }), + ...(typeof task.started_at === 'string' && Number.isSafeInteger(Date.parse(task.started_at)) + && { taskStartedAtMs: Date.parse(task.started_at) }), + handle: { + strategyType: 'lambda-microvm', + sessionId: task.session_id, + microvmId: metadata.microvmId, + endpoint: metadata.endpoint, + ...readMicrovmImageMetadata(metadata), + }, + }; +} + +export function intentMatchesGate(snapshot: MicrovmLifecycleSnapshot): boolean { + return snapshot.intent?.microvm_id === snapshot.handle.microvmId && snapshot.intent.request_id === snapshot.requestId; +} + +function eligible(snapshot: MicrovmLifecycleSnapshot, action: LifecycleAction, nowMs: number): boolean { + if (!LIVE_TASK_STATUSES.includes(snapshot.status)) return false; + if (action === 'resume') return true; + return supportsMicrovmLifecycle(snapshot.handle) + && snapshot.status === TaskStatus.AWAITING_APPROVAL && snapshot.requestId !== null + && snapshot.approval.kind === 'present' && snapshot.approval.status === 'PENDING' + && nowMs >= snapshot.approval.createdAtMs + && (snapshot.approval.deadlineMs === null || nowMs < snapshot.approval.deadlineMs) + && !(intentMatchesGate(snapshot) && (snapshot.intent?.action === 'resume' + || snapshot.intent?.deadline_ms !== snapshot.approval.deadlineMs)); +} + +/** + * Save before touching AWS compute. A wake is sticky for this gate, including + * while AWS still reports RUNNING: an older suspend may be in flight. + * Caller must recheck the gate/time before suspend and reconcile after commands. + */ +export async function saveMicrovmLifecycleIntent( + snapshot: MicrovmLifecycleSnapshot, action: LifecycleAction, nowMs = Date.now(), + options?: SessionControlOptions, +): Promise { + if (!timestamp(nowMs)) throw new Error('MicroVM lifecycle time must be epoch milliseconds'); + if (!eligible(snapshot, action, nowMs)) return { status: 'ineligible' }; + const sameGate = intentMatchesGate(snapshot); + // Repeated requests retain their original age/generation, so polling cannot + // reset the eventual recovery budget. A new gate/action gets a new generation. + const intent: MicrovmLifecycleIntent = sameGate && snapshot.intent?.action === action ? snapshot.intent : { + version: 1, + generation: randomUUID(), + microvm_id: snapshot.handle.microvmId, + request_id: snapshot.requestId, + action, + requested_at_ms: nowMs, + deadline_ms: snapshot.approval.kind === 'present' ? snapshot.approval.deadlineMs : sameGate ? snapshot.intent!.deadline_ms : null, + }; + const names: Record = { '#status': 'status' }; + const values: Record = { + ':intent': intent, + ':user': snapshot.userId, + ':status': snapshot.status, + ':type': 'lambda-microvm', + ':id': snapshot.handle.microvmId, + ':endpoint': snapshot.handle.endpoint, + }; + let condition = 'user_id = :user AND #status = :status AND compute_type = :type AND session_id = :id ' + + 'AND compute_metadata.microvmId = :id AND compute_metadata.endpoint = :endpoint'; + if (action === 'suspend') { + condition += ' AND compute_metadata.imageArn = :imageArn AND compute_metadata.imageVersion = :imageVersion' + + ' AND compute_metadata.lifecycleProtocol = :protocol'; + values[':imageArn'] = snapshot.handle.imageArn; + values[':imageVersion'] = snapshot.handle.imageVersion; + values[':protocol'] = snapshot.handle.lifecycleProtocol; + } + if (snapshot.requestId === null) { + condition += ' AND attribute_not_exists(awaiting_approval_request_id)'; + } else { + condition += ' AND awaiting_approval_request_id = :request'; + values[':request'] = snapshot.requestId; + } + if (!snapshot.intent) { + condition += ' AND attribute_not_exists(microvm_lifecycle)'; + } else { + condition += ' AND microvm_lifecycle.generation = :generation'; + values[':generation'] = snapshot.intent.generation; + } + const command = new TransactWriteCommand({ + // Separate from the persistent generation: a repeated save has a different + // condition shape, so reusing that generation as an AWS token would conflict. + ClientRequestToken: randomUUID(), + TransactItems: [ + { + Update: { + TableName: TASK_TABLE, + Key: { task_id: snapshot.taskId }, + UpdateExpression: 'SET microvm_lifecycle = :intent', + ConditionExpression: condition, + ExpressionAttributeNames: names, + ExpressionAttributeValues: values, + }, + }, + ...(action === 'suspend' && snapshot.approval.kind === 'present' ? [{ + ConditionCheck: { + TableName: APPROVALS_TABLE, + Key: { task_id: snapshot.taskId, request_id: snapshot.requestId }, + ConditionExpression: '#status = :pending AND user_id = :user AND created_at = :created AND timeout_s = :timeout', + ExpressionAttributeNames: { '#status': 'status' }, + ExpressionAttributeValues: { + ':pending': 'PENDING', ':user': snapshot.userId, ':created': snapshot.approval.created_at, ':timeout': snapshot.approval.timeout_s, + }, + }, + }] : []), + ], + }); + try { + const abortSignal = storeSignal(options); + await ddb.send(command, { abortSignal }); + abortSignal.throwIfAborted(); + return { status: 'saved', intent }; + } catch (error) { + const failure = error as { name?: string; CancellationReasons?: { Code?: string }[] }; + if (failure?.name === 'TransactionCanceledException' + && failure.CancellationReasons?.some(reason => reason.Code === 'ConditionalCheckFailed')) return { status: 'stale' }; + // Lost committed reply: observe exactly our generation and unchanged task + // identity before reporting success. Unknown/unreadable outcomes stay errors. + const current = await readMicrovmLifecycleSnapshot(snapshot.taskId, snapshot.userId, options); + if (current?.intent?.generation === intent.generation) { + if (current.status === snapshot.status && current.requestId === snapshot.requestId + && current.handle.microvmId === snapshot.handle.microvmId && current.handle.endpoint === snapshot.handle.endpoint + && eligible(current, action, Date.now())) return { status: 'saved', intent: current.intent }; + // Our write committed, but cancellation/a decision moved the task on. + return { status: 'stale' }; + } + throw error; + } +} diff --git a/cdk/src/handlers/shared/microvm-start.ts b/cdk/src/handlers/shared/microvm-start.ts new file mode 100644 index 000000000..c259e9d01 --- /dev/null +++ b/cdk/src/handlers/shared/microvm-start.ts @@ -0,0 +1,247 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { createHash } from 'node:crypto'; +import { GetCommand, TransactWriteCommand, UpdateCommand } from '@aws-sdk/lib-dynamodb'; +import { canonicalJson } from './canonical-json'; +import type { SessionHandle } from './compute-strategy'; +import type { ContinuationRecord } from './microvm-continuation-types'; +import { MICROVM_IMAGE_CAPABILITY_REQUEST_TIMEOUT_MS, readMicrovmImageMetadata, supportsMicrovmLifecycle } from './microvm-image-capability'; +import { continuationEnabled, ensureWorkerLease, leaseHandleUpdate } from './microvm-worker-lease'; +import { makeDocClient } from './ua'; +import { TaskStatus, TERMINAL_STATUSES } from '../../constructs/task-status'; + +type MicrovmHandle = Extract; + +/** + * Internal TaskTable attribute, deliberately not part of the task API. + * Each authorized worker attempt owns one logical start. Only a coordinator + * continuation assignment may replace it; retries reuse the existing token. + */ +interface StartReceipt { + readonly clientToken: string; + readonly requestHash: string; + readonly createdAt: string; + readonly expiresAt: number; + readonly handle?: MicrovmHandle; +} + +interface StartRecord { + readonly user_id: string; + readonly status: string; + readonly session_id?: string; + readonly compute_type?: string; + readonly compute_metadata?: Record; + readonly microvm_start?: StartReceipt; + readonly repo?: string; + readonly continuation?: ContinuationRecord; +} + +export interface MicrovmStartClaim { + readonly clientToken: string; + readonly handle?: MicrovmHandle; + readonly closed: boolean; +} + +// A local retry limit, NOT a claim about AWS's undocumented token retention. +// After this window an unknown outcome requires investigation, never a fresh VM. +export const MICROVM_START_REPLAY_WINDOW_MS = 120_000; + +const ddb = makeDocClient(); +const TABLE_NAME = process.env.TASK_TABLE_NAME!; +const ACTIVE = new Set([TaskStatus.HYDRATING, TaskStatus.RUNNING, TaskStatus.AWAITING_APPROVAL]); + +/** Include the full S3 content, not just its URI; ignore object-key ordering. */ +export function microvmStartRequestHash(request: unknown, payload: unknown): string { + const canonical = canonicalJson([request, payload]); + return createHash('sha256').update(canonical).digest('hex'); +} + +function recordedHandle(record: StartRecord): MicrovmHandle | undefined { + const handle = record.microvm_start?.handle; + if (handle?.strategyType === 'lambda-microvm' && handle.microvmId && handle.endpoint + && handle.sessionId === handle.microvmId) return handle; + // Read handles written before start receipts existed, without starting again. + const metadata = record.compute_metadata; + if (record.compute_type === 'lambda-microvm' && record.session_id + && metadata?.microvmId === record.session_id && metadata.endpoint) { + return { + strategyType: 'lambda-microvm', + sessionId: record.session_id, + microvmId: metadata.microvmId, + endpoint: metadata.endpoint, + ...readMicrovmImageMetadata(metadata), + }; + } + return undefined; +} + +async function readStartRecord(taskId: string, userId: string): Promise { + const result = await ddb.send(new GetCommand({ + TableName: TABLE_NAME, + Key: { task_id: taskId }, + ConsistentRead: true, + })); + const record = result.Item as StartRecord | undefined; + if (!record || record.user_id !== userId) { + throw new Error('MICROVM_START_STATE_INVALID: task is missing or its owner does not match'); + } + return record; +} + +/** + * Establish identity and input immutability before S3 writes or RunMicrovm. + * Conditional creation makes competing invocations use the winning receipt. + */ +export async function claimMicrovmStart( + taskId: string, + userId: string, + requestHash: string, + attemptId: string = taskId, +): Promise { + for (let attempt = 0; attempt < 2; attempt++) { + const record = await readStartRecord(taskId, userId); + const handle = recordedHandle(record); + if (TERMINAL_STATUSES.some(status => status === record.status)) { + return { clientToken: attemptId, handle, closed: true }; + } + if (!ACTIVE.has(record.status)) { + throw new Error(`MICROVM_START_STATE_INVALID: cannot start a task in ${record.status}`); + } + if (attemptId !== taskId && record.continuation?.attempt_id !== attemptId) { + throw new Error('MICROVM_START_STATE_INVALID: replacement has no coordinator assignment'); + } + if (record.microvm_start && record.microvm_start.clientToken !== attemptId) { + throw new Error('MICROVM_START_STATE_INVALID: start belongs to another worker attempt'); + } + if (handle) return { clientToken: attemptId, handle, closed: false }; + const receipt = record.microvm_start; + if (receipt) { + if (receipt.clientToken !== attemptId || !Number.isFinite(receipt.expiresAt)) { + throw new Error('MICROVM_START_STATE_INVALID: invalid saved start receipt'); + } + if (receipt.requestHash !== requestHash) { + throw new Error('MICROVM_START_INPUT_CHANGED: refusing to overwrite or restart an earlier MicroVM request'); + } + if (Date.now() >= receipt.expiresAt) { + throw new Error('MICROVM_START_OUTCOME_UNKNOWN: replay window expired; inspect the original start before retrying'); + } + await ensureWorkerLease({ taskId, userId, repo: record.repo ?? '', attemptId, requestHash }); + return { clientToken: receipt.clientToken, closed: false }; + } + const replacement = attemptId !== taskId + && record.status === TaskStatus.AWAITING_APPROVAL && record.continuation?.state === 'STARTING'; + if (record.status !== TaskStatus.HYDRATING && !replacement) { + throw new Error('MICROVM_START_STATE_INVALID: active task has no recoverable start receipt or handle'); + } + const now = Date.now(); + const receiptToSave: StartReceipt = { + clientToken: attemptId, + requestHash, + createdAt: new Date(now).toISOString(), + expiresAt: now + MICROVM_START_REPLAY_WINDOW_MS, + }; + try { + await ddb.send(new UpdateCommand({ + TableName: TABLE_NAME, + Key: { task_id: taskId }, + UpdateExpression: 'SET microvm_start = :receipt', + ConditionExpression: '#status = :startingStatus AND user_id = :user AND attribute_not_exists(microvm_start)' + + (replacement ? ' AND continuation.#state = :starting AND continuation.attempt_id = :attempt' : ''), + ExpressionAttributeNames: { '#status': 'status', ...(replacement && { '#state': 'state' }) }, + ExpressionAttributeValues: { + ':receipt': receiptToSave, + ':startingStatus': replacement ? TaskStatus.AWAITING_APPROVAL : TaskStatus.HYDRATING, + ':user': userId, + ...(replacement && { ':starting': 'STARTING', ':attempt': attemptId }), + }, + })); + await ensureWorkerLease({ taskId, userId, repo: record.repo ?? '', attemptId, requestHash }); + return { clientToken: attemptId, closed: false }; + } catch (err) { + if ((err as { name?: string }).name !== 'ConditionalCheckFailedException') throw err; + // Re-read the winner, or observe cancellation, before any side effect. + } + } + throw new Error('MICROVM_START_STATE_INVALID: start receipt changed while being claimed'); +} + +/** Retain a known ID even if cancellation won while RunMicrovm was in flight. */ +export async function saveMicrovmStartHandle( + taskId: string, + clientToken: string, + handle: MicrovmHandle, +): Promise { + const replacement = clientToken !== taskId; + const update = { + TableName: TABLE_NAME, + Key: { task_id: taskId }, + UpdateExpression: 'SET microvm_start.#handle = :handle, session_id = :id, ' + + 'compute_type = :type, compute_metadata = :metadata' + + (replacement ? ', continuation.worker_id = :id, continuation.#state = :restoring' : ''), + ConditionExpression: 'microvm_start.clientToken = :token AND ' + + '(attribute_not_exists(microvm_start.#handle) OR microvm_start.#handle.microvmId = :id) AND ' + + '(attribute_not_exists(session_id) OR session_id = :id)' + + (replacement ? ' AND continuation.attempt_id = :token' : ''), + ExpressionAttributeNames: { '#handle': 'handle', ...(replacement && { '#state': 'state' }) }, + ExpressionAttributeValues: { + ':token': clientToken, + ':id': handle.microvmId, + ':handle': handle, + ':type': 'lambda-microvm', + ':metadata': { + microvmId: handle.microvmId, endpoint: handle.endpoint, ...readMicrovmImageMetadata(handle), + }, + ...(replacement && { ':restoring': 'RESTORING' }), + }, + }; + if (continuationEnabled()) { + await ddb.send(new TransactWriteCommand({ + TransactItems: [ + { Update: update }, leaseHandleUpdate(taskId, clientToken, handle.microvmId), + ], + })); + } else { + await ddb.send(new UpdateCommand(update)); + } +} + +/** Enrich only the same durably saved launch; never replace its identity or task state. */ +export async function saveMicrovmImageCapability( + taskId: string, clientToken: string, handle: MicrovmHandle, +): Promise { + if (!supportsMicrovmLifecycle(handle)) throw new Error('MicroVM image capability is incomplete'); + await ddb.send(new UpdateCommand({ + TableName: TABLE_NAME, + Key: { task_id: taskId }, + UpdateExpression: 'SET microvm_start.#handle.lifecycleProtocol = :protocol, compute_metadata.lifecycleProtocol = :protocol', + ConditionExpression: 'microvm_start.clientToken = :token AND session_id = :id AND ' + + 'microvm_start.#handle.microvmId = :id AND compute_metadata.microvmId = :id AND ' + + 'microvm_start.#handle.imageArn = :arn AND compute_metadata.imageArn = :arn AND ' + + 'microvm_start.#handle.imageVersion = :version AND compute_metadata.imageVersion = :version', + ExpressionAttributeNames: { '#handle': 'handle' }, + ExpressionAttributeValues: { + ':token': clientToken, + ':id': handle.microvmId, + ':arn': handle.imageArn, + ':version': handle.imageVersion, + ':protocol': handle.lifecycleProtocol, + }, + }), { abortSignal: AbortSignal.timeout(MICROVM_IMAGE_CAPABILITY_REQUEST_TIMEOUT_MS) }); +} diff --git a/cdk/src/handlers/shared/microvm-supervisor.ts b/cdk/src/handlers/shared/microvm-supervisor.ts new file mode 100644 index 000000000..b59f18616 --- /dev/null +++ b/cdk/src/handlers/shared/microvm-supervisor.ts @@ -0,0 +1,469 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +// SPDX-License-Identifier: MIT-0 + +import { evaluateAgentHeartbeat } from './agent-heartbeat'; +import type { ComputeStrategy, SessionControlOptions, SessionHandle, SessionStatus } from './compute-strategy'; +import { logger } from './logger'; +import { microvmErrorIdentity } from './microvm-control'; +import { + intentMatchesGate, readMicrovmLifecycleSnapshot, saveMicrovmLifecycleIntent, + type MicrovmLifecycleSnapshot, +} from './microvm-lifecycle'; +import { decideMicrovmLifecycle, MICROVM_TRANSITION_POLL_MS } from './microvm-lifecycle-policy'; +import { readMicrovmSuspendEnabled } from './microvm-suspend-config'; +import { MICROVM_MAX_DURATION_SECONDS } from './strategies/lambda-microvm-strategy'; +import { MICROVM_SLEEP_AFTER_S_DEFAULT, MICROVM_SLEEP_AFTER_S_MAX } from './types'; +import { TaskStatus, TERMINAL_STATUSES, type TaskStatusType } from '../../constructs/task-status'; + +type MicrovmHandle = Extract; +export const MICROVM_SUPERVISOR_CYCLE_MS = 45_000; +export const MICROVM_MAX_POLL_FAILURES = 3; +export const MICROVM_RECOVERY_TIMEOUT_MS = 120_000; +export const MICROVM_STARTUP_TIMEOUT_MS = 300_000; +export const MICROVM_CLEANUP_TIMEOUT_MS = 25_000; + +interface Recovery { + readonly kind: 'starting' | 'unconfirmed' | 'suspend' | 'wake'; + readonly sinceMs: number; +} + +/** JSON-only state retained by waitForCondition; a fresh Lambda must reuse it. */ +export interface MicrovmSupervisorState { + readonly version: 1; + readonly microvmId: string; + readonly firstObservedAtMs: number; + readonly sessionDeadlineMs: number; + readonly lifetimeVerified: boolean; + /** False until AWS confirms a post-startup state; absent in older saved state. */ + readonly startupConfirmed?: boolean; + readonly consecutivePollFailures: number; + readonly consecutiveResumeFailures: number; + readonly recovery?: Recovery; + readonly anomalyReported: boolean; + readonly nextPollInMs: number; + /** Persisted across replay so unchanged healthy polls do not repeat a log record. */ + readonly diagnosticSignature?: string; +} + +export interface MicrovmSupervisorInput { + readonly taskId: string; + readonly userId: string; + readonly handle: MicrovmHandle; + readonly strategy: ComputeStrategy; + readonly previous?: MicrovmSupervisorState; + readonly pollIntervalMs: number; + readonly suspendEnabled: boolean; + /** Optional shared cycle deadline when retirement precedes ordinary supervision. */ + readonly abortSignal?: AbortSignal; + /** Implementations must respect the supplied signal; event failure is best-effort. */ + readonly emitEvent?: (eventType: string, metadata: Record, options: SessionControlOptions) => Promise; +} + +type SupervisorOutcome = + | { readonly kind: 'continue' } + | { readonly kind: 'closed'; readonly status: TaskStatusType } + | { readonly kind: 'substrate-terminal' } + | { readonly kind: 'failure' | 'ownership-lost'; readonly reason: string }; + +export type MicrovmSupervisorResult = { + readonly state: MicrovmSupervisorState; + readonly snapshot?: MicrovmLifecycleSnapshot; + readonly substrate?: SessionStatus; + readonly deferHeartbeat: boolean; + readonly heartbeatUnhealthy: boolean; +} & SupervisorOutcome; + +function permanent(error: unknown): boolean { + return ['AccessDeniedException', 'ValidationException', 'UnrecognizedClientException', 'InvalidSignatureException'] + .includes(microvmErrorIdentity(error).error_type); +} + +function closed(status: TaskStatusType): boolean { + return TERMINAL_STATUSES.includes(status) || status === TaskStatus.FINALIZING; +} + +function sameWorker(snapshot: MicrovmLifecycleSnapshot | undefined, input: MicrovmSupervisorInput): snapshot is MicrovmLifecycleSnapshot { + return snapshot?.handle.microvmId === input.handle.microvmId + && snapshot.handle.sessionId === input.handle.sessionId && snapshot.handle.endpoint === input.handle.endpoint; +} + +/** + * One bounded poll cycle. Stores intent before control calls and rechecks after + * every outcome. Task finalization and compute cleanup remain the durable caller's + * responsibility; no task status or capacity reservation is changed here. + */ +export async function superviseMicrovm(input: MicrovmSupervisorInput): Promise { + const now = Date.now(); + if (!Number.isSafeInteger(input.pollIntervalMs) || input.pollIntervalMs <= 0) { + throw new Error('MicroVM supervision requires a positive poll interval'); + } + if (input.previous && (input.previous.version !== 1 || input.previous.microvmId !== input.handle.microvmId)) { + throw new Error('MicroVM supervisor state belongs to another worker or protocol'); + } + const prior = input.previous; + let state: MicrovmSupervisorState = prior ?? { + version: 1, + microvmId: input.handle.microvmId, + firstObservedAtMs: now, + sessionDeadlineMs: now + MICROVM_MAX_DURATION_SECONDS * 1000, + lifetimeVerified: false, + startupConfirmed: false, + consecutivePollFailures: 0, + consecutiveResumeFailures: 0, + anomalyReported: false, + nextPollInMs: input.pollIntervalMs, + }; + const cycleTimeout = AbortSignal.timeout(MICROVM_SUPERVISOR_CYCLE_MS); + const options: SessionControlOptions = { + abortSignal: input.abortSignal ? AbortSignal.any([input.abortSignal, cycleTimeout]) : cycleTimeout, + }; + let snapshot: MicrovmLifecycleSnapshot | undefined; + let substrate: SessionStatus | undefined; + let stage = 'task-read'; + let readFailed = false; + let suspendRequested = false; + let policyReason: string | undefined; + const result = (outcome: SupervisorOutcome): MicrovmSupervisorResult => { + if (outcome.kind === 'continue') { + if (!readFailed) {state = { ...state, consecutivePollFailures: 0 };} else if (state.consecutivePollFailures >= MICROVM_MAX_POLL_FAILURES) { + outcome = { kind: 'failure', reason: `${stage}-failed-repeatedly` }; + } + } + const deferHeartbeat = snapshot?.status === TaskStatus.AWAITING_APPROVAL || state.recovery !== undefined; + const diagnostic = { + observed_state: substrate?.microvmState ?? null, + task_status: snapshot?.status ?? null, + request_id: snapshot?.requestId ?? null, + approval_state: snapshot?.approval.kind === 'present' ? snapshot.approval.status : snapshot?.approval.kind ?? null, + intent_action: snapshot?.intent?.action ?? null, + generation: snapshot?.intent?.generation ?? null, + recovery_kind: state.recovery?.kind ?? null, + recovery_since_ms: state.recovery?.sinceMs ?? null, + outcome: outcome.kind, + reason: 'reason' in outcome ? outcome.reason : null, + policy_reason: policyReason ?? null, + sleep_after_s: snapshot?.sleepAfterSeconds === undefined ? MICROVM_SLEEP_AFTER_S_DEFAULT + : Number.isInteger(snapshot.sleepAfterSeconds) && snapshot.sleepAfterSeconds >= 0 + && snapshot.sleepAfterSeconds <= MICROVM_SLEEP_AFTER_S_MAX ? snapshot.sleepAfterSeconds : null, + consecutive_poll_failures: state.consecutivePollFailures, + consecutive_resume_failures: state.consecutiveResumeFailures, + }; + const signature = JSON.stringify(diagnostic); + if (signature !== state.diagnosticSignature) { + const metadata = { + task_id: input.taskId, + microvm_id: input.handle.microvmId, + image_arn: input.handle.imageArn, + image_version: input.handle.imageVersion, + ...diagnostic, + intent_requested_at_ms: snapshot?.intent?.requested_at_ms, + approval_deadline_ms: snapshot?.approval.kind === 'present' ? snapshot.approval.deadlineMs : undefined, + session_deadline_ms: state.sessionDeadlineMs, + first_observed_at_ms: state.firstObservedAtMs, + startup_confirmed: state.startupConfirmed, + cycle_elapsed_ms: Date.now() - now, + }; + if (outcome.kind === 'failure' || outcome.kind === 'ownership-lost' || outcome.kind === 'substrate-terminal') { + logger.warn('MicroVM supervisor observation changed', metadata); + } else { + logger.info('MicroVM supervisor observation changed', metadata); + } + state = { ...state, diagnosticSignature: signature }; + } + return { + ...outcome, + state, + snapshot, + substrate, + deferHeartbeat, + heartbeatUnhealthy: !deferHeartbeat && snapshot?.status === TaskStatus.RUNNING + && evaluateAgentHeartbeat(snapshot.taskStartedAtMs, snapshot.heartbeatAtMs, Date.now()) !== undefined, + }; + }; + const retry = () => { + state = { ...state, nextPollInMs: Math.min(MICROVM_TRANSITION_POLL_MS, input.pollIntervalMs) }; + return result({ kind: 'continue' }); + }; + const report = async (eventType: string, detail: Record) => { + const metadata = { + task_id: input.taskId, + microvm_id: input.handle.microvmId, + ...(snapshot && { request_id: snapshot.requestId }), + ...detail, + }; + logger.warn(eventType, metadata); + if (input.emitEvent) { + try { await input.emitEvent(eventType, metadata, options); } catch (error) { + logger.warn('MicroVM lifecycle audit failed', { ...metadata, ...microvmErrorIdentity(error) }); + } + } + }; + const fresh = async () => { + options.abortSignal!.throwIfAborted(); + const current = await readMicrovmLifecycleSnapshot(input.taskId, input.userId, options); + options.abortSignal!.throwIfAborted(); + return current; + }; + const beginRecovery = (kind: Recovery['kind'], sinceMs = Date.now()) => { + state = { + ...state, + recovery: state.recovery?.kind === kind ? state.recovery : { kind, sinceMs }, + }; + }; + const recoveryExpired = () => state.recovery !== undefined + && Date.now() - state.recovery.sinceMs >= (state.recovery.kind === 'starting' + ? MICROVM_STARTUP_TIMEOUT_MS : MICROVM_RECOVERY_TIMEOUT_MS); + + try { + snapshot = await fresh(); + if (!sameWorker(snapshot, input)) return result({ kind: 'ownership-lost', reason: 'worker-record-changed-or-missing' }); + if (closed(snapshot.status)) { + if (snapshot.status === TaskStatus.FINALIZING && Date.now() >= state.sessionDeadlineMs) { + return result({ kind: 'failure', reason: 'session-deadline' }); + } + state = { ...state, nextPollInMs: Math.max(1, Math.min(input.pollIntervalMs, state.sessionDeadlineMs - Date.now())) }; + return result({ kind: 'closed', status: snapshot.status }); + } + stage = 'substrate-read'; + substrate = await input.strategy.pollSession(input.handle, options); + options.abortSignal!.throwIfAborted(); + const startedAt = substrate.microvmStartedAtMs; + const maximum = substrate.microvmMaximumDurationSeconds; + if (Number.isSafeInteger(startedAt) && startedAt! >= 0 + && Number.isSafeInteger(maximum) && maximum! > 0) { + const observedDeadline = startedAt! + Math.min(maximum!, MICROVM_MAX_DURATION_SECONDS) * 1000; + if (Number.isSafeInteger(observedDeadline)) { + state = { ...state, lifetimeVerified: true, sessionDeadlineMs: Math.min(state.sessionDeadlineMs, observedDeadline) }; + } + } + const approvalFailure = snapshot!.approval.kind === 'unavailable'; + readFailed = approvalFailure; + state = { + ...state, + consecutivePollFailures: approvalFailure ? state.consecutivePollFailures + 1 : state.consecutivePollFailures, + }; + const observed = substrate.microvmState ?? 'UNKNOWN'; + if (observed === 'RUNNING' || observed === 'SUSPENDING' || observed === 'SUSPENDED') { + state = { ...state, startupConfirmed: true }; + } + const wakeRepairWanted = state.recovery?.kind === 'wake' + && (!intentMatchesGate(snapshot!) || snapshot!.intent?.action !== 'resume'); + const awaitingDecisionConsumption = snapshot.status === TaskStatus.AWAITING_APPROVAL + && (snapshot.approval.kind !== 'present' || snapshot.approval.status !== 'PENDING' + || (snapshot.approval.deadlineMs !== null && Date.now() >= snapshot.approval.deadlineMs)); + if (observed === 'RUNNING') { + // AWS RUNNING does not prove the guest consumed a decided/expired gate. + // Recovery ends after fresh guest liveness, or an intentional early wake + // where the original human decision is still pending. + if (state.recovery?.kind !== 'wake' || (!wakeRepairWanted && ( + (snapshot.status === TaskStatus.RUNNING && (snapshot.heartbeatAtMs ?? -1) >= state.recovery.sinceMs) + || (snapshot.status === TaskStatus.AWAITING_APPROVAL && !awaitingDecisionConsumption) + ))) { + state = { ...state, recovery: undefined, consecutiveResumeFailures: 0 }; + } + } else if (observed === 'PENDING' || observed === 'UNKNOWN') { + // An uncertain observation cannot reset an in-flight wake/suspend clock. + if (!state.recovery) { + if (intentMatchesGate(snapshot) && snapshot.intent?.action === 'resume') { + // AWS can report PENDING while restoring an already-running worker. + // An API-triggered wake may arrive between supervisor polls; retain its + // saved start time instead of reusing the worker's original boot clock. + beginRecovery('wake', Math.min(Date.now(), snapshot.intent.requested_at_ms)); + } else if (observed === 'PENDING' + && (state.startupConfirmed === false || snapshot.status === TaskStatus.HYDRATING)) { + // The coordinator marks the task RUNNING before AWS finishes startup. + // Retain first-observation age across failed initial reads and replay. + beginRecovery('starting', state.firstObservedAtMs); + } else { + beginRecovery('unconfirmed'); + } + } + } + + let suspendEnabled = input.suspendEnabled && state.lifetimeVerified; + const policy = () => decideMicrovmLifecycle({ + snapshot: snapshot!, + substrate: substrate!, + nowMs: Date.now(), + sessionDeadlineMs: state.sessionDeadlineMs, + pollIntervalMs: input.pollIntervalMs, + suspendEnabled, + }); + let decision = policy(); + if ((wakeRepairWanted || (state.recovery?.kind === 'wake' + && (observed === 'SUSPENDING' || observed === 'SUSPENDED'))) + && (decision.action === 'wait' || decision.action === 'suspend') + && (observed === 'RUNNING' || observed === 'SUSPENDING' || observed === 'SUSPENDED')) { + // A previous cycle may have lost the wake-intent write. Its durable recovery + // state still forbids leaving the worker asleep while that write is retried. + decision = { + action: 'resume', + requestReady: observed === 'SUSPENDED', + reason: 'wake-recovery', + nextPollInMs: MICROVM_TRANSITION_POLL_MS, + }; + } + if (decision.action === 'suspend') { + suspendEnabled = await readMicrovmSuspendEnabled(options); + decision = policy(); + } + policyReason = decision.reason; + state = { ...state, nextPollInMs: decision.nextPollInMs }; + if (decision.action === 'reconcile-terminal') return result({ kind: 'substrate-terminal' }); + if (decision.action === 'terminate') return result({ kind: 'failure', reason: decision.reason }); + if (snapshot.status === TaskStatus.HYDRATING && Date.now() - state.firstObservedAtMs >= MICROVM_STARTUP_TIMEOUT_MS) { + return result({ kind: 'failure', reason: 'startup-deadline' }); + } + + const anomaly = (observed === 'SUSPENDING' || observed === 'SUSPENDED') + && state.recovery?.kind !== 'wake' + && !(snapshot.status === TaskStatus.AWAITING_APPROVAL && intentMatchesGate(snapshot) && snapshot.intent?.action === 'resume') + && (snapshot.status !== TaskStatus.AWAITING_APPROVAL || !intentMatchesGate(snapshot)); + if (anomaly && !state.anomalyReported) await report('microvm_suspend_anomaly', { reason: decision.reason }); + state = { ...state, anomalyReported: anomaly }; + + if (decision.action === 'wait') { + if (observed === 'SUSPENDING') beginRecovery('suspend', snapshot!.intent?.requested_at_ms ?? Date.now()); + if (observed === 'SUSPENDED') state = { ...state, recovery: undefined }; + if (recoveryExpired()) return result({ kind: 'failure', reason: 'recovery-deadline' }); + if (approvalFailure && state.consecutivePollFailures >= MICROVM_MAX_POLL_FAILURES) { + return result({ kind: 'failure', reason: 'approval-read-failed-repeatedly' }); + } + return result({ kind: 'continue' }); + } + + stage = `${decision.action}-intent`; + if (decision.action === 'resume' + && (observed === 'SUSPENDING' || observed === 'SUSPENDED' || awaitingDecisionConsumption || wakeRepairWanted)) { + beginRecovery('wake'); + } + const saved = await saveMicrovmLifecycleIntent(snapshot!, decision.action, Date.now(), options); + if (saved.status !== 'saved') return retry(); + snapshot = { ...snapshot!, intent: saved.intent }; + + if (decision.action === 'resume') { + if (observed === 'SUSPENDING' || observed === 'SUSPENDED') beginRecovery('wake'); + if (recoveryExpired()) return result({ kind: 'failure', reason: 'wake-deadline' }); + if (!decision.requestReady) return retry(); + } else { + beginRecovery('suspend', saved.intent.requested_at_ms); + // An uncompleted attempt while still RUNNING is a lost saving opportunity. + // Fence it with wake intent, including a delayed service-side suspension. + if (recoveryExpired()) { + beginRecovery('wake'); + await saveMicrovmLifecycleIntent(snapshot, 'resume', Date.now(), options); + return retry(); + } + } + + // Recheck the live switch after saving intent, then refresh the gate. An + // immutable Lambda environment alone cannot disable an existing execution. + if (decision.action === 'suspend') suspendEnabled = await readMicrovmSuspendEnabled(options); + // Database success is not a lock over the next AWS request. + stage = 'pre-command-read'; + snapshot = await fresh(); + if (!sameWorker(snapshot, input)) return result({ kind: 'ownership-lost', reason: 'worker-record-changed-or-missing' }); + if (closed(snapshot!.status)) return result({ kind: 'closed', status: snapshot!.status }); + if (snapshot!.intent?.generation !== saved.intent.generation || snapshot.requestId !== saved.intent.request_id) return retry(); + const latestDecision = policy(); + policyReason = latestDecision.reason; + if (decision.action === 'suspend' && latestDecision.action !== 'suspend') { + // Approval/deadline/disable can win after intent was saved but before the call. + beginRecovery('wake'); + await saveMicrovmLifecycleIntent(snapshot!, 'resume', Date.now(), options); + return retry(); + } + + stage = `${decision.action}-request`; + suspendRequested = decision.action === 'suspend'; + let commandError: unknown; + try { + const acknowledgement = decision.action === 'suspend' + ? await input.strategy.suspendSession(input.handle, options) + : await input.strategy.resumeSession(input.handle, options); + if (!acknowledgement.supported) commandError = new Error('Lifecycle request is unsupported'); + } catch (error) { commandError = error; } + if (commandError) { + await report(`microvm_${decision.action}_request_failed`, { stage, ...microvmErrorIdentity(commandError) }); + } + + stage = 'post-command-read'; + snapshot = await fresh(); + if (!sameWorker(snapshot, input)) return result({ kind: 'ownership-lost', reason: 'worker-record-changed-or-missing' }); + if (closed(snapshot!.status)) return result({ kind: 'closed', status: snapshot!.status }); + if (decision.action === 'suspend' && (commandError || policy().action !== 'suspend')) { + // Even a failed command can have committed. Retain wake until fresh AWS + // observations confirm recovery; never erase it after an acknowledgment. + beginRecovery('wake'); + await saveMicrovmLifecycleIntent(snapshot!, 'resume', Date.now(), options); + } + if (decision.action === 'resume') { + state = { ...state, consecutiveResumeFailures: commandError ? state.consecutiveResumeFailures + 1 : 0 }; + if (commandError && (permanent(commandError) || state.consecutiveResumeFailures >= MICROVM_MAX_POLL_FAILURES)) { + return result({ kind: 'failure', reason: 'resume-request-failed-repeatedly' }); + } + } + return retry(); + } catch (error) { + // A lost post-command read/write cannot prove the worker stayed awake. + // Persist this recovery obligation even when the wake-intent write failed. + if (suspendRequested) beginRecovery('wake'); + state = { ...state, consecutivePollFailures: state.consecutivePollFailures + (readFailed ? 0 : 1) }; + readFailed = true; + await report('microvm_supervisor_request_failed', { + stage, consecutive_failures: state.consecutivePollFailures, ...microvmErrorIdentity(error), + }); + if (permanent(error) || state.consecutivePollFailures >= MICROVM_MAX_POLL_FAILURES + || recoveryExpired() || Date.now() >= state.sessionDeadlineMs) { + return result({ kind: 'failure', reason: `${stage}-failed` }); + } + return retry(); + } +} + +/** Bounded best-effort cleanup; an unconfirmed outcome stays visible with its handle. */ +export async function stopMicrovmWithDiagnostics( + input: Pick, +): Promise { + const options: SessionControlOptions = { abortSignal: AbortSignal.timeout(MICROVM_CLEANUP_TIMEOUT_MS) }; + let failure = { error_type: 'NoStopEvidence' } as ReturnType; + for (let attempt = 0; attempt < 2; attempt++) { + try { + options.abortSignal!.throwIfAborted(); + const stopped = await input.strategy.stopSession(input.handle, options); + if (stopped && stopped.outcome !== 'unconfirmed') return; + if (stopped?.outcome === 'unconfirmed') failure = stopped; + } catch (error) { failure = microvmErrorIdentity(error); } + if (['AccessDeniedException', 'ValidationException'].includes(failure.error_type)) break; + } + const metadata = { + task_id: input.taskId, + microvm_id: input.handle.microvmId, + error_type: failure.error_type, + ...(failure.aws_request_id && { aws_request_id: failure.aws_request_id }), + }; + logger.error('MicroVM cleanup unconfirmed; retained handle requires recovery', metadata); + try { + await input.emitEvent?.('microvm_cleanup_unconfirmed', metadata, options); + } catch (error) { + logger.warn('MicroVM cleanup audit failed', { ...metadata, ...microvmErrorIdentity(error) }); + } +} diff --git a/cdk/src/handlers/shared/microvm-suspend-config.ts b/cdk/src/handlers/shared/microvm-suspend-config.ts new file mode 100644 index 000000000..bb828817e --- /dev/null +++ b/cdk/src/handlers/shared/microvm-suspend-config.ts @@ -0,0 +1,50 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { GetParameterCommand, SSMClient } from '@aws-sdk/client-ssm'; +import type { SessionControlOptions } from './compute-strategy'; +import { logger } from './logger'; +import { microvmErrorIdentity } from './microvm-control'; +import { makeClient } from './ua'; + +export const MICROVM_SUSPEND_CONFIG_TIMEOUT_MS = 3_000; +let client: SSMClient | undefined; + +/** + * Durable executions pin their function version, including its environment. + * Read a stable shared parameter without caching so existing executions can + * observe disable. An unavailable setting loses savings, not a healthy worker. + */ +export async function readMicrovmSuspendEnabled(options: SessionControlOptions): Promise { + const name = process.env.MICROVM_APPROVAL_SUSPEND_PARAMETER_NAME; + if (!name) return false; + const localSignal = AbortSignal.timeout(MICROVM_SUSPEND_CONFIG_TIMEOUT_MS); + const abortSignal = options.abortSignal + ? AbortSignal.any([options.abortSignal, localSignal]) : localSignal; + try { + abortSignal.throwIfAborted(); + client ??= makeClient(SSMClient, { maxAttempts: 1 }); + const response = await client.send(new GetParameterCommand({ Name: name }), { abortSignal }); + abortSignal.throwIfAborted(); + return response.Parameter?.Name === name && response.Parameter.Value === 'true'; + } catch (error) { + logger.warn('MicroVM suspension setting unavailable; new suspension disabled', microvmErrorIdentity(error)); + return false; + } +} diff --git a/cdk/src/handlers/shared/microvm-task-poll.ts b/cdk/src/handlers/shared/microvm-task-poll.ts new file mode 100644 index 000000000..c4cbf04a2 --- /dev/null +++ b/cdk/src/handlers/shared/microvm-task-poll.ts @@ -0,0 +1,106 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +const RETIREMENT_EVENT_TIMEOUT_MS = 3000; +import { formatMicrovmTerminalFailure } from './error-classifier'; +import { logger } from './logger'; +import { retireCheckpointedMicrovm, type RetirementResult } from './microvm-continuation-retirement'; +import { microvmErrorIdentity } from './microvm-control'; +import { MICROVM_SUPERVISOR_CYCLE_MS, superviseMicrovm, type MicrovmSupervisorInput } from './microvm-supervisor'; +import type { PollState } from './orchestrator'; +import { TaskStatus } from '../../constructs/task-status'; + +function retirementState(state: PollState, result: RetirementResult): PollState | undefined { + if (result === 'not-due') return undefined; + return { + attempts: state.attempts + 1, + lastStatus: TaskStatus.AWAITING_APPROVAL, + microvmSupervisor: state.microvmSupervisor, + microvmParked: result === 'parked', + microvmRetiring: result === 'stopping', + microvmOwnershipLost: result === 'ownership-lost', + }; +} + +async function retirementCycle( + input: Omit, state: PollState, force = false, +): Promise { + try { + const retirement = await retireCheckpointedMicrovm({ + ...input, sessionDeadlineMs: state.microvmSupervisor?.sessionDeadlineMs ?? Infinity, force, + }); + return retirementState(state, retirement); + } catch (error) { + // Fencing/termination may have committed before a lost reply. Keep the + // reservation held and retry the persisted transition; never resume tools + // or mark the task complete based on an uncertain control outcome. + const identity = microvmErrorIdentity(error); + logger.warn('MicroVM continuation retirement needs reconciliation', { + task_id: input.taskId, microvm_id: input.handle.microvmId, ...identity, + }); + if (state.microvmRetirementError !== identity.error_type) { + try { + await input.emitEvent?.('continuation_retirement_delayed', { + microvm_id: input.handle.microvmId, + error_id: identity.error_type, + detail: 'The saved task is waiting for worker shutdown or storage confirmation. Its capacity reservation remains held.', + }, { abortSignal: AbortSignal.timeout(RETIREMENT_EVENT_TIMEOUT_MS) }); + } catch (eventError) { + logger.warn('Could not publish continuation reconciliation feedback', { + task_id: input.taskId, ...microvmErrorIdentity(eventError), + }); + } + } + return { + attempts: state.attempts + 1, + lastStatus: state.lastStatus, + microvmSupervisor: state.microvmSupervisor, + microvmRetiring: true, + microvmRetirementError: identity.error_type, + }; + } +} + +/** Shared by initial and replacement durable executions. Never follow another handle. */ +export async function pollMicrovmTask(input: Omit, state: PollState): Promise { + const timeout = AbortSignal.timeout(MICROVM_SUPERVISOR_CYCLE_MS); + input = { ...input, abortSignal: input.abortSignal ? AbortSignal.any([input.abortSignal, timeout]) : timeout }; + const retiring = await retirementCycle(input, state); + if (retiring) return retiring; + const supervised = await superviseMicrovm({ ...input, previous: state.microvmSupervisor }); + if (supervised.kind === 'substrate-terminal' || supervised.kind === 'failure') { + // A complete pending checkpoint can outlive a dead worker. No tool ran + // beyond that barrier, and the original human request/decision is retained. + const recovered = await retirementCycle(input, { ...state, microvmSupervisor: supervised.state }, true); + if (recovered) return recovered; + } + const failure = supervised.kind === 'failure' ? supervised.reason + : supervised.kind === 'substrate-terminal' ? 'substrate-terminal' : undefined; + return { + attempts: state.attempts + 1, + lastStatus: supervised.snapshot?.status ?? state.lastStatus, + sessionUnhealthy: supervised.heartbeatUnhealthy, + microvmSupervisor: supervised.state, + microvmFailureReason: failure, + microvmFailureMessage: supervised.kind === 'substrate-terminal' ? formatMicrovmTerminalFailure( + `substrate state ${supervised.substrate!.status}`, supervised.substrate!.reason, + ) : undefined, + microvmOwnershipLost: supervised.kind === 'ownership-lost', + }; +} diff --git a/cdk/src/handlers/shared/microvm-worker-lease.ts b/cdk/src/handlers/shared/microvm-worker-lease.ts new file mode 100644 index 000000000..75c138b47 --- /dev/null +++ b/cdk/src/handlers/shared/microvm-worker-lease.ts @@ -0,0 +1,93 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { GetCommand, TransactWriteCommand } from '@aws-sdk/lib-dynamodb'; +import { workerLeaseKey, type WorkerLease } from './microvm-continuation-types'; +import { makeDocClient } from './ua'; + +const ddb = makeDocClient(); +const TABLE = process.env.TASK_TABLE_NAME!; + +export function continuationEnabled(): boolean { + return Boolean(process.env.CONTINUATION_BUCKET_NAME); +} + +/** Create authority once. Replays must never reactivate a fenced worker. */ +export async function ensureWorkerLease(input: { + taskId: string; userId: string; repo: string; attemptId: string; requestHash: string; +}): Promise { + if (!continuationEnabled()) return; + const lease: WorkerLease = { + ...workerLeaseKey(input.taskId), + lease_attempt_id: input.attemptId, + lease_state: 'ACTIVE', + lease_user_id: input.userId, + lease_repo: input.repo, + }; + try { + await ddb.send(new TransactWriteCommand({ + TransactItems: [ + { + ConditionCheck: { + TableName: TABLE, + Key: { task_id: input.taskId }, + ConditionExpression: 'user_id = :user AND microvm_start.clientToken = :attempt ' + + 'AND microvm_start.requestHash = :hash AND #status IN (:hydrating, :awaiting)', + ExpressionAttributeNames: { '#status': 'status' }, + ExpressionAttributeValues: { + ':user': input.userId, + ':attempt': input.attemptId, + ':hash': input.requestHash, + ':hydrating': 'HYDRATING', + ':awaiting': 'AWAITING_APPROVAL', + }, + }, + }, + { + Put: { + TableName: TABLE, + Item: lease, + ConditionExpression: 'attribute_not_exists(task_id)', + }, + }, + ], + })); + } catch (error) { + const result = await ddb.send(new GetCommand({ + TableName: TABLE, Key: workerLeaseKey(input.taskId), ConsistentRead: true, + })); + const current = result.Item as WorkerLease | undefined; + if (current?.lease_state !== 'ACTIVE' || current.lease_attempt_id !== input.attemptId + || current.lease_user_id !== input.userId || current.lease_repo !== input.repo) throw error; + } +} + +/** Bind a returned physical handle without altering whether the lease is fenced. */ +export function leaseHandleUpdate(taskId: string, attemptId: string, microvmId: string) { + return { + Update: { + TableName: TABLE, + Key: workerLeaseKey(taskId), + UpdateExpression: 'SET lease_microvm_id = :id', + ConditionExpression: 'lease_attempt_id = :attempt AND ' + + '(attribute_not_exists(lease_microvm_id) OR lease_microvm_id = :id)', + ExpressionAttributeValues: { ':attempt': attemptId, ':id': microvmId }, + }, + }; +} diff --git a/cdk/src/handlers/shared/orchestrator.ts b/cdk/src/handlers/shared/orchestrator.ts index 81685e6d7..23047fed5 100644 --- a/cdk/src/handlers/shared/orchestrator.ts +++ b/cdk/src/handlers/shared/orchestrator.ts @@ -20,10 +20,14 @@ import { S3Client } from '@aws-sdk/client-s3'; import { GetCommand, PutCommand, UpdateCommand } from '@aws-sdk/lib-dynamodb'; import { ulid } from 'ulid'; -import type { SessionHandle, SessionStatus } from './compute-strategy'; +import { evaluateAgentHeartbeat } from './agent-heartbeat'; +import { closeTaskApprovals } from './close-task-approvals'; +import type { SessionControlOptions, SessionHandle } from './compute-strategy'; import { AttachmentBudgetExceededError, AttachmentConfigurationError, AttachmentResolutionError, hydrateContext, resolveGitHubToken } from './context-hydration'; import { logger, type Logger } from './logger'; import { writeMinimalEpisode } from './memory'; +import { readMicrovmImageMetadata } from './microvm-image-capability'; +import type { MicrovmSupervisorState } from './microvm-supervisor'; import { coerceNumericOrNull } from './numeric'; import { computePromptVersion } from './prompt-version'; import { makeRegistryClient } from './registry/factory'; @@ -31,6 +35,7 @@ import { parseRef } from './registry/ref'; import { RegistryResolutionError, type ResolvedAsset } from './registry/types'; import { loadRepoConfig, type BlueprintConfig, type ComputeType } from './repo-config'; import { resolveUrlAttachments } from './resolve-url-attachments'; +import { acquireTaskSlot, releaseTaskSlot } from './task-concurrency'; import { APPROVAL_GATE_CAP_MAX, APPROVAL_GATE_CAP_MIN, type AgentAttachmentPayload, type AttachmentRecord, type TaskRecord } from './types'; import { makeClient, makeDocClient } from './ua'; import { computeTtlEpoch, DEFAULT_MAX_TURNS } from './validation'; @@ -40,7 +45,6 @@ const ddb = makeDocClient(); const TABLE_NAME = process.env.TASK_TABLE_NAME!; const EVENTS_TABLE_NAME = process.env.TASK_EVENTS_TABLE_NAME!; -const CONCURRENCY_TABLE_NAME = process.env.USER_CONCURRENCY_TABLE_NAME!; const RUNTIME_ARN = process.env.RUNTIME_ARN!; const MAX_CONCURRENT = Number(process.env.MAX_CONCURRENT_TASKS_PER_USER ?? '3'); const TASK_RETENTION_DAYS = Number(process.env.TASK_RETENTION_DAYS ?? '90'); @@ -62,25 +66,17 @@ export interface PollState { readonly consecutiveEcsPollFailures?: number; /** Consecutive polls where ECS reports completed but DDB is not terminal — escalated after 5. */ readonly consecutiveEcsCompletedPolls?: number; - /** - * True once `microvm_suspend_anomaly` has been emitted for the CURRENT anomaly - * episode, so the event fires once per episode instead of on every ~30 s poll - * (an 8-hour suspended task would otherwise write ~960 identical events). - * - * Re-armed (set back to false) by any non-anomalous observation — see - * {@link reconcileMicrovmSubstrateState}. Kept as a plain boolean rather than a - * counter/timestamp on purpose: P3's suspend policy will reshape this area - * anyway, and one flag is the smallest thing that fixes the duplication without - * pre-committing to a shape that work will have to undo. - */ - readonly microvmSuspendAnomalyReported?: boolean; + readonly microvmSupervisor?: MicrovmSupervisorState; + readonly microvmFailureReason?: string; + readonly microvmFailureMessage?: string; + /** A stale execution must clean up only its own handle, leaving the replacement alone. */ + readonly microvmOwnershipLost?: boolean; + /** The worker is stopped and its reservation released; the approval remains open. */ + readonly microvmParked?: boolean; + readonly microvmRetiring?: boolean; + readonly microvmRetirementError?: string; } -/** After RUNNING this long, we expect `agent_heartbeat_at` from the agent (if ever set). */ -const AGENT_HEARTBEAT_GRACE_SEC = 120; -/** If `agent_heartbeat_at` exists and is older than this, the session is treated as lost. */ -const AGENT_HEARTBEAT_STALE_SEC = 240; - /** * Whether a backend's liveness is (partly) inferred from `agent_heartbeat_at`. * @@ -96,14 +92,12 @@ const AGENT_HEARTBEAT_STALE_SEC = 240; * `AgentCoreComputeStrategy.pollSession` is an explicit stub that always * reports `running`, so a crashed container is invisible without the heartbeat. * - `lambda-microvm` — yes, and it is the SECOND of two complementary signals - * (ADR-021 P2). The substrate `GetMicrovm` check catches a VM that DIED; it - * cannot catch a VM that is alive and healthy while the in-guest pipeline is - * hung, deadlocked, or OOM-killed inside the guest — nothing self-terminates on - * this substrate (live-verified: a MicroVM with a broken hook sat in `RUNNING` - * indefinitely with no `stateReason`). Without the heartbeat check such a task - * would burn the full ~8.5 h poll window, billing an 8-hour MicroVM - * reservation, before the safety net fired. So liveness here is substrate state - * AND agent heartbeat. + * (ADR-021 P2). `GetMicrovm` reports VM state; a stale heartbeat detects loss + * of the in-guest heartbeat writer even if the VM still reports RUNNING. + * The writer runs in a separate thread, so it can continue during a pipeline + * deadlock: neither signal proves that coding work is progressing. A failed + * run hook can cause service teardown; after an accepted hook, explicit + * termination and the maximum session duration remain cleanup safeguards. * - `ecs` — no, and this is a HARD CORRECTNESS CONSTRAINT, not a tuning * preference. The ECS boot command (`ecs-strategy.ts`) invokes * `run_task_from_payload` directly and "bypasses the uvicorn server entirely", @@ -176,10 +170,11 @@ function substrateNoun(computeType: ComputeType | undefined): string { * @returns the task record. * @throws Error if the task is not found. */ -export async function loadTask(taskId: string): Promise { +export async function loadTask(taskId: string, consistentRead = false): Promise { const result = await ddb.send(new GetCommand({ TableName: TABLE_NAME, Key: { task_id: taskId }, + ...(consistentRead && { ConsistentRead: true }), })); if (!result.Item) { throw new Error(`Task ${taskId} not found`); @@ -188,31 +183,12 @@ export async function loadTask(taskId: string): Promise { } /** - * Admission control: check user concurrency and increment counter. + * Admission control: acquire or recover this task's capacity reservation. * @param task - the task record. * @returns true if admitted, false if concurrency limit reached. */ export async function admissionControl(task: TaskRecord): Promise { - try { - await ddb.send(new UpdateCommand({ - TableName: CONCURRENCY_TABLE_NAME, - Key: { user_id: task.user_id }, - UpdateExpression: 'SET active_count = if_not_exists(active_count, :zero) + :one, updated_at = :now', - ConditionExpression: 'attribute_not_exists(active_count) OR active_count < :max', - ExpressionAttributeValues: { - ':zero': 0, - ':one': 1, - ':max': MAX_CONCURRENT, - ':now': new Date().toISOString(), - }, - })); - return true; - } catch (err: unknown) { - if (err && typeof err === 'object' && 'name' in err && err.name === 'ConditionalCheckFailedException') { - return false; - } - throw err; - } + return acquireTaskSlot(task.task_id, task.user_id, MAX_CONCURRENT); } /** @@ -227,6 +203,7 @@ export async function transitionTask( fromStatus: TaskStatusType, toStatus: TaskStatusType, extraAttrs?: Record, + expectedMicrovmId?: string, ): Promise { const validTargets = VALID_TRANSITIONS[fromStatus]; if (!validTargets.includes(toStatus)) { @@ -263,11 +240,19 @@ export async function transitionTask( } } + let condition = fromStatus === TaskStatus.SUBMITTED && toStatus === TaskStatus.QUEUED + ? '#status = :fromStatus AND attribute_not_exists(concurrency_slot)' + : '#status = :fromStatus'; + if (expectedMicrovmId) { + condition += ' AND compute_type = :microvmType AND session_id = :microvmId AND compute_metadata.microvmId = :microvmId'; + expressionValues[':microvmType'] = 'lambda-microvm'; + expressionValues[':microvmId'] = expectedMicrovmId; + } await ddb.send(new UpdateCommand({ TableName: TABLE_NAME, Key: { task_id: taskId }, UpdateExpression: updateExpression, - ConditionExpression: '#status = :fromStatus', + ConditionExpression: condition, ExpressionAttributeNames: expressionNames, ExpressionAttributeValues: expressionValues, })); @@ -312,7 +297,9 @@ export async function emitTaskEvent( eventType: string, metadata?: Record, correlation?: EventCorrelation, + options?: SessionControlOptions, ): Promise { + options?.abortSignal?.throwIfAborted(); await ddb.send(new PutCommand({ TableName: EVENTS_TABLE_NAME, Item: { @@ -325,7 +312,7 @@ export async function emitTaskEvent( ...(correlation?.repo && { repo: correlation.repo }), ...(metadata && { metadata }), }, - })); + }), options); } /** Minimum allowed poll interval (5 seconds). */ @@ -336,8 +323,8 @@ const MAX_POLL_INTERVAL_MS = 300_000; /** * Build the ``compute_metadata`` map persisted on the task row at session start, * so a later handler can act on the right backend without re-deriving anything: - * ``cancel-task.ts`` reads ``clusterArn``/``taskArn`` from it today, and ADR-021 - * sub-decision 2 has the approve/deny Lambdas read ``microvmId`` from it in P3. + * cancellation uses the backend's worker identity, and MicroVM approval wake + * uses ``microvmId`` plus the saved image identity. * * Kept as an exhaustive switch (not a ternary + spread) so a fourth backend is a * compile error here rather than a silently empty metadata map — the field is @@ -357,7 +344,7 @@ export function buildComputeMetadata(handle: SessionHandle): Record { - const { - taskId, ddbStatus, substrate, microvmId, userId, correlation, log, repo, - suspendAnomalyReported = false, - } = args; - - if (substrate.status === 'running') { - // Healthy — and it also ENDS any anomaly episode, so the next one reports. - return { taskFailed: false, suspendAnomalyReported: false }; - } - - if (substrate.status === 'suspended') { - if (ddbStatus === TaskStatus.AWAITING_APPROVAL) { - // Orchestrator-intended suspend during an approval wait — the whole point - // of this backend. Nothing to report, and the anomaly is re-armed: if the - // task later leaves AWAITING_APPROVAL while still suspended, that is a new - // and genuinely reportable episode. - return { taskFailed: false, suspendAnomalyReported: false }; - } - // Suspended outside an approval wait. Nothing in ABCA suspends a MicroVM - // except the orchestrator's (P3) approval-wait policy, so this means either - // an out-of-band SuspendMicrovm call or a substrate-side suspend we did not - // ask for. Surface it — do NOT fail-fast (ADR-021: "an anomaly to surface, - // not fail-fast"); the VM's state is intact and resumable. - log.warn('MicroVM is suspended while the task is not awaiting approval', { - microvm_id: microvmId, - task_status: ddbStatus, - anomaly_already_reported: suspendAnomalyReported, - }); - if (!suspendAnomalyReported) { - await emitTaskEvent(taskId, 'microvm_suspend_anomaly', { - microvm_id: microvmId, - task_status: ddbStatus, - reason: 'suspended_outside_approval_wait', - }, correlation); - } - return { taskFailed: false, suspendAnomalyReported: true }; - } - - // Terminal substrate report (`completed` or `failed`). `pollSession` reports - // TERMINATING/TERMINATED/NotFound as `completed` because it cannot see an exit - // code; `failed` only reaches here if a future mapping adds one. - // - // `substrate.reason` is `GetMicrovm`'s `stateReason`, carried through verbatim. - // Appending it is what makes this string true on the dominant failure: without - // it a `/run` hook 4xx (which self-terminates the VM in ~12 s) rendered as the - // bare "substrate state completed", and the classifier's remedy then named a - // session duration cap, a host fault, or an external terminate — none of which - // happened. With it the operator gets "substrate state completed (Run lifecycle - // hook returned HTTP status 400…)", which points at the guest logs where the - // agent's own structured 4xx body already is. - const substrateReason = substrate.reason ? ` (${substrate.reason})` : ''; - const detail = substrate.status === 'failed' - ? `${substrate.error}${substrateReason}` - : `substrate state ${substrate.status}${substrateReason}`; - - const reread = await loadTask(taskId); - if (TERMINAL_STATUSES.includes(reread.status)) { - // The agent wrote its terminal status between this cycle's status read and - // now — the normal shutdown ordering. Not a failure. - log.info('MicroVM terminated after the agent wrote a terminal status', { - microvm_id: microvmId, - task_status: reread.status, - }); - // Terminal either way, so the flag no longer matters; carried through - // unchanged rather than reset so the value never lies about what happened. - return { taskFailed: false, suspendAnomalyReported }; - } - - log.error('MicroVM reached a terminal state before the agent wrote a terminal status', { - microvm_id: microvmId, - task_status: reread.status, - detail, - }); - // `releaseConcurrency: false` — the finalize step sees the now-terminal task - // and decrements, matching the ECS substrate-failure branch in orchestrate-task. - await failTask( - taskId, - reread.status, - `MicroVM substrate terminated before the agent wrote a terminal status: ${detail}`, - userId, - false, - repo, - ); - return { taskFailed: true, suspendAnomalyReported }; -} - /** * Load blueprint configuration for a task's repository and merge with platform defaults. * @param task - the task record (needs task.repo). @@ -1103,36 +914,24 @@ export async function pollTaskStatus( && item?.session_id && typeof item.started_at === 'string' ) { - const startedMs = Date.parse(item.started_at); const now = Date.now(); - if (!Number.isNaN(startedMs)) { - const runningAgeSec = (now - startedMs) / 1000; - - if (typeof item.agent_heartbeat_at === 'string') { - // Agent has sent at least one heartbeat — check staleness - const hbMs = Date.parse(item.agent_heartbeat_at); - if (!Number.isNaN(hbMs)) { - const hbAgeSec = (now - hbMs) / 1000; - if (runningAgeSec > AGENT_HEARTBEAT_GRACE_SEC && hbAgeSec > AGENT_HEARTBEAT_STALE_SEC) { - sessionUnhealthy = true; - logger.warn('Agent heartbeat stale while task RUNNING', { - task_id: taskId, - compute_type: computeType, - agent_heartbeat_at: item.agent_heartbeat_at, - heartbeat_age_sec: Math.round(hbAgeSec), - }); - } - } - } else if (runningAgeSec > AGENT_HEARTBEAT_GRACE_SEC + AGENT_HEARTBEAT_STALE_SEC) { - // Agent never sent a heartbeat and task has been RUNNING well past - // the grace period — likely early crash before pipeline started. - sessionUnhealthy = true; - logger.warn('Agent never sent heartbeat while task RUNNING past grace period', { - task_id: taskId, - compute_type: computeType, - running_age_sec: Math.round(runningAgeSec), - }); - } + const startedMs = Date.parse(item.started_at); + const heartbeatMs = typeof item.agent_heartbeat_at === 'string' ? Date.parse(item.agent_heartbeat_at) : undefined; + const health = evaluateAgentHeartbeat(startedMs, heartbeatMs, now); + sessionUnhealthy = health !== undefined; + if (health === 'stale') { + logger.warn('Agent heartbeat stale while task RUNNING', { + task_id: taskId, + compute_type: computeType, + agent_heartbeat_at: item.agent_heartbeat_at, + heartbeat_age_sec: Math.round((now - heartbeatMs!) / 1000), + }); + } else if (health === 'missing') { + logger.warn('Agent never sent heartbeat while task RUNNING past grace period', { + task_id: taskId, + compute_type: computeType, + running_age_sec: Math.round((now - startedMs) / 1000), + }); } } @@ -1153,26 +952,63 @@ export async function finalizeTask( taskId: string, pollState: PollState, userId: string, -): Promise { - const task = await loadTask(taskId); +): Promise { + const current = await loadTask(taskId, true); + const expectedId = pollState.microvmSupervisor?.microvmId; + if (expectedId && (current.user_id !== userId || current.compute_type !== 'lambda-microvm' + || current.session_id !== expectedId || current.compute_metadata?.microvmId !== expectedId)) { + logger.warn('MicroVM finalization skipped after worker ownership changed', { task_id: taskId, microvm_id: expectedId }); + return false; + } + try { + await finalizeTaskOutcome(taskId, pollState, current); + return true; + } finally { + // The marker makes this safe after a crash, an event failure, or a competing + // cleaner. A still-active task keeps its reservation. + await releaseTaskSlot(taskId, userId); + await closeTaskApprovals(taskId, userId); + } +} + +async function finalizeTaskOutcome(taskId: string, pollState: PollState, task: TaskRecord): Promise { + // Finalization can immediately follow a committed start failure/cancellation. + // A stale active state would emit the wrong terminal event. const currentStatus = task.status; // Correlation envelope on this function's own log lines too, not just the // events it emits — admission→terminal logs must join by {user_id, repo}. const { log, correlation } = envelopeFor(task); - // Lost session: RUNNING but agent heartbeats stopped (crash/OOM) — fail fast. - // - // FINALIZING is in the guard DEFENSIVELY, and is currently unreachable: the - // only writer of `sessionUnhealthy` is `pollTaskStatus`, which computes it - // under `currentStatus === TaskStatus.RUNNING`, so a FINALIZING task can never - // arrive here with the flag set. It is kept rather than removed because the - // reachability depends on a predicate in ANOTHER function: the day - // `pollTaskStatus` widens its own status gate (P3's suspend policy already has - // to revisit that block), a heartbeat-stale FINALIZING task must fail rather - // than fall through to the normal terminal path and be reported as a success. - // Dropping the arm would make that a silent behaviour change instead of a - // no-op. Do NOT "simplify" it away without also pinning `pollTaskStatus`'s - // RUNNING-only gate with a test. + if (pollState.microvmFailureReason && !TERMINAL_STATUSES.includes(currentStatus)) { + // Approval/HYDRATING permit FAILED, not TIMED_OUT. This is infrastructure + // failure and must not invent a TIMED_OUT decision on the approval row. + const timeout = pollState.microvmFailureReason === 'session-deadline' + && (currentStatus === TaskStatus.RUNNING || currentStatus === TaskStatus.FINALIZING); + const terminal = timeout ? TaskStatus.TIMED_OUT : TaskStatus.FAILED; + try { + await transitionTask(taskId, currentStatus, terminal, { + completed_at: new Date().toISOString(), + error_message: pollState.microvmFailureMessage ?? `MicroVM supervisor: ${pollState.microvmFailureReason}`, + }, pollState.microvmSupervisor?.microvmId); + } catch (error) { + const winner = await loadTask(taskId, true); + if (!TERMINAL_STATUSES.includes(winner.status)) throw error; + await emitTaskEvent(taskId, `task_${winner.status.toLowerCase()}`, { + final_status: winner.status, poll_attempts: pollState.attempts, + }, correlation); + return; + } + await emitTaskEvent(taskId, timeout ? 'task_timed_out' : 'task_failed', { + reason: 'microvm_supervisor', + detail: pollState.microvmFailureReason, + microvm_id: pollState.microvmSupervisor?.microvmId, + poll_attempts: pollState.attempts, + }, correlation); + return; + } + + // A heartbeat failure is detected while RUNNING. The strong read above may + // already observe FINALIZING; neither active status is a successful outcome. if ( pollState.sessionUnhealthy && (currentStatus === TaskStatus.RUNNING || currentStatus === TaskStatus.FINALIZING) @@ -1184,11 +1020,11 @@ export async function finalizeTask( error_message: 'Agent session lost: no recent heartbeat from the agent ' + `(${substrateNoun(task.compute_type)} may have crashed, been OOM-killed, or stopped)`, - }); + }, pollState.microvmSupervisor?.microvmId); transitioned = true; } catch (err) { // Task may have transitioned concurrently (e.g. agent wrote terminal status). - // Re-read to avoid double-decrement or contradictory events. + // Re-read to report the terminal status that actually won. log.warn('Finalization transition to FAILED (heartbeat) failed, task may have transitioned concurrently', { error: err instanceof Error ? err.message : String(err), }); @@ -1198,21 +1034,19 @@ export async function finalizeTask( reason: 'agent_heartbeat_stale', poll_attempts: pollState.attempts, }, correlation); - await decrementConcurrency(userId); } else { // Transition failed — re-read task to determine actual state. // If already terminal the block below will handle TTL + concurrency. - const reread = await loadTask(taskId); + const reread = await loadTask(taskId, true); if (TERMINAL_STATUSES.includes(reread.status)) { log.info('Heartbeat path: task already terminal after failed transition', { status: reread.status }); await emitTaskEvent(taskId, `task_${reread.status.toLowerCase()}`, { final_status: reread.status, poll_attempts: pollState.attempts, }, correlation); - await decrementConcurrency(userId); } else { - log.warn('Heartbeat path: task in unexpected state after failed transition, releasing concurrency', { status: reread.status }); - await decrementConcurrency(userId); + log.warn('Heartbeat path: task in unexpected state after failed transition, keeping its reservation until terminal', { status: reread.status }); + throw new Error(`Heartbeat finalization left task ${taskId} active in ${reread.status}`); } } return; @@ -1279,39 +1113,43 @@ export async function finalizeTask( final_status: currentStatus, poll_attempts: pollState.attempts, }, correlation); - await decrementConcurrency(userId); return; } // If still RUNNING / FINALIZING / AWAITING_APPROVAL after the poll - // window closes, transition to TIMED_OUT. AWAITING_APPROVAL uses the - // same transition — the stranded-approval reconciler is a secondary - // safety net with a longer timeout for tasks the orchestrator already - // lost track of. + // window closes, terminate the task through an allowed transition. Approval + // waits permit FAILED, not TIMED_OUT. Task closure cancels pending approvals; + // reaching the execution limit is not a human denial. if ( currentStatus === TaskStatus.RUNNING || currentStatus === TaskStatus.FINALIZING || currentStatus === TaskStatus.AWAITING_APPROVAL ) { - const terminalStatus = TaskStatus.TIMED_OUT; + const terminalStatus = currentStatus === TaskStatus.AWAITING_APPROVAL ? TaskStatus.FAILED : TaskStatus.TIMED_OUT; try { await transitionTask(taskId, currentStatus, terminalStatus, { completed_at: new Date().toISOString(), error_message: currentStatus === TaskStatus.AWAITING_APPROVAL - ? 'Orchestrator poll timeout exceeded while awaiting approval' + ? 'Task execution limit reached while waiting for approval. ' + + 'The task has closed; submit a new task to continue.' : 'Orchestrator poll timeout exceeded', }); } catch (err) { - // Task may have transitioned concurrently — re-read and accept + // Accept a committed terminal winner; otherwise retry finalization. log.warn('Finalization transition failed, task may have transitioned concurrently', { error: err instanceof Error ? err.message : String(err) }); + const current = await loadTask(taskId, true); + if (!TERMINAL_STATUSES.includes(current.status)) throw err; + await emitTaskEvent(taskId, `task_${current.status.toLowerCase()}`, { + final_status: current.status, poll_attempts: pollState.attempts, + }, correlation); + return; } - await emitTaskEvent(taskId, 'task_timed_out', { + await emitTaskEvent(taskId, terminalStatus === TaskStatus.FAILED ? 'task_failed' : 'task_timed_out', { reason: currentStatus === TaskStatus.AWAITING_APPROVAL ? 'approval_poll_timeout' : 'poll_timeout', poll_attempts: pollState.attempts, }, correlation); - await decrementConcurrency(userId); return; } @@ -1325,18 +1163,23 @@ export async function finalizeTask( }); } catch (err) { log.warn('Finalization transition from HYDRATING failed, task may have transitioned concurrently', { error: err instanceof Error ? err.message : String(err) }); + const current = await loadTask(taskId, true); + if (!TERMINAL_STATUSES.includes(current.status)) throw err; + await emitTaskEvent(taskId, `task_${current.status.toLowerCase()}`, { + final_status: current.status, poll_attempts: pollState.attempts, + }, correlation); + return; } await emitTaskEvent(taskId, 'task_failed', { reason: 'session_never_started', poll_attempts: pollState.attempts, }, correlation); - await decrementConcurrency(userId); return; } - // Unexpected state — log and release concurrency + // Unexpected active state — retain its reservation until terminal. log.error('Unexpected task state during finalization', { status: currentStatus }); - await decrementConcurrency(userId); + throw new Error(`Cannot finalize task ${taskId} in ${currentStatus}`); } /** @@ -1392,7 +1235,7 @@ export async function queueTask(task: TaskRecord): Promise { * @param fromStatus - the current status. * @param errorMessage - the error reason. * @param userId - the user who owns the task. - * @param releaseConcurrency - whether to decrement the concurrency counter. + * @param releaseConcurrency - whether to release this task's held reservation. * @param repo - optional target repo (`owner/repo`) for the correlation * envelope; omit for repo-less workflows. */ @@ -1416,41 +1259,17 @@ export async function failTask( log.warn('Failed to transition task to FAILED', { error: err instanceof Error ? err.message : String(err), }); + const current = await loadTask(taskId, true); + if (!TERMINAL_STATUSES.includes(current.status)) throw err; } - // Only emit / release concurrency after a successful transition. Callers such as - // orchestrate-task rethrow after failTask; Durable Execution retries the step and - // would otherwise re-run emit + decrement while the task is already FAILED. - if (transitioned) { - await emitTaskEvent(taskId, 'task_failed', { error_message: errorMessage }, { user_id: userId, repo }); - if (releaseConcurrency) { - await decrementConcurrency(userId); - } - } -} - -/** - * Decrement the user's concurrency counter (best-effort). - * @param userId - the user ID. - */ -async function decrementConcurrency(userId: string): Promise { + // A replay may find the FAILED write already committed. Events remain + // transition-owned; reservation release is independently safe to repeat. try { - await ddb.send(new UpdateCommand({ - TableName: CONCURRENCY_TABLE_NAME, - Key: { user_id: userId }, - UpdateExpression: 'SET active_count = active_count - :one, updated_at = :now', - ConditionExpression: 'active_count > :zero', - ExpressionAttributeValues: { - ':one': 1, - ':zero': 0, - ':now': new Date().toISOString(), - }, - })); - } catch (err: unknown) { - if (err && typeof err === 'object' && 'name' in err && err.name === 'ConditionalCheckFailedException') { - logger.info('Concurrency counter already at zero, nothing to decrement', { user_id: userId }); - } else { - logger.warn('Failed to decrement concurrency counter', { user_id: userId, error: err instanceof Error ? err.message : String(err) }); + if (transitioned) { + await emitTaskEvent(taskId, 'task_failed', { error_message: errorMessage }, { user_id: userId, repo }); } + } finally { + if (releaseConcurrency) await releaseTaskSlot(taskId, userId); } } diff --git a/cdk/src/handlers/shared/payload-bootstrap.ts b/cdk/src/handlers/shared/payload-bootstrap.ts new file mode 100644 index 000000000..c0dc95c94 --- /dev/null +++ b/cdk/src/handlers/shared/payload-bootstrap.ts @@ -0,0 +1,191 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { createHash } from 'node:crypto'; +import { DeleteObjectCommand, GetObjectCommand, PutObjectCommand, S3Client } from '@aws-sdk/client-s3'; +import { getSignedUrl } from '@aws-sdk/s3-request-presigner'; +import { canonicalJson } from './canonical-json'; +import { logger } from './logger'; +import { makeClient } from './ua'; +import constants from '../../../../contracts/constants.json'; + +export const PAYLOAD_BOOTSTRAP = constants.payload_bootstrap; +type Backend = 'ecs' | 'lambda-microvm'; + +export interface PayloadReference { + version: number; + task_id: string; + attempt_id?: string; + bootstrap_s3_uri: string; + payload_url: string; + expires_at: number; +} + +interface LaunchRecord { + fingerprint: string; + reference: PayloadReference; +} + +let client: S3Client | undefined; +function s3(): S3Client { + return client ??= makeClient(S3Client); +} + +function sha256(value: string): string { + return createHash('sha256').update(value).digest('hex'); +} + +/** Never let SDK errors echo bearer URLs into task records or logs. */ +export function redactPayloadUrls(message: string): string { + return message.replace(/https?:\/\/[^\s"'<>\\]+/gi, url => + /X-Amz-/i.test(url) ? '[redacted payload URL]' : url); +} + +async function readObject(bucket: string, key: string): Promise { + try { + const response = await s3().send(new GetObjectCommand({ Bucket: bucket, Key: key })); + if (!response.Body) throw new Error('PAYLOAD_BOOTSTRAP_UNREADABLE: empty object response'); + return await response.Body.transformToString(); + } catch (error) { + if ((error as { name?: string }).name === 'NoSuchKey') return undefined; + throw error; + } +} + +/** Conditional creation, including recovery after a committed write loses its reply. */ +async function createOnce(bucket: string, key: string, body: string): Promise { + try { + await s3().send(new PutObjectCommand({ + Bucket: bucket, Key: key, Body: body, ContentType: 'application/json', IfNoneMatch: '*', + })); + return body; + } catch (error) { + const saved = await readObject(bucket, key); + if (saved !== undefined) return saved; + throw error; + } +} + +/** + * Save a single-object capability outside the agent-readable task table. + * Replays read the same launch.json; re-signing would change Run's request + * while reusing its clientToken. The worker can read only bootstrap/* using + * its own credentials, never payload.json or launch.json. + */ +export async function preparePayloadReference(input: { + bucket: string; + taskId: string; + attemptId?: string; + backend: Backend; + payload: Record; + platformConfig?: Record; +}): Promise { + const { bucket, taskId, backend, payload } = input; + if (!/^[A-Za-z0-9_-]{1,128}$/.test(taskId) || payload.task_id !== taskId) { + throw new Error('PAYLOAD_BOOTSTRAP_INVALID: task identity does not match the payload'); + } + if (input.attemptId !== undefined && ( + backend !== 'lambda-microvm' || !/^[A-Za-z0-9_-]{1,128}$/.test(input.attemptId) + || payload.attempt_id !== input.attemptId + )) throw new Error('PAYLOAD_BOOTSTRAP_INVALID: worker attempt does not match the payload'); + const objectPrefix = input.attemptId ? `${taskId}/${input.attemptId}` : taskId; + const manifest = canonicalJson({ + version: PAYLOAD_BOOTSTRAP.version, backend, platform_config: input.platformConfig ?? {}, + }); + const manifestKey = `${PAYLOAD_BOOTSTRAP.manifest_prefix}${sha256(manifest)}.json`; + const payloadBody = canonicalJson({ + version: PAYLOAD_BOOTSTRAP.version, + task_id: taskId, + agent_payload: payload, + platform_config: input.platformConfig ?? {}, + }); + if (Buffer.byteLength(manifest) > PAYLOAD_BOOTSTRAP.max_manifest_bytes + || Buffer.byteLength(payloadBody) > PAYLOAD_BOOTSTRAP.max_payload_bytes) { + throw new Error('PAYLOAD_BOOTSTRAP_TOO_LARGE: bootstrap manifest or task payload exceeds its byte limit'); + } + const fingerprint = sha256(canonicalJson({ bucket, backend, manifest, payloadBody })); + const launchKey = `${objectPrefix}/${PAYLOAD_BOOTSTRAP.launch_filename}`; + const accept = (saved: string): PayloadReference => { + const record = JSON.parse(saved) as LaunchRecord; + if (record.fingerprint !== fingerprint || record.reference?.task_id !== taskId + || record.reference.attempt_id !== input.attemptId) { + throw new Error('PAYLOAD_BOOTSTRAP_CONFLICT: task already has different launch instructions'); + } + if (record.reference.expires_at <= Date.now()) { + throw new Error('PAYLOAD_BOOTSTRAP_EXPIRED: inspect the existing launch; do not start a replacement task'); + } + return record.reference; + }; + + // Refresh only identical, public deployment settings so bucket lifecycle + // expiry cannot reap an old manifest just as a new task starts using it. + await s3().send(new PutObjectCommand({ + Bucket: bucket, Key: manifestKey, Body: manifest, ContentType: 'application/json', + })); + const existing = await readObject(bucket, launchKey); + if (existing !== undefined) return accept(existing); + + const payloadKey = `${objectPrefix}/payload.json`; + const savedPayload = await createOnce(bucket, payloadKey, payloadBody); + if (savedPayload !== payloadBody) { + throw new Error('PAYLOAD_BOOTSTRAP_CONFLICT: task already has different stored instructions'); + } + const now = Date.now(); + const credentials = await s3().config.credentials(); + const lifetime = Math.min( + PAYLOAD_BOOTSTRAP.url_ttl_seconds, + credentials.expiration + ? Math.floor((credentials.expiration.getTime() - now) / 1000) + : PAYLOAD_BOOTSTRAP.url_ttl_seconds, + ); + if (lifetime < PAYLOAD_BOOTSTRAP.minimum_url_lifetime_seconds) { + throw new Error('PAYLOAD_BOOTSTRAP_CREDENTIALS_EXPIRING: refresh coordinator credentials before launch'); + } + const url = await getSignedUrl(s3(), new GetObjectCommand({ + Bucket: bucket, Key: payloadKey, + }), { expiresIn: lifetime }); + const reference: PayloadReference = { + version: PAYLOAD_BOOTSTRAP.version, + task_id: taskId, + ...(input.attemptId && { attempt_id: input.attemptId }), + bootstrap_s3_uri: `s3://${bucket}/${manifestKey}`, + payload_url: url, + expires_at: now + lifetime * 1000, + }; + const record = canonicalJson({ fingerprint, reference }); + return accept(await createOnce(bucket, launchKey, record)); +} + +/** Delete both task instructions and their saved capability; shared manifests expire by lifecycle. */ +export async function deletePayloadReference(bucket: string, taskId: string, attemptId?: string): Promise { + if (!/^[A-Za-z0-9_-]{1,128}$/.test(taskId) + || (attemptId !== undefined && !/^[A-Za-z0-9_-]{1,128}$/.test(attemptId))) { + throw new Error('PAYLOAD_BOOTSTRAP_INVALID: cleanup identity is invalid'); + } + const objectPrefix = attemptId ? `${taskId}/${attemptId}` : taskId; + for (const filename of ['payload.json', PAYLOAD_BOOTSTRAP.launch_filename]) { + try { + await s3().send(new DeleteObjectCommand({ Bucket: bucket, Key: `${objectPrefix}/${filename}` })); + } catch (error) { + logger.warn('Payload bootstrap cleanup failed (non-fatal)', { + task_id: taskId, filename, error: (error as { name?: string }).name ?? 'UnknownError', + }); + } + } +} diff --git a/cdk/src/handlers/shared/response.ts b/cdk/src/handlers/shared/response.ts index bab4d5b4c..9a9235c40 100644 --- a/cdk/src/handlers/shared/response.ts +++ b/cdk/src/handlers/shared/response.ts @@ -30,6 +30,7 @@ export const ErrorCode = { TRACE_NOT_AVAILABLE: 'TRACE_NOT_AVAILABLE', DUPLICATE_TASK: 'DUPLICATE_TASK', TASK_ALREADY_TERMINAL: 'TASK_ALREADY_TERMINAL', + TASK_STATE_CONFLICT: 'TASK_STATE_CONFLICT', RATE_LIMIT_EXCEEDED: 'RATE_LIMIT_EXCEEDED', WEBHOOK_NOT_FOUND: 'WEBHOOK_NOT_FOUND', WEBHOOK_ALREADY_REVOKED: 'WEBHOOK_ALREADY_REVOKED', @@ -131,6 +132,7 @@ export function errorResponse( code: string, message: string, requestId: string, + details?: Record, ): APIGatewayProxyResult { return { statusCode, @@ -140,6 +142,7 @@ export function errorResponse( code, message, request_id: requestId, + ...(details ? { details } : {}), }, }), }; diff --git a/cdk/src/handlers/shared/session-start-retry.ts b/cdk/src/handlers/shared/session-start-retry.ts index 2f7156e39..4b72fe802 100644 --- a/cdk/src/handlers/shared/session-start-retry.ts +++ b/cdk/src/handlers/shared/session-start-retry.ts @@ -20,22 +20,23 @@ /** * Session-start transient auto-retry (once), extracted from the durable * ``orchestrate-task`` handler so the four retry branches are unit-testable in - * isolation (the handler's inline ``start-session`` step is never invoked by - * the test suite). See #599 review B1/B2. + * isolation, with MicroVM receipt/replay behavior also covered through the + * strategy and durable-handler tests. See #599 review B1/B2. * - * session-start is the ONE place a retry is idempotent by construction — no repo - * clone, no commits, no PR have happened yet, so re-invoking - * RunTask/InvokeAgentRuntime can't double-run work. A transient hiccup here (an + * This retries the start API, before the caller has a session handle. That does + * not guarantee idempotency: a lost success response can leave a live session + * behind. MicroVM uses its saved start receipt and token; each other backend + * still needs its own idempotency policy. A transient hiccup here (an * ECS deploy-race "TaskDefinition is inactive", ENI/capacity delay, a * Bedrock/agentcore throttle) usually clears on a second attempt, so the first * transient failure is swallowed and retried once. A NON-transient failure (bad * config, missing ECS substrate) is re-thrown immediately — retrying it just - * wastes ~a minute. Mid-run crashes are NOT handled here (that's a later step; + * wastes another request. Mid-run crashes are NOT handled here (that's a later step; * the agent may have pushed commits). */ import type { ComputeStrategy, SessionHandle } from './compute-strategy'; -import { classifyError, isTransientError } from './error-classifier'; +import { classifyError, isTransientError, MicrovmStartUncertainError } from './error-classifier'; /** Emit a ``session_start_retry`` telemetry event. Matches the shape of * ``emitTaskEvent`` bound at the call site (best-effort — see below). */ @@ -114,17 +115,20 @@ export async function startSessionWithRetry( // classifier's transient patterns — e.g. ThrottlingException — is a separate // classifier-completeness concern, not this retry gate.) const classification = classifyError(String(firstErr)); - if (!isTransientError(classification)) { + if (!isTransientError(classification) && !(firstErr instanceof MicrovmStartUncertainError)) { throw firstErr; // service/user error — a retry won't help; surface now. } - deps.logger.warn('Session start hit a transient error — auto-retrying once', { + const uncertain = firstErr instanceof MicrovmStartUncertainError; + deps.logger.warn(uncertain + ? 'MicroVM start response is unknown — recovering the same request once' + : 'Session start hit a transient error — auto-retrying once', { task_id: deps.taskId, error: firstErr instanceof Error ? firstErr.message : String(firstErr), }); // Best-effort telemetry — a PutItem fault here must never abort the retry // or mis-report the outcome (B1). try { - await deps.emitRetryEvent(classification?.title ?? 'transient'); + await deps.emitRetryEvent(uncertain ? 'MicroVM start response unknown' : classification?.title ?? 'transient'); } catch (emitErr) { deps.logger.warn('session_start_retry event emit failed (non-fatal)', { task_id: deps.taskId, @@ -135,14 +139,19 @@ export async function startSessionWithRetry( const handle = await strategy.startSession(input); return { handle, autoRetried: true }; } catch (retryErr) { + // Even a definite rejection on the second call cannot prove that the + // first, unanswered call created nothing. Preserve that uncertainty. + const finalError = firstErr instanceof MicrovmStartUncertainError || retryErr instanceof MicrovmStartUncertainError + ? new MicrovmStartUncertainError(String(retryErr), { cause: retryErr }) + : retryErr; // Branch 4: the retry ALSO failed. Pin the retry fact onto the thrown error // so the caller can stamp ``[auto-retried]`` (N1) — otherwise a double- // transient failure is indistinguishable from a first-attempt failure and // the user is wrongly told "reply to retry" instead of "I already retried". - if (typeof retryErr === 'object' && retryErr !== null) { - (retryErr as Record)[AUTO_RETRIED] = true; + if (typeof finalError === 'object' && finalError !== null) { + (finalError as Record)[AUTO_RETRIED] = true; } - throw retryErr; + throw finalError; } } } diff --git a/cdk/src/handlers/shared/slack-blocks.ts b/cdk/src/handlers/shared/slack-blocks.ts index 86f116cc1..eaf5f47a4 100644 --- a/cdk/src/handlers/shared/slack-blocks.ts +++ b/cdk/src/handlers/shared/slack-blocks.ts @@ -39,7 +39,7 @@ interface PlainText { /** Section block: a single line/paragraph of mrkdwn content. */ export interface SectionBlock { readonly type: 'section'; - readonly text: MrkdwnText; + readonly text: MrkdwnText | PlainText; readonly block_id?: string; } diff --git a/cdk/src/handlers/shared/strategies/agentcore-strategy.ts b/cdk/src/handlers/shared/strategies/agentcore-strategy.ts index 762387abe..9835b9ec5 100644 --- a/cdk/src/handlers/shared/strategies/agentcore-strategy.ts +++ b/cdk/src/handlers/shared/strategies/agentcore-strategy.ts @@ -19,7 +19,7 @@ import { randomUUID } from 'crypto'; import { BedrockAgentCoreClient, InvokeAgentRuntimeCommand, StopRuntimeSessionCommand } from '@aws-sdk/client-bedrock-agentcore'; -import type { ComputeStrategy, SessionHandle, SessionStatus } from '../compute-strategy'; +import type { ComputeStrategy, SessionControlOptions, SessionHandle, SessionLifecycleResult, SessionStatus } from '../compute-strategy'; import { logger } from '../logger'; import type { BlueprintConfig } from '../repo-config'; import { makeClient } from '../ua'; @@ -49,7 +49,7 @@ export class AgentCoreComputeStrategy implements ComputeStrategy { // injection: when set, AgentCore exchanges the caller's identity for // a workload token and delivers it to the agent container via the // `WorkloadAccessToken` request header (read by - // `BedrockAgentCoreContext.set_workload_access_token` in app.py). + // `BedrockAgentCoreContext.set_workload_access_token` in server.py). // Without it, the agent's `resolve_linear_api_token()` short-circuits // before reaching the Identity SDK call. Requires the orchestrator // role to have `bedrock-agentcore:InvokeAgentRuntimeForUser` in @@ -79,21 +79,33 @@ export class AgentCoreComputeStrategy implements ComputeStrategy { }; } - async pollSession(_handle: SessionHandle): Promise { + async pollSession(_handle: SessionHandle, options?: SessionControlOptions): Promise { + options?.abortSignal?.throwIfAborted(); return { status: 'running' }; } - async stopSession(handle: SessionHandle): Promise { + async suspendSession(handle: SessionHandle): Promise { + if (handle.strategyType !== 'agentcore') throw new Error('suspendSession called with non-agentcore handle'); + return { supported: false }; + } + + async resumeSession(handle: SessionHandle): Promise { + if (handle.strategyType !== 'agentcore') throw new Error('resumeSession called with non-agentcore handle'); + return { supported: false }; + } + + async stopSession(handle: SessionHandle, options?: SessionControlOptions): Promise { if (handle.strategyType !== 'agentcore') { throw new Error('stopSession called with non-agentcore handle'); } const { runtimeArn } = handle; try { + options?.abortSignal?.throwIfAborted(); await getClient().send(new StopRuntimeSessionCommand({ agentRuntimeArn: runtimeArn, runtimeSessionId: handle.sessionId, - })); + }), options); logger.info('AgentCore session stopped', { session_id: handle.sessionId }); } catch (err) { const errName = err instanceof Error ? err.name : undefined; diff --git a/cdk/src/handlers/shared/strategies/ecs-strategy.ts b/cdk/src/handlers/shared/strategies/ecs-strategy.ts index 75853ee39..acbd91b39 100644 --- a/cdk/src/handlers/shared/strategies/ecs-strategy.ts +++ b/cdk/src/handlers/shared/strategies/ecs-strategy.ts @@ -18,9 +18,9 @@ */ import { ECSClient, RunTaskCommand, DescribeTasksCommand, StopTaskCommand } from '@aws-sdk/client-ecs'; -import { S3Client, PutObjectCommand, DeleteObjectCommand } from '@aws-sdk/client-s3'; -import type { ComputeStrategy, SessionHandle, SessionStatus } from '../compute-strategy'; +import type { ComputeStrategy, SessionControlOptions, SessionHandle, SessionLifecycleResult, SessionStatus } from '../compute-strategy'; import { logger } from '../logger'; +import { deletePayloadReference, preparePayloadReference, redactPayloadUrls } from '../payload-bootstrap'; import type { BlueprintConfig } from '../repo-config'; import { makeClient } from '../ua'; import { DEFAULT_MAX_TURNS } from '../validation'; @@ -33,14 +33,6 @@ function getClient(): ECSClient { return sharedClient; } -let sharedS3Client: S3Client | undefined; -function getS3Client(): S3Client { - if (!sharedS3Client) { - sharedS3Client = makeClient(S3Client); - } - return sharedS3Client; -} - const ECS_CLUSTER_ARN = process.env.ECS_CLUSTER_ARN; const ECS_TASK_DEFINITION_ARN = process.env.ECS_TASK_DEFINITION_ARN; /** @@ -79,50 +71,38 @@ export function toTaskDefinitionFamily(ref: string): string { const ECS_SECURITY_GROUP = process.env.ECS_SECURITY_GROUP; const ECS_CONTAINER_NAME = process.env.ECS_CONTAINER_NAME ?? 'AgentContainer'; const ECS_PAYLOAD_BUCKET = process.env.ECS_PAYLOAD_BUCKET; +const ECS_OVERRIDES_LIMIT_BYTES = 8192; -/** - * Inline-payload size (bytes) above which we warn that RunTask will likely - * reject the call when no payload bucket is configured. ECS caps the TOTAL - * containerOverrides blob at 8192 bytes; the other env vars + command consume - * some of that, so 6 KB of payload is the practical danger line. - */ -const INLINE_PAYLOAD_WARN_BYTES = 6144; - -/** - * S3 object key for a task's ECS payload. One object per task under its own - * task-id prefix; deleted by the orchestrator at finalize (see - * ``deleteEcsPayload``), with the bucket's 1-day lifecycle rule as a backstop. - */ -export function ecsPayloadKey(taskId: string): string { - return `${taskId}/payload.json`; +function safeLaunchError(error: unknown): Error { + const safe = new Error(redactPayloadUrls(error instanceof Error ? error.message : String(error))); + safe.name = error instanceof Error ? error.name : 'Error'; + return safe; } /** - * Delete a task's ECS payload object. Best-effort: a failed delete must never + * Delete a task's ECS payload and private launch reference. A failed delete must never * fail the task — the bucket's 1-day lifecycle rule reaps it regardless. Called * from the orchestrator's ``finalize`` step once the task is terminal. No-ops * when the payload bucket isn't configured (AgentCore-only deployments). */ export async function deleteEcsPayload(taskId: string): Promise { if (!ECS_PAYLOAD_BUCKET) return; - try { - await getS3Client().send(new DeleteObjectCommand({ - Bucket: ECS_PAYLOAD_BUCKET, - Key: ecsPayloadKey(taskId), - })); - logger.info('Deleted ECS payload object', { task_id: taskId }); - } catch (err) { - // Non-fatal — the lifecycle rule is the backstop. - logger.warn('Failed to delete ECS payload object (non-fatal)', { - task_id: taskId, - error: err instanceof Error ? err.message : String(err), - }); - } + await deletePayloadReference(ECS_PAYLOAD_BUCKET, taskId); } export class EcsComputeStrategy implements ComputeStrategy { readonly type = 'ecs'; + async suspendSession(handle: SessionHandle): Promise { + if (handle.strategyType !== 'ecs') throw new Error('suspendSession called with non-ecs handle'); + return { supported: false }; + } + + async resumeSession(handle: SessionHandle): Promise { + if (handle.strategyType !== 'ecs') throw new Error('resumeSession called with non-ecs handle'); + return { supported: false }; + } + async startSession(input: { taskId: string; /** Accepted to satisfy the ComputeStrategy interface; ECS doesn't @@ -167,46 +147,21 @@ export class EcsComputeStrategy implements ComputeStrategy { // The ECS container's default CMD starts the FastAPI server (uvicorn) which // waits for HTTP POST to /invocations — but in standalone ECS nobody sends - // that request. We override the container command to invoke run_task() + // that request. We override the container command to invoke run_task_from_payload() // directly with the full orchestrator payload (including hydrated_context). // This avoids the server entirely and runs the agent in batch mode. - const payloadJson = JSON.stringify(payload); - - // The payload (especially hydrated_context) routinely exceeds the 8192-byte - // cap that ECS RunTask enforces on the TOTAL containerOverrides blob, which - // rejected the call with InvalidParameterException. Write the payload to S3 - // and pass only a small pointer (AGENT_PAYLOAD_S3_URI); the container fetches - // it on boot. The inline AGENT_PAYLOAD remains as a fallback for small - // payloads / deployments without a payload bucket configured. - let payloadS3Uri: string | undefined; - if (ECS_PAYLOAD_BUCKET) { - const key = ecsPayloadKey(taskId); - await getS3Client().send(new PutObjectCommand({ - Bucket: ECS_PAYLOAD_BUCKET, - Key: key, - Body: payloadJson, - ContentType: 'application/json', - })); - payloadS3Uri = `s3://${ECS_PAYLOAD_BUCKET}/${key}`; - logger.info('Wrote ECS payload to S3', { - task_id: taskId, - bytes: payloadJson.length, - uri: payloadS3Uri, - }); - } else if (payloadJson.length > INLINE_PAYLOAD_WARN_BYTES) { - // No bucket configured AND the payload is large enough that the inline - // path will almost certainly blow the 8192-byte overrides cap. Surface a - // clear cause rather than a raw InvalidParameterException from RunTask. - logger.warn('ECS payload is large but ECS_PAYLOAD_BUCKET is not set — RunTask may reject it (see #502)', { - task_id: taskId, - bytes: payloadJson.length, - }); + if (!ECS_PAYLOAD_BUCKET) { + throw new Error('PAYLOAD_BOOTSTRAP_INVALID: ECS_PAYLOAD_BUCKET is required; deploy matching infrastructure'); } + const reference = await preparePayloadReference({ + bucket: ECS_PAYLOAD_BUCKET, taskId, backend: 'ecs', payload, + }).catch((error: unknown) => { + throw safeLaunchError(error); + }); const containerEnv = [ { name: 'TASK_ID', value: taskId }, { name: 'REPO_URL', value: String(payload.repo_url ?? '') }, - ...(payload.prompt ? [{ name: 'TASK_DESCRIPTION', value: String(payload.prompt) }] : []), ...(payload.issue_number ? [{ name: 'ISSUE_NUMBER', value: String(payload.issue_number) }] : []), // Single source of truth with the hydrate path in `orchestrator.ts`, which // resolves the same default via `DEFAULT_MAX_TURNS`. A literal here would @@ -217,47 +172,34 @@ export class EcsComputeStrategy implements ComputeStrategy { ...(blueprintConfig.model_id ? [{ name: 'ANTHROPIC_MODEL', value: blueprintConfig.model_id }] : []), ...(blueprintConfig.system_prompt_overrides ? [{ name: 'SYSTEM_PROMPT_OVERRIDES', value: blueprintConfig.system_prompt_overrides }] : []), { name: 'CLAUDE_CODE_USE_BEDROCK', value: '1' }, - // Prefer the S3 pointer; fall back to the inline payload when no bucket is - // configured (keeps small-payload / AgentCore-only deployments working with - // no behavior change). - ...(payloadS3Uri - ? [{ name: 'AGENT_PAYLOAD_S3_URI', value: payloadS3Uri }] - : [{ name: 'AGENT_PAYLOAD', value: payloadJson }]), + { name: 'AGENT_PAYLOAD_REF', value: JSON.stringify(reference) }, ...(payload.github_token_secret_arn ? [{ name: 'GITHUB_TOKEN_SECRET_ARN', value: String(payload.github_token_secret_arn) }] : []), ...(payload.memory_id ? [{ name: 'MEMORY_ID', value: String(payload.memory_id) }] : []), ]; - // Override the container command to run a Python one-liner that: - // 1. Loads the payload — from S3 (AGENT_PAYLOAD_S3_URI) when set, else the - // inline AGENT_PAYLOAD env var (fallback). - // 2. Calls entrypoint.run_task_from_payload(p), which maps the WHOLE payload - // dict to run_task's signature (rename prompt→task_description / - // model_id→anthropic_model, filter to accepted params, coerce str/int). - // This replaces an older hand-listed kwarg subset that silently dropped - // fields such as channel_source/channel_metadata (which meant no - // Linear/Jira reactions or channel MCP on ECS), build_command, - // cedar_policies, base_branch/merge_branches, attachments, trace, user_id, - // etc. Single source of truth in the agent, unit-tested (see - // test_run_task_from_payload). - // 3. Exits with code 0 on success, 1 on failure. - // This bypasses the uvicorn server entirely — no HTTP, no OTEL noise. + // Consume the one-object capability before importing/running the pipeline. + // The helper removes it from os.environ so repo subprocesses do not inherit it. const bootCommand = [ 'python', '-c', - 'import json, os, sys; ' - + 'sys.path.insert(0, "/app/src"); ' + 'import sys; sys.path.insert(0, "/app/src"); ' + + 'from payload_bootstrap import load_ecs_payload; p = load_ecs_payload(); ' + 'from entrypoint import run_task_from_payload; ' - + '_uri = os.environ.get("AGENT_PAYLOAD_S3_URI"); ' - + 'p = (' - + 'json.loads(__import__("boto3").client("s3").get_object(' - + 'Bucket=_uri.split("/",3)[2], Key=_uri.split("/",3)[3])["Body"].read()) ' - + 'if _uri else json.loads(os.environ["AGENT_PAYLOAD"])' - + '); ' + 'r = run_task_from_payload(p); ' + 'sys.exit(0 if r.get("status")=="success" else 1)', ]; + const overrides = { + containerOverrides: [{ + name: ECS_CONTAINER_NAME, + environment: containerEnv, + command: bootCommand, + }], + }; + if (Buffer.byteLength(JSON.stringify(overrides), 'utf8') > ECS_OVERRIDES_LIMIT_BYTES) { + throw new Error('PAYLOAD_BOOTSTRAP_TOO_LARGE: ECS container overrides exceed 8192 bytes'); + } const command = new RunTaskCommand({ cluster: ECS_CLUSTER_ARN, taskDefinition, @@ -278,21 +220,17 @@ export class EcsComputeStrategy implements ComputeStrategy { assignPublicIp: 'DISABLED', }, }, - overrides: { - containerOverrides: [{ - name: ECS_CONTAINER_NAME, - environment: containerEnv, - command: bootCommand, - }], - }, + overrides, }); - const result = await getClient().send(command); + const result = await getClient().send(command).catch((error: unknown) => { + throw safeLaunchError(error); + }); const ecsTask = result.tasks?.[0]; if (!ecsTask?.taskArn) { const failures = result.failures?.map(f => `${f.arn}: ${f.reason}`).join('; ') ?? 'unknown'; - throw new Error(`ECS RunTask returned no task: ${failures}`); + throw new Error(`ECS RunTask returned no task: ${redactPayloadUrls(failures)}`); } logger.info('ECS Fargate task started', { @@ -312,16 +250,17 @@ export class EcsComputeStrategy implements ComputeStrategy { }; } - async pollSession(handle: SessionHandle): Promise { + async pollSession(handle: SessionHandle, options?: SessionControlOptions): Promise { if (handle.strategyType !== 'ecs') { throw new Error('pollSession called with non-ecs handle'); } const { clusterArn, taskArn } = handle; + options?.abortSignal?.throwIfAborted(); const result = await getClient().send(new DescribeTasksCommand({ cluster: clusterArn, tasks: [taskArn], - })); + }), options); const ecsTask = result.tasks?.[0]; if (!ecsTask) { @@ -348,18 +287,19 @@ export class EcsComputeStrategy implements ComputeStrategy { return { status: 'running' }; } - async stopSession(handle: SessionHandle): Promise { + async stopSession(handle: SessionHandle, options?: SessionControlOptions): Promise { if (handle.strategyType !== 'ecs') { throw new Error('stopSession called with non-ecs handle'); } const { clusterArn, taskArn } = handle; try { + options?.abortSignal?.throwIfAborted(); await getClient().send(new StopTaskCommand({ cluster: clusterArn, task: taskArn, reason: 'Stopped by orchestrator', - })); + }), options); logger.info('ECS task stopped', { task_arn: taskArn }); } catch (err) { const errName = err instanceof Error ? err.name : undefined; diff --git a/cdk/src/handlers/shared/strategies/lambda-microvm-strategy.ts b/cdk/src/handlers/shared/strategies/lambda-microvm-strategy.ts index 6e3e8a0ef..0ce6020ba 100644 --- a/cdk/src/handlers/shared/strategies/lambda-microvm-strategy.ts +++ b/cdk/src/handlers/shared/strategies/lambda-microvm-strategy.ts @@ -19,18 +19,29 @@ import { GetMicrovmCommand, + GetMicrovmImageVersionCommand, LambdaMicrovmsClient, MicrovmState, RunMicrovmCommand, TerminateMicrovmCommand, + SuspendMicrovmCommand, + ResumeMicrovmCommand, } from '@aws-sdk/client-lambda-microvms'; -import { DeleteObjectCommand, PutObjectCommand, S3Client } from '@aws-sdk/client-s3'; // Cross-language contract (S9): `microvm_platform_config` is read by BOTH this // producer and `agent/src/server.py`'s `/run` consumer. Imported (not copied) so // `tsc` fails on a renamed field — see `contracts/constants.md`. import sharedConstants from '../../../../../contracts/constants.json'; -import type { ComputeStrategy, SessionHandle, SessionStatus } from '../compute-strategy'; +import type { ComputeStrategy, SessionControlOptions, SessionHandle, SessionLifecycleResult, SessionStatus, SessionStopResult } from '../compute-strategy'; +import { MicrovmStartUncertainError } from '../error-classifier'; import { logger } from '../logger'; +import { validAttemptId } from '../microvm-continuation-types'; +import { microvmErrorIdentity, microvmRequestIdentity } from '../microvm-control'; +import { + MICROVM_IMAGE_CAPABILITY_REQUEST_TIMEOUT_MS, MICROVM_LIFECYCLE_PROTOCOL, + readMicrovmImageMetadata, verifyMicrovmImageLifecycle, +} from '../microvm-image-capability'; +import { claimMicrovmStart, microvmStartRequestHash, saveMicrovmStartHandle, saveMicrovmImageCapability } from '../microvm-start'; +import { deletePayloadReference, preparePayloadReference, redactPayloadUrls } from '../payload-bootstrap'; import type { BlueprintConfig } from '../repo-config'; import { makeClient } from '../ua'; @@ -42,12 +53,14 @@ function getClient(): LambdaMicrovmsClient { return sharedClient; } -let sharedS3Client: S3Client | undefined; -function getS3Client(): S3Client { - if (!sharedS3Client) { - sharedS3Client = makeClient(S3Client); - } - return sharedS3Client; +/** Bound a control request, not the transition itself. A timeout needs reconciliation. */ +export const MICROVM_LIFECYCLE_REQUEST_TIMEOUT_MS = 10_000; + +function controlSignal(options?: SessionControlOptions): AbortSignal { + const limit = AbortSignal.timeout(MICROVM_LIFECYCLE_REQUEST_TIMEOUT_MS); + const signal = options?.abortSignal ? AbortSignal.any([limit, options.abortSignal]) : limit; + signal.throwIfAborted(); + return signal; } /** @@ -80,6 +93,7 @@ const MICROVM_EGRESS_CONNECTOR_ARNS = process.env.MICROVM_EGRESS_CONNECTOR_ARNS; */ const MICROVM_INGRESS_CONNECTOR_ARNS = process.env.MICROVM_INGRESS_CONNECTOR_ARNS; const MICROVM_PAYLOAD_BUCKET = process.env.MICROVM_PAYLOAD_BUCKET; +const HTTP_REQUEST_TIMEOUT = 408; /** * Session wall-clock ceiling passed on EVERY ``RunMicrovm`` call, pinned to the @@ -94,7 +108,7 @@ const MICROVM_PAYLOAD_BUCKET = process.env.MICROVM_PAYLOAD_BUCKET; * so a Blueprint override would be policy without a driver; add one only if a * real need appears. */ -export const MICROVM_MAX_DURATION_SECONDS = 28_800; +export const MICROVM_MAX_DURATION_SECONDS = sharedConstants.microvm_lifecycle.maximum_duration_seconds; /** * Hard service cap on ``runHookPayload`` (bytes), measured live rather than read @@ -112,43 +126,15 @@ export const MICROVM_MAX_DURATION_SECONDS = 28_800; * old 16 384 threshold would have inlined every envelope between 4 097 and * 16 384 bytes and had the service reject all of them. * - * This is the EXACT branch point for the inline/S3-pointer decision, with no - * safety margin — deliberately unlike ``ecs-strategy``, which keeps its inline - * warn line at 6 144 of ECS's 8 192-byte cap. That margin exists because ECS - * counts the *whole* ``containerOverrides`` blob (env vars, command, and payload - * share one budget), so the strategy cannot know how much of the 8 192 the - * payload actually gets. ``runHookPayload`` is a single standalone string, so - * the counted size is exactly what we measure and the boundary is computable. - * - * Consequence worth stating plainly: at 4 KB the **S3-pointer path is the - * dominant one**. A hydrated task payload (prompt + issue thread + repo context) - * essentially always exceeds 4 KB, so the inline branch is the exception (tiny - * repo-less prompts), not the common case. + * V2 always sends a signed payload reference. Enforce this byte limit on the + * final serialized reference; payload/config bytes live in S3, not in the hook. */ const RUN_HOOK_PAYLOAD_LIMIT_BYTES = 4_096; /** - * The ``GetMicrovm`` ``stateReason`` value that means "nothing to report". - * - * Live-observed, not guessed: an orchestrator-initiated ``TerminateMicrovm`` on the - * SUCCESS path leaves the MicroVM ``TERMINATED`` with exactly - * ``stateReason: "Success."`` — trailing period included. Recorded three times in - * ``docs/verification/645-p2-smoke-runbook.md``: **§5.1** ("Finalization called - * `TerminateMicrovm`", the verbatim CLI output), **§6.2** (the suspend/resume - * latency table) and **§2.9** ("Lifecycle — PASS", run 2). A healthy ``RUNNING`` - * MicroVM reports no reason at all (``None``, same §6.2 table). - * - * Normalized away in {@link LambdaMicrovmComputeStrategy.pollSession} so it never - * reaches the reconcile ``detail`` string, where it would append noise to every - * cleanly-finished task. - * - * A bare literal comparison is deliberate and the brittleness is bounded: this is - * a service-owned display string, so an exact match can only fail OPEN — a future - * ``"Success"`` without the period, or a different capitalisation, would leak one - * benign phrase into an operator-facing string. It cannot suppress a real reason, - * which is the direction that would matter. Left out of - * ``contracts/constants.json`` for the same reason: nothing in the agent reads it, - * so it is not a cross-language contract. + * The service reports "Success." after normal termination. Suppress only that + * exact benign display string; preserve other reasons for operator diagnosis. + * A future wording change may add harmless detail but cannot hide a failure. */ const MICROVM_BENIGN_STATE_REASON = 'Success.'; @@ -186,9 +172,10 @@ const MICROVM_BENIGN_STATE_REASON = 'Success.'; * ## What may and may not go in here * * NON-SECRET IDENTIFIERS ONLY — table names, bucket names, log-group names, and - * secret/role **ARNs**. Never a token, never a secret *value*: the envelope is - * written to an S3 object and echoed into MicroVM logs on a hook failure, and - * the agent resolves an ARN itself through its own (SessionRole / + * secret/role **ARNs**. Never a token, never a secret *value*: configuration is + * stored in the worker-readable deployment manifest and task object. Hook + * diagnostics omit payloads and redact signed URLs. The agent resolves an ARN + * itself through its own (SessionRole / * execution-role) credentials. The producer below is a map over exactly the * contract's keys, so a value can only reach the wire by being added to the * contract — an unrelated `process.env` entry (`GITHUB_TOKEN`, @@ -198,7 +185,7 @@ const MICROVM_BENIGN_STATE_REASON = 'Success.'; * * The contract's declaration order is the emission order (`JSON.stringify` * preserves insertion order for string keys), which keeps the serialized - * envelope — and therefore the 4 KB inline/S3 branch decision — deterministic + * manifest serialization deterministic * for a given environment. */ const PLATFORM_CONFIG_CONTRACT = sharedConstants.microvm_platform_config; @@ -351,91 +338,38 @@ export function buildMicrovmPlatformConfig( export const MICROVM_ERROR_MARKER = 'MicroVM'; /** - * Wrap an error escaping a MicroVM control-plane (or payload-upload) call so it + * Wrap an error escaping a MicroVM control-plane or payload-bootstrap call so it * carries {@link MICROVM_ERROR_MARKER} plus the originating operation. * * The AWS exception NAME is spliced into the message explicitly because * ``err.message`` alone omits it (``String(err)`` would include it, but the * classifier is handed the *wrapped* error) and the classifier keys on that - * name. ``cause`` retains the original for anyone who needs ``err.name``. + * name. ``cause`` retains a sanitized name/message copy: SDK errors and their + * nested causes can contain signed URLs or request metadata. * * The wrapper's own ``name`` is intentionally left as ``Error`` so * ``String(wrapped)`` reads ``Error: MicroVM failed: : `` — * marker first, which is the order the classifier patterns document. */ function wrapMicrovmError(operation: string, err: unknown): Error { - const name = err instanceof Error ? err.name : undefined; - const message = err instanceof Error ? err.message : String(err); + const name = err instanceof Error ? redactPayloadUrls(err.name) : undefined; + const message = redactPayloadUrls(err instanceof Error ? err.message : String(err)); const detail = name && name !== 'Error' && !message.includes(name) ? `${name}: ${message}` : message; - return new Error(`${MICROVM_ERROR_MARKER} ${operation} failed: ${detail}`, { cause: err }); -} - -/** - * S3 object key for a task's MicroVM ``/run`` payload. Same key shape as the - * ECS payload bucket (``ecsPayloadKey``): one object per task under its own - * task-id prefix, deleted by the orchestrator at finalize (see - * {@link deleteMicrovmPayload}), with the payload bucket's lifecycle-expiry rule - * (ADR-021 sub-decision 3, ``MICROVM_PAYLOAD_TTL_DAYS``) as the backstop. - */ -export function microvmPayloadKey(taskId: string): string { - return `${taskId}/payload.json`; + const safeCause = new Error(message); + safeCause.name = name ?? 'Error'; + const requestId = microvmErrorIdentity(err).aws_request_id; + if (requestId) Object.assign(safeCause, { $metadata: { requestId } }); + return new Error(`${MICROVM_ERROR_MARKER} ${operation} failed: ${detail}`, { cause: safeCause }); } -/** - * Delete a task's MicroVM ``/run`` payload object. Best-effort: a failed delete - * must never fail the task — the bucket's 1-day lifecycle rule reaps it - * regardless. Called from the orchestrator's ``finalize`` step once the task is - * terminal. No-ops when the payload bucket isn't configured. - * - * WHY this exists rather than leaning on the TTL alone (ECS parity, and a real - * exposure delta): the execution role's payload-bucket grant is - * ``grantRead`` on the WHOLE bucket — it cannot be per-task scoped, because the - * MicroVM has to read its own object before any tenant identity is installed. - * The key shape is ``/payload.json``, so any running MicroVM that knows - * (or guesses) another task's id can read that task's HYDRATED PROMPT — issue - * body, comment thread, repo context. The MicroVM also runs untrusted repo code. - * Relying only on ``MICROVM_PAYLOAD_TTL_DAYS = 1`` left that window open for up - * to ~24 h; deleting at finalize closes it to the task's own lifetime, which is - * exactly the posture ``deleteEcsPayload`` already gives the ECS backend. - * - * ISSUED UNCONDITIONALLY, including for a task whose payload went INLINE (under - * the {@link RUN_HOOK_PAYLOAD_LIMIT_BYTES} cap, so no object was ever written). - * That is on purpose: ``DeleteObject`` on a missing key succeeds, so the call is - * harmless and idempotent, whereas *deciding* to skip it would mean trusting a - * per-task record of the delivery mode — and if that record were ever wrong or - * absent, the skip would leave a real payload behind for the full TTL. Attempting - * always is the fail-safe direction. - * - * The consequence is that a successful call proves a delete was ISSUED, never that - * an object existed — S3 returns nothing that distinguishes the two on an - * unversioned bucket. The log line below says exactly that and no more; an earlier - * "Deleted MicroVM payload object" asserted a deletion that never happened on - * every inline task. +/** Remove task instructions and their saved download capability after finalization. + * Best-effort; bucket lifecycle reaps leftovers. Deployment manifests are shared. */ -export async function deleteMicrovmPayload(taskId: string): Promise { +export async function deleteMicrovmPayload(taskId: string, attemptId?: string): Promise { if (!MICROVM_PAYLOAD_BUCKET) return; - const key = microvmPayloadKey(taskId); - try { - await getS3Client().send(new DeleteObjectCommand({ - Bucket: MICROVM_PAYLOAD_BUCKET, - Key: key, - })); - // "issued", not "deleted": see the docstring. An inline-delivered task has no - // object here and the call still succeeds. - logger.info('MicroVM payload delete issued', { - task_id: taskId, - bucket: MICROVM_PAYLOAD_BUCKET, - key, - }); - } catch (err) { - // Non-fatal — the lifecycle rule is the backstop. - logger.warn('Failed to delete MicroVM payload object (non-fatal)', { - task_id: taskId, - error: err instanceof Error ? err.message : String(err), - }); - } + await deletePayloadReference(MICROVM_PAYLOAD_BUCKET, taskId, attemptId); } /** Split a comma-separated env-var list into trimmed, non-empty entries. */ @@ -509,23 +443,21 @@ function assertImageArn(identifier: string): void { * control-plane state machine the orchestrator can observe through * {@link LambdaMicrovmComputeStrategy.pollSession}. * - * P1 scope is start / poll / stop only. ``suspendSession`` / ``resumeSession`` - * (the interface widening across all three strategies) land in P3 — do NOT add - * them here piecemeal, ADR-021 sub-decision 1 requires them in one commit so - * the exhaustive-``never`` switch culture forces every backend to make a - * compile-checked decision about its suspend semantics. + * Suspend/resume submit service commands; the other two strategies return + * explicit unsupported results. microvm-supervisor supplies gate policy, + * durable intent and state reconciliation. Automatic suspension additionally + * requires compatible image hooks and enabled static/live rollout settings. */ export class LambdaMicrovmComputeStrategy implements ComputeStrategy { readonly type = 'lambda-microvm'; async startSession(input: { taskId: string; - /** Accepted to satisfy the ComputeStrategy interface. MicroVMs have no - * workload-token-injecting runtime (they inherit the ECS env-var identity - * posture until #249 / ADR-016 redesign the seam), so this is unused. */ + /** Checked against the stored task owner before claiming a start receipt. */ userId: string; payload: Record; blueprintConfig: BlueprintConfig; + microvmImage?: { readonly imageArn: string; readonly imageVersion: string }; }): Promise { if (!MICROVM_IMAGE_IDENTIFIER || !MICROVM_EXECUTION_ROLE_ARN || !MICROVM_EGRESS_CONNECTOR_ARNS || !MICROVM_PAYLOAD_BUCKET) { // Config/deploy mismatch: this repo is compute_type=lambda-microvm but the @@ -545,101 +477,28 @@ export class LambdaMicrovmComputeStrategy implements ComputeStrategy { } const { taskId, payload } = input; + const attemptId = payload.attempt_id ?? taskId; + if (!validAttemptId(attemptId)) throw new Error('MICROVM_START_STATE_INVALID: invalid worker attempt'); // An identifier that is not an ARN cannot launch anything — check before the // payload upload so a misconfiguration never leaves an orphan S3 object. assertImageArn(MICROVM_IMAGE_IDENTIFIER); + if (input.microvmImage && (input.microvmImage.imageArn !== MICROVM_IMAGE_IDENTIFIER + || !/^\d+\.\d+$/.test(input.microvmImage.imageVersion))) { + throw new Error('MICROVM_CONTINUATION_IMAGE_INVALID: recovery must use a version of the configured image'); + } + const imageVersion = input.microvmImage?.imageVersion ?? MICROVM_IMAGE_VERSION; - // Payload delivery (ADR-021 sub-decision 3): the `/run` lifecycle hook - // receives `runHookPayload` as its request body, capped at 4 KB by the - // service. The hydrated_context essentially always blows that, so the - // S3-pointer path (mirroring ECS #502) is the DOMINANT one here and the - // inline branch is the exception. The MicroVM EXECUTION role holds the read - // grant, exactly as the ECS task role does today. - // - // Three keys, deliberately mirroring the ECS container env contract - // (AGENT_PAYLOAD / AGENT_PAYLOAD_S3_URI) so the agent's `/run` hook has one - // self-describing shape to branch on: - // { "agent_payload": {...}, "platform_config": {...} } — inline - // { "agent_payload_s3_uri": "…", "platform_config": {...} } — pointer - // - // `platform_config` (see MICROVM_PLATFORM_CONFIG_KEYS) rides in BOTH forms, - // and is ALSO merged into the S3 object on the pointer path: - // s3://…//payload.json = { ...agent_payload, "platform_config": {…} } - // The duplication is deliberate and cheap (a few hundred bytes). It is the - // agent's env-block substitute — nothing else delivers it, because the - // snapshot must not bake it in — so it must be reachable whether the agent - // reads it off the hook body before fetching S3 or out of the fetched object. + // The manifest authenticates deployment settings through the worker's IAM + // grant. Payload access uses a single-object URL, saved outside TaskTable. const platformConfig = buildMicrovmPlatformConfig(); - const inlineEnvelope = JSON.stringify({ agent_payload: payload, platform_config: platformConfig }); - // Measure the SERIALIZED envelope, not the bare payload: the envelope is - // what the service counts against the 4 KB cap, and `platform_config` is part - // of it — which is precisely why nearly everything lands on the S3 path. - // Byte length (not String.length) because a multi-byte prompt/diff makes - // chars an undercount. - const inlineBytes = Buffer.byteLength(inlineEnvelope, 'utf8'); - - let runHookPayload: string; - let payloadS3Uri: string | undefined; - // EXACT boundary: `<= limit` inlines, `> limit` uploads. The service accepts - // 4 096 bytes and rejects 4 097 (measured), so 4 096 must still go inline. - if (inlineBytes <= RUN_HOOK_PAYLOAD_LIMIT_BYTES) { - runHookPayload = inlineEnvelope; - } else { - const key = microvmPayloadKey(taskId); - const uri = `s3://${MICROVM_PAYLOAD_BUCKET}/${key}`; - const pointerEnvelope = JSON.stringify({ - agent_payload_s3_uri: uri, - platform_config: platformConfig, - }); - // The pointer envelope is the LAST RESORT — there is no smaller shape to - // fall back to — so check it BEFORE the upload (an upload followed by a - // throw would leave an orphan object for the lifecycle rule to reap) and - // name the one thing an operator can actually act on. Unreachable in - // practice: the pointer plus all thirteen identifiers is well under 4 KB. - const pointerBytes = Buffer.byteLength(pointerEnvelope, 'utf8'); - if (pointerBytes > RUN_HOOK_PAYLOAD_LIMIT_BYTES) { - throw new Error( - `The MicroVM /run pointer envelope is ${pointerBytes} bytes, over the service's ` - + `${RUN_HOOK_PAYLOAD_LIMIT_BYTES}-byte runHookPayload cap, with the payload already moved ` - + 'to S3. The remaining size is the S3 URI plus the platform_config identifiers, so a ' - + 'pathologically long table/bucket/ARN name is the only possible cause — shorten the ' - + 'stack name (physical resource names derive from it) and redeploy.', - ); - } - // The S3 object carries the payload with `platform_config` merged in at the - // top level, so an agent that fetches the object gets the config with it. - // Platform config wins on a key collision — the payload has no - // `platform_config` key today, and if one ever appeared the platform's - // value is the authoritative one. - const payloadJson = JSON.stringify({ ...payload, platform_config: platformConfig }); - try { - await getS3Client().send(new PutObjectCommand({ - Bucket: MICROVM_PAYLOAD_BUCKET, - Key: key, - Body: payloadJson, - ContentType: 'application/json', - })); - } catch (err) { - // Marked so the classifier attributes an upload failure to this backend - // rather than letting a bare S3 exception name fall through to UNKNOWN. - throw wrapMicrovmError('payload upload', err); - } - payloadS3Uri = uri; - runHookPayload = pointerEnvelope; - logger.info('Wrote MicroVM run-hook payload to S3', { - task_id: taskId, - bytes: Buffer.byteLength(payloadJson, 'utf8'), - inline_bytes: inlineBytes, - inline_limit_bytes: RUN_HOOK_PAYLOAD_LIMIT_BYTES, - uri: payloadS3Uri, - }); - } // Explicit ingress control (F7, live 2026-07-31): `RunMicrovm` does NOT // default to "no ingress" — omitting the field attaches the AWS-managed - // PUBLIC `HTTP_INGRESS` connector and mints a public - // `*.lambda-microvm..on.aws` endpoint. So the field is ALWAYS sent. + // PUBLIC `HTTP_INGRESS` connector. So the field is ALWAYS sent. + // A service endpoint URL is also returned with `NO_INGRESS`; the URL alone + // does not establish guest reachability. Requests require a MicroVM auth + // token, and an unauthenticated 403 proves only that authentication check. // // The env var is unconditional in every CDK-deployed stack (its prop is // required), so in practice this always takes the `configuredIngress` branch @@ -651,9 +510,9 @@ export class LambdaMicrovmComputeStrategy implements ComputeStrategy { ? configuredIngress : [noIngressConnectorArn()]; - const command = new RunMicrovmCommand({ + const request = { imageIdentifier: MICROVM_IMAGE_IDENTIFIER, - ...(MICROVM_IMAGE_VERSION && { imageVersion: MICROVM_IMAGE_VERSION }), + ...(imageVersion && { imageVersion }), executionRoleArn: MICROVM_EXECUTION_ROLE_ARN, // Egress rides the platform VPC through an egress network connector so the // DNS Firewall / security-group / flow-log stack applies unchanged @@ -662,7 +521,7 @@ export class LambdaMicrovmComputeStrategy implements ComputeStrategy { // Never omitted — see the comment above. `NO_INGRESS` is the suppression // mechanism, not an empty list. ingressNetworkConnectors, - runHookPayload, + runHookPayload: 'payload-bootstrap-v2', maximumDurationInSeconds: MICROVM_MAX_DURATION_SECONDS, // `idlePolicy` is OMITTED — never passed, in any phase (ADR-021 // sub-decision 1, asserted by an invariant unit test). MicroVM idle @@ -673,15 +532,39 @@ export class LambdaMicrovmComputeStrategy implements ComputeStrategy { // is present, so omission is the unambiguous disabled state. Suspension is // orchestrator-owned (P3) — do NOT reintroduce this field. // - // `clientToken` is also deliberately omitted. It is an idempotency token, - // and the one place a MicroVM start is retried is `startSessionWithRetry`, - // which retries precisely BECAUSE the first attempt FAILED. Passing a - // task-derived token there would ask the service to dedupe against that - // failed attempt and could replay its outcome instead of genuinely - // retrying — turning the auto-retry into a no-op. Session start is already - // idempotent by construction at the ABCA level (no clone, commit, or PR has - // happened yet), so the token buys nothing and risks the retry. - }); + // The receipt below supplies a task-stable token across new SDK commands. + }; + + const requestHash = microvmStartRequestHash( + { ...request, payloadBucket: MICROVM_PAYLOAD_BUCKET }, { ...payload, platform_config: platformConfig }, + ); + const claim = await claimMicrovmStart(taskId, input.userId, requestHash, attemptId); + if (claim.closed) { + if (claim.handle) await this.stopSession(claim.handle); + throw new Error('MICROVM_START_TASK_CLOSED: task became terminal before session start'); + } + if (claim.handle) return claim.handle; + const reference = await preparePayloadReference({ + bucket: MICROVM_PAYLOAD_BUCKET, + taskId, + backend: 'lambda-microvm', + payload, + platformConfig, + ...(attemptId !== taskId && { attemptId }), + }).catch((error: unknown) => { throw wrapMicrovmError('payload bootstrap', error); }); + const runHookPayload = JSON.stringify(reference); + if (Buffer.byteLength(runHookPayload, 'utf8') > RUN_HOOK_PAYLOAD_LIMIT_BYTES) { + throw new Error('PAYLOAD_BOOTSTRAP_TOO_LARGE: launch reference exceeds the MicroVM hook limit'); + } + // Uploads can take time. Observe cancellation/another saved handle again + // immediately before the service call, using the same immutable request. + const latest = await claimMicrovmStart(taskId, input.userId, requestHash, attemptId); + if (latest.closed) { + if (latest.handle) await this.stopSession(latest.handle); + throw new Error('MICROVM_START_TASK_CLOSED: task became terminal before session start'); + } + if (latest.handle) return latest.handle; + const command = new RunMicrovmCommand({ ...request, runHookPayload, clientToken: latest.clientToken }); let result; try { @@ -691,16 +574,29 @@ export class LambdaMicrovmComputeStrategy implements ComputeStrategy { // / `ResourceNotFoundException` from THIS backend classify as MicroVM // faults, while identically-named AgentCore/ECS errors keep their existing // classification. See MICROVM_ERROR_MARKER. - throw wrapMicrovmError('RunMicrovm', err); + const wrapped = wrapMicrovmError('RunMicrovm', err); + const serviceError = err as { name?: string; $metadata?: { httpStatusCode?: number } }; + const httpStatus = serviceError?.$metadata?.httpStatusCode; + // A service timeout can carry a 4xx status without proving that creation + // never happened. Preserve uncertainty across any subsequent rejection. + const timedOut = httpStatus === HTTP_REQUEST_TIMEOUT + || ['TimeoutError', 'RequestTimeout', 'RequestTimeoutException'].includes(serviceError?.name ?? ''); + const knownRejection = !timedOut && (httpStatus !== undefined + ? httpStatus >= 400 && httpStatus < 500 + : ['AccessDeniedException', 'UnauthorizedException', 'ValidationException', + 'InvalidParameterValueException', 'ResourceNotFoundException', 'ThrottlingException', + 'TooManyRequestsException', 'ServiceQuotaExceededException', 'ConflictException'] + .includes(serviceError?.name ?? '')); + if (!knownRejection) throw new MicrovmStartUncertainError(wrapped.message, { cause: wrapped }); + throw wrapped; } const { microvmId, endpoint } = result; if (!microvmId || !endpoint) { - // A malformed response means a MicroVM may ALREADY BE RUNNING (and billing) - // that no caller will ever receive a handle for — nothing self-terminates on - // this substrate. Reap it here, best-effort, before failing: this is the one - // orphan window the orchestrator's own catch cannot cover, because - // `startSession` never returned a handle to it. + // A malformed response may describe an existing VM without a usable handle. + // The eight-hour service lifetime is a backstop, not prompt cleanup. + // Clean up any available ID here: the caller will not receive a handle + // it can use to stop this VM. if (microvmId) { await this.terminateBestEffort(microvmId, 'incomplete RunMicrovm response'); } @@ -709,28 +605,76 @@ export class LambdaMicrovmComputeStrategy implements ComputeStrategy { // failed` bucket with "Check AgentCore Runtime or ECS cluster health" — // advice that names the wrong substrate entirely. `RunMicrovm` is the // operation because that is the call whose response is malformed. - throw wrapMicrovmError( + const incomplete = wrapMicrovmError( 'RunMicrovm', new Error( `RunMicrovm returned an incomplete response (microvmId=${microvmId ?? 'missing'}, ` + `endpoint=${endpoint ? 'present' : 'missing'}, state=${result.state ?? 'unknown'})`, ), ); + if (!microvmId) throw new MicrovmStartUncertainError(incomplete.message, { cause: incomplete }); + throw incomplete; } - // Image ARN/version is logged, NOT carried in the handle (ADR-021 - // sub-decision 1) — it is deployment-time config, and this line is the - // diagnostic record of which snapshot a given session actually booted. + let handle: Extract = { + sessionId: microvmId, + strategyType: 'lambda-microvm', + microvmId, + endpoint, + ...readMicrovmImageMetadata({ imageArn: result.imageArn, imageVersion: result.imageVersion }), + }; + try { + await saveMicrovmStartHandle(taskId, latest.clientToken, handle); + } catch (err) { + // The write may have committed before its response was lost. Recover that + // receipt before destroying a computer whose handle is already durable. + try { + const saved = await claimMicrovmStart(taskId, input.userId, requestHash, attemptId); + if (!saved.closed && saved.handle?.microvmId === microvmId) return saved.handle; + } catch (readErr) { + logger.warn('Could not reconcile the MicroVM start receipt', { task_id: taskId, error: String(readErr) }); + } + await this.terminateBestEffort(microvmId, 'start receipt could not save handle'); + throw new Error(`MICROVM_START_RECEIPT_SAVE_FAILED: ${String(err)}`, { cause: err }); + } + + // Persist the known worker above BEFORE optional image discovery. A crash or + // failed lookup must not widen the orphan window or break ordinary coding. + // Never infer capability from a requested pin or the deployment's latest image. + if (handle.imageArn === MICROVM_IMAGE_IDENTIFIER && handle.imageVersion) { + try { + const identity = { imageArn: handle.imageArn, imageVersion: handle.imageVersion }; + const version = await getClient().send(new GetMicrovmImageVersionCommand({ + imageIdentifier: identity.imageArn, imageVersion: identity.imageVersion, + }), { abortSignal: AbortSignal.timeout(MICROVM_IMAGE_CAPABILITY_REQUEST_TIMEOUT_MS) }); + if (verifyMicrovmImageLifecycle(identity, version)) { + const capable = { ...handle, lifecycleProtocol: MICROVM_LIFECYCLE_PROTOCOL }; + await saveMicrovmImageCapability(taskId, latest.clientToken, capable); + handle = capable; + } + } catch (error) { + // Explicit degraded mode: the saved worker remains usable, with new + // suspension disabled. Do not expose image environment or AWS error text. + const name = (error as { name?: unknown })?.name; + logger.warn('MicroVM image capability unavailable; automatic suspension remains disabled', { + task_id: taskId, + microvm_id: microvmId, + error_type: typeof name === 'string' && /^[A-Za-z0-9_]{1,100}$/.test(name) ? name : 'Error', + }); + } + } + + // The durable handle carries actual identity/capability for later decisions. logger.info('Lambda MicroVM session started', { task_id: taskId, microvm_id: microvmId, state: result.state, image_identifier: MICROVM_IMAGE_IDENTIFIER, image_arn: result.imageArn, - image_version: result.imageVersion ?? MICROVM_IMAGE_VERSION, + image_version: handle.imageVersion ?? null, + lifecycle_protocol: handle.lifecycleProtocol ?? 'unverified', maximum_duration_seconds: MICROVM_MAX_DURATION_SECONDS, - payload_delivery: payloadS3Uri ? 's3_pointer' : 'inline', - ...(payloadS3Uri && { payload_s3_uri: payloadS3Uri }), + payload_delivery: 'signed_reference', // KEY NAMES only, never values: this is the one operator-visible record of // which optional platform identifiers a given session actually received, and // "the agent said ARTIFACTS_BUCKET_NAME is not configured" is otherwise a @@ -738,27 +682,15 @@ export class LambdaMicrovmComputeStrategy implements ComputeStrategy { platform_config_keys: Object.keys(platformConfig), }); - return { - // sessionId = microvmId, mirroring the ECS variant's "sessionId = the - // substrate identifier" precedent (ECS uses the task ARN) rather than - // AgentCore's fresh UUID — AgentCore only needs a UUID because - // `runtimeSessionId` is a caller-minted value that must be ≥ 33 chars. - // Here the substrate mints the id, every lifecycle API keys on it, and it - // is what an operator needs to correlate `TaskRecord.session_id` with the - // MicroVM in logs/console. A second synthetic UUID would add a - // non-actionable identifier and leave `session_id` un-joinable. - sessionId: microvmId, - strategyType: 'lambda-microvm', - microvmId, - endpoint, - }; + // Use AWS's identifier for both lifecycle calls and TaskRecord.session_id. + return handle; } /** * Report the substrate's view of the session — MECHANICALLY. No task-state * interpretation happens here (ADR-021 sub-decision 1): this method sees only * the handle, so the health rules that need the task's DynamoDB status live in - * the orchestrator (``reconcileMicrovmSubstrateState``). + * the durable supervisor (``superviseMicrovm``). * * State mapping: * - ``PENDING`` / ``RUNNING`` → ``running`` (PENDING is still booting, the @@ -766,33 +698,28 @@ export class LambdaMicrovmComputeStrategy implements ComputeStrategy { * - ``SUSPENDING`` / ``SUSPENDED`` → ``suspended``. SUSPENDING is folded in * because the VM is already on its way to frozen; reporting ``running`` * would tell the orchestrator compute is still progressing when it is not. - * Both map to a state the orchestrator treats as benign-or-anomalous - * depending on the task status, never as a failure. (``SUSPENDING`` was - * never observable live — suspend reaches ``SUSPENDED`` in under a second - * — so nothing may WAIT for it; it is mapped for completeness only.) + * The supervisor checks task/gate state to distinguish an expected wait + * from an anomaly. When a wake is needed, it saves wake intent and waits + * for SUSPENDED before issuing ResumeMicrovm. Unresolved transitions have + * a bounded recovery deadline. * - ``TERMINATING`` / ``TERMINATED`` → ``completed``. Both are terminal or - * terminal-bound and carry no exit code, so "the substrate is gone" is all - * the strategy can honestly say; whether that is success or failure is the - * orchestrator's call (it cross-references the DynamoDB status). This is + * terminal-bound and carry no exit code. TERMINATING confirms shutdown is + * underway, not that it has finished. The orchestrator determines task + * success or failure by cross-referencing DynamoDB status. This is * the load-bearing terminal signal: a terminated MicroVM stays observable * as ``TERMINATED`` for at least ~10 minutes (live-measured), so a poller * that waited for NotFound would spin on a finished VM. - * - anything else (an unrecognized future state) → ``running``, so a service - * enum addition can never fail a healthy task. + * - anything else (an unrecognized future state) → ``running`` with explicit + * UNKNOWN state. The supervisor applies bounded recovery instead of + * immediately declaring the task finished. + * ``microvmState`` also reports the explicit observed state (or local UNKNOWN / + * NOT_FOUND). P3 uses it to distinguish readiness from the coarse status. * - * ``stateReason`` is carried through on every mapped state as - * ``SessionStatus.reason``, VERBATIM and uninterpreted. It is the substrate's - * own account of WHY, and mapping it away is what made the dominant runtime - * failure unreadable: a ``/run`` hook 4xx self-terminates the VM within ~12 s - * (``docs/verification/645-p2-smoke-runbook.md`` §6.1) with - * ``stateReason = "Run lifecycle hook returned HTTP status 400. Please check - * your hook endpoint and application logs for more details."`` — and because - * ``TERMINATED → completed`` has no error slot, the orchestrator's reconcile - * detail read ``"substrate state completed"``, naming none of the three causes - * its remedy suggested. Reporting the reason keeps this method mechanical (no - * branch reads it) while giving the orchestrator something true to say. + * Service failure reasons are preserved in SessionStatus.reason so a failed + * hook is not reduced to "substrate completed". Suppress only the known benign + * success string; the orchestrator decides whether the task itself succeeded. */ - async pollSession(handle: SessionHandle): Promise { + async pollSession(handle: SessionHandle, options?: SessionControlOptions): Promise { if (handle.strategyType !== 'lambda-microvm') { throw new Error('pollSession called with non-lambda-microvm handle'); } @@ -800,11 +727,22 @@ export class LambdaMicrovmComputeStrategy implements ComputeStrategy { let state: string | undefined; let stateReason: string | undefined; + let requestIdentity: ReturnType = {}; + let lifetime: Pick = {}; try { const result = await getClient().send(new GetMicrovmCommand({ microvmIdentifier: microvmId, - })); + }), { abortSignal: controlSignal(options) }); state = result.state; + requestIdentity = microvmRequestIdentity(result); + const startedAtMs = result.startedAt instanceof Date ? result.startedAt.getTime() : NaN; + if (Number.isSafeInteger(startedAtMs) && startedAtMs >= 0 + && Number.isSafeInteger(result.maximumDurationInSeconds) && result.maximumDurationInSeconds! > 0) { + lifetime = { + microvmStartedAtMs: startedAtMs, + microvmMaximumDurationSeconds: result.maximumDurationInSeconds, + }; + } // `Success.` is the service's own "nothing to report" value on a clean // termination — carrying it would append noise to every healthy task's // detail string, so it is normalized away here rather than filtered at @@ -838,8 +776,9 @@ export class LambdaMicrovmComputeStrategy implements ComputeStrategy { if (err instanceof Error && err.name === 'ResourceNotFoundException') { logger.info('MicroVM not found on poll — treating as terminal', { microvm_id: microvmId, + ...microvmErrorIdentity(err), }); - return { status: 'completed' }; + return { status: 'completed', microvmState: 'NOT_FOUND' }; } throw wrapMicrovmError('GetMicrovm', err); } @@ -847,10 +786,10 @@ export class LambdaMicrovmComputeStrategy implements ComputeStrategy { switch (state) { case MicrovmState.PENDING: case MicrovmState.RUNNING: - return { status: 'running', ...(stateReason && { reason: stateReason }) }; + return { status: 'running', microvmState: state, ...lifetime, ...(stateReason && { reason: stateReason }) }; case MicrovmState.SUSPENDING: case MicrovmState.SUSPENDED: - return { status: 'suspended', ...(stateReason && { reason: stateReason }) }; + return { status: 'suspended', microvmState: state, ...lifetime, ...(stateReason && { reason: stateReason }) }; case MicrovmState.TERMINATING: case MicrovmState.TERMINATED: if (stateReason) { @@ -860,18 +799,22 @@ export class LambdaMicrovmComputeStrategy implements ComputeStrategy { // where the operator happens to be looking. logger.warn('MicroVM reached a terminal state with a substrate reason', { microvm_id: microvmId, + image_arn: handle.imageArn, + image_version: handle.imageVersion, + ...requestIdentity, state, state_reason: stateReason, }); } - return { status: 'completed', ...(stateReason && { reason: stateReason }) }; + return { status: 'completed', microvmState: state, ...lifetime, ...(stateReason && { reason: stateReason }) }; default: logger.warn('Unrecognized MicroVM state — reporting running', { microvm_id: microvmId, + ...requestIdentity, state, ...(stateReason && { state_reason: stateReason }), }); - return { status: 'running', ...(stateReason && { reason: stateReason }) }; + return { status: 'running', microvmState: 'UNKNOWN', ...lifetime, ...(stateReason && { reason: stateReason }) }; } } @@ -883,18 +826,67 @@ export class LambdaMicrovmComputeStrategy implements ComputeStrategy { * MicroVMs may be leaking) from "something else" (worth a warning). * * ADR-021: termination is the active cleanup path — it must not rely on - * ``maximumDurationInSeconds`` expiring, which would keep paying for an - * 8-hour reservation after the task is done. Live verification made that - * mandatory rather than belt-and-braces: a hook-less MicroVM reached - * ``RUNNING`` in 12 s and stayed ``RUNNING`` indefinitely with no - * ``stateReason`` — nothing self-terminates, so nothing cleans up if the - * orchestrator does not. + * ``maximumDurationInSeconds`` expiring. The service enforces an eight-hour + * lifetime (including suspended time), but task completion does not itself + * terminate the VM. Explicit termination avoids paying for unused running time + * until that limit. */ - async stopSession(handle: SessionHandle): Promise { + async stopSession(handle: SessionHandle, options?: SessionControlOptions): Promise { if (handle.strategyType !== 'lambda-microvm') { throw new Error('stopSession called with non-lambda-microvm handle'); } - await this.terminateBestEffort(handle.microvmId, 'session stop'); + return this.terminateBestEffort(handle.microvmId, 'session stop', options); + } + + /** Submit a suspend request; the caller owns gate checks and state reconciliation. */ + async suspendSession(handle: SessionHandle, options?: SessionControlOptions): Promise { + return this.requestLifecycle('suspendSession', handle, options); + } + + /** Submit a resume request; acknowledgement alone does not establish RUNNING. */ + async resumeSession(handle: SessionHandle, options?: SessionControlOptions): Promise { + return this.requestLifecycle('resumeSession', handle, options); + } + + private async requestLifecycle( + operation: 'suspendSession' | 'resumeSession', + handle: SessionHandle, + options?: SessionControlOptions, + ): Promise { + if (handle.strategyType !== 'lambda-microvm') { + throw new Error(`${operation} called with non-lambda-microvm handle`); + } + if (typeof handle.microvmId !== 'string' || !handle.microvmId.trim()) { + throw new Error(`${operation} requires a non-empty MicroVM identifier`); + } + const suspend = operation === 'suspendSession'; + const request = { microvmIdentifier: handle.microvmId }; + const startedAt = Date.now(); + const diagnostic = { + operation: suspend ? 'SuspendMicrovm' : 'ResumeMicrovm', + microvm_id: handle.microvmId, + image_arn: handle.imageArn, + image_version: handle.imageVersion, + }; + logger.info('MicroVM lifecycle request started', diagnostic); + try { + const response = await getClient().send( + suspend ? new SuspendMicrovmCommand(request) : new ResumeMicrovmCommand(request), + { abortSignal: controlSignal(options) }, + ); + // An accepted command is distinct from the observed transition/guest acknowledgment. + logger.info('MicroVM lifecycle request acknowledged', { + ...diagnostic, elapsed_ms: Date.now() - startedAt, ...microvmRequestIdentity(response), + }); + } catch (error) { + // Includes Conflict/NotFound: neither proves the desired state was reached. + // Even a timeout may have committed; the durable caller must observe again. + logger.warn('MicroVM lifecycle request failed', { + ...diagnostic, elapsed_ms: Date.now() - startedAt, ...microvmErrorIdentity(error), + }); + throw wrapMicrovmError(suspend ? 'SuspendMicrovm' : 'ResumeMicrovm', error); + } + return { supported: true }; } /** @@ -911,40 +903,63 @@ export class LambdaMicrovmComputeStrategy implements ComputeStrategy { * @param reason - why we are terminating, for the log line (the orphan-reap and * the ordinary finalize path are worth telling apart in CloudWatch). */ - private async terminateBestEffort(microvmId: string, reason: string): Promise { + private async terminateBestEffort(microvmId: string, reason: string, options?: SessionControlOptions): Promise { + const startedAt = Date.now(); try { - await getClient().send(new TerminateMicrovmCommand({ + const response = await getClient().send(new TerminateMicrovmCommand({ microvmIdentifier: microvmId, - })); - logger.info('Lambda MicroVM terminated', { microvm_id: microvmId, reason }); + }), { abortSignal: controlSignal(options) }); + logger.info('Lambda MicroVM termination requested', { + microvm_id: microvmId, reason, elapsed_ms: Date.now() - startedAt, ...microvmRequestIdentity(response), + }); + return { outcome: 'requested' }; } catch (err) { - const errName = err instanceof Error ? err.name : undefined; - if (errName === 'ResourceNotFoundException' || errName === 'ConflictException') { - // Already terminated (reaped) or already TERMINATING — the desired end - // state either way. ConflictException joins the info branch because a - // concurrent terminate (orchestrator finalize racing a user cancel) is - // routine here, and warning on it would train operators to ignore warns. - logger.info('MicroVM already terminated or terminating', { + const identity = microvmErrorIdentity(err); + const errName = identity.error_type; + if (errName === 'ConflictException') { + try { + const observed = await getClient().send(new GetMicrovmCommand({ + microvmIdentifier: microvmId, + }), { abortSignal: controlSignal(options) }); + if (observed?.state === 'TERMINATED') { + logger.info('MicroVM already terminated during concurrent cleanup', { microvm_id: microvmId, reason }); + return { outcome: 'terminated' }; + } + } catch (confirmationError) { + if (microvmErrorIdentity(confirmationError).error_type === 'ResourceNotFoundException') { + logger.info('MicroVM no longer found during concurrent cleanup', { microvm_id: microvmId, reason }); + return { outcome: 'not-found' }; + } + logger.warn('MicroVM termination conflict could not be confirmed', { + microvm_id: microvmId, ...microvmErrorIdentity(confirmationError), + }); + } + } + if (errName === 'ResourceNotFoundException') { + logger.info('MicroVM no longer found during termination', { microvm_id: microvmId, reason, error_type: errName, }); + return { outcome: 'not-found' }; } else if (errName === 'ThrottlingException' || errName === 'AccessDeniedException') { // A throttle or a missing lambda:TerminateMicrovm grant means the VM is // probably STILL RUNNING and billing — escalate. logger.error('Failed to terminate MicroVM', { microvm_id: microvmId, reason, - error_type: errName, - error: err instanceof Error ? err.message : String(err), + ...identity, }); } else { logger.warn('Failed to terminate MicroVM (best-effort)', { microvm_id: microvmId, reason, - error: err instanceof Error ? err.message : String(err), + ...identity, }); } + // Conflict can mean another lifecycle operation is in flight. It does not + // prove termination; retain that uncertainty for caller orphan reporting. + return { outcome: 'unconfirmed', ...identity }; } } } @@ -952,7 +967,7 @@ export class LambdaMicrovmComputeStrategy implements ComputeStrategy { /** * Re-exported so tests and future callers can assert the documented cap without * duplicating the literal. This is BOTH the service's limit and our exact - * inline/S3-pointer branch point — there is no separate threshold. + * maximum serialized v2 launch-reference size. */ export const MICROVM_RUN_HOOK_PAYLOAD_LIMIT_BYTES = RUN_HOOK_PAYLOAD_LIMIT_BYTES; diff --git a/cdk/src/handlers/shared/task-cancellation.ts b/cdk/src/handlers/shared/task-cancellation.ts new file mode 100644 index 000000000..ad3017f27 --- /dev/null +++ b/cdk/src/handlers/shared/task-cancellation.ts @@ -0,0 +1,175 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { randomUUID } from 'node:crypto'; +import { GetCommand, TransactWriteCommand, UpdateCommand } from '@aws-sdk/lib-dynamodb'; +import { ulid } from 'ulid'; +import { logger } from './logger'; +import type { TaskRecord } from './types'; +import { makeDocClient } from './ua'; +import { computeTtlEpoch } from './validation'; +import { TaskStatus, TERMINAL_STATUSES } from '../../constructs/task-status'; + +const ddb = makeDocClient(); +const MAX_ATTEMPTS = 3; +const STATE_TIMEOUT_MS = 5_000; + +export class TaskCancellationError extends Error { + constructor(public readonly reason: 'missing' | 'forbidden' | 'terminal' | 'conflict') { + super(`Task cancellation ${reason}`); + this.name = 'TaskCancellationError'; + } +} + +interface CancellationOptions { + readonly userId: string; + readonly taskTable: string; + readonly approvalsTable?: string; + readonly eventsTable?: string; + readonly retentionDays: number; +} + +/** The latest pre-cancel record supplies the compute handle for active cleanup. */ +export interface TaskCancellationResult { + readonly task: TaskRecord; + readonly cancelledAt: string; + readonly cancelledRequestId?: string; +} + +function conditionalConflict(error: unknown): boolean { + const value = error as { name?: string; CancellationReasons?: Array<{ Code?: string }> }; + return value?.name === 'ConditionalCheckFailedException' + || (value?.name === 'TransactionCanceledException' + && Boolean(value.CancellationReasons?.some(reason => reason.Code === 'ConditionalCheckFailed'))); +} + +/** + * Cancel exactly the observed task/gate. Approval, timeout and gate changes race + * through conditional writes; a conflict reloads state before another attempt. + * A decision that already committed remains a decision, never a cancellation. + */ +export async function cancelTaskState( + initialTask: TaskRecord, + options: CancellationOptions, +): Promise { + const abortSignal = AbortSignal.timeout(STATE_TIMEOUT_MS); + let task = initialTask; + for (let attempt = 0; attempt < MAX_ATTEMPTS; attempt++) { + if (task.user_id !== options.userId) throw new TaskCancellationError('forbidden'); + if (TERMINAL_STATUSES.includes(task.status)) throw new TaskCancellationError('terminal'); + const requestId = task.awaiting_approval_request_id; + let closeApproval = false; + if (requestId) { + if (!options.approvalsTable || !options.eventsTable) { + throw new Error('Cancellation of an approval wait requires approvals and events tables'); + } + const response = await ddb.send(new GetCommand({ + TableName: options.approvalsTable, + Key: { task_id: task.task_id, request_id: requestId }, + ConsistentRead: true, + }), { abortSignal }); + const approval = response.Item; + closeApproval = approval?.status === 'PENDING' && approval.user_id === options.userId; + if (approval && approval.user_id !== options.userId) { + // A malformed approval must not prevent the owner stopping their task. + // Do not change another user's approval record. + logger.error('Cancellation found an approval owner mismatch', { + event: 'approval_cancel_owner_mismatch', task_id: task.task_id, request_id: requestId, + }); + } + } + const now = new Date().toISOString(); + const update = { + TableName: options.taskTable, + Key: { task_id: task.task_id }, + UpdateExpression: 'SET #status = :cancelled, updated_at = :now, completed_at = :now, status_created_at = :sca, #ttl = :ttl', + ConditionExpression: 'attribute_exists(task_id) AND user_id = :user AND #status = :observed AND ' + + (requestId + ? 'awaiting_approval_request_id = :request' + : '(attribute_not_exists(awaiting_approval_request_id) OR awaiting_approval_request_id = :request)'), + ExpressionAttributeNames: { '#status': 'status', '#ttl': 'ttl' }, + ExpressionAttributeValues: { + ':cancelled': TaskStatus.CANCELLED, + ':now': now, + ':sca': `${TaskStatus.CANCELLED}#${now}`, + ':ttl': computeTtlEpoch(options.retentionDays), + ':user': options.userId, + ':observed': task.status, + ':request': requestId ?? null, + }, + }; + try { + if (closeApproval) { + await ddb.send(new TransactWriteCommand({ + ClientRequestToken: randomUUID(), + TransactItems: [ + { Update: update }, + { + Update: { + TableName: options.approvalsTable!, + Key: { task_id: task.task_id, request_id: requestId! }, + UpdateExpression: 'SET #status = :cancelled, decided_at = :now, cancellation_reason = :reason', + ConditionExpression: '#status = :pending AND user_id = :user', + ExpressionAttributeNames: { '#status': 'status' }, + ExpressionAttributeValues: { + ':cancelled': 'CANCELLED', + ':pending': 'PENDING', + ':user': options.userId, + ':now': now, + ':reason': 'Task cancelled by its owner', + }, + }, + }, + { + Put: { + TableName: options.eventsTable!, + Item: { + task_id: task.task_id, + event_id: ulid(), + event_type: 'approval_cancelled', + timestamp: now, + ttl: computeTtlEpoch(options.retentionDays), + metadata: { + request_id: requestId, + status: 'CANCELLED', + reason: 'Task cancelled by its owner', + }, + }, + }, + }, + ], + }), { abortSignal }); + } else { + await ddb.send(new UpdateCommand(update), { abortSignal }); + } + return { task, cancelledAt: now, ...(closeApproval && { cancelledRequestId: requestId }) }; + } catch (error) { + if (!conditionalConflict(error)) throw error; + logger.info('Task cancellation raced with a state change; refreshing', { + event: 'task_cancel_retry', task_id: task.task_id, attempt: attempt + 1, + }); + const fresh = await ddb.send(new GetCommand({ + TableName: options.taskTable, Key: { task_id: task.task_id }, ConsistentRead: true, + }), { abortSignal }); + if (!fresh.Item) throw new TaskCancellationError('missing'); + task = fresh.Item as TaskRecord; + } + } + throw new TaskCancellationError('conflict'); +} diff --git a/cdk/src/handlers/shared/task-concurrency.ts b/cdk/src/handlers/shared/task-concurrency.ts new file mode 100644 index 000000000..fdfcdedcb --- /dev/null +++ b/cdk/src/handlers/shared/task-concurrency.ts @@ -0,0 +1,255 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { randomUUID } from 'node:crypto'; +import { setTimeout as delay } from 'node:timers/promises'; +import { GetCommand, TransactWriteCommand } from '@aws-sdk/lib-dynamodb'; +import { logger } from './logger'; +import { workerLeaseKey } from './microvm-continuation-types'; +import { makeDocClient } from './ua'; +import { ACTIVE_STATUSES, TaskStatus, TERMINAL_STATUSES } from '../../constructs/task-status'; + +/** Internal TaskTable data; a released task cannot acquire another reservation. */ +export interface ReservationTask { + readonly task_id: string; + readonly user_id: string; + readonly status: string; + readonly continuation_launch?: unknown; + readonly microvm_start?: { readonly clientToken: string }; + readonly concurrency_slot?: { + readonly state: 'held' | 'released'; + readonly acquired_at: string; + readonly released_at?: string; + }; +} + +const ddb = makeDocClient(); +const TASK_TABLE = process.env.TASK_TABLE_NAME!; +const COUNTER_TABLE = process.env.USER_CONCURRENCY_TABLE_NAME!; + +function terminal(status: string): boolean { + return TERMINAL_STATUSES.some(value => value === status); +} + +async function readTask(taskId: string, userId: string): Promise { + const result = await ddb.send(new GetCommand({ + TableName: TASK_TABLE, Key: { task_id: taskId }, ConsistentRead: true, + })); + const task = result.Item as ReservationTask | undefined; + if (task && task.user_id !== userId) throw new Error('Concurrency reservation owner does not match task owner'); + return task; +} + +function conditionalFailure(error: unknown, index: number): boolean { + const failure = error as { name?: string; CancellationReasons?: { Code?: string }[] }; + return failure?.name === 'TransactionCanceledException' + && failure.CancellationReasons?.[index]?.Code === 'ConditionalCheckFailed'; +} + +/** + * The SDK does not retry TransactionCanceledException when another task is + * updating the same user's counter. Retry only explicit transaction conflicts; + * conditional failures still belong to the admission/release state machine. + */ +async function sendReservationTransaction( + command: TransactWriteCommand, + taskId: string, + userId: string, +): Promise { + const maxAttempts = 5; + const baseDelayMs = 50; + for (let attempt = 1; attempt <= maxAttempts; attempt++) { + try { + // Keep the exact transaction and client token across retries, including + // the worker-lease condition. A retry cannot weaken the shutdown fence. + await ddb.send(command); + return; + } catch (error) { + const failure = error as { name?: string; CancellationReasons?: { Code?: string }[] }; + const codes = failure?.CancellationReasons?.map(reason => reason.Code); + const conflict = failure?.name === 'TransactionCanceledException' + && codes?.includes('TransactionConflict') + && codes.every(code => code === 'None' || code === 'TransactionConflict'); + if (!conflict) throw error; + if (attempt === maxAttempts) { + logger.warn('Capacity transaction contention exhausted bounded retries', { + task_id: taskId, + user_id: userId, + attempts: attempt, + error_id: 'CONCURRENCY_TRANSACTION_CONFLICT', + cancellation_codes: codes, + }); + throw error; + } + await delay(Math.floor(Math.random() * baseDelayMs * 2 ** (attempt - 1))); + } + } +} + +/** Reserve once per task, including after a lost transaction acknowledgement. */ +export async function acquireTaskSlot(taskId: string, userId: string, limit: number): Promise { + const current = await readTask(taskId, userId); + if (!current) throw new Error(`Cannot reserve capacity for missing task ${taskId}`); + if (current.concurrency_slot?.state === 'held') { + return ACTIVE_STATUSES.some(status => status === current.status); + } + if (current.status !== TaskStatus.SUBMITTED || current.concurrency_slot) return false; + + const now = new Date().toISOString(); + const revision = randomUUID(); + try { + await sendReservationTransaction(new TransactWriteCommand({ + ClientRequestToken: revision, + TransactItems: [ + { + Update: { + TableName: TASK_TABLE, + Key: { task_id: taskId }, + UpdateExpression: 'SET concurrency_slot = :slot', + ConditionExpression: 'user_id = :user AND #status = :submitted AND attribute_not_exists(concurrency_slot)', + ExpressionAttributeNames: { '#status': 'status' }, + ExpressionAttributeValues: { + ':slot': { state: 'held', acquired_at: now }, + ':user': userId, + ':submitted': TaskStatus.SUBMITTED, + }, + }, + }, + { + Update: { + TableName: COUNTER_TABLE, + Key: { user_id: userId }, + UpdateExpression: 'SET active_count = if_not_exists(active_count, :zero) + :one, ' + + 'updated_at = :now, reservation_version = :version', + ConditionExpression: 'attribute_not_exists(active_count) OR active_count < :max', + ExpressionAttributeValues: { + ':zero': 0, ':one': 1, ':max': limit, ':now': now, ':version': revision, + }, + }, + }, + ], + }), taskId, userId); + return true; + } catch (error) { + // A competing invocation or a lost successful response may have reserved it. + const latest = await readTask(taskId, userId); + if (latest?.concurrency_slot?.state === 'held') { + return ACTIVE_STATUSES.some(status => status === latest.status); + } + if (!latest || latest.status !== TaskStatus.SUBMITTED || latest.concurrency_slot) return false; + if (conditionalFailure(error, 1)) return false; // Capacity was full. + throw error; // Throttling/conflict/outage is not evidence that capacity is full. + } +} + +/** + * Return a terminal task's reservation atomically with its counter update. + * Never infer ownership from status alone: legacy/unadmitted tasks have no + * marker and cannot return another task's seat. Reconciliation handles drift. + */ +export async function releaseTaskSlot(taskId: string, userId: string): Promise { + const current = await readTask(taskId, userId); + if (!current || !terminal(current.status) || current.concurrency_slot?.state !== 'held') return false; + + // A failed RunMicrovm response may still have created a worker. The coordinator + // confirms shutdown in its readonly lease before returning that worker's seat. + const workerAttempt = current.continuation_launch && current.microvm_start?.clientToken; + if (workerAttempt) { + const lease = await ddb.send(new GetCommand({ + TableName: TASK_TABLE, Key: workerLeaseKey(taskId), ConsistentRead: true, + })); + if (lease.Item?.lease_user_id !== userId || lease.Item.lease_attempt_id !== workerAttempt + || lease.Item.lease_state !== 'CLOSED') return false; + } + + let emptyCounter = false; + for (let attempt = 0; attempt < 2; attempt++) { + const now = new Date().toISOString(); + const revision = randomUUID(); + try { + await sendReservationTransaction(new TransactWriteCommand({ + ClientRequestToken: revision, + TransactItems: [ + { + Update: { + TableName: TASK_TABLE, + Key: { task_id: taskId }, + UpdateExpression: 'SET concurrency_slot.#state = :released, concurrency_slot.released_at = :now', + ConditionExpression: 'user_id = :user AND concurrency_slot.#state = :held ' + + 'AND #status IN (:completed, :failed, :cancelled, :timedOut)', + ExpressionAttributeNames: { '#state': 'state', '#status': 'status' }, + ExpressionAttributeValues: { + ':user': userId, + ':held': 'held', + ':released': 'released', + ':now': now, + ':completed': TaskStatus.COMPLETED, + ':failed': TaskStatus.FAILED, + ':cancelled': TaskStatus.CANCELLED, + ':timedOut': TaskStatus.TIMED_OUT, + }, + }, + }, + { + Update: { + TableName: COUNTER_TABLE, + Key: { user_id: userId }, + UpdateExpression: emptyCounter + ? 'SET active_count = if_not_exists(active_count, :zero), updated_at = :now, reservation_version = :version' + : 'SET active_count = active_count - :one, updated_at = :now, reservation_version = :version', + ConditionExpression: emptyCounter + ? 'attribute_not_exists(active_count) OR active_count >= :zero' + : 'active_count > :zero', + ExpressionAttributeValues: { + ':zero': 0, ...(!emptyCounter && { ':one': 1 }), ':now': now, ':version': revision, + }, + }, + }, + ...(workerAttempt ? [{ + ConditionCheck: { + TableName: TASK_TABLE, + Key: workerLeaseKey(taskId), + ConditionExpression: 'lease_user_id = :user AND lease_attempt_id = :attempt AND lease_state = :closed', + ExpressionAttributeValues: { ':user': userId, ':attempt': workerAttempt, ':closed': 'CLOSED' }, + }, + }] : []), + ], + }), taskId, userId); + if (emptyCounter) { + logger.warn('Released task reservation whose counter was already empty', { + task_id: taskId, user_id: userId, error_id: 'CONCURRENCY_EMPTY_COUNTER', + }); + } + return true; + } catch (error) { + const latest = await readTask(taskId, userId); + if (latest?.concurrency_slot?.state === 'released') return false; + if (!latest || !terminal(latest.status) || latest.concurrency_slot?.state !== 'held') return false; + if (!emptyCounter && conditionalFailure(error, 1)) { + // This attempt observed no positive count. Close its marker without + // subtracting from seats reserved since then. A concurrent repair may + // leave an overcount, which a later revision-guarded sweep can correct. + emptyCounter = true; + continue; + } + throw error; + } + } + throw new Error(`Could not release capacity reservation for task ${taskId}`); +} diff --git a/cdk/src/handlers/shared/types.ts b/cdk/src/handlers/shared/types.ts index dd0f1a3bb..2faf71e4f 100644 --- a/cdk/src/handlers/shared/types.ts +++ b/cdk/src/handlers/shared/types.ts @@ -19,6 +19,7 @@ import { classifyError, type ErrorClassification } from './error-classifier'; import { logger } from './logger'; +import type { ContinuationLaunchReceipt, ContinuationRecord } from './microvm-continuation-types'; import { coerceNumericOrNull } from './numeric'; import type { ComputeType } from './repo-config'; // Cross-language constants — see ``contracts/constants.md``. Imported at @@ -162,6 +163,9 @@ export type ChannelSource = 'api' | 'webhook' | 'slack' | 'linear' | 'jira'; export interface TaskRecord { readonly task_id: string; readonly user_id: string; + /** Internal immutable input pointer and current durable worker handoff. */ + readonly continuation_launch?: ContinuationLaunchReceipt; + readonly continuation?: ContinuationRecord; /** Cognito group names captured at admission for team-budget rollups. */ readonly team_ids?: readonly string[]; readonly status: TaskStatusType; @@ -285,6 +289,8 @@ export interface TaskRecord { readonly prompt_version?: string; readonly memory_written?: boolean; readonly compute_type?: ComputeType; + /** Approval-wait seconds before MicroVM sleep; 0 keeps it awake. Captured at submission. */ + readonly microvm_sleep_after_s?: number; readonly compute_metadata?: Record; readonly ttl?: number; /** @@ -337,10 +343,9 @@ export interface TaskRecord { readonly attachments?: AttachmentRecord[]; /** * Cedar HITL: per-task default approval timeout (design §10.2). - * Default 300s when absent. The engine clamps to - * ``[APPROVAL_TIMEOUT_S_MIN, APPROVAL_TIMEOUT_S_MAX]`` at task - * start; min-wins against per-rule ``@approval_timeout_s`` at - * gate-firing time. + * Zero (the default) means no automatic expiry. Positive explicit values use + * ``[APPROVAL_TIMEOUT_S_MIN, APPROVAL_TIMEOUT_S_MAX]``; the shortest positive + * task/rule deadline applies when a gate fires. */ readonly approval_timeout_s?: number; /** @@ -446,6 +451,8 @@ export interface TaskNotificationsConfig { * Strips internal fields not exposed in the API. */ export interface TaskDetail { + /** Configured MicroVM approval-wait delay; 0 disables sleep. Absent on legacy records. */ + readonly microvm_sleep_after_s?: number; readonly task_id: string; readonly status: TaskStatusType; /** ``null`` for a repo-less workflow (#248 Phase 3). */ @@ -739,6 +746,8 @@ export interface GetTaskEventsQuery { * Keep in sync with ``cli/src/types.ts``. */ export interface CreateTaskRequest { + /** MicroVM approval-wait seconds before sleep (0 = off, omitted = 600). Does not extend approval deadlines. */ + readonly microvm_sleep_after_s?: number; /** Target repository (``owner/repo``). Optional since #248 Phase 3: a * repo-less workflow (``requires_repo: false``) is submitted without it. * Required-ness is enforced conditionally in ``createTaskCore`` based on @@ -970,6 +979,9 @@ export function toTaskDetail( ): TaskDetail { const ctx = { task_id: record.task_id }; return { + ...(typeof record.microvm_sleep_after_s === 'number' && Number.isInteger(record.microvm_sleep_after_s) + && record.microvm_sleep_after_s >= MICROVM_SLEEP_AFTER_S_MIN && record.microvm_sleep_after_s <= MICROVM_SLEEP_AFTER_S_MAX + && { microvm_sleep_after_s: record.microvm_sleep_after_s }), task_id: record.task_id, status: record.status, repo: record.repo ?? null, @@ -1303,14 +1315,14 @@ export type ApprovalStatus = | 'PENDING' | 'APPROVED' | 'DENIED' + | 'CANCELLED' | 'TIMED_OUT' | 'STRANDED'; /** - * Cedar HITL severity, surfaced in the CLI approval prompt and used - * for severity-gated channel routing (§11.2: high-severity rules - * skip Slack-button auto-approval). Shared alias so the same literal - * union is not redefined inline in `ApprovalRecord`, + * Cedar HITL severity, surfaced in approval prompts and notifications. + * Native Slack approval buttons remain unimplemented. Shared alias so the + * same literal union is not redefined inline in `ApprovalRecord`, * `PendingApprovalSummary`, `PolicyRuleSummary`, etc. */ export type Severity = 'low' | 'medium' | 'high'; @@ -1338,7 +1350,10 @@ interface ApprovalRecordBase { readonly matching_rule_ids: readonly string[]; readonly created_at: string; readonly timeout_s: number; - readonly ttl: number; + /** Optional explicit deadline. Zero timeout has no automatic deadline. */ + readonly deadline_epoch?: number; + /** Retention cleanup is stamped after the owning task closes. */ + readonly ttl?: number; readonly user_id: string; readonly repo: string; } @@ -1362,14 +1377,22 @@ export interface DeniedApprovalRecord extends ApprovalRecordBase { readonly deny_reason?: string; } -/** TIMED_OUT approval row — decided_at required (server-set). */ +/** CANCELLED approval row — its owning task was cancelled or otherwise closed. */ +export interface CancelledApprovalRecord extends ApprovalRecordBase { + readonly status: 'CANCELLED'; + readonly decided_at: string; + readonly cancellation_reason: string; +} + +/** TIMED_OUT approval row — the guest's older timeout writer omits decided_at. */ export interface TimedOutApprovalRecord extends ApprovalRecordBase { readonly status: 'TIMED_OUT'; - readonly decided_at: string; + readonly decided_at?: string; + readonly deny_reason?: string; } -/** STRANDED approval row — decided_at required (set by the - * stranded-task reconciler). No user decision was ever recorded. */ +/** STRANDED approval row, when explicitly recorded. The current stranded-task + * reconciler closes the owning task and cancels its unanswered approval rows. */ export interface StrandedApprovalRecord extends ApprovalRecordBase { readonly status: 'STRANDED'; readonly decided_at: string; @@ -1388,6 +1411,7 @@ export type ApprovalRecord = | PendingApprovalRecord | ApprovedApprovalRecord | DeniedApprovalRecord + | CancelledApprovalRecord | TimedOutApprovalRecord | StrandedApprovalRecord; @@ -1406,8 +1430,8 @@ export interface PendingApprovalSummary { readonly reason: string; readonly created_at: string; readonly timeout_s: number; - /** Derived: `created_at + timeout_s` in ISO 8601 UTC. */ - readonly expires_at: string; + /** Null when the request has no automatic deadline. */ + readonly expires_at: string | null; /** Cedar rule ids that matched this request (design §10.1). Surfaced * so `bgagent pending` can show _why_ a gate fired without the user * spelunking TaskEventsTable. Empty array on pre-Cedar-HITL rows. */ @@ -1521,8 +1545,8 @@ export interface ApprovalDecisionRecordedEvent { * Old callers continue to work — every field is optional. New callers * can pre-approve common scopes (`tool_type:Read`, `bash_pattern:git * status*`) to avoid hitting gates for trusted operations, and can - * raise the per-task default approval timeout above the 300s default - * within the `[30, min(3600, maxLifetime - 300)]` bound. + * choose an explicit approval deadline of 30–3600 seconds. The default, zero, + * leaves unanswered requests available while their owning task remains open. * * Keep in sync with ``cli/src/types.ts``. */ @@ -1541,8 +1565,7 @@ export const INITIAL_APPROVALS_MAX_ENTRY_LENGTH = 128; * Sourced from ``contracts/constants.json`` (S9). */ export const APPROVAL_TIMEOUT_S_MIN = sharedConstants.approval_timeout_s.min; -/** Absolute ceiling for `approval_timeout_s` before the - * `maxLifetime - 300` clip is applied (§7.3). +/** Maximum positive explicit approval timeout. * Sourced from ``contracts/constants.json`` (S9). */ export const APPROVAL_TIMEOUT_S_MAX = sharedConstants.approval_timeout_s.max; @@ -1550,6 +1573,11 @@ export const APPROVAL_TIMEOUT_S_MAX = sharedConstants.approval_timeout_s.max; * Sourced from ``contracts/constants.json`` (S9). */ export const APPROVAL_TIMEOUT_S_DEFAULT = sharedConstants.approval_timeout_s.default; +/** Per-task MicroVM sleep delay bounds; zero disables automatic sleep. */ +export const MICROVM_SLEEP_AFTER_S_MIN = sharedConstants.microvm_sleep_after_s.min; +export const MICROVM_SLEEP_AFTER_S_MAX = sharedConstants.microvm_sleep_after_s.max; +export const MICROVM_SLEEP_AFTER_S_DEFAULT = sharedConstants.microvm_sleep_after_s.default; + /** * Cedar HITL: bounds + platform default for the per-task approval-gate cap * (design decision #13, §4 step 5). Blueprints may override via diff --git a/cdk/src/handlers/slack-interactions.ts b/cdk/src/handlers/slack-interactions.ts index 85a1e552e..9040fa312 100644 --- a/cdk/src/handlers/slack-interactions.ts +++ b/cdk/src/handlers/slack-interactions.ts @@ -17,16 +17,20 @@ * SOFTWARE. */ -import { GetCommand, UpdateCommand } from '@aws-sdk/lib-dynamodb'; +import { GetCommand } from '@aws-sdk/lib-dynamodb'; import type { APIGatewayProxyEvent, APIGatewayProxyResult } from 'aws-lambda'; import { logger } from './shared/logger'; import { getSlackSecret, SLACK_SECRET_PREFIX, verifySlackRequest } from './shared/slack-verify'; +import { cancelTaskState, TaskCancellationError } from './shared/task-cancellation'; +import type { TaskRecord } from './shared/types'; import { makeDocClient } from './shared/ua'; const ddb = makeDocClient(); const SIGNING_SECRET_ARN = process.env.SLACK_SIGNING_SECRET_ARN!; const TASK_TABLE = process.env.TASK_TABLE_NAME!; +const APPROVALS_TABLE = process.env.TASK_APPROVALS_TABLE_NAME; +const EVENTS_TABLE = process.env.TASK_EVENTS_TABLE_NAME; const USER_MAPPING_TABLE = process.env.SLACK_USER_MAPPING_TABLE_NAME!; interface SlackInteractionPayload { @@ -119,6 +123,7 @@ async function handleCancelAction(payload: SlackInteractionPayload, actionId: st const taskResult = await ddb.send(new GetCommand({ TableName: TASK_TABLE, Key: { task_id: taskId }, + ConsistentRead: true, })); if (!taskResult.Item) { @@ -131,32 +136,19 @@ async function handleCancelAction(payload: SlackInteractionPayload, actionId: st return; } - // Attempt to cancel. QUEUED (#441) is cancellable — it removes the - // task from the admission queue (no compute or concurrency to release). - const CANCELLABLE_STATUSES = ['PENDING_UPLOADS', 'QUEUED', 'SUBMITTED', 'HYDRATING', 'RUNNING', 'AWAITING_APPROVAL', 'FINALIZING']; + // Share the REST path's atomic task/approval cancellation and race handling. try { - await ddb.send(new UpdateCommand({ - TableName: TASK_TABLE, - Key: { task_id: taskId }, - UpdateExpression: 'SET #s = :cancelled, updated_at = :now', - ConditionExpression: '#s IN (:s1, :s2, :s3, :s4, :s5, :s6, :s7)', - ExpressionAttributeNames: { '#s': 'status' }, - ExpressionAttributeValues: { - ':cancelled': 'CANCELLED', - ':now': new Date().toISOString(), - ':s1': CANCELLABLE_STATUSES[0], - ':s2': CANCELLABLE_STATUSES[1], - ':s3': CANCELLABLE_STATUSES[2], - ':s4': CANCELLABLE_STATUSES[3], - ':s5': CANCELLABLE_STATUSES[4], - ':s6': CANCELLABLE_STATUSES[5], - ':s7': CANCELLABLE_STATUSES[6], - }, - })); + const cancelled = await cancelTaskState(taskResult.Item as TaskRecord, { + userId: platformUserId, + taskTable: TASK_TABLE, + approvalsTable: APPROVALS_TABLE, + eventsTable: EVENTS_TABLE, + retentionDays: Number(process.env.TASK_RETENTION_DAYS ?? '90'), + }); // Instant feedback: replace the Cancel button message with "Cancelling..." // then clean up all intermediate messages. - const channelMeta = taskResult.Item.channel_metadata as Record | undefined; + const channelMeta = cancelled.task.channel_metadata as Record | undefined; const channelId = payload.channel?.id ?? channelMeta?.slack_channel_id; if (channelMeta && channelId) { const botToken = await getSlackSecret(`${SLACK_SECRET_PREFIX}${teamId}`); @@ -172,8 +164,10 @@ async function handleCancelAction(payload: SlackInteractionPayload, actionId: st } } } catch (err) { - if ((err as Error)?.name === 'ConditionalCheckFailedException') { - await postToResponseUrl(payload.response_url, ':warning: Task is already in a terminal state.'); + if (err instanceof TaskCancellationError) { + await postToResponseUrl(payload.response_url, err.reason === 'conflict' + ? ':warning: Task state changed. Please retry cancellation.' + : ':warning: Task is no longer available for cancellation.'); } else { throw err; } diff --git a/cdk/src/handlers/slack-notify.ts b/cdk/src/handlers/slack-notify.ts index 4b391ac70..61242f83c 100644 --- a/cdk/src/handlers/slack-notify.ts +++ b/cdk/src/handlers/slack-notify.ts @@ -51,8 +51,9 @@ import { type DynamoDBDocumentClient, GetCommand, UpdateCommand } from '@aws-sdk // so importing the type back creates no runtime cycle. ``import type`` // is erased after compile, so the bundler sees a one-way dep. import type { FanOutEvent } from './fanout-task-events'; +import { APPROVAL_NOTIFICATION_EVENTS, isApprovalNotification, loadApprovalNotification, markApprovalNotificationDelivered } from './shared/approval-notifications'; import { logger } from './shared/logger'; -import { renderSlackBlocks } from './shared/slack-blocks'; +import { renderSlackBlocks, type SlackMessage } from './shared/slack-blocks'; import { getSlackSecret, SLACK_SECRET_PREFIX } from './shared/slack-verify'; import type { TaskRecord } from './shared/types'; @@ -97,8 +98,8 @@ const SLACK_DEDUP_ATTRIBUTE: Record = { * Slack entries in ``CHANNEL_DEFAULTS`` (see fanout-task-events.ts) — * drift means the router subscribes Slack to events that the * dispatcher silently ignores, which lies in batch telemetry - * (issue #64 review Cat 7). Forward-compat ``approval_required`` and - * ``status_response`` are deliberately absent until their emitters + * (issue #64 review Cat 7). Forward-compat ``status_response`` is + * deliberately absent until its emitter * ship; until then they fall through and are dropped at this gate. * ``pr_created`` is intentionally omitted from Slack — the * ``task_completed`` block already carries the View PR button, so a @@ -114,6 +115,7 @@ export const NOTIFIABLE_EVENTS = new Set([ 'task_timed_out', 'task_stranded', 'agent_error', + ...APPROVAL_NOTIFICATION_EVENTS, ]); /** @@ -229,10 +231,15 @@ export async function dispatchSlackEvent( const taskResult = await ddb.send(new GetCommand({ TableName: tableName, Key: { task_id: taskId }, + ConsistentRead: true, })); const task = taskResult.Item as TaskRecord | undefined; if (!task || task.channel_source !== 'slack') return; + const approval = isApprovalNotification(eventType) + ? await loadApprovalNotification(ddb, task, eventType, event.metadata ?? {}, 'slack') : null; + if (isApprovalNotification(eventType) && !approval) return; + // Dedup any event that should only ever post once per task even // under partial-batch retry (terminals, agent_error). The orchestrator // can also write multiple events of the same kind (retries, @@ -290,7 +297,10 @@ export async function dispatchSlackEvent( // ``metadata: { S: ... }`` shape itself from the raw stream record. const eventMetadata = event.metadata; - const message = renderSlackBlocks(eventType, task, eventMetadata); + const message: SlackMessage = approval ? { + text: `${approval.title} for task ${taskId}`, + blocks: [{ type: 'section', text: { type: 'plain_text', text: approval.text } }], + } : renderSlackBlocks(eventType, task, eventMetadata); const threadTs = channelMeta.slack_thread_ts; @@ -348,6 +358,8 @@ export async function dispatchSlackEvent( throw new SlackApiError(failureMessage); } + if (approval) await markApprovalNotificationDelivered(ddb, approval); + // Reactions always use the real channel id even for DMs. const reactionChannel = channelMeta.slack_channel_id; const reactionTarget = threadTs ?? result.ts; diff --git a/cdk/src/main.ts b/cdk/src/main.ts index 7637e5fea..c3371ae05 100644 --- a/cdk/src/main.ts +++ b/cdk/src/main.ts @@ -58,7 +58,8 @@ export interface BuildAppOptions { * Async because AgentCore-supported availability zones are resolved from the * account's zone mapping at synth time (live `DescribeAvailabilityZones` + * `sts:GetCallerIdentity`) when a concrete account/region is bound. Env-agnostic - * synth and the validated context override never touch AWS. + * synth does not call AWS; explicit overrides in supported regions are checked + * against the account's EC2 zone mapping. */ export async function buildApp(options: BuildAppOptions = {}): Promise { const app = new App(options.appProps); diff --git a/cdk/src/stacks/agent.ts b/cdk/src/stacks/agent.ts index b580ff2f1..fecac6c58 100644 --- a/cdk/src/stacks/agent.ts +++ b/cdk/src/stacks/agent.ts @@ -19,7 +19,7 @@ import * as path from 'path'; import * as bedrock from '@aws-cdk/aws-bedrock-alpha'; -import { ArnFormat, AspectPriority, Aspects, Stack, StackProps, RemovalPolicy, CfnOutput, CfnResource, Duration, Fn, Lazy } from 'aws-cdk-lib'; +import { ArnFormat, AspectPriority, Aspects, Stack, StackProps, NestedStack, RemovalPolicy, CfnOutput, CfnResource, Duration, Fn, Lazy } from 'aws-cdk-lib'; import * as agentcore from 'aws-cdk-lib/aws-bedrockagentcore'; import * as ec2 from 'aws-cdk-lib/aws-ec2'; import * as ecr_assets from 'aws-cdk-lib/aws-ecr-assets'; @@ -27,7 +27,7 @@ import * as iam from 'aws-cdk-lib/aws-iam'; import * as logs from 'aws-cdk-lib/aws-logs'; import * as secretsmanager from 'aws-cdk-lib/aws-secretsmanager'; import * as cr from 'aws-cdk-lib/custom-resources'; -import { NagSuppressions } from 'cdk-nag'; +import { NagSuppressions, type NagPackSuppression } from 'cdk-nag'; import { Construct, IConstruct } from 'constructs'; import { AdmissionQueuePickup } from '../constructs/admission-queue-pickup'; import { AgentMemory } from '../constructs/agent-memory'; @@ -35,6 +35,7 @@ import { AgentSessionRole } from '../constructs/agent-session-role'; import { AgentVpc } from '../constructs/agent-vpc'; import { ApiKeyTable } from '../constructs/api-key-table'; import { ApprovalMetricsPublisherConsumer } from '../constructs/approval-metrics-publisher-consumer'; +import { ApprovalRequestService } from '../constructs/approval-request-service'; import { AttachmentsBucket } from '../constructs/attachments-bucket'; import { PLATFORM_DEFAULT_AUX_MODEL_ID, @@ -48,6 +49,7 @@ import { BudgetAlerts } from '../constructs/budget-alerts'; import { BudgetTable } from '../constructs/budget-table'; import { CedarWasmLayer } from '../constructs/cedar-wasm-layer'; import { ConcurrencyReconciler } from '../constructs/concurrency-reconciler'; +import { ContinuationBucket } from '../constructs/continuation-bucket'; import { DnsFirewall } from '../constructs/dns-firewall'; import { EcsAgentCluster, resolveEcsTaskSizing } from '../constructs/ecs-agent-cluster'; import { EcsPayloadBucket } from '../constructs/ecs-payload-bucket'; @@ -57,12 +59,15 @@ import { IterationHeartbeat } from '../constructs/iteration-heartbeat'; import { JiraIntegration } from '../constructs/jira-integration'; import { LambdaMicrovmCompute, + createMicrovmExecutionRole, isLambdaMicrovmImageConfigured, type LambdaMicrovmImageInputs, } from '../constructs/lambda-microvm-compute'; +import { LambdaMicrovmStack } from '../constructs/lambda-microvm-stack'; import { LinearIdentityVault } from '../constructs/linear-identity-vault'; import { LinearIntegration } from '../constructs/linear-integration'; import { LinearVaultConsentPageStack } from '../constructs/linear-vault-consent-page'; +import { MicrovmContinuationManager } from '../constructs/microvm-continuation-manager'; import { OperationalAlerts } from '../constructs/operational-alerts'; import { OrchestrationReconciler } from '../constructs/orchestration-reconciler'; import { OrchestrationTable } from '../constructs/orchestration-table'; @@ -348,6 +353,35 @@ export class AgentStack extends Stack { // MicroVM termination grant (ADR-021 sub-decision 4). const computeType = this.node.tryGetContext('compute_type') ?? 'agentcore'; const lambdaMicrovmEnabled = computeType === 'lambda-microvm'; + const microvmNestedContext = this.node.tryGetContext('microvm_nested_stack'); + if (microvmNestedContext !== undefined && ![true, false, 'true', 'false'].includes(microvmNestedContext)) { + throw new Error('microvm_nested_stack must be true or false'); + } + // Require an explicit layout until flat-to-nested migration is verified. + // A synth-time AWS lookup cannot protect deployments of saved assemblies. + const microvmNested = microvmNestedContext !== false && microvmNestedContext !== 'false'; + if (lambdaMicrovmEnabled && microvmNestedContext === undefined) { + throw new Error( + 'microvm_nested_stack must be explicitly selected: use --context microvm_nested_stack=true ' + + 'for new or already-nested deployments, or --context microvm_nested_stack=false for existing flat deployments. ' + + 'Changing an existing flat deployment to nested can erase artifacts and pending payloads; ' + + 'setting true does not perform a migration. See docs/verification/645-p3-nested-stack.md.', + ); + } + const microvmResourceNamePrefix = this.node.tryGetContext('microvm_resource_name_prefix'); + if (microvmResourceNamePrefix !== undefined) { + if (typeof microvmResourceNamePrefix !== 'string') { + throw new Error('microvm_resource_name_prefix must be a string'); + } + if (!microvmNested) { + throw new Error('microvm_resource_name_prefix cannot be used with microvm_nested_stack=false'); + } + } + const suspendContext = this.node.tryGetContext('microvm_approval_suspend_enabled'); + if (suspendContext !== undefined && ![true, false, 'true', 'false'].includes(suspendContext)) { + throw new Error('microvm_approval_suspend_enabled must be true or false'); + } + const microvmApprovalSuspendEnabled = suspendContext === true || suspendContext === 'true'; // --- Tool-federation Gateway deploy gate (ADR-019 P1) --- // Whether to provision the AgentCore Gateway that federates the agent's MCP @@ -371,20 +405,6 @@ export class AgentStack extends Stack { // second chance to disagree. const linearVaultWorkload = linearVaultWorkloadName(this); - // Fail here, naming both flags, rather than 500 resources later. The two features - // together synthesize 505 resources against CloudFormation's hard 500 limit (MicroVM - // alone 496, the vault alone 488), so the combination is not deployable today. Left to - // the resource counter, the operator gets a per-type census and no hint that two - // context flags are the cause. - if (linearIdentityVaultEnabled && computeType === 'lambda-microvm') { - throw new Error( - 'enableLinearIdentityVault cannot be combined with compute_type=lambda-microvm: the two ' - + 'together exceed CloudFormation\'s 500-resource limit for this stack (505). Deploy the ' - + 'vault on the agentcore or ecs substrate, or omit enableLinearIdentityVault. See ' - + 'docs/design/ADR-016 and the LINEAR_SETUP_GUIDE.', - ); - } - // The operator-supplied MicroVM image inputs, resolved HERE (pure context // reads, no construct dependency) rather than at the construct's call site // below, because TaskApi — created well before the MicroVM construct — needs @@ -402,6 +422,8 @@ export class AgentStack extends Stack { const microvmImageInputs: LambdaMicrovmImageInputs = { baseImageArn: this.node.tryGetContext('microvm_base_image_arn'), baseImageVersion: this.node.tryGetContext('microvm_base_image_version'), + artifactSha256: this.node.tryGetContext('microvm_artifact_sha256'), + managedImageVersion: this.node.tryGetContext('microvm_managed_image_version'), externalImageIdentifier: this.node.tryGetContext('microvm_image_identifier'), externalImageVersion: this.node.tryGetContext('microvm_image_version'), }; @@ -409,7 +431,7 @@ export class AgentStack extends Stack { && isLambdaMicrovmImageConfigured(microvmImageInputs); // MicroVM image ARN placeholder — the image is created AFTER TaskApi, but the - // cancel Lambda's grant must be scoped to it. Same Lazy.string cycle-break as + // cancel and decision-handler grants must be scoped to it. Same Lazy.string cycle-break as // the runtime / orchestrator / SessionRole ARNs below. let microvmImageArnHolder: string | undefined; const lazyMicrovmImageArn = Lazy.string({ @@ -474,6 +496,13 @@ export class AgentStack extends Stack { }); inputGuardrail.createVersion('Initial version'); + // A retained durable Lambda version also pins GUARDRAIL_VERSION. Preserve + // that immutable dependency across updates, even when CDK creates a new one. + for (const child of inputGuardrail.node.findAll()) { + if (child instanceof CfnResource && child.cfnResourceType === 'AWS::Bedrock::GuardrailVersion') { + child.applyRemovalPolicy(RemovalPolicy.RETAIN); + } + } // --- TaskApi is constructed before the orchestrator (which it needs the // ARN of) and before the Runtime (which it needs the ARN of, for the @@ -540,6 +569,13 @@ export class AgentStack extends Stack { ...(microvmImageConfigured && { lambdaMicrovmImageArn: lazyMicrovmImageArn }), }); + const approvalRequests = new ApprovalRequestService(this, 'ApprovalRequests', { + taskTable: taskTable.table, approvalsTable: taskApprovalsTable.table, + }); + // Reuse the parent's regional API Gateway logging configuration. + const apiLoggingAccount = taskApi.api.node.tryFindChild('Account'); + if (apiLoggingAccount) approvalRequests.node.addDependency(apiLoggingAccount); + // Agent asset registry API (#246) in its own NestedStack + RestApi so its // ~35 resources don't count against this root stack's 500-resource limit. // It authorizes against the SHARED Cognito user pool, so a caller's JWT works @@ -613,6 +649,7 @@ export class AgentStack extends Stack { // AWAITING_APPROVAL; absent → hook fails closed with // ``approval_write_failed`` (the `ApprovalTablesUnavailable` path). TASK_APPROVALS_TABLE_NAME: taskApprovalsTable.table.tableName, + APPROVAL_REQUESTS_API_URL: approvalRequests.api.url, // Hint for the hook's remaining-maxLifetime calculation (§6.5 // pseudocode line 793). Kept in sync with the AgentCore // lifecycle configuration below so drift is visible. 8 hours. @@ -749,12 +786,9 @@ export class AgentStack extends Stack { // {user_id, repo, task_id}, and that role carries the tenant-data grants // constrained by aws:PrincipalTag conditions. The runtime role keeps only // non-tenant / shared access: - // - UserConcurrencyTable: user-scoped counter (agent path does not write - // it today; left here for the reconciler/orchestrator parity). // - GitHub PAT secret: read once at startup, before the agent assumes the // SessionRole. // - CloudWatch Logs + AgentCore Memory: shared/non-tenant. - userConcurrencyTable.table.grantReadWriteData(runtime); githubTokenSecret.grantRead(runtime); applicationLogGroup.grantWrite(runtime); agentMemory.grantReadWrite(runtime); @@ -822,17 +856,19 @@ export class AgentStack extends Stack { // --- Per-task SessionRole --- // Holds the tenant-data grants (the four task_id-partitioned tables, plus // per-user-prefixed trace writes and attachment reads), each constrained - // by aws:PrincipalTag conditions so a compromised session reaches only its - // own task's data. The agent assumes this with refreshable credentials + // by aws:PrincipalTag conditions for the existing session credentials. + // The compute role chooses the tags; that choice is not independently + // authenticated by this trust policy. Main-task writes are restricted to + // reporting attributes. The agent assumes this with refreshable credentials // (1h role-chaining cap, tasks run to 8h). Trust admits the runtime // ExecutionRole as the assuming principal; the ECS task role is added in // the ECS block below when that backend is enabled. const agentSessionRole = new AgentSessionRole(this, 'AgentSessionRole', { assumingRoles: [runtime.role], + taskTable: taskTable.table, + approvalsTable: taskApprovalsTable.table, taskScopedTables: [ - taskTable.table, taskEventsTable.table, - taskApprovalsTable.table, taskNudgesTable.table, ], traceArtifactsBucket: traceArtifactsBucket.bucket, @@ -843,6 +879,7 @@ export class AgentStack extends Stack { invokableModels: invokableBedrockModels, }); sessionRoleArnHolder = agentSessionRole.role.roleArn; + approvalRequests.grantRequests(agentSessionRole.role); // X-Ray tracing disabled — requires account-level UpdateTraceSegmentDestination // which needs CloudWatch Logs resource policy propagation. Re-enable via @@ -855,24 +892,15 @@ export class AgentStack extends Stack { }, ], true); - // Chunk 10 deploy-prep: the Cedar HITL additions (TaskApprovalsTable - // grant + extra env vars) pushed the runtime - // execution role past CDK's per-inline-policy size limit, causing CDK - // to auto-split excess statements into ``OverflowPolicy1`` / etc. - // Those overflow policies inherit the same wildcard - // ``bedrock:InvokeModel*`` / CloudWatch / cross-region-inference - // actions as the base policy but live at paths that any suppression - // placed at constructor time does NOT reach (CDK creates the - // overflow policies lazily during synth ``prepare()``, after the - // construct tree has been frozen). Use an Aspect that visits every - // node during synth and matches overflow-policy children of the - // runtime ExecutionRole so any present or future overflow is - // suppressed automatically without hardcoding - // ``OverflowPolicy`` indices. - // Roles known to overflow, with the evidence for each. Keyed by a path - // fragment rather than an `OverflowPolicy` index so future splits are - // covered automatically. - const OVERFLOW_SUPPRESSIONS: readonly { readonly pathFragment: string; readonly reason: string }[] = [ + // CDK splits large role policies during synth, after constructor-time + // suppressions have visited the existing children. Apply the documented + // exceptions to those later policies before cdk-nag inspects them, matching + // the owning role without relying on a particular OverflowPolicy index. + const OVERFLOW_SUPPRESSIONS: readonly { + readonly pathFragment: string; + readonly reason: string; + readonly appliesTo?: NagPackSuppression['appliesTo']; + }[] = [ { pathFragment: '/Runtime/ExecutionRole/OverflowPolicy', reason: @@ -888,14 +916,41 @@ export class AgentStack extends Stack { reason: 'CDK-generated overflow policy on the Linear webhook processor role carries the kms:GenerateDataKey* that SNS Topic.grantPublish emits for the CMK-encrypted operational-alerts topic. Scoped to that single topic key; the wildcard only spans the GenerateDataKey/GenerateDataKeyWithoutPlaintext pair.', }, + { + // Image lifecycle grants can push existing channel secret grants into + // an overflow policy. Exempt only those resource patterns; other wildcard + // grants in this role's future overflow documents still require review. + pathFragment: '/TaskOrchestrator/OrchestratorFn/ServiceRole/OverflowPolicy', + reason: + 'Channel setup creates workspace-specific OAuth secrets and Linear vault providers after deployment. These grants retain the configured account and Region and match only the Linear/Jira secret prefixes and Linear vault provider/client-secret prefixes.', + appliesTo: [{ + regex: '/^Resource::arn:.*:secretsmanager:.*:secret:bgagent-jira-oauth-\\*$/', + }, { + regex: '/^Resource::arn:.*:secretsmanager:.*:secret:bgagent-linear-oauth-\\*$/', + }, { + regex: '/^Resource::arn:.*:bedrock-agentcore:.*:token-vault/default/oauth2credentialprovider/bgagent-linear-oauth-\\*$/', + }, { + regex: '/^Resource::arn:.*:secretsmanager:.*:secret:bedrock-agentcore-identity!default/oauth2/bgagent-linear-oauth-\\*$/', + }], + }, + { + pathFragment: '/OrchestrationReconciler/ReconcilerFn/ServiceRole/OverflowPolicy', + reason: + 'The reconciler mints Linear feedback tokens for workspace providers created during channel setup. The grant matches only Linear vault provider/client-secret prefixes in this account and Region.', + appliesTo: [{ + regex: '/^Resource::arn:.*:bedrock-agentcore:.*:token-vault/default/oauth2credentialprovider/bgagent-linear-oauth-\\*$/', + }, { + regex: '/^Resource::arn:.*:secretsmanager:.*:secret:bedrock-agentcore-identity!default/oauth2/bgagent-linear-oauth-\\*$/', + }], + }, ]; const overflowSuppressionAspect = { visit(node: IConstruct) { const nodePath = node.node.path; if (!nodePath.endsWith('/Resource')) return; - for (const { pathFragment, reason } of OVERFLOW_SUPPRESSIONS) { + for (const { pathFragment, reason, appliesTo } of OVERFLOW_SUPPRESSIONS) { if (nodePath.includes(pathFragment)) { - NagSuppressions.addResourceSuppressions(node, [{ id: 'AwsSolutions-IAM5', reason }]); + NagSuppressions.addResourceSuppressions(node, [{ id: 'AwsSolutions-IAM5', reason, appliesTo }]); return; } } @@ -997,10 +1052,10 @@ export class AgentStack extends Stack { // agentcore-only, matching how other optional constructs are context-gated. // (``computeType`` is read near the top of the constructor — TaskApi needs it // for the conditional MicroVM cancel grant.) - // Ephemeral bucket for ECS task payloads — the orchestrator writes the - // payload here (it exceeds the 8 KB RunTask containerOverrides limit) and - // passes only an S3 URI pointer; the container fetches it on boot, the - // orchestrator deletes it at finalize. Only synthesized under the ecs gate. + // ECS v2 bootstrap storage: deployment manifests, task instructions and + // private launch references. Workers read manifests with IAM and download + // their task through a signed one-object URL. Finalize deletes task objects; + // the one-day lifecycle reaps leftovers. Synthesized only under the ECS gate. const ecsPayloadBucket = computeType === 'ecs' ? new EcsPayloadBucket(this, 'EcsPayloadBucket') : undefined; @@ -1008,7 +1063,7 @@ export class AgentStack extends Stack { NagSuppressions.addResourceSuppressions(ecsPayloadBucket.bucket, [ { id: 'AwsSolutions-S1', - reason: 'Ephemeral per-task payloads with a 1-day TTL; writes confined to the orchestrator IAM role by grantPut, reads to the ECS task role by grantRead, both scoped to this bucket. Object deleted at finalize. Object-level audit intentionally omitted — CloudTrail data events / a log bucket are not justified for transient boot payloads.', + reason: 'Ephemeral bootstrap storage with a 1-day TTL. The coordinator writes manifests and task objects, signs task reads and deletes task objects at finalize. Workers read only bootstrap/* with IAM and use single-object signed URLs for payloads; other object reads and bucket listing are explicitly denied. Object-level audit intentionally omitted — CloudTrail data events / a log bucket are not justified for transient boot payloads.', }, ]); } @@ -1092,6 +1147,8 @@ export class AgentStack extends Stack { }), taskTable: taskTable.table, taskEventsTable: taskEventsTable.table, + taskApprovalsTable: taskApprovalsTable.table, + approvalRequestsApiUrl: approvalRequests.api.url, userConcurrencyTable: userConcurrencyTable.table, githubTokenSecret, memoryId: agentMemory.memory.memoryId, @@ -1101,7 +1158,7 @@ export class AgentStack extends Stack { // without this grant. The AgentCore runtime gets the equivalent grant // where it is created above. agentMemory, - // Read-only grant so the container can fetch its payload from S3. + // The task role reads bootstrap manifests; a signed URL delivers its payload. payloadBucket: ecsPayloadBucket!.bucket, // ECS parity: the same bucket the runtime uses for ARTIFACTS_BUCKET_NAME — // a repo-bound artifact workflow delivers here. Wires the @@ -1132,40 +1189,57 @@ export class AgentStack extends Stack { // agentcore-only. The construct itself enforces the ADR's Region gate, so a // deploy into a Region without Lambda MicroVMs fails at synth rather than on // the first task. - const lambdaMicrovm = lambdaMicrovmEnabled - ? new LambdaMicrovmCompute(this, 'LambdaMicrovmCompute', { - vpc: agentVpc.vpc, - // Per-session IAM scoping (#209): the MicroVM execution role is admitted - // to the same per-task SessionRole the AgentCore runtime and the Fargate - // task role use, so tenant-data access is tag-scoped on every substrate. - agentSessionRole, - // ADR-021 P2 runtime parity on the MicroVM execution role. Same two props - // EcsAgentCluster takes, for the same reasons: the PAT is read at startup - // before the SessionRole is assumed, and MEMORY_ID (already delivered in - // agent_payload) makes the agent ATTEMPT a memory write that fails closed - // without the grant. The remaining parity grants (channel OAuth, Bedrock, - // AZ describe) need no stack input and are wired inside the construct. - githubTokenSecret, - agentMemory, - // ADR-021 P2-F4: the SAME log group whose name travels to the guest in - // `agentPlatformConfig.logGroupName` below (→ `LOG_GROUP_NAME`). P2 - // delivered the name without the grant, so the agent's structured per-task - // lines and its METRICS_REPORT were AccessDenied on - // logs:CreateLogStream and the platform's canonical observability streams - // were empty on this backend. Passing the construct (not the name) keeps the - // grant and the delivered value derived from one object. - applicationLogGroup, - // Resolved above TaskApi — see `microvmImageInputs`. - ...microvmImageInputs, - }) - : undefined; + const microvmProps = { + vpc: agentVpc.vpc, + // Per-session IAM scoping (#209): the MicroVM execution role is admitted + // to the same per-task SessionRole the AgentCore runtime and the Fargate + // task role use, so tenant-data access is tag-scoped on every substrate. + agentSessionRole, + // ADR-021 P2 runtime parity on the MicroVM execution role. Same two props + // EcsAgentCluster takes, for the same reasons: the PAT is read at startup + // before the SessionRole is assumed, and MEMORY_ID (already delivered in + // agent_payload) makes the agent ATTEMPT a memory write that fails closed + // without the grant. The remaining parity grants (channel OAuth, Bedrock, + // AZ describe) need no stack input and are wired inside the construct. + githubTokenSecret, + agentMemory, + // ADR-021 P2-F4: the SAME log group whose name travels to the guest in + // `agentPlatformConfig.logGroupName` below (→ `LOG_GROUP_NAME`). P2 + // delivered the name without the grant, so the agent's structured per-task + // lines and its METRICS_REPORT were AccessDenied on + // logs:CreateLogStream and the platform's canonical observability streams + // were empty on this backend. Passing the construct (not the name) keeps the + // grant and the delivered value derived from one object. + applicationLogGroup, + // Resolved above TaskApi — see `microvmImageInputs`. + ...microvmImageInputs, + }; + let lambdaMicrovm: LambdaMicrovmCompute | undefined; + if (lambdaMicrovmEnabled) { + if (microvmNested) { + // Preserve the execution role's original parent path and therefore its + // logical ID. Moving it with the image creates a SessionRole trust cycle. + const roleScope = new Construct(this, 'LambdaMicrovmCompute'); + const executionRole = createMicrovmExecutionRole(roleScope, 'ExecutionRole'); + lambdaMicrovm = new LambdaMicrovmStack(this, 'Microvm', { + ...microvmProps, + deploymentName: this.stackName, + resourceNamePrefix: microvmResourceNamePrefix, + executionRole, + }).compute; + } else { + lambdaMicrovm = new LambdaMicrovmCompute(this, 'LambdaMicrovmCompute', microvmProps); + } + } - // Resolve the Lazy TaskApi's cancel grant is scoped by. The invariant the + // Resolve the image ARN used by TaskApi's cancel and wake grants. The invariant the // Lazy's `produce` guards: `microvmImageConfigured` (computed from the same // inputs, via the same predicate) is true exactly when the construct sets // `imageArn`, so a configured deployment always has an ARN to resolve and an // unconfigured one never asks for it. microvmImageArnHolder = lambdaMicrovm?.imageArn; + const continuationBucket = lambdaMicrovm ? new ContinuationBucket(this, 'ContinuationBucket') : undefined; + continuationBucket?.grantWorker(agentSessionRole.role); // Advertise which compute substrate this deploy actually provisioned, so the // CLI can refuse to onboard a repo as ``compute_type: ecs`` when the ECS gate @@ -1210,6 +1284,10 @@ export class AgentStack extends Stack { value: lambdaMicrovm.artifactObjectKey, description: 'S3 key the Lambda MicroVMs artifact must be uploaded to (matches the build role\'s s3:GetObject scope)', }); + new CfnOutput(this, 'MicrovmArtifactBaseObjectKey', { + value: lambdaMicrovm.artifactBaseObjectKey, + description: 'Base artifact key; managed packaging adds the ZIP SHA-256, manual builds use this key', + }); new CfnOutput(this, 'MicrovmBuildRoleArn', { value: lambdaMicrovm.buildRole.roleArn, description: 'IAM role for `aws lambda-microvms create-microvm-image --build-role-arn`', @@ -1268,6 +1346,7 @@ export class AgentStack extends Stack { // of the resources they identify. agentPlatformConfig: { taskApprovalsTableName: taskApprovalsTable.table.tableName, + approvalRequestsApiUrl: approvalRequests.api.url, nudgesTableName: taskNudgesTable.table.tableName, logGroupName: applicationLogGroup.logGroupName, // INTENTIONAL, not a wiring bug: both keys resolve to the SAME bucket @@ -1289,29 +1368,11 @@ export class AgentStack extends Stack { // Same helper, same resolved geography as the AgentCore runtime env // above (#764) — the two substrates cannot be told to call different // inference profiles. - // `inferenceProfileId(geo, AUX)` rather than main's `haikuInferenceProfileId(geo)`: - // that helper was removed on this branch when the two duplicate haiku paths were - // collapsed into one, so main's call site no longer resolves. anthropicDefaultHaikuModel: inferenceProfileId(bedrockGeoRegion, PLATFORM_DEFAULT_AUX_MODEL_ID), - // The MAIN model, delivered the same way for the same reason. Only the auxiliary - // one was, so on this substrate the main model came from a literal in - // agent/src/config.py that a geography change does not touch: a non-default - // `bedrockGeoRegion` granted one geography while the agent asked for another, - // and every task with no per-repo override failed at turn 0 with AccessDenied. anthropicModel: inferenceProfileId(bedrockGeoRegion, PLATFORM_DEFAULT_MODEL_ID), - // Substrate parity for the Identity vault: the AgentCore runtime gets these - // as env and the ECS container via EcsAgentCluster, so a MicroVM guest must - // receive them too or its agent skips vault minting and falls back to a - // Secrets-Manager token a vault-managed workspace does not have — losing - // reactions and state transitions on work that otherwise succeeds. Forwarded - // as platform_config because a snapshot must not bake configuration in. - ...(linearIdentityVault - ? { - linearVaultEnabled: 'true', - linearWorkloadIdentityName: linearVaultWorkload, - } - : {}), }, + ...(continuationBucket && { continuationBucket }), + taskApprovalsTable: taskApprovalsTable.table, // Route ``compute_type: 'ecs'`` repos to the Fargate cluster above — // only when the cluster was synthesized (deploy --context compute_type=ecs). ...(ecsCluster && { @@ -1344,6 +1405,8 @@ export class AgentStack extends Stack { imageIdentifier: lambdaMicrovm.imageIdentifier, imageArn: lambdaMicrovm.imageArn, imageVersion: lambdaMicrovm.imageVersion, + approvalsTable: taskApprovalsTable.table, + approvalSuspendEnabled: microvmApprovalSuspendEnabled, executionRoleArn: lambdaMicrovm.executionRole.roleArn, egressConnectorArns: lambdaMicrovm.egressConnectorArns, // Explicit NO_INGRESS, not an omission: RunMicrovm attaches a PUBLIC @@ -1358,13 +1421,31 @@ export class AgentStack extends Stack { // Now that the orchestrator exists, resolve the Lazy used by TaskApi at synth. orchestratorArnHolder = orchestrator.alias.functionArn; + // Stateless scheduled jobs share a nested stack to leave room in both + // flat and nested MicroVM layouts. Upgrades replace their generated-name + // functions/schedules; task tables, buckets and compute resources stay put. + const concurrencyMaintenance = new NestedStack(this, 'ConcurrencyMaintenance'); + if (continuationBucket && lambdaMicrovm?.imageArn) { + taskApi.enableMicrovmContinuations( + continuationBucket.bucket.bucketName, orchestrator.fn.functionArn, userConcurrencyTable.table, + maxConcurrentTasksPerUser, + ); + new MicrovmContinuationManager(concurrencyMaintenance, 'MicrovmContinuationManager', { + taskTable: taskTable.table, + approvalsTable: taskApprovalsTable.table, + userConcurrencyTable: userConcurrencyTable.table, + continuationBucket, + orchestratorFunctionArn: orchestrator.fn.functionArn, + imageArn: lambdaMicrovm.imageArn, + }); + } // Grant the orchestrator Lambda read+write access to memory // (reads during context hydration, writes for fallback episodes) agentMemory.grantReadWrite(orchestrator.fn); // --- Concurrency counter reconciler (drift correction) --- - new ConcurrencyReconciler(this, 'ConcurrencyReconciler', { + new ConcurrencyReconciler(concurrencyMaintenance, 'ConcurrencyReconciler', { taskTable: taskTable.table, userConcurrencyTable: userConcurrencyTable.table, }); @@ -1374,7 +1455,7 @@ export class AgentStack extends Stack { // concurrency cap is hit) in FIFO order as slots free up: flips // QUEUED -> SUBMITTED and re-invokes the orchestrator, whose atomic // admissionControl remains the single writer of the counter. - new AdmissionQueuePickup(this, 'AdmissionQueuePickup', { + new AdmissionQueuePickup(concurrencyMaintenance, 'AdmissionQueuePickup', { taskTable: taskTable.table, taskEventsTable: taskEventsTable.table, userConcurrencyTable: userConcurrencyTable.table, @@ -1386,9 +1467,10 @@ export class AgentStack extends Stack { // (orchestrator Lambda crash between TaskTable write and InvokeAgentRuntime, // container crash during startup, etc.). Transitions to FAILED with a // `task_stranded` event. - new StrandedTaskReconciler(this, 'StrandedTaskReconciler', { + new StrandedTaskReconciler(concurrencyMaintenance, 'StrandedTaskReconciler', { taskTable: taskTable.table, taskEventsTable: taskEventsTable.table, + taskApprovalsTable: taskApprovalsTable.table, userConcurrencyTable: userConcurrencyTable.table, }); @@ -1396,7 +1478,7 @@ export class AgentStack extends Stack { // Auto-cancels PENDING_UPLOADS tasks that were never confirmed within // 30 minutes (client crash, abandoned session, network failure). // Cleans up orphaned S3 objects under the task's attachment prefix. - new PendingUploadCleanup(this, 'PendingUploadCleanup', { + new PendingUploadCleanup(concurrencyMaintenance, 'PendingUploadCleanup', { taskTable: taskTable.table, taskEventsTable: taskEventsTable.table, attachmentsBucket: attachmentsBucket.bucket, @@ -1428,6 +1510,7 @@ export class AgentStack extends Stack { userPool: taskApi.userPool, taskTable: taskTable.table, taskEventsTable: taskEventsTable.table, + taskApprovalsTable: taskApprovalsTable.table, budgetTable: budgetTable.table, repoTable: repoTable.table, orchestratorFunctionArn: orchestrator.alias.functionArn, @@ -1521,6 +1604,9 @@ export class AgentStack extends Stack { userPool: taskApi.userPool, taskTable: taskTable.table, taskEventsTable: taskEventsTable.table, + taskApprovalsTable: taskApprovalsTable.table, + lambdaMicrovmImageArn: lambdaMicrovm?.imageArn, + continuationBucketName: continuationBucket?.bucket.bucketName, budgetTable: budgetTable.table, repoTable: repoTable.table, // Enables the webhook processor's orchestration path @@ -1926,6 +2012,7 @@ export class AgentStack extends Stack { const fanOutConsumer = new FanOutConsumer(this, 'FanOutConsumer', { taskEventsTable: taskEventsTable.table, taskTable: taskTable.table, + taskApprovalsTable: taskApprovalsTable.table, repoTable: repoTable.table, githubTokenSecret, // Slack bot-token grant is guarded on this prop — pass the @@ -1988,24 +2075,12 @@ export class AgentStack extends Stack { taskTable: taskTable.table, }); - // --- Vault parity for every Lambda that talks to Linear ----------------- - // The webhook processor was granted vault access when the vault landed and - // nothing else was, on the assumption that widening could wait. It could not: - // these three post the PR-opened comment, the terminal comment, the epic - // rollup, and the GitHub-side issue updates. Without the grant each falls back - // to a Secrets-Manager token that a vault-onboarded workspace does not - // maintain, so the task succeeds and the Linear issue shows nothing after the - // opening comment — no reaction, no state change, no PR link. Live-caught as - // 401s in the fan-out log while the task itself completed and opened its PR. + // Every Linear writer needs the vault grant; otherwise its feedback may fail + // while the coding task succeeds. The webhook processor is wired by + // LinearIntegration. The coordinator also forwards these identifiers to + // MicroVM guests in authenticated platform_config. if (linearIdentityVault) { - // The full set, derived from the transitive import graph rather than from the - // handlers that import a Linear module DIRECTLY — which is how the reconciler - // and the heartbeat were missed: both reach a minting resolver through - // orchestration-channel-factory, two hops away. A source-level test now - // recomputes this set and fails if a handler joins it without being wired. - // - // Deliberately NOT here: the webhook RECEIVER (verifies signatures, never - // mints) and the stranded reconciler (reaches no channel that mints). + // A source-graph test includes indirect callers and guards this inventory. for (const linearWriter of [ fanOutConsumer.fn, orchestrator.fn, @@ -2241,8 +2316,8 @@ interface PinnedLogResource { * constructed — the hash in each one is not reproducible from the construct path * alone, which is precisely why they have to be written down. * - * An entry stays until its stack is gone. Removing one while the stack still - * exists re-introduces the rename and the failed update that comes with it. + * Removing an entry requires checking the deployed ids and completing any + * necessary log-delivery migration; otherwise the rename can fail the update. */ const PINNED_LOG_DELIVERY_BY_STACK: Record = { 'backgroundagent-dev': [ @@ -2279,7 +2354,7 @@ const PINNED_LOG_DELIVERY_BY_STACK: Record }; /** - * Pin the auto-created log-delivery resources to stable logical ids, ALWAYS. + * Apply captured legacy log-delivery ids for entries in the pin table. * * These resources are created for us by the AgentCore Runtime and named after * whatever construct path the library uses internally, so a library-side rename @@ -2290,15 +2365,12 @@ const PINNED_LOG_DELIVERY_BY_STACK: Record * update rolls the whole stack back. Owning the ids ourselves decouples us from * the library's internal naming. * - * Applied unconditionally rather than behind a flag. Three cases, all safe: - * - * - An existing stack in the account that owns these resources: the ids match - * what CloudFormation already recorded, so it updates them in place. This is - * the case that was broken. - * - A fresh stack or account: nothing owns these names yet, so they create - * normally. The ids are ours rather than the library's, which is the point; - * the values themselves carry no meaning beyond being stable. - * - Any other name: the ids embed the stack name, so each stack gets its own. + * Known limitation (#703): stack name does not identify deployment history. + * An existing same-named stack that already uses the library's current ids can + * collide when these pins rename its resources. Fresh stacks have no existing + * resources to collide with. Other stack names keep the library's naming. + * Removing the pins also requires migration for stacks still using these ids; + * the account-agnostic fix and migration are tracked in PR #705. * * The values were read off a stack deployed before the rename. Do not "tidy" * them — they are a record of what CloudFormation already has, and editing one @@ -2307,9 +2379,7 @@ const PINNED_LOG_DELIVERY_BY_STACK: Record function pinLogDeliveryLogicalIds(runtime: agentcore.Runtime): void { const stack = Stack.of(runtime); const pins = PINNED_LOG_DELIVERY_BY_STACK[stack.stackName]; - // Only the stack these ids were recorded from can use them: they embed that - // stack's name. Any other stack keeps the library's own naming, which is - // correct for it — it has no pre-rename resources to line up with. + // This lookup checks only the name, not the account or deployed ids (#703). if (!pins) return; for (const pin of pins) { diff --git a/cdk/test/bootstrap/__snapshots__/version.test.ts.snap b/cdk/test/bootstrap/__snapshots__/version.test.ts.snap index d4fe819d0..cf7c721d9 100644 --- a/cdk/test/bootstrap/__snapshots__/version.test.ts.snap +++ b/cdk/test/bootstrap/__snapshots__/version.test.ts.snap @@ -1,3 +1,3 @@ // Jest Snapshot v1, https://jestjs.io/docs/snapshot-testing -exports[`bootstrap version module hash is stable 1`] = `"d30eb8e63c8f6fd5551de03a72acb342231175aaf417bd98811d6a07b0be3afc"`; +exports[`bootstrap version module hash is stable 1`] = `"dc6301b65558c8fb51200e7b712a754e1b85adda7a512b14ec4f647dc954f541"`; diff --git a/cdk/test/bootstrap/bootstrap-template.test.ts b/cdk/test/bootstrap/bootstrap-template.test.ts index d4f8dc9d7..5be0d5790 100644 --- a/cdk/test/bootstrap/bootstrap-template.test.ts +++ b/cdk/test/bootstrap/bootstrap-template.test.ts @@ -132,6 +132,8 @@ describe('Bootstrap template', () => { expect(passRole!.Resource).toEqual([ 'arn:aws:iam::*:role/backgroundagent-dev-LambdaMicrovmComputeBuild*', 'arn:aws:iam::*:role/backgroundagent-dev-LambdaMicrovmComputeConnector*', + 'arn:aws:iam::*:role/backgroundagent-dev-MicrovmBuildRole', + 'arn:aws:iam::*:role/backgroundagent-dev-MicrovmConnectorRole', ]); }); @@ -154,6 +156,25 @@ describe('Bootstrap template', () => { }); describe('CloudFormationExecutionRole', () => { + it('can pass only itself to CloudFormation for nested stacks', () => { + const role = template.Resources.CloudFormationExecutionRole.Properties; + const policy = role.Policies?.find( + (item: any) => item.PolicyName === 'PassExecutionRoleToCloudFormation', + ); + expect(policy?.PolicyDocument).toEqual({ + Version: '2012-10-17', + Statement: [{ + Sid: 'PassSelfToCloudFormation', + Effect: 'Allow', + Action: 'iam:PassRole', + Resource: { + 'Fn::Sub': `arn:\${AWS::Partition}:iam::\${AWS::AccountId}:role/${role.RoleName['Fn::Sub']}`, + }, + Condition: { StringEquals: { 'iam:PassedToService': 'cloudformation.amazonaws.com' } }, + }], + }); + }); + it('exists and is an IAM Role', () => { expect(template.Resources.CloudFormationExecutionRole).toBeDefined(); expect(template.Resources.CloudFormationExecutionRole.Type).toBe('AWS::IAM::Role'); diff --git a/cdk/test/bootstrap/policies.test.ts b/cdk/test/bootstrap/policies.test.ts index 3ea8184b0..ceafbcd2b 100644 --- a/cdk/test/bootstrap/policies.test.ts +++ b/cdk/test/bootstrap/policies.test.ts @@ -500,7 +500,7 @@ describe('computeLambdaMicrovmPolicy', () => { const resolvedDoc = stack.resolve(doc); const statements = resolvedDoc.Statement as Array<{ Sid: string }>; - expect(statements.map((s) => s.Sid)).toEqual(['LambdaMicrovms', 'MicrovmPassRoles']); + expect(statements.map((s) => s.Sid)).toEqual(['LambdaMicrovms', 'MicrovmPassRoles', 'MicrovmSuspendConfiguration']); }); it('covers the expected service prefixes', () => { @@ -511,8 +511,16 @@ describe('computeLambdaMicrovmPolicy', () => { ); const prefixes = new Set(allActions.map((a) => a.split(':')[0])); - // `iam` joins `lambda` as of the MicrovmPassRoles statement (ADR-021 P2r2-F9). - expect(prefixes).toEqual(new Set(['lambda', 'iam'])); + expect(prefixes).toEqual(new Set(['lambda', 'iam', 'ssm'])); + }); + + it('limits parameter lifecycle and tagging to the ABCA MicroVM suspension switch', () => { + const statement = stack.resolve(doc).Statement.find((s: { Sid: string }) => s.Sid === 'MicrovmSuspendConfiguration'); + expect(statement.Resource).toBe('arn:aws:ssm:*:*:parameter/backgroundagent-*/microvm-approval-suspend-enabled'); + expect(statement.Action).toEqual([ + 'ssm:GetParameters', 'ssm:PutParameter', 'ssm:DeleteParameter', + 'ssm:AddTagsToResource', 'ssm:RemoveTagsFromResource', 'ssm:ListTagsForResource', + ]); }); describe('MicrovmPassRoles (ADR-021 P2r2-F9)', () => { @@ -548,6 +556,8 @@ describe('computeLambdaMicrovmPolicy', () => { expect(resources).toEqual([ 'arn:aws:iam::*:role/backgroundagent-dev-LambdaMicrovmComputeBuild*', 'arn:aws:iam::*:role/backgroundagent-dev-LambdaMicrovmComputeConnector*', + 'arn:aws:iam::*:role/backgroundagent-dev-MicrovmBuildRole', + 'arn:aws:iam::*:role/backgroundagent-dev-MicrovmConnectorRole', ]); // NOT the stack-wide role prefix the conditioned statement uses: an // unconditioned pass on `backgroundagent-dev-*` would drop the @@ -593,6 +603,22 @@ describe('computeLambdaMicrovmPolicy', () => { expect(matched).toBe(true); } }); + + it('limits nested PassRole to the two explicit build and operator names', () => { + const resources = passRoleStatement().Resource as string[]; + const matches = (name: string) => resources.some(pattern => + new RegExp(`^${pattern.replace(/[.+?^${}()|[\]\\]/g, '\\$&').replace(/\*/g, '.*')}$`) + .test(`arn:aws:iam::123456789012:role/${name}`)); + expect(matches('backgroundagent-dev-MicrovmBuildRole')).toBe(true); + expect(matches('backgroundagent-dev-MicrovmConnectorRole')).toBe(true); + for (const name of [ + 'backgroundagent-dev-MicrovmExecutionRole', + 'backgroundagent-dev-MicrovmBuildRoleOther', + 'backgroundagent-dev-OtherBuildRole', + ]) { + expect(matches(name)).toBe(false); + } + }); }); it('covers both CFN resource types the construct synthesizes', () => { diff --git a/cdk/test/bootstrap/synth-coverage.test.ts b/cdk/test/bootstrap/synth-coverage.test.ts index 5e3e08b66..8c1c760f6 100644 --- a/cdk/test/bootstrap/synth-coverage.test.ts +++ b/cdk/test/bootstrap/synth-coverage.test.ts @@ -34,6 +34,7 @@ describe('Bootstrap policy synth coverage', () => { let template: Template; let registryTemplate: Template; let allowedActions: Set; + let allTemplates: Template[]; beforeAll(() => { const app = new App(); @@ -41,6 +42,9 @@ describe('Bootstrap policy synth coverage', () => { env: { account: '123456789012', region: 'us-east-1' }, }); template = Template.fromStack(stack); + allTemplates = [template, ...stack.node.findAll() + .filter((node): node is NestedStack => node instanceof NestedStack) + .map(child => Template.fromStack(child))]; const registryStack = stack.node.tryFindChild('AgentRegistryStack') as AgentRegistryStack | undefined; if (!registryStack) { @@ -64,8 +68,8 @@ describe('Bootstrap policy synth coverage', () => { }); it('maps every synthesized CFN type (that needs IAM) to bootstrap actions', () => { - const resources = template.toJSON().Resources as Record; - const typesInTemplate = new Set(Object.values(resources).map((r) => r.Type)); + const typesInTemplate = new Set(allTemplates.flatMap(child => + Object.values(child.toJSON().Resources as Record).map(resource => resource.Type))); const unmapped: string[] = []; const missingByType: Record = {}; @@ -88,6 +92,33 @@ describe('Bootstrap policy synth coverage', () => { expect(missingByType).toEqual({}); }); + it.each([false, true])('covers managed MicroVM and nested resources (microvm_nested_stack=%s)', microvmNested => { + const app = new App({ + context: { + compute_type: 'lambda-microvm', + microvm_nested_stack: microvmNested, + microvm_base_image_arn: 'arn:aws:lambda:us-east-1:aws:microvm-image:al2023-1', + microvm_base_image_version: '1', + microvm_artifact_sha256: 'a'.repeat(64), + }, + }); + const stack = new AgentStack(app, 'backgroundagent-microvm-coverage', { + env: { account: '123456789012', region: 'us-east-1' }, + }); + const root = Template.fromStack(stack); + const children = stack.node.findAll().filter((node): node is NestedStack => node instanceof NestedStack); + expect(children.length).toBeGreaterThan(0); + const types = new Set([root, ...children.map(child => Template.fromStack(child))].flatMap(child => + Object.values(child.toJSON().Resources as Record).map(resource => resource.Type))); + expect(types).toContain('AWS::SSM::Parameter'); + expect(types).toContain('AWS::Lambda::MicrovmImage'); + for (const type of types) { + if (CFN_TYPES_WITHOUT_EXEC_ROLE_IAM.has(type)) continue; + expect({ type, mapped: type in RESOURCE_ACTION_MAP }).toEqual({ type, mapped: true }); + expect({ type, missing: findMissingBootstrapActions(type, allowedActions) }).toEqual({ type, missing: [] }); + } + }); + it('maps the context-gated tool-gateway CFN types (ADR-019, not in default synth)', () => { // The default synth above never instantiates the ToolGateway construct, so // its two CFN types would slip past the coverage loop. Synthesize the gated diff --git a/cdk/test/bootstrap/version.test.ts b/cdk/test/bootstrap/version.test.ts index 1c1bce5cb..d7995fd42 100644 --- a/cdk/test/bootstrap/version.test.ts +++ b/cdk/test/bootstrap/version.test.ts @@ -17,9 +17,63 @@ * SOFTWARE. */ +import { PolicyDocument } from 'aws-cdk-lib/aws-iam'; +import * as nestedStackPolicy from '../../src/bootstrap/nested-stack-policy'; +import * as bootstrapPolicies from '../../src/bootstrap/policies'; import { BOOTSTRAP_VERSION, computeBootstrapHash } from '../../src/bootstrap/version'; describe('bootstrap version module', () => { + afterEach(() => { + jest.restoreAllMocks(); + }); + + function hashDocument(document: Record): string { + const policy = new PolicyDocument(); + jest.spyOn(policy, 'toJSON').mockReturnValue(document); + jest.spyOn(bootstrapPolicies, 'allPolicies').mockReturnValue([policy]); + return computeBootstrapHash(); + } + + const statement = { + Effect: 'Allow', + Action: 'iam:PassRole', + Resource: 'arn:aws:iam::123456789012:role/example', + Condition: { StringEquals: { 'iam:PassedToService': 'cloudformation.amazonaws.com' } }, + }; + + it.each([ + { Action: 'iam:CreateRole' }, + { Effect: 'Deny' }, + { Resource: 'arn:aws:iam::123456789012:role/other' }, + { Condition: { StringEquals: { 'iam:PassedToService': 'lambda.amazonaws.com' } } }, + ])('hash changes when nested permissions change: %j', (change) => { + const before = hashDocument({ Version: '2012-10-17', Statement: [statement] }); + const after = hashDocument({ Version: '2012-10-17', Statement: [{ ...statement, ...change }] }); + expect(after).not.toBe(before); + }); + + it('ignores object key order at every depth', () => { + const before = hashDocument({ Version: '2012-10-17', Statement: [statement] }); + const after = hashDocument({ + Statement: [{ + Condition: statement.Condition, + Resource: statement.Resource, + Action: statement.Action, + Effect: statement.Effect, + }], + Version: '2012-10-17', + }); + expect(after).toBe(before); + }); + + it('includes the generated inline execution policy in the hash', () => { + const before = computeBootstrapHash(); + const policy = nestedStackPolicy.nestedStackExecutionPolicy(); + policy.Statement[0].Action = 'iam:GetRole'; + jest.spyOn(nestedStackPolicy, 'nestedStackExecutionPolicy').mockReturnValue(policy); + expect(computeBootstrapHash()).not.toBe(before); + }); + it('BOOTSTRAP_VERSION matches semver format', () => { expect(BOOTSTRAP_VERSION).toMatch(/^\d+\.\d+\.\d+$/); }); diff --git a/cdk/test/constructs/agent-session-role.test.ts b/cdk/test/constructs/agent-session-role.test.ts index 676b34f65..a84e86aa9 100644 --- a/cdk/test/constructs/agent-session-role.test.ts +++ b/cdk/test/constructs/agent-session-role.test.ts @@ -24,6 +24,7 @@ import * as dynamodb from 'aws-cdk-lib/aws-dynamodb'; import * as iam from 'aws-cdk-lib/aws-iam'; import * as s3 from 'aws-cdk-lib/aws-s3'; import { AgentSessionRole } from '../../src/constructs/agent-session-role'; +import taskWriteAttributes from '../../src/constructs/agent-task-write-attributes.json'; function createStack() { const app = new App(); @@ -51,10 +52,10 @@ function createStack() { const sessionRole = new AgentSessionRole(stack, 'AgentSessionRole', { assumingRoles: [computeRole], + taskTable, + approvalsTable: taskApprovalsTable, taskScopedTables: [ - taskTable, taskEventsTable, - taskApprovalsTable, taskNudgesTable, ], traceArtifactsBucket, @@ -107,15 +108,18 @@ describe('AgentSessionRole construct', () => { const actions = Array.isArray(s.Action) ? s.Action : [s.Action]; return actions.some((a: string) => a.startsWith('dynamodb:')); }); - // Four task-scoped tables → four conditioned statements. - expect(ddbStatements).toHaveLength(4); + // Main task read/update, approval read, and two supporting table grants. + expect(ddbStatements).toHaveLength(5); for (const s of ddbStatements) { - expect(s.Condition).toEqual({ - 'ForAllValues:StringEquals': { - 'dynamodb:LeadingKeys': ['${aws:PrincipalTag/task_id}'], - }, - }); const actions = Array.isArray(s.Action) ? s.Action : [s.Action]; + const readonlyTask = JSON.stringify(s.Resource).includes('TaskTable') + && !JSON.stringify(s.Resource).includes('Approvals') && actions.includes('dynamodb:GetItem'); + expect(s.Condition['ForAllValues:StringEquals']['dynamodb:LeadingKeys']) + .toEqual(readonlyTask + ? ['${aws:PrincipalTag/task_id}', 'worker-lease#${aws:PrincipalTag/task_id}'] + : ['${aws:PrincipalTag/task_id}']); + expect(s.Condition.Null['dynamodb:LeadingKeys']).toBe('false'); + if (readonlyTask) expect(actions).not.toContain('dynamodb:UpdateItem'); // Scan must NOT be granted — it ignores leading-keys. expect(actions).not.toContain('dynamodb:Scan'); } @@ -140,6 +144,65 @@ describe('AgentSessionRole construct', () => { expect(tracePut).toBeDefined(); }); + test('workers cannot forge decisions or notification flags through any approval-table write', () => { + const statements = Object.values(template.findResources('AWS::IAM::Policy')) + .flatMap(policy => policy.Properties.PolicyDocument.Statement) + .filter((statement: { Resource: unknown }) => JSON.stringify(statement.Resource).includes('TaskApprovalsTable')); + expect(statements).toHaveLength(1); + expect(statements[0].Action).toEqual([ + 'dynamodb:GetItem', 'dynamodb:BatchGetItem', 'dynamodb:Query', 'dynamodb:ConditionCheckItem', + ]); + expect(statements[0].Condition['ForAllValues:StringEquals']['dynamodb:LeadingKeys']) + .toEqual(['${aws:PrincipalTag/task_id}']); + }); + + test('task records cannot be replaced, deleted or updated without an attribute allowlist', () => { + const policy = Object.entries(template.findResources('AWS::IAM::Policy')) + .find(([id]) => id.includes('AgentSessionRole'))![1]; + const statements = policy.Properties.PolicyDocument.Statement.filter( + (s: { Resource: unknown }) => JSON.stringify(s.Resource).includes('TaskTable'), + ); + expect(statements).toHaveLength(2); + for (const s of statements) { + const actions: string[] = Array.isArray(s.Action) ? s.Action : [s.Action]; + expect(actions.every((action) => [ + 'dynamodb:GetItem', 'dynamodb:BatchGetItem', 'dynamodb:Query', + 'dynamodb:ConditionCheckItem', 'dynamodb:UpdateItem', + ].includes(action))).toBe(true); + if (actions.includes('dynamodb:UpdateItem')) { + const attrs = s.Condition?.['ForAllValues:StringEquals']?.['dynamodb:Attributes']; + expect(attrs).toEqual(taskWriteAttributes); + expect(s.Condition.Null['dynamodb:Attributes']).toBe('false'); + for (const protectedAttribute of [ + 'microvm_start', 'microvm_lifecycle', 'microvm_sleep_after_s', 'concurrency_slot', 'user_id', 'created_at', + 'session_id', 'compute_type', 'compute_metadata', 'agent_runtime_arn', + ]) { + expect(attrs).not.toContain(protectedAttribute); + } + } + } + }); + + test('rejects granting unrestricted supporting-table access to the main task table', () => { + const stack = new Stack(new App(), 'DuplicateTable'); + const table = new dynamodb.Table(stack, 'Tasks', { + partitionKey: { name: 'task_id', type: dynamodb.AttributeType.STRING }, + }); + const computeRole = new iam.Role(stack, 'Compute', { + assumedBy: new iam.ServicePrincipal('ecs-tasks.amazonaws.com'), + }); + expect(() => new AgentSessionRole(stack, 'Session', { + assumingRoles: [computeRole], + taskTable: table, + approvalsTable: new dynamodb.Table(stack, 'ApprovalReadTable', { + partitionKey: { name: 'task_id', type: dynamodb.AttributeType.STRING }, + }), + taskScopedTables: [table], + traceArtifactsBucket: new s3.Bucket(stack, 'Traces'), + attachmentsBucket: new s3.Bucket(stack, 'Attachments'), + })).toThrow('taskTable and approvalsTable must not appear in taskScopedTables'); + }); + test('S3 artifact writes are scoped to the per-task_id prefix (#248 Phase 3)', () => { const policies = template.findResources('AWS::IAM::Policy'); const sessionPolicy = Object.entries(policies).find(([id]) => @@ -208,7 +271,11 @@ describe('AgentSessionRole construct', () => { }); new AgentSessionRole(stack, 'SR', { assumingRoles: [computeRole], - taskScopedTables: [table], + taskTable: table, + approvalsTable: new dynamodb.Table(stack, 'ApprovalReadTable', { + partitionKey: { name: 'task_id', type: dynamodb.AttributeType.STRING }, + }), + taskScopedTables: [], traceArtifactsBucket: new s3.Bucket(stack, 'TB'), attachmentsBucket: new s3.Bucket(stack, 'AB'), invokableModels: [model], @@ -231,8 +298,7 @@ describe('AgentSessionRole construct', () => { }); test('omitting invokableModels grants no bedrock action (isolated tests)', () => { - const { template: t } = createStack(); - const policies = t.findResources('AWS::IAM::Policy'); + const policies = template.findResources('AWS::IAM::Policy'); const sessionPolicy = Object.entries(policies).find(([id]) => id.includes('AgentSessionRole'), )![1]; @@ -254,7 +320,11 @@ describe('AgentSessionRole construct', () => { }); const sessionRole = new AgentSessionRole(stack, 'SR', { assumingRoles: [agentcoreRole], - taskScopedTables: [table], + taskTable: table, + approvalsTable: new dynamodb.Table(stack, 'ApprovalReadTable', { + partitionKey: { name: 'task_id', type: dynamodb.AttributeType.STRING }, + }), + taskScopedTables: [], traceArtifactsBucket: new s3.Bucket(stack, 'TB'), attachmentsBucket: new s3.Bucket(stack, 'AB'), }); diff --git a/cdk/test/constructs/approval-request-service.test.ts b/cdk/test/constructs/approval-request-service.test.ts new file mode 100644 index 000000000..9a89bf32e --- /dev/null +++ b/cdk/test/constructs/approval-request-service.test.ts @@ -0,0 +1,66 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { App, Stack } from 'aws-cdk-lib'; +import { Template } from 'aws-cdk-lib/assertions'; +import * as dynamodb from 'aws-cdk-lib/aws-dynamodb'; +import * as iam from 'aws-cdk-lib/aws-iam'; +import { ApprovalRequestService } from '../../src/constructs/approval-request-service'; + +let root: Template; +let child: Template; +beforeAll(() => { + const stack = new Stack(new App(), 'ApprovalFixture', { + env: { account: '123456789012', region: 'us-east-1' }, + }); + const table = (id: string) => new dynamodb.Table(stack, id, { + partitionKey: { name: 'task_id', type: dynamodb.AttributeType.STRING }, + }); + const service = new ApprovalRequestService(stack, 'ApprovalRequests', { + taskTable: table('Tasks'), approvalsTable: table('Approvals'), + }); + const worker = new iam.Role(stack, 'WorkerSession', { + assumedBy: new iam.ServicePrincipal('ecs-tasks.amazonaws.com'), + }); + service.grantRequests(worker); + root = Template.fromStack(stack); + child = Template.fromStack(service); +}); + +test('requires IAM authentication on the only writable endpoint', () => { + child.resourceCountIs('AWS::ApiGateway::Method', 1); + child.hasResourceProperties('AWS::ApiGateway::Method', { + HttpMethod: 'POST', AuthorizationType: 'AWS_IAM', + }); +}); + +test('binds invocation to the session task tag and grants no direct Lambda invocation', () => { + const statements = Object.values(root.findResources('AWS::IAM::Policy')) + .flatMap(policy => policy.Properties.PolicyDocument.Statement); + expect(statements).toHaveLength(1); + expect(statements[0].Action).toBe('execute-api:Invoke'); + expect(JSON.stringify(statements[0].Resource)).toContain('/v1/POST/tasks/${aws:PrincipalTag/task_id}'); + expect(statements[0].Condition).toEqual({ Null: { 'aws:PrincipalTag/task_id': 'false' } }); +}); + +test('keeps the trusted writer and API infrastructure out of the parent resource budget', () => { + root.resourceCountIs('AWS::Lambda::Function', 0); + child.resourceCountIs('AWS::Lambda::Function', 1); + child.resourceCountIs('AWS::DynamoDB::Table', 0); +}); diff --git a/cdk/test/constructs/concurrency-reconciler.test.ts b/cdk/test/constructs/concurrency-reconciler.test.ts index 34e1e88e4..b8e7403db 100644 --- a/cdk/test/constructs/concurrency-reconciler.test.ts +++ b/cdk/test/constructs/concurrency-reconciler.test.ts @@ -43,8 +43,10 @@ function createStack(): Template { } describe('ConcurrencyReconciler construct', () => { + let template: Template; + beforeAll(() => { template = createStack(); }); + test('creates a Lambda function', () => { - const template = createStack(); template.hasResourceProperties('AWS::Lambda::Function', { Runtime: 'nodejs24.x', Timeout: 300, @@ -52,14 +54,12 @@ describe('ConcurrencyReconciler construct', () => { }); test('creates an EventBridge rule with rate schedule', () => { - const template = createStack(); template.hasResourceProperties('AWS::Events::Rule', { ScheduleExpression: 'rate(15 minutes)', }); }); test('Lambda has correct environment variables', () => { - const template = createStack(); template.hasResourceProperties('AWS::Lambda::Function', { Environment: { Variables: Match.objectLike({ @@ -69,4 +69,15 @@ describe('ConcurrencyReconciler construct', () => { }, }); }); + + test('can read reservations and update their markers without deleting or creating tasks', () => { + const tableId = Object.keys(template.findResources('AWS::DynamoDB::Table')).find(id => id.startsWith('TaskTable'))!; + const statements = Object.values(template.findResources('AWS::IAM::Policy')) + .flatMap(policy => policy.Properties.PolicyDocument.Statement) + .filter(statement => JSON.stringify(statement.Resource).includes(tableId)); + const actions = statements.flatMap(statement => [statement.Action].flat()); + expect(actions).toEqual(expect.arrayContaining(['dynamodb:GetItem', 'dynamodb:Scan', 'dynamodb:UpdateItem'])); + expect(actions).not.toEqual(expect.arrayContaining(['dynamodb:PutItem'])); + expect(actions).not.toEqual(expect.arrayContaining(['dynamodb:DeleteItem'])); + }); }); diff --git a/cdk/test/constructs/continuation-bucket.test.ts b/cdk/test/constructs/continuation-bucket.test.ts new file mode 100644 index 000000000..2a99f19f2 --- /dev/null +++ b/cdk/test/constructs/continuation-bucket.test.ts @@ -0,0 +1,56 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { App, Stack } from 'aws-cdk-lib'; +import { Template, Match } from 'aws-cdk-lib/assertions'; +import * as iam from 'aws-cdk-lib/aws-iam'; +import { ContinuationBucket } from '../../src/constructs/continuation-bucket'; + +test('active version-pinned checkpoints survive time and stack removal; workers cannot list/delete', () => { + const stack = new Stack(new App(), 'Test'); + const storage = new ContinuationBucket(stack, 'Continuation'); + const worker = new iam.Role(stack, 'Worker', { assumedBy: new iam.ServicePrincipal('lambda.amazonaws.com') }); + storage.grantWorker(worker); + const template = Template.fromStack(stack); + template.hasResource('AWS::S3::Bucket', { + DeletionPolicy: 'Retain', + Properties: { + VersioningConfiguration: { Status: 'Enabled' }, + LifecycleConfiguration: { + Rules: [{ + Id: 'abandoned-multipart-uploads', + Status: 'Enabled', + AbortIncompleteMultipartUpload: { DaysAfterInitiation: 1 }, + }], + }, + }, + }); + template.hasResourceProperties('AWS::IAM::Policy', { + PolicyDocument: { + Statement: [Match.objectLike({ + Action: ['s3:GetObject', 's3:GetObjectVersion', 's3:PutObject'], + Resource: Match.anyValue(), + })], + }, + }); + const policies = JSON.stringify(template.findResources('AWS::IAM::Policy')); + expect(policies).toContain('continuations/${aws:PrincipalTag/task_id}/*'); + expect(policies).not.toContain('s3:List'); + expect(policies).not.toContain('s3:Delete'); +}); diff --git a/cdk/test/constructs/ecs-agent-cluster.test.ts b/cdk/test/constructs/ecs-agent-cluster.test.ts index 3c5a95688..77b285131 100644 --- a/cdk/test/constructs/ecs-agent-cluster.test.ts +++ b/cdk/test/constructs/ecs-agent-cluster.test.ts @@ -38,6 +38,7 @@ function createStack(overrides?: { bedrockGeoRegion?: string; withMemory?: boolean; withLinearVault?: boolean; + withApprovals?: boolean; taskSizing?: { buildTaskCpu?: number; buildTaskMemoryMiB?: number; @@ -73,6 +74,12 @@ function createStack(overrides?: { const userConcurrencyTable = new dynamodb.Table(stack, 'UserConcurrencyTable', { partitionKey: { name: 'user_id', type: dynamodb.AttributeType.STRING }, }); + const taskApprovalsTable = overrides?.withApprovals + ? new dynamodb.Table(stack, 'TaskApprovalsTable', { + partitionKey: { name: 'task_id', type: dynamodb.AttributeType.STRING }, + sortKey: { name: 'request_id', type: dynamodb.AttributeType.STRING }, + }) + : undefined; const githubTokenSecret = new secretsmanager.Secret(stack, 'GitHubTokenSecret'); @@ -90,6 +97,7 @@ function createStack(overrides?: { agentImageAsset, taskTable, taskEventsTable, + taskApprovalsTable, userConcurrencyTable, githubTokenSecret, memoryId: overrides?.memoryId, @@ -596,6 +604,32 @@ describe('EcsAgentCluster construct', () => { }); }); + test('legacy direct access cannot replace task records, edit coordinator fields or access capacity', () => { + const policies = Object.entries(baseTemplate.findResources('AWS::IAM::Policy')) + .filter(([id]) => id.includes('TaskRole')); + expect(policies).toHaveLength(1); + const statements = policies[0][1].Properties.PolicyDocument.Statement; + expect(JSON.stringify(statements)).not.toContain('UserConcurrencyTable'); + const taskStatements = statements.filter( + (s: { Resource: unknown }) => JSON.stringify(s.Resource).includes('TaskTable'), + ); + expect(taskStatements).toHaveLength(2); + for (const s of taskStatements) { + const actions = Array.isArray(s.Action) ? s.Action : [s.Action]; + expect(actions.every((a: string) => [ + 'dynamodb:GetItem', 'dynamodb:BatchGetItem', 'dynamodb:Query', + 'dynamodb:ConditionCheckItem', 'dynamodb:UpdateItem', + ].includes(a))).toBe(true); + if (actions.includes('dynamodb:UpdateItem')) { + const attrs = s.Condition['ForAllValues:StringEquals']['dynamodb:Attributes']; + expect(attrs).toContain('agent_heartbeat_at'); + expect(attrs).not.toContain('microvm_start'); + expect(attrs).not.toContain('concurrency_slot'); + expect(s.Condition.Null['dynamodb:Attributes']).toBe('false'); + } + } + }); + test('build def caps build parallelism to prevent OOM (K14 / ABCA-691)', () => { // The build task def serializes the mise DAG (MISE_JOBS=1) and pins the jest // fleet (JEST_MAX_WORKERS=4) so the cross-package build storm can't OOM the @@ -710,6 +744,7 @@ describe('EcsAgentCluster construct', () => { }); const taskTable = mk('TaskTable'); const taskEventsTable = mk('TaskEventsTable'); + const taskApprovalsTable = mk('TaskApprovalsTable'); const userConcurrencyTable = new dynamodb.Table(stack, 'UserConcurrencyTable', { partitionKey: { name: 'user_id', type: dynamodb.AttributeType.STRING }, }); @@ -720,7 +755,9 @@ describe('EcsAgentCluster construct', () => { assumedBy: new iam.ServicePrincipal('bedrock-agentcore.amazonaws.com'), }), ], - taskScopedTables: [taskTable, taskEventsTable], + taskTable, + approvalsTable: taskApprovalsTable, + taskScopedTables: [taskEventsTable], traceArtifactsBucket: new s3.Bucket(stack, 'TraceBucket'), attachmentsBucket: new s3.Bucket(stack, 'AttachmentsBucket'), }); @@ -730,15 +767,20 @@ describe('EcsAgentCluster construct', () => { agentImageAsset, taskTable, taskEventsTable, + taskApprovalsTable, userConcurrencyTable, githubTokenSecret, agentSessionRole: sessionRole, + approvalRequestsApiUrl: 'https://approval.execute-api.us-east-1.amazonaws.com/v1', }); return Template.fromStack(stack); } + let sessionTemplate: Template; + beforeAll(() => { sessionTemplate = createWithSessionRole(); }); + test('injects AGENT_SESSION_ROLE_ARN into the container', () => { - createWithSessionRole().hasResourceProperties('AWS::ECS::TaskDefinition', { + sessionTemplate.hasResourceProperties('AWS::ECS::TaskDefinition', { ContainerDefinitions: Match.arrayWith([ Match.objectLike({ Environment: Match.arrayWith([ @@ -750,7 +792,7 @@ describe('EcsAgentCluster construct', () => { }); test('task role gets sts:AssumeRole on the SessionRole, not direct task-table DDB grants', () => { - const template = createWithSessionRole(); + const template = sessionTemplate; const policies = template.findResources('AWS::IAM::Policy'); // Identify the task role's own inline policy: it is the one carrying the @@ -773,28 +815,10 @@ describe('EcsAgentCluster construct', () => { expect(taskRolePolicies).toHaveLength(1); const taskRoleStatements = taskRolePolicies[0][1].Properties.PolicyDocument.Statement; - // No unconditioned dynamodb item grant on the task role (the only DDB the - // task role may touch directly is UserConcurrencyTable — assert that any - // DDB statement present is NOT a leading-key-less task-table grant by - // checking none grant dynamodb write actions without a condition beyond - // the concurrency table). Simplest robust check: the task role carries no - // dynamodb:GetItem/Query/BatchWriteItem statement at all for the task - // tables — grantReadWriteData on a removed table would have produced one. - const ddbItemStatements = taskRoleStatements.filter((s: { Action: string | string[] }) => { - const actions = Array.isArray(s.Action) ? s.Action : [s.Action]; - return actions.some((a: string) => - ['dynamodb:GetItem', 'dynamodb:Query', 'dynamodb:BatchWriteItem'].includes(a), - ); - }); - // The only permitted DDB item access on the task role is the - // UserConcurrencyTable grant. The two task-scoped tables (TaskTable, - // TaskEventsTable) must NOT appear — assert no statement references them. - const serialized = JSON.stringify(ddbItemStatements); - expect(serialized).not.toContain('TaskTable'); - expect(serialized).not.toContain('TaskEventsTable'); - - // The conditioned (SessionRole) DDB statements still exist — exactly two - // task-scoped tables, each leading-key gated. + // All tenant data stays on SessionRole; the agent has no counter access. + expect(JSON.stringify(taskRoleStatements)).not.toContain('dynamodb:'); + + // Main-task read/update statements plus events and approval grants. let conditioned = 0; for (const policy of Object.values(policies)) { for (const s of policy.Properties.PolicyDocument.Statement) { @@ -803,11 +827,45 @@ describe('EcsAgentCluster construct', () => { } } } - expect(conditioned).toBe(2); + expect(conditioned).toBe(4); + }); + + test('both task definitions point at the approval table authorized by the SessionRole', () => { + const tables = sessionTemplate.findResources('AWS::DynamoDB::Table'); + const approvalId = Object.keys(tables).find(id => id.startsWith('TaskApprovalsTable')); + expect(approvalId).toBeDefined(); + const definitions = Object.values(sessionTemplate.findResources('AWS::ECS::TaskDefinition')); + expect(definitions).toHaveLength(2); + for (const definition of definitions) { + expect(definition.Properties.ContainerDefinitions[0].Environment).toContainEqual({ + Name: 'TASK_APPROVALS_TABLE_NAME', + Value: { Ref: approvalId }, + }); + } }); }); }); +describe('EcsAgentCluster approval wiring without a SessionRole', () => { + test('rejects approval wiring without task-scoped credentials before synthesis', () => { + expect(() => createStack({ withApprovals: true })).toThrow('ECS approvals require agentSessionRole and approvalRequestsApiUrl'); + }); + + test('rejects a build setting that would erase approval-table wiring', () => { + const node = new Stack(new App({ + context: { ecsExtraBuildEnv: { TASK_APPROVALS_TABLE_NAME: '' } }, + }), 'S').node; + expect(() => resolveEcsTaskSizing(node)).toThrow('TASK_APPROVALS_TABLE_NAME'); + }); + + test('rejects a build setting that would replace the approval service', () => { + const node = new Stack(new App({ + context: { ecsExtraBuildEnv: { APPROVAL_REQUESTS_API_URL: 'https://other.example' } }, + }), 'S').node; + expect(() => resolveEcsTaskSizing(node)).toThrow('APPROVAL_REQUESTS_API_URL'); + }); +}); + describe('EcsAgentCluster payload bucket (#502)', () => { function createWithPayloadBucket(): Template { const app = new App(); @@ -853,7 +911,7 @@ describe('EcsAgentCluster payload bucket (#502)', () => { }); }); - test('grants the task role READ on the payload bucket, never write/delete', () => { + test('allows only bootstrap reads and explicitly denies other objects and listing', () => { const template = createWithPayloadBucket(); const policies = template.findResources('AWS::IAM::Policy'); const s3Actions = new Set(); @@ -865,7 +923,13 @@ describe('EcsAgentCluster payload bucket (#502)', () => { } } } - // Read actions present... + const statements = Object.values(policies).flatMap(p => p.Properties.PolicyDocument.Statement); + const allow = statements.find(s => s.Effect === 'Allow' && s.Action === 's3:GetObject'); + expect(JSON.stringify(allow.Resource)).toContain('/bootstrap/*'); + expect(statements).toEqual(expect.arrayContaining([ + expect.objectContaining({ Effect: 'Deny', Action: 's3:GetObject*', NotResource: allow.Resource }), + expect.objectContaining({ Effect: 'Deny', Action: 's3:List*' }), + ])); expect([...s3Actions].some(a => a === 's3:GetObject' || a === 's3:GetObject*')).toBe(true); // ...and NO write/delete on the payload bucket from the task role. expect(s3Actions.has('s3:PutObject')).toBe(false); diff --git a/cdk/test/constructs/fanout-consumer.test.ts b/cdk/test/constructs/fanout-consumer.test.ts index 9a7b15149..52a6ae14a 100644 --- a/cdk/test/constructs/fanout-consumer.test.ts +++ b/cdk/test/constructs/fanout-consumer.test.ts @@ -39,6 +39,26 @@ function makeTaskEventsTable(stack: Stack): dynamodb.Table { } describe('FanOutConsumer', () => { + test('binds approval notifications to only GetItem and UpdateItem on the approvals table', () => { + const stack = new Stack(new App(), 'ApprovalNotifications'); + const table = new dynamodb.Table(stack, 'Approvals', { + partitionKey: { name: 'task_id', type: dynamodb.AttributeType.STRING }, + }); + new FanOutConsumer(stack, 'FanOut', { + taskEventsTable: makeTaskEventsTable(stack), taskApprovalsTable: table, + }); + const template = Template.fromStack(stack); + template.hasResourceProperties('AWS::Lambda::Function', { + Environment: { Variables: Match.objectLike({ TASK_APPROVALS_TABLE_NAME: stack.resolve(table.tableName) }) }, + }); + template.hasResourceProperties('AWS::IAM::Policy', { + PolicyDocument: { + Statement: Match.arrayWith([Match.objectLike({ + Action: ['dynamodb:GetItem', 'dynamodb:UpdateItem'], Resource: [stack.resolve(table.tableArn)], + })]), + }, + }); + }); test('attaches a single DynamoEventSource on the TaskEventsTable stream', () => { const app = new App(); const stack = new Stack(app, 'TestStack'); diff --git a/cdk/test/constructs/lambda-microvm-compute.test.ts b/cdk/test/constructs/lambda-microvm-compute.test.ts index 133816282..b39990f4a 100644 --- a/cdk/test/constructs/lambda-microvm-compute.test.ts +++ b/cdk/test/constructs/lambda-microvm-compute.test.ts @@ -53,6 +53,7 @@ import { LAMBDA_MICROVM_SUPPORTED_REGIONS } from '../../src/handlers/shared/micr // keep passing after someone lowered the real budget. const BASE_IMAGE_ARN = 'arn:aws:lambda:us-east-1:aws:microvm-image:al2023-1'; +const ARTIFACT_SHA256 = 'a'.repeat(64); const GITHUB_TOKEN_SECRET_ARN = 'arn:aws:secretsmanager:us-east-1:123456789012:secret:abca/github-token-AbCdEf'; /** @@ -67,6 +68,9 @@ interface BuildOptions { readonly region?: string; readonly context?: Record; readonly withImage?: boolean; + readonly imageEnvironmentVariables?: Record; + readonly artifactSha256?: string | null; + readonly managedImageVersion?: string; readonly externalImageIdentifier?: string; readonly externalImageVersion?: string; readonly withSessionRole?: boolean; @@ -113,11 +117,13 @@ function instantiate(options: BuildOptions = {}): Omit { }); agentSessionRole = new AgentSessionRole(stack, 'AgentSessionRole', { assumingRoles: [runtimeRole], - taskScopedTables: [ - new dynamodb.Table(stack, 'TaskTable', { - partitionKey: { name: 'task_id', type: dynamodb.AttributeType.STRING }, - }), - ], + taskTable: new dynamodb.Table(stack, 'TaskTable', { + partitionKey: { name: 'task_id', type: dynamodb.AttributeType.STRING }, + }), + approvalsTable: new dynamodb.Table(stack, 'ApprovalReadTable', { + partitionKey: { name: 'task_id', type: dynamodb.AttributeType.STRING }, + }), + taskScopedTables: [], traceArtifactsBucket: new s3.Bucket(stack, 'TraceBucket'), attachmentsBucket: new s3.Bucket(stack, 'AttachmentsBucket'), }); @@ -141,9 +147,12 @@ function instantiate(options: BuildOptions = {}): Omit { ...(options.withImage && { baseImageArn: BASE_IMAGE_ARN, baseImageVersion: '1', + artifactSha256: options.artifactSha256 === null ? undefined : options.artifactSha256 ?? ARTIFACT_SHA256, }), externalImageIdentifier: options.externalImageIdentifier, externalImageVersion: options.externalImageVersion, + managedImageVersion: options.managedImageVersion, + imageEnvironmentVariables: options.imageEnvironmentVariables, minimumMemoryInMiB: options.minimumMemoryInMiB, }); @@ -165,6 +174,33 @@ describe('LambdaMicrovmCompute — image provisioned from a managed base image', template = built.template; }); + describe('explicit managed runtime version', () => { + let pinned: Built; + + beforeAll(() => { + pinned = build({ + withImage: true, withSessionRole: true, withRuntimeParity: true, managedImageVersion: '7.0', + }); + }); + + test('pins new workers without changing the managed image build or ownership', () => { + expect(pinned.construct.imageVersion).toBe('7.0'); + expect(pinned.template.findResources('AWS::Lambda::MicrovmImage')) + .toEqual(template.findResources('AWS::Lambda::MicrovmImage')); + }); + + // Constructor-validation cases deliberately have no successful template to cache. + test.each(['', 'latest', '0', '1.a'])('rejects invalid runtime version %p', managedImageVersion => { + expect(() => build({ withImage: true, managedImageVersion })) + .toThrow('microvm_managed_image_version'); + }); + + test('requires managed image inputs', () => { + expect(() => build({ managedImageVersion: '7.0' })) + .toThrow('microvm_managed_image_version'); + }); + }); + test('synthesizes exactly one MicroVM image from the artifact bucket object', () => { template.resourceCountIs('AWS::Lambda::MicrovmImage', 1); template.hasResourceProperties('AWS::Lambda::MicrovmImage', { @@ -175,26 +211,36 @@ describe('LambdaMicrovmCompute — image provisioned from a managed base image', 'Fn::Join': ['', [ 's3://', { Ref: Match.stringLikeRegexp('LambdaMicrovmComputeArtifactBucket') }, - `/${MICROVM_ARTIFACT_OBJECT_KEY}`, + `/microvm-images/agent-artifact-${ARTIFACT_SHA256}.zip`, ]], }, }, }); }); + test('a changed artifact updates the URI without replacing the image identity', () => { + const next = build({ withImage: true, artifactSha256: 'b'.repeat(64) }); + const firstImages = template.findResources('AWS::Lambda::MicrovmImage'); + const nextImages = next.template.findResources('AWS::Lambda::MicrovmImage'); + expect(Object.keys(nextImages)).toEqual(Object.keys(firstImages)); + const first = Object.values(firstImages)[0]!.Properties; + const second = Object.values(nextImages)[0]!.Properties; + expect(second.Name).toEqual(first.Name); + expect(second.CodeArtifact.Uri).not.toEqual(first.CodeArtifact.Uri); + expect(JSON.stringify(second.CodeArtifact.Uri)).toContain(`agent-artifact-${'b'.repeat(64)}.zip`); + }); + + test.each([null, '', 'not-a-sha256', 'A'.repeat(64), 'a'.repeat(63)])( + 'rejects managed builds without an exact artifact digest: %p', + artifactSha256 => { + expect(() => instantiate({ withImage: true, artifactSha256 })) + .toThrow(/microvm_artifact_sha256/); + }, + ); + test('builds an ARM64 image at the largest ACCEPTED BASELINE (8 GiB)', () => { - // 32768 was rejected live: "The requested memory size of 32768 MiB is not - // supported by base MicroVM image …al2023-1. Supported memory sizes in MiB - // are: [512, 1024, 2048, 4096, 8192]." Note this configures the BASELINE — - // the service scales vertically to a 32 GiB / 16 vCPU peak on its own, which - // is why nothing here asks for the peak. - // - // `ARM_64`, not `arm64`: the CDK L1 types Architecture as a plain string and - // documents no allowed values, and CloudFormation rejected the lowercase - // spelling at change-set early validation — "arm64 is not a valid enum value. - // Supported values: [ARM_64]" (ADR-021 P2-F2). The literal is spelled out here - // rather than imported from the construct so the test fails if the constant is - // "corrected" back to Docker's spelling. + // Assert the accepted baseline and API enum spelling independently of source + // constants; this test does not measure runtime memory or scaling. template.hasResourceProperties('AWS::Lambda::MicrovmImage', { CpuConfigurations: [{ Architecture: 'ARM_64' }], Resources: [{ MinimumMemoryInMiB: 8192 }], @@ -204,13 +250,9 @@ describe('LambdaMicrovmCompute — image provisioned from a managed base image', expect(Math.max(...MICROVM_SUPPORTED_MEMORY_MIB)).toBe(DEFAULT_MINIMUM_MEMORY_MIB); }); - test('declares EXACTLY the four hooks the agent serves (P2), and no more', () => { - // `toEqual` on the whole object, not per-key assertions: the invariant runs in - // BOTH directions. A hook the agent serves but the image does not declare is - // never called (the P2 R2 regression this replaces — the agent gained - // /validate and /terminate while the construct still advertised two hooks); - // a hook the image declares but the agent does not serve fails the - // corresponding build or lifecycle transition. Only an exact set catches both. + test('declares all six served hooks with the shared P3 lifecycle budget', () => { + // Compare the whole set; declaration and runtime support must move together. + // The supervisor still gates automatic sleep on each worker's saved capability. const images = template.findResources('AWS::Lambda::MicrovmImage'); const hooks = Object.values(images)[0]!.Properties.Hooks; expect(hooks.Port).toBe(8080); @@ -226,10 +268,14 @@ describe('LambdaMicrovmCompute — image provisioned from a managed base image', RunTimeoutInSeconds: 60, Terminate: 'ENABLED', // Near the BOTTOM of the service's 1–60 s window on purpose: the handler is - // a log-and-acknowledge with nothing to drain (progress writes are already - // durable per event), and the budget bounds how long teardown waits on a + // a barrier close plus log-and-acknowledge with no progress queue to drain + // (ordinary event writes are best effort), and the budget bounds teardown on a // WEDGED guest that is still holding admission-gating memory quota. TerminateTimeoutInSeconds: 15, + Suspend: 'ENABLED', + SuspendTimeoutInSeconds: sharedConstants.microvm_hook_budgets.lifecycle_hook_timeout_seconds, + Resume: 'ENABLED', + ResumeTimeoutInSeconds: sharedConstants.microvm_hook_budgets.lifecycle_hook_timeout_seconds, }); // BUILD (image) hooks. /ready is MANDATORY: create-microvm-image refuses ANY @@ -303,27 +349,25 @@ describe('LambdaMicrovmCompute — image provisioned from a managed base image', }; const image = Object.values(template.findResources('AWS::Lambda::MicrovmImage'))[0]!; - // Hooks: all four states AND all four timeouts, in one comparison — which is + // Hooks: all six states AND all six timeouts, in one comparison — which is // precisely the assertion the missing one would have been. expect(toCfnKeys(flagJson('hooks'))).toEqual(image.Properties.Hooks); // ...and the architecture enum, the other half of P2-F2. expect(toCfnKeys(flagJson('cpu-configurations'))).toEqual(image.Properties.CpuConfigurations); + expect(Object.entries(flagJson('environment-variables') as Record) + .map(([Key, Value]) => ({ Key, Value }))).toEqual(image.Properties.EnvironmentVariables); }); - test('does NOT declare /suspend or /resume — they are P3 and nothing answers them yet', () => { - // The remaining half of the exactness rule, called out separately because it - // is the one that must survive P3 landing suspend/resume in ONE commit across - // all three strategies: until then, declaring either fails the corresponding - // lifecycle transition on a real suspend attempt. + test('pause/wake handlers leave response headroom inside declared service timeouts', () => { const images = template.findResources('AWS::Lambda::MicrovmImage'); const hooks = Object.values(images)[0]!.Properties.Hooks; - expect(hooks.MicrovmHooks.Suspend).toBeUndefined(); - expect(hooks.MicrovmHooks.SuspendTimeoutInSeconds).toBeUndefined(); - expect(hooks.MicrovmHooks.Resume).toBeUndefined(); - expect(hooks.MicrovmHooks.ResumeTimeoutInSeconds).toBeUndefined(); + expect(hooks.MicrovmHooks.SuspendTimeoutInSeconds) + .toBeGreaterThan(sharedConstants.microvm_hook_budgets.lifecycle_handler_budget_seconds); + expect(hooks.MicrovmHooks.ResumeTimeoutInSeconds) + .toBeGreaterThan(sharedConstants.microvm_hook_budgets.lifecycle_handler_budget_seconds); }); - test('the agent hook routes are exactly the four the service calls, under one prefix', () => { + test('the six declared agent hook routes share the service-owned prefix', () => { // The cross-package contract that used to be checked against the rendered // template. It cannot be any more: the template carries `ENABLED`, not a path // (P2-F2), so the routes now have a dedicated source — MICROVM_AGENT_HOOK_ROUTES @@ -334,13 +378,15 @@ describe('LambdaMicrovmCompute — image provisioned from a managed base image', // service POSTs to exactly these paths ("POST /aws/lambda-microvms/runtime/v1/ // ready HTTP/1.1" 200 OK, and the same for the other three). const routes = Object.values(MICROVM_AGENT_HOOK_ROUTES); - expect(routes).toHaveLength(4); + expect(routes).toHaveLength(6); for (const route of routes) { - expect(route).toMatch(/^\/aws\/lambda-microvms\/runtime\/v1\/(ready|validate|run|terminate)$/); + expect(route).toMatch(/^\/aws\/lambda-microvms\/runtime\/v1\/(ready|validate|run|terminate|suspend|resume)$/); } expect([...routes].sort()).toEqual([ '/aws/lambda-microvms/runtime/v1/ready', + '/aws/lambda-microvms/runtime/v1/resume', '/aws/lambda-microvms/runtime/v1/run', + '/aws/lambda-microvms/runtime/v1/suspend', '/aws/lambda-microvms/runtime/v1/terminate', '/aws/lambda-microvms/runtime/v1/validate', ]); @@ -348,7 +394,7 @@ describe('LambdaMicrovmCompute — image provisioned from a managed base image', // Hooks properties use — so "the agent serves every hook the image enables" // stays checkable from one place. expect(Object.keys(MICROVM_AGENT_HOOK_ROUTES).sort()) - .toEqual(['ready', 'run', 'terminate', 'validate']); + .toEqual(['ready', 'resume', 'run', 'suspend', 'terminate', 'validate']); }); test('every declared hook timeout sits inside the service window for its kind', () => { @@ -395,11 +441,20 @@ describe('LambdaMicrovmCompute — image provisioned from a managed base image', .toBeGreaterThanOrEqual(30); }); - test('bakes NO environment variables into the snapshot (ADR-021: no secrets in the image)', () => { + test('bakes only the non-secret image protocol marker by default', () => { template.hasResourceProperties('AWS::Lambda::MicrovmImage', { - EnvironmentVariables: [], + EnvironmentVariables: [{ + Key: sharedConstants.microvm_lifecycle.image_protocol_env, + Value: String(sharedConstants.microvm_lifecycle.protocol_version), + }], }); }); + test('rejects an override of the image source protocol', () => { + expect(() => instantiate({ + withImage: true, + imageEnvironmentVariables: { [sharedConstants.microvm_lifecycle.image_protocol_env]: '999' }, + })).toThrow('protocol marker is owned by the image source'); + }); test('routes image build-time egress through the BUILD connector, not the runtime one', () => { // The runtime connector is 443-only, and `agent/Dockerfile` runs `apt-get` @@ -575,7 +630,7 @@ describe('LambdaMicrovmCompute — image provisioned from a managed base image', } }); - test('payload bucket expires objects (the ONLY reaper on this backend)', () => { + test('payload bucket expires objects as a fallback for failed finalization', () => { template.hasResourceProperties('AWS::S3::Bucket', { LifecycleConfiguration: { Rules: Match.arrayWith([ @@ -622,29 +677,10 @@ describe('LambdaMicrovmCompute — image provisioned from a managed base image', }); test('NO source-key condition on any MicroVM-facing role trust (P2-F1/F3)', () => { - // The sharpest IAM assertion in this file, and the one most likely to be - // "fixed" back by a reviewer applying the standard service-principal - // confused-deputy pattern. It must not be. - // - // The Lambda MicroVMs service presents NO source condition key when it assumes - // these roles, so a trust policy carrying one is unassumable. Live 2026-08-06/07 - // (evidence inlined in ADR-021 §4; `docs/verification/645-p2-smoke-runbook.md` - // is the raw session log, additional detail rather than the sole proof), one - // root cause, two symptoms: - // both network connectors CREATE_FAILED deterministically on a freshly deleted - // stack (P2-F1), and RunMicrovm reported a MISLEADING caller-side - // `iam:PassRole` AccessDenied on the orchestrator (P2-F3) — with the grant - // present, `simulate-principal-policy` returning `allowed`, no permissions - // boundary, and an unconditioned PassRole ALSO denied. Removing the execution - // role's trust conditions made the next submission reach RUNNING in 6 s. - // - // What compensates is asserted elsewhere in this file and in - // `test/constructs/task-orchestrator.test.ts`: the EXECUTION role is passable by - // the orchestrator only, scoped to its EXACT ARN — and with NO - // `iam:PassedToService` condition either, because the same missing-context-key - // root cause blocks that path too (P2r2-F10), which is why the exact ARN is the - // whole of the scoping. Every resource these roles reach is account-scoped by - // ARN apart from two justified `Resource: '*'` statements. + // Recorded service calls rejected source-conditioned role trust; removing + // those conditions restored connector creation and worker launch. Keep exact + // resource grants and verify service support before adding conditions again. + // See ADR-021 §4 for the tested trust-policy limitations. const roles = Object.entries(template.findResources('AWS::IAM::Role')) .filter(([id]) => id.includes('LambdaMicrovmComputeBuildRole') || id.includes('LambdaMicrovmComputeExecutionRole') @@ -663,7 +699,7 @@ describe('LambdaMicrovmCompute — image provisioned from a managed base image', } }); - test('build role reads exactly the one artifact object and writes MicroVM logs', () => { + test('build role reads exactly the selected and manual artifacts and writes MicroVM logs', () => { const policies = Object.entries(template.findResources('AWS::IAM::Policy')) .filter(([id]) => id.includes('LambdaMicrovmComputeBuildRole')); expect(policies).toHaveLength(1); @@ -679,21 +715,31 @@ describe('LambdaMicrovmCompute — image provisioned from a managed base image', // Object-scoped, not bucket-scoped. const s3Statement = statements.find((s: { Action: string }) => s.Action === 's3:GetObject'); expect(JSON.stringify(s3Statement.Resource)).toContain(MICROVM_ARTIFACT_OBJECT_KEY); + expect(s3Statement.Resource).toEqual([ + built.stack.resolve(built.construct.artifactBucket.arnForObjects( + `microvm-images/agent-artifact-${ARTIFACT_SHA256}.zip`, + )), + built.stack.resolve(built.construct.artifactBucket.arnForObjects(MICROVM_ARTIFACT_OBJECT_KEY)), + ]); }); - test('execution role gets READ-ONLY on the payload bucket and no write/delete', () => { + test('execution role gets only bootstrap reads, explicit payload/list denies and no writes', () => { const policies = Object.entries(template.findResources('AWS::IAM::Policy')) .filter(([id]) => id.includes('LambdaMicrovmComputeExecutionRole')); const statements = policies.flatMap(([, p]) => p.Properties.PolicyDocument.Statement); const actions: string[] = statements.flatMap((s: { Action: string | string[] }) => Array.isArray(s.Action) ? s.Action : [s.Action]); - // CDK's grantRead renders the Get*/List* read set. - expect(actions).toContain('s3:GetObject*'); + const allowed = statements.filter((s: { Effect: string }) => s.Effect === 'Allow' && s.Action === 's3:GetObject'); + expect(allowed).toHaveLength(1); + expect(JSON.stringify(allowed[0].Resource)).toContain('/bootstrap/*'); + const denied = statements.find((s: { NotResource?: unknown }) => s.NotResource); + expect(denied.Effect).toBe('Deny'); + expect(denied.NotResource).toEqual(allowed[0].Resource); // The MicroVM runs untrusted repo code — it must not be able to clobber // another task's payload, so nothing mutating may appear. const s3Actions = actions.filter(a => a.startsWith('s3:')); - expect(s3Actions).toEqual(['s3:GetObject*', 's3:GetBucket*', 's3:List*']); + expect(s3Actions).toEqual(['s3:GetObject', 's3:GetObject*', 's3:List*']); for (const action of s3Actions) { expect(action).not.toMatch(/Put|Delete|Abort|Write|^s3:\*$/); } @@ -745,7 +791,7 @@ describe('LambdaMicrovmCompute — image provisioned from a managed base image', // ecs-agent-cluster). Without it a Linear/Jira task's 👀→✅ reaction and the // channel MCP silently no-op. const prefixStatement = executionRoleStatements().find( - statement => JSON.stringify(statement.Resource).includes('bgagent-linear-oauth-*'), + statement => JSON.stringify(statement.Resource ?? '').includes('bgagent-linear-oauth-*'), )!; expect(prefixStatement).toBeDefined(); expect(prefixStatement.Action).toBe('secretsmanager:GetSecretValue'); @@ -800,7 +846,7 @@ describe('LambdaMicrovmCompute — image provisioned from a managed base image', // why it survived P2: the substrate looked fine while the platform's canonical // observability streams were empty. const statements = executionRoleStatements().filter((statement) => { - const resource = JSON.stringify(statement.Resource); + const resource = JSON.stringify(statement.Resource ?? ''); return resource.includes('ApplicationLogGroup'); }); expect(statements).toHaveLength(1); @@ -893,7 +939,7 @@ describe('LambdaMicrovmCompute — image provisioned from a managed base image', const s3Resources = executionRoleStatements() .flatMap(st => (Array.isArray(st.Action) ? st.Action : [st.Action])) .filter(action => action.startsWith('s3:')); - expect(s3Resources).toEqual(['s3:GetObject*', 's3:GetBucket*', 's3:List*']); + expect(s3Resources).toEqual(['s3:GetObject', 's3:GetObject*', 's3:List*']); const rendered = JSON.stringify( executionRoleStatements().filter((st) => { const actions = Array.isArray(st.Action) ? st.Action : [st.Action]; @@ -946,11 +992,7 @@ describe('LambdaMicrovmCompute — image provisioned from a managed base image', expect(JSON.stringify(template.toJSON())).not.toContain('CreateMicrovmAuthToken'); }); - test('warns that a configured image has no smoke-parity guarantee (hook phasing)', () => { - // ADR-021 sub-decision 3, as corrected by the live P1 run and completed in P2: - // all four served hooks are declared, so the image is creatable, launchable and - // payload-deliverable — which makes it look even MORE like a working backend, - // while nothing has exercised clone → change → PR on it. + test('describes configured-image verification and rollback without installation-specific history', () => { const warnings = built.construct.node.metadata.filter(m => m.type === 'aws:cdk:warning'); const message = warnings.map(w => String(w.data)).join('\n'); expect(JSON.stringify(built.construct.node.metadata)) @@ -958,18 +1000,14 @@ describe('LambdaMicrovmCompute — image provisioned from a managed base image', // The superseded id must be gone, not merely reworded — operators grep for it. expect(JSON.stringify(built.construct.node.metadata)) .not.toContain('abca:microvm-image-p1-not-runnable'); - expect(message).toContain('smoke'); - expect(message).toContain('P2'); - // It must state what IS true now, or it reads as the old (wrong) claim — and - // the hook list here is what an operator compares against a failed build or a - // failed lifecycle transition, so all four have to be named. - for (const hook of ['/ready', '/validate', '/run', '/terminate']) { + expect(message).toContain('verify the configured image and coordinator together'); + for (const hook of ['/ready', '/validate', '/run', '/terminate', '/suspend', '/resume']) { expect(message).toContain(hook); } - // ...and it must still say which two are NOT declared, or the enumeration - // above reads as "everything is wired". - expect(message).toContain('/suspend'); - expect(message).toContain('/resume'); + expect(message).toContain('actual launched image version before allowing sleep'); + expect(message).toContain('Nested deployments require bundle 1.9.0'); + expect(message).toContain('explicit image version for rollback'); + expect(message).not.toContain('Image 6.0'); }); test('enables every hook the agent serves, and only those (rendered form)', () => { @@ -978,14 +1016,9 @@ describe('LambdaMicrovmCompute — image provisioned from a managed base image', // the shape CloudFormation validates. `"ENABLED"`, never a path (P2-F2). const images = template.findResources('AWS::Lambda::MicrovmImage'); const rendered = JSON.stringify(Object.values(images)[0]!.Properties.Hooks); - for (const hook of ['Run', 'Terminate', 'Ready', 'Validate']) { + for (const hook of ['Run', 'Terminate', 'Ready', 'Validate', 'Suspend', 'Resume']) { expect(rendered).toContain(`"${hook}":"ENABLED"`); } - // P3, and nothing answers them yet. OMITTED rather than "DISABLED", so the - // absence assertion stays meaningful. - for (const hook of ['Suspend', 'Resume']) { - expect(rendered).not.toContain(hook); - } expect(rendered).not.toContain('DISABLED'); }); @@ -1186,12 +1219,8 @@ describe('LambdaMicrovmCompute — memory sizing', () => { expect(error).toBeDefined(); expect(error!.message).toContain('32768'); expect(error!.message).toContain('512, 1024, 2048, 4096, 8192'); - // The message must say BASELINE, or an operator reads the rejection as "this - // backend caps at 8 GiB" and moves a repo to ECS it did not need to. + // Distinguish baseline validation from runtime capacity. expect(error!.message).toContain('BASELINE'); - expect(error!.message).toContain('32 GiB'); - // ...and points at the backend that DOES have the SUSTAINED capacity. - expect(error!.message).toContain('compute_type=ecs'); }); test.each([0, 256, 6144, 16384, 8193])('rejects the unsupported baseline %i MiB', (mib) => { diff --git a/cdk/test/constructs/lambda-microvm-stack.test.ts b/cdk/test/constructs/lambda-microvm-stack.test.ts new file mode 100644 index 000000000..6d3d581cf --- /dev/null +++ b/cdk/test/constructs/lambda-microvm-stack.test.ts @@ -0,0 +1,206 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { App, CfnOutput, NestedStack, Stack } from 'aws-cdk-lib'; +import { Match, Template } from 'aws-cdk-lib/assertions'; +import * as dynamodb from 'aws-cdk-lib/aws-dynamodb'; +import * as ec2 from 'aws-cdk-lib/aws-ec2'; +import * as iam from 'aws-cdk-lib/aws-iam'; +import * as s3 from 'aws-cdk-lib/aws-s3'; +import { Construct } from 'constructs'; +import { AgentSessionRole } from '../../src/constructs/agent-session-role'; +import { + createMicrovmExecutionRole, + LambdaMicrovmCompute, + type LambdaMicrovmImageInputs, +} from '../../src/constructs/lambda-microvm-compute'; +import { LambdaMicrovmStack } from '../../src/constructs/lambda-microvm-stack'; + +const ENV = { account: '123456789012', region: 'us-east-1' }; +const IMAGE_INPUTS: Array<[string, LambdaMicrovmImageInputs]> = [ + ['bootstrap', {}], + ['imported', { externalImageIdentifier: 'existing-image', externalImageVersion: '6.0' }], + ['managed', { + baseImageArn: 'arn:aws:lambda:us-east-1:aws:microvm-image:al2023-1', + baseImageVersion: '1', + artifactSha256: 'a'.repeat(64), + }], +]; + +describe.each(IMAGE_INPUTS)('LambdaMicrovmStack %s', (mode, imageInputs) => { + let parent: Template; + let child: Template; + let compute: LambdaMicrovmCompute; + + beforeAll(() => { + const stack = new Stack(new App(), 'backgroundagent-dev', { env: ENV }); + const vpc = new ec2.Vpc(stack, 'Vpc', { maxAzs: 2 }); + const executionRole = createMicrovmExecutionRole( + new Construct(stack, 'LambdaMicrovmCompute'), 'ExecutionRole', + ); + const sessionRole = new AgentSessionRole(stack, 'AgentSessionRole', { + assumingRoles: [new iam.Role(stack, 'RuntimeRole', { + assumedBy: new iam.ServicePrincipal('bedrock-agentcore.amazonaws.com'), + })], + taskTable: new dynamodb.Table(stack, 'Tasks', { + partitionKey: { name: 'task_id', type: dynamodb.AttributeType.STRING }, + }), + approvalsTable: new dynamodb.Table(stack, 'ApprovalReadTable', { + partitionKey: { name: 'task_id', type: dynamodb.AttributeType.STRING }, + }), + taskScopedTables: [], + traceArtifactsBucket: new s3.Bucket(stack, 'TraceBucket'), + attachmentsBucket: new s3.Bucket(stack, 'AttachmentsBucket'), + }); + const nested = new LambdaMicrovmStack(stack, 'Microvm', { + ...imageInputs, + vpc, + executionRole, + agentSessionRole: sessionRole, + deploymentName: stack.stackName, + }); + compute = nested.compute; + // Exercise actual parent consumers and child outputs, not just a detached child. + new CfnOutput(stack, 'ArtifactBucket', { value: compute.artifactBucket.bucketName }); + new CfnOutput(stack, 'PayloadBucket', { value: compute.payloadBucket.bucketArn }); + if (compute.imageArn) new CfnOutput(stack, 'ImageArn', { value: compute.imageArn }); + parent = Template.fromStack(stack); // Also rejects parent/child dependency cycles. + child = Template.fromStack(nested); + }); + + test('moves backend resources into the child, preserving parent-owned runtime trust', () => { + parent.resourceCountIs('AWS::Lambda::NetworkConnector', 0); + parent.resourceCountIs('AWS::Lambda::MicrovmImage', 0); + child.resourceCountIs('AWS::Lambda::NetworkConnector', 2); + child.resourceCountIs('AWS::S3::Bucket', 2); + child.resourceCountIs('AWS::Lambda::MicrovmImage', mode === 'managed' ? 1 : 0); + const parentRoles = parent.findResources('AWS::IAM::Role'); + const executionId = Object.keys(parentRoles).find(id => id.startsWith('LambdaMicrovmComputeExecutionRole'))!; + expect(executionId).toBeDefined(); + parent.hasResourceProperties('AWS::IAM::Role', { + AssumeRolePolicyDocument: { + Statement: Match.arrayWith([Match.objectLike({ + Principal: { AWS: { 'Fn::GetAtt': [executionId, 'Arn'] } }, + })]), + }, + }); + expect(Object.keys(child.findResources('AWS::IAM::Role')).some(id => id.includes('ExecutionRole'))).toBe(false); + }); + + test('uses concrete parent-derived service and narrowly scoped bootstrap role names', () => { + expect(compute.imageName).toBe('backgroundagent-dev-abca-agent'); + child.hasResourceProperties('AWS::Lambda::NetworkConnector', { Name: 'backgroundagent-dev-microvm-egress' }); + child.hasResourceProperties('AWS::Lambda::NetworkConnector', { Name: 'backgroundagent-dev-microvm-build-egress' }); + child.hasResourceProperties('AWS::IAM::Role', { RoleName: 'backgroundagent-dev-MicrovmBuildRole' }); + child.hasResourceProperties('AWS::IAM::Role', { RoleName: 'backgroundagent-dev-MicrovmConnectorRole' }); + expect(JSON.stringify(child.toJSON())).not.toContain('Token-TOKEN'); + }); + + test('retains build/runtime network separation and backend tags', () => { + const groups = Object.values(child.findResources('AWS::EC2::SecurityGroup')); + const egressPorts = groups.map(group => + group.Properties.SecurityGroupEgress.map((rule: { FromPort: number }) => rule.FromPort).sort()); + expect(egressPorts).toEqual(expect.arrayContaining([[443], [443, 80]])); + child.hasResourceProperties('AWS::Lambda::NetworkConnector', { + Tags: Match.arrayWith([{ Key: 'abca:compute-backend', Value: 'lambda-microvm' }]), + }); + // CDK's raw cleanup-provider resources inherit tags from CloudFormation; + // they do not expose a CDK TagManager that emits resource-level Tags. + parent.hasResourceProperties('AWS::CloudFormation::Stack', { + Tags: Match.arrayWith([{ Key: 'abca:compute-backend', Value: 'lambda-microvm' }]), + }); + const executionPolicies = Object.entries(parent.findResources('AWS::IAM::Policy')) + .filter(([id]) => id.startsWith('LambdaMicrovmComputeExecutionRole')); + expect(executionPolicies).toHaveLength(1); + expect(JSON.stringify(executionPolicies)).not.toContain('dynamodb:'); + }); + + test('preserves image selection semantics across the boundary', () => { + if (mode === 'bootstrap') { + expect(compute.imageIdentifier).toBeUndefined(); + } else if (mode === 'imported') { + expect(compute.imageArn).toContain(':microvm-image:existing-image'); + expect(compute.imageVersion).toBe('6.0'); + } else { + child.hasResourceProperties('AWS::Lambda::MicrovmImage', { + Name: 'backgroundagent-dev-abca-agent', + Resources: [{ MinimumMemoryInMiB: 8192 }], + Hooks: { MicrovmHooks: Match.objectLike({ Suspend: 'ENABLED', Resume: 'ENABLED' }) }, + }); + expect(compute.imageVersion).toBeUndefined(); + } + }); +}); + +test('rejects nested token sanitization and an execution role in the wrong stack', () => { + const app = new App(); + const parent = new Stack(app, 'Parent', { env: ENV }); + const vpc = new ec2.Vpc(parent, 'Vpc', { maxAzs: 2 }); + const child = new NestedStack(parent, 'UnconfiguredChild'); + expect(() => new LambdaMicrovmCompute(child, 'Compute', { vpc })) + .toThrow(/concrete deploymentName/); + const sibling = new Stack(app, 'Sibling', { env: ENV }); + expect(() => new LambdaMicrovmStack(parent, 'WrongRole', { + vpc, + deploymentName: 'Parent', + executionRole: createMicrovmExecutionRole(sibling, 'ExecutionRole'), + })).toThrow(/owned by its parent/); +}); + +test('rejects a deployment name that would exceed IAM role-name limits', () => { + const parent = new Stack(new App(), 'LongParent', { env: ENV }); + expect(() => new LambdaMicrovmStack(parent, 'Microvm', { + vpc: new ec2.Vpc(parent, 'Vpc', { maxAzs: 2 }), + deploymentName: 'a'.repeat(64), + executionRole: createMicrovmExecutionRole(parent, 'ExecutionRole'), + })).toThrow(/at most 64 characters/); +}); + +test('allows overlapping service names while preserving parent role identity and bootstrap role names', () => { + const parent = new Stack(new App(), 'backgroundagent-dev', { env: ENV }); + const executionRole = createMicrovmExecutionRole( + new Construct(parent, 'LambdaMicrovmCompute'), 'ExecutionRole', + ); + const nested = new LambdaMicrovmStack(parent, 'Microvm', { + ...IMAGE_INPUTS[2][1], + vpc: new ec2.Vpc(parent, 'Vpc', { maxAzs: 2 }), + deploymentName: parent.stackName, + resourceNamePrefix: 'backgroundagent-dev-p3', + executionRole, + }); + const child = Template.fromStack(nested); + child.hasResourceProperties('AWS::Lambda::MicrovmImage', { Name: 'backgroundagent-dev-p3-abca-agent' }); + child.hasResourceProperties('AWS::Lambda::NetworkConnector', { Name: 'backgroundagent-dev-p3-microvm-egress' }); + child.hasResourceProperties('AWS::Lambda::NetworkConnector', { Name: 'backgroundagent-dev-p3-microvm-build-egress' }); + child.hasResourceProperties('AWS::Logs::LogGroup', { LogGroupName: '/aws/lambda-microvms/backgroundagent-dev-p3-abca-agent' }); + child.hasResourceProperties('AWS::IAM::Role', { RoleName: 'backgroundagent-dev-MicrovmBuildRole' }); + child.hasResourceProperties('AWS::IAM::Role', { RoleName: 'backgroundagent-dev-MicrovmConnectorRole' }); + expect(Object.keys(Template.fromStack(parent).findResources('AWS::IAM::Role'))) + .toContain('LambdaMicrovmComputeExecutionRoleAA0C4A0D'); +}); + +test.each(['', '-invalid', 'has/slash', 'a'.repeat(41)])('rejects an unsafe migration name prefix %p', resourceNamePrefix => { + const parent = new Stack(new App(), 'backgroundagent-dev', { env: ENV }); + expect(() => new LambdaMicrovmStack(parent, 'Microvm', { + vpc: new ec2.Vpc(parent, 'Vpc', { maxAzs: 2 }), + deploymentName: parent.stackName, + resourceNamePrefix, + executionRole: createMicrovmExecutionRole(parent, 'ExecutionRole'), + })).toThrow(/microvm_resource_name_prefix/); +}); diff --git a/cdk/test/constructs/linear-integration.test.ts b/cdk/test/constructs/linear-integration.test.ts index 9638d24e1..98a719aac 100644 --- a/cdk/test/constructs/linear-integration.test.ts +++ b/cdk/test/constructs/linear-integration.test.ts @@ -404,3 +404,54 @@ describe('revoked-authorization recording (#812)', () => { expect(JSON.stringify(registryStatements.map((s) => s.Action))).toContain('dynamodb:UpdateItem'); }); }); + +describe('Linear approval decision permissions', () => { + let template: Template; + beforeAll(() => { + const app = new App(); + const stack = new Stack(app, 'ApprovalStack'); + const table = (id: string) => new dynamodb.Table(stack, id, { + partitionKey: { name: 'id', type: dynamodb.AttributeType.STRING }, + }); + new LinearIntegration(stack, 'Linear', { + api: new apigw.RestApi(stack, 'Api'), + userPool: new cognito.UserPool(stack, 'Users'), + taskTable: table('Tasks'), + taskEventsTable: table('Events'), + taskApprovalsTable: table('Approvals'), + userConcurrencyTable: table('Counters'), + continuationBucketName: 'continuations', + orchestratorFunctionArn: 'arn:aws:lambda:us-east-1:123456789012:function:coordinator:live', + lambdaMicrovmImageArn: 'arn:aws:lambda:us-east-1:123456789012:microvm-image:test', + }); + template = Template.fromStack(stack); + }); + test('configures the processor even without orchestration', () => { + template.hasResourceProperties('AWS::Lambda::Function', { + Environment: { + Variables: Match.objectLike({ + TASK_APPROVALS_TABLE_NAME: Match.anyValue(), + USER_CONCURRENCY_TABLE_NAME: Match.anyValue(), + CONTINUATION_BUCKET_NAME: 'continuations', + }), + }, + }); + }); + test('grants wake only for the configured image and continuation dispatch for its coordinator', () => { + const functions = template.findResources('AWS::Lambda::Function'); + const processor = Object.values(functions).find(fn => fn.Properties.Environment?.Variables?.TASK_APPROVALS_TABLE_NAME); + const role = processor!.Properties.Role['Fn::GetAtt'][0]; + const policies = Object.values(template.findResources('AWS::IAM::Policy')) + .filter(p => p.Properties.Roles?.some((r: { Ref?: string }) => r.Ref === role)); + const statements = policies.flatMap(p => p.Properties.PolicyDocument.Statement); + const wake = statements.find(s => Array.isArray(s.Action) && s.Action.includes('lambda:ResumeMicrovm')); + expect(wake.Action).toEqual(['lambda:GetMicrovm', 'lambda:ResumeMicrovm']); + expect(wake.Resource).toEqual(['arn:aws:lambda:us-east-1:123456789012:microvm-image:test', + 'arn:aws:lambda:us-east-1:123456789012:microvm-image:test:*']); + expect(statements).toContainEqual(expect.objectContaining({ + Action: 'lambda:InvokeFunction', Resource: 'arn:aws:lambda:us-east-1:123456789012:function:coordinator:*', + })); + expect(statements.some(s => (Array.isArray(s.Action) ? s.Action : [s.Action]).includes('dynamodb:UpdateItem') + && JSON.stringify(s.Resource).includes('Counters'))).toBe(true); + }); +}); diff --git a/cdk/test/constructs/microvm-continuation-manager.test.ts b/cdk/test/constructs/microvm-continuation-manager.test.ts new file mode 100644 index 000000000..e8a1e5bdc --- /dev/null +++ b/cdk/test/constructs/microvm-continuation-manager.test.ts @@ -0,0 +1,69 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { App, Stack } from 'aws-cdk-lib'; +import { Template } from 'aws-cdk-lib/assertions'; +import * as dynamodb from 'aws-cdk-lib/aws-dynamodb'; +import { ContinuationBucket } from '../../src/constructs/continuation-bucket'; +import { MicrovmContinuationManager } from '../../src/constructs/microvm-continuation-manager'; +import { TaskApi } from '../../src/constructs/task-api'; + +test('scheduled recovery and decisions invoke retained versions of only the original coordinator', () => { + const stack = new Stack(new App(), 'Continuations'); + const table = (id: string) => new dynamodb.Table(stack, id, { + partitionKey: { name: 'task_id', type: dynamodb.AttributeType.STRING }, + }); + const taskTable = table('Tasks'); + const approvalsTable = table('Approvals'); + const userConcurrencyTable = table('Counters'); + const bucket = new ContinuationBucket(stack, 'Saved'); + const coordinator = 'arn:aws:lambda:us-west-2:123456789012:function:coordinator'; + const imageArn = 'arn:aws:lambda:us-west-2:123456789012:microvm-image:agent'; + new MicrovmContinuationManager(stack, 'Manager', { + taskTable, + approvalsTable, + userConcurrencyTable, + continuationBucket: bucket, + orchestratorFunctionArn: coordinator, + imageArn, + }); + const api = new TaskApi(stack, 'Api', { taskTable, taskEventsTable: table('Events'), taskApprovalsTable: approvalsTable }); + api.enableMicrovmContinuations(bucket.bucket.bucketName, coordinator, userConcurrencyTable, 7); + const template = Template.fromStack(stack); + const policies = Object.entries(template.findResources('AWS::IAM::Policy')); + for (const name of ['ManagerReconcilerFn', 'ApproveTaskFn', 'DenyTaskFn']) { + const statements = policies.filter(([id]) => id.includes(name)) + .flatMap(([, policy]) => policy.Properties.PolicyDocument.Statement); + expect(statements).toContainEqual(expect.objectContaining({ + Action: 'lambda:InvokeFunction', Resource: `${coordinator}:*`, + })); + expect(statements.some(statement => JSON.stringify(statement.Action).includes('RunMicrovm'))).toBe(false); + } + const functions = Object.entries(template.findResources('AWS::Lambda::Function')); + for (const name of ['ManagerReconcilerFn', 'ApproveTaskFn', 'DenyTaskFn']) { + const fn = functions.find(([id]) => id.includes(name))![1]; + expect(fn.Properties.Environment.Variables.ORCHESTRATOR_FUNCTION_ARN).toBe(coordinator); + expect(fn.Properties.Environment.Variables.CONTINUATION_BUCKET_NAME).toBeDefined(); + expect(fn.Properties.Environment.Variables.USER_CONCURRENCY_TABLE_NAME) + .toEqual(stack.resolve(userConcurrencyTable.tableName)); + if (name !== 'ManagerReconcilerFn') { + expect(fn.Properties.Environment.Variables.MAX_CONCURRENT_TASKS_PER_USER).toBe('7'); + } + } +}); diff --git a/cdk/test/constructs/payload-bootstrap-permissions.test.ts b/cdk/test/constructs/payload-bootstrap-permissions.test.ts new file mode 100644 index 000000000..20bfddfba --- /dev/null +++ b/cdk/test/constructs/payload-bootstrap-permissions.test.ts @@ -0,0 +1,67 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { App, Stack } from 'aws-cdk-lib'; +import { Template } from 'aws-cdk-lib/assertions'; +import * as iam from 'aws-cdk-lib/aws-iam'; +import * as s3 from 'aws-cdk-lib/aws-s3'; +import { grantCoordinatorPayloads, grantWorkerBootstrap } from '../../src/constructs/payload-bootstrap-permissions'; + +describe('payload bootstrap permission boundary', () => { + let template: Template; + beforeAll(() => { + const stack = new Stack(new App(), 'PayloadPermissions'); + const bucket = new s3.Bucket(stack, 'Payloads'); + const worker = new iam.Role(stack, 'Worker', { assumedBy: new iam.ServicePrincipal('lambda.amazonaws.com') }); + const coordinator = new iam.Role(stack, 'Coordinator', { assumedBy: new iam.ServicePrincipal('lambda.amazonaws.com') }); + grantWorkerBootstrap(bucket, worker); + grantCoordinatorPayloads(bucket, coordinator); + template = Template.fromStack(stack); + }); + + test('worker can read only deployment manifests, including against another public bucket', () => { + const policy = Object.entries(template.findResources('AWS::IAM::Policy')) + .find(([id]) => id.startsWith('Worker'))![1].Properties.PolicyDocument.Statement; + expect(policy).toHaveLength(3); + const allow = policy.find((s: { Effect: string }) => s.Effect === 'Allow'); + expect(allow.Action).toBe('s3:GetObject'); + expect(JSON.stringify(allow.Resource)).toContain('/bootstrap/*'); + const deny = policy.find((s: { NotResource?: unknown }) => s.NotResource); + expect(deny.Effect).toBe('Deny'); + expect(deny.Action).toBe('s3:GetObject*'); + expect(deny.NotResource).toEqual(allow.Resource); + expect(policy.find((s: { Action: string }) => s.Action === 's3:List*').Effect).toBe('Deny'); + expect(JSON.stringify(policy)).not.toContain('s3:Put'); + expect(JSON.stringify(policy)).not.toContain('s3:Delete'); + }); + + test('coordinator publishes manifests and privately persists, signs and deletes task references', () => { + const policy = Object.entries(template.findResources('AWS::IAM::Policy')) + .find(([id]) => id.startsWith('Coordinator'))![1].Properties.PolicyDocument.Statement; + const writes = policy.find((s: { Action: string }) => s.Action === 's3:PutObject'); + expect(JSON.stringify(writes.Resource)).toContain('/bootstrap/*'); + const reads = policy.find((s: { Action: string[] }) => Array.isArray(s.Action) && s.Action.includes('s3:GetObject')); + expect(reads.Action).toEqual(['s3:GetObject', 's3:DeleteObject']); + expect(JSON.stringify(reads.Resource)).toContain('*/payload.json'); + expect(JSON.stringify(reads.Resource)).toContain('*/launch.json'); + expect(JSON.stringify(reads.Resource)).not.toContain('/bootstrap/*'); + const list = policy.find((s: { Action: string }) => s.Action === 's3:ListBucket'); + expect(list.Resource).toEqual({ 'Fn::GetAtt': [expect.stringMatching(/^Payloads/), 'Arn'] }); + }); +}); diff --git a/cdk/test/constructs/registry.test.ts b/cdk/test/constructs/registry.test.ts index 63f84ad72..2cc2992bf 100644 --- a/cdk/test/constructs/registry.test.ts +++ b/cdk/test/constructs/registry.test.ts @@ -19,7 +19,8 @@ import { App, Stack } from 'aws-cdk-lib'; import { Template, Match } from 'aws-cdk-lib/assertions'; -import { AgentRegistry } from '../../src/constructs/registry'; +import { applicationPolicy } from '../../src/bootstrap/policies/application'; +import { AgentRegistry, AgentRegistryStack } from '../../src/constructs/registry'; function createStack(): Template { const app = new App(); @@ -153,3 +154,49 @@ describe('AgentRegistry construct', () => { expect(JSON.stringify(statement?.Resource)).toContain('registry/*'); }); }); + +describe.each([ + 'backgroundagent-dev', + `backgroundagent-dev-${'x'.repeat(108)}`, +])('registry waiter in parent stack %s', stackName => { + let names: string[]; + + beforeAll(() => { + const app = new App(); + const parent = new Stack(app, 'Parent', { + stackName, + env: { account: '123456789012', region: 'us-west-2' }, + }); + const registries = ['FirstRegistry', 'SecondRegistry'].map( + id => new AgentRegistryStack(parent, id, { registryName: id }), + ); + names = registries.map(nested => { + const waiters = Object.values( + Template.fromStack(nested).findResources('AWS::StepFunctions::StateMachine'), + ); + expect(waiters).toHaveLength(1); + return waiters[0].Properties.StateMachineName as string; + }); + }); + + test('uses names permitted by the deployed bootstrap policy', () => { + const statement = applicationPolicy().toJSON().Statement.find( + (candidate: { Sid: string }) => candidate.Sid === 'StepFunctions', + ); + const resourcePattern = new RegExp( + `^${(statement.Resource as string).split('*').join('.*')}$`, + ); + for (const name of names) { + expect(name).toBeDefined(); + expect(`arn:aws:states:us-west-2:123456789012:stateMachine:${name}`) + .toMatch(resourcePattern); + } + }); + + test('keeps names distinct and within the Step Functions name limit', () => { + expect(new Set(names).size).toBe(2); + for (const name of names) { + expect(name).toMatch(/^[A-Za-z0-9-]{1,80}$/); + } + }); +}); diff --git a/cdk/test/constructs/slack-integration.test.ts b/cdk/test/constructs/slack-integration.test.ts index 702e9bf30..12a02f575 100644 --- a/cdk/test/constructs/slack-integration.test.ts +++ b/cdk/test/constructs/slack-integration.test.ts @@ -47,11 +47,36 @@ describe('SlackIntegration construct', () => { userPool, taskTable, taskEventsTable, + taskApprovalsTable: dynamodb.Table.fromTableArn( + stack, 'Approvals', 'arn:aws:dynamodb:us-east-1:123456789012:table/Approvals', + ), }); template = Template.fromStack(stack); }); + test('interaction cancellation can settle approvals and publish their closure event', () => { + const functions = Object.entries(template.findResources('AWS::Lambda::Function')); + const interaction = functions.find(([id]) => id.includes('SlackInteractionsFn'))![1]; + expect(interaction.Properties.Environment.Variables).toMatchObject({ + TASK_APPROVALS_TABLE_NAME: 'Approvals', + TASK_EVENTS_TABLE_NAME: expect.anything(), + TASK_RETENTION_DAYS: expect.any(String), + }); + const policy = Object.entries(template.findResources('AWS::IAM::Policy')) + .find(([id]) => id.includes('SlackInteractionsFn'))![1]; + expect(policy.Properties.PolicyDocument.Statement).toEqual(expect.arrayContaining([ + expect.objectContaining({ + Action: ['dynamodb:GetItem', 'dynamodb:UpdateItem'], + Resource: ['arn:aws:dynamodb:us-east-1:123456789012:table/Approvals'], + }), + expect.objectContaining({ + Action: 'dynamodb:PutItem', + Resource: expect.arrayContaining([expect.objectContaining({ 'Fn::GetAtt': expect.arrayContaining([expect.stringContaining('TaskEventsTable')]) })]), + }), + ])); + }); + test('creates three Slack DynamoDB tables (installation + user mapping + channel mapping)', () => { // TaskTable + TaskEventsTable + SlackInstallation + SlackUserMapping + SlackChannelMapping = 5 template.resourceCountIs('AWS::DynamoDB::Table', 5); diff --git a/cdk/test/constructs/task-api-microvm-wake.test.ts b/cdk/test/constructs/task-api-microvm-wake.test.ts new file mode 100644 index 000000000..4b45e7c37 --- /dev/null +++ b/cdk/test/constructs/task-api-microvm-wake.test.ts @@ -0,0 +1,102 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +// SPDX-License-Identifier: MIT-0 + +import { App, Stack } from 'aws-cdk-lib'; +import { Template } from 'aws-cdk-lib/assertions'; +import * as dynamodb from 'aws-cdk-lib/aws-dynamodb'; +import { TaskApi } from '../../src/constructs/task-api'; + +const IMAGE_ARN = 'arn:aws:lambda:us-east-1:123456789012:microvm-image:agent'; +let configured: Template; +let unconfigured: Template; +function fixture(imageArn?: string): Template { + const stack = new Stack(new App(), 'ApiTest'); + const table = (id: string, sortKey?: string) => new dynamodb.Table(stack, id, { + partitionKey: { name: 'task_id', type: dynamodb.AttributeType.STRING }, + ...(sortKey && { sortKey: { name: sortKey, type: dynamodb.AttributeType.STRING } }), + }); + new TaskApi(stack, 'Api', { + taskTable: table('Tasks'), + taskEventsTable: table('Events', 'event_id'), + taskApprovalsTable: table('Approvals', 'request_id'), + lambdaMicrovmImageArn: imageArn, + }); + return Template.fromStack(stack); +} +type Statement = { Action: string | string[]; Resource: unknown }; +function grants(template: Template, functionName: string): Statement[] { + return Object.entries(template.findResources('AWS::IAM::Policy')) + .filter(([id]) => id.includes(functionName)) + .flatMap(([, policy]) => policy.Properties.PolicyDocument.Statement as Statement[]) + .filter(statement => JSON.stringify(statement.Action).includes('Microvm')); +} +beforeAll(() => { + configured = fixture(IMAGE_ARN); + unconfigured = fixture(); +}); + +test.each(['ApproveTaskFn', 'DenyTaskFn'])('%s gets only Get/Resume against this image', name => { + expect(grants(configured, name)).toEqual([expect.objectContaining({ + Action: ['lambda:GetMicrovm', 'lambda:ResumeMicrovm'], + Resource: [IMAGE_ARN, `${IMAGE_ARN}:*`], + })]); +}); + +test('cancel keeps only Terminate against the same image', () => { + expect(grants(configured, 'CancelTaskFn')).toEqual([expect.objectContaining({ + Action: 'lambda:TerminateMicrovm', Resource: [IMAGE_ARN, `${IMAGE_ARN}:*`], + })]); +}); + +test.each(['ApproveTaskFn', 'DenyTaskFn', 'CancelTaskFn'])('%s gets no lifecycle grant without an image', name => { + expect(grants(unconfigured, name)).toEqual([]); +}); + +test('the decision APIs keep their15s Lambda budget and read worker IDs from saved task metadata', () => { + const functions = Object.entries(configured.findResources('AWS::Lambda::Function')) + .filter(([id]) => id.includes('ApproveTaskFn') || id.includes('DenyTaskFn')); + expect(functions).toHaveLength(2); + for (const [, resource] of functions) { + expect(resource.Properties.Timeout).toBe(15); + expect(resource.Properties.Environment.Variables.MICROVM_IMAGE_IDENTIFIER).toBeUndefined(); + } +}); + +test('cancellation and pending listing receive their approval workflow permissions', () => { + const functions = configured.findResources('AWS::Lambda::Function'); + const cancel = Object.entries(functions).find(([id]) => id.includes('CancelTaskFn'))![1]; + expect(cancel.Properties.Environment.Variables.TASK_APPROVALS_TABLE_NAME).toBeDefined(); + const policies = Object.entries(configured.findResources('AWS::IAM::Policy')); + const cancelPolicy = policies.find(([id]) => id.includes('CancelTaskFn'))![1]; + expect(cancelPolicy.Properties.PolicyDocument.Statement).toEqual(expect.arrayContaining([ + expect.objectContaining({ + Action: ['dynamodb:GetItem', 'dynamodb:UpdateItem'], + Resource: expect.arrayContaining([expect.objectContaining({ 'Fn::GetAtt': expect.arrayContaining([expect.stringContaining('Approvals')]) })]), + }), + ])); + const pendingPolicy = policies.find(([id]) => id.includes('GetPendingFn'))![1]; + expect(pendingPolicy.Properties.PolicyDocument.Statement).toEqual(expect.arrayContaining([ + expect.objectContaining({ + Action: 'dynamodb:BatchGetItem', + Resource: expect.arrayContaining([expect.objectContaining({ 'Fn::GetAtt': expect.arrayContaining([expect.stringContaining('Tasks')]) })]), + }), + ])); +}); diff --git a/cdk/test/constructs/task-api.test.ts b/cdk/test/constructs/task-api.test.ts index 4cb03b1ce..48737bd25 100644 --- a/cdk/test/constructs/task-api.test.ts +++ b/cdk/test/constructs/task-api.test.ts @@ -23,7 +23,7 @@ import * as dynamodb from 'aws-cdk-lib/aws-dynamodb'; import * as s3 from 'aws-cdk-lib/aws-s3'; import { TaskApi, type TaskApiProps } from '../../src/constructs/task-api'; -function createStack(overrides?: Partial): { stack: Stack; template: Template } { +function createStack(overrides?: Partial, withUploads = false): { stack: Stack; template: Template } { const app = new App(); const stack = new Stack(app, 'TestStack'); @@ -39,6 +39,12 @@ function createStack(overrides?: Partial): { stack: Stack; templat new TaskApi(stack, 'TaskApi', { taskTable, taskEventsTable, + ...(withUploads && { + attachmentsBucket: new s3.Bucket(stack, 'UploadsBucket'), + userConcurrencyTable: new dynamodb.Table(stack, 'CapacityTable', { + partitionKey: { name: 'user_id', type: dynamodb.AttributeType.STRING }, + }), + }), ...overrides, }); @@ -131,11 +137,26 @@ describe('TaskApi construct', () => { let baseTemplate: Template; let budgetTemplate: Template; let webhookTemplate: Template; + let uploadsTemplate: Template; beforeAll(() => { baseTemplate = createStack().template; budgetTemplate = createStackWithBudget().template; webhookTemplate = createStackWithWebhooks().template; + uploadsTemplate = createStack(undefined, true).template; + }); + + test('upload confirmation can inspect capacity but cannot reserve or return a seat', () => { + const policy = Object.entries(uploadsTemplate.findResources('AWS::IAM::Policy')) + .find(([id]) => id.includes('ConfirmUploadsFn'))![1]; + const tableId = Object.keys(uploadsTemplate.findResources('AWS::DynamoDB::Table')).find(id => id.startsWith('CapacityTable'))!; + const actions = policy.Properties.PolicyDocument.Statement + .filter((statement: any) => JSON.stringify(statement.Resource).includes(tableId)) + .flatMap((statement: any) => [statement.Action].flat()); + expect(actions).toContain('dynamodb:GetItem'); + expect(actions).not.toContain('dynamodb:UpdateItem'); + expect(actions).not.toContain('dynamodb:PutItem'); + expect(actions).not.toContain('dynamodb:DeleteItem'); }); test('creates a Cognito User Pool', () => { diff --git a/cdk/test/constructs/task-orchestrator.test.ts b/cdk/test/constructs/task-orchestrator.test.ts index a7b2e9cf2..b59d83d7e 100644 --- a/cdk/test/constructs/task-orchestrator.test.ts +++ b/cdk/test/constructs/task-orchestrator.test.ts @@ -36,6 +36,7 @@ interface StackOverrides { /** ADR-021 P2: the identifiers the orchestrator forwards as `platform_config`. */ agentPlatformConfig?: { taskApprovalsTableName: string; + approvalRequestsApiUrl?: string; nudgesTableName: string; logGroupName: string; artifactsBucketName: string; @@ -187,6 +188,14 @@ describe('TaskOrchestrator construct', () => { }); }); + test('retains published versions so in-flight durable executions can replay after deployment', () => { + baseTemplate.resourceCountIs('AWS::Lambda::Version', 1); + baseTemplate.hasResource('AWS::Lambda::Version', { + DeletionPolicy: 'Retain', + UpdateReplacePolicy: 'Retain', + }); + }); + test('grants AgentCore runtime invocation permissions with wildcard sub-resource', () => { baseTemplate.hasResourceProperties('AWS::IAM::Policy', { PolicyDocument: { @@ -608,6 +617,7 @@ describe('TaskOrchestrator with the Lambda MicroVMs backend (ADR-021)', () => { imageIdentifier?: string; imageArn?: string; imageVersion?: string; + approvalSuspendEnabled?: boolean; ingressConnectorArns?: string[]; }): { template: Template } { const app = new App(); @@ -627,6 +637,8 @@ describe('TaskOrchestrator with the Lambda MicroVMs backend (ADR-021)', () => { }), runtimeArn: 'arn:aws:bedrock-agentcore:us-east-1:123456789012:runtime/test-runtime', microvmConfig: { + approvalsTable: mkTable('MicrovmApprovalsTable', 'request_id'), + approvalSuspendEnabled: config?.approvalSuspendEnabled, imageIdentifier: config?.imageIdentifier ?? IMAGE_ARN, imageArn: config?.imageArn ?? IMAGE_ARN, imageVersion: config?.imageVersion, @@ -672,12 +684,13 @@ describe('TaskOrchestrator with the Lambda MicroVMs backend (ADR-021)', () => { template = createMicrovmStack().template; pinnedTemplate = createMicrovmStack({ imageVersion: '4', + approvalSuspendEnabled: true, ingressConnectorArns: ['arn:aws:lambda:us-east-1:aws:network-connector:x', 'arn:y'], }).template; // What LambdaMicrovmCompute passes when the operator gave a bare image NAME: // the exact ARN it derived, never a wildcard. nameDerivedTemplate = createMicrovmStack({ - imageIdentifier: 'abca-agent', + imageIdentifier: NAME_DERIVED_IMAGE_ARN, imageArn: NAME_DERIVED_IMAGE_ARN, }).template; noMicrovmTemplate = createStack().template; @@ -692,6 +705,34 @@ describe('TaskOrchestrator with the Lambda MicroVMs backend (ADR-021)', () => { expect(env.MICROVM_PAYLOAD_BUCKET).toBeDefined(); }); + test('new sleep defaults off while explicit opt-in enables it', () => { + expect(orchestratorEnv(template).MICROVM_APPROVAL_SUSPEND_ENABLED).toBe('false'); + expect(orchestratorEnv(pinnedTemplate).MICROVM_APPROVAL_SUSPEND_ENABLED).toBe('true'); + template.hasResourceProperties('AWS::SSM::Parameter', { + Name: '/TestStack/microvm-approval-suspend-enabled', Type: 'String', Value: 'false', + }); + pinnedTemplate.hasResourceProperties('AWS::SSM::Parameter', { + Name: '/TestStack/microvm-approval-suspend-enabled', Type: 'String', Value: 'true', + }); + }); + + test('existing executions receive a stable parameter name with exact read-only permission', () => { + const [parameterId] = Object.keys(template.findResources('AWS::SSM::Parameter')); + expect(orchestratorEnv(template).MICROVM_APPROVAL_SUSPEND_PARAMETER_NAME).toEqual({ Ref: parameterId }); + const statement = microvmStatements(template).find(s => s.Sid === 'MicrovmSuspendConfiguration')!; + expect(statement.Action).toBe('ssm:GetParameter'); + expect(statement.Resource).toEqual({ + 'Fn::Join': ['', ['arn:', { Ref: 'AWS::Partition' }, ':ssm:us-east-1:123456789012:parameter', { Ref: parameterId }]], + }); + }); + + test('approval observation grants read and transaction condition checks, never approval writes', () => { + const statement = microvmStatements(template).find(s => s.Sid === 'MicrovmApprovalObservation')!; + expect(statement.Action).toEqual(['dynamodb:GetItem', 'dynamodb:ConditionCheckItem']); + expect(JSON.stringify(statement.Resource)).toContain('MicrovmApprovalsTable'); + expect(orchestratorEnv(template).TASK_APPROVALS_TABLE_NAME).toEqual({ Ref: expect.stringMatching(/^MicrovmApprovalsTable/) }); + }); + test('ALWAYS injects the ingress var, carrying the NO_INGRESS control', () => { // OUTCOME assertion, not an omission assertion. This var used to be injected // only when non-empty, and the test asserted it was `undefined` — which is @@ -715,6 +756,8 @@ describe('TaskOrchestrator with the Lambda MicroVMs backend (ADR-021)', () => { // ingress included, is present whenever microvmConfig is. const env = orchestratorEnv(template); expect(Object.keys(env).filter((k) => k.startsWith('MICROVM_')).sort()).toEqual([ + 'MICROVM_APPROVAL_SUSPEND_ENABLED', + 'MICROVM_APPROVAL_SUSPEND_PARAMETER_NAME', 'MICROVM_EGRESS_CONNECTOR_ARNS', 'MICROVM_EXECUTION_ROLE_ARN', 'MICROVM_IMAGE_IDENTIFIER', @@ -733,14 +776,17 @@ describe('TaskOrchestrator with the Lambda MicroVMs backend (ADR-021)', () => { expect(env.MICROVM_INGRESS_CONNECTOR_ARNS).not.toContain('NO_INGRESS'); }); - test('grants exactly the four P1 lifecycle actions and nothing more', () => { + test('grants the lifecycle actions used by the supervisor and image discovery', () => { const actions = microvmStatements(template) .flatMap(s => Array.isArray(s.Action) ? s.Action : [s.Action]) .filter(a => a.startsWith('lambda:')); expect(actions.sort()).toEqual([ 'lambda:GetMicrovm', + 'lambda:GetMicrovmImageVersion', 'lambda:PassNetworkConnector', + 'lambda:ResumeMicrovm', 'lambda:RunMicrovm', + 'lambda:SuspendMicrovm', 'lambda:TerminateMicrovm', ]); }); @@ -753,11 +799,9 @@ describe('TaskOrchestrator with the Lambda MicroVMs backend (ADR-021)', () => { }); test('a name-derived image ARN is scoped to that exact name, never a wildcard', () => { - // An out-of-band image referenced by bare NAME is a valid RunMicrovm - // identifier but not an IAM resource. LambdaMicrovmCompute resolves it to the - // exact `microvmImage` ARN, so ADR-021's "scoped to platform-created images" - // holds here too — the account/Region-wide `microvm-image:*` widening this - // test previously accepted is a compliance violation, not a fallback. + // LambdaMicrovmCompute resolves a configured image name to an exact ARN + // for both RunMicrovm and IAM. Neither path may fall back to an account-wide + // image wildcard; the service does not accept a bare RunMicrovm image name. const lifecycle = microvmStatements(nameDerivedTemplate).find(s => s.Sid === 'MicrovmLifecycle')!; expect(lifecycle.Resource).toEqual([ NAME_DERIVED_IMAGE_ARN, @@ -791,19 +835,19 @@ describe('TaskOrchestrator with the Lambda MicroVMs backend (ADR-021)', () => { expect(JSON.stringify(passRole.Resource)).not.toContain('*'); }); - test('grants NO suspend/resume (P3) and NO auth-token minting (never)', () => { + test('grants supervisor control but no auth-token minting', () => { const actions = new Set( Object.values(template.findResources('AWS::IAM::Policy')) .flatMap(p => p.Properties.PolicyDocument.Statement as Array<{ Action: string | string[] }>) .flatMap(s => Array.isArray(s.Action) ? s.Action : [s.Action]), ); - expect(actions.has('lambda:SuspendMicrovm')).toBe(false); - expect(actions.has('lambda:ResumeMicrovm')).toBe(false); + expect(actions.has('lambda:SuspendMicrovm')).toBe(true); + expect(actions.has('lambda:ResumeMicrovm')).toBe(true); expect(actions.has('lambda:CreateMicrovmAuthToken')).toBe(false); expect(actions.has('lambda:CreateMicrovmShellAuthToken')).toBe(false); }); - test('gets write on the payload bucket but NOT delete (lifecycle rule is the reaper)', () => { + test('can upload payloads and delete only task payload objects at finalize', () => { const payloadStatements = Object.values(template.findResources('AWS::IAM::Policy')) .flatMap(p => p.Properties.PolicyDocument.Statement as Array<{ Action: string | string[]; @@ -813,12 +857,25 @@ describe('TaskOrchestrator with the Lambda MicroVMs backend (ADR-021)', () => { const actions = payloadStatements.flatMap(s => Array.isArray(s.Action) ? s.Action : [s.Action]); expect(actions).toContain('s3:PutObject'); - expect(actions).not.toContain('s3:DeleteObject'); + expect(actions).toContain('s3:DeleteObject'); + expect(actions).toContain('s3:GetObject'); + expect(actions.filter(action => action.startsWith('s3:List'))).toEqual(['s3:ListBucket']); + const list = payloadStatements.find(s => s.Action === 's3:ListBucket'); + expect(list!.Resource).toEqual({ 'Fn::GetAtt': [expect.stringMatching(/^MicrovmPayloadBucket/), 'Arn'] }); + const deletes = payloadStatements.filter(s => + (Array.isArray(s.Action) ? s.Action : [s.Action]).some(action => action.startsWith('s3:Delete'))); + expect(deletes).toHaveLength(1); + expect(deletes[0]!.Action).toEqual(['s3:GetObject', 's3:DeleteObject']); + expect(JSON.stringify(deletes[0]!.Resource)).toContain('/*/payload.json'); + expect(JSON.stringify(deletes[0]!.Resource)).toContain('/*/launch.json'); + expect(JSON.stringify(deletes[0]!.Resource)).not.toContain('/bootstrap/*'); }); test('adds no MicroVM statements when microvmConfig is omitted', () => { expect(microvmStatements(noMicrovmTemplate)).toEqual([]); expect(orchestratorEnv(noMicrovmTemplate).MICROVM_IMAGE_IDENTIFIER).toBeUndefined(); + expect(orchestratorEnv(noMicrovmTemplate).MICROVM_APPROVAL_SUSPEND_PARAMETER_NAME).toBeUndefined(); + noMicrovmTemplate.resourceCountIs('AWS::SSM::Parameter', 0); }); }); @@ -835,7 +892,10 @@ describe('TaskOrchestrator agentPlatformConfig (ADR-021 P2 platform_config trans * NAME, grants nothing" property can be asserted against actual logical IDs * rather than string literals. */ - function createPlatformConfigStack(withConfig: boolean): { template: Template } { + function createPlatformConfigStack( + withConfig: boolean, + approvalRequestsApiUrl = 'https://approval.execute-api.us-east-1.amazonaws.com/v1', + ): { template: Template } { const app = new App(); const stack = new Stack(app, 'TestStack', { env: { account: '123456789012', region: 'us-east-1' }, @@ -857,6 +917,7 @@ describe('TaskOrchestrator agentPlatformConfig (ADR-021 P2 platform_config trans ...(withConfig && { agentPlatformConfig: { taskApprovalsTableName: approvalsTable.tableName, + approvalRequestsApiUrl, nudgesTableName: nudgesTable.tableName, logGroupName: '/aws/abca/application', artifactsBucketName: traceBucket.bucketName, @@ -885,7 +946,11 @@ describe('TaskOrchestrator agentPlatformConfig (ADR-021 P2 platform_config trans withoutConfigTemplate = createPlatformConfigStack(false).template; }); - test('injects the eight forwarded identifiers under the names the strategy reads', () => { + test('rejects platform approval configuration without the service URL', () => { + expect(() => createPlatformConfigStack(true, '')).toThrow('requires approvalRequestsApiUrl'); + }); + + test('injects forwarded identifiers under the names the strategy reads', () => { // These names are a CONTRACT with // `handlers/shared/strategies/lambda-microvm-strategy.ts`'s // PLATFORM_CONFIG_ENV_VARS map, and with the AgentCore runtime env block in @@ -893,6 +958,7 @@ describe('TaskOrchestrator agentPlatformConfig (ADR-021 P2 platform_config trans // side silently strips a key from every MicroVM task's platform_config. const env = orchestratorEnvVars(template); expect(env.TASK_APPROVALS_TABLE_NAME).toEqual({ Ref: expect.stringMatching(/^TaskApprovalsTable/) }); + expect(env.APPROVAL_REQUESTS_API_URL).toBe('https://approval.execute-api.us-east-1.amazonaws.com/v1'); expect(env.NUDGES_TABLE_NAME).toEqual({ Ref: expect.stringMatching(/^TaskNudgesTable/) }); expect(env.LOG_GROUP_NAME).toBe('/aws/abca/application'); expect(env.ARTIFACTS_BUCKET_NAME).toEqual({ Ref: expect.stringMatching(/^TraceArtifactsBucket/) }); @@ -909,14 +975,13 @@ describe('TaskOrchestrator agentPlatformConfig (ADR-021 P2 platform_config trans expect(env.ANTHROPIC_MODEL).toBe(MAIN_PROFILE); }); - test('carries the four REQUIRED platform_config sources together (never a partial set)', () => { - // The strategy refuses to start a lambda-microvm session without these four. - // Three come from the orchestrator's own wiring and one from this block, so - // this is the assertion that they are all reachable from ONE deploy. + test('carries all required platform configuration, including the approval service', () => { + // One deployment must supply the complete configuration required at launch. const env = orchestratorEnvVars(createStack({ githubTokenSecretArn: 'arn:aws:secretsmanager:us-east-1:123456789012:secret:github-token-abc123', agentPlatformConfig: { taskApprovalsTableName: 'approvals', + approvalRequestsApiUrl: 'https://approval.execute-api.us-east-1.amazonaws.com/v1', nudgesTableName: 'nudges', logGroupName: '/aws/abca/application', artifactsBucketName: 'artifacts', @@ -926,6 +991,7 @@ describe('TaskOrchestrator agentPlatformConfig (ADR-021 P2 platform_config trans anthropicModel: MAIN_PROFILE, }, }).template); + expect(env.APPROVAL_REQUESTS_API_URL).toBe('https://approval.execute-api.us-east-1.amazonaws.com/v1'); expect(env.TASK_TABLE_NAME).toBeDefined(); expect(env.TASK_EVENTS_TABLE_NAME).toBeDefined(); expect(env.GITHUB_TOKEN_SECRET_ARN).toBeDefined(); diff --git a/cdk/test/handlers/approve-task.test.ts b/cdk/test/handlers/approve-task.test.ts index f6067059b..c8c12f411 100644 --- a/cdk/test/handlers/approve-task.test.ts +++ b/cdk/test/handlers/approve-task.test.ts @@ -21,6 +21,11 @@ import type { APIGatewayProxyEvent } from 'aws-lambda'; // --- Mocks --- const mockSend = jest.fn(); +const mockWake = jest.fn(); +jest.mock('../../src/handlers/shared/microvm-approval-wake', () => ({ + ...jest.requireActual('../../src/handlers/shared/microvm-approval-wake'), + wakeMicrovmAfterApproval: (...args: unknown[]) => mockWake(...args), +})); // Construct a stub TransactionCanceledException that has the // `err.name` + `CancellationReasons` the handler reads, plus makes @@ -56,7 +61,7 @@ process.env.TASK_APPROVALS_TABLE_NAME = 'Approvals'; process.env.TASK_EVENTS_TABLE_NAME = 'Events'; process.env.APPROVE_RATE_LIMIT_PER_MINUTE = '30'; -import { handler } from '../../src/handlers/approve-task'; +import { handler, recordApprovalForUser } from '../../src/handlers/approve-task'; function makeEvent(overrides: Partial = {}): APIGatewayProxyEvent { return { @@ -92,6 +97,7 @@ function makeEvent(overrides: Partial = {}): APIGatewayPro beforeEach(() => { mockSend.mockReset(); + mockWake.mockReset().mockResolvedValue(undefined); ulidCounter = 0; }); @@ -235,3 +241,81 @@ describe('approve-task — error classification', () => { expect(res.statusCode).toBe(202); }); }); + +describe('postcommit MicroVM wake', () => { + test('wake diagnostics are task-bound and keep the supplied remaining-time signal', async () => { + mockSend.mockResolvedValue({}); + mockWake.mockImplementationOnce(async input => { + await input.emitEvent('microvm_resume_orphan', { stage: 'resume-request' }, input.options); + }); + expect((await handler(makeEvent())).statusCode).toBe(202); + const emitted = mockSend.mock.calls.find(([command]) => command.input.Item?.event_type === 'microvm_resume_orphan'); + expect(emitted?.[0].input.Item).toMatchObject({ + task_id: 'task-1', user_id: 'user-alice', metadata: { stage: 'resume-request' }, + }); + expect(emitted?.[1]).toBe(mockWake.mock.calls[0][0].options); + }); + + test('wake runs after the committed transaction and keeps its decision identity', async () => { + mockSend.mockResolvedValue({}); + const response = await handler(makeEvent()); + expect(response.statusCode).toBe(202); + const transaction = mockSend.mock.calls.findIndex(([command]) => command._type === 'TransactWrite'); + expect(mockSend.mock.invocationCallOrder[transaction]).toBeLessThan(mockWake.mock.invocationCallOrder[0]); + expect(mockWake.mock.calls[0][0]).toMatchObject({ + taskId: 'task-1', + userId: 'user-alice', + decision: 'APPROVED', + options: { abortSignal: expect.any(AbortSignal) }, + }); + }); + test('an unexpected wake failure cannot change the committed202 response', async () => { + mockSend.mockResolvedValue({}); + mockWake.mockRejectedValue(new Error('private wake failure')); + const response = await handler(makeEvent()); + expect(response.statusCode).toBe(202); + expect(response.body).not.toContain('private wake failure'); + expect(mockSend.mock.calls.filter(([command]) => command._type === 'TransactWrite')).toHaveLength(1); + }); + test('transaction failure cannot wake a worker', async () => { + mockSend.mockResolvedValueOnce({}).mockRejectedValueOnce(new Error('transaction failed')); + expect((await handler(makeEvent())).statusCode).toBe(500); + expect(mockWake).not.toHaveBeenCalled(); + }); + test('near Lambda timeout, optional postcommit work gets an already-expired budget', async () => { + mockSend.mockResolvedValue({}); + const response = await handler(makeEvent(), { getRemainingTimeInMillis: () => 500 }); + expect(response.statusCode).toBe(202); + expect(mockSend.mock.calls.map(([command]) => command._type)).toEqual(['Update', 'TransactWrite']); + expect(mockWake.mock.calls[0][0].options.abortSignal.aborted).toBe(true); + }); + test('audit failure still permits wake and does not fail the decision', async () => { + mockSend.mockResolvedValueOnce({}).mockResolvedValueOnce({}).mockRejectedValueOnce(new Error('audit failed')); + expect((await handler(makeEvent())).statusCode).toBe(202); + expect(mockWake).toHaveBeenCalledTimes(1); + }); +}); + +describe('trusted channel decision source', () => { + test('stores the source inside the guarded decision transaction', async () => { + mockSend.mockResolvedValue({}); + const result = await recordApprovalForUser({ + userId: 'user-alice', + taskId: 'task-1', + body: JSON.stringify({ request_id: 'gate', decision: 'approve' }), + decisionSource: 'linear-source', + }); + expect(result.statusCode).toBe(202); + const update = mockSend.mock.calls.find(([cmd]) => cmd._type === 'TransactWrite')![0].input.TransactItems[0].Update; + expect(update.UpdateExpression).toContain('decision_source = :source'); + expect(update.ExpressionAttributeValues[':source']).toBe('linear-source'); + expect(update.ConditionExpression).toContain('#status = :pending'); + expect(update.ConditionExpression).toContain('deadline_epoch > :epoch'); + }); + test('does not trust a source supplied in the HTTP body', async () => { + mockSend.mockResolvedValue({}); + await handler(makeEvent({ body: JSON.stringify({ request_id: 'gate', decision: 'approve', decisionSource: 'forged' }) })); + const update = mockSend.mock.calls.find(([cmd]) => cmd._type === 'TransactWrite')![0].input.TransactItems[0].Update; + expect(update.UpdateExpression).not.toContain('decision_source'); + }); +}); diff --git a/cdk/test/handlers/cancel-task.test.ts b/cdk/test/handlers/cancel-task.test.ts index b64465766..515fce01b 100644 --- a/cdk/test/handlers/cancel-task.test.ts +++ b/cdk/test/handlers/cancel-task.test.ts @@ -42,12 +42,14 @@ jest.mock('@aws-sdk/lib-dynamodb', () => ({ GetCommand: jest.fn((input: unknown) => ({ _type: 'Get', input })), UpdateCommand: jest.fn((input: unknown) => ({ _type: 'Update', input })), PutCommand: jest.fn((input: unknown) => ({ _type: 'Put', input })), + TransactWriteCommand: jest.fn((input: unknown) => ({ _type: 'TransactWrite', input })), })); jest.mock('ulid', () => ({ ulid: jest.fn(() => 'REQ-ULID') })); process.env.TASK_TABLE_NAME = 'Tasks'; process.env.TASK_EVENTS_TABLE_NAME = 'TaskEvents'; +process.env.TASK_APPROVALS_TABLE_NAME = 'Approvals'; process.env.TASK_RETENTION_DAYS = '90'; process.env.RUNTIME_ARN = 'arn:aws:bedrock-agentcore:us-east-1:123456789012:runtime/default'; process.env.ECS_CLUSTER_ARN = 'arn:aws:ecs:us-east-1:123456789012:cluster/agent-cluster'; @@ -127,7 +129,32 @@ beforeEach(() => { }); describe('cancel-task handler', () => { - test('cancels a running task successfully', async () => { + test('closes an unanswered approval atomically before stopping its worker', async () => { + mockSend.mockReset(); + mockSend.mockResolvedValueOnce({ + Item: { + ...RUNNING_TASK, status: 'AWAITING_APPROVAL', awaiting_approval_request_id: 'request-1', + }, + }) + .mockResolvedValueOnce({ Item: { status: 'PENDING', user_id: RUNNING_TASK.user_id } }) + .mockResolvedValueOnce({}) + .mockResolvedValueOnce({}); + const response = await handler(makeEvent()); + expect(response.statusCode).toBe(200); + const transaction = mockSend.mock.calls.find(([command]) => command._type === 'TransactWrite')![0].input; + expect(transaction.TransactItems[0].Update.ExpressionAttributeValues[':cancelled']).toBe('CANCELLED'); + expect(transaction.TransactItems[1].Update.Key.request_id).toBe('request-1'); + expect(transaction.TransactItems[1].Update.ExpressionAttributeValues[':cancelled']).toBe('CANCELLED'); + expect(transaction.TransactItems[2].Put.Item.event_type).toBe('approval_cancelled'); + expect(mockAgentCoreSend).toHaveBeenCalledTimes(1); + }); + + test.each(['HYDRATING', 'RUNNING', 'AWAITING_APPROVAL', 'FINALIZING'])('stops an AgentCore session when cancelling %s', async (status) => { + mockSend.mockReset(); + mockSend + .mockResolvedValueOnce({ Item: { ...RUNNING_TASK, status } }) + .mockResolvedValueOnce({}) + .mockResolvedValueOnce({}); const result = await handler(makeEvent()); expect(result.statusCode).toBe(200); @@ -187,19 +214,26 @@ describe('cancel-task handler', () => { expect(result.statusCode).toBe(409); expect(JSON.parse(result.body).error.code).toBe('TASK_ALREADY_TERMINAL'); + expect(mockAgentCoreSend).not.toHaveBeenCalled(); + expect(mockEcsSend).not.toHaveBeenCalled(); + expect(mockMicrovmSend).not.toHaveBeenCalled(); }); - test('returns 409 on ConditionalCheckFailedException (race condition)', async () => { + test.each(['RUNNING', 'AWAITING_APPROVAL'])('does not stop %s compute when cancellation loses a terminal race', async (status) => { mockSend.mockReset(); - mockSend.mockResolvedValueOnce({ Item: RUNNING_TASK }); + mockSend.mockResolvedValueOnce({ Item: { ...RUNNING_TASK, status, compute_type: 'lambda-microvm' } }); const condError = new Error('Condition not met'); condError.name = 'ConditionalCheckFailedException'; mockSend.mockRejectedValueOnce(condError); + mockSend.mockResolvedValueOnce({ Item: { ...RUNNING_TASK, status: 'COMPLETED' } }); const result = await handler(makeEvent()); expect(result.statusCode).toBe(409); expect(JSON.parse(result.body).error.code).toBe('TASK_ALREADY_TERMINAL'); + expect(mockAgentCoreSend).not.toHaveBeenCalled(); + expect(mockEcsSend).not.toHaveBeenCalled(); + expect(mockMicrovmSend).not.toHaveBeenCalled(); }); test('returns 500 on unexpected DynamoDB error', async () => { @@ -232,21 +266,23 @@ describe('cancel-task handler', () => { expect(typeof eventCall.input.Item.ttl).toBe('number'); }); - test('can cancel tasks in SUBMITTED state', async () => { + test.each(['SUBMITTED', 'QUEUED', 'PENDING_UPLOADS'])('cancels %s without stopping compute', async (status) => { mockSend.mockReset(); mockSend - .mockResolvedValueOnce({ Item: { ...RUNNING_TASK, status: 'SUBMITTED' } }) + .mockResolvedValueOnce({ Item: { ...RUNNING_TASK, status } }) .mockResolvedValueOnce({}) .mockResolvedValueOnce({}); const result = await handler(makeEvent()); expect(result.statusCode).toBe(200); expect(mockAgentCoreSend).not.toHaveBeenCalled(); + expect(mockEcsSend).not.toHaveBeenCalled(); + expect(mockMicrovmSend).not.toHaveBeenCalled(); }); - test('does not call StopRuntimeSession when RUNNING but session_id is missing', async () => { + test.each(['RUNNING', 'AWAITING_APPROVAL'])('does not call StopRuntimeSession when %s has no session_id', async (status) => { mockSend.mockReset(); - const noSession = { ...RUNNING_TASK }; + const noSession = { ...RUNNING_TASK, status }; delete (noSession as { session_id?: string }).session_id; mockSend .mockResolvedValueOnce({ Item: noSession }) @@ -287,10 +323,11 @@ describe('cancel-task handler', () => { expect(mockAgentCoreSend).toHaveBeenCalled(); }); - test('cancels ECS-backed running task via StopTask', async () => { + test.each(['HYDRATING', 'RUNNING', 'AWAITING_APPROVAL', 'FINALIZING'])('stops an ECS session when cancelling %s', async (status) => { mockSend.mockReset(); const ecsTask = { ...RUNNING_TASK, + status, compute_type: 'ecs', compute_metadata: { clusterArn: 'arn:aws:ecs:us-east-1:123456789012:cluster/agent-cluster', @@ -332,7 +369,7 @@ describe('cancel-task handler', () => { expect(mockAgentCoreSend).not.toHaveBeenCalled(); }); - test('cancels a lambda-microvm task via TerminateMicrovm, not AgentCore or ECS', async () => { + test.each(['HYDRATING', 'RUNNING', 'AWAITING_APPROVAL', 'FINALIZING'])('terminates only the MicroVM session when cancelling %s', async (status) => { mockSend.mockReset(); // NOTE: RUNNING_TASK carries `agent_runtime_arn`, and RUNTIME_ARN is also set // in this suite's env — exactly the mixed-deployment shape that would send a @@ -340,6 +377,7 @@ describe('cancel-task handler', () => { // first. This test is the regression guard for that branch ordering. const microvmTask = { ...RUNNING_TASK, + status, compute_type: 'lambda-microvm', session_id: 'mvm-0123456789abcdef', compute_metadata: { @@ -378,10 +416,11 @@ describe('cancel-task handler', () => { expect(mockAgentCoreSend).not.toHaveBeenCalled(); }); - test('a TerminateMicrovm failure still returns 200 (the CANCELLED write stands)', async () => { + test.each(['RUNNING', 'AWAITING_APPROVAL'])('keeps cancellation committed when TerminateMicrovm fails for %s', async (status) => { mockSend.mockReset(); const microvmTask = { ...RUNNING_TASK, + status, compute_type: 'lambda-microvm', compute_metadata: { microvmId: 'mvm-0123456789abcdef', endpoint: 'https://x' }, }; diff --git a/cdk/test/handlers/confirm-uploads.test.ts b/cdk/test/handlers/confirm-uploads.test.ts index dfa211fc6..83b37c9ef 100644 --- a/cdk/test/handlers/confirm-uploads.test.ts +++ b/cdk/test/handlers/confirm-uploads.test.ts @@ -226,9 +226,8 @@ describe('confirm-uploads handler', () => { switch (ddbCallCount) { case 1: return Promise.resolve({ Item: PENDING_TASK }); // GetCommand (task) case 2: return Promise.resolve({ Item: { active_count: 1 } }); // GetCommand (concurrency pre-check) - case 3: return Promise.resolve({}); // UpdateCommand (checkConcurrency) - case 4: return Promise.resolve({}); // UpdateCommand (status transition) - case 5: return Promise.resolve({}); // PutCommand (event) + case 3: return Promise.resolve({}); // UpdateCommand (status transition) + case 4: return Promise.resolve({}); // PutCommand (event) default: return Promise.resolve({}); } }); @@ -265,6 +264,9 @@ describe('confirm-uploads handler', () => { const body = JSON.parse(result.body); expect(body.data.status).toBe('SUBMITTED'); expect(lambdaSend).toHaveBeenCalled(); + const capacityCalls = ddbSend.mock.calls.filter(([command]) => command.input.TableName === 'Concurrency'); + expect(capacityCalls).toHaveLength(1); + expect(capacityCalls[0][0]._type).toBe('Get'); }); test('returns 429 when concurrency pre-check fails', async () => { diff --git a/cdk/test/handlers/deny-task.test.ts b/cdk/test/handlers/deny-task.test.ts index 4fabae92e..96004f043 100644 --- a/cdk/test/handlers/deny-task.test.ts +++ b/cdk/test/handlers/deny-task.test.ts @@ -20,6 +20,11 @@ import type { APIGatewayProxyEvent } from 'aws-lambda'; const mockSend = jest.fn(); +const mockWake = jest.fn(); +jest.mock('../../src/handlers/shared/microvm-approval-wake', () => ({ + ...jest.requireActual('../../src/handlers/shared/microvm-approval-wake'), + wakeMicrovmAfterApproval: (...args: unknown[]) => mockWake(...args), +})); class MockTransactionCanceledException extends Error { name = 'TransactionCanceledException'; @@ -48,7 +53,7 @@ process.env.TASK_TABLE_NAME = 'Tasks'; process.env.TASK_APPROVALS_TABLE_NAME = 'Approvals'; process.env.TASK_EVENTS_TABLE_NAME = 'Events'; -import { handler } from '../../src/handlers/deny-task'; +import { handler, recordDenialForUser } from '../../src/handlers/deny-task'; // Secret fixtures assembled at runtime so the source file itself // never holds a contiguous secret literal (Code Defender pre-commit @@ -90,6 +95,7 @@ function makeEvent(overrides: Partial = {}): APIGatewayPro beforeEach(() => { mockSend.mockReset(); + mockWake.mockReset().mockResolvedValue(undefined); ulidCounter = 0; }); @@ -204,3 +210,81 @@ describe('deny-task — error classification', () => { expect(body.data.request_id).toBe('01KREQ'); }); }); + +describe('postcommit MicroVM wake', () => { + test('wake diagnostics are task-bound and keep the supplied remaining-time signal', async () => { + mockSend.mockResolvedValue({}); + mockWake.mockImplementationOnce(async input => { + await input.emitEvent('microvm_resume_orphan', { stage: 'resume-request' }, input.options); + }); + expect((await handler(makeEvent())).statusCode).toBe(202); + const emitted = mockSend.mock.calls.find(([command]) => command.input.Item?.event_type === 'microvm_resume_orphan'); + expect(emitted?.[0].input.Item).toMatchObject({ + task_id: 'task-1', user_id: 'user-alice', metadata: { stage: 'resume-request' }, + }); + expect(emitted?.[1]).toBe(mockWake.mock.calls[0][0].options); + }); + + test('wake runs after the committed transaction and keeps its decision identity', async () => { + mockSend.mockResolvedValue({}); + const response = await handler(makeEvent()); + expect(response.statusCode).toBe(202); + const transaction = mockSend.mock.calls.findIndex(([command]) => command._type === 'TransactWrite'); + expect(mockSend.mock.invocationCallOrder[transaction]).toBeLessThan(mockWake.mock.invocationCallOrder[0]); + expect(mockWake.mock.calls[0][0]).toMatchObject({ + taskId: 'task-1', + userId: 'user-alice', + decision: 'DENIED', + options: { abortSignal: expect.any(AbortSignal) }, + }); + }); + test('an unexpected wake failure cannot change the committed202 response', async () => { + mockSend.mockResolvedValue({}); + mockWake.mockRejectedValue(new Error('private wake failure')); + const response = await handler(makeEvent()); + expect(response.statusCode).toBe(202); + expect(response.body).not.toContain('private wake failure'); + expect(mockSend.mock.calls.filter(([command]) => command._type === 'TransactWrite')).toHaveLength(1); + }); + test('transaction failure cannot wake a worker', async () => { + mockSend.mockResolvedValueOnce({}).mockRejectedValueOnce(new Error('transaction failed')); + expect((await handler(makeEvent())).statusCode).toBe(500); + expect(mockWake).not.toHaveBeenCalled(); + }); + test('near Lambda timeout, optional postcommit work gets an already-expired budget', async () => { + mockSend.mockResolvedValue({}); + const response = await handler(makeEvent(), { getRemainingTimeInMillis: () => 500 }); + expect(response.statusCode).toBe(202); + expect(mockSend.mock.calls.map(([command]) => command._type)).toEqual(['Update', 'TransactWrite']); + expect(mockWake.mock.calls[0][0].options.abortSignal.aborted).toBe(true); + }); + test('audit failure still permits wake and does not fail the decision', async () => { + mockSend.mockResolvedValueOnce({}).mockResolvedValueOnce({}).mockRejectedValueOnce(new Error('audit failed')); + expect((await handler(makeEvent())).statusCode).toBe(202); + expect(mockWake).toHaveBeenCalledTimes(1); + }); +}); + +describe('trusted channel decision source', () => { + test('stores the source inside the guarded decision transaction', async () => { + mockSend.mockResolvedValue({}); + const result = await recordDenialForUser({ + userId: 'user-alice', + taskId: 'task-1', + body: JSON.stringify({ request_id: 'gate', decision: 'deny' }), + decisionSource: 'linear-source', + }); + expect(result.statusCode).toBe(202); + const update = mockSend.mock.calls.find(([cmd]) => cmd._type === 'TransactWrite')![0].input.TransactItems[0].Update; + expect(update.UpdateExpression).toContain('decision_source = :source'); + expect(update.ExpressionAttributeValues[':source']).toBe('linear-source'); + expect(update.ConditionExpression).toContain('#status = :pending'); + expect(update.ConditionExpression).toContain('deadline_epoch > :epoch'); + }); + test('does not trust a source supplied in the HTTP body', async () => { + mockSend.mockResolvedValue({}); + await handler(makeEvent({ body: JSON.stringify({ request_id: 'gate', decision: 'deny', decisionSource: 'forged' }) })); + const update = mockSend.mock.calls.find(([cmd]) => cmd._type === 'TransactWrite')![0].input.TransactItems[0].Update; + expect(update.UpdateExpression).not.toContain('decision_source'); + }); +}); diff --git a/cdk/test/handlers/fanout-task-events.test.ts b/cdk/test/handlers/fanout-task-events.test.ts index 78125b2bf..9a0ea81fa 100644 --- a/cdk/test/handlers/fanout-task-events.test.ts +++ b/cdk/test/handlers/fanout-task-events.test.ts @@ -100,6 +100,7 @@ jest.mock('../../src/handlers/slack-notify', () => { // + GraphQL path. Default ``{ ok: true }`` so a test that forgets to // script the mock still drives the happy path (postIssueComment returns // a LinearPostResult, not a bare boolean). +const mockPostIdentifiedComment: jest.Mock = jest.fn().mockResolvedValue({ ok: true }); const mockPostIssueComment: jest.Mock = jest.fn().mockResolvedValue({ ok: true }); // Standalone comment-triggered iterations get a threaded reply to // the human's @bgagent comment, on top of the metrics comment. replyToComment @@ -110,6 +111,7 @@ const mockReplyToComment: jest.Mock = jest.fn().mockResolvedValue('reply-id'); // rather than posting a fresh replyToComment. const mockUpsertThreadedReply: jest.Mock = jest.fn().mockResolvedValue('reply-id'); jest.mock('../../src/handlers/shared/linear-feedback', () => ({ + postIdentifiedComment: (...args: unknown[]) => mockPostIdentifiedComment(...args), postIssueComment: ( ctx: { linearWorkspaceId: string; registryTableName: string }, issueId: string, @@ -181,6 +183,7 @@ jest.mock('../../src/handlers/shared/jira-feedback', () => ({ process.env.TASK_TABLE_NAME = 'Tasks'; process.env.GITHUB_TOKEN_SECRET_ARN = 'arn:aws:secretsmanager:us-east-1:0:secret:platform'; process.env.LINEAR_WORKSPACE_REGISTRY_TABLE_NAME = 'LinearWorkspaceRegistry'; +process.env.TASK_APPROVALS_TABLE_NAME = 'Approvals'; process.env.JIRA_WORKSPACE_REGISTRY_TABLE_NAME = 'JiraWorkspaceRegistry'; /** Flatten the stubbed ADF (`{ _adf: paragraphs }`) back to a newline-joined @@ -208,6 +211,8 @@ import { renderJiraFinishedPointer, } from '../../src/handlers/shared/jira-status-comment'; +let streamSequence = 1; + function mkRecord( eventName: 'INSERT' | 'MODIFY' | 'REMOVE', newImage: Record }> | undefined, @@ -216,7 +221,7 @@ function mkRecord( eventID: `evt-${Math.random().toString(36).slice(2)}`, eventName, eventSource: 'aws:dynamodb', - dynamodb: newImage ? { NewImage: newImage as never } : {}, + dynamodb: { SequenceNumber: String(streamSequence++), ...(newImage ? { NewImage: newImage as never } : {}) }, } as unknown as DynamoDBRecord; } @@ -324,8 +329,11 @@ describe('fanout-task-events: per-channel filter contract (design §6.2)', () => const f = CHANNEL_DEFAULTS.slack; expect([...f].sort()).toEqual([ 'agent_error', + 'approval_cancelled', + 'approval_decision_recorded', 'approval_requested', 'approval_stranded', + 'approval_timed_out', 'session_started', 'status_response', 'task_cancelled', @@ -344,23 +352,12 @@ describe('fanout-task-events: per-channel filter contract (design §6.2)', () => }); test('every Slack-default event the dispatcher actually renders today is in NOTIFIABLE_EVENTS (drift guard)', () => { - // The router subscribes Slack to events the dispatcher must - // render. ``approval_requested``, ``approval_stranded``, and - // ``status_response`` are forward-compat (no Slack-side renderer - // today — the CLI surfaces approval UX; Slack is only in the - // channel-defaults set so a future Slack-button renderer can - // light up without changing the router filter). They're allowed - // to be in CHANNEL_DEFAULTS.slack but absent from - // NOTIFIABLE_EVENTS — when their emitters land, this test will - // start failing and force the dispatcher update at the same time. - // Every OTHER Slack default must be renderable, otherwise - // telemetry lies. Use ``requireActual`` to bypass the - // slack-notify mock and read the real exported NOTIFIABLE_EVENTS - // set. + // Approval emitters are live and must have a renderer. Only status_response + // remains a placeholder. Read the actual module despite the dispatcher mock. const real = jest.requireActual( '../../src/handlers/slack-notify', ); - const forwardCompat = new Set(['approval_requested', 'approval_stranded', 'status_response']); + const forwardCompat = new Set(['status_response']); const expectedRenderable = [...CHANNEL_DEFAULTS.slack].filter( e => !forwardCompat.has(e), ); @@ -395,11 +392,16 @@ describe('fanout-task-events: per-channel filter contract (design §6.2)', () => ]); }); - test('Linear subscribes to pr_created + terminal events + task_timed_out (ADR-016 P4.5 courtesy comment + post-once final-status)', () => { + test('Linear subscribes to approvals, pr_created and terminal events', () => { // review should-fix: task_timed_out added so a Linear standalone iteration // that times out still settles (matches Jira/Slack, which already had it). const f = CHANNEL_DEFAULTS.linear; expect([...f].sort()).toEqual([ + 'approval_cancelled', + 'approval_decision_recorded', + 'approval_requested', + 'approval_stranded', + 'approval_timed_out', 'pr_created', 'task_cancelled', 'task_completed', @@ -780,6 +782,7 @@ describe('fanout-task-events: GitHub dispatcher (Chunk J)', () => { // task short-circuits inside the dispatcher (channel_source === // 'api' / 'github'). Pre-existing tests don't assert on it. mockPostIssueComment.mockReset().mockResolvedValue({ ok: true }); + mockPostIdentifiedComment.mockReset().mockResolvedValue({ ok: true }); }); test('first terminal event POSTs a new comment and persists the comment_id to TaskTable', async () => { @@ -906,7 +909,7 @@ describe('fanout-task-events: GitHub dispatcher (Chunk J)', () => { const event = { Records: [mkEvent('task_completed', 't-gh')] } as DynamoDBStreamEvent; const result = await handler(event); expect(result.batchItemFailures).toHaveLength(1); - expect(result.batchItemFailures[0].itemIdentifier).toBe(event.Records[0].eventID); + expect(result.batchItemFailures[0].itemIdentifier).toBe(event.Records[0].dynamodb!.SequenceNumber); // No UpdateCommand fires (no id to persist from a failed upsert). const updateCalls = mockDdbSend.mock.calls.filter( c => (c[0] as { _type?: string })._type === 'Update', @@ -1071,7 +1074,7 @@ describe('fanout-task-events: GitHub dispatcher (Chunk J)', () => { // Record IS in batchItemFailures — Lambda will replay until // the rate-limit window opens. Critical: the swallow-as-terminal // path would have produced an empty array (silent drop). - expect(result.batchItemFailures).toEqual([{ itemIdentifier: record.eventID }]); + expect(result.batchItemFailures).toEqual([{ itemIdentifier: record.dynamodb!.SequenceNumber }]); }, ); @@ -1399,7 +1402,7 @@ describe('fanout-task-events: Slack dispatcher', () => { const record = mkEvent('task_completed', 't-slack-fail'); const result = await handler({ Records: [record] }); - expect(result.batchItemFailures).toEqual([{ itemIdentifier: record.eventID }]); + expect(result.batchItemFailures).toEqual([{ itemIdentifier: record.dynamodb!.SequenceNumber }]); }); test('Slack dispatcher SlackApiError swallow does NOT escalate to retry', async () => { @@ -1474,6 +1477,7 @@ describe('fanout-task-events: Linear dispatcher', () => { beforeEach(() => { mockDdbSend.mockReset().mockResolvedValue({ Item: undefined }); mockPostIssueComment.mockReset().mockResolvedValue({ ok: true }); + mockPostIdentifiedComment.mockReset().mockResolvedValue({ ok: true }); mockReplyToComment.mockReset().mockResolvedValue('reply-id'); mockUpsertThreadedReply.mockReset().mockResolvedValue('reply-id'); // Slack/GitHub mocks aren't asserted here but leaving them @@ -1499,6 +1503,77 @@ describe('fanout-task-events: Linear dispatcher', () => { }); }; + test.each([false, true])('routes wrapped approval to Linear without settling the task (post failure=%s)', async fails => { + mockDdbSend.mockImplementation(async command => { + if (command._type !== 'Get') return {}; + if (command.input.TableName === 'Approvals') { + return { + Item: { + user_id: 'u-1', + status: 'PENDING', + tool_name: 'Bash', + severity: 'medium', + reason: 'Review this', + tool_input_preview: 'git push', + timeout_s: 1800, + created_at: '2026-09-16T12:00:00Z', + }, + }; + } + return { + Item: { + ...TASK_RECORD_LINEAR, + status: 'AWAITING_APPROVAL', + awaiting_approval_request_id: 'g1', + channel_metadata: { ...TASK_RECORD_LINEAR.channel_metadata, trigger_comment_id: 'trigger', iteration_reply_comment_id: 'reply' }, + }, + }; + }); + if (fails) mockPostIdentifiedComment.mockResolvedValueOnce({ ok: false, retryable: true }); + const outcome = await routeEvent({ + task_id: 't-lin', + event_id: 'approval-event', + event_type: 'agent_milestone', + timestamp: '2026-09-16T12:00:00Z', + metadata: { milestone: 'approval_requested', request_id: 'g1' }, + }); + expect(mockPostIdentifiedComment).toHaveBeenCalledTimes(1); + expect(mockPostIdentifiedComment.mock.calls[0][1].body).toContain('bgagent approve t-lin g1 --scope this_call'); + expect(mockUpsertThreadedReply).not.toHaveBeenCalled(); + const updates = mockDdbSend.mock.calls.filter(([c]) => c._type === 'Update'); + if (fails) { + expect(outcome.infraRejections).toHaveLength(1); + expect(updates).toHaveLength(1); + } else { + expect(outcome.infraRejections).toHaveLength(0); + expect(updates).toHaveLength(2); + expect(updates[1][0].input.ExpressionAttributeNames).toEqual({ '#marker': 'notified_linear_approval_requested' }); + } + }); + + test('starts binding retention when a closed-request notice cannot be delivered', async () => { + mockDdbSend.mockImplementation(async command => { + if (command._type !== 'Get') return {}; + return { + Item: command.input.TableName === 'Approvals' + ? { user_id: 'u-1', status: 'CANCELLED', cancellation_reason: 'Cancelled by owner' } + : { ...TASK_RECORD_LINEAR, status: 'CANCELLED', awaiting_approval_request_id: 'g1' }, + }; + }); + mockPostIssueComment.mockResolvedValueOnce({ ok: false, retryable: false }); + await routeEvent({ + task_id: 't-lin', + event_id: 'closed', + event_type: 'approval_cancelled', + timestamp: '2026-09-18T12:00:00Z', + metadata: { request_id: 'g1' }, + }); + const updates = mockDdbSend.mock.calls.filter(([c]) => c._type === 'Update'); + expect(updates).toHaveLength(1); + expect(updates[0][0].input.Key.task_id).toContain('LINEAR_COMMENT#'); + expect(updates[0][0].input.UpdateExpression).toContain('if_not_exists(#ttl'); + }); + test('task_completed posts ✅ comment with cost / turns / duration on linked Linear issue', async () => { mockGet(TASK_RECORD_LINEAR); @@ -1628,7 +1703,7 @@ describe('fanout-task-events: Linear dispatcher', () => { // Transient failure → record enters batchItemFailures so Lambda retries. expect(result.batchItemFailures).toHaveLength(1); - expect(result.batchItemFailures[0].itemIdentifier).toBe(event.Records[0].eventID); + expect(result.batchItemFailures[0].itemIdentifier).toBe(event.Records[0].dynamodb!.SequenceNumber); // Marker NOT persisted — the retry will re-post. const updateCalls = mockDdbSend.mock.calls.filter((c) => (c[0] as { _type?: string })._type === 'Update'); expect(updateCalls).toHaveLength(0); @@ -1733,7 +1808,7 @@ describe('fanout-task-events: Linear dispatcher', () => { expect(mockPostIssueComment).toHaveBeenCalledTimes(1); expect(result.batchItemFailures).toHaveLength(1); - expect(result.batchItemFailures[0]).toEqual({ itemIdentifier: records[0].eventID }); + expect(result.batchItemFailures[0]).toEqual({ itemIdentifier: records[0].dynamodb!.SequenceNumber }); // And no marker write: the retry must be allowed to post. const updates = mockDdbSend.mock.calls @@ -1918,8 +1993,8 @@ describe('fanout-task-events: Linear dispatcher', () => { // task — so the thumbnail renders without depending on the racy comment edit. const PNG = 'https://cdn.example/screenshots/iter.png'; const DEPLOY = 'https://app.vercel.app'; - mockDdbSend.mockReset().mockImplementation((cmd: { _type?: string; input?: { ConsistentRead?: boolean } }) => { - if (cmd?._type === 'Get' && cmd.input?.ConsistentRead) { + mockDdbSend.mockReset().mockImplementation((cmd: { _type?: string; input?: { ConsistentRead?: boolean; ProjectionExpression?: string } }) => { + if (cmd?._type === 'Get' && cmd.input?.ConsistentRead && cmd.input.ProjectionExpression === 'screenshot_url, screenshot_preview_url') { // the late re-read: screenshot has landed durably by now return Promise.resolve({ Item: { screenshot_url: PNG, screenshot_preview_url: DEPLOY } }); } @@ -2248,6 +2323,7 @@ describe('fanout-task-events: Jira dispatcher', () => { mockLoadRepoConfig.mockReset().mockResolvedValue(null); mockResolveGitHubToken.mockReset().mockResolvedValue('ghp_fake'); mockPostIssueComment.mockReset().mockResolvedValue({ ok: true }); + mockPostIdentifiedComment.mockReset().mockResolvedValue({ ok: true }); }); const mockGet = (item: unknown) => { @@ -2476,7 +2552,7 @@ describe('fanout-task-events: Jira dispatcher', () => { const result = await handler({ Records: [record] }); expect(result.batchItemFailures).toEqual([ - { itemIdentifier: record.eventID }, + { itemIdentifier: record.dynamodb!.SequenceNumber }, ]); const releases = mockDdbSend.mock.calls .map(([command]) => command as { @@ -2541,7 +2617,7 @@ describe('fanout-task-events: Jira dispatcher', () => { const result = await handler({ Records: [record] }); - expect(result.batchItemFailures).toEqual([{ itemIdentifier: record.eventID }]); + expect(result.batchItemFailures).toEqual([{ itemIdentifier: record.dynamodb!.SequenceNumber }]); expect(mockUpdateIssueCommentAdf).toHaveBeenCalledTimes(1); const releases = mockDdbSend.mock.calls .map(([command]) => command as { @@ -2635,7 +2711,7 @@ describe('fanout-task-events: Jira dispatcher', () => { expect(mockPostIssueCommentAdf).toHaveBeenCalledTimes(1); expect(result.batchItemFailures).toHaveLength(1); - expect(result.batchItemFailures[0]).toEqual({ itemIdentifier: records[0].eventID }); + expect(result.batchItemFailures[0]).toEqual({ itemIdentifier: records[0].dynamodb!.SequenceNumber }); // No marker write — the retry must be allowed to post. const updates = mockDdbSend.mock.calls @@ -3035,10 +3111,8 @@ describe('fanout-task-events: agent_milestone routing (effective event type)', ( // --------------------------------------------------------------------------- /** - * Stream record with a caller-supplied ``eventID`` so the test can - * assert which record surfaces in ``batchItemFailures``. ``mkEvent`` - * uses ``Math.random()`` for the id which is fine for parse tests but - * useless when we need to cross-reference the failure identifier. + * Stream record with distinct event and sequence identifiers. The event ID + * identifies log entries; only the numeric sequence is a valid retry cursor. */ function mkEventWithId(type: string, eventID: string, taskId = 't-fail'): DynamoDBRecord { return { @@ -3046,6 +3120,7 @@ function mkEventWithId(type: string, eventID: string, taskId = 't-fail'): Dynamo eventName: 'INSERT', eventSource: 'aws:dynamodb', dynamodb: { + SequenceNumber: String(streamSequence++), NewImage: { task_id: { S: taskId }, event_id: { S: `01ABC${type}` }, @@ -3061,9 +3136,8 @@ describe('fanout-task-events: partial-batch response', () => { // The construct sets ``reportBatchItemFailures: true`` on // the event-source-mapping, so the handler must return a batch response // rather than ``void``. Returning ``void`` makes Lambda retry the WHOLE - // batch on any unhandled throw — replaying every sibling event and - // defeating the per-task ordering guarantee promised upstream by - // ``ParallelizationFactor: 1``. + // batch on an unhandled throw. A sequence cursor avoids replaying + // successful earlier records; later records still need deduplication. // // The architecturally reachable poison-pill path is a // throw that bypasses ``routeEvent``'s ``Promise.allSettled``. The @@ -3079,6 +3153,7 @@ describe('fanout-task-events: partial-batch response', () => { beforeEach(() => { mockDdbSend.mockReset().mockResolvedValue({ Item: undefined }); + mockDispatchSlackEvent.mockReset().mockResolvedValue(undefined); mockUpsertTaskComment.mockReset(); mockRenderCommentBody.mockReset().mockReturnValue('rendered body'); mockLoadRepoConfig.mockReset().mockResolvedValue(null); @@ -3086,6 +3161,25 @@ describe('fanout-task-events: partial-batch response', () => { mockClearTokenCache.mockReset(); }); + test.each([true, false])('returns the AWS sequence cursor when eventID is present=%s', async hasEventId => { + const record = mkEventWithId('task_created', 'opaque-event-id'); + record.dynamodb!.SequenceNumber = '400000000000000000000001'; + if (!hasEventId) delete record.eventID; + mockDispatchSlackEvent.mockRejectedValueOnce(new Error('temporary delivery failure')); + await expect(handler({ Records: [record] })).resolves.toEqual({ + batchItemFailures: [{ itemIdentifier: '400000000000000000000001' }], + }); + }); + + test.each([undefined, ''])('rejects the batch when a failed record has an invalid sequence (%s)', async sequence => { + const record = mkEventWithId('task_created', 'opaque-event-id'); + record.dynamodb!.SequenceNumber = sequence; + mockDispatchSlackEvent.mockRejectedValueOnce(new Error('temporary delivery failure')); + await expect(handler({ Records: [record] })).rejects.toThrow( + 'Failed DynamoDB record is missing its sequence number', + ); + }); + test('AccessDeniedException from resolveTokenSecretArn lands in infraRejections and flags the record for retry', async () => { // An earlier version swallowed this rejection inside // ``Promise.allSettled``, which silently dropped transient infra @@ -3125,8 +3219,8 @@ describe('fanout-task-events: partial-batch response', () => { const result = await handler(event); // Record is flagged for partial-batch retry — Lambda will replay - // this single eventID, leaving siblings alone. - expect(result.batchItemFailures).toEqual([{ itemIdentifier: poisonId }]); + // from its sequence onward, including any successful later records. + expect(result.batchItemFailures).toEqual([{ itemIdentifier: event.Records[0].dynamodb!.SequenceNumber }]); // The rejection is observable through the dispatcher-rejected // warn so operators can alarm distinctly from the generic @@ -3147,11 +3241,8 @@ describe('fanout-task-events: partial-batch response', () => { // loop throws past ``routeEvent``'s containment (simulated here by // making ``logger.warn`` throw on the rate-limit path — the // closest real non-``routeEvent`` code path), the handler's - // per-record try/catch must push the record's ``eventID`` into - // ``batchItemFailures`` so Lambda retries ONLY that record. A handler - // that returned void would make Lambda retry the ENTIRE - // batch, replaying every sibling event and defeating per-task - // ordering. + // per-record try/catch must return its ``SequenceNumber`` so Lambda + // retries from the failed record onward. const loggerModule = await import('../../src/handlers/shared/logger'); // Rate-limit warn on the 21st event throws; earlier events succeed. let warnCalls = 0; @@ -3191,7 +3282,7 @@ describe('fanout-task-events: partial-batch response', () => { // from the handler's perspective (``routeEvent`` short-circuits // on "task not found" since the shared DDB mock returns no Item). expect(result.batchItemFailures).toEqual([ - { itemIdentifier: 'evt-20' }, + { itemIdentifier: records[20].dynamodb!.SequenceNumber }, ]); } finally { warnSpy.mockRestore(); @@ -3202,7 +3293,7 @@ describe('fanout-task-events: partial-batch response', () => { // Mixed batch: one record throws past routeEvent (via the same // rate-limit-warn trick as above but in a simpler shape — we make // the second record specifically trigger the throw), the other - // routes cleanly. The response must list ONLY the failing eventID. + // routes cleanly. The response must list only the failed sequence. const loggerModule = await import('../../src/handlers/shared/logger'); const warnSpy = jest.spyOn(loggerModule.logger, 'warn').mockImplementation( (_msg: string, meta?: Record) => { @@ -3222,9 +3313,9 @@ describe('fanout-task-events: partial-batch response', () => { const result = await handler({ Records: records }); expect(result.batchItemFailures).toHaveLength(1); - expect(result.batchItemFailures[0]).toEqual({ itemIdentifier: 'evt-chatty-20' }); + expect(result.batchItemFailures[0]).toEqual({ itemIdentifier: records[21].dynamodb!.SequenceNumber }); // Specifically NOT the successful record. - expect(result.batchItemFailures.map(f => f.itemIdentifier)).not.toContain('evt-ok'); + expect(result.batchItemFailures.map(f => f.itemIdentifier)).not.toContain(records[0].dynamodb!.SequenceNumber); } finally { warnSpy.mockRestore(); } diff --git a/cdk/test/handlers/get-pending.test.ts b/cdk/test/handlers/get-pending.test.ts index b19b438c5..cb19b6908 100644 --- a/cdk/test/handlers/get-pending.test.ts +++ b/cdk/test/handlers/get-pending.test.ts @@ -27,6 +27,7 @@ jest.mock('@aws-sdk/client-dynamodb', () => ({ jest.mock('@aws-sdk/lib-dynamodb', () => ({ DynamoDBDocumentClient: { from: jest.fn(() => ({ send: mockSend })) }, QueryCommand: jest.fn((input: unknown) => ({ _type: 'Query', input })), + BatchGetCommand: jest.fn((input: unknown) => ({ _type: 'BatchGet', input })), UpdateCommand: jest.fn((input: unknown) => ({ _type: 'Update', input })), })); @@ -34,6 +35,7 @@ let ulidCounter = 0; jest.mock('ulid', () => ({ ulid: jest.fn(() => `ULID${ulidCounter++}`) })); process.env.TASK_APPROVALS_TABLE_NAME = 'Approvals'; +process.env.TASK_TABLE_NAME = 'Tasks'; process.env.PENDING_RATE_LIMIT_PER_MINUTE = '10'; import { handler } from '../../src/handlers/get-pending'; @@ -76,21 +78,138 @@ beforeEach(() => { }); /** - * Two-mock chain used by every test that returns a non-empty pending - * list: (1) rate-limit Update passes, (2) GSI Query returns ``items``. - * - * ``matching_rule_ids`` is projected onto the ``user_id-status-index`` - * GSI directly (see the construct), so the handler maps each field - * from the Query result without a second read — keeping the mock - * chain short. + * GSI metadata plus strongly consistent reads of the owning task states. */ -function setupPendingMocks(items: ReadonlyArray>): void { +function setupPendingMocks( + items: ReadonlyArray>, + tasks = items.map(row => ({ + task_id: row.task_id, + user_id: 'user-alice', + status: 'AWAITING_APPROVAL', + awaiting_approval_request_id: row.request_id, + })), +): void { mockSend .mockResolvedValueOnce({}) // rate-limit - .mockResolvedValueOnce({ Items: items }); + .mockResolvedValueOnce({ Items: items }) + .mockResolvedValueOnce({ Responses: { Tasks: tasks } }); } describe('get-pending', () => { + test('finds a current request after a full page of cancelled legacy requests', async () => { + const oldRows = Array.from({ length: 100 }, (_, i) => ({ task_id: `old-${i}`, request_id: 'r' })); + const lastKey = { task_id: 'old-99', request_id: 'r', user_id: 'user-alice', status: 'PENDING' }; + mockSend.mockResolvedValueOnce({}) + .mockResolvedValueOnce({ Items: oldRows, LastEvaluatedKey: lastKey }) + .mockResolvedValueOnce({ + Responses: { + Tasks: oldRows.map(row => ({ + ...row, user_id: 'user-alice', status: 'CANCELLED', + })), + }, + }) + .mockResolvedValueOnce({ Items: [{ task_id: 'current', request_id: 'new' }] }) + .mockResolvedValueOnce({ + Responses: { + Tasks: [{ + task_id: 'current', + user_id: 'user-alice', + status: 'AWAITING_APPROVAL', + awaiting_approval_request_id: 'new', + }], + }, + }); + + const response = await handler(makeEvent()); + expect(response.statusCode).toBe(200); + expect(JSON.parse(response.body).data.pending).toEqual([ + expect.objectContaining({ task_id: 'current', request_id: 'new' }), + ]); + const queries = mockSend.mock.calls.filter(([command]) => command._type === 'Query'); + expect(queries).toHaveLength(2); + expect(queries[1][0].input.ExclusiveStartKey).toEqual(lastKey); + expect(queries[0][1].abortSignal).toBe(queries[1][1].abortSignal); + }); + + test('stops paging at the display limit even when the index has more rows', async () => { + const rows = Array.from({ length: 100 }, (_, i) => ({ task_id: `live-${i}`, request_id: 'r' })); + mockSend.mockResolvedValueOnce({}) + .mockResolvedValueOnce({ Items: rows, LastEvaluatedKey: { task_id: 'more' } }) + .mockResolvedValueOnce({ + Responses: { + Tasks: rows.map(row => ({ + task_id: row.task_id, + user_id: 'user-alice', + status: 'AWAITING_APPROVAL', + awaiting_approval_request_id: 'r', + })), + }, + }); + const response = await handler(makeEvent()); + expect(response.statusCode).toBe(200); + expect(JSON.parse(response.body).data.pending).toHaveLength(100); + expect(mockSend.mock.calls.filter(([command]) => command._type === 'Query')).toHaveLength(1); + }); + + test('reports a later-page read failure instead of returning a misleading empty list', async () => { + mockSend.mockResolvedValueOnce({}) + .mockResolvedValueOnce({ Items: [], LastEvaluatedKey: { task_id: 'more' } }) + .mockRejectedValueOnce(new Error('Later page unavailable')); + const response = await handler(makeEvent()); + expect(response.statusCode).toBe(500); + }); + + test.each(['CANCELLED', 'COMPLETED', 'FAILED', 'TIMED_OUT', 'RUNNING'])( + 'omits a legacy pending approval when its task is %s', + async (status) => { + setupPendingMocks([{ task_id: 't', request_id: 'r' }], [{ + task_id: 't', user_id: 'user-alice', status, awaiting_approval_request_id: 'r', + }]); + const response = await handler(makeEvent()); + expect(response.statusCode).toBe(200); + expect(JSON.parse(response.body).data.pending).toEqual([]); + const batch = mockSend.mock.calls.find(([command]) => command._type === 'BatchGet')![0].input; + expect(batch.RequestItems.Tasks.ConsistentRead).toBe(true); + }, + ); + + test('omits a replaced gate, missing task and mismatched owner', async () => { + setupPendingMocks([ + { task_id: 'old', request_id: 'r' }, + { task_id: 'missing', request_id: 'r' }, + { task_id: 'foreign', request_id: 'r' }, + ], [ + { task_id: 'old', user_id: 'user-alice', status: 'AWAITING_APPROVAL', awaiting_approval_request_id: 'new' }, + { task_id: 'foreign', user_id: 'someone-else', status: 'AWAITING_APPROVAL', awaiting_approval_request_id: 'r' }, + ]); + const response = await handler(makeEvent()); + expect(JSON.parse(response.body).data.pending).toEqual([]); + }); + + test('retries unprocessed task reads instead of dropping the request', async () => { + mockSend.mockResolvedValueOnce({}) + .mockResolvedValueOnce({ Items: [{ task_id: 't', request_id: 'r' }] }) + .mockResolvedValueOnce({ UnprocessedKeys: { Tasks: { Keys: [{ task_id: 't' }] } } }) + .mockResolvedValueOnce({ + Responses: { + Tasks: [{ + task_id: 't', user_id: 'user-alice', status: 'AWAITING_APPROVAL', awaiting_approval_request_id: 'r', + }], + }, + }); + const response = await handler(makeEvent()); + expect(response.statusCode).toBe(200); + expect(JSON.parse(response.body).data.pending).toHaveLength(1); + }); + + test('reports incomplete task reads as an error, not an empty pending list', async () => { + mockSend.mockResolvedValueOnce({}) + .mockResolvedValueOnce({ Items: [{ task_id: 't', request_id: 'r' }] }) + .mockResolvedValue({ UnprocessedKeys: { Tasks: { Keys: [{ task_id: 't' }] } } }); + const response = await handler(makeEvent()); + expect(response.statusCode).toBe(500); + }); + test('401 when no Cognito claims', async () => { const event = makeEvent(); (event.requestContext.authorizer as { claims: Record }).claims = {}; @@ -174,7 +293,7 @@ describe('get-pending', () => { expect(body.data.pending[0].severity).toBe('medium'); }); - test('expires_at falls back to created_at when timeout is missing', async () => { + test('expires_at is null when no timeout is configured', async () => { setupPendingMocks([ { task_id: 't', @@ -189,7 +308,7 @@ describe('get-pending', () => { ]); const res = await handler(makeEvent()); const body = JSON.parse(res.body); - expect(body.data.pending[0].expires_at).toBe('2026-05-07T00:00:00Z'); + expect(body.data.pending[0].expires_at).toBeNull(); }); test('500 on DDB error after rate-limit passes', async () => { diff --git a/cdk/test/handlers/get-policies.test.ts b/cdk/test/handlers/get-policies.test.ts index 7b0157f01..e4f008960 100644 --- a/cdk/test/handlers/get-policies.test.ts +++ b/cdk/test/handlers/get-policies.test.ts @@ -249,7 +249,7 @@ describe('get-policies', () => { expect(hardRule.summary).toBeDefined(); }); - test('soft rules carry severity + approval_timeout_s', async () => { + test('built-in soft rules carry severity without an implicit approval deadline', async () => { mockSend.mockResolvedValue({}); mockLoadRepoConfig.mockResolvedValue(null); const res = await handler(makeEvent('soft%2Fshape')); @@ -258,6 +258,6 @@ describe('get-policies', () => { (r: { rule_id: string }) => r.rule_id === 'force_push_any', ); expect(soft.severity).toBe('medium'); - expect(soft.approval_timeout_s).toBe(300); + expect(soft.approval_timeout_s).toBeUndefined(); }); }); diff --git a/cdk/test/handlers/linear-webhook-processor.test.ts b/cdk/test/handlers/linear-webhook-processor.test.ts index 45f8b8c61..7d60cc254 100644 --- a/cdk/test/handlers/linear-webhook-processor.test.ts +++ b/cdk/test/handlers/linear-webhook-processor.test.ts @@ -20,6 +20,11 @@ import * as fs from 'fs'; import * as path from 'path'; +const approvalReplyMock = jest.fn(); +jest.mock('../../src/handlers/shared/linear-approval-reply', () => ({ + handleLinearApprovalReply: (...args: unknown[]) => approvalReplyMock(...args), +})); + const ddbSend = jest.fn(); jest.mock('@aws-sdk/client-dynamodb', () => ({ DynamoDBClient: jest.fn(() => ({})) })); jest.mock('@aws-sdk/lib-dynamodb', () => ({ @@ -1189,3 +1194,33 @@ describe('every channel_metadata builder carries the vault fields', () => { expect(count).toBeGreaterThanOrEqual(4); }); }); + +describe('native approval reply routing', () => { + const oldTable = process.env.TASK_APPROVALS_TABLE_NAME; + afterEach(() => { + if (oldTable === undefined) delete process.env.TASK_APPROVALS_TABLE_NAME; + else process.env.TASK_APPROVALS_TABLE_NAME = oldTable; + approvalReplyMock.mockReset(); + }); + test('consumes approval replies before mention or new-task routing', async () => { + process.env.TASK_APPROVALS_TABLE_NAME = 'Approvals'; + approvalReplyMock.mockResolvedValue(true); + createTaskCoreMock.mockClear(); + const payload = { + type: 'Comment', + action: 'create', + organizationId: 'org-1', + actor: { id: 'user-1' }, + data: { id: 'reply', parentId: 'approval-root', issueId: 'issue-1', body: 'approve' }, + }; + await handler(eventWith(payload)); + expect(approvalReplyMock).toHaveBeenCalledWith(payload, expect.objectContaining({ approvalsTable: 'Approvals' })); + expect(createTaskCoreMock).not.toHaveBeenCalled(); + }); + test('lets transient approval failures reach the async retry mechanism', async () => { + process.env.TASK_APPROVALS_TABLE_NAME = 'Approvals'; + approvalReplyMock.mockRejectedValue(new Error('approval unavailable')); + await expect(handler(eventWith({ type: 'Comment', action: 'create', data: { id: 'reply', body: 'approve' } }))) + .rejects.toThrow('approval unavailable'); + }); +}); diff --git a/cdk/test/handlers/orchestrate-task-microvm.test.ts b/cdk/test/handlers/orchestrate-task-microvm.test.ts index 1023238fd..dfb90358e 100644 --- a/cdk/test/handlers/orchestrate-task-microvm.test.ts +++ b/cdk/test/handlers/orchestrate-task-microvm.test.ts @@ -39,10 +39,17 @@ const MICROVM_ID = 'mvm-0123456789abcdef'; const ENDPOINT = 'https://mvm-0123456789abcdef.microvm.lambda.us-east-1.amazonaws.com'; const mockMicrovmSend = jest.fn(); +const mockSsmSend = jest.fn(); +jest.mock('@aws-sdk/client-ssm', () => ({ + SSMClient: jest.fn(() => ({ send: mockSsmSend })), + GetParameterCommand: jest.fn((input: unknown) => ({ input })), +})); jest.mock('@aws-sdk/client-lambda-microvms', () => ({ LambdaMicrovmsClient: jest.fn(() => ({ send: mockMicrovmSend })), RunMicrovmCommand: jest.fn((input: unknown) => ({ _type: 'RunMicrovm', input })), GetMicrovmCommand: jest.fn((input: unknown) => ({ _type: 'GetMicrovm', input })), + SuspendMicrovmCommand: jest.fn((input: unknown) => ({ _type: 'SuspendMicrovm', input })), + ResumeMicrovmCommand: jest.fn((input: unknown) => ({ _type: 'ResumeMicrovm', input })), TerminateMicrovmCommand: jest.fn((input: unknown) => ({ _type: 'TerminateMicrovm', input })), MicrovmState: { PENDING: 'PENDING', @@ -58,12 +65,36 @@ jest.mock('@aws-sdk/client-lambda-microvms', () => ({ // PUT and the finalize DELETE are both assertions this file needs to make, and a // per-instance mock silently discards them. const mockS3Send = jest.fn().mockResolvedValue({}); +const mockDdbSend = jest.fn().mockResolvedValue({}); +jest.mock('@aws-sdk/client-dynamodb', () => ({ DynamoDBClient: jest.fn(() => ({})) })); +jest.mock('@aws-sdk/lib-dynamodb', () => ({ + DynamoDBDocumentClient: { from: jest.fn(() => ({ send: mockDdbSend })) }, + GetCommand: jest.fn((input: unknown) => ({ _type: 'Get', input })), + PutCommand: jest.fn((input: unknown) => ({ _type: 'Put', input })), + UpdateCommand: jest.fn((input: unknown) => ({ _type: 'Update', input })), +})); +const mockResolveAsset = jest.fn(); +jest.mock('../../src/handlers/shared/registry/factory', () => ({ + makeRegistryClient: () => ({ resolve: mockResolveAsset }), +})); +jest.mock('../../src/handlers/shared/context-hydration', () => ({ + hydrateContext: jest.fn().mockResolvedValue({ sources: [], token_estimate: 0, truncated: false }), +})); +jest.mock('@aws-sdk/s3-request-presigner', () => ({ getSignedUrl: async () => 'https://payloads.s3.us-east-1.amazonaws.com/task/payload.json?X-Amz-Signature=' + Date.now() })); +const mockObjects = new Map(); jest.mock('@aws-sdk/client-s3', () => ({ - S3Client: jest.fn(() => ({ send: mockS3Send })), + GetObjectCommand: jest.fn((input: unknown) => ({ _type: 'GetObject', input })), + S3Client: jest.fn(() => ({ send: mockS3Send, config: { credentials: async () => ({ accessKeyId: 'EXAMPLE', secretAccessKey: 'unused' }) } })), PutObjectCommand: jest.fn((input: unknown) => ({ _type: 'PutObject', input })), DeleteObjectCommand: jest.fn((input: unknown) => ({ _type: 'DeleteObject', input })), })); +const mockReleaseTaskSlot = jest.fn().mockResolvedValue(false); +jest.mock('../../src/handlers/shared/task-concurrency', () => ({ + acquireTaskSlot: jest.fn().mockResolvedValue(true), + releaseTaskSlot: (...args: unknown[]) => mockReleaseTaskSlot(...args), +})); + // Real orchestrator helpers would talk to DynamoDB; stub them and assert on the // calls. `buildComputeMetadata` is kept REAL (re-exported from the actual module) // so the persisted metadata shape is genuinely exercised, not mirrored. @@ -72,8 +103,24 @@ const mockTransitionTask = jest.fn(); const mockEmitTaskEvent = jest.fn(); const mockFinalizeTask = jest.fn(); const mockPollTaskStatus = jest.fn(); -const mockReconcile = jest.fn(); +const mockReadLifecycle = jest.fn(); +const mockSaveIntent = jest.fn(); +jest.mock('../../src/handlers/shared/microvm-lifecycle', () => ({ + ...jest.requireActual('../../src/handlers/shared/microvm-lifecycle'), + readMicrovmLifecycleSnapshot: (...args: unknown[]) => mockReadLifecycle(...args), + saveMicrovmLifecycleIntent: (...args: unknown[]) => mockSaveIntent(...args), +})); const mockFailTask = jest.fn(); +const mockLoadTask = jest.fn(); +const mockClaimStart = jest.fn(); +const mockSaveHandle = jest.fn(); +const mockHydrateAndTransition = jest.fn(); +const mockLoadBlueprint = jest.fn(); +jest.mock('../../src/handlers/shared/microvm-start', () => ({ + ...jest.requireActual('../../src/handlers/shared/microvm-start'), + claimMicrovmStart: (...args: unknown[]) => mockClaimStart(...args), + saveMicrovmStartHandle: (...args: unknown[]) => mockSaveHandle(...args), +})); jest.mock('../../src/handlers/shared/orchestrator', () => ({ admissionControl: jest.fn().mockResolvedValue(true), emitTaskEvent: (...a: unknown[]) => mockEmitTaskEvent(...a), @@ -83,13 +130,10 @@ jest.mock('../../src/handlers/shared/orchestrator', () => ({ }), failTask: (...a: unknown[]) => mockFailTask(...a), finalizeTask: (...a: unknown[]) => mockFinalizeTask(...a), - hydrateAndTransition: jest.fn().mockResolvedValue({ repo_url: 'org/repo', task_id: 'TASK001' }), - loadBlueprintConfig: jest.fn().mockResolvedValue({ compute_type: 'lambda-microvm', runtime_arn: '' }), - loadTask: jest.fn().mockResolvedValue({ - task_id: 'TASK001', user_id: 'user-1', status: 'SUBMITTED', repo: 'org/repo', - }), + hydrateAndTransition: (...a: unknown[]) => mockHydrateAndTransition(...a), + loadBlueprintConfig: (...a: unknown[]) => mockLoadBlueprint(...a), + loadTask: (...args: unknown[]) => mockLoadTask(...args), pollTaskStatus: (...a: unknown[]) => mockPollTaskStatus(...a), - reconcileMicrovmSubstrateState: (...a: unknown[]) => mockReconcile(...a), transitionTask: (...a: unknown[]) => mockTransitionTask(...a), buildComputeMetadata: realOrchestrator.buildComputeMetadata, })); @@ -110,22 +154,20 @@ process.env.MICROVM_EXECUTION_ROLE_ARN = 'arn:aws:iam::123456789012:role/AbcaMic process.env.MICROVM_EGRESS_CONNECTOR_ARNS = 'arn:aws:lambda:us-east-1:123456789012:network-connector/egress-1'; process.env.MICROVM_PAYLOAD_BUCKET = 'test-microvm-payload-bucket'; process.env.TASK_TABLE_NAME = 'Tasks'; +process.env.APPROVAL_REQUESTS_API_URL = 'https://approval.execute-api.us-east-1.amazonaws.com/v1'; process.env.TASK_EVENTS_TABLE_NAME = 'TaskEvents'; process.env.USER_CONCURRENCY_TABLE_NAME = 'UserConcurrency'; process.env.TASK_RETENTION_DAYS = '90'; -// platform_config (ADR-021 P2): the four REQUIRED identifiers the MicroVM -// strategy refuses to start a session without — they are the agent's only -// channel for them, since a snapshot must not bake configuration in. Read at -// call time by `buildMicrovmPlatformConfig`, but set here alongside the rest -// for clarity. +// Required platform configuration is supplied at launch, never baked into the image. process.env.GITHUB_TOKEN_SECRET_ARN = 'arn:aws:secretsmanager:us-east-1:123456789012:secret:abca/github-token-AbCdEf'; process.env.AGENT_SESSION_ROLE_ARN = 'arn:aws:iam::123456789012:role/AbcaAgentSessionRole'; import { TaskStatus } from '../../src/constructs/task-status'; import { handler } from '../../src/handlers/orchestrate-task'; -import { LambdaMicrovmComputeStrategy } from '../../src/handlers/shared/strategies/lambda-microvm-strategy'; +import type { MicrovmLifecycleSnapshot } from '../../src/handlers/shared/microvm-lifecycle'; +import { LambdaMicrovmComputeStrategy, MICROVM_RUN_HOOK_PAYLOAD_LIMIT_BYTES } from '../../src/handlers/shared/strategies/lambda-microvm-strategy'; /** * Minimal stand-in for the durable-execution context: `step` runs its body @@ -134,11 +176,14 @@ import { LambdaMicrovmComputeStrategy } from '../../src/handlers/shared/strategi */ function fakeContext(opts: { pollOnce?: boolean } = {}) { const steps: string[] = []; + const stepConfigs: Record = {}; return { steps, + stepConfigs, ctx: { - step: async (name: string, fn: () => Promise) => { + step: async (name: string, fn: () => Promise, config?: unknown) => { steps.push(name); + stepConfigs[name] = config; return fn(); }, waitForCondition: async ( @@ -216,17 +261,313 @@ function s3CommandsOfType(type: string) { return mockS3Send.mock.calls.map(c => c[0]).filter(c => c._type === type); } +function failedTransition() { + return mockTransitionTask.mock.calls.find(([, , to]) => to === TaskStatus.FAILED)!; +} + +function lifecycle(status: MicrovmLifecycleSnapshot['status'] = 'RUNNING'): MicrovmLifecycleSnapshot { + return { + sleepAfterSeconds: 30, + taskId: 'TASK001', + userId: 'user-1', + status, + requestId: null, + approval: { kind: 'none' }, + taskStartedAtMs: Date.now() - 600_000, + heartbeatAtMs: Date.now(), + handle: { strategyType: 'lambda-microvm', sessionId: MICROVM_ID, microvmId: MICROVM_ID, endpoint: ENDPOINT }, + }; +} + beforeEach(() => { jest.clearAllMocks(); - mockS3Send.mockReset(); - mockS3Send.mockResolvedValue({}); + mockDdbSend.mockReset().mockResolvedValue({}); + mockResolveAsset.mockReset(); + mockHydrateAndTransition.mockReset().mockResolvedValue({ repo_url: 'org/repo', task_id: 'TASK001' }); + mockLoadBlueprint.mockReset().mockResolvedValue({ compute_type: 'lambda-microvm', runtime_arn: '' }); + mockTransitionTask.mockReset().mockResolvedValue(undefined); + mockEmitTaskEvent.mockReset().mockResolvedValue(undefined); + mockFinalizeTask.mockReset().mockResolvedValue(undefined); + mockFailTask.mockReset().mockResolvedValue(undefined); + mockLoadTask.mockReset().mockImplementation(async (_id: string, consistentRead?: boolean) => ({ + task_id: 'TASK001', + user_id: 'user-1', + status: consistentRead ? 'HYDRATING' : 'SUBMITTED', + repo: 'org/repo', + })); + mockClaimStart.mockReset().mockImplementation(async (taskId: string) => ({ clientToken: taskId, closed: false })); + mockSaveHandle.mockReset().mockResolvedValue(undefined); + mockObjects.clear(); + mockS3Send.mockReset().mockImplementation(async ({ _type: type, input: command }) => { + if (type === 'GetObject') { + if (!mockObjects.has(command.Key)) throw Object.assign(new Error('missing'), { name: 'NoSuchKey' }); + return { Body: { transformToString: async () => mockObjects.get(command.Key) } }; + } + if (type === 'DeleteObject') { mockObjects.delete(command.Key); return {}; } + if (command.IfNoneMatch === '*' && mockObjects.has(command.Key)) throw Object.assign(new Error('exists'), { name: 'PreconditionFailed' }); + mockObjects.set(command.Key, command.Body); + return {}; + }); mockMicrovmSend.mockReset(); + mockSsmSend.mockReset(); mockPollTaskStatus.mockResolvedValue({ attempts: 1, lastStatus: TaskStatus.COMPLETED }); - mockReconcile.mockResolvedValue({ taskFailed: false }); + mockReadLifecycle.mockReset().mockResolvedValue(lifecycle('COMPLETED')); + mockSaveIntent.mockReset().mockImplementation(async (snapshot, action) => { + const intent = { + version: 1, + generation: 'generation', + microvm_id: MICROVM_ID, + request_id: snapshot.requestId, + action, + requested_at_ms: Date.now(), + deadline_ms: snapshot.approval.kind === 'present' ? snapshot.approval.deadlineMs : null, + }; + mockReadLifecycle.mockResolvedValue({ ...snapshot, intent }); + return { status: 'saved', intent }; + }); + delete process.env.MICROVM_APPROVAL_SUSPEND_ENABLED; + delete process.env.MICROVM_APPROVAL_SUSPEND_PARAMETER_NAME; }); describe('orchestrate-task for a lambda-microvm task', () => { - test('persists microvmId and endpoint in compute_metadata on the RUNNING transition', async () => { + test.each(['true', 'false', 'unavailable'])('real supervisor checks live setting %s before Suspend', async value => { + runMicrovmOk(); + const now = Date.now(); + const snapshot = lifecycle('AWAITING_APPROVAL'); + mockReadLifecycle.mockResolvedValue({ + ...snapshot, + requestId: 'gate', + handle: { + ...snapshot.handle, + imageArn: 'arn:aws:lambda:us-east-1:123456789012:microvm-image:agent', + imageVersion: '3.0', + lifecycleProtocol: '1', + }, + approval: { + kind: 'present', + status: 'PENDING', + created_at: new Date(now - 45_000).toISOString(), + timeout_s: 600, + createdAtMs: now - 45_000, + deadlineMs: now + 555_000, + }, + }); + const parameterName = '/backgroundagent-dev/microvm-approval-suspend-enabled'; + process.env.MICROVM_APPROVAL_SUSPEND_ENABLED = 'true'; + process.env.MICROVM_APPROVAL_SUSPEND_PARAMETER_NAME = parameterName; + if (value === 'unavailable') mockSsmSend.mockRejectedValue(new Error('unavailable')); + else mockSsmSend.mockResolvedValue({ Parameter: { Name: parameterName, Value: value } }); + mockMicrovmSend.mockResolvedValueOnce({ + microvmId: MICROVM_ID, + state: 'RUNNING', + startedAt: new Date(now - 60_000), + maximumDurationInSeconds: 28_800, + }).mockResolvedValue({}); + await handler({ task_id: 'TASK001' }, fakeContext().ctx as never); + expect(mockSsmSend).toHaveBeenCalledWith({ input: { Name: parameterName } }, { abortSignal: expect.any(AbortSignal) }); + expect(commandsOfType('SuspendMicrovm')).toHaveLength(value === 'true' ? 1 : 0); + expect(mockFinalizeTask.mock.calls[0][1].microvmFailureReason).toBeUndefined(); + expect(commandsOfType('TerminateMicrovm')).toHaveLength(1); + }); + + test('registry assets alone exceed the hook cap and survive real hydration and v2 S3 delivery (#818)', async () => { + const runtime = { + transport: 'http', + url: 'https://mcp.example.com/tools', + headers: { 'X-Registry-Context': 'x'.repeat(6000) }, + }; + const asset = { kind: 'mcp_server', namespace: 'acme', name: 'large', version: '1.0.0', runtime }; + mockResolveAsset.mockResolvedValue({ ...asset, warnings: [] }); + mockLoadBlueprint.mockResolvedValue({ + compute_type: 'lambda-microvm', + runtime_arn: '', + mcp_servers: ['registry://mcp_server/acme/large@1.0.0'], + }); + // Exercise the real registry resolution, payload assembly and audit writes; + // only external service boundaries are stubbed on this hydration path. + mockHydrateAndTransition.mockImplementationOnce(realOrchestrator.hydrateAndTransition); + runMicrovmOk(); + + await handler({ task_id: 'TASK001' }, fakeContext().ctx as never); + + expect(mockResolveAsset).toHaveBeenCalledTimes(1); + const upload = s3CommandsOfType('PutObject').find(c => c.input.Key === 'TASK001/payload.json'); + const document = JSON.parse(upload.input.Body); + expect(document.agent_payload.resolved_assets).toEqual([asset]); + const { resolved_assets: assets, ...basePayload } = document.agent_payload; + expect(Buffer.byteLength(JSON.stringify(basePayload))).toBeLessThan(MICROVM_RUN_HOOK_PAYLOAD_LIMIT_BYTES); + expect(Buffer.byteLength(JSON.stringify(assets))).toBeGreaterThan(MICROVM_RUN_HOOK_PAYLOAD_LIMIT_BYTES); + const run = commandsOfType('RunMicrovm'); + expect(run).toHaveLength(1); + const wire = run[0].input.runHookPayload; + expect(Buffer.byteLength(wire)).toBeLessThanOrEqual(MICROVM_RUN_HOOK_PAYLOAD_LIMIT_BYTES); + expect(JSON.parse(wire)).toMatchObject({ version: 2, task_id: 'TASK001', payload_url: expect.any(String) }); + expect(wire).not.toContain('X-Registry-Context'); + const audit = mockDdbSend.mock.calls.find(([c]) => c.input.ExpressionAttributeValues?.[':ra']); + expect(audit![0].input.ExpressionAttributeValues[':ra']).toEqual([ + { kind: 'mcp_server', id: 'acme/large', version: '1.0.0' }, + ]); + }); + + test.each([2, 3])('cancellation on task read %s checks release before leaving the pipeline', async (cancelOnRead) => { + let reads = 0; + mockLoadTask.mockImplementation(async () => ({ + task_id: 'TASK001', + user_id: 'user-1', + repo: 'org/repo', + status: ++reads >= cancelOnRead ? TaskStatus.CANCELLED : TaskStatus.SUBMITTED, + })); + await handler({ task_id: 'TASK001' }, fakeContext().ctx as never); + expect(commandsOfType('RunMicrovm')).toHaveLength(0); + expect(mockReleaseTaskSlot).toHaveBeenCalledWith('TASK001', 'user-1'); + }); + + test('replaying a persisted start failure reaches finalization once without an earlier slot release', async () => { + let failed = false; + mockMicrovmSend.mockRejectedValue(Object.assign(new Error('denied'), { name: 'AccessDeniedException' })); + mockClaimStart.mockImplementation(async () => ({ clientToken: 'TASK001', closed: failed })); + mockTransitionTask.mockImplementation(async (_id, _from, to) => { failed = to === TaskStatus.FAILED; }); + mockLoadTask.mockImplementation(async (_id, consistentRead) => ({ + task_id: 'TASK001', + user_id: 'user-1', + repo: 'org/repo', + status: consistentRead ? (failed ? TaskStatus.FAILED : TaskStatus.HYDRATING) : TaskStatus.SUBMITTED, + })); + const { ctx } = fakeContext(); + await handler({ task_id: 'TASK001' }, { + ...ctx, + step: async (name: string, fn: () => Promise, config?: unknown) => { + if (name === 'start-session') await fn(); + return ctx.step(name, fn, config); + }, + } as never); + expect(commandsOfType('RunMicrovm')).toHaveLength(1); + expect(mockTransitionTask).toHaveBeenCalledTimes(1); + expect(mockFailTask).not.toHaveBeenCalled(); + expect(mockFinalizeTask).toHaveBeenCalledTimes(1); + }); + + test('durable replay after registration recovers one computer without a second transition', async () => { + runMicrovmOk(); + let savedHandle: unknown; + let registered = false; + mockClaimStart.mockImplementation(async () => ({ + clientToken: 'TASK001', closed: false, ...(savedHandle ? { handle: savedHandle } : {}), + })); + mockSaveHandle.mockImplementation(async (_taskId, _token, handle) => { savedHandle = handle; }); + mockTransitionTask.mockImplementationOnce(async () => { registered = true; }); + mockLoadTask.mockImplementation(async (_id, consistentRead) => ({ + task_id: 'TASK001', + user_id: 'user-1', + repo: 'org/repo', + status: consistentRead ? (registered ? TaskStatus.RUNNING : TaskStatus.HYDRATING) : TaskStatus.SUBMITTED, + ...(registered && { session_id: MICROVM_ID }), + })); + const { ctx, stepConfigs } = fakeContext(); + const replayContext = { + ...ctx, + step: async (name: string, fn: () => Promise, config?: unknown) => { + if (name === 'start-session') await fn(); // response/checkpoint lost + return ctx.step(name, fn, config); + }, + }; + await handler({ task_id: 'TASK001' }, replayContext as never); + expect(commandsOfType('RunMicrovm')).toHaveLength(1); + expect(mockTransitionTask).toHaveBeenCalledTimes(1); + expect(stepConfigs['start-session'].retryStrategy(new Error('failed'), 1)).toEqual({ shouldRetry: false }); + expect(mockFailTask).not.toHaveBeenCalled(); + }); + + test('a lost task-transition response is reconciled before stopping a registered computer', async () => { + runMicrovmOk(); + let registered = false; + mockTransitionTask.mockImplementationOnce(async () => { + registered = true; + throw new Error('DynamoDB response lost'); + }); + mockLoadTask.mockImplementation(async (_id, consistentRead) => ({ + task_id: 'TASK001', + user_id: 'user-1', + repo: 'org/repo', + status: consistentRead ? (registered ? TaskStatus.RUNNING : TaskStatus.HYDRATING) : TaskStatus.SUBMITTED, + ...(registered && { session_id: MICROVM_ID }), + })); + await handler({ task_id: 'TASK001' }, fakeContext().ctx as never); + expect(mockFailTask).not.toHaveBeenCalled(); + expect(mockFinalizeTask).toHaveBeenCalledTimes(1); + expect(commandsOfType('RunMicrovm')).toHaveLength(1); + }); + + test('cancellation after creation terminates and finalizes without restoring RUNNING', async () => { + runMicrovmOk(); + mockLoadTask.mockImplementation(async (_id, consistentRead) => ({ + task_id: 'TASK001', + user_id: 'user-1', + repo: 'org/repo', + status: consistentRead && commandsOfType('RunMicrovm').length > 0 ? TaskStatus.CANCELLED : TaskStatus.SUBMITTED, + })); + const { ctx, steps } = fakeContext(); + await handler({ task_id: 'TASK001' }, ctx as never); + expect(mockTransitionTask).not.toHaveBeenCalled(); + expect(mockFailTask).not.toHaveBeenCalled(); + expect(commandsOfType('TerminateMicrovm')).toHaveLength(1); + expect(mockFinalizeTask).toHaveBeenCalledWith('TASK001', { attempts: 0 }, 'user-1'); + expect(steps).toContain('finalize-before-session'); + expect(mockPollTaskStatus).not.toHaveBeenCalled(); + }); + + test('a session_started audit-event failure does not destroy a registered session', async () => { + runMicrovmOk(); + mockEmitTaskEvent.mockImplementationOnce(async (_id, eventType) => { + if (eventType === 'session_started') throw new Error('event write failed'); + }); + await handler({ task_id: 'TASK001' }, fakeContext().ctx as never); + expect(mockFailTask).not.toHaveBeenCalled(); + const terminateIndex = mockMicrovmSend.mock.calls.findIndex(([command]) => command._type === 'TerminateMicrovm'); + expect(mockMicrovmSend.mock.invocationCallOrder[terminateIndex]) + .toBeGreaterThan(mockFinalizeTask.mock.invocationCallOrder[0]); + }); + + test('two unanswered starts persist an unknown outcome instead of inviting another task', async () => { + const timeout = Object.assign(new Error('response lost'), { name: 'TimeoutError' }); + mockMicrovmSend.mockRejectedValue(timeout); + await handler({ task_id: 'TASK001' }, fakeContext().ctx as never); + expect(commandsOfType('RunMicrovm')).toHaveLength(2); + expect(failedTransition()[3].error_message).toContain('MICROVM_START_OUTCOME_UNKNOWN:'); + }); + + test('an unanswered start remains visible if the agent already changed the task to RUNNING', async () => { + mockMicrovmSend.mockRejectedValue(Object.assign(new Error('response lost'), { name: 'TimeoutError' })); + mockLoadTask.mockImplementation(async (_id, consistentRead) => ({ + task_id: 'TASK001', + user_id: 'user-1', + repo: 'org/repo', + status: consistentRead ? TaskStatus.RUNNING : TaskStatus.SUBMITTED, + })); + await handler({ task_id: 'TASK001' }, fakeContext().ctx as never); + expect(failedTransition()[1]).toBe(TaskStatus.RUNNING); + expect(failedTransition()[3].error_message).toContain('MICROVM_START_OUTCOME_UNKNOWN:'); + }); + + test('cancellation after a lost response records the unknown computer without starting another', async () => { + mockClaimStart.mockResolvedValueOnce({ clientToken: 'TASK001', closed: false }) + .mockResolvedValueOnce({ clientToken: 'TASK001', closed: false }) + .mockResolvedValue({ clientToken: 'TASK001', closed: true }); + mockMicrovmSend.mockRejectedValue(Object.assign(new Error('response lost'), { name: 'TimeoutError' })); + mockLoadTask.mockImplementation(async (_id, consistentRead) => ({ + task_id: 'TASK001', + user_id: 'user-1', + repo: 'org/repo', + status: consistentRead && commandsOfType('RunMicrovm').length > 0 ? TaskStatus.CANCELLED : TaskStatus.SUBMITTED, + })); + await handler({ task_id: 'TASK001' }, fakeContext().ctx as never); + expect(commandsOfType('RunMicrovm')).toHaveLength(1); + expect(mockEmitTaskEvent).toHaveBeenCalledWith('TASK001', 'microvm_start_outcome_unknown', + { client_token: 'TASK001', task_status: TaskStatus.CANCELLED }, expect.any(Object)); + expect(mockFailTask).not.toHaveBeenCalled(); + }); + + test('persists the worker and actual image identity on the RUNNING transition', async () => { runMicrovmOk(); const { ctx } = fakeContext(); @@ -241,7 +582,9 @@ describe('orchestrate-task for a lambda-microvm task', () => { expect(to).toBe(TaskStatus.RUNNING); expect(attrs.compute_type).toBe('lambda-microvm'); // ADR-021: the P3 approve/deny Lambdas resume from exactly these two keys. - expect(attrs.compute_metadata).toEqual({ microvmId: MICROVM_ID, endpoint: ENDPOINT }); + expect(attrs.compute_metadata).toEqual({ + microvmId: MICROVM_ID, endpoint: ENDPOINT, imageArn: 'arn:image', imageVersion: '7', + }); // sessionId is the microvmId (substrate identifier, mirroring ECS). expect(attrs.session_id).toBe(MICROVM_ID); // agent_runtime_arn is an AgentCore-only attribute and must not appear. @@ -299,21 +642,23 @@ describe('orchestrate-task for a lambda-microvm task', () => { // --- finalize-time payload delete (review NB3) --- - test('deletes its OWN payload object on finalize, closing the cross-task read window', async () => { - // The execution role's payload-bucket grant is bucket-wide `grantRead`, so a - // TTL-only reaper left every finished task's hydrated prompt readable by any - // running MicroVM for ~24 h. ECS already deleted at finalize; this is the parity - // that matters. + test('deletes its own payload and saved capability on finalize', async () => { + // Revoke the signed payload link and remove its private replay record. + // The one-day lifecycle remains a cleanup backstop. runMicrovmOk(); await handler({ task_id: 'TASK001' }, fakeContext().ctx as never); const deletes = s3CommandsOfType('DeleteObject'); - expect(deletes).toHaveLength(1); + expect(deletes).toHaveLength(2); expect(deletes[0].input).toEqual({ Bucket: 'test-microvm-payload-bucket', Key: 'TASK001/payload.json', }); + expect(deletes[1].input).toEqual({ + Bucket: 'test-microvm-payload-bucket', + Key: 'TASK001/launch.json', + }); // ...and it did not reach for the ECS deleter. expect(mockDeleteEcsPayload).not.toHaveBeenCalled(); }); @@ -326,162 +671,121 @@ describe('orchestrate-task for a lambda-microvm task', () => { // billed. `fakeContext` runs one iteration and ignores `waitStrategy`, so this // needed `loopingContext` to be assertable at all. - test('a stale agent heartbeat stops the poll AND reclaims the MicroVM', async () => { + test('a stale agent heartbeat stops the real supervisor and reclaims the MicroVM', async () => { runMicrovmOk(); - // The substrate stays healthy throughout — this is the blind spot. mockMicrovmSend.mockResolvedValue({ microvmId: MICROVM_ID, state: 'RUNNING' }); - mockReconcile.mockResolvedValue({ taskFailed: false, suspendAnomalyReported: false }); - mockPollTaskStatus.mockResolvedValue({ - attempts: 1, - lastStatus: TaskStatus.RUNNING, - sessionUnhealthy: true, - }); - + mockReadLifecycle.mockResolvedValue({ ...lifecycle(), heartbeatAtMs: Date.now() - 300_000 }); const looping = loopingContext(); await handler({ task_id: 'TASK001' }, looping.ctx as never); - - // Exited on the FIRST unhealthy observation — not after burning the 8.5 h window. expect(looping.iterations).toHaveLength(1); - // finalize ran, and it saw the unhealthy flag (that is what writes FAILED). - expect(mockFinalizeTask).toHaveBeenCalledTimes(1); expect(mockFinalizeTask.mock.calls[0][1]).toMatchObject({ sessionUnhealthy: true }); - // ...and the reservation was actually reclaimed. This is the billing outcome. expect(commandsOfType('TerminateMicrovm')).toHaveLength(1); - expect(commandsOfType('TerminateMicrovm')[0].input) - .toEqual({ microvmIdentifier: MICROVM_ID }); + expect(mockPollTaskStatus).not.toHaveBeenCalled(); }); test('a healthy heartbeat keeps polling until a terminal task status', async () => { - // The other side of the same predicate: without this, a `sessionUnhealthy: true` - // hard-coded into the poll would pass the test above. runMicrovmOk(); mockMicrovmSend.mockResolvedValue({ microvmId: MICROVM_ID, state: 'RUNNING' }); - mockReconcile.mockResolvedValue({ taskFailed: false, suspendAnomalyReported: false }); - mockPollTaskStatus - .mockResolvedValueOnce({ attempts: 1, lastStatus: TaskStatus.RUNNING, sessionUnhealthy: false }) - .mockResolvedValueOnce({ attempts: 2, lastStatus: TaskStatus.RUNNING, sessionUnhealthy: false }) - .mockResolvedValue({ attempts: 3, lastStatus: TaskStatus.COMPLETED, sessionUnhealthy: false }); - + mockReadLifecycle.mockResolvedValueOnce(lifecycle()).mockResolvedValueOnce(lifecycle()); const looping = loopingContext(); await handler({ task_id: 'TASK001' }, looping.ctx as never); - expect(looping.iterations).toHaveLength(3); expect(mockFinalizeTask.mock.calls[0][1]).toMatchObject({ lastStatus: TaskStatus.COMPLETED }); - // Terminate happens on every finalize, healthy or not — the VM does not - // self-terminate on this substrate. expect(commandsOfType('TerminateMicrovm')).toHaveLength(1); }); - test('cross-checks the substrate through reconcileMicrovmSubstrateState while non-terminal', async () => { + test('unexpected suspension uses durable wake intent and requests Resume', async () => { runMicrovmOk(); - mockPollTaskStatus.mockResolvedValue({ attempts: 1, lastStatus: TaskStatus.RUNNING }); - // GetMicrovm during the poll, then TerminateMicrovm on finalize. + mockReadLifecycle.mockResolvedValue(lifecycle()); mockMicrovmSend.mockResolvedValueOnce({ microvmId: MICROVM_ID, state: 'SUSPENDED' }); - await handler({ task_id: 'TASK001' }, fakeContext().ctx as never); - expect(commandsOfType('GetMicrovm')).toHaveLength(1); - expect(mockReconcile).toHaveBeenCalledTimes(1); - const args = mockReconcile.mock.calls[0][0]; - expect(args.microvmId).toBe(MICROVM_ID); - expect(args.ddbStatus).toBe(TaskStatus.RUNNING); - // The strategy's mechanical mapping is what the orchestrator interprets. - expect(args.substrate).toEqual({ status: 'suspended' }); + expect(mockSaveIntent.mock.calls[0][1]).toBe('resume'); + expect(commandsOfType('ResumeMicrovm')).toHaveLength(1); + expect(mockFinalizeTask.mock.calls[0][1].microvmSupervisor.recovery.kind).toBe('wake'); + expect(mockEmitTaskEvent.mock.calls.filter(c => c[1] === 'microvm_suspend_anomaly')).toHaveLength(1); }); - test('returns a failed poll state when reconciliation fails the task', async () => { + test('terminal substrate carries its classified reason into strong finalization', async () => { runMicrovmOk(); - mockPollTaskStatus.mockResolvedValue({ attempts: 3, lastStatus: TaskStatus.RUNNING }); - mockMicrovmSend.mockResolvedValueOnce({ microvmId: MICROVM_ID, state: 'TERMINATED' }); - mockReconcile.mockResolvedValue({ taskFailed: true }); - - await handler({ task_id: 'TASK001' }, fakeContext().ctx as never); - - expect(mockFinalizeTask).toHaveBeenCalledWith( - 'TASK001', - { attempts: 3, lastStatus: TaskStatus.FAILED }, - 'user-1', - ); + mockReadLifecycle.mockResolvedValue(lifecycle()); + mockMicrovmSend.mockResolvedValueOnce({ microvmId: MICROVM_ID, state: 'TERMINATED', stateReason: 'Run lifecycle hook returned HTTP status 400.' }); + const looping = loopingContext(); + await handler({ task_id: 'TASK001' }, looping.ctx as never); + expect(looping.iterations).toHaveLength(1); + expect(mockFinalizeTask.mock.calls[0][1]).toMatchObject({ + microvmFailureReason: 'substrate-terminal', + microvmFailureMessage: expect.stringMatching(/^MICROVM_RUN_HOOK_REJECTED:/), + }); }); - test('skips the substrate cross-check once the DDB status is terminal', async () => { + test('a terminal task skips Get but still terminates its original handle', async () => { runMicrovmOk(); - mockPollTaskStatus.mockResolvedValue({ attempts: 1, lastStatus: TaskStatus.COMPLETED }); - await handler({ task_id: 'TASK001' }, fakeContext().ctx as never); - expect(commandsOfType('GetMicrovm')).toHaveLength(0); - expect(mockReconcile).not.toHaveBeenCalled(); - // Finalize still terminates. expect(commandsOfType('TerminateMicrovm')).toHaveLength(1); }); - test('a GetMicrovm poll failure is non-fatal and finalize still terminates', async () => { + test('three Get failures stop the real durable loop and reclaim the VM', async () => { runMicrovmOk(); - mockPollTaskStatus.mockResolvedValue({ attempts: 1, lastStatus: TaskStatus.RUNNING }); - mockMicrovmSend.mockRejectedValueOnce(new Error('transient')); - - await expect(handler({ task_id: 'TASK001' }, fakeContext().ctx as never)).resolves.toBeUndefined(); - - expect(mockReconcile).not.toHaveBeenCalled(); - expect(commandsOfType('TerminateMicrovm')).toHaveLength(1); - }); - - test('threads suspendAnomalyReported into the reconcile call so the event fires once', async () => { - runMicrovmOk(); - mockPollTaskStatus.mockResolvedValue({ attempts: 1, lastStatus: TaskStatus.RUNNING }); - mockMicrovmSend.mockResolvedValueOnce({ microvmId: MICROVM_ID, state: 'SUSPENDED' }); - mockReconcile.mockResolvedValue({ taskFailed: false, suspendAnomalyReported: true }); - - await handler({ task_id: 'TASK001' }, fakeContext().ctx as never); - - // First poll of the task: nothing reported yet. - expect(mockReconcile.mock.calls[0][0].suspendAnomalyReported).toBe(false); - // ...and the reconciler's answer is carried into the state the next poll reads. - expect(mockFinalizeTask).toHaveBeenCalledWith( - 'TASK001', - expect.objectContaining({ microvmSuspendAnomalyReported: true }), - 'user-1', - ); + mockReadLifecycle.mockResolvedValue(lifecycle()); + mockMicrovmSend.mockRejectedValue(new Error('transient')); + const looping = loopingContext(); + await handler({ task_id: 'TASK001' }, looping.ctx as never); + expect(looping.iterations).toHaveLength(3); + expect(mockFinalizeTask.mock.calls[0][1]).toMatchObject({ + microvmFailureReason: 'substrate-read-failed', microvmSupervisor: { consecutivePollFailures: 3 }, + }); + expect(commandsOfType('TerminateMicrovm')).toHaveLength(2); + expect(mockEmitTaskEvent.mock.calls.filter(c => c[1] === 'microvm_cleanup_unconfirmed')).toHaveLength(1); }); - test('a MicroVM poll failure carries the anomaly flag forward rather than re-arming it', async () => { - // A GetMicrovm hiccup is not evidence that the anomaly ended, so it must not - // silently re-arm the event and produce a duplicate on the next cycle. + test('fast polls cannot exhaust the old attempt cap while the session deadline remains', async () => { runMicrovmOk(); - mockPollTaskStatus.mockResolvedValue({ attempts: 1, lastStatus: TaskStatus.RUNNING }); - mockMicrovmSend.mockRejectedValueOnce(new Error('transient')); - + mockReadLifecycle.mockResolvedValue(lifecycle()); + mockMicrovmSend.mockResolvedValue({ microvmId: MICROVM_ID, state: 'RUNNING' }); const { ctx } = fakeContext(); - // Seed the poll state as if a previous cycle had already reported. - const seededCtx = { + let continued: boolean | undefined; + await handler({ task_id: 'TASK001' }, { ...ctx, - waitForCondition: async ( - _name: string, - fn: (state: Record) => Promise, - ) => fn({ attempts: 1, microvmSuspendAnomalyReported: true }), - }; + waitForCondition: async (_name: string, fn: (state: Record) => Promise>, cfg: any) => { + const state = JSON.parse(JSON.stringify(await fn({ attempts: 2000 }))); + continued = cfg.waitStrategy(state).shouldContinue; + return state; + }, + } as never); + expect(continued).toBe(true); + }); - await handler({ task_id: 'TASK001' }, seededCtx as never); + test('lost worker ownership skips task finalization and payload deletion, but reaps only the old VM', async () => { + runMicrovmOk(); + mockReadLifecycle.mockResolvedValue({ ...lifecycle(), handle: { ...lifecycle().handle, microvmId: 'other' } }); + const looping = loopingContext(); + await handler({ task_id: 'TASK001' }, looping.ctx as never); + expect(looping.iterations).toHaveLength(1); + expect(mockFinalizeTask).not.toHaveBeenCalled(); + expect(s3CommandsOfType('DeleteObject')).toHaveLength(0); + expect(commandsOfType('GetMicrovm')).toHaveLength(0); + expect(commandsOfType('TerminateMicrovm')[0].input.microvmIdentifier).toBe(MICROVM_ID); + }); - expect(mockFinalizeTask).toHaveBeenCalledWith( - 'TASK001', - expect.objectContaining({ microvmSuspendAnomalyReported: true }), - 'user-1', - ); + test('database finalization failure still requests termination and propagates for durable retry', async () => { + runMicrovmOk(); + mockFinalizeTask.mockRejectedValue(new Error('database unavailable')); + await expect(handler({ task_id: 'TASK001' }, fakeContext().ctx as never)).rejects.toThrow('database unavailable'); + expect(commandsOfType('TerminateMicrovm')).toHaveLength(1); + expect(s3CommandsOfType('DeleteObject')).toHaveLength(0); }); describe('orphan reap when session start fails AFTER RunMicrovm succeeded', () => { - test('terminates the MicroVM from the in-memory handle when the persist write fails', async () => { - // The MicroVM is already RUNNING and billing, and its id exists ONLY in this - // Lambda's memory — no poll or finalize step will ever see it. Nothing - // self-terminates on this substrate, so without the reap it bills for the - // full 8 h cap while holding admission-gating memory quota. + test('terminates the known MicroVM when registration fails before committing', async () => { + // The service created the computer, but registration did not commit. + // Its start receipt retains the ID; normal task polling has not started. runMicrovmOk(); mockTransitionTask.mockRejectedValueOnce(new Error('ConditionalCheckFailedException')); - await expect(handler({ task_id: 'TASK001' }, fakeContext().ctx as never)) - .rejects.toThrow('ConditionalCheckFailedException'); + await handler({ task_id: 'TASK001' }, fakeContext().ctx as never); + expect(failedTransition()[3].error_message).toContain('ConditionalCheckFailedException'); const terminates = commandsOfType('TerminateMicrovm'); expect(terminates).toHaveLength(1); @@ -492,11 +796,13 @@ describe('orchestrate-task for a lambda-microvm task', () => { runMicrovmOk(); mockTransitionTask.mockRejectedValueOnce(new Error('persist exploded')); - await expect(handler({ task_id: 'TASK001' }, fakeContext().ctx as never)) - .rejects.toThrow('persist exploded'); + await handler({ task_id: 'TASK001' }, fakeContext().ctx as never); + expect(failedTransition()[3].error_message).toContain('persist exploded'); - expect(mockFailTask).toHaveBeenCalledTimes(1); - const [, fromStatus, reason] = mockFailTask.mock.calls[0]; + expect(mockFailTask).not.toHaveBeenCalled(); + expect(mockFinalizeTask).toHaveBeenCalledTimes(1); + const [, fromStatus, , attrs] = failedTransition(); + const reason = attrs.error_message; expect(fromStatus).toBe(TaskStatus.HYDRATING); expect(reason).toContain('Session start failed'); expect(reason).toContain('persist exploded'); @@ -512,8 +818,8 @@ describe('orchestrate-task for a lambda-microvm task', () => { // stopSession is internally best-effort (it logs AccessDenied at error level // and returns), so the reap is a no-op here — the user must still see why // session start actually failed. - await expect(handler({ task_id: 'TASK001' }, fakeContext().ctx as never)) - .rejects.toThrow('persist exploded'); + await handler({ task_id: 'TASK001' }, fakeContext().ctx as never); + expect(failedTransition()[3].error_message).toContain('persist exploded'); }); test('even a stopSession that BREAKS its no-throw contract cannot mask the original error', async () => { @@ -527,8 +833,8 @@ describe('orchestrate-task for a lambda-microvm task', () => { .mockRejectedValueOnce(new Error('stopSession itself threw')); try { - await expect(handler({ task_id: 'TASK001' }, fakeContext().ctx as never)) - .rejects.toThrow('persist exploded'); + await handler({ task_id: 'TASK001' }, fakeContext().ctx as never); + expect(failedTransition()[3].error_message).toContain('persist exploded'); } finally { stopSpy.mockRestore(); } @@ -539,7 +845,7 @@ describe('orchestrate-task for a lambda-microvm task', () => { err.name = 'ThrottlingException'; mockMicrovmSend.mockRejectedValueOnce(err); - await expect(handler({ task_id: 'TASK001' }, fakeContext().ctx as never)).rejects.toThrow(); + await handler({ task_id: 'TASK001' }, fakeContext().ctx as never); expect(commandsOfType('TerminateMicrovm')).toHaveLength(0); }); diff --git a/cdk/test/handlers/orchestrate-task.test.ts b/cdk/test/handlers/orchestrate-task.test.ts index be9901081..81aef8e35 100644 --- a/cdk/test/handlers/orchestrate-task.test.ts +++ b/cdk/test/handlers/orchestrate-task.test.ts @@ -18,6 +18,12 @@ */ // --- Mocks --- +const mockAcquireTaskSlot = jest.fn(); +const mockReleaseTaskSlot = jest.fn(); +jest.mock('../../src/handlers/shared/task-concurrency', () => ({ + acquireTaskSlot: (...args: unknown[]) => mockAcquireTaskSlot(...args), + releaseTaskSlot: (...args: unknown[]) => mockReleaseTaskSlot(...args), +})); const mockDdbSend = jest.fn(); jest.mock('@aws-sdk/client-dynamodb', () => ({ DynamoDBClient: jest.fn(() => ({})) })); jest.mock('@aws-sdk/lib-dynamodb', () => ({ @@ -106,6 +112,9 @@ const baseTask = { beforeEach(() => { jest.clearAllMocks(); + mockDdbSend.mockReset(); + mockAcquireTaskSlot.mockReset().mockResolvedValue(true); + mockReleaseTaskSlot.mockReset().mockResolvedValue(true); ulidCounter = 0; mockLoadRepoConfig.mockResolvedValue(null); }); @@ -124,22 +133,16 @@ describe('loadTask', () => { }); describe('admissionControl', () => { - test('returns true when concurrency slot is available', async () => { - mockDdbSend.mockResolvedValueOnce({}); - const result = await admissionControl(baseTask as any); - expect(result).toBe(true); + test('reserves using the task identity and configured user limit', async () => { + expect(await admissionControl(baseTask as any)).toBe(true); + expect(mockAcquireTaskSlot).toHaveBeenCalledWith('TASK001', 'user-123', 3); }); - - test('returns false when concurrency limit reached', async () => { - const condErr = new Error('Conditional check failed'); - condErr.name = 'ConditionalCheckFailedException'; - mockDdbSend.mockRejectedValueOnce(condErr); - const result = await admissionControl(baseTask as any); - expect(result).toBe(false); + test('returns false when no reservation was acquired', async () => { + mockAcquireTaskSlot.mockResolvedValue(false); + expect(await admissionControl(baseTask as any)).toBe(false); }); - - test('throws on unexpected DDB errors', async () => { - mockDdbSend.mockRejectedValueOnce(new Error('DynamoDB error')); + test('propagates an outage so durable execution can retry', async () => { + mockAcquireTaskSlot.mockRejectedValue(new Error('DynamoDB error')); await expect(admissionControl(baseTask as any)).rejects.toThrow('DynamoDB error'); }); }); @@ -548,6 +551,37 @@ describe('hydrateAndTransition — Cedar HITL payload threading', () => { }); describe('pollTaskStatus', () => { + test.each(['agentcore', 'lambda-microvm'] as const)( + '%s stays healthy immediately after a long approval wait resumes', + async (computeType) => { + const beforeApproval = new Date(Date.now() - 600_000).toISOString(); + mockDdbSend.mockResolvedValueOnce({ + Item: { + status: 'AWAITING_APPROVAL', + session_id: 'sess-1', + started_at: beforeApproval, + agent_heartbeat_at: beforeApproval, + }, + }); + const waiting = await pollTaskStatus('TASK001', { attempts: 5 }, computeType); + expect(waiting.sessionUnhealthy).toBe(false); + + // The agent's conditional resume transaction refreshes this timestamp with + // status RUNNING. Poll before the independent 45-second worker ticks. + mockDdbSend.mockResolvedValueOnce({ + Item: { + status: 'RUNNING', + session_id: 'sess-1', + started_at: beforeApproval, + agent_heartbeat_at: new Date().toISOString(), + }, + }); + const resumed = await pollTaskStatus('TASK001', waiting, computeType); + expect(resumed.lastStatus).toBe('RUNNING'); + expect(resumed.sessionUnhealthy).toBe(false); + }, + ); + test('increments attempt count and reads status', async () => { mockDdbSend.mockResolvedValueOnce({ Item: { status: 'RUNNING' } }); const result = await pollTaskStatus('TASK001', { attempts: 5 }, 'agentcore'); @@ -1192,10 +1226,115 @@ describe('hydrateAndTransition — registry asset resolution (#246)', () => { }); describe('finalizeTask', () => { + const supervisor = { + version: 1 as const, + microvmId: 'vm', + firstObservedAtMs: 1, + sessionDeadlineMs: 28_800_001, + lifetimeVerified: true, + consecutivePollFailures: 3, + consecutiveResumeFailures: 0, + anomalyReported: false, + nextPollInMs: 5_000, + }; + const vmTask = () => ({ + ...baseTask, + status: 'AWAITING_APPROVAL', + compute_type: 'lambda-microvm', + session_id: 'vm', + compute_metadata: { microvmId: 'vm', endpoint: 'https://vm.example' }, + }); + + test.each(['AWAITING_APPROVAL', 'RUNNING', 'HYDRATING'])('supervisor failure strongly reads %s and conditionally fails only its worker', async status => { + mockDdbSend.mockResolvedValueOnce({ Item: { ...vmTask(), status } }).mockResolvedValue({}); + await finalizeTask('TASK001', { + attempts: 3, microvmSupervisor: supervisor, microvmFailureReason: 'substrate-read-failed', + }, 'user-123'); + const read = mockDdbSend.mock.calls[0][0]; + expect(read.input.ConsistentRead).toBe(true); + const write = mockDdbSend.mock.calls[1][0]; + expect(write.input.ConditionExpression).toContain('compute_metadata.microvmId = :microvmId'); + expect(write.input.ExpressionAttributeValues).toMatchObject({ + ':fromStatus': status, + ':toStatus': 'FAILED', + ':microvmId': 'vm', + ':attr_error_message': 'MicroVM supervisor: substrate-read-failed', + }); + expect(mockReleaseTaskSlot).toHaveBeenCalledTimes(1); + }); + + test('supervisor finalization preserves a committed cancellation', async () => { + mockDdbSend.mockResolvedValueOnce({ Item: { ...vmTask(), status: 'CANCELLED', memory_written: true } }) + .mockResolvedValue({}); + await finalizeTask('TASK001', { + attempts: 3, microvmSupervisor: supervisor, microvmFailureReason: 'resume-request-failed-repeatedly', + }, 'user-123'); + expect(mockDdbSend.mock.calls.some(([command]) => command.input.ExpressionAttributeValues?.[':toStatus'])).toBe(false); + expect(mockDdbSend.mock.calls.find(([command]) => command._type === 'Put')?.[0].input.Item.event_type).toBe('task_cancelled'); + }); + + test('changed worker ownership skips both task mutation and reservation release', async () => { + mockDdbSend.mockResolvedValueOnce({ Item: { ...vmTask(), session_id: 'replacement' } }); + await expect(finalizeTask('TASK001', { + attempts: 3, microvmSupervisor: supervisor, microvmFailureReason: 'substrate-read-failed', + }, 'user-123')).resolves.toBe(false); + expect(mockDdbSend).toHaveBeenCalledTimes(1); + expect(mockReleaseTaskSlot).not.toHaveBeenCalled(); + }); + + test.each([['RUNNING', 'TIMED_OUT'], ['AWAITING_APPROVAL', 'FAILED'], ['HYDRATING', 'FAILED']])( + 'absolute session expiry in %s uses the allowed %s task outcome', async (status, expected) => { + mockDdbSend.mockResolvedValueOnce({ Item: { ...vmTask(), status } }).mockResolvedValue({}); + await finalizeTask('TASK001', { + attempts: 2000, microvmSupervisor: supervisor, microvmFailureReason: 'session-deadline', + }, 'user-123'); + expect(mockDdbSend.mock.calls[1][0].input.ExpressionAttributeValues[':toStatus']).toBe(expected); + }, + ); + + test('the legacy approval poll timeout also uses FAILED without changing the approval decision', async () => { + mockDdbSend.mockResolvedValueOnce({ Item: vmTask() }).mockResolvedValue({}); + await finalizeTask('TASK001', { attempts: 1020 }, 'user-123'); + expect(mockDdbSend.mock.calls[1][0].input.ExpressionAttributeValues[':toStatus']).toBe('FAILED'); + expect(mockDdbSend.mock.calls[2][0].input.Item.event_type).toBe('task_failed'); + }); + + test.each([ + ['FAILED', 'HYDRATING'], + ['CANCELLED', 'RUNNING'], + ])('honors committed %s instead of stale %s during immediate finalization', async (committed, stale) => { + mockDdbSend.mockImplementation(async (command) => { + if (command._type === 'Get') { + return { + Item: { + ...baseTask, + status: command.input.ConsistentRead ? committed : stale, + memory_written: true, + }, + }; + } + // A stale transition loses to the terminal status already in DynamoDB. + if (command.input.ExpressionAttributeValues?.[':fromStatus']) { + throw Object.assign(new Error('terminal state already committed'), { name: 'ConditionalCheckFailedException' }); + } + return {}; + }); + await finalizeTask('TASK001', { attempts: 0 }, 'user-123'); + const events = mockDdbSend.mock.calls + .filter(([command]) => command._type === 'Put') + .map(([command]) => command.input.Item); + expect(events).toEqual([expect.objectContaining({ + event_type: `task_${committed.toLowerCase()}`, + metadata: expect.objectContaining({ final_status: committed }), + })]); + expect(mockDdbSend.mock.calls.some(([command]) => + command.input.ExpressionAttributeValues?.[':fromStatus'])).toBe(false); + }); + test('handles already-terminal task', async () => { mockDdbSend .mockResolvedValueOnce({ Item: { ...baseTask, status: 'COMPLETED' } }) // loadTask - .mockResolvedValue({}); // emitTaskEvent + decrementConcurrency + .mockResolvedValue({}); // emitTaskEvent await finalizeTask('TASK001', { attempts: 10, lastStatus: 'COMPLETED' }, 'user-123'); // Verify emitTaskEvent was called (PutCommand) expect(mockDdbSend).toHaveBeenCalled(); @@ -1204,7 +1343,7 @@ describe('finalizeTask', () => { test('transitions RUNNING to FAILED when pollState.sessionUnhealthy', async () => { mockDdbSend .mockResolvedValueOnce({ Item: { ...baseTask, status: 'RUNNING' } }) // loadTask - .mockResolvedValue({}); // transitionTask + emitTaskEvent + decrementConcurrency + .mockResolvedValue({}); // transitionTask + emitTaskEvent await finalizeTask( 'TASK001', { attempts: 12, lastStatus: 'RUNNING', sessionUnhealthy: true }, @@ -1296,7 +1435,7 @@ describe('finalizeTask', () => { test('transitions RUNNING to TIMED_OUT on poll timeout', async () => { mockDdbSend .mockResolvedValueOnce({ Item: { ...baseTask, status: 'RUNNING' } }) // loadTask - .mockResolvedValue({}); // transitionTask + emitTaskEvent + decrementConcurrency + .mockResolvedValue({}); // transitionTask + emitTaskEvent await finalizeTask('TASK001', { attempts: 1020, lastStatus: 'RUNNING' }, 'user-123'); expect(mockDdbSend).toHaveBeenCalled(); }); @@ -1304,7 +1443,7 @@ describe('finalizeTask', () => { test('transitions HYDRATING to FAILED when session never started', async () => { mockDdbSend .mockResolvedValueOnce({ Item: { ...baseTask, status: 'HYDRATING' } }) // loadTask - .mockResolvedValue({}); // transitionTask + emitTaskEvent + decrementConcurrency + .mockResolvedValue({}); // transitionTask + emitTaskEvent await finalizeTask('TASK001', { attempts: 15, lastStatus: 'HYDRATING' }, 'user-123'); // First call: loadTask, second call: transitionTask (HYDRATING -> FAILED) const transitionCall = mockDdbSend.mock.calls[1][0]; @@ -1316,34 +1455,32 @@ describe('finalizeTask', () => { expect(eventCall.input.Item.metadata.reason).toBe('session_never_started'); }); - test('releases concurrency for unexpected task state', async () => { - mockDdbSend - .mockResolvedValueOnce({ Item: { ...baseTask, status: 'SUBMITTED' } }) // loadTask - .mockResolvedValue({}); // decrementConcurrency - await finalizeTask('TASK001', { attempts: 5, lastStatus: 'SUBMITTED' }, 'user-123'); - // Should still call decrementConcurrency (UpdateCommand for user concurrency) - const lastCall = mockDdbSend.mock.calls[mockDdbSend.mock.calls.length - 1][0]; - expect(lastCall.input.Key).toEqual({ user_id: 'user-123' }); + test('an unexpected active state makes finalization retry instead of reporting success', async () => { + mockDdbSend.mockResolvedValueOnce({ Item: { ...baseTask, status: 'SUBMITTED' } }); + mockReleaseTaskSlot.mockResolvedValue(false); + await expect(finalizeTask('TASK001', { attempts: 5 }, 'user-123')).rejects.toThrow('Cannot finalize'); + expect(mockDdbSend.mock.calls.some(([command]) => command._type === 'Put')).toBe(false); }); - test('resolves without throwing when decrementConcurrency hits ConditionalCheckFailedException', async () => { - const condErr = new Error('Conditional check failed'); - condErr.name = 'ConditionalCheckFailedException'; - mockDdbSend - .mockResolvedValueOnce({ Item: { ...baseTask, status: 'COMPLETED', memory_written: true } }) // loadTask - .mockResolvedValueOnce({}) // TTL stamp - .mockResolvedValueOnce({}) // emitTaskEvent - .mockRejectedValueOnce(condErr); // decrementConcurrency CCF - await expect(finalizeTask('TASK001', { attempts: 10, lastStatus: 'COMPLETED' }, 'user-123')).resolves.toBeUndefined(); + test('already-released reservations allow finalization to finish', async () => { + mockDdbSend.mockResolvedValueOnce({ Item: { ...baseTask, status: 'COMPLETED', memory_written: true } }) + .mockResolvedValue({}); + mockReleaseTaskSlot.mockResolvedValue(false); + await expect(finalizeTask('TASK001', { attempts: 10 }, 'user-123')).resolves.toBe(true); }); - test('resolves without throwing when decrementConcurrency hits a non-CCF error', async () => { - mockDdbSend - .mockResolvedValueOnce({ Item: { ...baseTask, status: 'COMPLETED', memory_written: true } }) // loadTask - .mockResolvedValueOnce({}) // TTL stamp - .mockResolvedValueOnce({}) // emitTaskEvent - .mockRejectedValueOnce(new Error('DDB timeout')); // decrementConcurrency non-CCF - await expect(finalizeTask('TASK001', { attempts: 10, lastStatus: 'COMPLETED' }, 'user-123')).resolves.toBeUndefined(); + test('a reservation release outage propagates for retry', async () => { + mockDdbSend.mockResolvedValueOnce({ Item: { ...baseTask, status: 'COMPLETED', memory_written: true } }) + .mockResolvedValue({}); + mockReleaseTaskSlot.mockRejectedValue(new Error('DDB timeout')); + await expect(finalizeTask('TASK001', { attempts: 10 }, 'user-123')).rejects.toThrow('DDB timeout'); + }); + + test('a terminal-event failure still attempts reservation release', async () => { + mockDdbSend.mockResolvedValueOnce({ Item: { ...baseTask, status: 'COMPLETED', memory_written: true } }) + .mockResolvedValueOnce({}).mockRejectedValueOnce(new Error('event unavailable')); + await expect(finalizeTask('TASK001', { attempts: 10 }, 'user-123')).rejects.toThrow('event unavailable'); + expect(mockReleaseTaskSlot).toHaveBeenCalledWith('TASK001', 'user-123'); }); }); @@ -1357,8 +1494,7 @@ describe('failTask', () => { test('releases concurrency when requested', async () => { mockDdbSend.mockResolvedValue({}); await failTask('TASK001', 'HYDRATING', 'hydration error', 'user-123', true); - // transitionTask + emitTaskEvent + decrementConcurrency = 3 calls - expect(mockDdbSend).toHaveBeenCalledTimes(3); + expect(mockReleaseTaskSlot).toHaveBeenCalledWith('TASK001', 'user-123'); }); test('transitions from HYDRATING to FAILED when called with HYDRATING status', async () => { @@ -1370,24 +1506,29 @@ describe('failTask', () => { expect(transitionCall.input.ExpressionAttributeValues[':toStatus']).toBe('FAILED'); }); - test('handles transition failure gracefully without emitting when not transitioned', async () => { - mockDdbSend.mockRejectedValueOnce(new Error('Condition failed')); // transitionTask only + test('a committed terminal winner is accepted without another failure event', async () => { + mockDdbSend.mockRejectedValueOnce(new Error('Condition failed')) + .mockResolvedValueOnce({ Item: { ...baseTask, status: 'CANCELLED' } }); await expect(failTask('TASK001', 'SUBMITTED', 'error', 'user-123', false)).resolves.toBeUndefined(); - expect(mockDdbSend).toHaveBeenCalledTimes(1); + expect(mockDdbSend.mock.calls.some(([command]) => command._type === 'Put')).toBe(false); + }); + + test('failure to persist a terminal state propagates without releasing an active task', async () => { + mockDdbSend.mockRejectedValueOnce(new Error('write unavailable')) + .mockResolvedValueOnce({ Item: { ...baseTask, status: 'HYDRATING' } }); + await expect(failTask('TASK001', 'HYDRATING', 'error', 'user-123', true)).rejects.toThrow('write unavailable'); + expect(mockReleaseTaskSlot).not.toHaveBeenCalled(); }); - test('second failTask does not re-emit or re-decrement when transition fails (idempotent under step retry)', async () => { + test('failure replay rechecks release without duplicating the failure event', async () => { mockDdbSend.mockResolvedValue({}); await failTask('TASK001', 'HYDRATING', 'first failure', 'user-123', true); - expect(mockDdbSend).toHaveBeenCalledTimes(3); // transition + emit + decrement - mockDdbSend.mockClear(); - const condErr = new Error('The conditional request failed'); - condErr.name = 'ConditionalCheckFailedException'; - mockDdbSend.mockRejectedValueOnce(condErr); // already FAILED — transition no-ops - - await expect(failTask('TASK001', 'HYDRATING', 'durable replay', 'user-123', true)).resolves.toBeUndefined(); - expect(mockDdbSend).toHaveBeenCalledTimes(1); // transition attempt only; no Put, no concurrency Update + mockDdbSend.mockRejectedValueOnce(new Error('already FAILED')) + .mockResolvedValueOnce({ Item: { ...baseTask, status: 'FAILED' } }); + await expect(failTask('TASK001', 'HYDRATING', 'replay', 'user-123', true)).resolves.toBeUndefined(); + expect(mockDdbSend.mock.calls.some(([command]) => command._type === 'Put')).toBe(false); + expect(mockReleaseTaskSlot).toHaveBeenCalledTimes(2); }); }); @@ -1485,7 +1626,7 @@ describe('finalizeTask TTL stamping', () => { test('stamps TTL on task already in terminal state', async () => { mockDdbSend .mockResolvedValueOnce({ Item: { ...baseTask, status: 'COMPLETED' } }) // loadTask - .mockResolvedValue({}); // UpdateCommand (TTL stamp) + emitTaskEvent + decrementConcurrency + .mockResolvedValue({}); // UpdateCommand (TTL stamp) + emitTaskEvent await finalizeTask('TASK001', { attempts: 10, lastStatus: 'COMPLETED' }, 'user-123'); // Second call should be the TTL stamp UpdateCommand const ttlStampCall = mockDdbSend.mock.calls[1][0]; @@ -1498,7 +1639,7 @@ describe('finalizeTask TTL stamping', () => { mockDdbSend .mockResolvedValueOnce({ Item: { ...baseTask, status: 'COMPLETED' } }) // loadTask .mockRejectedValueOnce(new Error('DDB error')) // TTL stamp fails - .mockResolvedValue({}); // emitTaskEvent + decrementConcurrency + .mockResolvedValue({}); // emitTaskEvent // Should not throw await finalizeTask('TASK001', { attempts: 10, lastStatus: 'COMPLETED' }, 'user-123'); expect(mockDdbSend).toHaveBeenCalled(); @@ -1597,9 +1738,7 @@ describe('finalizeTask — memory fallback', () => { // Should not throw — the try-catch in finalizeTask prevents crash await finalizeTask('TASK001', { attempts: 10, lastStatus: 'COMPLETED' }, 'user-123'); expect(mockWriteMinimalEpisode).toHaveBeenCalled(); - // decrementConcurrency should still be called (last UpdateCommand for user concurrency) - const lastCall = mockDdbSend.mock.calls[mockDdbSend.mock.calls.length - 1][0]; - expect(lastCall.input.Key).toEqual({ user_id: 'user-123' }); + expect(mockReleaseTaskSlot).toHaveBeenCalledWith('TASK001', 'user-123'); }); test('converts string duration_s and cost_usd to numbers', async () => { diff --git a/cdk/test/handlers/reconcile-admission-queue.test.ts b/cdk/test/handlers/reconcile-admission-queue.test.ts index 6d3983fe3..38f4894de 100644 --- a/cdk/test/handlers/reconcile-admission-queue.test.ts +++ b/cdk/test/handlers/reconcile-admission-queue.test.ts @@ -279,7 +279,7 @@ describe('reconcile-admission-queue — races and failures', () => { expect(updates[0].input.ConditionExpression).toBe('#s = :queued'); expect(updates[0].input.ExpressionAttributeValues[':submitted']).toEqual({ S: 'SUBMITTED' }); // Restore: guarded on SUBMITTED, sets QUEUED with a fresh QUEUED# status key. - expect(updates[1].input.ConditionExpression).toBe('#s = :submitted'); + expect(updates[1].input.ConditionExpression).toBe('#s = :submitted AND attribute_not_exists(concurrency_slot)'); expect(updates[1].input.ExpressionAttributeValues[':queued']).toEqual({ S: 'QUEUED' }); expect(updates[1].input.ExpressionAttributeValues[':sca'].S).toMatch(/^QUEUED#/); }); diff --git a/cdk/test/handlers/reconcile-concurrency.test.ts b/cdk/test/handlers/reconcile-concurrency.test.ts index d93e1bf9d..ae58dc687 100644 --- a/cdk/test/handlers/reconcile-concurrency.test.ts +++ b/cdk/test/handlers/reconcile-concurrency.test.ts @@ -17,214 +17,127 @@ * SOFTWARE. */ -// --- Mocks --- -const mockDdbSend = jest.fn(); -jest.mock('@aws-sdk/client-dynamodb', () => ({ - DynamoDBClient: jest.fn(() => ({ send: mockDdbSend })), - ScanCommand: jest.fn((input: unknown) => ({ _type: 'Scan', input })), - QueryCommand: jest.fn((input: unknown) => ({ _type: 'Query', input })), - UpdateItemCommand: jest.fn((input: unknown) => ({ _type: 'UpdateItem', input })), +const mockSend = jest.fn(); +const mockRelease = jest.fn(); +jest.mock('@aws-sdk/lib-dynamodb', () => ({ + ScanCommand: jest.fn((input: unknown) => ({ kind: 'scan', input })), + UpdateCommand: jest.fn((input: unknown) => ({ kind: 'update', input })), +})); +jest.mock('../../src/handlers/shared/ua', () => ({ makeDocClient: () => ({ send: mockSend }) })); +jest.mock('../../src/handlers/shared/task-concurrency', () => ({ + releaseTaskSlot: (...args: unknown[]) => mockRelease(...args), })); - -// Set env vars before importing process.env.TASK_TABLE_NAME = 'Tasks'; -process.env.USER_CONCURRENCY_TABLE_NAME = 'UserConcurrency'; - +process.env.USER_CONCURRENCY_TABLE_NAME = 'Counters'; import { handler } from '../../src/handlers/reconcile-concurrency'; +function held(id: string, user = 'user', status = 'RUNNING') { + return { task_id: id, user_id: user, status, concurrency_slot: { state: 'held' } }; +} +function seed(counters: object[], tasks: object[]) { + mockSend.mockResolvedValueOnce({ Items: counters }).mockResolvedValueOnce({ Items: tasks }).mockResolvedValue({}); +} +function updates() { + return mockSend.mock.calls.filter(([command]) => command.kind === 'update').map(([command]) => command.input); +} beforeEach(() => { - jest.clearAllMocks(); + mockSend.mockReset(); + mockRelease.mockReset().mockResolvedValue(true); }); -describe('reconcile-concurrency handler', () => { - test('completes without errors when no users exist', async () => { - mockDdbSend.mockResolvedValueOnce({ Items: [], LastEvaluatedKey: undefined }); - await handler(); - expect(mockDdbSend).toHaveBeenCalledTimes(1); - }); - - test('no update when stored count matches actual count', async () => { - // Scan returns one user with active_count=2 - mockDdbSend - .mockResolvedValueOnce({ - Items: [{ user_id: { S: 'user-1' }, active_count: { N: '2' } }], - LastEvaluatedKey: undefined, - }) - // Query returns count=2 (matching) - .mockResolvedValueOnce({ Count: 2, LastEvaluatedKey: undefined }); - - await handler(); - - // 1 scan + 1 query = 2 calls, no UpdateItemCommand - expect(mockDdbSend).toHaveBeenCalledTimes(2); - const calls = mockDdbSend.mock.calls; - const updateCalls = calls.filter((c: any[]) => c[0]._type === 'UpdateItem'); - expect(updateCalls).toHaveLength(0); - }); - - test('fires UpdateItemCommand with ConditionExpression when count drifts', async () => { - // Scan returns one user with active_count=5 - mockDdbSend - .mockResolvedValueOnce({ - Items: [{ user_id: { S: 'user-1' }, active_count: { N: '5' } }], - LastEvaluatedKey: undefined, - }) - // Query returns actual count=2 (drift) - .mockResolvedValueOnce({ Count: 2, LastEvaluatedKey: undefined }) - // Update succeeds - .mockResolvedValueOnce({}); - - await handler(); - - // 1 scan + 1 query + 1 update = 3 calls - expect(mockDdbSend).toHaveBeenCalledTimes(3); - const updateCall = mockDdbSend.mock.calls[2][0]; - expect(updateCall._type).toBe('UpdateItem'); - // Verify ConditionExpression for TOCTOU protection - expect(updateCall.input.ConditionExpression).toBe('active_count = :stored'); - expect(updateCall.input.ExpressionAttributeValues[':stored']).toEqual({ N: '5' }); - expect(updateCall.input.ExpressionAttributeValues[':count']).toEqual({ N: '2' }); - }); - - test('continues to next user on ConditionalCheckFailedException', async () => { - const condErr = new Error('Conditional check failed'); - condErr.name = 'ConditionalCheckFailedException'; - - mockDdbSend - // Scan returns two users - .mockResolvedValueOnce({ - Items: [ - { user_id: { S: 'user-1' }, active_count: { N: '5' } }, - { user_id: { S: 'user-2' }, active_count: { N: '3' } }, - ], - LastEvaluatedKey: undefined, - }) - // User-1: query returns 2 (drift) - .mockResolvedValueOnce({ Count: 2, LastEvaluatedKey: undefined }) - // User-1: update fails with CCF - .mockRejectedValueOnce(condErr) - // User-2: query returns 1 (drift) - .mockResolvedValueOnce({ Count: 1, LastEvaluatedKey: undefined }) - // User-2: update succeeds - .mockResolvedValueOnce({}); - - await handler(); - - // 1 scan + 2 queries + 2 updates = 5 calls - expect(mockDdbSend).toHaveBeenCalledTimes(5); - }); - - test('continues to next user when query fails', async () => { - mockDdbSend - // Scan returns two users - .mockResolvedValueOnce({ - Items: [ - { user_id: { S: 'user-1' }, active_count: { N: '2' } }, - { user_id: { S: 'user-2' }, active_count: { N: '3' } }, - ], - LastEvaluatedKey: undefined, - }) - // User-1: query throws - .mockRejectedValueOnce(new Error('DynamoDB timeout')) - // User-2: query returns 3 (matches) - .mockResolvedValueOnce({ Count: 3, LastEvaluatedKey: undefined }); - - await handler(); - - // 1 scan + 2 queries (one failed) = 3 calls - expect(mockDdbSend).toHaveBeenCalledTimes(3); - }); +test('empty tables complete without writes', async () => { + seed([], []); + await handler(); + expect(updates()).toEqual([]); + expect(mockRelease).not.toHaveBeenCalled(); +}); - test('handles scan pagination', async () => { - mockDdbSend - // First scan page - .mockResolvedValueOnce({ - Items: [{ user_id: { S: 'user-1' }, active_count: { N: '1' } }], - LastEvaluatedKey: { user_id: { S: 'user-1' } }, - }) - // User-1 query - .mockResolvedValueOnce({ Count: 1, LastEvaluatedKey: undefined }) - // Second scan page - .mockResolvedValueOnce({ - Items: [{ user_id: { S: 'user-2' }, active_count: { N: '2' } }], - LastEvaluatedKey: undefined, - }) - // User-2 query - .mockResolvedValueOnce({ Count: 2, LastEvaluatedKey: undefined }); +test('approval waits count as held seats and a matching counter needs no repair', async () => { + seed([{ user_id: 'user', active_count: 2, reservation_version: 'v1' }], + [held('one'), held('two', 'user', 'AWAITING_APPROVAL')]); + await handler(); + expect(updates()).toEqual([]); + expect(mockRelease).not.toHaveBeenCalled(); +}); - await handler(); +test('repairs drift only while the observed count and reservation revision still match', async () => { + seed([{ user_id: 'user', active_count: 5, reservation_version: 'v1' }], [held('one'), held('two')]); + await handler(); + expect(updates()).toEqual([expect.objectContaining({ + Key: { user_id: 'user' }, + ConditionExpression: 'attribute_exists(user_id) AND active_count = :stored AND reservation_version = :observed', + ExpressionAttributeValues: expect.objectContaining({ ':count': 2, ':stored': 5, ':observed': 'v1' }), + })]); +}); - // 2 scans + 2 queries = 4 calls - expect(mockDdbSend).toHaveBeenCalledTimes(4); +test('missing counters are recreated only if no reservation writer has created them meanwhile', async () => { + seed([], [held('one')]); + await handler(); + expect(updates()[0]).toMatchObject({ + ConditionExpression: 'attribute_not_exists(user_id)', + ExpressionAttributeValues: { ':count': 1 }, }); +}); - test('handles query pagination for countActiveTasks', async () => { - mockDdbSend - // Scan - .mockResolvedValueOnce({ - Items: [{ user_id: { S: 'user-1' }, active_count: { N: '0' } }], - LastEvaluatedKey: undefined, - }) - // First query page: count=3, has more - .mockResolvedValueOnce({ - Count: 3, - LastEvaluatedKey: { user_id: { S: 'user-1' }, task_id: { S: 'T3' } }, - }) - // Second query page: count=2, done - .mockResolvedValueOnce({ Count: 2, LastEvaluatedKey: undefined }) - // Update (drift: stored=0 vs actual=5) - .mockResolvedValueOnce({}); - - await handler(); +test('legacy counter repair checks that a version has not been installed concurrently', async () => { + seed([{ user_id: 'user', active_count: 4 }], [held('one')]); + await handler(); + expect(updates()[0].ConditionExpression).toContain('attribute_not_exists(reservation_version)'); +}); - // 1 scan + 2 queries + 1 update = 4 calls - expect(mockDdbSend).toHaveBeenCalledTimes(4); - const updateCall = mockDdbSend.mock.calls[3][0]; - expect(updateCall._type).toBe('UpdateItem'); - expect(updateCall.input.ExpressionAttributeValues[':count']).toEqual({ N: '5' }); - }); +test('ambiguous legacy active tasks prevent guessing a count', async () => { + seed([{ user_id: 'user', active_count: 5 }], [{ task_id: 'legacy', user_id: 'user', status: 'AWAITING_APPROVAL' }]); + await handler(); + expect(updates()).toEqual([]); +}); - test('skips items without user_id', async () => { - mockDdbSend.mockResolvedValueOnce({ - Items: [{ active_count: { N: '1' } }], // no user_id - LastEvaluatedKey: undefined, - }); +test('queued and unadmitted terminal tasks do not inflate the count', async () => { + seed([{ user_id: 'user', active_count: 3 }], [ + { task_id: 'queued', user_id: 'user', status: 'QUEUED' }, + { task_id: 'failed', user_id: 'user', status: 'FAILED' }, + ]); + await handler(); + expect(updates()[0].ExpressionAttributeValues[':count']).toBe(0); +}); - await handler(); +test('counts terminal held markers before asking shared cleanup to release them', async () => { + seed([{ user_id: 'user', active_count: 0, reservation_version: 'v1' }], [held('done', 'user', 'FAILED'), held('live')]); + await handler(); + expect(updates()[0].ExpressionAttributeValues[':count']).toBe(2); + expect(mockRelease).toHaveBeenCalledWith('done', 'user'); + expect(mockSend.mock.invocationCallOrder.at(-1)!).toBeLessThan(mockRelease.mock.invocationCallOrder[0]); +}); - // Only the scan call, no query or update - expect(mockDdbSend).toHaveBeenCalledTimes(1); - }); +test('a changed revision skips stale repair and still tries terminal cleanup', async () => { + seed([{ user_id: 'user', active_count: 3, reservation_version: 'v1' }], [held('done', 'user', 'FAILED')]); + mockSend.mockRejectedValueOnce(Object.assign(new Error('changed'), { name: 'ConditionalCheckFailedException' })); + await handler(); + expect(updates()).toHaveLength(1); + expect(mockRelease).toHaveBeenCalledWith('done', 'user'); +}); - test('continues to next user when UpdateItemCommand fails with non-CCF error', async () => { - mockDdbSend - // Scan returns two users with drift - .mockResolvedValueOnce({ - Items: [ - { user_id: { S: 'user-1' }, active_count: { N: '5' } }, - { user_id: { S: 'user-2' }, active_count: { N: '4' } }, - ], - LastEvaluatedKey: undefined, - }) - // User-1: query returns 2 (drift) - .mockResolvedValueOnce({ Count: 2, LastEvaluatedKey: undefined }) - // User-1: update fails with non-CCF error - .mockRejectedValueOnce(new Error('InternalServerError')) - // User-2: query returns 1 (drift) - .mockResolvedValueOnce({ Count: 1, LastEvaluatedKey: undefined }) - // User-2: update succeeds - .mockResolvedValueOnce({}); +test('one user repair failure does not stop the next user', async () => { + seed([{ user_id: 'one', active_count: 3 }, { user_id: 'two', active_count: 3 }], []); + mockSend.mockRejectedValueOnce(new Error('unavailable')); + await handler(); + expect(updates().map(update => update.Key.user_id)).toEqual(['one', 'two']); +}); - await handler(); +test('scans every counter page before strongly scanning every task page', async () => { + mockSend.mockResolvedValueOnce({ Items: [{ user_id: 'user', active_count: 2 }], LastEvaluatedKey: { user_id: 'user' } }) + .mockResolvedValueOnce({ Items: [] }) + .mockResolvedValueOnce({ Items: [held('one')], LastEvaluatedKey: { task_id: 'one' } }) + .mockResolvedValueOnce({ Items: [held('two')] }); + await handler(); + expect(mockSend.mock.calls.map(([command]) => [command.input.TableName, command.input.ConsistentRead])) + .toEqual([['Counters', true], ['Counters', true], ['Tasks', true], ['Tasks', true]]); + expect(mockSend.mock.calls[3][0].input.ExclusiveStartKey).toEqual({ task_id: 'one' }); +}); - // 1 scan + 2 queries + 2 updates = 5 calls - expect(mockDdbSend).toHaveBeenCalledTimes(5); - // Verify user-2's update was still attempted and succeeded - const updateCalls = mockDdbSend.mock.calls.filter((c: any[]) => c[0]._type === 'UpdateItem'); - expect(updateCalls).toHaveLength(2); - // User-2's update should have the correct values - const user2Update = updateCalls[1][0]; - expect(user2Update.input.Key).toEqual({ user_id: { S: 'user-2' } }); - expect(user2Update.input.ExpressionAttributeValues[':count']).toEqual({ N: '1' }); - }); +test('an incomplete task scan aborts before any partial count is installed', async () => { + mockSend.mockResolvedValueOnce({ Items: [{ user_id: 'user', active_count: 2 }] }) + .mockResolvedValueOnce({ Items: [held('one')], LastEvaluatedKey: { task_id: 'one' } }) + .mockRejectedValueOnce(new Error('scan unavailable')); + await expect(handler()).rejects.toThrow('scan unavailable'); + expect(updates()).toEqual([]); }); diff --git a/cdk/test/handlers/reconcile-microvm-continuations.test.ts b/cdk/test/handlers/reconcile-microvm-continuations.test.ts new file mode 100644 index 000000000..30a01a351 --- /dev/null +++ b/cdk/test/handlers/reconcile-microvm-continuations.test.ts @@ -0,0 +1,255 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +const mockSend = jest.fn(); +const mockStop = jest.fn(); +const mockPoll = jest.fn(); +const mockClose = jest.fn(); +const mockDelete = jest.fn(); +const mockRelease = jest.fn(); +const mockRetire = jest.fn(); +const mockDispatch = jest.fn(); +jest.mock('../../src/handlers/shared/ua', () => ({ makeDocClient: () => ({ send: mockSend }) })); +jest.mock('../../src/handlers/shared/strategies/lambda-microvm-strategy', () => ({ + LambdaMicrovmComputeStrategy: jest.fn(() => ({ stopSession: mockStop, pollSession: mockPoll })), + MICROVM_MAX_DURATION_SECONDS: 28800, +})); +jest.mock('../../src/handlers/shared/close-task-approvals', () => ({ closeTaskApprovals: (...args: unknown[]) => mockClose(...args) })); +jest.mock('../../src/handlers/shared/task-concurrency', () => ({ releaseTaskSlot: (...args: unknown[]) => mockRelease(...args) })); +jest.mock('../../src/handlers/shared/microvm-continuation-storage', () => ({ + deleteClosedTaskContinuations: (...args: unknown[]) => mockDelete(...args), +})); +jest.mock('../../src/handlers/shared/microvm-continuation-retirement', () => ({ + retireCheckpointedMicrovm: (...args: unknown[]) => mockRetire(...args), +})); +jest.mock('../../src/handlers/shared/microvm-continuation-dispatch', () => ({ + dispatchMicrovmContinuation: (...args: unknown[]) => mockDispatch(...args), +})); +jest.mock('../../src/handlers/shared/logger', () => ({ logger: { warn: jest.fn(), info: jest.fn() } })); + +import { handler, reconcileMicrovmContinuation } from '../../src/handlers/reconcile-microvm-continuations'; + +let task: any; +beforeEach(() => { + jest.resetAllMocks(); + task = { + task_id: 'task', + user_id: 'user', + status: 'AWAITING_APPROVAL', + awaiting_approval_request_id: 'request', + continuation: { state: 'PARKED' }, + continuation_launch: {}, + microvm_start: { clientToken: 'attempt', createdAt: new Date().toISOString() }, + }; + mockSend.mockResolvedValue({}); +}); + +test('a parked unanswered task is offered for admission without waking its retired worker', async () => { + await reconcileMicrovmContinuation(task); + expect(mockDispatch).toHaveBeenCalledWith('task', 'user', 'request', expect.any(Object)); + expect(mockStop).not.toHaveBeenCalled(); + expect(mockDelete).not.toHaveBeenCalled(); +}); + +test('a terminated source triggers verified retirement before dispatch', async () => { + task.continuation.state = 'READY'; + task.microvm_start.handle = { microvmId: 'old' }; + mockPoll.mockResolvedValue({ microvmState: 'TERMINATED' }); + await reconcileMicrovmContinuation(task); + expect(mockRetire).toHaveBeenCalledWith(expect.objectContaining({ force: true, handle: { microvmId: 'old' } })); + expect(mockRetire.mock.invocationCallOrder[0]).toBeLessThan(mockDispatch.mock.invocationCallOrder[0]); +}); + +test('unknown start keeps artifacts and capacity until its maximum service lifetime', async () => { + task.status = 'FAILED'; + mockSend.mockResolvedValue({ Item: task }); + await reconcileMicrovmContinuation(task); + expect(mockClose).toHaveBeenCalled(); + expect(mockRelease).not.toHaveBeenCalled(); + expect(mockDelete).not.toHaveBeenCalled(); + task.microvm_start.createdAt = new Date(Date.now() - 29_000_000).toISOString(); + await reconcileMicrovmContinuation(task); + expect(mockRelease).toHaveBeenCalledWith('task', 'user'); + expect(mockDelete).toHaveBeenCalled(); + const closed = mockSend.mock.calls.find(([command]) => command.constructor.name === 'UpdateCommand')![0].input; + expect(closed.ExpressionAttributeValues).toMatchObject({ ':closed': 'CLOSED', ':attempt': 'attempt' }); +}); + +test('unconfirmed termination retains artifacts and reservation', async () => { + task.status = 'CANCELLED'; + task.microvm_start.handle = { microvmId: 'vm' }; + mockSend.mockResolvedValue({ Item: task }); + mockPoll.mockResolvedValue({ microvmState: 'TERMINATING' }); + await reconcileMicrovmContinuation(task); + expect(mockStop).toHaveBeenCalledWith({ microvmId: 'vm' }, expect.any(Object)); + expect(mockRelease).not.toHaveBeenCalled(); + expect(mockDelete).not.toHaveBeenCalled(); +}); + +test('terminal scan rows cannot stop a task whose current record is active', async () => { + mockSend.mockResolvedValue({ Item: task }); + await reconcileMicrovmContinuation({ ...task, status: 'FAILED' }); + expect(mockStop).not.toHaveBeenCalled(); + expect(mockClose).not.toHaveBeenCalled(); +}); + +test('closes a retired source using its preserved lease token after microvm_start was removed', async () => { + task.status = 'CANCELLED'; + delete task.microvm_start; + mockSend.mockImplementation(async command => { + if (command.constructor.name === 'UpdateCommand') return {}; + return { + Item: command.input.Key.task_id === 'task' ? task : { + lease_user_id: 'user', lease_attempt_id: 'retired-launch-token', lease_state: 'PARKED', + }, + }; + }); + await reconcileMicrovmContinuation(task); + const closed = mockSend.mock.calls.find(([command]) => command.constructor.name === 'UpdateCommand')![0].input; + expect(closed.ExpressionAttributeValues[':attempt']).toBe('retired-launch-token'); + expect(closed.ExpressionAttributeValues[':observedState']).toBe('PARKED'); + expect(mockRelease).toHaveBeenCalledWith('task', 'user'); + expect(mockDelete).toHaveBeenCalled(); +}); + +test('releases a cancelled replacement admitted before any start receipt or AWS call', async () => { + task.status = 'CANCELLED'; + delete task.microvm_start; + task.continuation = { state: 'STARTING', attempt_id: 'new-token' }; + task.concurrency_slot = { state: 'held', attempt_id: 'new-token' }; + mockSend.mockImplementation(async command => { + if (command.constructor.name === 'UpdateCommand') return {}; + return { + Item: command.input.Key.task_id === 'task' ? task : { + lease_user_id: 'user', lease_attempt_id: 'new-token', lease_state: 'ACTIVE', + }, + }; + }); + await reconcileMicrovmContinuation(task); + const closed = mockSend.mock.calls.find(([command]) => command.constructor.name === 'UpdateCommand')![0].input; + expect(closed.ExpressionAttributeValues[':attempt']).toBe('new-token'); + expect(closed.ExpressionAttributeValues[':observedState']).toBe('ACTIVE'); + expect(closed.ConditionExpression).toContain('attribute_not_exists(lease_microvm_id)'); + expect(mockRelease).toHaveBeenCalledWith('task', 'user'); + expect(mockDelete).toHaveBeenCalled(); +}); + +test('does not release a live lease just because the start record is absent', async () => { + task.status = 'CANCELLED'; + delete task.microvm_start; + mockSend.mockImplementation(async command => ({ + Item: command.input.Key.task_id === 'task' ? task : { + lease_user_id: 'user', lease_attempt_id: 'live-token', lease_state: 'ACTIVE', + }, + })); + await expect(reconcileMicrovmContinuation(task)).rejects.toThrow('LEASE_INVALID'); + expect(mockRelease).not.toHaveBeenCalled(); + expect(mockDelete).not.toHaveBeenCalled(); +}); + +test.each(['held', 'released'])('cleans a fenced pre-launch failure with a %s slot', async state => { + task.status = 'FAILED'; + delete task.microvm_start; + task.continuation = { state: 'STARTING', attempt_id: 'new-token' }; + task.concurrency_slot = { state, attempt_id: 'new-token' }; + mockSend.mockImplementation(async command => { + if (command.constructor.name === 'UpdateCommand') return {}; + return { + Item: command.input.Key.task_id === 'task' ? task : { + lease_user_id: 'user', lease_attempt_id: 'new-token', lease_state: 'FENCED', + }, + }; + }); + await reconcileMicrovmContinuation(task); + const closed = mockSend.mock.calls.find(([command]) => command.constructor.name === 'UpdateCommand')![0].input; + expect(closed.ExpressionAttributeValues).toMatchObject({ ':attempt': 'new-token', ':observedState': 'FENCED' }); + expect(closed.ConditionExpression).toContain('attribute_not_exists(lease_microvm_id)'); + expect(mockDelete).toHaveBeenCalled(); +}); + +test.each([ + { lease_microvm_id: 'possibly-live-worker' }, + { lease_user_id: 'other-user' }, + { lease_attempt_id: 'other-attempt' }, +])('retains checkpoints for an inconsistent fenced attempt: %j', async mismatch => { + task.status = 'FAILED'; + delete task.microvm_start; + task.continuation = { state: 'STARTING', attempt_id: 'new-token' }; + task.concurrency_slot = { state: 'released', attempt_id: 'new-token' }; + mockSend.mockImplementation(async command => ({ + Item: command.input.Key.task_id === 'task' ? task : { + lease_user_id: 'user', lease_attempt_id: 'new-token', lease_state: 'FENCED', ...mismatch, + }, + })); + await expect(reconcileMicrovmContinuation(task)).rejects.toThrow('LEASE_INVALID'); + expect(mockRelease).not.toHaveBeenCalled(); + expect(mockDelete).not.toHaveBeenCalled(); +}); + +test('invalid unknown-start timestamp cannot be treated as proof of shutdown', async () => { + task.status = 'FAILED'; + task.microvm_start.createdAt = 'invalid'; + mockSend.mockResolvedValue({ Item: task }); + await expect(reconcileMicrovmContinuation(task)).rejects.toThrow('START_TIME_INVALID'); + expect(mockRelease).not.toHaveBeenCalled(); +}); + +test('persists the last completed batch when invocation time is low, then resumes from that key', async () => { + const rows = Array.from({ length: 6 }, (_, index) => ({ ...task, task_id: `task-${index}` })); + mockSend.mockImplementation(async command => { + if (command.constructor.name === 'GetCommand') return { Item: { cursor: { task_id: 'previous' } } }; + if (command.constructor.name === 'ScanCommand') return { Items: rows, LastEvaluatedKey: { task_id: 'page-end' } }; + return {}; + }); + const remaining = jest.fn().mockReturnValueOnce(100000).mockReturnValueOnce(100000).mockReturnValue(40000); + await handler({}, { getRemainingTimeInMillis: remaining }); + expect(mockDispatch).toHaveBeenCalledTimes(4); + expect(mockSend.mock.calls.find(([command]) => command.constructor.name === 'ScanCommand')![0].input.ExclusiveStartKey) + .toEqual({ task_id: 'previous' }); + expect(mockSend.mock.calls.at(-1)![0].input.ExpressionAttributeValues[':cursor']).toEqual({ task_id: 'task-3' }); + expect(mockSend.mock.calls.at(-1)![0].input).toMatchObject({ + ConditionExpression: '#cursor = :previous', + ExpressionAttributeValues: { ':previous': { task_id: 'previous' } }, + }); +}); + +test('a failed row does not prevent later rows or clearing the cursor after a complete scan', async () => { + mockSend.mockImplementation(async command => command.constructor.name === 'ScanCommand' + ? { Items: [task, { ...task, task_id: 'next' }] } : {}); + mockDispatch.mockRejectedValueOnce(new Error('temporary')).mockResolvedValueOnce(true); + await handler({}, { getRemainingTimeInMillis: () => 100000 }); + expect(mockDispatch).toHaveBeenCalledTimes(2); + expect(mockSend.mock.calls.at(-1)![0].input.UpdateExpression).toBe('REMOVE #cursor'); + expect(mockSend.mock.calls.at(-1)![0].input.ConditionExpression).toBe('attribute_not_exists(#cursor)'); +}); + +test.each(['ConditionalCheckFailedException', 'ProvisionedThroughputExceededException'])( + 'only an overlapping sweep may supersede the saved cursor: %s', + async name => { + const error = Object.assign(new Error('cursor write rejected'), { name }); + mockSend.mockImplementation(async command => { + if (command.constructor.name === 'UpdateCommand') throw error; + return {}; + }); + const run = handler({}, { getRemainingTimeInMillis: () => 100000 }); + if (name === 'ConditionalCheckFailedException') await expect(run).resolves.toBeUndefined(); + else await expect(run).rejects.toBe(error); + expect(mockSend.mock.calls.filter(([command]) => command.constructor.name === 'UpdateCommand')).toHaveLength(1); + }, +); diff --git a/cdk/test/handlers/reconcile-stranded-tasks.test.ts b/cdk/test/handlers/reconcile-stranded-tasks.test.ts index 8c2346c5c..f841017c7 100644 --- a/cdk/test/handlers/reconcile-stranded-tasks.test.ts +++ b/cdk/test/handlers/reconcile-stranded-tasks.test.ts @@ -19,6 +19,15 @@ // --- Mocks --- const mockDdbSend = jest.fn(); +const mockRelease = jest.fn(); +const mockCloseApprovals = jest.fn(); +jest.mock('../../src/handlers/shared/close-task-approvals', () => ({ + closeTaskApprovals: (...args: unknown[]) => mockCloseApprovals(...args), +})); +jest.mock('../../src/handlers/shared/task-concurrency', () => ({ + releaseTaskSlot: (...args: unknown[]) => mockRelease(...args), +})); +beforeEach(() => mockRelease.mockReset().mockResolvedValue(true)); jest.mock('@aws-sdk/client-dynamodb', () => ({ DynamoDBClient: jest.fn(() => ({ send: mockDdbSend })), QueryCommand: jest.fn((input: unknown) => ({ _type: 'Query', input })), @@ -89,7 +98,7 @@ describe('reconcile-stranded-tasks', () => { expect(mockDdbSend).toHaveBeenCalledTimes(3); }); - test('task older than 1200s → fails + emits events + decrements concurrency', async () => { + test('task older than 1200s → fails + emits events + releases the task reservation', async () => { const ancient = new Date(Date.now() - 25 * 60 * 1000).toISOString(); // 25 min ago primeResponses([ // Query SUBMITTED returns one stranded candidate. @@ -103,7 +112,6 @@ describe('reconcile-stranded-tasks', () => { {}, // conditional UpdateItem → FAILED {}, // PutItem task_stranded event {}, // PutItem task_failed event - {}, // UpdateItem decrement concurrency { Items: [] }, // Query HYDRATING { Items: [] }, // Query AWAITING_APPROVAL ]); @@ -132,10 +140,7 @@ describe('reconcile-stranded-tasks', () => { }); expect(eventTypes).toEqual(expect.arrayContaining(['task_stranded', 'task_failed'])); - // Concurrency decrement. - const decrementCall = (mockDdbSend.mock.calls as [{ _type: string; input: Record }][]) - .find(([c]) => c._type === 'UpdateItem' && String(c.input.UpdateExpression).includes('active_count')); - expect(decrementCall).toBeDefined(); + expect(mockRelease).toHaveBeenCalledWith('t-stranded', 'u-1'); }); test('#441: task with old created_at but FRESH status_created_at is NOT failed (freshly picked up from the queue)', async () => { @@ -180,7 +185,6 @@ describe('reconcile-stranded-tasks', () => { {}, // conditional UpdateItem → FAILED {}, // PutItem task_stranded event {}, // PutItem task_failed event - {}, // UpdateItem decrement concurrency { Items: [] }, // HYDRATING { Items: [] }, // AWAITING_APPROVAL ]); @@ -215,7 +219,6 @@ describe('reconcile-stranded-tasks', () => { {}, // conditional UpdateItem → FAILED {}, // PutItem task_stranded event {}, // PutItem task_failed event - {}, // UpdateItem decrement concurrency { Items: [] }, // HYDRATING { Items: [] }, // AWAITING_APPROVAL ]); @@ -246,7 +249,7 @@ describe('reconcile-stranded-tasks', () => { { Items: [] }, // AWAITING_APPROVAL query ]); - // Must NOT throw; no events written, no concurrency decrement. + // Must NOT throw; no events written, no reservation release. await handler(); const writes = (mockDdbSend.mock.calls as [{ _type: string; input: Record }][]) @@ -264,7 +267,7 @@ describe('reconcile-stranded-tasks', () => { // (``event_type`` == their own names) never reach it. This test // pins the extra ``agent_milestone`` / ``approval_stranded`` emit // that makes the stranded case visible on §11.3 widgets. - const ancient = new Date(Date.now() - 2 * 3600 * 1000).toISOString(); + const ancient = new Date(Date.now() - 10 * 3600 * 1000).toISOString(); primeResponses([ { Items: [] }, // SUBMITTED { Items: [] }, // HYDRATING @@ -274,7 +277,6 @@ describe('reconcile-stranded-tasks', () => { {}, // PutItem task_stranded event {}, // PutItem task_failed event {}, // PutItem approval_stranded milestone (Chunk 10 new) - {}, // UpdateItem decrement concurrency ]); await handler(); @@ -312,7 +314,6 @@ describe('reconcile-stranded-tasks', () => { {}, // conditional UpdateItem → FAILED {}, // PutItem task_stranded {}, // PutItem task_failed - {}, // UpdateItem concurrency { Items: [] }, // HYDRATING { Items: [] }, // AWAITING_APPROVAL ]); @@ -441,7 +442,6 @@ describe('reconcile-stranded-tasks', () => { {}, // UpdateItem t-ok (transition) → success {}, // PutItem task_stranded event {}, // PutItem task_failed event - {}, // UpdateItem decrement concurrency ddbErr, // UpdateItem t-fail (transition) → throws { Items: [] }, // HYDRATING query { Items: [] }, // AWAITING_APPROVAL query @@ -467,8 +467,8 @@ describe('reconcile-stranded-tasks', () => { mockTaskRow({ task_id: 't-2', user_id: 'u-b', created_at: ancient }), ], }, - {}, {}, {}, {}, // t-1: transition + 2 events + decrement - {}, {}, {}, {}, // t-2: transition + 2 events + decrement + {}, {}, {}, // t-1: transition + 2 events + {}, {}, {}, // t-2: transition + 2 events { Items: [] }, // HYDRATING { Items: [] }, // AWAITING_APPROVAL ]); diff --git a/cdk/test/handlers/request-approval-local.test.ts b/cdk/test/handlers/request-approval-local.test.ts new file mode 100644 index 000000000..ee15a5aa0 --- /dev/null +++ b/cdk/test/handlers/request-approval-local.test.ts @@ -0,0 +1,179 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +/** Real transaction conditions; loopback endpoint and dummy credentials only. */ +import { randomUUID } from 'node:crypto'; +import { CreateTableCommand, DeleteTableCommand, DynamoDBClient } from '@aws-sdk/client-dynamodb'; +import { DynamoDBDocumentClient, GetCommand, PutCommand, TransactWriteCommand, UpdateCommand } from '@aws-sdk/lib-dynamodb'; + +const endpoint = process.env.ABCA_DDB_LOCAL_ENDPOINT; +if (process.env.CI === 'true' && !endpoint) throw new Error('CI requires ABCA_DDB_LOCAL_ENDPOINT'); +if (endpoint && (new URL(endpoint).hostname !== '127.0.0.1' || new URL(endpoint).protocol !== 'http:')) { + throw new Error('Approval integration tests require an http://127.0.0.1 endpoint'); +} +const mockBeforeTransaction = jest.fn(); +let mockClient: DynamoDBDocumentClient; +jest.mock('../../src/handlers/shared/ua', () => ({ + makeDocClient: () => ({ + send: async (command: unknown) => { + if (command instanceof TransactWriteCommand) await mockBeforeTransaction(); + return mockClient.send(command as GetCommand); + }, + }), +})); +const suffix = randomUUID(); +const tasks = `approval-tasks-${suffix}`; +const approvals = `approval-requests-${suffix}`; +Object.assign(process.env, { TASK_TABLE_NAME: tasks, TASK_APPROVALS_TABLE_NAME: approvals }); +import { recordWorkerRequest } from '../../src/handlers/request-approval'; + +const raw = new DynamoDBClient({ + endpoint: endpoint ?? 'http://127.0.0.1:1', + region: 'us-east-1', + credentials: { accessKeyId: 'local', secretAccessKey: 'local' }, +}); +mockClient = DynamoDBDocumentClient.from(raw); +const local = endpoint ? describe : describe.skip; +jest.setTimeout(30_000); + +local('Trusted approval writer against DynamoDB Local', () => { + let taskId: string; + const gate = 'gate'; + const row = () => ({ + task_id: taskId, + request_id: gate, + user_id: 'owner', + repo: 'owner/repo', + tool_name: 'Bash', + tool_input_preview: '{"command":"git push"}', + tool_input_sha256: 'a'.repeat(64), + reason: 'Protected operation', + severity: 'high', + matching_rule_ids: ['protected'], + status: 'PENDING', + created_at: new Date().toISOString(), + timeout_s: 0, + }); + const create = () => recordWorkerRequest({ + operation: 'create', task_id: taskId, request_id: gate, worker_attempt_id: 'worker', approval: row(), + }); + const timeout = () => recordWorkerRequest({ + operation: 'timeout', task_id: taskId, request_id: gate, worker_attempt_id: 'worker', + }); + const read = async (table: string) => (await mockClient.send(new GetCommand({ + TableName: table, + Key: { task_id: taskId, ...(table === approvals ? { request_id: gate } : {}) }, + ConsistentRead: true, + }))).Item; + const change = async (table: string, field: string, value: string, lease = false) => + mockClient.send(new UpdateCommand({ + TableName: table, + Key: { task_id: lease ? `worker-lease#${taskId}` : taskId, ...(table === approvals ? { request_id: gate } : {}) }, + UpdateExpression: 'SET #field = :value', + ExpressionAttributeNames: { '#field': field }, + ExpressionAttributeValues: { ':value': value }, + })); + beforeAll(async () => { + for (const table of [tasks, approvals]) { + await raw.send(new CreateTableCommand({ + TableName: table, + BillingMode: 'PAY_PER_REQUEST', + AttributeDefinitions: [{ AttributeName: 'task_id', AttributeType: 'S' }, + ...(table === approvals ? [{ AttributeName: 'request_id', AttributeType: 'S' as const }] : [])], + KeySchema: [{ AttributeName: 'task_id', KeyType: 'HASH' }, + ...(table === approvals ? [{ AttributeName: 'request_id', KeyType: 'RANGE' as const }] : [])], + })); + } + }); + afterAll(async () => { + try { for (const table of [tasks, approvals]) await raw.send(new DeleteTableCommand({ TableName: table })); } finally { raw.destroy(); } + }); + beforeEach(async () => { + taskId = randomUUID(); + mockBeforeTransaction.mockReset(); + await mockClient.send(new PutCommand({ + TableName: tasks, + Item: { + task_id: taskId, status: 'RUNNING', user_id: 'owner', repo: 'owner/repo', compute_type: 'lambda-microvm', + }, + })); + await mockClient.send(new PutCommand({ + TableName: tasks, + Item: { + task_id: `worker-lease#${taskId}`, lease_state: 'ACTIVE', lease_attempt_id: 'worker', lease_user_id: 'owner', + }, + })); + }); + + test('creates a durable pending request and its task pointer atomically', async () => { + expect(await create()).toEqual({ ok: true }); + expect(await read(approvals)).toMatchObject({ status: 'PENDING', timeout_s: 0 }); + expect(await read(approvals)).not.toHaveProperty('ttl'); + expect(await read(tasks)).toMatchObject({ status: 'AWAITING_APPROVAL', awaiting_approval_request_id: gate }); + }); + test.each(['APPROVED', 'DENIED', 'CANCELLED'])('timeout preserves an existing %s decision', async status => { + await create(); + await change(approvals, 'status', status); + const before = await read(approvals); + expect(await timeout()).toMatchObject({ + ok: false, + code: 'TransactionCanceledException', + cancellation_reasons: [{ Code: 'ConditionalCheckFailed' }, { Code: 'None' }, { Code: 'None' }], + }); + expect(await read(approvals)).toEqual(before); + }); + test('only pending requests can transition to a non-human timeout', async () => { + await create(); + expect(await timeout()).toEqual({ ok: true }); + expect(await read(approvals)).toMatchObject({ status: 'TIMED_OUT' }); + expect(await read(approvals)).not.toHaveProperty('decision_source'); + }); + test('cancellation between read and write prevents both request and task changes', async () => { + mockBeforeTransaction.mockImplementationOnce(() => change(tasks, 'status', 'CANCELLED')); + expect(await create()).toMatchObject({ ok: false, code: 'TransactionCanceledException' }); + expect(await read(approvals)).toBeUndefined(); + expect(await read(tasks)).toMatchObject({ status: 'CANCELLED' }); + }); + test.each(['create', 'timeout'])('retirement revokes a stale worker during %s', async operation => { + if (operation === 'timeout') await create(); + const before = await read(approvals); + mockBeforeTransaction.mockImplementationOnce(() => change(tasks, 'lease_state', 'PARKED', true)); + const result = await (operation === 'create' ? create() : timeout()); + expect(result).toMatchObject({ ok: false, code: 'TransactionCanceledException' }); + expect(result.cancellation_reasons).toEqual([ + { Code: 'None' }, { Code: 'None' }, { Code: 'ConditionalCheckFailed' }, + ]); + expect(await read(approvals)).toEqual(before); + }); + test('a second request cannot replace the first while the task waits', async () => { + await create(); + const result = await recordWorkerRequest({ + operation: 'create', + task_id: taskId, + request_id: 'replacement', + worker_attempt_id: 'worker', + approval: { ...row(), request_id: 'replacement' }, + }); + expect(result).toMatchObject({ ok: false, code: 'TransactionCanceledException' }); + expect(await read(tasks)).toMatchObject({ awaiting_approval_request_id: gate }); + expect((await mockClient.send(new GetCommand({ + TableName: approvals, Key: { task_id: taskId, request_id: 'replacement' }, + }))).Item).toBeUndefined(); + }); +}); diff --git a/cdk/test/handlers/request-approval.test.ts b/cdk/test/handlers/request-approval.test.ts new file mode 100644 index 000000000..80d682958 --- /dev/null +++ b/cdk/test/handlers/request-approval.test.ts @@ -0,0 +1,139 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import type { APIGatewayProxyEvent } from 'aws-lambda'; + +const send = jest.fn(); +jest.mock('../../src/handlers/shared/ua', () => ({ makeDocClient: () => ({ send }) })); +import { handler, recordWorkerRequest } from '../../src/handlers/request-approval'; + +const approval = { + task_id: 'task', + request_id: 'gate', + user_id: 'owner', + repo: 'owner/repo', + tool_name: 'Bash', + tool_input_preview: '{"command":"git push"}', + tool_input_sha256: 'a'.repeat(64), + reason: 'Protected operation', + severity: 'high', + matching_rule_ids: ['protected'], + status: 'PENDING', + created_at: '2026-09-18T00:00:00Z', + timeout_s: 0, +}; +const input = { operation: 'create' as const, task_id: 'task', request_id: 'gate', approval }; +let task: Record; +beforeEach(() => { + task = { user_id: 'owner', repo: 'owner/repo', compute_type: 'ecs', status: 'RUNNING' }; + send.mockReset().mockImplementation(async command => + command.constructor.name === 'GetCommand' ? { Item: task } : {}); +}); + +test.each([ + { status: 'APPROVED' }, { status: 'DENIED' }, + { notified_linear_approval_requested: true }, { decision_source: 'forged' }, + { decided_at: 'now' }, { scope: 'all' }, { ttl: 1 }, { user_id: 'other' }, { repo: 'other/repo' }, +])('rejects forged decision/notification fields and ownership: %j', async changes => { + expect(await recordWorkerRequest({ ...input, approval: { ...approval, ...changes } })) + .toMatchObject({ ok: false, code: 'APPROVAL_REQUEST_INVALID' }); + expect(send).toHaveBeenCalledTimes(1); +}); + +test('records only the pending request, guarded by current task ownership and status', async () => { + expect(await recordWorkerRequest(input)).toEqual({ ok: true }); + const items = send.mock.calls[1][0].input.TransactItems; + expect(items[0].Put.Item).toEqual(approval); + expect(items[0].Put.ConditionExpression).toBe('attribute_not_exists(request_id)'); + expect(items[1].Update.ConditionExpression).toBe('#status = :running AND user_id = :user'); + expect(items[1].Update.ExpressionAttributeValues[':user']).toBe('owner'); +}); + +test('fences stale MicroVM writers in the same transaction', async () => { + task.compute_type = 'lambda-microvm'; + await recordWorkerRequest({ ...input, worker_attempt_id: 'worker-token' }); + const lease = send.mock.calls[1][0].input.TransactItems[2].ConditionCheck; + expect(lease.Key).toEqual({ task_id: 'worker-lease#task' }); + expect(lease.ConditionExpression).toBe('lease_state = :active AND lease_attempt_id = :attempt AND lease_user_id = :user'); + expect(lease.ExpressionAttributeValues).toEqual({ + ':active': 'ACTIVE', ':attempt': 'worker-token', ':user': 'owner', + }); +}); + +test('timeout cannot overwrite a human decision or resume the task', async () => { + await recordWorkerRequest({ operation: 'timeout', task_id: 'task', request_id: 'gate' }); + const items = send.mock.calls[1][0].input.TransactItems; + expect(items[0].Update.ExpressionAttributeValues[':timeout']).toBe('TIMED_OUT'); + expect(items[0].Update.ConditionExpression).toBe('#status = :pending AND user_id = :user'); + expect(items[1].ConditionCheck.ConditionExpression).toContain('awaiting_approval_request_id = :request'); + expect(items[1].Update).toBeUndefined(); +}); + +test('preserves cancellation reasons without returning request contents', async () => { + send.mockRejectedValueOnce({ + name: 'TransactionCanceledException', + CancellationReasons: [ + { Code: 'ConditionalCheckFailed', Item: { secret: 'must not return' } }, + ], + }); + expect(await recordWorkerRequest(input)).toEqual({ + ok: false, code: 'TransactionCanceledException', cancellation_reasons: [{ Code: 'ConditionalCheckFailed' }], + }); +}); + +test('requires IAM caller identity and refuses body/path task substitution', async () => { + const event = { + body: JSON.stringify(input), + pathParameters: { task_id: 'other-task' }, + requestContext: { identity: { userArn: 'arn:aws:sts::123456789012:assumed-role/Session/task' } }, + } as unknown as APIGatewayProxyEvent; + expect((await handler(event)).statusCode).toBe(400); + expect(send).not.toHaveBeenCalled(); + expect((await handler({ ...event, requestContext: {} } as APIGatewayProxyEvent)).statusCode).toBe(403); +}); + +test('accepts the signed path for the matching request', async () => { + const result = await handler({ + body: JSON.stringify(input), + pathParameters: { task_id: 'task' }, + requestContext: { identity: { userArn: 'arn:aws:sts::123456789012:assumed-role/Session/task' } }, + } as unknown as APIGatewayProxyEvent); + expect(result.statusCode).toBe(200); + expect(JSON.parse(result.body)).toEqual({ data: { ok: true } }); +}); + +test('reports service failures as unavailable with a request ID, not invalid input', async () => { + send.mockRejectedValueOnce({ name: 'ProvisionedThroughputExceededException' }); + const result = await handler({ + body: JSON.stringify(input), + pathParameters: { task_id: 'task' }, + requestContext: { requestId: 'api-request', identity: { userArn: 'worker' } }, + } as unknown as APIGatewayProxyEvent); + expect(result.statusCode).toBe(503); + expect(JSON.parse(result.body)).toMatchObject({ + error: { code: 'ProvisionedThroughputExceededException', request_id: 'api-request' }, + }); +}); + +test.each(['*', 'task/*', '../task', 'task?x', 'task#lease', 'task\n'])('rejects unsafe task/request/worker identifiers %j before database access', async id => { + for (const key of ['task_id', 'request_id', 'worker_attempt_id']) { + expect(await recordWorkerRequest({ ...input, [key]: id })).toEqual({ ok: false, code: 'APPROVAL_REQUEST_INVALID' }); + } + expect(send).not.toHaveBeenCalled(); +}); diff --git a/cdk/test/handlers/shared/agent-heartbeat.test.ts b/cdk/test/handlers/shared/agent-heartbeat.test.ts new file mode 100644 index 000000000..17832d6b3 --- /dev/null +++ b/cdk/test/handlers/shared/agent-heartbeat.test.ts @@ -0,0 +1,36 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +// SPDX-License-Identifier: MIT-0 + +import { evaluateAgentHeartbeat } from '../../../src/handlers/shared/agent-heartbeat'; + +test.each([ + [undefined, undefined, 1_000_000, undefined], + [NaN, undefined, 1_000_000, undefined], + [0, undefined, 360_000, undefined], + [0, undefined, 360_001, 'missing'], + [0, -200_000, 120_000, undefined], + [0, -200_000, 120_001, 'stale'], + [0, 120_000, 360_000, undefined], + [0, 120_000, 360_001, 'stale'], + [0, 400_000, 360_001, undefined], +] as const)('heartbeat boundary: start=%s heartbeat=%s now=%s yields %s', (start, heartbeat, now, expected) => { + expect(evaluateAgentHeartbeat(start, heartbeat, now)).toBe(expected); +}); diff --git a/cdk/test/handlers/shared/approval-notifications.test.ts b/cdk/test/handlers/shared/approval-notifications.test.ts new file mode 100644 index 000000000..a99200238 --- /dev/null +++ b/cdk/test/handlers/shared/approval-notifications.test.ts @@ -0,0 +1,205 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import type { DynamoDBDocumentClient } from '@aws-sdk/lib-dynamodb'; +import { TaskStatus } from '../../../src/constructs/task-status'; +import { approvalNotificationMarkdown, loadApprovalNotification, markApprovalNotificationDelivered } from '../../../src/handlers/shared/approval-notifications'; +import type { TaskRecord } from '../../../src/handlers/shared/types'; + +const send = jest.fn(); +const ddb = { send } as unknown as DynamoDBDocumentClient; +const task = { + task_id: 'task', + user_id: 'owner', + status: TaskStatus.AWAITING_APPROVAL, + awaiting_approval_request_id: 'gate', +} as TaskRecord; +const row = { + status: 'PENDING', + user_id: 'owner', + tool_name: 'Bash', + severity: 'high', + reason: 'Destructive command', + tool_input_preview: 'git push --force', + created_at: '2026-09-16T12:00:00Z', + timeout_s: 1800, +}; +beforeEach(() => { + process.env.TASK_APPROVALS_TABLE_NAME = 'Approvals'; + send.mockReset().mockResolvedValue({ Item: row }); +}); + +test('shows the saved action and exact deadline, with one-call CLI approval', async () => { + const message = await loadApprovalNotification(ddb, task, 'approval_requested', { request_id: 'gate', reason: 'wrong' }, 'slack'); + expect(message?.text).toContain('Destructive command'); + expect(message?.text).toContain('git push --force'); + expect(message?.text).toContain('2026-09-16T12:30:00.000Z'); + expect(message?.text).toContain('bgagent approve task gate --scope this_call'); + expect(message?.text).toContain('bgagent deny task gate'); + expect(send.mock.calls[0][0].input.ConsistentRead).toBe(true); +}); + +test('explains the README decision before exposing internal policy and request IDs', async () => { + send.mockResolvedValue({ + Item: { + ...row, + tool_name: 'Read', + reason: 'Soft-deny: legacy_extra_0', + tool_input_preview: JSON.stringify({ file_path: '/workspace/task/README.md' }), + }, + }); + const message = (await loadApprovalNotification(ddb, { ...task, repo: 'owner/repo' }, 'approval_requested', { request_id: 'gate' }, 'linear'))!; + expect(message.text).toContain('The agent wants to read the file "README.md".'); + expect(message.text).toContain('Repository: owner/repo'); + expect(message.text).toContain('your configured policy requires a human decision'); + expect(message.text).toContain('Approve: allow this action once.'); + expect(message.text).toContain('Deny: block this action and return the decision to the agent.'); + const primary = message.text.split('Technical details / CLI alternative')[0]; + expect(primary).toContain('Reply approve or deny to this comment'); + expect(primary).not.toMatch(/legacy_extra|Task:|Request:|severity:/); + expect(message.text).toContain('Policy detail: Soft-deny: legacy_extra_0'); +}); + +test.each([ + ['Write', { file_path: '/etc/config', content: 'replacement' }, 'creating the file or replacing its contents.'], + ['Edit', { file_path: '/workspace/task/app.py', old_string: 'before', new_string: 'after' }, 'change text in "app.py".'], + ['Bash', { command: 'git push --force', description: 'Publish the branch' }, 'Command: "git push --force"'], + ['WebFetch', { url: 'https://example.com' }, 'fetch content from "https://example.com".'], + ['Grep', { pattern: 'password', path: '/etc' }, 'search file contents.'], + ['mcp__custom__publish', { target: 'production' }, 'call the tool "mcp__custom__publish".'], +])('describes %s without hiding its saved arguments', async (tool, input, expected) => { + const preview = JSON.stringify(input); + send.mockResolvedValue({ Item: { ...row, tool_name: tool, tool_input_preview: preview } }); + const message = (await loadApprovalNotification(ddb, task, 'approval_requested', { request_id: 'gate' }, 'linear'))!; + expect(message.text).toContain(expected); + expect(message.text).toContain(`Saved arguments: ${preview}`); + if (tool === 'Bash') expect(message.text).toContain("Agent's explanation: Publish the branch"); + expect(message.text).not.toContain('incomplete'); +}); + +test('does not invent a read target from truncated arguments', async () => { + send.mockResolvedValue({ Item: { ...row, tool_name: 'Read', tool_input_preview: '{"file_path": "/workspace/task/secret...' } }); + const message = (await loadApprovalNotification(ddb, task, 'approval_requested', { request_id: 'gate' }, 'linear'))!; + expect(message.text).toContain('saved arguments are incomplete'); + expect(message.text).not.toContain('wants to read the file'); +}); + +test('redacts extracted command descriptions and makes display truncation explicit', async () => { + const token = `ghp_${'a'.repeat(36)}`; + send.mockResolvedValue({ + Item: { + ...row, + tool_input_preview: JSON.stringify({ + command: `echo ${token} ${'x'.repeat(600)}`, description: `Use ${token}`, + }), + }, + }); + const message = (await loadApprovalNotification(ddb, task, 'approval_requested', { request_id: 'gate' }, 'linear'))!; + expect(message.text).not.toContain(token); + expect(message.text).toContain('[REDACTED-GITHUB_TOKEN]'); + expect(message.text).toContain('[shortened]'); + expect(message.text).toContain('may not show the full action'); +}); + +test.each([ + ['cancelled task', { ...task, status: TaskStatus.CANCELLED }, row], + ['new gate', { ...task, awaiting_approval_request_id: 'new' }, row], + ['already decided', task, { ...row, status: 'APPROVED' }], + ['foreign owner', task, { ...row, user_id: 'other' }], + ['already delivered', task, { ...row, notified_slack_approval_requested: '2026-09-16' }], + ['missing approval', task, undefined], +])('suppresses a delayed request: %s', async (_label, owningTask, approval) => { + send.mockResolvedValue({ Item: approval }); + expect(await loadApprovalNotification(ddb, owningTask, 'approval_requested', { request_id: 'gate' }, 'slack')).toBeNull(); +}); + +test.each([ + ['APPROVED', 'approval_decision_recorded', 'Approval recorded'], + ['DENIED', 'approval_decision_recorded', 'Denial recorded'], + ['TIMED_OUT', 'approval_timed_out', 'Approval request timed out'], + ['CANCELLED', 'approval_cancelled', 'Approval request cancelled'], + ['STRANDED', 'approval_stranded', 'Approval wait could not continue'], +])('reports saved %s even when the task is now terminal', async (status, event, title) => { + send.mockResolvedValue({ Item: { ...row, status, cancellation_reason: 'Task cancelled by its owner' } }); + const message = await loadApprovalNotification(ddb, { ...task, status: TaskStatus.CANCELLED }, event, { request_id: 'gate' }, 'linear'); + expect(message?.title).toBe(title); + expect(message?.text).not.toContain('bgagent approve'); + if (status === 'CANCELLED') expect(message?.text).toContain('Task cancelled by its owner'); + if (status === 'APPROVED' || status === 'DENIED') { + expect(message?.text).toContain('Task status: CANCELLED. The task has already ended.'); + expect(message?.text).not.toContain('when its worker is ready'); + } +}); + +test.each([ + [undefined, 'The configured decision deadline was reached.'], + ['poll failed 3 consecutive times', 'poll failed 3 consecutive times'], +])('reports timeout closure instead of the original policy reason (%s)', async (denyReason, expected) => { + send.mockResolvedValue({ Item: { ...row, status: 'TIMED_OUT', deny_reason: denyReason } }); + const message = await loadApprovalNotification(ddb, task, 'approval_timed_out', { request_id: 'gate' }, 'slack'); + expect(message?.text).toContain(expected); + expect(message?.text).not.toContain(row.reason); +}); + +test('handles the actual stranded-reconciler event without changing the pending approval', async () => { + const failed = { ...task, status: TaskStatus.FAILED, error_message: 'Approval stranded: task paused for approval for 7200s with no resume transition.' }; + const message = await loadApprovalNotification(ddb, failed, 'approval_stranded', { reason: 'STRANDED_NO_HEARTBEAT' }, 'linear'); + expect(message).toMatchObject({ title: 'Approval wait could not continue', requestId: 'gate' }); + expect(message?.text).toContain(failed.error_message); + expect(message?.text).not.toContain('bgagent approve'); + expect(send).toHaveBeenCalledTimes(1); + expect(send.mock.calls[0][0].input.Key).toEqual({ task_id: 'task', request_id: 'gate' }); +}); + +test.each([ + [task, row], + [{ ...task, status: TaskStatus.FAILED, error_message: 'Unrelated failure' }, row], + [{ ...task, status: TaskStatus.FAILED, error_message: 'Approval stranded: old wait', awaiting_approval_request_id: undefined }, row], + [{ ...task, status: TaskStatus.FAILED, error_message: 'Approval stranded: old wait' }, { ...row, user_id: 'foreign' }], + [{ ...task, status: TaskStatus.FAILED, error_message: 'Approval stranded: old wait' }, { ...row, status: 'APPROVED' }], +])('does not infer a stranded request from an unrelated or decided task', async (owningTask, approval) => { + send.mockResolvedValue({ Item: approval }); + expect(await loadApprovalNotification(ddb, owningTask, 'approval_stranded', { reason: 'STRANDED_NO_HEARTBEAT' }, 'linear')).toBeNull(); +}); + +test('delivery receipts are separate for each request and channel and cannot recreate a deleted row', async () => { + const message = (await loadApprovalNotification(ddb, task, 'approval_requested', { request_id: 'gate' }, 'linear'))!; + await markApprovalNotificationDelivered(ddb, message); + expect(send.mock.calls[1][0].input).toMatchObject({ + Key: { task_id: 'task', request_id: 'gate' }, + ConditionExpression: 'attribute_exists(task_id) AND user_id = :user', + ExpressionAttributeNames: { '#marker': 'notified_linear_approval_requested' }, + }); +}); + +test('does not turn repository content into markdown or shell command substitutions', async () => { + send.mockResolvedValue({ Item: { ...row, tool_input_preview: '``` @everyone $(touch file)' } }); + const message = (await loadApprovalNotification(ddb, { ...task, awaiting_approval_request_id: '$(evil)' }, 'approval_requested', { request_id: '$(evil)' }, 'linear'))!; + expect(message.text).not.toContain('bgagent approve'); + expect(approvalNotificationMarkdown(message).match(/```/g)).toHaveLength(2); +}); + +test('redacts known credential shapes before posting a preview to a shared channel', async () => { + const token = `ghp_${'a'.repeat(36)}`; + send.mockResolvedValue({ Item: { ...row, tool_input_preview: `\u001b[31mTOKEN=${token}\u202eecho` } }); + const message = (await loadApprovalNotification(ddb, task, 'approval_requested', { request_id: 'gate' }, 'slack'))!; + expect(message.text).toContain('[REDACTED-GITHUB_TOKEN]'); + expect(message.text).not.toContain(token); + expect(message.text).not.toMatch(/[\u001b\u202e]/); +}); diff --git a/cdk/test/handlers/shared/canonical-json.test.ts b/cdk/test/handlers/shared/canonical-json.test.ts new file mode 100644 index 000000000..58d62f656 --- /dev/null +++ b/cdk/test/handlers/shared/canonical-json.test.ts @@ -0,0 +1,31 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { canonicalJson } from '../../../src/handlers/shared/canonical-json'; + +test('preserves the existing receipt byte format across nested key insertion orders', () => { + const expected = '{"a":[{"b":2,"z":1},3],"z":null}'; + expect(canonicalJson({ z: null, a: [{ z: 1, b: 2 }, 3] })).toBe(expected); + expect(canonicalJson({ a: [{ b: 2, z: 1 }, 3], z: null })).toBe(expected); +}); + +test('keeps array order and distinct primitive values significant', () => { + expect(canonicalJson([1, '1', false, null])).toBe('[1,"1",false,null]'); + expect(canonicalJson([1, 2])).not.toBe(canonicalJson([2, 1])); +}); diff --git a/cdk/test/handlers/shared/close-task-approvals.test.ts b/cdk/test/handlers/shared/close-task-approvals.test.ts new file mode 100644 index 000000000..f6cf11e55 --- /dev/null +++ b/cdk/test/handlers/shared/close-task-approvals.test.ts @@ -0,0 +1,72 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +const mockSend = jest.fn(); +jest.mock('../../../src/handlers/shared/ua', () => ({ makeDocClient: () => ({ send: mockSend }) })); +import { closeTaskApprovals } from '../../../src/handlers/shared/close-task-approvals'; + +beforeEach(() => { + jest.clearAllMocks(); + process.env.TASK_APPROVALS_TABLE_NAME = 'approvals'; +}); +afterEach(() => { delete process.env.TASK_APPROVALS_TABLE_NAME; }); + +test('a parked task keeps its unanswered requests and has no retention timer', async () => { + mockSend.mockResolvedValue({ Item: { user_id: 'user', status: 'AWAITING_APPROVAL' } }); + await closeTaskApprovals('task', 'user'); + expect(mockSend).toHaveBeenCalledTimes(1); +}); + +test('task closure cancels unanswered requests, preserves answers, and stamps history retention', async () => { + mockSend.mockResolvedValueOnce({ Item: { user_id: 'user', status: 'FAILED' } }) + .mockResolvedValueOnce({ + Items: [ + { user_id: 'user', request_id: 'pending', status: 'PENDING' }, + { user_id: 'user', request_id: 'approved', status: 'APPROVED' }, + { user_id: 'other', request_id: 'wrong-owner', status: 'PENDING' }, + ], + }).mockResolvedValue({}); + await closeTaskApprovals('task', 'user'); + const updates = mockSend.mock.calls.slice(2).map(([command]) => command.input); + expect(updates).toHaveLength(2); + expect(updates[0].ExpressionAttributeValues).toMatchObject({ + ':cancelled': 'CANCELLED', ':observed': 'PENDING', ':reason': 'Owning task is failed.', + }); + expect(updates[0].ExpressionAttributeValues[':ttl']).toBeGreaterThan(Date.now() / 1000 + 86400); + expect(updates[1].UpdateExpression).toBe('SET #ttl = if_not_exists(#ttl, :ttl)'); + expect(updates[1].ExpressionAttributeValues[':cancelled']).toBeUndefined(); +}); + +test('a mismatched task owner cannot close another user’s approvals', async () => { + mockSend.mockResolvedValue({ Item: { user_id: 'other', status: 'CANCELLED' } }); + await closeTaskApprovals('task', 'user'); + expect(mockSend).toHaveBeenCalledTimes(1); +}); + +test('an answer racing task closure keeps its decision and still receives retention', async () => { + mockSend.mockResolvedValueOnce({ Item: { user_id: 'user', status: 'CANCELLED' } }) + .mockResolvedValueOnce({ Items: [{ user_id: 'user', request_id: 'request', status: 'PENDING' }] }) + .mockRejectedValueOnce(Object.assign(new Error('answered'), { name: 'ConditionalCheckFailedException' })) + .mockResolvedValue({}); + await closeTaskApprovals('task', 'user'); + expect(mockSend.mock.calls[3][0].input).toMatchObject({ + UpdateExpression: 'SET #ttl = if_not_exists(#ttl, :ttl)', + ConditionExpression: 'user_id = :user AND #status <> :pending', + }); +}); diff --git a/cdk/test/handlers/shared/create-task-core.test.ts b/cdk/test/handlers/shared/create-task-core.test.ts index 00de18b9e..954724037 100644 --- a/cdk/test/handlers/shared/create-task-core.test.ts +++ b/cdk/test/handlers/shared/create-task-core.test.ts @@ -854,6 +854,26 @@ describe('createTaskCore', () => { expect(result.body).toContain('pr_number'); }); + test.each([undefined, 0, 1, 600, 3600])('persists and returns the resolved MicroVM sleep delay %s', async value => { + const result = await createTaskCore( + { repo: 'org/repo', task_description: 'wait preference', ...(value !== undefined && { microvm_sleep_after_s: value }) }, + makeContext(), 'req-sleep', + ); + expect(result.statusCode).toBe(201); + expect(getPersistedTaskRecord().microvm_sleep_after_s).toBe(value ?? 600); + expect(JSON.parse(result.body).data.microvm_sleep_after_s).toBe(value ?? 600); + }); + + test.each([-1, 3601, 0.5, '600', null, false, {}, []])('rejects invalid MicroVM sleep delay %j before creating a task', async value => { + const result = await createTaskCore( + { repo: 'org/repo', task_description: 'invalid delay', microvm_sleep_after_s: value } as any, + makeContext(), 'req-sleep-invalid', + ); + expect(result.statusCode).toBe(400); + expect(JSON.parse(result.body).error.message).toContain('microvm_sleep_after_s'); + expect(mockSend.mock.calls.some(([command]) => command._type === 'Put')).toBe(false); + }); + // -- trace flag (design §10.1) -------------------------------------- test('trace: true persists on the task record and surfaces in the response', async () => { diff --git a/cdk/test/handlers/shared/error-classifier.test.ts b/cdk/test/handlers/shared/error-classifier.test.ts index e42a69886..3b734204f 100644 --- a/cdk/test/handlers/shared/error-classifier.test.ts +++ b/cdk/test/handlers/shared/error-classifier.test.ts @@ -17,7 +17,7 @@ * SOFTWARE. */ -import { classifyError, ErrorCategory, ErrorClass, isTransientError, retryGuidance, type ErrorClassification } from '../../../src/handlers/shared/error-classifier'; +import { classifyError, ErrorCategory, ErrorClass, formatMicrovmTerminalFailure, isTransientError, retryGuidance, type ErrorClassification } from '../../../src/handlers/shared/error-classifier'; import { LAMBDA_MICROVM_SUPPORTED_REGIONS } from '../../../src/handlers/shared/microvm-regions'; import { toTaskDetail, type TaskRecord } from '../../../src/handlers/shared/types'; @@ -480,6 +480,16 @@ describe('classifyError', () => { // --- Lambda MicroVMs (ADR-021) --- describe('Lambda MicroVMs errors', () => { + test('an unknown start requires investigation even when its detail describes a transient failure', () => { + const result = classifyError( + 'Session start failed: MICROVM_START_OUTCOME_UNKNOWN: MicroVM RunMicrovm failed: TimeoutError: response lost [auto-retried]', + )!; + expect(result.retryable).toBe(false); + expect(result.errorClass).toBe(ErrorClass.SERVICE); + expect(retryGuidance(result, true)).toMatch(/needs your ABCA admin/); + expect(result.remedy).toMatch(/before submitting another task/); + }); + test('classifies regional unavailability as a non-retryable CONFIG fault with the supported-Region list', () => { // ADR-021: "If startSession fails because the MicroVM service is unavailable // in the stack region, then the orchestrator shall classify the failure with @@ -575,8 +585,7 @@ describe('classifyError', () => { }); test('classifies a MicroVM substrate-failure reason written by the orchestrator', () => { - // Must stay in lockstep with the reason string - // `reconcileMicrovmSubstrateState` persists. + // Retain classification of legacy messages persisted by the P2 reconciler. const result = classifyError( 'MicroVM substrate terminated before the agent wrote a terminal status: substrate state completed', )!; @@ -589,7 +598,7 @@ describe('classifyError', () => { // --- lifecycle-hook 4xx: non-retryable, and it must OUTRANK the generic entry --- // // The service reports a guest 4xx in `stateReason`, which - // `reconcileMicrovmSubstrateState` appends to the persisted message — so BOTH + // the P2 reconciler appended to its persisted message — so BOTH // the generic `MicroVM substrate terminated` string and the hook-status string // are present in one `error_message` and ORDER decides the answer. These tests // exist because the wrong order is invisible to a message-only assertion. @@ -598,11 +607,94 @@ describe('classifyError', () => { const hookReason = (status: number) => `Run lifecycle hook returned HTTP status ${status}. ` + 'Please check your hook endpoint and application logs for more details.'; - /** ...as `reconcileMicrovmSubstrateState` persists it. */ + /** Legacy persisted form; current finalization also supplies stable MICROVM_* codes. */ const reconciled = (reason: string) => 'MicroVM substrate terminated before the agent wrote a terminal status: ' + `substrate state completed (${reason})`; + test.each([ + 'Resume lifecycle hook failed. Please check your hook endpoint and application logs for more details.', + 'Resume lifecycle hook connection was refused. Please check your hook endpoint and application logs for more details.', + 'Resume lifecycle hook timed out. Please check your hook endpoint and application logs for more details.', + 'Resume lifecycle hook returned HTTP status 503.', + 'Resume lifecycle hook returned HTTP status 409.', + ])('gives actionable wake diagnostics for %s', reason => { + const persisted = formatMicrovmTerminalFailure('substrate state completed', reason); + expect(persisted).toMatch(/^MICROVM_RESUME_HOOK_FAILED: /); + expect(persisted).toContain(reason); + for (const message of [persisted, reconciled(reason), reason]) { + const result = classifyError(message)!; + expect(result.title).toBe('The MicroVM could not wake after being paused'); + expect(result.errorClass).toBe(ErrorClass.SERVICE); + expect(result.retryable).toBe(false); + expect(result.remedy).toContain('AWS request ID'); + expect(result.remedy).toContain('before starting a replacement'); + expect(retryGuidance(result)).toMatch(/needs your ABCA admin/i); + } + }); + + test('keeps unknown wording generic and honors persisted codes ahead of diagnostic words', () => { + expect(formatMicrovmTerminalFailure('completed', 'diagnostic: Resume lifecycle hook connection was refused.')) + .toMatch(/^MICROVM_SUBSTRATE_TERMINATED: /); + expect(classifyError(`MICROVM_SUBSTRATE_TERMINATED: ${reconciled('Resume lifecycle hook connection was refused.')}`)!.errorClass) + .toBe(ErrorClass.TRANSIENT); + expect(classifyError('MICROVM_RESUME_HOOK_FAILED: concurrency limit; missing_secret')!.errorClass) + .toBe(ErrorClass.SERVICE); + }); + + test.each([ + 'MicroVM host unavailable.', + 'MicroVM capacity unavailable in this Availability Zone.', + 'MicroVM host unavailable in this region.', + ])('does not mistake "%s" for an unsupported region', (reason) => { + for (const message of [reason, reconciled(reason)]) { + const result = classifyError(message)!; + expect(result.category).toBe(ErrorCategory.COMPUTE); + expect(result.errorClass).toBe(ErrorClass.TRANSIENT); + expect(result.retryable).toBe(true); + expect(retryGuidance(result)).toMatch(/reply here to try again/i); + } + }); + + test.each([ + 'MicroVM host unavailable.', + 'MicroVM unavailable in this region.', + 'INSUFFICIENT_GITHUB_REPO_PERMISSIONS', + 'concurrency limit reached', + 'BLOCKED[missing_secret]: diagnostic text', + 'Run lifecycle hook returned HTTP status 400.', + ])('a stable terminal code cannot be reclassified by diagnostic text: %s', (reason) => { + const result = classifyError(`MICROVM_SUBSTRATE_TERMINATED: ${reconciled(reason)}`)!; + expect(result.title).toBe('The MicroVM stopped before the agent reported a result'); + expect(result.category).toBe(ErrorCategory.COMPUTE); + expect(result.errorClass).toBe(ErrorClass.TRANSIENT); + }); + + test('a stable hook-rejection code keeps configuration guidance despite other diagnostic words', () => { + const result = classifyError( + `MICROVM_RUN_HOOK_REJECTED: ${reconciled('MicroVM host unavailable; concurrency limit')}`, + )!; + expect(result.category).toBe(ErrorCategory.CONFIG); + expect(result.retryable).toBe(false); + expect(retryGuidance(result)).toMatch(/needs your ABCA admin/i); + }); + + test.each(['AccessDeniedException', 'UnauthorizedException'])( + 'does not retry a marked %s when starting a MicroVM', (name) => { + const result = classifyError(`Session start failed: MicroVM RunMicrovm failed: ${name}: denied`)!; + expect(result.category).toBe(ErrorCategory.AUTH); + expect(result.errorClass).toBe(ErrorClass.SERVICE); + expect(result.retryable).toBe(false); + }, + ); + + test.each(['ValidationException', 'InvalidParameterValueException'])('does not retry MicroVM %s', (name) => { + const result = classifyError(`Session start failed: MicroVM RunMicrovm failed: ${name}: invalid connector`)!; + expect(result.category).toBe(ErrorCategory.CONFIG); + expect(result.errorClass).toBe(ErrorClass.SERVICE); + expect(result.retryable).toBe(false); + }); + test.each([400, 403, 404, 422, 499])( 'classifies a lifecycle-hook %i as a NON-retryable config fault', (status) => { @@ -612,6 +704,7 @@ describe('classifyError', () => { // generic COMPUTE/TRANSIENT entry whose remedy is "reply here to try // again" — an invitation to loop forever on a version-skewed deployment. const result = classifyError(reconciled(hookReason(status)))!; + expect(classifyError(hookReason(status))).toEqual(result); expect(result.category).toBe(ErrorCategory.CONFIG); expect(result.retryable).toBe(false); expect(result.errorClass).toBe(ErrorClass.SERVICE); @@ -626,7 +719,7 @@ describe('classifyError', () => { test('a lifecycle-hook 5xx stays RETRYABLE — MICROVM_RUN_PAYLOAD_UNREADABLE is a 500', () => { // The scoping that makes the entry above safe. The agent answers 500 for a - // truncated/racing S3 payload, which a retry genuinely can fix, so the 4xx + // failed S3 read, which a retry may fix, so the 4xx // pattern must not swallow the 5xx family. const result = classifyError(reconciled(hookReason(500)))!; expect(result.category).toBe(ErrorCategory.COMPUTE); @@ -644,6 +737,35 @@ describe('classifyError', () => { expect(result.retryable).toBe(true); }); + test.each(['substrate-read-failed', 'substrate-read-failed-repeatedly'])( + 'gives operator guidance for supervisor %s without promising cleanup succeeded', (reason) => { + const result = classifyError(`MicroVM supervisor: ${reason}`)!; + expect(result).toMatchObject({ + category: ErrorCategory.COMPUTE, + title: 'The MicroVM status could not be checked', + retryable: false, + errorClass: ErrorClass.SERVICE, + }); + expect(result.remedy).toContain('microvm_supervisor_request_failed'); + expect(result.remedy).toContain('AWS request ID'); + expect(result.remedy).toContain('Confirm worker termination'); + }, + ); + + test.each([ + 'AgentCore supervisor: substrate-read-failed', + 'Agent output mentions MicroVM supervisor: substrate-read-failed', + 'MicroVM supervisor: substrate-read-failed-unrecognized', + ])('does not infer a supervisor failure from unrelated text: %s', (message) => { + expect(classifyError(message)!.category).toBe(ErrorCategory.UNKNOWN); + }); + + test('a persisted terminal code still takes precedence over supervisor wording', () => { + const result = classifyError('MICROVM_SUBSTRATE_TERMINATED: MicroVM supervisor: substrate-read-failed')!; + expect(result.title).toBe('The MicroVM stopped before the agent reported a result'); + expect(result.retryable).toBe(true); + }); + test('every new MicroVM classification carries a full, non-empty guidance shape', () => { const messages = [ 'Session start failed: UnknownEndpoint: Inaccessible host: `lambda.eu-central-1.amazonaws.com\'', @@ -651,6 +773,7 @@ describe('classifyError', () => { 'MicroVM RunMicrovm failed: ThrottlingException: Rate exceeded', 'MicroVM RunMicrovm failed: ResourceNotFoundException: image not found', 'MicroVM substrate terminated before the agent wrote a terminal status: substrate state completed', + 'MicroVM supervisor: substrate-read-failed', reconciled(hookReason(400)), ]; for (const msg of messages) { @@ -695,6 +818,9 @@ describe('classifyError', () => { ['ServiceQuotaExceededException: quota exceeded'], ['ResourceNotFoundException: Requested resource not found'], ['TooManyRequestsException: slow down'], + ['AccessDeniedException: denied'], + ['UnauthorizedException: denied'], + ['ValidationException: invalid'], ])('an unmarked "%s" still classifies as UNKNOWN, exactly as before', (message) => { const result = classifyError(message)!; expect(result.category).toBe(PRE_CHANGE_UNKNOWN.category); diff --git a/cdk/test/handlers/shared/linear-approval-reply.test.ts b/cdk/test/handlers/shared/linear-approval-reply.test.ts new file mode 100644 index 000000000..fe48da903 --- /dev/null +++ b/cdk/test/handlers/shared/linear-approval-reply.test.ts @@ -0,0 +1,193 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import type { DynamoDBDocumentClient } from '@aws-sdk/lib-dynamodb'; +import { handleLinearApprovalReply } from '../../../src/handlers/shared/linear-approval-reply'; +import { linearApprovalCommentId } from '../../../src/handlers/shared/linear-approval-thread'; + +const mockApprove = jest.fn(); +const mockDeny = jest.fn(); +const mockPost = jest.fn(); +const mockReadComment = jest.fn(); +jest.mock('../../../src/handlers/approve-task', () => ({ recordApprovalForUser: mockApprove })); +jest.mock('../../../src/handlers/deny-task', () => ({ recordDenialForUser: mockDeny })); +jest.mock('../../../src/handlers/shared/linear-feedback', () => ({ + postIdentifiedComment: (...args: unknown[]) => mockPost(...args), + readLinearApprovalComment: (...args: unknown[]) => mockReadComment(...args), +})); +const thread = { workspaceId: 'ws', issueId: 'issue', taskId: 'task', requestId: 'gate', userId: 'owner' }; +const root = linearApprovalCommentId(thread); +const event = { + action: 'create', + organizationId: 'ws', + actor: { id: 'actor' }, + data: { id: 'reply', body: 'approve', parentId: root, issueId: 'issue' }, +}; +const source = JSON.stringify(['linear', 'ws', 'reply']); +let approval: Record; +let task: Record; +const send = jest.fn(); +const lookupUser = jest.fn(); +const deps = { + ddb: { send } as unknown as DynamoDBDocumentClient, + approvalsTable: 'Approvals', + taskTable: 'Tasks', + registryTable: 'Registry', + lookupUser, +}; + +beforeEach(() => { + jest.clearAllMocks(); + approval = { user_id: 'owner', status: 'PENDING' }; + task = { + user_id: 'owner', + channel_source: 'linear', + status: 'AWAITING_APPROVAL', + channel_metadata: { linear_workspace_id: 'ws', linear_issue_id: 'issue' }, + }; + lookupUser.mockResolvedValue('owner'); + mockPost.mockResolvedValue({ ok: true }); + mockReadComment.mockResolvedValue({ + id: 'reply', + body: 'approve', + user: { id: 'real-author' }, + issue: { id: 'issue' }, + parent: { id: root }, + }); + mockApprove.mockResolvedValue({ statusCode: 202 }); + mockDeny.mockResolvedValue({ statusCode: 202 }); + send.mockImplementation(async command => { + if (command.input.UpdateExpression) return {}; + if (command.input.Key.task_id.startsWith('LINEAR_COMMENT#')) { + return { Item: { kind: 'linear_approval_thread', thread } }; + } + return { Item: command.input.TableName === 'Tasks' ? task : approval }; + }); +}); + +test.each(['approve', 'deny'])('records %s for the bound request and mapped owner', async body => { + mockReadComment.mockResolvedValue({ + id: 'reply', body, user: { id: 'real-author' }, issue: { id: 'issue' }, parent: { id: root }, + }); + expect(await handleLinearApprovalReply({ ...event, data: { ...event.data, body } }, deps)).toBe(true); + const record = body === 'approve' ? mockApprove : mockDeny; + expect(record).toHaveBeenCalledWith({ + userId: 'owner', + taskId: 'task', + decisionSource: source, + body: JSON.stringify({ request_id: 'gate', decision: body }), + }); + expect(mockPost).toHaveBeenCalledWith(expect.anything(), expect.objectContaining({ issueId: 'issue', parentId: root })); + expect(lookupUser).toHaveBeenCalledWith('ws', 'real-author'); +}); + +test.each([ + null, + { body: 'deny' }, + { user: null }, + { botActor: { id: 'app' } }, + { issue: { id: 'elsewhere' } }, + { parent: { id: 'other-gate' } }, + { id: 'other-reply' }, +])('rejects forged, edited, deleted or bot consent: %j', async change => { + const original = await mockReadComment(); + mockReadComment.mockResolvedValue(change === null ? null : { ...original, ...change }); + expect(await handleLinearApprovalReply(event, deps)).toBe(true); + expect(mockApprove).not.toHaveBeenCalled(); + expect(mockDeny).not.toHaveBeenCalled(); + expect(mockPost).not.toHaveBeenCalled(); +}); + +test('retries unavailable verification without trusting the webhook actor', async () => { + mockReadComment.mockRejectedValue(new Error('Linear unavailable')); + await expect(handleLinearApprovalReply(event, deps)).rejects.toThrow('Linear unavailable'); + expect(lookupUser).not.toHaveBeenCalled(); + expect(mockApprove).not.toHaveBeenCalled(); +}); + +test.each(['update', 'remove'])('ignores %s events', async action => { + expect(await handleLinearApprovalReply({ ...event, action }, deps)).toBe(false); + expect(send).not.toHaveBeenCalled(); +}); + +test.each(['I approve', '> approve', 'approve and delete it', '@bgagent approve'])('ignores prose: %s', async body => { + expect(await handleLinearApprovalReply({ ...event, data: { ...event.data, body } }, deps)).toBe(false); + expect(send).not.toHaveBeenCalled(); +}); + +test('ignores top-level replies and unknown threads', async () => { + expect(await handleLinearApprovalReply({ ...event, data: { ...event.data, parentId: undefined } }, deps)).toBe(false); + send.mockResolvedValue({}); + expect(await handleLinearApprovalReply(event, deps)).toBe(false); + expect(mockApprove).not.toHaveBeenCalled(); +}); + +test.each([null, 'different-owner'])('rejects unmapped or different owner %s', async user => { + lookupUser.mockResolvedValue(user); + expect(await handleLinearApprovalReply(event, deps)).toBe(true); + expect(mockApprove).not.toHaveBeenCalled(); + expect(mockPost.mock.calls[0][1].body).toContain('Only the task owner'); +}); + +test('does not act across workspaces or issues', async () => { + expect(await handleLinearApprovalReply({ ...event, organizationId: 'other' }, deps)).toBe(false); + expect(await handleLinearApprovalReply({ ...event, data: { ...event.data, issueId: 'other' } }, deps)).toBe(true); + expect(mockApprove).not.toHaveBeenCalled(); + expect(mockPost).not.toHaveBeenCalled(); +}); + +test.each(['user_id', 'channel_source', 'channel_metadata'])('rejects changed task binding: %s', async field => { + task[field] = 'changed'; + await handleLinearApprovalReply(event, deps); + expect(mockApprove).not.toHaveBeenCalled(); + expect(mockPost.mock.calls[0][1].body).toContain('no longer available'); +}); + +test('acknowledges duplicate deliveries without recording again even after the task advances', async () => { + approval = { user_id: 'owner', status: 'APPROVED', decision_source: source }; + task.status = 'SUCCEEDED'; + await handleLinearApprovalReply(event, deps); + await handleLinearApprovalReply(event, deps); + expect(mockApprove).not.toHaveBeenCalled(); + expect(mockPost.mock.calls[0]).toEqual(mockPost.mock.calls[1]); +}); + +test.each([404, 409])('reports a closed or superseded gate (HTTP %s)', async statusCode => { + mockApprove.mockResolvedValue({ statusCode }); + await handleLinearApprovalReply(event, deps); + expect(mockPost.mock.calls[0][1].body).toContain('No new decision'); + expect(JSON.parse(mockApprove.mock.calls[0][0].body).request_id).toBe('gate'); +}); + +test('recognizes a duplicate that committed concurrently', async () => { + mockApprove.mockImplementation(async () => { + approval = { user_id: 'owner', status: 'APPROVED', decision_source: source }; + return { statusCode: 404 }; + }); + await handleLinearApprovalReply(event, deps); + expect(mockPost.mock.calls[0][1].body).toContain('Approved for this action once'); +}); + +test('retries transient decision and acknowledgement failures', async () => { + mockApprove.mockResolvedValue({ statusCode: 503 }); + await expect(handleLinearApprovalReply(event, deps)).rejects.toThrow('HTTP 503'); + mockApprove.mockResolvedValue({ statusCode: 202 }); + mockPost.mockResolvedValue({ ok: false, retryable: true }); + await expect(handleLinearApprovalReply(event, deps)).rejects.toThrow('acknowledgement failure'); +}); diff --git a/cdk/test/handlers/shared/linear-approval-thread.test.ts b/cdk/test/handlers/shared/linear-approval-thread.test.ts new file mode 100644 index 000000000..8b6494510 --- /dev/null +++ b/cdk/test/handlers/shared/linear-approval-thread.test.ts @@ -0,0 +1,77 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import type { DynamoDBDocumentClient } from '@aws-sdk/lib-dynamodb'; +import { + closeLinearApprovalThread, linearApprovalCommentId, parseLinearApprovalReply, + readLinearApprovalThread, saveLinearApprovalThread, +} from '../../../src/handlers/shared/linear-approval-thread'; +const thread = { workspaceId: 'ws', issueId: 'issue', taskId: 'task', requestId: 'gate', userId: 'owner' }; +const send = jest.fn(); +const ddb = { send } as unknown as DynamoDBDocumentClient; +beforeEach(() => send.mockReset().mockResolvedValue({})); + +test('Linear-compatible UUIDv4 binds distinct workspace, issue, task and request identities', () => { + const id = linearApprovalCommentId(thread); + expect(id).toMatch(/^[0-9a-f]{8}-[0-9a-f]{4}-4[0-9a-f]{3}-a[0-9a-f]{3}-[0-9a-f]{12}$/); + expect(linearApprovalCommentId({ ...thread })).toBe(id); + for (const field of ['workspaceId', 'issueId', 'taskId', 'requestId']) { + expect(linearApprovalCommentId({ ...thread, [field]: 'other' })).not.toBe(id); + } +}); + +test('mapping writes preserve existing TTL and reject changed bindings', async () => { + await saveLinearApprovalThread(ddb, 'Approvals', thread); + const input = send.mock.calls[0][0].input; + expect(input.Key.task_id).toContain('LINEAR_COMMENT#ws#'); + expect(input.UpdateExpression).not.toContain('ttl'); + expect(input.ConditionExpression).toContain('#thread = :thread'); + expect(input.ExpressionAttributeNames).not.toHaveProperty('user_id'); +}); + +test('reads strongly and validates stored scope and shape', async () => { + const id = linearApprovalCommentId(thread); + send.mockResolvedValue({ Item: { kind: 'linear_approval_thread', thread } }); + expect(await readLinearApprovalThread(ddb, 'Approvals', 'ws', id)).toEqual(thread); + expect(send.mock.calls[0][0].input.ConsistentRead).toBe(true); + expect(await readLinearApprovalThread(ddb, 'Approvals', 'other', id)).toBeNull(); + send.mockResolvedValue({ Item: { kind: 'linear_approval_thread', thread: { ...thread, userId: undefined } } }); + expect(await readLinearApprovalThread(ddb, 'Approvals', 'ws', id)).toBeNull(); +}); + +test('closure starts retention once and never recreates a missing mapping', async () => { + await closeLinearApprovalThread(ddb, 'Approvals', thread); + const input = send.mock.calls[0][0].input; + expect(input.UpdateExpression).toContain('if_not_exists'); + expect(input.ConditionExpression).toBe('attribute_exists(task_id)'); + send.mockRejectedValue({ name: 'ConditionalCheckFailedException' }); + await expect(closeLinearApprovalThread(ddb, 'Approvals', thread)).resolves.toBeUndefined(); + send.mockRejectedValue(new Error('throttle')); + await expect(closeLinearApprovalThread(ddb, 'Approvals', thread)).rejects.toThrow('throttle'); +}); + +test.each(['approve', 'APPROVE!', ' approve. '])('accepts explicit answer %s', text => { + expect(parseLinearApprovalReply(text)).toBe('approve'); +}); +test.each(['deny', ' Deny! '])('accepts denial %s', text => { + expect(parseLinearApprovalReply(text)).toBe('deny'); +}); +test.each([undefined, 'I approve', '`approve`', 'approve\ndeny', 'approved'])('rejects ambiguous input %s', text => { + expect(parseLinearApprovalReply(text)).toBeNull(); +}); diff --git a/cdk/test/handlers/shared/linear-feedback.test.ts b/cdk/test/handlers/shared/linear-feedback.test.ts index 68f124349..526202821 100644 --- a/cdk/test/handlers/shared/linear-feedback.test.ts +++ b/cdk/test/handlers/shared/linear-feedback.test.ts @@ -33,6 +33,8 @@ import { deleteComment, fetchRecentComments, postIssueComment, + postIdentifiedComment, + readLinearApprovalComment, reactToComment, replyToComment, reportIssueFailure, @@ -73,6 +75,113 @@ describe('linear-feedback', () => { fetchMock.mockResolvedValue(jsonResponse({ data: { commentCreate: { success: true } } })); }); + describe('readLinearApprovalComment', () => { + test('reads the workspace, human author, decision and exact thread from Linear', async () => { + const comment = { + id: 'reply', + body: 'approve', + user: { id: 'human' }, + botActor: null, + issue: { id: ISSUE_ID }, + parent: { id: 'root' }, + }; + fetchMock.mockResolvedValue(jsonResponse({ data: { organization: { id: CTX.linearWorkspaceId }, viewer: { id: 'app-identity' }, comment } })); + expect(await readLinearApprovalComment(CTX, 'reply')).toEqual(comment); + const request = JSON.parse(fetchMock.mock.calls[0][1].body); + expect(request.variables).toEqual({ id: 'reply' }); + expect(request.query).toContain('user { id }'); + expect(request.query).toContain('botActor { id }'); + expect(request.query).toContain('viewer { id }'); + }); + test('rejects a genuine human comment made as the saved OAuth token identity', async () => { + fetchMock.mockResolvedValue(jsonResponse({ + data: { + organization: { id: CTX.linearWorkspaceId }, + viewer: { id: 'task-owner' }, + comment: { + id: 'reply', + body: 'approve', + user: { id: 'task-owner' }, + botActor: null, + issue: { id: ISSUE_ID }, + parent: { id: 'root' }, + }, + }, + })); + expect(await readLinearApprovalComment(CTX, 'reply')).toBeNull(); + }); + test('returns no consent for a deleted comment', async () => { + fetchMock.mockResolvedValue(jsonResponse({ data: { organization: { id: CTX.linearWorkspaceId }, viewer: { id: 'app-identity' }, comment: null } })); + expect(await readLinearApprovalComment(CTX, 'deleted')).toBeNull(); + }); + test.each([ + { errors: [{ message: 'Unavailable' }] }, + { data: { organization: { id: CTX.linearWorkspaceId }, comment: { user: { id: 'human' } } } }, + { data: { organization: { id: 'other-workspace' }, comment: {} } }, + ])('fails closed on lookup errors or incorrect workspace: %j', async response => { + fetchMock.mockResolvedValue(jsonResponse(response)); + await expect(readLinearApprovalComment(CTX, 'reply')).rejects.toThrow('verification'); + }); + test('fails closed without credentials', async () => { + resolveLinearOauthTokenMock.mockResolvedValue(null); + await expect(readLinearApprovalComment(CTX, 'reply')).rejects.toThrow('token unavailable'); + expect(fetchMock).not.toHaveBeenCalled(); + }); + }); + + describe('postIdentifiedComment', () => { + const input = { id: 'stable', issueId: ISSUE_ID, body: 'Approval needed', parentId: 'root' }; + test('passes the stable identity and exact thread to Linear', async () => { + expect(await postIdentifiedComment(CTX, input)).toEqual({ ok: true }); + expect(JSON.parse(fetchMock.mock.calls[0][1].body).variables).toEqual({ input }); + }); + test('recovers a lost creation response by verifying saved content and destination', async () => { + fetchMock.mockRejectedValueOnce(new Error('lost response')); + fetchMock.mockResolvedValueOnce(jsonResponse({ + data: { + comment: { + body: input.body, issue: { id: ISSUE_ID }, parent: { id: 'root' }, + }, + }, + })); + expect(await postIdentifiedComment(CTX, input)).toEqual({ ok: true }); + }); + test.each([ + ['Approval needed\r\nReply here\r\n', true], + ['Approval needed\nApprove a DIFFERENT action', false], + ['Approval needed\nReply here', false], + ['approval needed\nReply here', false], + ])('compares replay content conservatively: %j', async (body, accepted) => { + fetchMock.mockRejectedValueOnce(new Error('lost response')); + fetchMock.mockResolvedValueOnce(jsonResponse({ + data: { + comment: { body, issue: { id: ISSUE_ID }, parent: { id: 'root' } }, + }, + })); + expect(await postIdentifiedComment(CTX, { ...input, body: 'Approval needed\nReply here' })) + .toEqual(accepted ? { ok: true } : { ok: false, retryable: false }); + }); + test('does not accept an ID collision in another issue or thread', async () => { + fetchMock.mockResolvedValueOnce(jsonResponse({ errors: ['already exists'] })); + fetchMock.mockResolvedValueOnce(jsonResponse({ + data: { + comment: { + body: input.body, issue: { id: 'other' }, parent: { id: 'root' }, + }, + }, + })); + expect(await postIdentifiedComment(CTX, input)).toEqual({ ok: false, retryable: false }); + }); + test('does not treat success=false as delivery', async () => { + fetchMock.mockResolvedValue(jsonResponse({ data: { commentCreate: { success: false } } })); + expect(await postIdentifiedComment(CTX, input)).toEqual({ ok: false, retryable: false }); + }); + test('retries when creation and verification are both unavailable', async () => { + fetchMock.mockRejectedValue(new Error('outage')); + expect(await postIdentifiedComment(CTX, input)).toEqual({ ok: false, retryable: true }); + }); + }); + describe('postIssueComment', () => { test('POSTs the commentCreate mutation with the issue id and body', async () => { const result = await postIssueComment(CTX, ISSUE_ID, '❌ blocked'); diff --git a/cdk/test/handlers/shared/microvm-approval-wake.test.ts b/cdk/test/handlers/shared/microvm-approval-wake.test.ts new file mode 100644 index 000000000..43fb6b3ce --- /dev/null +++ b/cdk/test/handlers/shared/microvm-approval-wake.test.ts @@ -0,0 +1,256 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +// SPDX-License-Identifier: MIT-0 + +const mockRead = jest.fn(); +const mockSave = jest.fn(); +const mockSend = jest.fn(); +const mockEmit = jest.fn(); +const mockDispatch = jest.fn(); +jest.mock('../../../src/handlers/shared/microvm-continuation-dispatch', () => ({ + dispatchMicrovmContinuation: (...args: unknown[]) => mockDispatch(...args), +})); +const mockLogger = { warn: jest.fn(), info: jest.fn() }; +jest.mock('../../../src/handlers/shared/microvm-lifecycle', () => ({ + readMicrovmLifecycleSnapshot: (...args: unknown[]) => mockRead(...args), + saveMicrovmLifecycleIntent: (...args: unknown[]) => mockSave(...args), +})); +jest.mock('../../../src/handlers/shared/logger', () => ({ logger: mockLogger })); +jest.mock('@aws-sdk/client-lambda-microvms', () => ({ + LambdaMicrovmsClient: jest.fn(() => ({ send: mockSend })), + GetMicrovmCommand: jest.fn(input => ({ type: 'get', input })), + ResumeMicrovmCommand: jest.fn(input => ({ type: 'resume', input })), +})); + +import { approvalPostCommitOptions, wakeMicrovmAfterApproval } from '../../../src/handlers/shared/microvm-approval-wake'; +import type { MicrovmLifecycleSnapshot } from '../../../src/handlers/shared/microvm-lifecycle'; + +const NOW = Date.parse('2026-09-15T12:00:00Z'); +let row: MicrovmLifecycleSnapshot; +let controller: AbortController; +function wake(decision: 'APPROVED' | 'DENIED' = 'APPROVED') { + return wakeMicrovmAfterApproval({ + taskId: 'task', + userId: 'user', + requestId: 'gate', + decision, + options: { abortSignal: controller.signal }, + emitEvent: mockEmit, + }); +} +beforeEach(() => { + jest.clearAllMocks(); + mockDispatch.mockReset().mockResolvedValue(false); + jest.spyOn(Date, 'now').mockReturnValue(NOW); + controller = new AbortController(); + row = { + taskId: 'task', + userId: 'user', + status: 'AWAITING_APPROVAL', + requestId: 'gate', + handle: { strategyType: 'lambda-microvm', microvmId: 'vm', sessionId: 'vm', endpoint: 'https://vm.example' }, + approval: { + kind: 'present', + status: 'APPROVED', + created_at: new Date(NOW - 60_000).toISOString(), + timeout_s: 600, + createdAtMs: NOW - 60_000, + deadlineMs: NOW + 540_000, + }, + }; + mockRead.mockReset().mockImplementation(async () => structuredClone(row)); + mockSave.mockReset().mockImplementation(async (snapshot, action) => { + row = { + ...snapshot, + intent: { + version: 1, + microvm_id: 'vm', + request_id: snapshot.requestId, + generation: 'wake-generation', + action, + requested_at_ms: NOW, + deadline_ms: NOW + 540_000, + }, + }; + return { status: 'saved', intent: row.intent }; + }); + mockSend.mockReset().mockResolvedValue({ state: 'SUSPENDED' }); + mockEmit.mockReset().mockResolvedValue(undefined); +}); +afterEach(() => jest.restoreAllMocks()); + +test('dispatches a retired continuation without touching the fenced worker', async () => { + mockDispatch.mockResolvedValue(true); + await wake(); + expect(mockDispatch).toHaveBeenCalledWith('task', 'user', 'gate', expect.objectContaining({ + abortSignal: controller.signal, + })); + expect(mockRead).not.toHaveBeenCalled(); + expect(mockSave).not.toHaveBeenCalled(); + expect(mockSend).not.toHaveBeenCalled(); +}); + +test('does not fall back to the old worker when continuation dispatch is uncertain', async () => { + mockDispatch.mockRejectedValue(new Error('dispatch response lost')); + await wake(); + expect(mockRead).not.toHaveBeenCalled(); + expect(mockSave).not.toHaveBeenCalled(); + expect(mockSend).not.toHaveBeenCalled(); + expect(mockEmit).toHaveBeenCalledWith('microvm_resume_orphan', expect.objectContaining({ + reason: 'wake-reconciliation-failed', + }), expect.anything()); +}); + +test('ties the saved decision generation to the accepted AWS request without logging its body', async () => { + mockSend.mockResolvedValueOnce({ state: 'SUSPENDED' }).mockResolvedValueOnce({ + $metadata: { requestId: 'aws-inline-123' }, private: 'secret-response', + }); + await wake(); + expect(mockLogger.info).toHaveBeenCalledWith('MicroVM wake requested after approval decision', expect.objectContaining({ + task_id: 'task', + request_id: 'gate', + microvm_id: 'vm', + generation: 'wake-generation', + intent_requested_at_ms: NOW, + aws_request_id: 'aws-inline-123', + elapsed_ms: 0, + })); + expect(JSON.stringify(mockLogger.info.mock.calls)).not.toContain('secret-response'); + expect(mockRead).toHaveBeenCalledTimes(3); +}); + +test.each(['APPROVED', 'DENIED'] as const)('%s saves wake before Get, rechecks ownership, then requests Resume and reads again', async decision => { + if (row.approval.kind === 'present') row = { ...row, approval: { ...row.approval, status: decision } }; + await wake(decision); + expect(mockSend.mock.calls.map(([command]) => command.type)).toEqual(['get', 'resume']); + expect(mockSave.mock.calls[0][1]).toBe('resume'); + expect(mockSave.mock.invocationCallOrder[0]).toBeLessThan(mockSend.mock.invocationCallOrder[0]); + expect(mockRead).toHaveBeenCalledTimes(3); + expect(mockRead.mock.invocationCallOrder[1]).toBeLessThan(mockSend.mock.invocationCallOrder[1]); + expect(mockSend.mock.invocationCallOrder[1]).toBeLessThan(mockRead.mock.invocationCallOrder[2]); + expect(mockSend.mock.calls[1][0].input).toEqual({ microvmIdentifier: 'vm' }); + expect(row.intent?.action).toBe('resume'); +}); + +test.each(['RUNNING', 'SUSPENDING', 'PENDING', 'UNKNOWN', 'TERMINATED'])('%s retains wake intent without premature Resume', async state => { + mockSend.mockResolvedValue({ state }); + await wake(); + expect(row.intent?.action).toBe('resume'); + expect(mockSend.mock.calls.map(([command]) => command.type)).toEqual(['get']); +}); + +test.each(['missing', 'cancelled', 'new-gate', 'pending'])('%s snapshot cannot trigger compute control', async kind => { + if (kind === 'missing') mockRead.mockResolvedValue(undefined); + if (kind === 'cancelled') row = { ...row, status: 'CANCELLED' }; + if (kind === 'new-gate') row = { ...row, requestId: 'new-gate' }; + if (kind === 'pending' && row.approval.kind === 'present') row = { ...row, approval: { ...row.approval, status: 'PENDING' } }; + await wake(); + expect(mockSave).not.toHaveBeenCalled(); + expect(mockSend).not.toHaveBeenCalled(); +}); + +test.each(['stale', 'ineligible'])('a %s intent write cannot authorize Get or Resume', async status => { + mockSave.mockResolvedValue({ status }); + await wake(); + expect(mockSend).not.toHaveBeenCalled(); +}); + +test.each(['cancel', 'gate', 'worker', 'generation'])('%s winning after Get blocks Resume', async change => { + mockSend.mockImplementationOnce(async () => { + if (change === 'cancel') row = { ...row, status: 'CANCELLED' }; + if (change === 'gate') row = { ...row, requestId: 'new-gate' }; + if (change === 'worker') row = { ...row, handle: { ...row.handle, microvmId: 'other' } }; + if (change === 'generation') row = { ...row, intent: { ...row.intent!, generation: 'other' } }; + return { state: 'SUSPENDED' }; + }); + await wake(); + expect(mockSend).toHaveBeenCalledTimes(1); +}); + +test('a consumed decision can fence its old suspend while task RUNNING', async () => { + row = { + ...row, + status: 'RUNNING', + requestId: null, + approval: { kind: 'none' }, + intent: { + version: 1, + microvm_id: 'vm', + request_id: 'gate', + action: 'suspend', + generation: 'old', + requested_at_ms: NOW - 5_000, + deadline_ms: NOW + 540_000, + }, + }; + await wake(); + expect(row.intent?.request_id).toBeNull(); + expect(row.intent?.action).toBe('resume'); + expect(mockSend.mock.calls.map(([command]) => command.type)).toEqual(['get', 'resume']); +}); + +test('a failed Resume still reads again, retains wake and only logs safe identifiers', async () => { + mockSend.mockResolvedValueOnce({ state: 'SUSPENDED' }).mockRejectedValueOnce(Object.assign( + new Error('private SDK details'), { name: 'AccessDeniedException', $metadata: { requestId: 'safe-123' } }, + )); + await expect(wake()).resolves.toBeUndefined(); + expect(mockRead).toHaveBeenCalledTimes(3); + expect(row.intent?.action).toBe('resume'); + expect(mockLogger.warn.mock.calls[0][1]).toMatchObject({ error_type: 'AccessDeniedException', aws_request_id: 'safe-123' }); + expect(mockEmit).toHaveBeenCalledWith('microvm_resume_orphan', expect.objectContaining({ + task_id: 'task', + request_id: 'gate', + microvm_id: 'vm', + stage: 'resume-request', + reason: 'resume-request-failed', + error_type: 'AccessDeniedException', + aws_request_id: 'safe-123', + }), expect.objectContaining({ abortSignal: expect.any(AbortSignal) })); + expect(JSON.stringify(mockLogger.warn.mock.calls)).not.toContain('private SDK details'); +}); + +test('an orphan-event failure remains best-effort and never discards the wake intent', async () => { + mockSend.mockResolvedValueOnce({ state: 'SUSPENDED' }).mockRejectedValueOnce(new Error('request failed')); + mockEmit.mockRejectedValue(new Error('private audit error')); + await expect(wake()).resolves.toBeUndefined(); + expect(row.intent?.action).toBe('resume'); + expect(mockRead).toHaveBeenCalledTimes(3); + expect(JSON.stringify(mockLogger.warn.mock.calls)).not.toContain('private audit error'); +}); + +test.each(['before-read', 'after-read', 'after-get'])('expired parent budget %s blocks subsequent work', async stage => { + if (stage === 'before-read') controller.abort(); + if (stage === 'after-read') mockRead.mockImplementationOnce(async () => { controller.abort(); return row; }); + if (stage === 'after-get') mockSend.mockImplementationOnce(async () => { controller.abort(); return { state: 'SUSPENDED' }; }); + await wake(); + if (stage === 'before-read') expect(mockRead).not.toHaveBeenCalled(); + if (stage !== 'after-get') expect(mockSave).not.toHaveBeenCalled(); + expect(mockSend.mock.calls.filter(([command]) => command.type === 'resume')).toHaveLength(0); +}); + +test('postcommit budget reserves response time and does not restart an expired invocation', () => { + const timeout = jest.spyOn(AbortSignal, 'timeout'); + approvalPostCommitOptions(NOW, { getRemainingTimeInMillis: () => 14_000 }); + expect(timeout).toHaveBeenLastCalledWith(8_000); + approvalPostCommitOptions(NOW, { getRemainingTimeInMillis: () => 2_000 }); + expect(timeout).toHaveBeenLastCalledWith(1_000); + expect(approvalPostCommitOptions(NOW, { getRemainingTimeInMillis: () => 900 }).abortSignal?.aborted).toBe(true); + expect(approvalPostCommitOptions(NOW - 14_500).abortSignal?.aborted).toBe(true); +}); diff --git a/cdk/test/handlers/shared/microvm-continuation-dispatch.test.ts b/cdk/test/handlers/shared/microvm-continuation-dispatch.test.ts new file mode 100644 index 000000000..217eaf739 --- /dev/null +++ b/cdk/test/handlers/shared/microvm-continuation-dispatch.test.ts @@ -0,0 +1,110 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +const mockSend = jest.fn(); +const mockInvoke = jest.fn(); +const mockAdmit = jest.fn(); +const mockWarn = jest.fn(); +jest.mock('../../../src/handlers/shared/ua', () => ({ + makeDocClient: () => ({ send: mockSend }), makeClient: () => ({ send: mockInvoke }), +})); +jest.mock('../../../src/handlers/shared/microvm-continuation-start', () => ({ + admitContinuation: (...args: unknown[]) => mockAdmit(...args), +})); +jest.mock('../../../src/handlers/shared/logger', () => ({ logger: { info: jest.fn(), warn: mockWarn } })); + +import { dispatchMicrovmContinuation } from '../../../src/handlers/shared/microvm-continuation-dispatch'; + +let task: any; +beforeEach(() => { + jest.clearAllMocks(); + process.env.CONTINUATION_BUCKET_NAME = 'saved'; + process.env.ORCHESTRATOR_FUNCTION_ARN = 'arn:aws:lambda:us-west-2:123456789012:function:coordinator:live'; + task = { + task_id: 'task', + user_id: 'user', + compute_type: 'lambda-microvm', + status: 'AWAITING_APPROVAL', + awaiting_approval_request_id: 'request', + continuation: { state: 'PARKED', attempt_id: 'new-worker' }, + continuation_launch: { orchestrator_version: '42' }, + }; + mockSend.mockImplementation(async () => ({ Item: structuredClone(task) })); + mockAdmit.mockImplementation(async () => ({ kind: 'ready', task: structuredClone(task) })); + mockInvoke.mockResolvedValue({ StatusCode: 202 }); +}); +afterEach(() => { + delete process.env.CONTINUATION_BUCKET_NAME; + delete process.env.ORCHESTRATOR_FUNCTION_ARN; +}); + +test('uses the original published version and identical durable identity on repeated dispatch', async () => { + expect(await dispatchMicrovmContinuation('task', 'user', 'request')).toBe(true); + expect(await dispatchMicrovmContinuation('task', 'user', 'request')).toBe(true); + const [first, second] = mockInvoke.mock.calls.map(([command]) => command.input); + expect(first).toEqual(second); + expect(first.FunctionName).toBe('arn:aws:lambda:us-west-2:123456789012:function:coordinator:42'); + expect(first.DurableExecutionName).toMatch(/^[a-f0-9]{64}$/); + expect(first.DurableExecutionName.length).toBeLessThanOrEqual(64); + expect(JSON.parse(first.Payload.toString())).toEqual({ + task_id: 'task', continuation_request_id: 'request', continuation_attempt_id: 'new-worker', + }); +}); + +test('does not invoke a worker while capacity is unavailable', async () => { + mockAdmit.mockResolvedValue({ kind: 'capacity' }); + expect(await dispatchMicrovmContinuation('task', 'user', 'request')).toBe(true); + expect(mockInvoke).not.toHaveBeenCalled(); +}); + +test('fenced source must not be resumed while retirement is finishing', async () => { + task.continuation.state = 'FENCED'; + expect(await dispatchMicrovmContinuation('task', 'user', 'request')).toBe(true); + expect(mockAdmit).not.toHaveBeenCalled(); + expect(mockInvoke).not.toHaveBeenCalled(); +}); + +test.each(['READY', 'CONSUMED'])('leaves ordinary %s worker handling to the wake path', async state => { + task.continuation.state = state; + expect(await dispatchMicrovmContinuation('task', 'user', 'request')).toBe(false); + expect(mockInvoke).not.toHaveBeenCalled(); +}); + +test('an invocation failure preserves the assigned token for the scheduled retry', async () => { + mockInvoke.mockRejectedValue(new Error('lost invocation response')); + await expect(dispatchMicrovmContinuation('task', 'user', 'request')).rejects.toThrow('lost invocation'); + expect(task.continuation.attempt_id).toBe('new-worker'); +}); + +test('reports bounded validation detail and operation identity for dispatch diagnosis', async () => { + const error = Object.assign(new Error('durableExecutionName must have length less than or equal to 64'), { + name: 'ValidationException', + }); + mockInvoke.mockRejectedValue(error); + await expect(dispatchMicrovmContinuation('task', 'user', 'request')).rejects.toBe(error); + expect(mockWarn).toHaveBeenCalledWith( + 'Saved task continuation dispatch needs reconciliation', + expect.objectContaining({ + operation: 'Invoke', + coordinator_version: '42', + durable_execution_name_length: 64, + validation_detail: error.message, + }), + ); +}); diff --git a/cdk/test/handlers/shared/microvm-continuation-retirement.test.ts b/cdk/test/handlers/shared/microvm-continuation-retirement.test.ts new file mode 100644 index 000000000..c06e8b772 --- /dev/null +++ b/cdk/test/handlers/shared/microvm-continuation-retirement.test.ts @@ -0,0 +1,195 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +const mockSend = jest.fn(); +const mockVerify = jest.fn(); +const mockStopSession = jest.fn(); +jest.mock('../../../src/handlers/shared/ua', () => ({ makeDocClient: () => ({ send: mockSend }) })); +jest.mock('../../../src/handlers/shared/microvm-continuation-storage', () => ({ + verifyContinuationCheckpoint: (...args: unknown[]) => mockVerify(...args), +})); +jest.mock('@aws-sdk/lib-dynamodb', () => ({ + GetCommand: jest.fn(input => ({ kind: 'get', input })), + TransactWriteCommand: jest.fn(input => ({ kind: 'transact', input })), +})); +jest.mock('../../../src/handlers/shared/logger', () => ({ logger: { warn: jest.fn() } })); + +import type { ComputeStrategy } from '../../../src/handlers/shared/compute-strategy'; +import { retireCheckpointedMicrovm } from '../../../src/handlers/shared/microvm-continuation-retirement'; + +const handle = { + strategyType: 'lambda-microvm' as const, + microvmId: 'microvm-one', + sessionId: 'microvm-one', + endpoint: 'https://example.invalid', +}; +const identity = { task_id: 'task', attempt_id: 'microvm-one', request_id: 'request', user_id: 'user', repo: 'owner/repo' }; +const record = { + version: 1, + state: 'READY', + identity, + manifest: { + kind: 'manifest', + key: `continuations/task/microvm-one/request/manifest/${'a'.repeat(64)}.json`, + sha256: 'a'.repeat(64), + version_id: 'version-1', + size_bytes: 100, + }, +}; +let task: Record; +let strategy: ComputeStrategy; +let transactions: any[]; +let leaseState: string; + +beforeEach(() => { + jest.clearAllMocks(); + process.env.CONTINUATION_BUCKET_NAME = 'continuations'; + task = { + task_id: 'task', + user_id: 'user', + repo: 'owner/repo', + status: 'AWAITING_APPROVAL', + session_id: handle.microvmId, + awaiting_approval_request_id: 'request', + continuation: structuredClone(record), + microvm_start: { clientToken: 'task' }, + concurrency_slot: { state: 'held', acquired_at: '2026-09-17T00:00:00Z' }, + }; + transactions = []; + leaseState = 'ACTIVE'; + mockVerify.mockResolvedValue(undefined); + mockSend.mockImplementation(async ({ kind, input }) => { + if (kind === 'get' && input.Key.task_id === 'worker-lease#task') { + return { + Item: { + lease_state: leaseState, + lease_attempt_id: 'task', + lease_microvm_id: 'microvm-one', + lease_user_id: 'user', + lease_repo: 'owner/repo', + }, + }; + } + if (kind === 'get') { + return input.Key.request_id + ? { Item: { user_id: 'user', status: 'PENDING', created_at: new Date(Date.now() - 7200000).toISOString() } } + : { Item: structuredClone(task) }; + } + transactions.push(input.TransactItems); + const values = input.TransactItems[0].Update.ExpressionAttributeValues; + task.continuation = structuredClone(values[':fenced'] ?? values[':parked']); + leaseState = task.continuation.state; + if (leaseState === 'PARKED') task.concurrency_slot = { ...task.concurrency_slot, state: 'released' }; + return {}; + }); + strategy = { + type: 'lambda-microvm', + startSession: jest.fn(), + suspendSession: jest.fn(), + resumeSession: jest.fn(), + stopSession: mockStopSession.mockResolvedValue({ outcome: 'requested' }), + pollSession: jest.fn().mockResolvedValue({ status: 'running', microvmState: 'TERMINATING' }), + }; +}); +afterEach(() => { delete process.env.CONTINUATION_BUCKET_NAME; }); + +function run(force = false) { + return retireCheckpointedMicrovm({ + taskId: 'task', userId: 'user', handle, strategy, sessionDeadlineMs: Date.now() + 28800000, force, + }); +} + +test('fences before termination, holds capacity until terminal observation, then parks once', async () => { + expect(await run()).toBe('stopping'); + expect(mockVerify).toHaveBeenCalledWith(record, expect.objectContaining({ abortSignal: expect.any(AbortSignal) })); + expect(transactions).toHaveLength(1); + expect(transactions[0][1].Update.ExpressionAttributeValues).toMatchObject({ + ':fenced': 'FENCED', ':active': 'ACTIVE', ':attempt': 'task', ':vm': 'microvm-one', + }); + expect(transactions[0][0].Update.ConditionExpression).toContain('continuation = :record'); + expect(task.status).toBe('AWAITING_APPROVAL'); + (strategy.pollSession as jest.Mock).mockResolvedValue({ status: 'completed', microvmState: 'TERMINATED' }); + expect(await run()).toBe('parked'); + expect(transactions).toHaveLength(2); + expect(transactions[1][2].Update.UpdateExpression).toContain('active_count = active_count - :one'); + expect(transactions[1][0].Update.ExpressionAttributeValues[':slot'].state).toBe('held'); + expect(task.concurrency_slot.state).toBe('released'); + expect(task.awaiting_approval_request_id).toBe('request'); + expect(await run()).toBe('parked'); + expect(transactions).toHaveLength(2); +}); + +test('a human answer winning the fence race leaves the original worker running', async () => { + mockSend.mockImplementation(async ({ kind, input }) => { + if (kind === 'get') { + return input.Key.request_id + ? { Item: { created_at: new Date(Date.now() - 7200000).toISOString() } } + : { Item: structuredClone(task) }; + } + task.status = 'RUNNING'; + delete task.continuation; + throw new Error('fence condition lost'); + }); + expect(await run()).toBe('not-due'); + expect(mockStopSession).not.toHaveBeenCalled(); +}); + +test('missing durable objects cannot retire the only copy of the workspace', async () => { + mockVerify.mockRejectedValue(new Error('missing pinned workspace version')); + await expect(run()).rejects.toThrow('missing pinned workspace'); + expect(transactions).toHaveLength(0); + expect(mockStopSession).not.toHaveBeenCalled(); +}); + +test('sleep off retains the worker until its lifetime margin', async () => { + task.microvm_sleep_after_s = 0; + expect(await run()).toBe('not-due'); + expect(mockStopSession).not.toHaveBeenCalled(); + expect(await retireCheckpointedMicrovm({ + taskId: 'task', userId: 'user', handle, strategy, sessionDeadlineMs: Date.now() + 100000, + })).toBe('stopping'); +}); + +test('a stale coordinator never follows or stops a replacement worker', async () => { + task.session_id = 'microvm-new'; + expect(await run(true)).toBe('ownership-lost'); + expect(transactions).toHaveLength(0); + expect(mockStopSession).not.toHaveBeenCalled(); +}); + +test.each(['FENCED', 'PARKED'])('a worker-written %s label cannot replace coordinator lease authority', async state => { + task.continuation = { ...task.continuation, state, source_handle: handle }; + await expect(run()).rejects.toThrow('no coordinator authority'); + expect(mockStopSession).not.toHaveBeenCalled(); + expect(transactions).toHaveLength(0); +}); + +test('lost fence reply cannot accept a worker-written label while its readonly lease remains active', async () => { + const normal = mockSend.getMockImplementation()!; + mockSend.mockImplementation(async (command, options) => { + if (command.kind === 'transact') { + task.continuation = { ...task.continuation, state: 'FENCED', source_handle: handle }; + throw new Error('fence did not commit'); + } + return normal(command, options); + }); + expect(await run()).toBe('not-due'); + expect(mockStopSession).not.toHaveBeenCalled(); + expect(leaseState).toBe('ACTIVE'); +}); diff --git a/cdk/test/handlers/shared/microvm-continuation-runner.test.ts b/cdk/test/handlers/shared/microvm-continuation-runner.test.ts new file mode 100644 index 000000000..7f2fecb50 --- /dev/null +++ b/cdk/test/handlers/shared/microvm-continuation-runner.test.ts @@ -0,0 +1,186 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +const mockLoad = jest.fn(); +const mockLaunch = jest.fn(); +const mockAdmit = jest.fn(); +const mockStart = jest.fn(); +const mockPoll = jest.fn(); +const mockWorkerPoll = jest.fn(); +const mockStop = jest.fn(); +const mockFinalize = jest.fn(); +const mockDelete = jest.fn(); +const mockSend = jest.fn(); +jest.mock('../../../src/handlers/shared/ua', () => ({ makeDocClient: () => ({ send: mockSend }) })); +jest.mock('../../../src/handlers/shared/orchestrator', () => ({ + loadTask: (...args: unknown[]) => mockLoad(...args), + emitTaskEvent: jest.fn(), + envelopeFor: () => ({ correlation: {}, log: { error: jest.fn() } }), + finalizeTask: (...args: unknown[]) => mockFinalize(...args), +})); +jest.mock('../../../src/handlers/shared/compute-strategy', () => ({ + resolveComputeStrategy: () => ({ + startSession: mockStart, pollSession: mockWorkerPoll, + }), +})); +jest.mock('../../../src/handlers/shared/microvm-continuation-storage', () => ({ + loadContinuationLaunch: (...args: unknown[]) => mockLaunch(...args), +})); +jest.mock('../../../src/handlers/shared/microvm-continuation-start', () => ({ + admitContinuation: (...args: unknown[]) => mockAdmit(...args), +})); +jest.mock('../../../src/handlers/shared/microvm-task-poll', () => ({ + pollMicrovmTask: (...args: unknown[]) => mockPoll(...args), +})); +jest.mock('../../../src/handlers/shared/microvm-supervisor', () => ({ + stopMicrovmWithDiagnostics: (...args: unknown[]) => mockStop(...args), +})); +jest.mock('../../../src/handlers/shared/strategies/lambda-microvm-strategy', () => ({ + deleteMicrovmPayload: (...args: unknown[]) => mockDelete(...args), +})); + +import type { DurableContext } from '@aws/durable-execution-sdk-js'; +import type { ComputeStrategy } from '../../../src/handlers/shared/compute-strategy'; +import { + continuationWaitStrategy, failContinuationAttempt, pollContinuationRestore, runMicrovmContinuation, +} from '../../../src/handlers/shared/microvm-continuation-runner'; + +const event = { task_id: 'task', continuation_request_id: 'request', continuation_attempt_id: 'attempt-new' }; +const handle = { strategyType: 'lambda-microvm' as const, microvmId: 'vm-new', sessionId: 'vm-new', endpoint: 'https://worker.invalid' }; +let task: any; +const context = { + step: async (_name: string, action: () => Promise) => action(), + waitForCondition: async (_name: string, check: (state: any) => Promise, options: any) => { + let state = options.initialState; + for (let i = 0; i < 3; i++) { + state = await check(state); + if (!options.waitStrategy(state).shouldContinue) return state; + } + throw new Error('fixture wait did not complete'); + }, +} as unknown as DurableContext; + +beforeEach(() => { + jest.clearAllMocks(); + process.env.AWS_LAMBDA_FUNCTION_VERSION = '42'; + task = { + task_id: 'task', + user_id: 'user', + status: 'AWAITING_APPROVAL', + compute_type: 'lambda-microvm', + awaiting_approval_request_id: 'request', + continuation_launch: { orchestrator_version: '42' }, + continuation: { + state: 'STARTING', + attempt_id: 'attempt-new', + started_at: new Date().toISOString(), + source_handle: { imageArn: 'arn:image:original', imageVersion: '7.0' }, + }, + }; + mockLoad.mockImplementation(async () => structuredClone(task)); + mockLaunch.mockResolvedValue({ + orchestrator_version: '42', + payload: { task_id: 'task', user_id: 'user', message: 'original instructions' }, + blueprint: { compute_type: 'lambda-microvm' }, + }); + mockAdmit.mockImplementation(async () => ({ kind: 'ready', task: structuredClone(task) })); + mockStart.mockImplementation(async () => { + task.session_id = 'vm-new'; + task.compute_metadata = { microvmId: 'vm-new' }; + task.microvm_start = { clientToken: 'attempt-new', handle }; + task.continuation.state = 'CONSUMED'; + task.status = 'RUNNING'; + return handle; + }); + mockPoll.mockImplementation(async () => { + task.status = 'COMPLETED'; + return { attempts: 1, lastStatus: 'COMPLETED' }; + }); + mockWorkerPoll.mockResolvedValue({ status: 'running', microvmState: 'RUNNING' }); + mockSend.mockImplementation(async () => { task.status = 'FAILED'; return {}; }); +}); +afterEach(() => { delete process.env.AWS_LAMBDA_FUNCTION_VERSION; }); + +test('runs saved inputs on the pinned original image with a fresh worker lifetime', async () => { + await runMicrovmContinuation(event, context); + expect(mockStart).toHaveBeenCalledWith(expect.objectContaining({ + microvmImage: { imageArn: 'arn:image:original', imageVersion: '7.0' }, + payload: { + task_id: 'task', + user_id: 'user', + message: 'original instructions', + attempt_id: 'attempt-new', + task_started_at: task.continuation.started_at, + }, + })); + expect(mockFinalize).toHaveBeenCalledTimes(1); + expect(mockDelete).toHaveBeenCalledWith('task', 'attempt-new'); + expect(mockStop).toHaveBeenCalledWith(expect.objectContaining({ handle })); +}); + +test('a new parked approval ends this execution without finalizing the task', async () => { + mockPoll.mockResolvedValue({ attempts: 1, microvmParked: true }); + await runMicrovmContinuation(event, context); + expect(mockFinalize).not.toHaveBeenCalled(); + expect(mockStop).not.toHaveBeenCalled(); + expect(mockDelete).toHaveBeenCalledWith('task', 'attempt-new'); +}); + +test('a stale continuation event cannot launch or finalize a different assigned worker', async () => { + task.continuation.attempt_id = 'another'; + await runMicrovmContinuation(event, context); + expect(mockStart).not.toHaveBeenCalled(); + expect(mockFinalize).not.toHaveBeenCalled(); +}); + +test('a different coordinator version fails before launching a new worker', async () => { + process.env.AWS_LAMBDA_FUNCTION_VERSION = '43'; + await runMicrovmContinuation(event, context); + expect(mockStart).not.toHaveBeenCalled(); + const transaction = mockSend.mock.calls[0][0].input.TransactItems; + expect(transaction[0].Update.ExpressionAttributeValues[':detail']).toContain('VERSION_CHANGED'); + expect(transaction[1].Update.ExpressionAttributeValues[':fenced']).toBe('FENCED'); +}); + +test('restoration uses its own deadline and never issues /resume', async () => { + task.session_id = 'vm-new'; + task.compute_metadata = { microvmId: 'vm-new' }; + task.continuation.state = 'RESTORING'; + const resume = jest.fn(); + const strategy = { pollSession: mockWorkerPoll, resumeSession: resume } as unknown as ComputeStrategy; + const deadlineMs = Date.now() + 600_000; + const waiting = await pollContinuationRestore(event, 'user', handle, strategy, { deadlineMs }); + expect(waiting).toEqual({ deadlineMs, consecutivePollFailures: 0 }); + expect(resume).not.toHaveBeenCalled(); + const failed = await pollContinuationRestore(event, 'user', handle, strategy, { deadlineMs: Date.now() - 1 }); + expect(failed.failure).toContain('RESTORE_TIMEOUT'); +}); + +test('failed recovery fences the current attempt even after a new checkpoint replaced the old record', async () => { + task.microvm_start = { clientToken: 'attempt-new', handle }; + task.continuation = { state: 'READY', identity: { attempt_id: 'vm-new' } }; + await failContinuationAttempt(event, 'user', 'restore failed'); + expect(mockSend).toHaveBeenCalledTimes(1); + expect(mockSend.mock.calls[0][0].input.TransactItems[0].Update.ConditionExpression).toContain('microvm_start.clientToken = :attempt'); +}); + +test('retirement reconciliation delays do not trigger ordinary failure finalization', () => { + expect(continuationWaitStrategy({ attempts: 99, microvmRetiring: true, microvmRetirementError: 'Timeout' })) + .toEqual({ shouldContinue: true, delay: { seconds: 30 } }); +}); diff --git a/cdk/test/handlers/shared/microvm-continuation-start.test.ts b/cdk/test/handlers/shared/microvm-continuation-start.test.ts new file mode 100644 index 000000000..ad2a0893c --- /dev/null +++ b/cdk/test/handlers/shared/microvm-continuation-start.test.ts @@ -0,0 +1,145 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +const mockSend = jest.fn(); +jest.mock('../../../src/handlers/shared/ua', () => ({ makeDocClient: () => ({ send: mockSend }) })); +jest.mock('@aws-sdk/lib-dynamodb', () => ({ + GetCommand: jest.fn(input => ({ kind: 'get', input })), + TransactWriteCommand: jest.fn(input => ({ kind: 'transact', input })), +})); + +import { admitContinuation } from '../../../src/handlers/shared/microvm-continuation-start'; + +const identity = { task_id: 'task', user_id: 'user', repo: 'owner/repo', attempt_id: 'vm-old', request_id: 'request' }; +let task: Record; +let lease: Record; +let approval: Record; +let transactions: any[]; + +beforeEach(() => { + jest.clearAllMocks(); + task = { + task_id: 'task', + user_id: 'user', + repo: 'owner/repo', + status: 'AWAITING_APPROVAL', + session_id: 'vm-old', + awaiting_approval_request_id: 'request', + concurrency_slot: { state: 'released' }, + microvm_start: { clientToken: 'task' }, + continuation: { + version: 1, + state: 'PARKED', + identity, + source_handle: { microvmId: 'vm-old' }, + manifest: { + kind: 'manifest', + key: `continuations/task/vm-old/request/manifest/${'a'.repeat(64)}.json`, + sha256: 'a'.repeat(64), + version_id: 'v1', + size_bytes: 100, + }, + }, + }; + lease = { lease_state: 'PARKED', lease_attempt_id: 'task', lease_microvm_id: 'vm-old', lease_user_id: 'user', lease_repo: 'owner/repo' }; + approval = { user_id: 'user', status: 'APPROVED' }; + transactions = []; + mockSend.mockImplementation(async ({ kind, input }) => { + if (kind === 'get') { + return { + Item: structuredClone( + input.Key.request_id ? approval : input.Key.task_id.startsWith('worker-lease#') ? lease : task, + ), + }; + } + transactions.push(input.TransactItems); + const values = input.TransactItems[0].Update.ExpressionAttributeValues; + task.continuation = values[':starting']; + task.concurrency_slot = values[':slot']; + lease = input.TransactItems[2].Put.Item; + return {}; + }); +}); + +test('claims exactly one seat and worker token; a replay verifies the same readonly lease', async () => { + const first = await admitContinuation('task', 'user', 'request', 3); + const second = await admitContinuation('task', 'user', 'request', 3); + expect(first.kind).toBe('ready'); + expect(second.kind).toBe('ready'); + expect(transactions).toHaveLength(1); + expect(lease.lease_state).toBe('ACTIVE'); + expect(task.concurrency_slot.attempt_id).toBe(lease.lease_attempt_id); + expect(transactions[0][1].Update.ConditionExpression).toContain('active_count < :limit'); + expect(transactions[0][2].Put.ConditionExpression).toContain('lease_state = :parked'); + expect(transactions[0][3].ConditionCheck.ExpressionAttributeValues[':approved']).toBe('APPROVED'); +}); + +test.each(['PENDING', 'CANCELLED'])('does not allocate capacity for %s approval', async status => { + approval.status = status; + expect((await admitContinuation('task', 'user', 'request', 3)).kind).toBe(status === 'PENDING' ? 'waiting' : 'closed'); + expect(transactions).toHaveLength(0); +}); + +test('rejects worker-writable STARTING state without a matching active readonly lease', async () => { + task.continuation = { ...task.continuation, state: 'STARTING', attempt_id: 'new' }; + task.concurrency_slot = { state: 'held', attempt_id: 'new' }; + await expect(admitContinuation('task', 'user', 'request', 3)).rejects.toThrow('LEASE_INVALID'); + expect(transactions).toHaveLength(0); +}); + +test('recovers a transaction that committed but lost its reply without allocating again', async () => { + const normal = mockSend.getMockImplementation()!; + mockSend.mockImplementation(async (command, options) => { + const result = await normal(command, options); + if (command.kind === 'transact') throw new Error('reply lost'); + return result; + }); + expect((await admitContinuation('task', 'user', 'request', 3)).kind).toBe('ready'); + expect(transactions).toHaveLength(1); +}); + +test('capacity denial leaves the same saved request parked', async () => { + const normal = mockSend.getMockImplementation()!; + mockSend.mockImplementation(async (command, options) => { + if (command.kind === 'transact') { + throw Object.assign(new Error('capacity'), { + name: 'TransactionCanceledException', CancellationReasons: [{ Code: 'None' }, { Code: 'ConditionalCheckFailed' }], + }); + } + return normal(command, options); + }); + expect((await admitContinuation('task', 'user', 'request', 3)).kind).toBe('capacity'); + expect(task.continuation.state).toBe('PARKED'); + expect(lease.lease_state).toBe('PARKED'); +}); + +test('an explicit elapsed deadline is resolved atomically with replacement admission', async () => { + approval = { ...approval, status: 'PENDING', timeout_s: 30, created_at: new Date(Date.now() - 60_000).toISOString() }; + expect((await admitContinuation('task', 'user', 'request', 3)).kind).toBe('ready'); + expect(transactions[0][3].Update).toMatchObject({ + ExpressionAttributeValues: { ':timeout': 30, ':created': approval.created_at }, + }); + expect(transactions[0][3].Update.UpdateExpression).toContain('#status'); +}); + +test('timeout zero remains unanswered regardless of request age', async () => { + approval = { ...approval, status: 'PENDING', timeout_s: 0, created_at: '2020-01-01T00:00:00Z' }; + expect((await admitContinuation('task', 'user', 'request', 3)).kind).toBe('waiting'); + expect(transactions).toHaveLength(0); +}); diff --git a/cdk/test/handlers/shared/microvm-continuation-storage.test.ts b/cdk/test/handlers/shared/microvm-continuation-storage.test.ts new file mode 100644 index 000000000..615e7018e --- /dev/null +++ b/cdk/test/handlers/shared/microvm-continuation-storage.test.ts @@ -0,0 +1,215 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { createHash } from 'node:crypto'; +import { Readable } from 'node:stream'; + +const mockS3 = jest.fn(); +const mockDdb = jest.fn(); +jest.mock('../../../src/handlers/shared/ua', () => ({ + makeClient: () => ({ send: mockS3 }), + makeDocClient: () => ({ send: mockDdb }), +})); +import { + deleteClosedTaskContinuations, loadContinuationLaunch, saveContinuationLaunch, verifyContinuationCheckpoint, +} from '../../../src/handlers/shared/microvm-continuation-storage'; +import type { ContinuationRecord } from '../../../src/handlers/shared/microvm-continuation-types'; + +const digest = (bytes: Buffer) => createHash('sha256').update(bytes).digest('hex'); +const payload = { task_id: 'task', user_id: 'user', prompt: 'saved instruction' }; +const blueprint = { compute_type: 'lambda-microvm' as const }; +const inputs = { version: 1, task_id: 'task', user_id: 'user', payload, blueprint, orchestrator_version: '42' }; +function launch(value: unknown = inputs) { + const bytes = Buffer.from(JSON.stringify(value)); + const sha256 = digest(bytes); + return { + bytes, + receipt: { + version: 1, + key: `continuations/task/launch/${sha256}.json`, + sha256, + version_id: 'version-1', + size_bytes: bytes.length, + orchestrator_version: '42', + }, + }; +} +function object(bytes: Buffer, overrides: Record = {}) { + return { Body: Readable.from([bytes]), ContentLength: bytes.length, VersionId: 'version-1', ...overrides }; +} +beforeEach(() => { + mockS3.mockReset(); + mockDdb.mockReset(); + process.env.CONTINUATION_BUCKET_NAME = 'bucket'; + process.env.AWS_LAMBDA_FUNCTION_VERSION = '42'; +}); +afterEach(() => { + delete process.env.CONTINUATION_BUCKET_NAME; + delete process.env.AWS_LAMBDA_FUNCTION_VERSION; +}); + +test('loads the exact published launch version and rejects changed bytes', async () => { + const saved = launch(); + mockS3.mockResolvedValueOnce(object(saved.bytes)); + expect(await loadContinuationLaunch('task', 'user', saved.receipt)).toEqual(inputs); + expect(mockS3.mock.calls[0][0].input.VersionId).toBe('version-1'); + mockS3.mockResolvedValueOnce(object(Buffer.from('x'.repeat(saved.bytes.length)))); + await expect(loadContinuationLaunch('task', 'user', saved.receipt)).rejects.toThrow('checksum'); +}); + +test.each([ + { VersionId: 'null' }, { VersionId: 'different' }, { ContentLength: 0 }, { ContentLength: 100_000_000 }, +])('rejects incomplete or unpinned object metadata %j and closes the stream', async overrides => { + const saved = launch(); + const response = object(saved.bytes, overrides); + mockS3.mockResolvedValue(response); + await expect(loadContinuationLaunch('task', 'user', saved.receipt)).rejects.toThrow('STORAGE_INVALID'); + expect(response.Body.destroyed).toBe(true); +}); + +test('stops a response that sends more bytes than declared', async () => { + const saved = launch(); + const response = object(Buffer.concat([saved.bytes, Buffer.from('extra')]), { ContentLength: saved.bytes.length }); + mockS3.mockResolvedValue(response); + await expect(loadContinuationLaunch('task', 'user', saved.receipt)).rejects.toThrow('exceeded'); + expect(response.Body.destroyed).toBe(true); +}); + +test('aborting checkpoint verification closes a stalled response body', async () => { + const controller = new AbortController(); + const body = new Readable({ read() { /* Deliberately stalled transport. */ } }); + mockS3.mockImplementation(async () => { + setImmediate(() => controller.abort()); + return { Body: body, ContentLength: 10, VersionId: 'v1' }; + }); + const record = { manifest: { key: 'manifest', version_id: 'v1' } } as ContinuationRecord; + await expect(verifyContinuationCheckpoint(record, { abortSignal: controller.signal })).rejects.toThrow('TIMEOUT'); + expect(body.destroyed).toBe(true); +}); + +test.each([null, { ...inputs, user_id: 'other' }])('rejects saved input identity %j', async value => { + const saved = launch(value); + mockS3.mockResolvedValue(object(saved.bytes)); + await expect(loadContinuationLaunch('task', 'user', saved.receipt)).rejects.toThrow('INPUT_INVALID'); +}); + +test('recovers a lost Put reply only after exact versioned readback', async () => { + let published: Buffer; + mockS3.mockImplementation(async command => { + if (command.constructor.name === 'PutObjectCommand') { + published = command.input.Body; + throw new Error('reply lost'); + } + return object(published); + }); + mockDdb.mockResolvedValue({}); + await saveContinuationLaunch('task', 'user', payload, blueprint); + expect(mockS3.mock.calls[0][0].input).toMatchObject({ IfNoneMatch: '*', ServerSideEncryption: 'AES256' }); + const committed = mockDdb.mock.calls[0][0].input; + expect(committed.ExpressionAttributeValues[':receipt']).toMatchObject({ version_id: 'version-1', orchestrator_version: '42' }); + expect(committed.UpdateExpression).toContain('REMOVE #ttl'); +}); + +test('verifies both archive versions and checksums before permitting retirement', async () => { + const identity = { task_id: 'task', user_id: 'user', repo: 'owner/repo', attempt_id: 'vm', request_id: 'request' }; + const prefix = 'continuations/task/vm/request/'; + const conversation = { key: `${prefix}${'a'.repeat(64)}.json`, sha256: 'a'.repeat(64), size_bytes: 100, version_id: 'conversation-v' }; + const workspace = { key: `${prefix}workspace/${'b'.repeat(64)}.tar`, sha256: 'b'.repeat(64), size_bytes: 512, version_id: 'workspace-v' }; + const bytes = Buffer.from(JSON.stringify({ version: 1, identity, conversation, workspace })); + const record: ContinuationRecord = { + version: 1, + state: 'READY', + identity, + manifest: { kind: 'manifest', key: 'manifest', sha256: digest(bytes), size_bytes: bytes.length, version_id: 'version-1' }, + }; + mockS3.mockResolvedValueOnce(object(bytes)); + for (const receipt of [conversation, workspace]) { + mockS3.mockResolvedValueOnce({ + VersionId: receipt.version_id, + ContentLength: receipt.size_bytes, + ChecksumSHA256: Buffer.from(receipt.sha256, 'hex').toString('base64'), + }); + } + await verifyContinuationCheckpoint(record); + expect(mockS3.mock.calls.slice(1).map(([command]) => command.input.VersionId)).toEqual(['conversation-v', 'workspace-v']); + mockS3.mockResolvedValueOnce(object(bytes)).mockResolvedValueOnce({ VersionId: 'wrong' }); + await expect(verifyContinuationCheckpoint(record)).rejects.toThrow('version, length or checksum'); +}); + +test.each(['checksum', 'length', 'identity', 'version'])('rejects a manifest %s mismatch before checking archives', async mismatch => { + const identity = { task_id: 'task', user_id: 'user', repo: 'owner/repo', attempt_id: 'vm', request_id: 'request' }; + const bytes = Buffer.from(JSON.stringify({ + version: mismatch === 'version' ? 999 : 1, + identity: mismatch === 'identity' ? { ...identity, user_id: 'other' } : identity, + conversation: { + key: `continuations/task/vm/request/${'a'.repeat(64)}.json`, + sha256: 'a'.repeat(64), + size_bytes: 100, + version_id: 'conversation-v', + }, + workspace: { + key: `continuations/task/vm/request/workspace/${'b'.repeat(64)}.tar`, + sha256: 'b'.repeat(64), + size_bytes: 512, + version_id: 'workspace-v', + }, + })); + const record: ContinuationRecord = { + version: 1, + state: 'READY', + identity, + manifest: { + kind: 'manifest', + key: 'manifest', + version_id: 'version-1', + sha256: mismatch === 'checksum' ? '0'.repeat(64) : digest(bytes), + size_bytes: bytes.length + (mismatch === 'length' ? 1 : 0), + }, + }; + mockS3.mockResolvedValueOnce(object(bytes)); + await expect(verifyContinuationCheckpoint(record)).rejects.toThrow( + mismatch === 'checksum' || mismatch === 'length' + ? 'checkpoint manifest checksum does not match' + : 'checkpoint manifest identity does not match', + ); + expect(mockS3).toHaveBeenCalledTimes(1); +}); + +test('cleanup preserves pending data and removes every closed-task version before clearing its pointer', async () => { + const options = { abortSignal: AbortSignal.timeout(1000) }; + mockDdb.mockResolvedValueOnce({ Item: { user_id: 'user', status: 'AWAITING_APPROVAL' } }); + await deleteClosedTaskContinuations('task', 'user', options); + expect(mockS3).not.toHaveBeenCalled(); + mockDdb.mockResolvedValueOnce({ Item: { user_id: 'user', status: 'CANCELLED' } }).mockResolvedValue({}); + mockS3.mockResolvedValueOnce({ + Versions: [{ Key: 'continuations/task/object', VersionId: 'v1' }], + DeleteMarkers: [{ Key: 'continuations/task/object', VersionId: 'v2' }], + }).mockResolvedValueOnce({}).mockResolvedValueOnce({}); + await deleteClosedTaskContinuations('task', 'user', options); + expect(mockS3.mock.calls[1][0].input.Delete.Objects).toHaveLength(2); + expect(mockDdb.mock.calls.at(-1)![0].input.UpdateExpression).toContain('REMOVE continuation, continuation_launch'); +}); + +test('partial delete failure retains the cleanup marker for a later retry', async () => { + mockDdb.mockResolvedValue({ Item: { user_id: 'user', status: 'FAILED' } }); + mockS3.mockResolvedValueOnce({ Versions: [{ Key: 'continuations/task/object', VersionId: 'v1' }] }) + .mockResolvedValueOnce({ Errors: [{ Code: 'AccessDenied' }] }); + await expect(deleteClosedTaskContinuations('task', 'user', {})).rejects.toThrow('CLEANUP_FAILED'); + expect(mockDdb).toHaveBeenCalledTimes(1); +}); diff --git a/cdk/test/handlers/shared/microvm-control.test.ts b/cdk/test/handlers/shared/microvm-control.test.ts new file mode 100644 index 000000000..0d234e8f5 --- /dev/null +++ b/cdk/test/handlers/shared/microvm-control.test.ts @@ -0,0 +1,48 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +// SPDX-License-Identifier: MIT-0 + +import { microvmErrorIdentity, microvmRequestIdentity } from '../../../src/handlers/shared/microvm-control'; + +test.each([undefined, null, {}, { $metadata: {} }, { $metadata: { requestId: 'bad\nsecret' } }])( + 'omits absent or malformed request metadata: %j', response => { + expect(microvmRequestIdentity(response)).toEqual({}); + }, +); + +test('keeps only the request ID from a successful reply', () => { + expect(microvmRequestIdentity({ + $metadata: { requestId: 'aws-123', headers: { authorization: 'secret' } }, payload: 'secret', + })).toEqual({ aws_request_id: 'aws-123' }); +}); + +test('uses the original SDK identity behind a wrapper without copying messages', () => { + expect(microvmErrorIdentity(Object.assign(new Error('private wrapper'), { + cause: Object.assign(new Error('private SDK detail'), { + name: 'AccessDeniedException', $metadata: { requestId: 'aws-456' }, + }), + }))).toEqual({ error_type: 'AccessDeniedException', aws_request_id: 'aws-456' }); +}); + +test('rejects malformed error names and request IDs', () => { + expect(microvmErrorIdentity({ + name: 'private\ncontent', $metadata: { requestId: 'secret'.repeat(30) }, + })).toEqual({ error_type: 'Error' }); +}); diff --git a/cdk/test/handlers/shared/microvm-image-capability.test.ts b/cdk/test/handlers/shared/microvm-image-capability.test.ts new file mode 100644 index 000000000..68bc1dd7c --- /dev/null +++ b/cdk/test/handlers/shared/microvm-image-capability.test.ts @@ -0,0 +1,95 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +// SPDX-License-Identifier: MIT-0 + +import sharedConstants from '../../../../contracts/constants.json'; +import { + MICROVM_IMAGE_PROTOCOL_ENV, MICROVM_LIFECYCLE_PROTOCOL, readMicrovmImageMetadata, + supportsMicrovmLifecycle, verifyMicrovmImageLifecycle, +} from '../../../src/handlers/shared/microvm-image-capability'; + +const identity = { + imageArn: 'arn:aws:lambda:us-east-1:123456789012:microvm-image:test', + imageVersion: '3.0', +}; +type Version = Parameters[1]; +const version = (): Version => ({ + ...identity, + environmentVariables: { [MICROVM_IMAGE_PROTOCOL_ENV]: MICROVM_LIFECYCLE_PROTOCOL }, + hooks: { + port: sharedConstants.microvm_lifecycle.hook_port, + microvmImageHooks: { ready: 'ENABLED', validate: 'ENABLED' }, + microvmHooks: { + run: 'ENABLED', + terminate: 'ENABLED', + suspend: 'ENABLED', + resume: 'ENABLED', + suspendTimeoutInSeconds: sharedConstants.microvm_hook_budgets.lifecycle_hook_timeout_seconds, + resumeTimeoutInSeconds: sharedConstants.microvm_hook_budgets.lifecycle_hook_timeout_seconds, + }, + }, +}); + +test('verifies the exact launched version and its compatible hooks and marker', () => { + expect(verifyMicrovmImageLifecycle(identity, version())).toBe(true); + // Deactivation only blocks new launches; it does not rewrite an existing snapshot. + const inactive = { ...version(), status: 'INACTIVE' }; + expect(verifyMicrovmImageLifecycle(identity, inactive)).toBe(true); +}); + +test.each([ + ['different image', (v: Version) => { v.imageArn += '-other'; }], + ['different version', (v: Version) => { v.imageVersion = '4.0'; }], + ['no marker', (v: Version) => { v.environmentVariables = {}; }], + ['wrong protocol', (v: Version) => { v.environmentVariables![MICROVM_IMAGE_PROTOCOL_ENV] = '999'; }], + ['wrong port', (v: Version) => { v.hooks!.port = 8081; }], + ['no hooks', (v: Version) => { v.hooks = undefined; }], + ['no ready', (v: Version) => { v.hooks!.microvmImageHooks!.ready = 'DISABLED'; }], + ['no validate', (v: Version) => { v.hooks!.microvmImageHooks!.validate = 'DISABLED'; }], + ['no run', (v: Version) => { v.hooks!.microvmHooks!.run = 'DISABLED'; }], + ['no terminate', (v: Version) => { v.hooks!.microvmHooks!.terminate = 'DISABLED'; }], + ['no suspend', (v: Version) => { v.hooks!.microvmHooks!.suspend = 'DISABLED'; }], + ['no resume', (v: Version) => { v.hooks!.microvmHooks!.resume = 'DISABLED'; }], + ['short suspend', (v: Version) => { v.hooks!.microvmHooks!.suspendTimeoutInSeconds = 1; }], + ['short resume', (v: Version) => { v.hooks!.microvmHooks!.resumeTimeoutInSeconds = 1; }], + ['unknown suspend budget', (v: Version) => { v.hooks!.microvmHooks!.suspendTimeoutInSeconds = undefined; }], + ['fractional resume budget', (v: Version) => { v.hooks!.microvmHooks!.resumeTimeoutInSeconds = 30.5; }], +] as const)('rejects %s', (_label, mutate) => { + const changed = version(); + mutate(changed); + expect(verifyMicrovmImageLifecycle(identity, changed)).toBe(false); +}); + +test.each([undefined, null, [], {}, { lifecycleProtocol: '1' }, { ...identity, imageArn: 1 }, + { ...identity, imageVersion: ' 3.0' }, { ...identity, imageArn: 'a\nb' }])('malformed metadata never invents image capability: %j', value => { + expect(readMicrovmImageMetadata(value)).toEqual({}); + expect(supportsMicrovmLifecycle(value)).toBe(false); +}); + +test('legacy and unsupported metadata preserve identity without claiming lifecycle support', () => { + for (const lifecycleProtocol of [undefined, '999']) { + const metadata = { ...identity, lifecycleProtocol }; + expect(readMicrovmImageMetadata(metadata)).toEqual(identity); + expect(supportsMicrovmLifecycle(metadata)).toBe(false); + } + const capable = { ...identity, lifecycleProtocol: MICROVM_LIFECYCLE_PROTOCOL }; + expect(readMicrovmImageMetadata({ ...capable, untrustedExtra: 'drop-me' })).toEqual(capable); + expect(supportsMicrovmLifecycle(capable)).toBe(true); +}); diff --git a/cdk/test/handlers/shared/microvm-lifecycle-local.test.ts b/cdk/test/handlers/shared/microvm-lifecycle-local.test.ts new file mode 100644 index 000000000..c2940ba9e --- /dev/null +++ b/cdk/test/handlers/shared/microvm-lifecycle-local.test.ts @@ -0,0 +1,530 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +/** Opt-in DynamoDB Local: loopback endpoint and dummy credentials only. */ +import { randomUUID } from 'node:crypto'; +import { CreateTableCommand, DeleteTableCommand, DynamoDBClient } from '@aws-sdk/client-dynamodb'; +import { DeleteCommand, DynamoDBDocumentClient, GetCommand, PutCommand, ScanCommand, TransactWriteCommand, UpdateCommand } from '@aws-sdk/lib-dynamodb'; + +const endpoint = process.env.ABCA_DDB_LOCAL_ENDPOINT; +if (process.env.CI === 'true' && !endpoint) { + throw new Error('CI requires ABCA_DDB_LOCAL_ENDPOINT; lifecycle transaction tests must not skip'); +} +if (endpoint && (new URL(endpoint).hostname !== '127.0.0.1' || new URL(endpoint).protocol !== 'http:')) { + throw new Error('Lifecycle integration tests require an http://127.0.0.1 DynamoDB Local endpoint'); +} +const mockBeforeSend = jest.fn(); +const mockAfterSend = jest.fn(); +const mockClients: DynamoDBDocumentClient[] = []; +const mockDeleteContinuations = jest.fn(); +jest.mock('../../../src/handlers/shared/microvm-continuation-storage', () => ({ + ...jest.requireActual('../../../src/handlers/shared/microvm-continuation-storage'), + deleteClosedTaskContinuations: (...args: unknown[]) => mockDeleteContinuations(...args), +})); +jest.mock('../../../src/handlers/shared/microvm-suspend-config', () => ({ + readMicrovmSuspendEnabled: async () => true, +})); +jest.mock('../../../src/handlers/shared/ua', () => { + const actual = jest.requireActual('../../../src/handlers/shared/ua'); + return { + ...actual, + makeDocClient: () => { + const client = actual.makeDocClient({ + endpoint: process.env.ABCA_DDB_LOCAL_ENDPOINT ?? 'http://127.0.0.1:1', + region: 'us-east-1', + credentials: { accessKeyId: 'local', secretAccessKey: 'local' }, + }); + const send = client.send.bind(client); + client.send = async (command: unknown, options: unknown) => { + await mockBeforeSend(command); + const result = await send(command, options); + await mockAfterSend(command, result); + return result; + }; + mockClients.push(client); + return client; + }, + }; +}); +const suffix = randomUUID(); +const tasks = `lifecycle-tasks-${suffix}`; +const approvals = `lifecycle-approvals-${suffix}`; +Object.assign(process.env, { TASK_TABLE_NAME: tasks, TASK_APPROVALS_TABLE_NAME: approvals }); +import { reconcileMicrovmContinuation } from '../../../src/handlers/reconcile-microvm-continuations'; +import { failContinuationAttempt } from '../../../src/handlers/shared/microvm-continuation-runner'; +import { workerLeaseKey } from '../../../src/handlers/shared/microvm-continuation-types'; +import { readMicrovmLifecycleSnapshot, saveMicrovmLifecycleIntent } from '../../../src/handlers/shared/microvm-lifecycle'; +import { claimMicrovmStart, saveMicrovmImageCapability, saveMicrovmStartHandle } from '../../../src/handlers/shared/microvm-start'; +import { superviseMicrovm, type MicrovmSupervisorState } from '../../../src/handlers/shared/microvm-supervisor'; + +const raw = new DynamoDBClient({ + endpoint: endpoint ?? 'http://127.0.0.1:1', + region: 'us-east-1', + credentials: { accessKeyId: 'local', secretAccessKey: 'local' }, +}); +const admin = DynamoDBDocumentClient.from(raw); +const local = endpoint ? describe : describe.skip; +jest.setTimeout(30_000); + +local('MicroVM lifecycle against DynamoDB Local', () => { + beforeAll(async () => { + for (const name of [tasks, approvals]) { + await raw.send(new CreateTableCommand({ + TableName: name, + BillingMode: 'PAY_PER_REQUEST', + AttributeDefinitions: [{ AttributeName: 'task_id', AttributeType: 'S' }, + ...(name === approvals ? [{ AttributeName: 'request_id', AttributeType: 'S' as const }] : [])], + KeySchema: [{ AttributeName: 'task_id', KeyType: 'HASH' }, + ...(name === approvals ? [{ AttributeName: 'request_id', KeyType: 'RANGE' as const }] : [])], + })); + } + }); + beforeEach(async () => { + mockBeforeSend.mockReset(); + mockAfterSend.mockReset(); + mockDeleteContinuations.mockReset(); + for (const name of [tasks, approvals]) { + const result = await admin.send(new ScanCommand({ TableName: name })); + for (const item of result.Items ?? []) { + await admin.send(new DeleteCommand({ + TableName: name, + Key: { + task_id: item.task_id, + ...(name === approvals ? { request_id: item.request_id } : {}), + }, + })); + } + } + await admin.send(new PutCommand({ + TableName: tasks, + Item: { + task_id: 'task', + microvm_sleep_after_s: 30, + user_id: 'user', + status: 'AWAITING_APPROVAL', + compute_type: 'lambda-microvm', + session_id: 'vm', + compute_metadata: { + microvmId: 'vm', + endpoint: 'https://vm.example', + imageArn: 'arn:aws:lambda:us-east-1:123456789012:microvm-image:test', + imageVersion: '3.0', + lifecycleProtocol: '1', + }, + awaiting_approval_request_id: 'gate', + }, + })); + await approval('gate'); + }); + afterAll(async () => { + try { for (const name of [tasks, approvals]) await raw.send(new DeleteTableCommand({ TableName: name })); } finally { for (const client of mockClients) client.destroy(); raw.destroy(); } + }); + async function approval(requestId: string) { + await admin.send(new PutCommand({ + TableName: approvals, + Item: { + task_id: 'task', + request_id: requestId, + user_id: 'user', + status: 'PENDING', + created_at: new Date(Date.now() - 45_000).toISOString(), + timeout_s: 600, + }, + })); + } + async function current() { + const value = await readMicrovmLifecycleSnapshot('task', 'user'); + if (!value) throw new Error('Expected a MicroVM task'); + return value; + } + + test.each([false, true])('cleans an early failed replacement, preserving a concurrent worker ID (%s)', async lateWorker => { + const attempt = randomUUID(); + await admin.send(new PutCommand({ + TableName: tasks, + Item: { + task_id: 'task', + user_id: 'user', + status: 'AWAITING_APPROVAL', + compute_type: 'lambda-microvm', + continuation: { state: 'STARTING', attempt_id: attempt }, + concurrency_slot: { state: 'held', attempt_id: attempt }, + }, + })); + await admin.send(new PutCommand({ + TableName: tasks, + Item: { ...workerLeaseKey('task'), lease_user_id: 'user', lease_attempt_id: attempt, lease_state: 'ACTIVE' }, + })); + await failContinuationAttempt({ + task_id: 'task', continuation_request_id: 'gate', continuation_attempt_id: attempt, + }, 'user', 'MICROVM_CONTINUATION_VERSION_CHANGED'); + expect((await admin.send(new GetCommand({ TableName: tasks, Key: workerLeaseKey('task') }))).Item?.lease_state) + .toBe('FENCED'); + // Finalization can release a no-receipt attempt before the scheduled sweep. + await admin.send(new UpdateCommand({ + TableName: tasks, + Key: { task_id: 'task' }, + UpdateExpression: 'SET concurrency_slot.#state = :released', + ExpressionAttributeNames: { '#state': 'state' }, + ExpressionAttributeValues: { ':released': 'released' }, + })); + expect(await claimMicrovmStart('task', 'user', 'hash', attempt)).toMatchObject({ closed: true }); + if (lateWorker) { + mockBeforeSend.mockImplementation(async command => { + if (command instanceof UpdateCommand && command.input.UpdateExpression?.includes('lease_state = :closed')) { + await admin.send(new UpdateCommand({ + TableName: tasks, + Key: workerLeaseKey('task'), + UpdateExpression: 'SET lease_microvm_id = :id', + ExpressionAttributeValues: { ':id': 'late-worker' }, + })); + } + }); + } + const row = (await admin.send(new GetCommand({ TableName: tasks, Key: { task_id: 'task' } }))).Item!; + if (lateWorker) { + await expect(reconcileMicrovmContinuation(row as Parameters[0])) + .rejects.toThrow('The conditional request failed'); + expect(mockDeleteContinuations).not.toHaveBeenCalled(); + } else { + await reconcileMicrovmContinuation(row as Parameters[0]); + expect((await admin.send(new GetCommand({ TableName: tasks, Key: workerLeaseKey('task') }))).Item?.lease_state) + .toBe('CLOSED'); + expect(mockDeleteContinuations).toHaveBeenCalledWith('task', 'user', expect.any(Object)); + } + }); + + test('real supervisor and store preserve approval-during-suspend wake across serialized polls', async () => { + const handle = (await current()).handle; + let observed = 'RUNNING'; + const strategy = { + type: 'lambda-microvm' as const, + startSession: jest.fn(), + pollSession: jest.fn(async () => ({ + status: 'running' as const, + microvmState: observed as 'RUNNING' | 'SUSPENDING' | 'SUSPENDED', + microvmStartedAtMs: Date.now() - 60_000, + microvmMaximumDurationSeconds: 28_800, + })), + stopSession: jest.fn(), + suspendSession: jest.fn(async () => { + expect((await current()).intent?.action).toBe('suspend'); + await admin.send(new UpdateCommand({ + TableName: approvals, + Key: { task_id: 'task', request_id: 'gate' }, + UpdateExpression: 'SET #s = :s', + ExpressionAttributeNames: { '#s': 'status' }, + ExpressionAttributeValues: { ':s': 'APPROVED' }, + })); + return { supported: true as const }; + }), + resumeSession: jest.fn(async () => ({ supported: true as const })), + }; + const cycle = (previous?: MicrovmSupervisorState) => superviseMicrovm({ + taskId: 'task', + userId: 'user', + handle, + strategy, + suspendEnabled: true, + pollIntervalMs: 30_000, + previous: previous ? JSON.parse(JSON.stringify(previous)) : undefined, + }); + const first = await cycle(); + expect(first.kind).toBe('continue'); + const savedWake = (await current()).intent; + expect(savedWake?.action).toBe('resume'); + observed = 'SUSPENDING'; + const second = await cycle(first.state); + expect(strategy.resumeSession.mock.calls).toHaveLength(0); + observed = 'SUSPENDED'; + const third = await cycle(second.state); + expect(strategy.resumeSession.mock.calls).toHaveLength(1); + expect((await current()).intent).toEqual(savedWake); + await admin.send(new UpdateCommand({ + TableName: tasks, + Key: { task_id: 'task' }, + UpdateExpression: 'SET #s = :s, agent_heartbeat_at = :now REMOVE awaiting_approval_request_id', + ExpressionAttributeNames: { '#s': 'status' }, + ExpressionAttributeValues: { ':s': 'RUNNING', ':now': new Date().toISOString() }, + })); + observed = 'RUNNING'; + const restored = await cycle(third.state); + expect((await current()).intent).toMatchObject({ action: 'resume', request_id: null }); + expect((await cycle(restored.state)).state.recovery).toBeUndefined(); + expect(strategy.suspendSession.mock.calls).toHaveLength(1); + }); + + test('real pre-command read prevents Suspend when approval wins just after intent commit', async () => { + mockAfterSend.mockImplementation(async command => { + const intent = command instanceof TransactWriteCommand + ? command.input.TransactItems?.[0].Update?.ExpressionAttributeValues?.[':intent'] : undefined; + if (intent?.action === 'suspend') { + await admin.send(new UpdateCommand({ + TableName: approvals, + Key: { task_id: 'task', request_id: 'gate' }, + UpdateExpression: 'SET #s = :s', + ExpressionAttributeNames: { '#s': 'status' }, + ExpressionAttributeValues: { ':s': 'APPROVED' }, + })); + } + }); + const strategy = { + type: 'lambda-microvm' as const, + startSession: jest.fn(), + stopSession: jest.fn(), + pollSession: jest.fn(async () => ({ + status: 'running' as const, + microvmState: 'RUNNING' as const, + microvmStartedAtMs: Date.now() - 60_000, + microvmMaximumDurationSeconds: 28_800, + })), + suspendSession: jest.fn(), + resumeSession: jest.fn(), + }; + const result = await superviseMicrovm({ + taskId: 'task', + userId: 'user', + handle: (await current()).handle, + strategy, + suspendEnabled: true, + pollIntervalMs: 30_000, + }); + expect(result.kind).toBe('continue'); + expect(strategy.suspendSession.mock.calls).toHaveLength(0); + expect(strategy.resumeSession.mock.calls).toHaveLength(0); + expect(await current()).toMatchObject({ + status: 'AWAITING_APPROVAL', approval: { status: 'APPROVED' }, intent: { action: 'resume' }, + }); + }); + async function taskRow() { + return (await admin.send(new GetCommand({ TableName: tasks, Key: { task_id: 'task' }, ConsistentRead: true }))).Item!; + } + async function set(table: string, field: string, value: unknown) { + await admin.send(new UpdateCommand({ + TableName: table, + Key: { task_id: 'task', ...(table === approvals && { request_id: 'gate' }) }, + UpdateExpression: 'SET #field = :value', + ExpressionAttributeNames: { '#field': field }, + ExpressionAttributeValues: { ':value': value }, + })); + } + + test('suspend atomically records only coordinator intent and original deadline', async () => { + const observed = await current(); + const result = await saveMicrovmLifecycleIntent(observed, 'suspend'); + expect(result.status).toBe('saved'); + const saved = await taskRow(); + expect(saved.status).toBe('AWAITING_APPROVAL'); + expect(saved.compute_metadata).toEqual({ + microvmId: 'vm', + endpoint: 'https://vm.example', + imageArn: 'arn:aws:lambda:us-east-1:123456789012:microvm-image:test', + imageVersion: '3.0', + lifecycleProtocol: '1', + }); + expect(saved.microvm_lifecycle).toMatchObject({ + action: 'suspend', + request_id: 'gate', + microvm_id: 'vm', + deadline_ms: observed.approval.kind === 'present' ? observed.approval.deadlineMs : NaN, + }); + }); + test.each([ + ['imageArn', 'arn:aws:lambda:us-east-1:123456789012:microvm-image:replacement'], + ['imageVersion', '4.0'], ['lifecycleProtocol', '999'], ['lifecycleProtocol', undefined], + ])('changed image %s fences an already planned suspend', async (field, value) => { + const old = await current(); + const metadata = { ...(await taskRow()).compute_metadata, [field]: value }; + if (value === undefined) delete metadata[field]; + await set(tasks, 'compute_metadata', metadata); + expect(await saveMicrovmLifecycleIntent(old, 'suspend')).toEqual({ status: 'stale' }); + expect((await taskRow()).microvm_lifecycle).toBeUndefined(); + }); + test('legacy image cannot sleep but can be recovered with a wake', async () => { + await set(tasks, 'compute_metadata', { microvmId: 'vm', endpoint: 'https://vm.example' }); + const legacy = await current(); + expect(await saveMicrovmLifecycleIntent(legacy, 'suspend')).toEqual({ status: 'ineligible' }); + expect((await saveMicrovmLifecycleIntent(legacy, 'resume')).status).toBe('saved'); + }); + describe('image capability enrichment', () => { + const handle = { + strategyType: 'lambda-microvm' as const, + sessionId: 'vm', + microvmId: 'vm', + endpoint: 'https://vm.example', + imageArn: 'arn:aws:lambda:us-east-1:123456789012:microvm-image:test', + imageVersion: '3.0', + }; + const capable = { ...handle, lifecycleProtocol: '1' }; + beforeEach(async () => { + await set(tasks, 'status', 'HYDRATING'); + await set(tasks, 'microvm_start', { + clientToken: 'task', requestHash: 'original', expiresAt: Date.now() + 120_000, + }); + await saveMicrovmStartHandle('task', 'task', handle); + }); + test('persists support in both records without changing a concurrent terminal state', async () => { + await set(tasks, 'status', 'CANCELLED'); + await saveMicrovmImageCapability('task', 'task', capable); + const saved = await taskRow(); + expect(saved.status).toBe('CANCELLED'); + expect(saved.microvm_start.handle).toEqual(capable); + expect(saved.compute_metadata).toEqual({ + microvmId: handle.microvmId, + endpoint: handle.endpoint, + imageArn: handle.imageArn, + imageVersion: handle.imageVersion, + lifecycleProtocol: '1', + }); + expect(await claimMicrovmStart('task', 'user', 'changed-request')).toEqual({ + clientToken: 'task', closed: true, handle: capable, + }); + }); + test.each([ + 'session_id', 'microvm_start.clientToken', 'microvm_start.handle.microvmId', + 'microvm_start.handle.imageArn', 'microvm_start.handle.imageVersion', + 'compute_metadata.microvmId', 'compute_metadata.imageArn', 'compute_metadata.imageVersion', + ])('changed %s rejects the entire capability update', async path => { + const changed = await taskRow(); + const fields = path.split('.'); + let parent = changed; + for (const field of fields.slice(0, -1)) parent = parent[field]; + parent[fields[fields.length - 1]] = 'replacement'; + await admin.send(new PutCommand({ TableName: tasks, Item: changed })); + await expect(saveMicrovmImageCapability('task', 'task', capable)) + .rejects.toMatchObject({ name: 'ConditionalCheckFailedException' }); + expect(await taskRow()).toEqual(changed); + }); + test('lost committed reply is recovered by the original start receipt', async () => { + mockAfterSend.mockImplementationOnce(() => { + throw Object.assign(new Error('lost reply'), { name: 'TimeoutError' }); + }); + await expect(saveMicrovmImageCapability('task', 'task', capable)).rejects.toThrow('lost reply'); + expect(await claimMicrovmStart('task', 'user', 'original')).toEqual({ + clientToken: 'task', closed: false, handle: capable, + }); + }); + test('incomplete capability rejects before a database mutation', async () => { + const saved = await taskRow(); + mockBeforeSend.mockClear(); + await expect(saveMicrovmImageCapability('task', 'task', handle)).rejects.toThrow('incomplete'); + expect(mockBeforeSend).not.toHaveBeenCalled(); + expect(await taskRow()).toEqual(saved); + }); + }); + test('a wake blocks an older absent-record sleep and remains sticky on fresh reads', async () => { + const old = await current(); + await saveMicrovmLifecycleIntent(await current(), 'resume'); + expect(await saveMicrovmLifecycleIntent(old, 'suspend')).toEqual({ status: 'stale' }); + expect(await saveMicrovmLifecycleIntent(await current(), 'suspend')).toEqual({ status: 'ineligible' }); + expect((await taskRow()).microvm_lifecycle.action).toBe('resume'); + }); + test('a new wake generation fences a previously recorded suspend', async () => { + await saveMicrovmLifecycleIntent(await current(), 'suspend'); + const oldSleep = await current(); + await saveMicrovmLifecycleIntent(await current(), 'resume'); + expect(await saveMicrovmLifecycleIntent(oldSleep, 'suspend')).toEqual({ status: 'stale' }); + expect((await taskRow()).microvm_lifecycle.generation).not.toBe(oldSleep.intent?.generation); + }); + test('a later gate may sleep without allowing a late writer from the earlier gate', async () => { + await saveMicrovmLifecycleIntent(await current(), 'resume'); + const oldGate = await current(); + await approval('gate-two'); + await set(tasks, 'awaiting_approval_request_id', 'gate-two'); + expect(await saveMicrovmLifecycleIntent(oldGate, 'resume')).toEqual({ status: 'stale' }); + expect((await saveMicrovmLifecycleIntent(await current(), 'suspend')).status).toBe('saved'); + expect((await taskRow()).microvm_lifecycle.request_id).toBe('gate-two'); + }); + test.each([ + ['status', 'CANCELLED'], ['user_id', 'other-user'], ['session_id', 'other-vm'], + ['compute_type', 'ecs'], ['awaiting_approval_request_id', 'different-gate'], + ['compute_metadata', { microvmId: 'vm', endpoint: 'https://replacement.example' }], + ])('task %s changing after the read rejects intent', async (field, value) => { + const old = await current(); + await set(tasks, field as string, value); + expect(await saveMicrovmLifecycleIntent(old, 'suspend')).toEqual({ status: 'stale' }); + expect((await taskRow()).microvm_lifecycle).toBeUndefined(); + }); + test.each([ + ['status', 'APPROVED'], ['status', 'DENIED'], ['status', 'TIMED_OUT'], + ['created_at', '2020-01-01T00:00:00Z'], ['timeout_s', 10], ['user_id', 'other-user'], + ])('approval %s changing after the read rolls back the entire suspend write', async (field, value) => { + const old = await current(); + await set(approvals, field as string, value); + expect(await saveMicrovmLifecycleIntent(old, 'suspend')).toEqual({ status: 'stale' }); + expect((await taskRow()).microvm_lifecycle).toBeUndefined(); + }); + test('missing gate rejects sleep but permits conservative wake', async () => { + const old = await current(); + await admin.send(new DeleteCommand({ TableName: approvals, Key: { task_id: 'task', request_id: 'gate' } })); + expect(await saveMicrovmLifecycleIntent(old, 'suspend')).toEqual({ status: 'stale' }); + expect((await saveMicrovmLifecycleIntent(await current(), 'resume')).status).toBe('saved'); + }); + test('lost committed reply recovers the exact generation from the database', async () => { + const observed = await current(); + mockAfterSend.mockImplementationOnce(async command => { + expect(command).toBeInstanceOf(TransactWriteCommand); + throw new Error('lost committed response'); + }); + const result = await saveMicrovmLifecycleIntent(observed, 'suspend'); + expect(result).toEqual({ status: 'saved', intent: (await taskRow()).microvm_lifecycle }); + }); + test.each(['cancel', 'approve'])('%s after a committed write but before lost-reply recovery is stale', async change => { + const observed = await current(); + mockAfterSend.mockImplementationOnce(async command => { + expect(command).toBeInstanceOf(TransactWriteCommand); + await set(change === 'cancel' ? tasks : approvals, 'status', change === 'cancel' ? 'CANCELLED' : 'APPROVED'); + throw new Error('lost committed response'); + }); + expect(await saveMicrovmLifecycleIntent(observed, 'suspend')).toEqual({ status: 'stale' }); + }); + test('fresh module/client observes the saved wake without resetting age or generation', async () => { + await saveMicrovmLifecycleIntent(await current(), 'resume'); + const saved = (await taskRow()).microvm_lifecycle; + let restarted: typeof import('../../../src/handlers/shared/microvm-lifecycle'); + jest.isolateModules(() => { restarted = jest.requireActual('../../../src/handlers/shared/microvm-lifecycle'); }); + const observed = await restarted!.readMicrovmLifecycleSnapshot('task', 'user'); + expect(observed?.intent).toEqual(saved); + expect(await restarted!.saveMicrovmLifecycleIntent(observed!, 'resume')).toEqual({ status: 'saved', intent: saved }); + expect((await taskRow()).microvm_lifecycle).toEqual(saved); + }); + test('competing sleepers converge; a later wake fences both old snapshots', async () => { + const first = await current(); + const second = await current(); + const results = await Promise.allSettled([saveMicrovmLifecycleIntent(first, 'suspend'), saveMicrovmLifecycleIntent(second, 'suspend')]); + expect(results.filter(result => result.status === 'fulfilled' && result.value.status === 'saved')).toHaveLength(1); + for (const result of results) { + if (result.status === 'rejected') expect(result.reason.name).toBe('TransactionCanceledException'); + } + await saveMicrovmLifecycleIntent(await current(), 'resume'); + expect(await saveMicrovmLifecycleIntent(first, 'suspend')).toEqual({ status: 'stale' }); + expect(await saveMicrovmLifecycleIntent(second, 'suspend')).toEqual({ status: 'stale' }); + }); + test('approval-read failure is explicit and cannot prevent recording a wake', async () => { + mockBeforeSend.mockImplementation(async command => { + if (command instanceof GetCommand && command.input.TableName === approvals) { + throw Object.assign(new Error('simulated authorization failure'), { name: 'AccessDeniedException' }); + } + }); + const observed = await current(); + expect(observed.approval).toEqual({ kind: 'unavailable', errorType: 'AccessDeniedException' }); + expect(await saveMicrovmLifecycleIntent(observed, 'suspend')).toEqual({ status: 'ineligible' }); + expect((await saveMicrovmLifecycleIntent(observed, 'resume')).status).toBe('saved'); + }); +}); diff --git a/cdk/test/handlers/shared/microvm-lifecycle.test.ts b/cdk/test/handlers/shared/microvm-lifecycle.test.ts new file mode 100644 index 000000000..1270ec43c --- /dev/null +++ b/cdk/test/handlers/shared/microvm-lifecycle.test.ts @@ -0,0 +1,338 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +const mockSend = jest.fn(); +jest.mock('../../../src/handlers/shared/ua', () => ({ makeDocClient: () => ({ send: mockSend }) })); +process.env.TASK_TABLE_NAME = 'LifecycleTasks'; +process.env.TASK_APPROVALS_TABLE_NAME = 'LifecycleApprovals'; + +import { GetCommand, TransactWriteCommand } from '@aws-sdk/lib-dynamodb'; +import type { MicrovmObservedState } from '../../../src/handlers/shared/compute-strategy'; +import { + readMicrovmLifecycleSnapshot, saveMicrovmLifecycleIntent, type MicrovmLifecycleSnapshot, + type MicrovmLifecycleIntent, +} from '../../../src/handlers/shared/microvm-lifecycle'; +import { decideMicrovmLifecycle, type MicrovmLifecyclePolicyInput } from '../../../src/handlers/shared/microvm-lifecycle-policy'; + +const NOW = 1_800_000_000_000; +const CREATED = new Date(NOW - 45_000).toISOString(); +const DEADLINE = NOW + 555_000; +const imageMetadata = { + imageArn: 'arn:aws:lambda:us-east-1:123456789012:microvm-image:test', + imageVersion: '3.0', + lifecycleProtocol: '1', +}; +const task = { + microvm_sleep_after_s: 30, + task_id: 'task', + user_id: 'user', + status: 'AWAITING_APPROVAL', + compute_type: 'lambda-microvm', + session_id: 'vm', + compute_metadata: { microvmId: 'vm', endpoint: 'https://vm.example', ...imageMetadata }, + awaiting_approval_request_id: 'gate', +}; +const row = { task_id: 'task', user_id: 'user', request_id: 'gate', status: 'PENDING', created_at: CREATED, timeout_s: 600 }; +const intent = (action: 'suspend' | 'resume' = 'suspend'): MicrovmLifecycleIntent => ({ + version: 1, + generation: 'generation-one', + microvm_id: 'vm', + request_id: 'gate', + action, + requested_at_ms: NOW - 10_000, + deadline_ms: DEADLINE, +}); +const snapshot = (overrides: Partial = {}): MicrovmLifecycleSnapshot => ({ + sleepAfterSeconds: 30, + taskId: 'task', + userId: 'user', + status: 'AWAITING_APPROVAL', + requestId: 'gate', + handle: { strategyType: 'lambda-microvm', sessionId: 'vm', microvmId: 'vm', endpoint: 'https://vm.example', ...imageMetadata }, + approval: { kind: 'present', status: 'PENDING', created_at: CREATED, timeout_s: 600, createdAtMs: NOW - 45_000, deadlineMs: DEADLINE }, + ...overrides, +}); +const policy = (state: MicrovmObservedState = 'RUNNING', change: Partial = {}) => decideMicrovmLifecycle({ + snapshot: snapshot(), + substrate: { status: 'running', microvmState: state }, + nowMs: NOW, + sessionDeadlineMs: NOW + 3_600_000, + pollIntervalMs: 30_000, + suspendEnabled: true, + ...change, +}); + +beforeEach(() => { + mockSend.mockReset(); + jest.spyOn(Date, 'now').mockReturnValue(NOW); +}); +afterEach(() => jest.restoreAllMocks()); + +describe('MicroVM lifecycle policy', () => { + test('default waits ten minutes from the original gate creation before sleeping', () => { + const current = snapshot({ + sleepAfterSeconds: undefined, + approval: { + kind: 'present', + status: 'PENDING', + created_at: new Date(NOW).toISOString(), + createdAtMs: NOW, + timeout_s: 1800, + deadlineMs: NOW + 1_800_000, + }, + }); + expect(policy('RUNNING', { snapshot: current, nowMs: NOW + 599_000 })) + .toMatchObject({ action: 'wait', reason: 'suspend-grace', nextPollInMs: 1000 }); + expect(policy('RUNNING', { snapshot: current, nowMs: NOW + 600_000 })).toMatchObject({ action: 'suspend' }); + }); + test('a default five-minute gate stays awake and retains its original deadline', () => { + const current = snapshot({ + sleepAfterSeconds: undefined, + approval: { + kind: 'present', + status: 'PENDING', + created_at: new Date(NOW).toISOString(), + createdAtMs: NOW, + timeout_s: 300, + deadlineMs: NOW + 300_000, + }, + }); + expect(policy('RUNNING', { snapshot: current, nowMs: NOW + 200_000 })) + .toMatchObject({ action: 'wait', reason: 'suspend-grace', nextPollInMs: 30_000 }); + expect(policy('RUNNING', { snapshot: current, nowMs: NOW + 240_000 })) + .toMatchObject({ action: 'resume', reason: 'wake-deadline' }); + expect(current.approval).toMatchObject({ deadlineMs: NOW + 300_000 }); + }); + test.each([0, -1, NaN, Infinity, 0.5, 3601, '30', null])('off or malformed delay %s keeps awake and repairs existing sleep', value => { + const current = snapshot({ sleepAfterSeconds: value as number }); + expect(policy('RUNNING', { snapshot: current })).toMatchObject({ action: 'wait', reason: 'task-sleep-disabled' }); + expect(policy('SUSPENDED', { snapshot: { ...current, intent: intent() } })) + .toMatchObject({ action: 'resume', requestReady: true, reason: 'task-sleep-disabled' }); + expect(policy('SUSPENDING', { snapshot: { ...current, intent: intent() } })) + .toMatchObject({ action: 'resume', requestReady: false, reason: 'task-sleep-disabled' }); + expect(policy('SUSPENDED', { snapshot: { ...current, status: 'CANCELLED' } })).toMatchObject({ action: 'terminate' }); + }); + test('long pending gate may suspend only after grace and explicit RUNNING', () => { + expect(policy()).toEqual({ action: 'suspend', requestReady: true, reason: 'pending-long-gate', nextPollInMs: 5_000 }); + expect(mockSend).not.toHaveBeenCalled(); + }); + test.each(['PENDING', 'UNKNOWN'] as const)('%s is not proof of awake state', state => { + expect(policy(state)).toMatchObject({ action: 'wait', nextPollInMs: 5_000 }); + }); + test('legacy coarse running without a service observation cannot trigger suspend', () => { + expect(policy('RUNNING', { substrate: { status: 'running' } })).toMatchObject({ action: 'wait', reason: 'unconfirmed-state' }); + }); + test.each([ + { suspendEnabled: false }, + { snapshot: snapshot({ handle: { ...snapshot().handle, lifecycleProtocol: undefined } }) }, + ])('requires both enable and the actual worker capability: %j', change => { + expect(policy('RUNNING', change)).toMatchObject({ action: 'wait', reason: 'suspend-disabled' }); + }); + test('long poll setting is clamped to the end of grace', () => { + const gate = snapshot(); + expect(policy('RUNNING', { nowMs: NOW - 20_000, snapshot: gate, pollIntervalMs: 600_000 })) + .toMatchObject({ action: 'wait', reason: 'suspend-grace', nextPollInMs: 5_000 }); + }); + test('too little useful sleep keeps the VM awake', () => { + expect(policy('RUNNING', { nowMs: DEADLINE - 80_000 })) + .toMatchObject({ action: 'wait', reason: 'short-window', nextPollInMs: 20_000 }); + expect(policy('RUNNING', { nowMs: DEADLINE - 90_000 })).toMatchObject({ action: 'suspend' }); + }); + test('intended sleep waits until wake margin even with an oversized poll interval', () => { + expect(policy('SUSPENDED', { snapshot: snapshot({ intent: intent() }), pollIntervalMs: 900_000, suspendEnabled: false })) + .toEqual({ action: 'wait', reason: 'intentionally-suspended', nextPollInMs: DEADLINE - NOW - 60_000 }); + }); + test.each(['APPROVED', 'DENIED', 'TIMED_OUT', 'STRANDED'] as const)('%s always preserves wake intent', status => { + const pending = snapshot().approval; + if (pending.kind !== 'present') throw new Error('fixture'); + for (const state of ['RUNNING', 'SUSPENDING', 'SUSPENDED'] as const) { + expect(policy(state, { snapshot: snapshot({ approval: { ...pending, status }, intent: intent() }), suspendEnabled: false })) + .toMatchObject({ action: 'resume', requestReady: state === 'SUSPENDED', reason: 'approval-terminal' }); + } + }); + test.each([DEADLINE - 60_000, DEADLINE + 1000])('wakes at/past the original deadline margin (%s)', nowMs => { + expect(policy('SUSPENDED', { nowMs, snapshot: snapshot({ intent: intent() }) })) + .toMatchObject({ action: 'resume', requestReady: true, reason: 'wake-deadline' }); + }); + test('session lifetime also bounds the wake margin', () => { + expect(policy('SUSPENDED', { sessionDeadlineMs: NOW + 30_000, snapshot: snapshot({ intent: intent() }) })) + .toMatchObject({ action: 'resume', reason: 'wake-deadline' }); + expect(policy('SUSPENDED', { sessionDeadlineMs: NOW })).toMatchObject({ action: 'terminate', reason: 'session-deadline' }); + }); + test.each(['SUSPENDING', 'SUSPENDED'] as const)('unintended %s is repaired, but resume waits until SUSPENDED', state => { + expect(policy(state)).toMatchObject({ action: 'resume', requestReady: state === 'SUSPENDED', reason: 'unintended-suspension' }); + }); + test('acknowledged wake remains sticky even when a delayed suspend finishes later', () => { + const current = snapshot({ intent: intent('resume') }); + expect(policy('RUNNING', { snapshot: current })).toMatchObject({ action: 'wait', reason: 'wake-intent' }); + expect(policy('SUSPENDING', { snapshot: current })).toMatchObject({ action: 'resume', requestReady: false }); + expect(policy('SUSPENDED', { snapshot: current })).toMatchObject({ action: 'resume', requestReady: true }); + }); + test('another gate or changed deadline cannot inherit an old sleep intent', () => { + expect(policy('RUNNING', { snapshot: snapshot({ intent: { ...intent(), request_id: 'old-gate' } }) })) + .toMatchObject({ action: 'resume', reason: 'previous-gate-suspend' }); + expect(policy('SUSPENDED', { snapshot: snapshot({ intent: { ...intent(), deadline_ms: DEADLINE + 1 } }) })) + .toMatchObject({ action: 'resume', reason: 'approval-deadline-changed' }); + }); + test.each(['missing', 'invalid', 'unavailable'] as const)('%s approval forbids sleep and wakes a sleeping VM', kind => { + const current = snapshot({ approval: kind === 'unavailable' ? { kind, errorType: 'AccessDeniedException' } : { kind } }); + expect(policy('RUNNING', { snapshot: current })).toMatchObject({ action: 'wait' }); + expect(policy('SUSPENDED', { snapshot: current })).toMatchObject({ action: 'resume', requestReady: true }); + expect(policy('RUNNING', { snapshot: { ...current, intent: intent() } })).toMatchObject({ action: 'resume', requestReady: false }); + }); + test.each(['COMPLETED', 'FAILED', 'CANCELLED', 'TIMED_OUT', 'FINALIZING'] as const)('%s tasks are never revived', status => { + expect(policy('SUSPENDED', { snapshot: snapshot({ status }) })).toMatchObject({ action: 'terminate' }); + }); + test.each(['TERMINATING', 'TERMINATED', 'NOT_FOUND'] as const)('%s triggers task reconciliation', state => { + expect(policy(state)).toMatchObject({ action: 'reconcile-terminal' }); + expect(policy(state, { snapshot: snapshot({ status: 'COMPLETED' }) })).toMatchObject({ action: 'wait' }); + }); + test('a working task is never suspended for lack of traffic; unexpected sleep recovers', () => { + const current = snapshot({ status: 'RUNNING', requestId: null, approval: { kind: 'none' } }); + expect(policy('RUNNING', { snapshot: current })).toMatchObject({ action: 'wait', reason: 'working' }); + expect(policy('SUSPENDED', { snapshot: current })).toMatchObject({ action: 'resume', reason: 'suspended-outside-gate' }); + }); + test.each([0, -1, NaN, Infinity])('rejects invalid poll interval %s', pollIntervalMs => { + expect(() => policy('RUNNING', { pollIntervalMs })).toThrow('positive poll interval'); + }); +}); + +describe('MicroVM lifecycle store', () => { + test('an expired caller budget prevents reads, writes and lost-reply recovery', async () => { + const controller = new AbortController(); + controller.abort(new Error('caller deadline')); + const options = { abortSignal: controller.signal }; + await expect(readMicrovmLifecycleSnapshot('task', 'user', options)).rejects.toThrow('caller deadline'); + await expect(saveMicrovmLifecycleIntent(snapshot(), 'resume', NOW, options)).rejects.toThrow('caller deadline'); + expect(mockSend).not.toHaveBeenCalled(); + }); + test('a task read finishing after the caller budget cannot start the approval read', async () => { + const controller = new AbortController(); + mockSend.mockImplementationOnce(async () => { + controller.abort(new Error('caller deadline')); + return { Item: task }; + }); + await expect(readMicrovmLifecycleSnapshot('task', 'user', { abortSignal: controller.signal })) + .rejects.toThrow('caller deadline'); + expect(mockSend).toHaveBeenCalledTimes(1); + }); + test('a write finishing after the caller budget cannot start another recovery budget', async () => { + const controller = new AbortController(); + mockSend.mockImplementationOnce(async () => { + controller.abort(new Error('caller deadline')); + return {}; + }); + await expect(saveMicrovmLifecycleIntent(snapshot(), 'resume', NOW, { abortSignal: controller.signal })) + .rejects.toThrow('caller deadline'); + expect(mockSend).toHaveBeenCalledTimes(1); + }); + test('a cancelled task may retain its gate pointer and needs no approval read', async () => { + mockSend.mockResolvedValueOnce({ Item: { ...task, status: 'CANCELLED' } }); + const closed = await readMicrovmLifecycleSnapshot('task', 'user'); + expect(closed).toMatchObject({ status: 'CANCELLED', requestId: 'gate', approval: { kind: 'none' } }); + expect(await saveMicrovmLifecycleIntent(closed!, 'resume', NOW)).toEqual({ status: 'ineligible' }); + expect(mockSend).toHaveBeenCalledTimes(1); + }); + test('reads current task and only its current gate consistently with one bounded read budget', async () => { + mockSend.mockResolvedValueOnce({ Item: task }).mockResolvedValueOnce({ Item: row }); + expect(await readMicrovmLifecycleSnapshot('task', 'user')).toEqual(snapshot()); + expect(mockSend.mock.calls.map(([command]) => command.input)).toEqual([ + { TableName: 'LifecycleTasks', Key: { task_id: 'task' }, ConsistentRead: true }, + { TableName: 'LifecycleApprovals', Key: { task_id: 'task', request_id: 'gate' }, ConsistentRead: true }, + ]); + expect(mockSend.mock.calls[0][1].abortSignal).toBe(mockSend.mock.calls[1][1].abortSignal); + }); + test.each([undefined, { ...task, compute_type: 'ecs' }])('missing/non-MicroVM task is inapplicable', async item => { + mockSend.mockResolvedValueOnce({ Item: item }); + await expect(readMicrovmLifecycleSnapshot('task', 'user')).resolves.toBeUndefined(); + }); + test.each([ + { user_id: 'other' }, { session_id: 'other-vm' }, { compute_metadata: {} }, + { awaiting_approval_request_id: null }, { status: 'RUNNING' }, + { microvm_lifecycle: { ...intent(), version: 99 } }, + ])('invalid task/intent fails visibly: %j', async change => { + mockSend.mockResolvedValueOnce({ Item: { ...task, ...change } }); + await expect(readMicrovmLifecycleSnapshot('task', 'user')).rejects.toThrow('MicroVM lifecycle'); + expect(mockSend).toHaveBeenCalledTimes(1); + }); + test.each([ + { user_id: 'other' }, { request_id: 'old-gate' }, { status: 'MADE_UP' }, + { created_at: '2026-02-30T00:00:00Z' }, { created_at: 'not-a-date' }, + { timeout_s: -1 }, { timeout_s: Number.MAX_SAFE_INTEGER }, + ])('bad approval data remains explicitly invalid: %j', async change => { + mockSend.mockResolvedValueOnce({ Item: task }).mockResolvedValueOnce({ Item: { ...row, ...change } }); + expect((await readMicrovmLifecycleSnapshot('task', 'user'))?.approval).toEqual({ kind: 'invalid' }); + }); + test('missing approval and failed approval reads remain distinguishable', async () => { + mockSend.mockResolvedValueOnce({ Item: task }).mockResolvedValueOnce({}); + expect((await readMicrovmLifecycleSnapshot('task', 'user'))?.approval).toEqual({ kind: 'missing' }); + mockSend.mockResolvedValueOnce({ Item: task }).mockRejectedValueOnce(Object.assign(new Error('private data'), { name: 'AccessDeniedException' })); + expect((await readMicrovmLifecycleSnapshot('task', 'user'))?.approval).toEqual({ kind: 'unavailable', errorType: 'AccessDeniedException' }); + }); + test('suspend records intent with an atomic exact-gate condition', async () => { + mockSend.mockResolvedValueOnce({}); + const saved = await saveMicrovmLifecycleIntent(snapshot(), 'suspend', NOW); + expect(saved).toMatchObject({ status: 'saved', intent: { action: 'suspend', request_id: 'gate', deadline_ms: DEADLINE } }); + const command = mockSend.mock.calls[0][0]; + expect(command).toBeInstanceOf(TransactWriteCommand); + expect(command.input.TransactItems).toHaveLength(2); + expect(command.input.TransactItems[1].ConditionCheck.ExpressionAttributeValues) + .toEqual({ ':pending': 'PENDING', ':user': 'user', ':created': CREATED, ':timeout': 600 }); + expect(command.input.TransactItems[0].Update.UpdateExpression).toBe('SET microvm_lifecycle = :intent'); + }); + test('resume preserves its generation and original recovery age when replayed', async () => { + mockSend.mockResolvedValue({}); + const previous = intent('resume'); + expect(await saveMicrovmLifecycleIntent(snapshot({ intent: previous }), 'resume', NOW)).toEqual({ status: 'saved', intent: previous }); + const command = mockSend.mock.calls[0][0]; + expect(command.input.TransactItems).toHaveLength(1); + expect(command.input.ClientRequestToken).not.toBe(previous.generation); + expect(command.input.TransactItems[0].Update.ExpressionAttributeValues[':generation']).toBe(previous.generation); + }); + test('a wake cannot become sleep again within the same gate', async () => { + expect(await saveMicrovmLifecycleIntent(snapshot({ intent: intent('resume') }), 'suspend', NOW)).toEqual({ status: 'ineligible' }); + expect(mockSend).not.toHaveBeenCalled(); + }); + test('expired gate and closed task cannot receive suspend intent', async () => { + expect(await saveMicrovmLifecycleIntent(snapshot(), 'suspend', DEADLINE)).toEqual({ status: 'ineligible' }); + expect(await saveMicrovmLifecycleIntent(snapshot({ status: 'CANCELLED' }), 'resume', NOW)).toEqual({ status: 'ineligible' }); + expect(mockSend).not.toHaveBeenCalled(); + }); + test('conditional conflict asks caller to re-observe rather than overwrite', async () => { + mockSend.mockRejectedValueOnce({ name: 'TransactionCanceledException', CancellationReasons: [{ Code: 'ConditionalCheckFailed' }] }); + expect(await saveMicrovmLifecycleIntent(snapshot(), 'suspend', NOW)).toEqual({ status: 'stale' }); + }); + test('lost committed reply is recovered only by reading the exact saved generation', async () => { + let saved: MicrovmLifecycleIntent; + mockSend.mockImplementation(async command => { + if (command instanceof TransactWriteCommand) { + saved = command.input.TransactItems![0].Update!.ExpressionAttributeValues![':intent'] as MicrovmLifecycleIntent; + throw new Error('lost response'); + } + return { Item: command.input.TableName === 'LifecycleTasks' ? { ...task, microvm_lifecycle: saved! } : row }; + }); + expect(await saveMicrovmLifecycleIntent(snapshot(), 'suspend', NOW)).toMatchObject({ status: 'saved', intent: { generation: expect.any(String) } }); + expect(mockSend.mock.calls.filter(([command]) => command instanceof TransactWriteCommand)).toHaveLength(1); + expect(mockSend.mock.calls.filter(([command]) => command instanceof GetCommand)).toHaveLength(2); + }); + test('definite permission failure with no committed intent remains an error', async () => { + mockSend.mockRejectedValueOnce(new Error('access denied')).mockResolvedValueOnce({ Item: task }).mockResolvedValueOnce({ Item: row }); + await expect(saveMicrovmLifecycleIntent(snapshot(), 'suspend', NOW)).rejects.toThrow('access denied'); + }); +}); diff --git a/cdk/test/handlers/shared/microvm-start-recovery.test.ts b/cdk/test/handlers/shared/microvm-start-recovery.test.ts new file mode 100644 index 000000000..9f542c6ab --- /dev/null +++ b/cdk/test/handlers/shared/microvm-start-recovery.test.ts @@ -0,0 +1,320 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { TaskStatus } from '../../../src/constructs/task-status'; + +const mockDdbSend = jest.fn(); +const mockMicrovmSend = jest.fn(); +const mockS3Send = jest.fn(); +jest.mock('@aws-sdk/client-dynamodb', () => ({ DynamoDBClient: jest.fn(() => ({})) })); +jest.mock('@aws-sdk/lib-dynamodb', () => ({ + DynamoDBDocumentClient: { from: jest.fn(() => ({ send: mockDdbSend })) }, + GetCommand: jest.fn((input: unknown) => ({ kind: 'get', input })), + UpdateCommand: jest.fn((input: unknown) => ({ kind: 'update', input })), +})); +jest.mock('@aws-sdk/client-lambda-microvms', () => ({ + LambdaMicrovmsClient: jest.fn(() => ({ send: mockMicrovmSend })), + RunMicrovmCommand: jest.fn((input: unknown) => ({ kind: 'run', input })), + TerminateMicrovmCommand: jest.fn((input: unknown) => ({ kind: 'terminate', input })), + MicrovmState: {}, +})); +jest.mock('@aws-sdk/s3-request-presigner', () => ({ getSignedUrl: async () => 'https://payloads.s3.us-east-1.amazonaws.com/task/payload.json?X-Amz-Signature=' + Date.now() })); +const mockObjects = new Map(); +jest.mock('@aws-sdk/client-s3', () => ({ + DeleteObjectCommand: jest.fn((input: unknown) => ({ kind: 'delete', input })), + GetObjectCommand: jest.fn((input: unknown) => ({ kind: 'get', input })), + S3Client: jest.fn(() => ({ send: mockS3Send, config: { credentials: async () => ({ accessKeyId: 'EXAMPLE', secretAccessKey: 'unused' }) } })), + PutObjectCommand: jest.fn((input: unknown) => ({ kind: 'put', input })), +})); +jest.mock('../../../src/handlers/shared/logger', () => ({ + logger: { info: jest.fn(), warn: jest.fn(), error: jest.fn() }, +})); + +Object.assign(process.env, { + TASK_TABLE_NAME: 'tasks', + APPROVAL_REQUESTS_API_URL: 'https://approval.execute-api.us-east-1.amazonaws.com/v1', + TASK_EVENTS_TABLE_NAME: 'events', + MICROVM_IMAGE_IDENTIFIER: 'arn:aws:lambda:us-east-1:123456789012:microvm-image:test', + MICROVM_IMAGE_VERSION: '1', + MICROVM_EXECUTION_ROLE_ARN: 'arn:aws:iam::123456789012:role/execution', + MICROVM_EGRESS_CONNECTOR_ARNS: 'arn:aws:lambda:us-east-1:123456789012:network-connector:egress', + MICROVM_INGRESS_CONNECTOR_ARNS: 'arn:aws:lambda:us-east-1:aws:network-connector:aws-network-connector:NO_INGRESS', + MICROVM_PAYLOAD_BUCKET: 'payloads', + GITHUB_TOKEN_SECRET_ARN: 'arn:aws:secretsmanager:us-east-1:123456789012:secret:token', + AGENT_SESSION_ROLE_ARN: 'arn:aws:iam::123456789012:role/session', +}); + +import { MicrovmStartUncertainError } from '../../../src/handlers/shared/error-classifier'; +import { claimMicrovmStart, MICROVM_START_REPLAY_WINDOW_MS, microvmStartRequestHash } from '../../../src/handlers/shared/microvm-start'; +import { startSessionWithRetry } from '../../../src/handlers/shared/session-start-retry'; +import { LambdaMicrovmComputeStrategy } from '../../../src/handlers/shared/strategies/lambda-microvm-strategy'; + +const TASK_ID = '01K4YKCNV8P7WZBCSFDV2RNH49'; +const input = { + taskId: TASK_ID, + userId: 'user', + payload: { task_id: TASK_ID, prompt: 'x'.repeat(5_000) }, + blueprintConfig: { compute_type: 'lambda-microvm' as const, runtime_arn: '' }, +}; +const handle = { + strategyType: 'lambda-microvm' as const, + sessionId: 'mvm-one', + microvmId: 'mvm-one', + endpoint: 'https://example.invalid', +}; +let record: Record; +let lostResponses: number; +let created: Map>; +let now: number; + +beforeEach(() => { + jest.clearAllMocks(); + record = { user_id: 'user', status: TaskStatus.HYDRATING }; + lostResponses = 0; + created = new Map(); + now = Date.parse('2026-09-13T15:00:00Z'); + jest.spyOn(Date, 'now').mockImplementation(() => now); + mockObjects.clear(); + mockS3Send.mockReset().mockImplementation(async ({ kind: type, input: command }) => { + if (type === 'get') { + if (!mockObjects.has(command.Key)) throw Object.assign(new Error('missing'), { name: 'NoSuchKey' }); + return { Body: { transformToString: async () => mockObjects.get(command.Key) } }; + } + if (type === 'delete') { mockObjects.delete(command.Key); return {}; } + if (command.IfNoneMatch === '*' && mockObjects.has(command.Key)) throw Object.assign(new Error('exists'), { name: 'PreconditionFailed' }); + mockObjects.set(command.Key, command.Body); + return {}; + }); + mockDdbSend.mockReset().mockImplementation(async ({ kind, input: command }) => { + if (kind === 'get') return { Item: structuredClone(record) }; + const values = command.ExpressionAttributeValues; + if (values[':receipt']) { + if (record.microvm_start || record.status !== TaskStatus.HYDRATING) { + throw Object.assign(new Error('condition failed'), { name: 'ConditionalCheckFailedException' }); + } + record.microvm_start = structuredClone(values[':receipt']); + } else { + record.microvm_start.handle = values[':handle']; + record.session_id = values[':id']; + record.compute_type = values[':type']; + record.compute_metadata = values[':metadata']; + } + return {}; + }); + // Service emulator: same-token/same-request replay returns the original ID. + // This tests our use of the API contract; AWS retention still needs live proof. + mockMicrovmSend.mockReset().mockImplementation(async ({ kind, input: command }) => { + if (kind === 'terminate') return {}; + const token = command.clientToken ?? `sdk-generated-${created.size}`; + if (!created.has(token)) created.set(token, structuredClone(command)); + expect(created.get(token)).toEqual(command); + if (lostResponses-- > 0) { + throw Object.assign(new Error('response lost after creation'), { name: 'TimeoutError' }); + } + return { microvmId: handle.microvmId, endpoint: handle.endpoint, state: 'RUNNING' }; + }); +}); + +afterEach(() => jest.restoreAllMocks()); + +function runCalls() { + return mockMicrovmSend.mock.calls.filter(([command]) => command.kind === 'run'); +} + +test('a lost successful response retries one logical start and records the recovered handle', async () => { + lostResponses = 1; + const result = await startSessionWithRetry(new LambdaMicrovmComputeStrategy(), input, { + taskId: TASK_ID, emitRetryEvent: jest.fn(), logger: { warn: jest.fn() }, + }); + expect(result).toEqual({ handle, autoRetried: true }); + expect(runCalls()).toHaveLength(2); + expect(created.size).toBe(1); + expect(record.microvm_start.clientToken).toBe(TASK_ID); + expect(record.microvm_start.handle).toEqual(handle); + expect(record.compute_metadata).toEqual({ microvmId: handle.microvmId, endpoint: handle.endpoint }); + expect(mockDdbSend.mock.calls.filter(([command]) => command.kind === 'get') + .every(([command]) => command.input.ConsistentRead)).toBe(true); +}); + +test('a fresh strategy after a process restart reuses the persisted token', async () => { + lostResponses = 1; + await expect(new LambdaMicrovmComputeStrategy().startSession(input)).rejects.toThrow('response lost'); + const second = await new LambdaMicrovmComputeStrategy().startSession(input); + expect(second).toEqual(handle); + expect(created.size).toBe(1); + expect(runCalls()).toHaveLength(2); +}); + +test('a definite rejection after a lost response does not erase the first unknown outcome', async () => { + mockMicrovmSend + .mockRejectedValueOnce(Object.assign(new Error('response lost'), { name: 'TimeoutError' })) + .mockRejectedValueOnce(Object.assign(new Error('denied'), { name: 'AccessDeniedException' })); + await expect(startSessionWithRetry(new LambdaMicrovmComputeStrategy(), input, { + taskId: TASK_ID, emitRetryEvent: jest.fn(), logger: { warn: jest.fn() }, + })).rejects.toBeInstanceOf(MicrovmStartUncertainError); +}); + +test.each([ + { name: 'RequestTimeout', $metadata: { httpStatusCode: 400 } }, + { name: 'RequestTimeoutException', $metadata: { httpStatusCode: 400 } }, + { name: 'ServiceError', $metadata: { httpStatusCode: 408 } }, +])('a service timeout stays uncertain despite its HTTP status ($name)', async (timeout) => { + mockMicrovmSend + .mockRejectedValueOnce(Object.assign(new Error('start response timed out'), timeout)) + .mockRejectedValueOnce(Object.assign(new Error('denied'), { name: 'AccessDeniedException' })); + await expect(startSessionWithRetry(new LambdaMicrovmComputeStrategy(), input, { + taskId: TASK_ID, emitRetryEvent: jest.fn(), logger: { warn: jest.fn() }, + })).rejects.toBeInstanceOf(MicrovmStartUncertainError); + expect(runCalls().map(([command]) => command.input.clientToken)).toEqual([TASK_ID, TASK_ID]); +}); + +test('a confirmed rejection may retry the same token without creating an extra computer', async () => { + mockMicrovmSend.mockRejectedValueOnce(Object.assign(new Error('quota'), { name: 'ServiceQuotaExceededException' })); + const result = await startSessionWithRetry(new LambdaMicrovmComputeStrategy(), input, { + taskId: TASK_ID, emitRetryEvent: jest.fn(), logger: { warn: jest.fn() }, + }); + expect(result.handle).toEqual(handle); + expect(created.size).toBe(1); + expect(runCalls().map(([command]) => command.input.clientToken)).toEqual([TASK_ID, TASK_ID]); +}); + +test('different tasks receive different persisted tokens', async () => { + expect((await claimMicrovmStart(TASK_ID, 'user', 'hash')).clientToken).toBe(TASK_ID); + record = { user_id: 'user', status: TaskStatus.HYDRATING }; + expect((await claimMicrovmStart('another-task', 'user', 'hash')).clientToken).toBe('another-task'); +}); + +test('a task owner mismatch prevents all start side effects', async () => { + record.user_id = 'different-user'; + await expect(new LambdaMicrovmComputeStrategy().startSession(input)).rejects.toThrow('owner'); + expect(mockS3Send).not.toHaveBeenCalled(); + expect(mockMicrovmSend).not.toHaveBeenCalled(); +}); + +test('replay after saving a handle makes no further RunMicrovm or payload write', async () => { + await new LambdaMicrovmComputeStrategy().startSession(input); + now += MICROVM_START_REPLAY_WINDOW_MS * 10; + expect(await new LambdaMicrovmComputeStrategy().startSession(input)).toEqual(handle); + expect(runCalls()).toHaveLength(1); + expect(mockS3Send.mock.calls.filter(([c]) => c.kind === 'put' && c.input.Key.endsWith('/payload.json'))).toHaveLength(1); +}); + +test('changed input is refused before overwriting the first task payload', async () => { + lostResponses = 1; + await expect(new LambdaMicrovmComputeStrategy().startSession(input)).rejects.toThrow(); + await expect(new LambdaMicrovmComputeStrategy().startSession({ + ...input, payload: { ...input.payload, prompt: 'changed'.repeat(1_000) }, + })).rejects.toThrow('MICROVM_START_INPUT_CHANGED'); + expect(mockS3Send.mock.calls.filter(([c]) => c.kind === 'put' && c.input.Key.endsWith('/payload.json'))).toHaveLength(1); + expect(runCalls()).toHaveLength(1); +}); + +test('an expired unknown start never receives a new token or another RunMicrovm call', async () => { + lostResponses = 1; + await expect(new LambdaMicrovmComputeStrategy().startSession(input)).rejects.toThrow(); + now += MICROVM_START_REPLAY_WINDOW_MS; + await expect(new LambdaMicrovmComputeStrategy().startSession(input)) + .rejects.toThrow('MICROVM_START_OUTCOME_UNKNOWN'); + expect(runCalls()).toHaveLength(1); + expect(mockS3Send.mock.calls.filter(([c]) => c.kind === 'put' && c.input.Key.endsWith('/payload.json'))).toHaveLength(1); +}); + +test.each([TaskStatus.CANCELLED, TaskStatus.COMPLETED, TaskStatus.FAILED, TaskStatus.TIMED_OUT])( + '%s before starting never creates a MicroVM or writes a payload', async (status) => { + record.status = status; + await expect(new LambdaMicrovmComputeStrategy().startSession(input)).rejects.toThrow('MICROVM_START_TASK_CLOSED'); + expect(mockMicrovmSend).not.toHaveBeenCalled(); + expect(mockS3Send).not.toHaveBeenCalled(); + }, +); + +test('cancellation during payload upload prevents RunMicrovm', async () => { + mockS3Send.mockImplementationOnce(async () => { record.status = TaskStatus.CANCELLED; }); + await expect(new LambdaMicrovmComputeStrategy().startSession(input)).rejects.toThrow('MICROVM_START_TASK_CLOSED'); + expect(runCalls()).toHaveLength(0); +}); + +test('a cancelled task with a saved handle reaps that computer instead of starting another', async () => { + await new LambdaMicrovmComputeStrategy().startSession(input); + record.status = TaskStatus.CANCELLED; + await expect(new LambdaMicrovmComputeStrategy().startSession(input)).rejects.toThrow('MICROVM_START_TASK_CLOSED'); + expect(runCalls()).toHaveLength(1); + expect(mockMicrovmSend).toHaveBeenLastCalledWith({ + kind: 'terminate', input: { microvmIdentifier: handle.microvmId }, + }, { abortSignal: expect.any(AbortSignal) }); +}); + +test('failure to save a returned handle terminates the known computer', async () => { + const normal = mockDdbSend.getMockImplementation()!; + mockDdbSend.mockImplementation(async (command) => { + if (command.input.ExpressionAttributeValues?.[':handle']) throw new Error('DynamoDB unavailable'); + return normal(command); + }); + await expect(new LambdaMicrovmComputeStrategy().startSession(input)) + .rejects.toThrow('MICROVM_START_RECEIPT_SAVE_FAILED'); + expect(mockMicrovmSend).toHaveBeenLastCalledWith({ + kind: 'terminate', input: { microvmIdentifier: handle.microvmId }, + }, { abortSignal: expect.any(AbortSignal) }); +}); + +test('a lost receipt-write response recovers the committed handle without termination', async () => { + const normal = mockDdbSend.getMockImplementation()!; + mockDdbSend.mockImplementation(async (command) => { + const result = await normal(command); + if (command.input.ExpressionAttributeValues?.[':handle']) throw new Error('write response lost'); + return result; + }); + expect(await new LambdaMicrovmComputeStrategy().startSession(input)).toEqual(handle); + expect(mockMicrovmSend).toHaveBeenCalledTimes(1); + expect(record.microvm_start.handle).toEqual(handle); +}); + +test.each(['get', 'update'])( + 'a failed initial receipt %s prevents both payload upload and service creation', async (failedOperation) => { + const normal = mockDdbSend.getMockImplementation()!; + mockDdbSend.mockImplementation(async (command) => { + if (command.kind === failedOperation) throw new Error('DynamoDB unavailable'); + return normal(command); + }); + await expect(new LambdaMicrovmComputeStrategy().startSession(input)).rejects.toThrow('DynamoDB unavailable'); + expect(mockS3Send).not.toHaveBeenCalled(); + expect(mockMicrovmSend).not.toHaveBeenCalled(); + }, +); + +test('a competing receipt claim is re-read rather than overwritten', async () => { + const normal = mockDdbSend.getMockImplementation()!; + let raced = false; + mockDdbSend.mockImplementation(async (command) => { + if (command.input.ExpressionAttributeValues?.[':receipt'] && !raced) { + raced = true; + record.microvm_start = command.input.ExpressionAttributeValues[':receipt']; + } + return normal(command); + }); + expect(await claimMicrovmStart(TASK_ID, 'user', 'hash')).toEqual({ clientToken: TASK_ID, closed: false }); + expect(mockDdbSend.mock.calls.filter(([command]) => command.kind === 'update')).toHaveLength(1); +}); + +test('request fingerprints cover platform settings and full S3 content', () => { + const first = microvmStartRequestHash({ image: 'a', pointer: 's3://b/k' }, { prompt: 'one', repo: 'r' }); + expect(microvmStartRequestHash({ pointer: 's3://b/k', image: 'a' }, { repo: 'r', prompt: 'one' })).toBe(first); + expect(microvmStartRequestHash({ image: 'b', pointer: 's3://b/k' }, { prompt: 'one', repo: 'r' })).not.toBe(first); + expect(microvmStartRequestHash({ image: 'a', pointer: 's3://b/k' }, { prompt: 'two', repo: 'r' })).not.toBe(first); +}); diff --git a/cdk/test/handlers/shared/microvm-supervisor.test.ts b/cdk/test/handlers/shared/microvm-supervisor.test.ts new file mode 100644 index 000000000..51cb36ad6 --- /dev/null +++ b/cdk/test/handlers/shared/microvm-supervisor.test.ts @@ -0,0 +1,725 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +// SPDX-License-Identifier: MIT-0 + +import type { ComputeStrategy, SessionStatus } from '../../../src/handlers/shared/compute-strategy'; +import type { MicrovmLifecycleSnapshot } from '../../../src/handlers/shared/microvm-lifecycle'; + +const mockRead = jest.fn(); +const mockSave = jest.fn(); +const mockSuspendEnabled = jest.fn(); +jest.mock('../../../src/handlers/shared/microvm-suspend-config', () => ({ + readMicrovmSuspendEnabled: (...args: unknown[]) => mockSuspendEnabled(...args), +})); +const mockLogger = { warn: jest.fn(), info: jest.fn(), error: jest.fn() }; +jest.mock('../../../src/handlers/shared/microvm-lifecycle', () => ({ + ...jest.requireActual('../../../src/handlers/shared/microvm-lifecycle'), + readMicrovmLifecycleSnapshot: (...args: unknown[]) => mockRead(...args), + saveMicrovmLifecycleIntent: (...args: unknown[]) => mockSave(...args), +})); +jest.mock('../../../src/handlers/shared/logger', () => ({ logger: mockLogger })); + +import { + stopMicrovmWithDiagnostics, superviseMicrovm, MICROVM_RECOVERY_TIMEOUT_MS, + type MicrovmSupervisorInput, type MicrovmSupervisorState, +} from '../../../src/handlers/shared/microvm-supervisor'; + +const NOW = Date.parse('2026-09-15T10:00:00Z'); +const handle = { + strategyType: 'lambda-microvm' as const, + sessionId: 'vm', + microvmId: 'vm', + endpoint: 'https://vm.example', + imageArn: 'arn:aws:lambda:us-east-1:123456789012:microvm-image:agent', + imageVersion: '3.0', + lifecycleProtocol: '1', +}; +let time: number; +let row: MicrovmLifecycleSnapshot; +let strategy: { + type: ComputeStrategy['type']; + startSession: jest.Mock; + pollSession: jest.Mock; + stopSession: jest.Mock; + suspendSession: jest.Mock; + resumeSession: jest.Mock; +}; +let emitEvent: jest.Mock; +let generation: number; + +function observation(state: string = 'RUNNING'): SessionStatus { + return { + status: state === 'TERMINATED' ? 'completed' : state.startsWith('SUSPEND') ? 'suspended' : 'running', + microvmState: state as SessionStatus['microvmState'], + microvmStartedAtMs: NOW - 60_000, + microvmMaximumDurationSeconds: 28_800, + }; +} +function intent(action: 'suspend' | 'resume', requestedAt = NOW - 10_000) { + row = { + ...row, + intent: { + version: 1, + generation: `generation-${++generation}`, + microvm_id: 'vm', + request_id: row.requestId, + action, + requested_at_ms: requestedAt, + deadline_ms: row.approval.kind === 'present' ? row.approval.deadlineMs : null, + }, + }; +} +function working() { + row = { ...row, status: 'RUNNING', requestId: null, approval: { kind: 'none' } }; +} +function approve() { + if (row.approval.kind !== 'present') throw new Error('fixture has no approval'); + row = { ...row, approval: { ...row.approval, status: 'APPROVED' } }; +} +function run(previous?: MicrovmSupervisorState, change: Partial = {}) { + return superviseMicrovm({ + taskId: 'task', + userId: 'user', + handle, + strategy, + pollIntervalMs: 30_000, + suspendEnabled: true, + previous: previous ? JSON.parse(JSON.stringify(previous)) : undefined, + emitEvent, + ...change, + }); +} + +beforeEach(() => { + jest.clearAllMocks(); + time = NOW; + generation = 0; + jest.spyOn(Date, 'now').mockImplementation(() => time); + row = { + taskId: 'task', + sleepAfterSeconds: 30, + userId: 'user', + status: 'AWAITING_APPROVAL', + handle, + requestId: 'gate', + approval: { + kind: 'present', + status: 'PENDING', + created_at: new Date(NOW - 45_000).toISOString(), + timeout_s: 600, + createdAtMs: NOW - 45_000, + deadlineMs: NOW + 555_000, + }, + }; + mockRead.mockReset().mockImplementation(async () => structuredClone(row)); + mockSuspendEnabled.mockReset().mockResolvedValue(true); + mockSave.mockReset().mockImplementation(async (_snapshot, action) => { + if (row.intent?.action !== action || row.intent.request_id !== row.requestId) intent(action, time); + return { status: 'saved', intent: row.intent }; + }); + strategy = { + type: 'lambda-microvm', + startSession: jest.fn(), + pollSession: jest.fn().mockResolvedValue(observation()), + stopSession: jest.fn().mockResolvedValue(undefined), + suspendSession: jest.fn().mockResolvedValue({ supported: true }), + resumeSession: jest.fn().mockResolvedValue({ supported: true }), + }; + emitEvent = jest.fn().mockResolvedValue(undefined); +}); +afterEach(() => jest.restoreAllMocks()); + +test('logs changed observations once across durable replay without moving the wake deadline', async () => { + intent('resume', NOW - 10_000); + approve(); + strategy.pollSession.mockResolvedValue(observation('PENDING')); + const first = await run(); + expect(mockLogger.info).toHaveBeenCalledWith('MicroVM supervisor observation changed', expect.objectContaining({ + observed_state: 'PENDING', + task_id: 'task', + microvm_id: 'vm', + request_id: 'gate', + approval_state: 'APPROVED', + recovery_kind: 'wake', + recovery_since_ms: NOW - 10_000, + intent_requested_at_ms: NOW - 10_000, + image_version: '3.0', + })); + const signature = first.state.diagnosticSignature; + mockLogger.info.mockClear(); + time += 1_000; + const second = await run(first.state); + expect(second.state.diagnosticSignature).toBe(signature); + expect(mockLogger.info).not.toHaveBeenCalledWith('MicroVM supervisor observation changed', expect.anything()); + time += MICROVM_RECOVERY_TIMEOUT_MS; + const failed = await run(second.state); + expect(failed.kind).toBe('failure'); + expect(failed.state.recovery?.sinceMs).toBe(NOW - 10_000); + expect(mockLogger.warn).toHaveBeenCalledWith('MicroVM supervisor observation changed', expect.objectContaining({ + outcome: 'failure', + recovery_since_ms: NOW - 10_000, + observed_state: 'PENDING', + session_deadline_ms: first.state.sessionDeadlineMs, + })); +}); + +test('saves intent, rechecks the gate, requests suspend and rechecks the outcome', async () => { + const result = await run(); + expect(result.kind).toBe('continue'); + expect(result.state.recovery?.kind).toBe('suspend'); + expect(strategy.suspendSession).toHaveBeenCalledTimes(1); + expect(mockRead).toHaveBeenCalledTimes(3); + expect(mockSave.mock.invocationCallOrder[0]).toBeLessThan(mockRead.mock.invocationCallOrder[1]); + expect(mockRead.mock.invocationCallOrder[1]).toBeLessThan(strategy.suspendSession.mock.invocationCallOrder[0]); + expect(strategy.suspendSession.mock.invocationCallOrder[0]).toBeLessThan(mockRead.mock.invocationCallOrder[2]); + expect(strategy.stopSession).not.toHaveBeenCalled(); +}); + +test.each(['disabled', 'legacy', 'unverified-lifetime', 'short-window'])('%s keeps a working VM awake', async kind => { + if (kind === 'legacy') row = { ...row, handle: { ...handle, lifecycleProtocol: undefined } }; + if (kind === 'unverified-lifetime') strategy.pollSession.mockResolvedValue({ status: 'running', microvmState: 'RUNNING' }); + if (kind === 'short-window') time = NOW + 480_000; + expect((await run(undefined, { suspendEnabled: kind !== 'disabled' })).kind).toBe('continue'); + expect(strategy.suspendSession).not.toHaveBeenCalled(); + expect(strategy.resumeSession).not.toHaveBeenCalled(); + expect(mockSuspendEnabled).not.toHaveBeenCalled(); +}); + +test('per-task off overrides enabled deployment and records the reason', async () => { + row = { ...row, sleepAfterSeconds: 0 }; + expect((await run()).kind).toBe('continue'); + expect(strategy.suspendSession).not.toHaveBeenCalled(); + expect(mockLogger.info).toHaveBeenCalledWith('MicroVM supervisor observation changed', expect.objectContaining({ + sleep_after_s: 0, policy_reason: 'task-sleep-disabled', + })); +}); + +test('fresh task preference after intent commit prevents an in-flight suspend', async () => { + mockSave.mockImplementationOnce(async () => { + intent('suspend', time); + row = { ...row, sleepAfterSeconds: 0 }; + return { status: 'saved', intent: row.intent }; + }); + expect((await run()).kind).toBe('continue'); + expect(strategy.suspendSession).not.toHaveBeenCalled(); + expect(row.intent?.action).toBe('resume'); +}); + +test('an existing opt-in execution observes live disable without failing healthy compute', async () => { + working(); + const first = await run(); + row = { + ...row, + status: 'AWAITING_APPROVAL', + requestId: 'next-gate', + approval: { + kind: 'present', + status: 'PENDING', + created_at: new Date(NOW - 45_000).toISOString(), + timeout_s: 600, + createdAtMs: NOW - 45_000, + deadlineMs: NOW + 555_000, + }, + }; + mockSuspendEnabled.mockResolvedValue(false); + for (let attempt = 0; attempt < 4; attempt++) { + const result = await run(first.state); + expect(result.kind).toBe('continue'); + expect(result.state.consecutivePollFailures).toBe(0); + } + expect(strategy.suspendSession).not.toHaveBeenCalled(); + expect(mockSave).not.toHaveBeenCalled(); +}); + +test('disable after intent commit fences wake before any Suspend request', async () => { + mockSuspendEnabled.mockResolvedValueOnce(true).mockResolvedValue(false); + const result = await run(); + expect(result.state.recovery?.kind).toBe('wake'); + expect(row.intent?.action).toBe('resume'); + expect(strategy.suspendSession).not.toHaveBeenCalled(); + expect(mockSuspendEnabled).toHaveBeenCalledTimes(2); +}); + +test('wake recovery never depends on reading the suspension setting', async () => { + const first = await run(); + approve(); + strategy.pollSession.mockResolvedValue(observation('SUSPENDED')); + mockSuspendEnabled.mockClear().mockResolvedValue(false); + expect((await run(first.state)).kind).toBe('continue'); + expect(strategy.resumeSession).toHaveBeenCalledTimes(1); + expect(mockSuspendEnabled).not.toHaveBeenCalled(); +}); + +test('later polls and serialized restart cannot extend the original service deadline', async () => { + working(); + const first = await run(); + const deadline = NOW - 60_000 + 28_800_000; + expect(first.state.sessionDeadlineMs).toBe(deadline); + time += 60_000; + strategy.pollSession.mockResolvedValue({ ...observation(), microvmStartedAtMs: time, microvmMaximumDurationSeconds: 99_999 }); + const later = await run(first.state); + expect(later.state.sessionDeadlineMs).toBe(deadline); + time = deadline; + expect(await run(later.state)).toMatchObject({ kind: 'failure', reason: 'session-deadline' }); +}); + +test('unknown state has a fixed recovery window across serialized polls', async () => { + working(); + strategy.pollSession.mockResolvedValue(observation('UNKNOWN')); + const first = await run(); + time += MICROVM_RECOVERY_TIMEOUT_MS; + const expired = await run(first.state); + expect(expired).toMatchObject({ kind: 'failure', reason: 'recovery-deadline' }); + expect(expired.state.recovery?.sinceMs).toBe(first.state.recovery?.sinceMs); +}); + +test('a newly RUNNING task retains the full startup allowance while its first worker observation is PENDING', async () => { + working(); // The coordinator marks the task RUNNING before AWS reports readiness. + strategy.pollSession.mockResolvedValue(observation('PENDING')); + const first = await run(); + expect(first).toMatchObject({ + kind: 'continue', state: { recovery: { kind: 'starting', sinceMs: NOW } }, + }); + time += 120_000; + const later = await run(first.state); + expect(later.kind).toBe('continue'); + expect(later.state.recovery).toEqual(first.state.recovery); + time = NOW + 300_000; + expect((await run(later.state)).kind).toBe('failure'); +}); + +test('an initial read failure cannot shorten or restart the pending startup allowance', async () => { + working(); + mockRead.mockRejectedValueOnce(new Error('temporary read failure')); + const failed = await run(); + time += 60_000; + strategy.pollSession.mockResolvedValue(observation('PENDING')); + const pending = await run(failed.state); + expect(pending).toMatchObject({ + kind: 'continue', state: { recovery: { kind: 'starting', sinceMs: NOW } }, + }); + time = NOW + 300_000; + expect((await run(pending.state)).kind).toBe('failure'); +}); + +test.each(['current', 'legacy'])('%s saved state does not restart startup after a confirmed running worker', async version => { + working(); + const running = await run(); + const previous = { ...running.state }; + if (version === 'legacy') delete previous.startupConfirmed; + time += 60_000; + strategy.pollSession.mockResolvedValue(observation('PENDING')); + const pending = await run(previous); + expect(pending).toMatchObject({ + kind: 'continue', state: { recovery: { kind: 'unconfirmed', sinceMs: time } }, + }); + time += MICROVM_RECOVERY_TIMEOUT_MS; + expect((await run(pending.state)).kind).toBe('failure'); +}); + +test('repeated Get errors retain their count even when task reads succeed', async () => { + strategy.pollSession.mockRejectedValue(Object.assign(new Error('network'), { name: 'TimeoutError' })); + const first = await run(); + const second = await run(first.state); + const third = await run(second.state); + expect(first.kind).toBe('continue'); + expect(second.state.consecutivePollFailures).toBe(2); + expect(third).toMatchObject({ kind: 'failure', state: { consecutivePollFailures: 3 } }); +}); + +test('a complete successful observation clears prior poll failures', async () => { + working(); + strategy.pollSession.mockRejectedValueOnce(new Error('network')); + const failed = await run(); + expect(failed.state.consecutivePollFailures).toBe(1); + expect((await run(failed.state)).state.consecutivePollFailures).toBe(0); +}); + +test('permanent Get denial escalates without waiting for more requests', async () => { + strategy.pollSession.mockRejectedValue(Object.assign(new Error('private-detail'), { name: 'AccessDeniedException' })); + expect((await run()).kind).toBe('failure'); + expect(JSON.stringify(emitEvent.mock.calls)).not.toContain('private-detail'); +}); + +test.each(['missing', 'replaced'])('%s worker ownership is never followed to another VM', async kind => { + if (kind === 'missing') mockRead.mockResolvedValue(undefined); + else row = { ...row, handle: { ...handle, microvmId: 'other-vm', sessionId: 'other-vm' } }; + expect((await run()).kind).toBe('ownership-lost'); + expect(strategy.pollSession).not.toHaveBeenCalled(); + expect(strategy.suspendSession).not.toHaveBeenCalled(); +}); + +test('a cancelled task is returned for cleanup without any wake or status change', async () => { + row = { ...row, status: 'CANCELLED' }; + expect(await run()).toMatchObject({ kind: 'closed', status: 'CANCELLED' }); + expect(strategy.pollSession).not.toHaveBeenCalled(); + expect(mockSave).not.toHaveBeenCalled(); +}); + +test('terminal substrate observations are handed back for strong task reconciliation', async () => { + strategy.pollSession.mockResolvedValue(observation('TERMINATED')); + expect((await run()).kind).toBe('substrate-terminal'); + expect(mockSave).not.toHaveBeenCalled(); +}); + +test('an intended long sleep stays asleep and does not fail heartbeat liveness', async () => { + intent('suspend'); + strategy.pollSession.mockResolvedValue(observation('SUSPENDED')); + const result = await run(); + expect(result).toMatchObject({ kind: 'continue', deferHeartbeat: true }); + expect(result.state.recovery).toBeUndefined(); + expect(strategy.resumeSession).not.toHaveBeenCalled(); +}); + +test('a service stuck in SUSPENDING has a bounded transition window', async () => { + intent('suspend', NOW - MICROVM_RECOVERY_TIMEOUT_MS); + strategy.pollSession.mockResolvedValue(observation('SUSPENDING')); + expect((await run()).kind).toBe('failure'); +}); + +test('approval between intent save and the pre-command read prevents suspend', async () => { + mockRead.mockImplementationOnce(async () => structuredClone(row)) + .mockImplementationOnce(async () => { approve(); return structuredClone(row); }); + expect((await run()).kind).toBe('continue'); + expect(strategy.suspendSession).not.toHaveBeenCalled(); + expect(row.intent?.action).toBe('resume'); +}); + +test('approval during suspend saves sticky wake, then waits for SUSPENDED before Resume', async () => { + strategy.suspendSession.mockImplementationOnce(async () => { approve(); return { supported: true }; }); + const first = await run(); + expect(row.intent?.action).toBe('resume'); + expect(strategy.resumeSession).not.toHaveBeenCalled(); + strategy.pollSession.mockResolvedValue(observation('SUSPENDING')); + const suspending = await run(first.state); + expect(strategy.resumeSession).not.toHaveBeenCalled(); + strategy.pollSession.mockResolvedValue(observation('SUSPENDED')); + const waking = await run(suspending.state); + expect(strategy.resumeSession).toHaveBeenCalledTimes(1); + expect(waking.state.recovery?.kind).toBe('wake'); + strategy.pollSession.mockResolvedValue(observation()); + const awake = await run(waking.state); + expect(row.intent?.action).toBe('resume'); + expect(awake.state.recovery?.kind).toBe('wake'); + working(); + row = { ...row, heartbeatAtMs: time }; + const restored = await run(awake.state); + expect(row.intent).toMatchObject({ action: 'resume', request_id: null }); + expect((await run(restored.state)).state.recovery).toBeUndefined(); +}); + +test('cancellation during a command cannot trigger a post-command resume', async () => { + strategy.suspendSession.mockImplementationOnce(async () => { + row = { ...row, status: 'CANCELLED' }; + return { supported: true }; + }); + expect(await run()).toMatchObject({ kind: 'closed', status: 'CANCELLED' }); + expect(strategy.resumeSession).not.toHaveBeenCalled(); +}); + +test('an uncertain suspend failure leaves coding recoverable and fences late suspension', async () => { + strategy.suspendSession.mockRejectedValueOnce(Object.assign(new Error('private-detail'), { name: 'TimeoutError' })); + const result = await run(); + expect(result.kind).toBe('continue'); + expect(row.intent?.action).toBe('resume'); + expect(result.state.recovery?.kind).toBe('wake'); + expect(JSON.stringify(emitEvent.mock.calls)).not.toContain('private-detail'); +}); + +test('repeated resume failures escalate despite successful state reads', async () => { + approve(); + intent('suspend'); + strategy.pollSession.mockResolvedValue(observation('SUSPENDED')); + strategy.resumeSession.mockRejectedValue(Object.assign(new Error('retry'), { name: 'TimeoutError' })); + const first = await run(); + const second = await run(first.state); + expect(await run(second.state)).toMatchObject({ + kind: 'failure', reason: 'resume-request-failed-repeatedly', state: { consecutiveResumeFailures: 3 }, + }); +}); + +test.each(['PENDING', 'UNKNOWN'])('%s cannot reset a wake recovery clock', async observed => { + approve(); + intent('suspend'); + strategy.pollSession.mockResolvedValue(observation('SUSPENDED')); + const first = await run(); + time += 60_000; + strategy.pollSession.mockResolvedValue(observation(observed)); + const unknown = await run(first.state); + expect(unknown.state.recovery).toEqual(first.state.recovery); + time += 60_000; + strategy.pollSession.mockResolvedValue(observation('SUSPENDED')); + expect((await run(unknown.state)).kind).toBe('failure'); + expect(strategy.resumeSession).toHaveBeenCalledTimes(1); +}); + +test('PENDING after an API wake uses the wake intent instead of an expired startup clock', async () => { + intent('suspend'); + strategy.pollSession.mockResolvedValue(observation('SUSPENDED')); + const asleep = await run(); + expect(asleep.state.recovery).toBeUndefined(); + + time += 360_000; + approve(); + intent('resume', time); + const wakeStartedAt = time; + strategy.pollSession.mockResolvedValue(observation('PENDING')); + const pending = await run(asleep.state); + expect(pending).toMatchObject({ + kind: 'continue', + deferHeartbeat: true, + state: { recovery: { kind: 'wake', sinceMs: wakeStartedAt } }, + }); + expect(pending.state.firstObservedAtMs).toBe(asleep.state.firstObservedAtMs); + expect(pending.state.sessionDeadlineMs).toBe(asleep.state.sessionDeadlineMs); + expect(strategy.resumeSession).not.toHaveBeenCalled(); + + time += 10_000; + const replayed = await run(pending.state); + expect(replayed.state.recovery).toEqual(pending.state.recovery); + strategy.pollSession.mockResolvedValue(observation('RUNNING')); + const consuming = await run(replayed.state); + expect(consuming.state.recovery).toEqual(pending.state.recovery); + working(); + row = { ...row, heartbeatAtMs: time }; + const rebound = await run(consuming.state); + expect(rebound.state.recovery).toEqual(pending.state.recovery); + expect(row.intent?.request_id).toBeNull(); + expect((await run(rebound.state)).state.recovery).toBeUndefined(); +}); + +test.each(['PENDING', 'UNKNOWN'])('%s does not extend an already-expired API wake intent', async observed => { + approve(); + intent('resume', NOW - MICROVM_RECOVERY_TIMEOUT_MS); + strategy.pollSession.mockResolvedValue(observation(observed)); + expect(await run()).toMatchObject({ + kind: 'failure', + state: { recovery: { kind: 'wake', sinceMs: NOW - MICROVM_RECOVERY_TIMEOUT_MS } }, + }); + expect(strategy.resumeSession).not.toHaveBeenCalled(); +}); + +test('unexpected PENDING after coding starts receives one bounded recovery window', async () => { + working(); + row = { ...row, heartbeatAtMs: time }; + const running = await run(); + time += 360_000; + strategy.pollSession.mockResolvedValue(observation('PENDING')); + const pending = await run(running.state); + expect(pending).toMatchObject({ + kind: 'continue', + state: { recovery: { kind: 'unconfirmed', sinceMs: time } }, + }); + time += MICROVM_RECOVERY_TIMEOUT_MS; + expect((await run(pending.state)).kind).toBe('failure'); +}); + +test('durable wake recovery repairs a missing wake write even after a RUNNING observation', async () => { + intent('suspend'); + const initial = await run(undefined, { suspendEnabled: false }); + const previous = { ...initial.state, recovery: { kind: 'wake' as const, sinceMs: NOW } }; + const result = await run(previous); + expect(result.kind).toBe('continue'); + expect(row.intent?.action).toBe('resume'); + expect(strategy.suspendSession).not.toHaveBeenCalled(); +}); + +test('recovered RUNNING gets heartbeat grace only until a fresh heartbeat', async () => { + working(); + row = { ...row, heartbeatAtMs: NOW - 300_000 }; + strategy.pollSession.mockResolvedValue(observation('SUSPENDED')); + const waking = await run(); + strategy.pollSession.mockResolvedValue(observation()); + time += 5_000; + const waiting = await run(waking.state); + expect(waiting.deferHeartbeat).toBe(true); + time += 45_000; + row = { ...row, heartbeatAtMs: time }; + const healthy = await run(waiting.state); + expect(healthy.deferHeartbeat).toBe(false); + expect(healthy.state.recovery).toBeUndefined(); + time += 300_000; + expect((await run(healthy.state)).deferHeartbeat).toBe(false); +}); + +test('a resumed RUNNING worker with no fresh heartbeat cannot retain grace forever', async () => { + working(); + row = { ...row, heartbeatAtMs: NOW - 300_000 }; + strategy.pollSession.mockResolvedValue(observation('SUSPENDED')); + const waking = await run(); + strategy.pollSession.mockResolvedValue(observation()); + time += MICROVM_RECOVERY_TIMEOUT_MS; + expect((await run(waking.state)).kind).toBe('failure'); +}); + +test('intent storage failures remain bounded when all observations succeed', async () => { + approve(); + strategy.pollSession.mockResolvedValue(observation('SUSPENDED')); + mockSave.mockRejectedValue(new Error('store timeout')); + const first = await run(); + const second = await run(first.state); + expect((await run(second.state)).kind).toBe('failure'); + expect(strategy.resumeSession).not.toHaveBeenCalled(); +}); + +test('an audit failure does not discard the wake recovery outcome', async () => { + working(); + strategy.pollSession.mockResolvedValue(observation('SUSPENDED')); + emitEvent.mockRejectedValue(new Error('private audit detail')); + expect((await run()).kind).toBe('continue'); + expect(strategy.resumeSession).toHaveBeenCalledTimes(1); + expect(JSON.stringify(mockLogger.warn.mock.calls)).not.toContain('private audit detail'); +}); + +test.each(['post-read', 'wake-write'])('lost %s after Suspend retains the wake obligation across restart', async failure => { + if (failure === 'post-read') { + mockRead.mockResolvedValueOnce(structuredClone(row)) + .mockImplementationOnce(async () => structuredClone(row)) + .mockRejectedValueOnce(new Error('lost read')); + } else { + strategy.suspendSession.mockImplementationOnce(async () => { + approve(); + mockSave.mockRejectedValueOnce(new Error('lost compensating write')); + return { supported: true }; + }); + } + const uncertain = await run(); + expect(uncertain.state.recovery?.kind).toBe('wake'); + const repaired = await run(uncertain.state); + expect(repaired.kind).toBe('continue'); + expect(row.intent?.action).toBe('resume'); + expect(strategy.suspendSession).toHaveBeenCalledTimes(1); +}); + +test('a normal wake in progress is not reported as an unintended suspension', async () => { + approve(); + intent('resume'); + strategy.pollSession.mockResolvedValue(observation('SUSPENDING')); + await run(); + expect(emitEvent).not.toHaveBeenCalled(); +}); + +test.each(['PENDING', 'RUNNING'])('HYDRATING startup is bounded when AWS reports %s', async observed => { + working(); + row = { ...row, status: 'HYDRATING' }; + strategy.pollSession.mockResolvedValue(observation(observed)); + const starting = await run(); + time += 300_000; + expect(await run(starting.state)).toMatchObject({ kind: 'failure', reason: 'startup-deadline' }); +}); + +test('FINALIZING can finish normally but cannot exceed the original lifetime', async () => { + working(); + const active = await run(); + row = { ...row, status: 'FINALIZING' }; + expect((await run(active.state)).kind).toBe('closed'); + time = active.state.sessionDeadlineMs; + expect(await run(active.state)).toMatchObject({ kind: 'failure', reason: 'session-deadline' }); +}); + +test('ordinary RUNNING heartbeat failures remain visible after recovery ends', async () => { + working(); + row = { ...row, taskStartedAtMs: NOW - 600_000, heartbeatAtMs: NOW - 300_000 }; + expect((await run()).heartbeatUnhealthy).toBe(true); + row = { ...row, heartbeatAtMs: NOW }; + expect((await run()).heartbeatUnhealthy).toBe(false); + row = { ...row, heartbeatAtMs: undefined }; + expect((await run()).heartbeatUnhealthy).toBe(true); +}); + +test('AWS RUNNING cannot hide a worker that never consumes a committed decision', async () => { + approve(); + const first = await run(); + expect(first.state.recovery?.kind).toBe('wake'); + time += MICROVM_RECOVERY_TIMEOUT_MS; + expect(await run(first.state)).toMatchObject({ kind: 'failure', reason: 'wake-deadline' }); +}); + +test('an expired gate stuck PENDING is bounded even when AWS stays RUNNING', async () => { + time += 600_000; + const first = await run(); + expect(first.state.recovery?.kind).toBe('wake'); + time += MICROVM_RECOVERY_TIMEOUT_MS; + expect((await run(first.state)).kind).toBe('failure'); +}); + +test('multiple failed requests within one cycle count as one failed cycle', async () => { + row = { ...row, approval: { kind: 'unavailable', errorType: 'TimeoutError' } }; + strategy.pollSession.mockResolvedValue(observation('SUSPENDED')); + mockSave.mockRejectedValue(new Error('write timeout')); + const first = await run(); + expect(first.state.consecutivePollFailures).toBe(1); + const second = await run(first.state); + expect(second.state.consecutivePollFailures).toBe(2); + expect((await run(second.state)).kind).toBe('failure'); +}); + +test('a new gate between wake intent and dispatch prevents that stale Resume', async () => { + approve(); + strategy.pollSession.mockResolvedValue(observation('SUSPENDED')); + mockRead.mockImplementationOnce(async () => structuredClone(row)) + .mockImplementationOnce(async () => { row = { ...row, requestId: 'next-gate' }; return structuredClone(row); }); + await run(); + expect(strategy.resumeSession).not.toHaveBeenCalled(); +}); + +test('anomaly reporting survives a failed poll and rearms after a fresh recovered heartbeat', async () => { + working(); + row = { ...row, heartbeatAtMs: time }; + strategy.pollSession.mockResolvedValue(observation('SUSPENDED')); + const first = await run(); + expect(emitEvent.mock.calls.filter(([type]) => type === 'microvm_suspend_anomaly')).toHaveLength(1); + strategy.pollSession.mockRejectedValueOnce(new Error('transient')); + const failed = await run(first.state); + expect(failed.state.anomalyReported).toBe(true); + const recovering = await run(failed.state); + expect(emitEvent.mock.calls.filter(([type]) => type === 'microvm_suspend_anomaly')).toHaveLength(1); + strategy.pollSession.mockResolvedValue(observation()); + const recovered = await run(recovering.state); + expect(recovered.state.recovery).toBeUndefined(); + strategy.pollSession.mockResolvedValue(observation('SUSPENDED')); + await run(recovered.state); + expect(emitEvent.mock.calls.filter(([type]) => type === 'microvm_suspend_anomaly')).toHaveLength(2); +}); + +describe('cleanup evidence', () => { + test.each(['requested', 'not-found'] as const)('%s ends cleanup without claiming more evidence', async outcome => { + strategy.stopSession.mockResolvedValue({ outcome }); + await stopMicrovmWithDiagnostics({ taskId: 'task', handle, strategy, emitEvent }); + expect(strategy.stopSession).toHaveBeenCalledTimes(1); + expect(emitEvent).not.toHaveBeenCalled(); + }); + test('a conflict gets one bounded retry, then reports uncertainty with the retained handle', async () => { + strategy.stopSession.mockResolvedValue({ outcome: 'unconfirmed', error_type: 'ConflictException', aws_request_id: 'safe-123' }); + await stopMicrovmWithDiagnostics({ taskId: 'task', handle, strategy, emitEvent }); + expect(strategy.stopSession).toHaveBeenCalledTimes(2); + expect(strategy.stopSession.mock.calls[1][1]).toBe(strategy.stopSession.mock.calls[0][1]); + expect(emitEvent).toHaveBeenCalledWith('microvm_cleanup_unconfirmed', { + task_id: 'task', microvm_id: 'vm', error_type: 'ConflictException', aws_request_id: 'safe-123', + }, expect.objectContaining({ abortSignal: expect.any(AbortSignal) })); + }); + test('denied cleanup does not keep retrying and failed audit never masks finalization', async () => { + strategy.stopSession.mockResolvedValue({ outcome: 'unconfirmed', error_type: 'AccessDeniedException' }); + emitEvent.mockRejectedValue(new Error('private audit details')); + await expect(stopMicrovmWithDiagnostics({ taskId: 'task', handle, strategy, emitEvent })).resolves.toBeUndefined(); + expect(strategy.stopSession).toHaveBeenCalledTimes(1); + expect(JSON.stringify(mockLogger.warn.mock.calls)).not.toContain('private audit details'); + }); +}); diff --git a/cdk/test/handlers/shared/microvm-suspend-config.test.ts b/cdk/test/handlers/shared/microvm-suspend-config.test.ts new file mode 100644 index 000000000..1dd37bb39 --- /dev/null +++ b/cdk/test/handlers/shared/microvm-suspend-config.test.ts @@ -0,0 +1,94 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +const mockSend = jest.fn(); +const mockWarn = jest.fn(); +jest.mock('@aws-sdk/client-ssm', () => ({ + SSMClient: jest.fn(() => ({ send: mockSend })), + GetParameterCommand: jest.fn((input: unknown) => ({ input })), +})); +jest.mock('../../../src/handlers/shared/logger', () => ({ logger: { warn: mockWarn } })); + +import { readMicrovmSuspendEnabled } from '../../../src/handlers/shared/microvm-suspend-config'; + +const parameterName = '/backgroundagent-dev/microvm-approval-suspend-enabled'; +const originalName = process.env.MICROVM_APPROVAL_SUSPEND_PARAMETER_NAME; +beforeEach(() => { + process.env.MICROVM_APPROVAL_SUSPEND_PARAMETER_NAME = parameterName; + mockSend.mockReset().mockResolvedValue({ Parameter: { Name: parameterName, Value: 'true' } }); + mockWarn.mockClear(); +}); +afterEach(() => { + jest.restoreAllMocks(); + if (originalName === undefined) delete process.env.MICROVM_APPROVAL_SUSPEND_PARAMETER_NAME; + else process.env.MICROVM_APPROVAL_SUSPEND_PARAMETER_NAME = originalName; +}); + +test('rereads the exact parameter and observes disable without an environment change', async () => { + expect(await readMicrovmSuspendEnabled({})).toBe(true); + mockSend.mockResolvedValue({ Parameter: { Name: parameterName, Value: 'false' } }); + expect(await readMicrovmSuspendEnabled({})).toBe(false); + expect(mockSend).toHaveBeenCalledTimes(2); + expect(mockSend).toHaveBeenLastCalledWith({ input: { Name: parameterName } }, { abortSignal: expect.any(AbortSignal) }); +}); + +test.each(['false', 'TRUE', ' true ', '', undefined])('does not enable on value %s', async value => { + mockSend.mockResolvedValue({ Parameter: { Name: parameterName, Value: value } }); + expect(await readMicrovmSuspendEnabled({})).toBe(false); +}); + +test.each([{}, { Parameter: { Name: '/other', Value: 'true' } }])('requires the requested parameter identity', async response => { + mockSend.mockResolvedValue(response); + expect(await readMicrovmSuspendEnabled({})).toBe(false); +}); + +test('missing configuration never starts a request', async () => { + delete process.env.MICROVM_APPROVAL_SUSPEND_PARAMETER_NAME; + expect(await readMicrovmSuspendEnabled({})).toBe(false); + expect(mockSend).not.toHaveBeenCalled(); +}); + +test('read failure disables savings and logs only safe error identifiers', async () => { + mockSend.mockRejectedValue(Object.assign(new Error('signed-url-secret'), { + name: 'AccessDeniedException', $metadata: { requestId: 'request-123' }, + })); + expect(await readMicrovmSuspendEnabled({})).toBe(false); + expect(JSON.stringify(mockWarn.mock.calls)).not.toContain('signed-url-secret'); + expect(mockWarn).toHaveBeenCalledWith(expect.any(String), { + error_type: 'AccessDeniedException', aws_request_id: 'request-123', + }); +}); + +test('an exhausted parent budget starts no request', async () => { + expect(await readMicrovmSuspendEnabled({ abortSignal: AbortSignal.abort() })).toBe(false); + expect(mockSend).not.toHaveBeenCalled(); +}); + +test.each(['parent', 'local'])('%s expiry aborts the actual SDK request and rejects a late true response', async source => { + const parent = new AbortController(); + const local = new AbortController(); + const timeout = jest.spyOn(AbortSignal, 'timeout').mockReturnValue(local.signal); + mockSend.mockImplementation(async (_command, options) => { + (source === 'parent' ? parent : local).abort(); + expect(options.abortSignal.aborted).toBe(true); + return { Parameter: { Name: parameterName, Value: 'true' } }; + }); + expect(await readMicrovmSuspendEnabled({ abortSignal: parent.signal })).toBe(false); + expect(timeout).toHaveBeenCalledWith(3_000); +}); diff --git a/cdk/test/handlers/shared/microvm-worker-lease.test.ts b/cdk/test/handlers/shared/microvm-worker-lease.test.ts new file mode 100644 index 000000000..100d009f0 --- /dev/null +++ b/cdk/test/handlers/shared/microvm-worker-lease.test.ts @@ -0,0 +1,83 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +const mockSend = jest.fn(); +jest.mock('../../../src/handlers/shared/ua', () => ({ makeDocClient: () => ({ send: mockSend }) })); +jest.mock('@aws-sdk/lib-dynamodb', () => ({ + GetCommand: jest.fn(input => ({ kind: 'get', input })), + TransactWriteCommand: jest.fn(input => ({ kind: 'transact', input })), +})); + +import { ensureWorkerLease, leaseHandleUpdate } from '../../../src/handlers/shared/microvm-worker-lease'; + +const input = { taskId: 'task', userId: 'user', repo: 'owner/repo', attemptId: 'attempt-1', requestHash: 'hash' }; +const lease = { + task_id: 'worker-lease#task', + lease_attempt_id: 'attempt-1', + lease_state: 'ACTIVE', + lease_user_id: 'user', + lease_repo: 'owner/repo', +}; + +beforeEach(() => { + jest.clearAllMocks(); + process.env.CONTINUATION_BUCKET_NAME = 'continuations'; + mockSend.mockResolvedValue({}); +}); +afterEach(() => { delete process.env.CONTINUATION_BUCKET_NAME; }); + +test('creates authority only alongside the same task owner and immutable start receipt', async () => { + await ensureWorkerLease(input); + const transaction = mockSend.mock.calls[0][0].input.TransactItems; + expect(transaction[0].ConditionCheck.ExpressionAttributeValues).toMatchObject({ + ':user': 'user', ':attempt': 'attempt-1', ':hash': 'hash', + }); + expect(transaction[1].Put).toMatchObject({ + Item: lease, ConditionExpression: 'attribute_not_exists(task_id)', + }); + // Reserved rows cannot enter ordinary task status/user indexes. + expect(transaction[1].Put.Item).not.toHaveProperty('status'); + expect(transaction[1].Put.Item).not.toHaveProperty('user_id'); +}); + +test('lost successful acknowledgement accepts the same active lease with a registered handle', async () => { + mockSend.mockRejectedValueOnce(new Error('reply lost')).mockResolvedValueOnce({ + Item: { ...lease, lease_microvm_id: 'microvm-one' }, + }); + await expect(ensureWorkerLease(input)).resolves.toBeUndefined(); + expect(mockSend.mock.calls[1][0].input.ConsistentRead).toBe(true); +}); + +test.each([ + { lease_state: 'FENCED' }, { lease_state: 'PARKED' }, { lease_attempt_id: 'new-attempt' }, + { lease_user_id: 'other' }, { lease_repo: 'other/repo' }, +])('a replay cannot replace or reactivate changed authority: %j', async changed => { + mockSend.mockRejectedValueOnce(new Error('conditional failure')).mockResolvedValueOnce({ + Item: { ...lease, ...changed }, + }); + await expect(ensureWorkerLease(input)).rejects.toThrow('conditional failure'); + expect(mockSend).toHaveBeenCalledTimes(2); +}); + +test('handle registration never changes the lease state', () => { + const update = leaseHandleUpdate('task', 'attempt-1', 'microvm-one').Update; + expect(update.Key).toEqual({ task_id: 'worker-lease#task' }); + expect(update.UpdateExpression).toBe('SET lease_microvm_id = :id'); + expect(update.ConditionExpression).toContain('lease_attempt_id = :attempt'); +}); diff --git a/cdk/test/handlers/shared/orchestrator.test.ts b/cdk/test/handlers/shared/orchestrator.test.ts index b53e1671a..71578f6f2 100644 --- a/cdk/test/handlers/shared/orchestrator.test.ts +++ b/cdk/test/handlers/shared/orchestrator.test.ts @@ -42,8 +42,10 @@ import { TaskStatus } from '../../../src/constructs/task-status'; import type { SessionHandle, SessionStatus } from '../../../src/handlers/shared/compute-strategy'; // The real classifier: the reason-append must not break the anchor the substrate // -failure classification keys on. -import { classifyError } from '../../../src/handlers/shared/error-classifier'; -import { buildComputeMetadata, reconcileMicrovmSubstrateState } from '../../../src/handlers/shared/orchestrator'; +import { classifyError, formatMicrovmTerminalFailure } from '../../../src/handlers/shared/error-classifier'; +import { renderFailureReply, renderPanelFailureReason } from '../../../src/handlers/shared/failure-reply'; +import { buildComputeMetadata, finalizeTask } from '../../../src/handlers/shared/orchestrator'; +import { toTaskDetail, type TaskRecord } from '../../../src/handlers/shared/types'; const MICROVM_ID = 'mvm-0123456789abcdef'; const ENDPOINT = 'https://mvm-0123456789abcdef.microvm.lambda.us-east-1.amazonaws.com'; @@ -57,40 +59,51 @@ function commandsOfType(type: string): Array<{ _type: string; input: Record c._type === type); } -/** - * Prime the mocked doc client: the FIRST Get returns a task row with - * ``rereadStatus``; every Put/Update resolves empty. Mirrors the single re-read - * `reconcileMicrovmSubstrateState` performs before failing a task. - */ -function primeReread(rereadStatus: string): void { - mockDdbSend.mockImplementation((cmd: { _type: string }) => { - if (cmd._type === 'Get') { - return Promise.resolve({ - Item: { task_id: 'TASK001', user_id: 'user-1', repo: 'org/repo', status: rereadStatus }, - }); - } - return Promise.resolve({}); - }); +/** Strong finalization observes this committed task before choosing an outcome. */ +function primeReread(status: string): void { + mockDdbSend.mockImplementation((cmd: { _type: string }) => Promise.resolve(cmd._type === 'Get' ? { + Item: { + task_id: 'TASK001', + user_id: 'user-1', + repo: 'org/repo', + status, + memory_written: true, + compute_type: 'lambda-microvm', + session_id: MICROVM_ID, + compute_metadata: { microvmId: MICROVM_ID, endpoint: ENDPOINT }, + }, + } : {})); } - -const CORRELATION = { user_id: 'user-1', repo: 'org/repo' }; - -function reconcile(substrate: SessionStatus, ddbStatus: string, suspendAnomalyReported?: boolean) { - return reconcileMicrovmSubstrateState({ - taskId: 'TASK001', - ddbStatus: ddbStatus as never, - substrate, - microvmId: MICROVM_ID, - userId: 'user-1', - correlation: CORRELATION, - log: mockLogger, - repo: 'org/repo', - ...(suspendAnomalyReported !== undefined && { suspendAnomalyReported }), - }); +const mockRelease = jest.fn(); +jest.mock('../../../src/handlers/shared/task-concurrency', () => ({ + acquireTaskSlot: jest.fn(), releaseTaskSlot: (...args: unknown[]) => mockRelease(...args), +})); +function finish(substrate: SessionStatus, polledStatus = TaskStatus.RUNNING) { + return finalizeTask('TASK001', { + attempts: 1, + lastStatus: polledStatus, + microvmFailureReason: 'substrate-terminal', + microvmFailureMessage: formatMicrovmTerminalFailure( + substrate.status === 'failed' ? substrate.error : `substrate state ${substrate.status}`, substrate.reason, + ), + microvmSupervisor: { + version: 1, + microvmId: MICROVM_ID, + firstObservedAtMs: 1, + sessionDeadlineMs: 28_800_001, + lifetimeVerified: true, + consecutivePollFailures: 0, + consecutiveResumeFailures: 0, + anomalyReported: false, + nextPollInMs: 5_000, + }, + }, 'user-1'); } beforeEach(() => { jest.clearAllMocks(); + mockLogger.child.mockReturnValue(mockLogger); + mockRelease.mockResolvedValue(false); mockDdbSend.mockReset(); mockDdbSend.mockResolvedValue({}); }); @@ -132,14 +145,23 @@ describe('buildComputeMetadata', () => { expect(buildComputeMetadata(handle)).toEqual({ microvmId: MICROVM_ID, endpoint: ENDPOINT }); }); - test('never carries the MicroVM image ARN (deployment config, not session state)', () => { + test('preserves actual image identity and verified capability for later policy decisions', () => { const metadata = buildComputeMetadata({ sessionId: MICROVM_ID, strategyType: 'lambda-microvm', microvmId: MICROVM_ID, endpoint: ENDPOINT, + imageArn: 'arn:aws:lambda:us-east-1:123456789012:microvm-image:test', + imageVersion: '3.0', + lifecycleProtocol: '1', + }); + expect(metadata).toEqual({ + microvmId: MICROVM_ID, + endpoint: ENDPOINT, + imageArn: 'arn:aws:lambda:us-east-1:123456789012:microvm-image:test', + imageVersion: '3.0', + lifecycleProtocol: '1', }); - expect(Object.keys(metadata).sort()).toEqual(['endpoint', 'microvmId']); }); test('produces only string values (compute_metadata is Record in DDB)', () => { @@ -161,329 +183,60 @@ describe('buildComputeMetadata', () => { }); }); -describe('reconcileMicrovmSubstrateState', () => { - describe('running substrate', () => { - test('is a no-op: no DDB reads, no events, task not failed', async () => { - const result = await reconcile({ status: 'running' }, TaskStatus.RUNNING); - - // `suspendAnomalyReported: false` RE-ARMS the once-per-episode event: a VM - // that resumed and is later suspended again earns a fresh anomaly event. - expect(result).toEqual({ taskFailed: false, suspendAnomalyReported: false }); - expect(mockDdbSend).not.toHaveBeenCalled(); - }); - }); - - describe('suspended substrate', () => { - test('is healthy while the task is AWAITING_APPROVAL — no event, no failure', async () => { - const result = await reconcile({ status: 'suspended' }, TaskStatus.AWAITING_APPROVAL); - - // The orchestrator-intended suspend during an approval wait is the whole - // economic point of the backend: it must be silent — and it re-arms the - // anomaly event, because leaving AWAITING_APPROVAL while still suspended - // would be a new, genuinely reportable episode. - expect(result).toEqual({ taskFailed: false, suspendAnomalyReported: false }); - expect(mockDdbSend).not.toHaveBeenCalled(); - expect(mockLogger.warn).not.toHaveBeenCalled(); - }); - - test('writes an anomaly event and does NOT fail the task when the status is RUNNING', async () => { - const result = await reconcile({ status: 'suspended' }, TaskStatus.RUNNING); - - expect(result).toEqual({ taskFailed: false, suspendAnomalyReported: true }); - - const puts = commandsOfType('Put'); - expect(puts).toHaveLength(1); - expect(puts[0].input.TableName).toBe('TaskEvents'); - const item = puts[0].input.Item as Record; - expect(item.event_type).toBe('microvm_suspend_anomaly'); - expect(item.task_id).toBe('TASK001'); - // Correlation envelope (#245) stamped as top-level fields. - expect(item.user_id).toBe('user-1'); - expect(item.repo).toBe('org/repo'); - expect(item.metadata).toEqual({ - microvm_id: MICROVM_ID, - task_status: TaskStatus.RUNNING, - reason: 'suspended_outside_approval_wait', - }); - - // Crucially: no status transition — a suspended VM is resumable, so - // failing the task would destroy recoverable work. - expect(commandsOfType('Update')).toHaveLength(0); - expect(mockLogger.warn).toHaveBeenCalled(); - }); - - test.each([ - TaskStatus.HYDRATING, - TaskStatus.RUNNING, - TaskStatus.FINALIZING, - ])('treats suspended + %s as an anomaly rather than a failure', async (status) => { - const result = await reconcile({ status: 'suspended' }, status); - - expect(result).toEqual({ taskFailed: false, suspendAnomalyReported: true }); - expect(commandsOfType('Put')[0].input.Item).toMatchObject({ - event_type: 'microvm_suspend_anomaly', - metadata: { task_status: status }, - }); - }); - - test('emits the anomaly event ONCE across repeated polls of the same episode', async () => { - // The poll runs every ~30 s for up to 8.5 h; without the flag an - // out-of-band suspend would write ~960 identical TaskEvents, burying the - // first informative one. The caller threads the returned flag back in. - let reported: boolean | undefined; - for (let poll = 0; poll < 5; poll += 1) { - const result = await reconcile({ status: 'suspended' }, TaskStatus.RUNNING, reported); - reported = result.suspendAnomalyReported; - // The no-fail-fast behaviour is unchanged on EVERY poll — that is the - // property the suppression must not break. - expect(result.taskFailed).toBe(false); - expect(result.suspendAnomalyReported).toBe(true); - } - - expect(commandsOfType('Put')).toHaveLength(1); - expect((commandsOfType('Put')[0].input.Item as Record).event_type) - .toBe('microvm_suspend_anomaly'); - // The WARN log is deliberately NOT suppressed: per-poll evidence is what a - // timeline investigation needs, and CloudWatch is not a user-facing surface. - expect(mockLogger.warn).toHaveBeenCalledTimes(5); - }); - - test('the repeat-suppressed polls record that the event was already reported', async () => { - await reconcile({ status: 'suspended' }, TaskStatus.RUNNING, true); - - expect(commandsOfType('Put')).toHaveLength(0); - expect(mockLogger.warn).toHaveBeenCalledWith( - expect.stringContaining('suspended while the task is not awaiting approval'), - expect.objectContaining({ anomaly_already_reported: true }), - ); - }); - - test('RE-ARMS after the VM resumes, so a second episode emits again', async () => { - // Recovery genuinely re-arms (documented decision): a flapping suspend loop - // is the pathology an operator most needs to see, and latching forever - // would hide it after the first occurrence. - const first = await reconcile({ status: 'suspended' }, TaskStatus.RUNNING, false); - expect(first.suspendAnomalyReported).toBe(true); - - const recovered = await reconcile({ status: 'running' }, TaskStatus.RUNNING, first.suspendAnomalyReported); - expect(recovered.suspendAnomalyReported).toBe(false); - - const second = await reconcile({ status: 'suspended' }, TaskStatus.RUNNING, recovered.suspendAnomalyReported); - expect(second.suspendAnomalyReported).toBe(true); - - // Two episodes → two events. - expect(commandsOfType('Put')).toHaveLength(2); - }); - - test('RE-ARMS when the task enters AWAITING_APPROVAL, so a later out-of-band suspend reports', async () => { - const first = await reconcile({ status: 'suspended' }, TaskStatus.RUNNING, false); - expect(first.suspendAnomalyReported).toBe(true); - - // The gate opened: this suspend is now the intended one. - const intended = await reconcile( - { status: 'suspended' }, TaskStatus.AWAITING_APPROVAL, first.suspendAnomalyReported); - expect(intended.suspendAnomalyReported).toBe(false); - expect(commandsOfType('Put')).toHaveLength(1); - - // The gate closed but the VM is still suspended — a new anomaly. - const third = await reconcile( - { status: 'suspended' }, TaskStatus.RUNNING, intended.suspendAnomalyReported); - expect(third.suspendAnomalyReported).toBe(true); - expect(commandsOfType('Put')).toHaveLength(2); - }); - - test('defaults to NOT-yet-reported when the caller omits the flag', async () => { - // Back-compat for any caller (and the first poll of every task) that has no - // prior state: the event must fire, not be suppressed by an undefined flag. - const result = await reconcile({ status: 'suspended' }, TaskStatus.RUNNING); - expect(result.suspendAnomalyReported).toBe(true); - expect(commandsOfType('Put')).toHaveLength(1); - }); +describe('MicroVM terminal finalization', () => { + test.each([ + ['MicroVM host unavailable.', 'MICROVM_SUBSTRATE_TERMINATED', 'compute', true], + ['capacity unavailable in this Availability Zone.', 'MICROVM_SUBSTRATE_TERMINATED', 'compute', true], + ['MicroVM unavailable in this region.', 'MICROVM_SUBSTRATE_TERMINATED', 'compute', true], + ['INSUFFICIENT_GITHUB_REPO_PERMISSIONS', 'MICROVM_SUBSTRATE_TERMINATED', 'compute', true], + ['BLOCKED[missing_secret]: diagnostic text', 'MICROVM_SUBSTRATE_TERMINATED', 'compute', true], + ["agent_status='success', build_ok=False", 'MICROVM_SUBSTRATE_TERMINATED', 'compute', true], + ["agent_status='success', build_ok=timeout [auto-retried]", 'MICROVM_SUBSTRATE_TERMINATED', 'compute', true], + ['Run lifecycle hook returned HTTP status 400.', 'MICROVM_RUN_HOOK_REJECTED', 'config', false], + ['Run lifecycle hook returned HTTP status 500.', 'MICROVM_SUBSTRATE_TERMINATED', 'compute', true], + ['Resume lifecycle hook failed. Please check your hook endpoint and application logs for more details.', 'MICROVM_RESUME_HOOK_FAILED', 'compute', false], + ])('persists stable classification and consistent user guidance for %s', async (reason, code, category, retryable) => { + primeReread(TaskStatus.RUNNING); + await finish({ status: 'completed', reason }); + const values = commandsOfType('Update')[0].input.ExpressionAttributeValues as Record; + const errorMessage = String(values[':attr_error_message']); + expect(errorMessage).toMatch(new RegExp(`^${code}: `)); + expect(errorMessage).toContain(reason); + expect(toTaskDetail({ + task_id: 'TASK001', status: TaskStatus.FAILED, error_message: errorMessage, + } as TaskRecord).error_classification).toMatchObject({ category, retryable }); + const input = { status: TaskStatus.FAILED, errorMessage, taskId: 'TASK001' }; + for (const reply of [renderFailureReply(input), renderPanelFailureReason(input)]) { + expect(reply).toMatch(retryable ? /reply here to try again/i : /needs your ABCA admin/i); + expect(reply).not.toContain('Lambda MicroVMs is not available in this Region'); + expect(reply).not.toContain('I automatically tried again'); + } }); - describe('terminal substrate', () => { - test('fails the task when the re-read status is still non-terminal', async () => { - primeReread(TaskStatus.RUNNING); - - const result = await reconcile({ status: 'completed' }, TaskStatus.RUNNING); - - expect(result).toEqual({ taskFailed: true, suspendAnomalyReported: false }); - - // Re-read before acting (guards the "agent wrote terminal, VM torn down" - // race), then the FAILED transition. - expect(commandsOfType('Get')).toHaveLength(1); - const updates = commandsOfType('Update'); - expect(updates).toHaveLength(1); - expect(updates[0].input.TableName).toBe('Tasks'); - const values = updates[0].input.ExpressionAttributeValues as Record; - expect(values[':toStatus']).toBe(TaskStatus.FAILED); - expect(values[':fromStatus']).toBe(TaskStatus.RUNNING); - // The reason string is what error-classifier keys the substrate-failure - // classification on — keep the two in lockstep. - expect(values[':attr_error_message']).toBe( - 'MicroVM substrate terminated before the agent wrote a terminal status: substrate state completed', - ); - - // Plus the task_failed audit event. - const puts = commandsOfType('Put'); - expect(puts).toHaveLength(1); - expect((puts[0].input.Item as Record).event_type).toBe('task_failed'); - }); - - test('does NOT fail the task when the re-read shows the agent already wrote a terminal status', async () => { - primeReread(TaskStatus.COMPLETED); - - const result = await reconcile({ status: 'completed' }, TaskStatus.RUNNING); - - // Normal shutdown ordering: agent writes COMPLETED, exits, VM terminates. - // Without the re-read this would have failed a successful task. - expect(result).toEqual({ taskFailed: false, suspendAnomalyReported: false }); - expect(commandsOfType('Update')).toHaveLength(0); - expect(commandsOfType('Put')).toHaveLength(0); - }); - - test.each([ - TaskStatus.COMPLETED, - TaskStatus.FAILED, - TaskStatus.CANCELLED, - TaskStatus.TIMED_OUT, - ])('accepts a re-read terminal status of %s without failing the task', async (status) => { + test.each([TaskStatus.COMPLETED, TaskStatus.FAILED, TaskStatus.CANCELLED, TaskStatus.TIMED_OUT])( + 'preserves a committed %s winner after a stale active poll', async status => { primeReread(status); - - const result = await reconcile({ status: 'completed' }, TaskStatus.RUNNING); - - expect(result).toEqual({ taskFailed: false, suspendAnomalyReported: false }); - expect(commandsOfType('Update')).toHaveLength(0); - }); - - test('carries the substrate error detail into the failure reason', async () => { - primeReread(TaskStatus.RUNNING); - - const result = await reconcile({ status: 'failed', error: 'host fault' }, TaskStatus.RUNNING); - - expect(result).toEqual({ taskFailed: true, suspendAnomalyReported: false }); - const values = commandsOfType('Update')[0].input.ExpressionAttributeValues as Record; - expect(values[':attr_error_message']).toBe( - 'MicroVM substrate terminated before the agent wrote a terminal status: host fault', - ); - }); - - // --- stateReason in the detail (review B1) --- - - test('appends the substrate reason so the DOMINANT failure names its real cause', async () => { - // The exact live shape: a /run hook 4xx self-terminates the VM in ~12 s - // (645-p2-smoke-runbook.md §6.1). Without the reason this read "substrate - // state completed" and the classifier's remedy named a session duration cap, a - // host fault, or an external terminate — none of which happened. - primeReread(TaskStatus.RUNNING); - const reason = 'Run lifecycle hook returned HTTP status 400. Please check your hook endpoint ' - + 'and application logs for more details.'; - - await reconcile({ status: 'completed', reason }, TaskStatus.RUNNING); - - const values = commandsOfType('Update')[0].input.ExpressionAttributeValues as Record; - expect(values[':attr_error_message']).toBe( - 'MicroVM substrate terminated before the agent wrote a terminal status: ' - + `substrate state completed (${reason})`, - ); - }); - - test('appends the reason to a failed substrate report too, without losing the error', async () => { - primeReread(TaskStatus.RUNNING); - - await reconcile( - { status: 'failed', error: 'host fault', reason: 'hypervisor evicted the guest' }, - TaskStatus.RUNNING, - ); - - const values = commandsOfType('Update')[0].input.ExpressionAttributeValues as Record; - expect(values[':attr_error_message']).toBe( - 'MicroVM substrate terminated before the agent wrote a terminal status: ' - + 'host fault (hypervisor evicted the guest)', - ); - }); - - test('renders unchanged when the substrate supplies no reason', async () => { - // A live-verified hung MicroVM reports no `stateReason` at all, so the - // reason-less string stays the baseline — and stays the one the classifier's - // `MicroVM substrate terminated…` pattern is anchored on. - primeReread(TaskStatus.RUNNING); - - await reconcile({ status: 'completed' }, TaskStatus.RUNNING); - - const values = commandsOfType('Update')[0].input.ExpressionAttributeValues as Record; - expect(values[':attr_error_message']).toBe( - 'MicroVM substrate terminated before the agent wrote a terminal status: substrate state completed', - ); - }); - - test('a reason-carrying message still classifies, and a hook 4xx outranks the generic entry', async () => { - // The append must not break classification — that would trade a misleading - // remedy for no remedy. It now does BETTER than preserve the generic anchor: - // a hook-4xx reason reaches a dedicated NON-retryable entry, because every - // 4xx the guest can answer is a wiring fault an identical retry cannot fix. - // Both strings live in one `error_message`, so this is really an assertion - // about classifier ORDER. - primeReread(TaskStatus.RUNNING); - await reconcile({ status: 'completed', reason: 'Run lifecycle hook returned HTTP status 400.' }, TaskStatus.RUNNING); - const values = commandsOfType('Update')[0].input.ExpressionAttributeValues as Record; - - const classification = classifyError(String(values[':attr_error_message'])); - - expect(classification!.title).toBe('The MicroVM rejected its own run payload'); - expect(classification!.retryable).toBe(false); - }); - - test('a NON-hook reason keeps the generic retryable substrate-failure entry', async () => { - // The other half: the `MicroVM substrate terminated…` anchor must still be - // the answer for the reasons it was written for (duration cap, host fault, - // external terminate), so the entry above must not have swallowed them. - primeReread(TaskStatus.RUNNING); - await reconcile( - { status: 'completed', reason: 'host fault (hypervisor evicted the guest)' }, - TaskStatus.RUNNING, - ); - const values = commandsOfType('Update')[0].input.ExpressionAttributeValues as Record; - - const classification = classifyError(String(values[':attr_error_message'])); - - expect(classification!.title).toBe('The MicroVM stopped before the agent reported a result'); - expect(classification!.retryable).toBe(true); - }); - - test('fails from AWAITING_APPROVAL too — a terminated VM cannot resume the gate', async () => { - primeReread(TaskStatus.AWAITING_APPROVAL); - - const result = await reconcile({ status: 'completed' }, TaskStatus.AWAITING_APPROVAL); - - expect(result).toEqual({ taskFailed: true, suspendAnomalyReported: false }); - const values = commandsOfType('Update')[0].input.ExpressionAttributeValues as Record; - expect(values[':fromStatus']).toBe(TaskStatus.AWAITING_APPROVAL); - expect(values[':toStatus']).toBe(TaskStatus.FAILED); - }); - - test('transitions from the RE-READ status, not the stale polled status', async () => { - // Task moved HYDRATING → RUNNING between the poll read and the re-read; the - // conditional transition must use the fresh value or it fails its own - // ConditionExpression and the task is left stuck. - primeReread(TaskStatus.RUNNING); - - await reconcile({ status: 'completed' }, TaskStatus.HYDRATING); - - const values = commandsOfType('Update')[0].input.ExpressionAttributeValues as Record; - expect(values[':fromStatus']).toBe(TaskStatus.RUNNING); - }); - - test('does not decrement concurrency — the finalize step owns the release', async () => { - primeReread(TaskStatus.RUNNING); - - await reconcile({ status: 'completed' }, TaskStatus.RUNNING); - - // Matches the ECS substrate-failure branch: failTask(..., releaseConcurrency=false). - const concurrencyWrites = commandsOfType('Update').filter( - c => c.input.TableName === 'Concurrency', - ); - expect(concurrencyWrites).toHaveLength(0); - }); + await finish({ status: 'completed' }); + expect(commandsOfType('Get')[0].input.ConsistentRead).toBe(true); + expect(commandsOfType('Update')).toHaveLength(1); + for (const command of commandsOfType('Update')) { + expect(command.input.ConditionExpression).toBeUndefined(); // terminal TTL stamp only + } + expect((commandsOfType('Put')[0].input.Item as Record).event_type).toBe(`task_${status.toLowerCase()}`); + expect(mockRelease).toHaveBeenCalledTimes(1); + }, + ); + + test('transitions from the strong current status, retaining the original service error', async () => { + primeReread(TaskStatus.AWAITING_APPROVAL); + await finish({ status: 'failed', error: 'host fault', reason: 'hypervisor evicted the guest' }); + const values = commandsOfType('Update')[0].input.ExpressionAttributeValues as Record; + expect(values[':fromStatus']).toBe(TaskStatus.AWAITING_APPROVAL); + expect(values[':toStatus']).toBe(TaskStatus.FAILED); + expect(values[':attr_error_message']).toBe( + 'MICROVM_SUBSTRATE_TERMINATED: MicroVM substrate terminated before the agent wrote a terminal status: host fault (hypervisor evicted the guest)', + ); + expect(classifyError(String(values[':attr_error_message']))?.retryable).toBe(true); + expect(mockRelease).toHaveBeenCalledWith('TASK001', 'user-1'); }); }); diff --git a/cdk/test/handlers/shared/payload-bootstrap.test.ts b/cdk/test/handlers/shared/payload-bootstrap.test.ts new file mode 100644 index 000000000..c14f5ec8d --- /dev/null +++ b/cdk/test/handlers/shared/payload-bootstrap.test.ts @@ -0,0 +1,185 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { createHash } from 'node:crypto'; + +const mockSend = jest.fn(); +const mockSign = jest.fn(); +const mockCredentials = jest.fn(); +jest.mock('@aws-sdk/client-s3', () => ({ + S3Client: jest.fn(() => ({ send: mockSend, config: { credentials: mockCredentials } })), + PutObjectCommand: jest.fn(input => ({ kind: 'put', input })), + GetObjectCommand: jest.fn(input => ({ kind: 'get', input })), + DeleteObjectCommand: jest.fn(input => ({ kind: 'delete', input })), +})); +jest.mock('@aws-sdk/s3-request-presigner', () => ({ + getSignedUrl: (...args: unknown[]) => mockSign(...args), +})); + +import { deletePayloadReference, PAYLOAD_BOOTSTRAP, preparePayloadReference, redactPayloadUrls } from '../../../src/handlers/shared/payload-bootstrap'; + +const objects = new Map(); +const input = { + bucket: 'payload-bucket', + taskId: 'task-1', + backend: 'lambda-microvm' as const, + payload: { task_id: 'task-1', prompt: 'hello' }, + platformConfig: { github_token_secret_arn: 'arn:aws:secretsmanager:us-west-2:123456789012:secret:github' }, +}; +let loseReplyFor: string | undefined; + +beforeEach(() => { + jest.clearAllMocks(); + objects.clear(); + loseReplyFor = undefined; + mockCredentials.mockResolvedValue({ accessKeyId: 'EXAMPLE', secretAccessKey: 'unused' }); + mockSign.mockImplementation(async () => `https://payload-bucket.s3.us-east-1.amazonaws.com/task-1/payload.json?X-Amz-Security-Token=secret&X-Amz-Signature=${mockSign.mock.calls.length}`); + mockSend.mockImplementation(async ({ kind, input: request }) => { + if (kind === 'get') { + if (!objects.has(request.Key)) throw Object.assign(new Error('missing'), { name: 'NoSuchKey' }); + return { Body: { transformToString: async () => objects.get(request.Key)! } }; + } + if (kind === 'delete') { objects.delete(request.Key); return {}; } + if (request.IfNoneMatch === '*' && objects.has(request.Key)) { + throw Object.assign(new Error('exists'), { name: 'PreconditionFailed' }); + } + objects.set(request.Key, request.Body); + if (loseReplyFor === request.Key) { + loseReplyFor = undefined; + throw new Error('lost committed reply'); + } + return {}; + }); +}); + +test('separates public manifest from private payload and saved capability', async () => { + const ref = await preparePayloadReference(input); + const key = ref.bootstrap_s3_uri.split('/').slice(3).join('/'); + const manifest = objects.get(key)!; + expect(key).toBe(`bootstrap/${createHash('sha256').update(manifest).digest('hex')}.json`); + expect(JSON.parse(manifest)).toEqual({ + version: 2, backend: 'lambda-microvm', platform_config: input.platformConfig, + }); + expect(manifest).not.toContain('hello'); + expect(manifest).not.toContain('X-Amz'); + expect(JSON.parse(objects.get('task-1/payload.json')!).agent_payload).toEqual(input.payload); + expect(JSON.parse(objects.get('task-1/launch.json')!).reference).toEqual(ref); + expect(mockSign).toHaveBeenCalledWith(expect.anything(), expect.objectContaining({ + input: { Bucket: input.bucket, Key: 'task-1/payload.json' }, + }), { expiresIn: 900 }); +}); + +test('replay and concurrent preparation return identical launch bytes without resigning on replay', async () => { + const first = await preparePayloadReference(input); + const replay = await preparePayloadReference({ ...input, payload: { prompt: 'hello', task_id: 'task-1' } }); + expect(JSON.stringify(replay)).toBe(JSON.stringify(first)); + expect(mockSign).toHaveBeenCalledTimes(1); + // Fixed pair, covering conditional object publication by competing callers. + objects.clear(); + // eslint-disable-next-line @cdklabs/promiseall-no-unbounded-parallelism + const [a, b] = await Promise.all([preparePayloadReference(input), preparePayloadReference(input)]); + expect(JSON.stringify(a)).toBe(JSON.stringify(b)); +}); + +test.each(['task-1/payload.json', 'task-1/launch.json'])('recovers a lost committed response for %s', async key => { + loseReplyFor = key; + const first = await preparePayloadReference(input); + expect(await preparePayloadReference(input)).toEqual(first); +}); + +test('changed instructions cannot overwrite an existing payload or launch reference', async () => { + await preparePayloadReference(input); + const before = objects.get('task-1/payload.json'); + await expect(preparePayloadReference({ ...input, payload: { ...input.payload, prompt: 'changed' } })) + .rejects.toThrow('PAYLOAD_BOOTSTRAP_CONFLICT'); + expect(objects.get('task-1/payload.json')).toBe(before); +}); + +test('replacement attempts have independent immutable objects and scoped cleanup', async () => { + await preparePayloadReference(input); + const original = objects.get('task-1/payload.json'); + const replacement = { + ...input, attemptId: 'replacement-2', payload: { ...input.payload, attempt_id: 'replacement-2' }, + }; + const reference = await preparePayloadReference(replacement); + expect(reference.attempt_id).toBe('replacement-2'); + expect(objects.has('task-1/replacement-2/launch.json')).toBe(true); + expect(await preparePayloadReference(replacement)).toEqual(reference); + await expect(preparePayloadReference({ + ...replacement, payload: { ...replacement.payload, prompt: 'changed' }, + })).rejects.toThrow('CONFLICT'); + expect(objects.get('task-1/payload.json')).toBe(original); + await deletePayloadReference(input.bucket, input.taskId, replacement.attemptId); + expect(objects.has('task-1/replacement-2/payload.json')).toBe(false); + expect(objects.get('task-1/payload.json')).toBe(original); +}); + +test.each(['../other', 'different-attempt'])('rejects invalid or mismatched attempt %s before S3', async attemptId => { + await expect(preparePayloadReference({ ...input, attemptId, payload: { ...input.payload, attempt_id: 'replacement-2' } })) + .rejects.toThrow('worker attempt'); + expect(mockSend).not.toHaveBeenCalled(); +}); + +test('orphaned payload after a crash cannot be overwritten with changed instructions', async () => { + await preparePayloadReference(input); + objects.delete('task-1/launch.json'); + await expect(preparePayloadReference({ ...input, payload: { ...input.payload, prompt: 'changed' } })) + .rejects.toThrow('PAYLOAD_BOOTSTRAP_CONFLICT'); + expect(JSON.parse(objects.get('task-1/payload.json')!).agent_payload.prompt).toBe('hello'); +}); + +test('expired references fail closed without signing a replacement', async () => { + await preparePayloadReference(input); + const record = JSON.parse(objects.get('task-1/launch.json')!); + record.reference.expires_at = Date.now() - 1; + objects.set('task-1/launch.json', JSON.stringify(record)); + await expect(preparePayloadReference(input)).rejects.toThrow('PAYLOAD_BOOTSTRAP_EXPIRED'); + expect(mockSign).toHaveBeenCalledTimes(1); +}); + +test('credential lifetime bounds the link and refuses credentials expiring too soon', async () => { + mockCredentials.mockResolvedValue({ expiration: new Date(Date.now() + 20_000) }); + await expect(preparePayloadReference(input)).rejects.toThrow('CREDENTIALS_EXPIRING'); + expect(mockSign).not.toHaveBeenCalled(); +}); + +test.each(['../bootstrap', 'task/other', ''])('invalid task key %s performs no S3 writes', async taskId => { + await expect(preparePayloadReference({ ...input, taskId })).rejects.toThrow('identity'); + expect(mockSend).not.toHaveBeenCalled(); +}); + +test('oversized payload fails before uploading or signing', async () => { + await expect(preparePayloadReference({ + ...input, payload: { ...input.payload, prompt: 'x'.repeat(PAYLOAD_BOOTSTRAP.max_payload_bytes) }, + })).rejects.toThrow('TOO_LARGE'); + expect(mockSend).not.toHaveBeenCalled(); +}); + +test('cleanup removes the capability and payload, preserving shared configuration', async () => { + const ref = await preparePayloadReference(input); + await deletePayloadReference(input.bucket, input.taskId); + expect(objects.has('task-1/payload.json')).toBe(false); + expect(objects.has('task-1/launch.json')).toBe(false); + expect(objects.has(ref.bootstrap_s3_uri.split('/').slice(3).join('/'))).toBe(true); +}); + +test('error redaction removes the whole signed URL including the session token', () => { + expect(redactPayloadUrls('failed https://b.s3.us-east-1.amazonaws.com/key?X-Amz-Security-Token=secret&X-Amz-Signature=abc end')) + .toBe('failed [redacted payload URL] end'); +}); diff --git a/cdk/test/handlers/shared/session-lifecycle.test.ts b/cdk/test/handlers/shared/session-lifecycle.test.ts new file mode 100644 index 000000000..58f420140 --- /dev/null +++ b/cdk/test/handlers/shared/session-lifecycle.test.ts @@ -0,0 +1,244 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { inspect } from 'node:util'; + +const mockMicrovmSend = jest.fn(); +const mockEcsSend = jest.fn(); +const mockAgentcoreSend = jest.fn(); +const mockLifecycleLogger = { info: jest.fn(), warn: jest.fn(), error: jest.fn() }; +jest.mock('../../../src/handlers/shared/logger', () => ({ logger: mockLifecycleLogger })); +jest.mock('@aws-sdk/client-lambda-microvms', () => ({ + ...jest.requireActual('@aws-sdk/client-lambda-microvms'), + LambdaMicrovmsClient: jest.fn(() => ({ send: mockMicrovmSend })), +})); +jest.mock('@aws-sdk/client-ecs', () => ({ + ...jest.requireActual('@aws-sdk/client-ecs'), + ECSClient: jest.fn(() => ({ send: mockEcsSend })), +})); +jest.mock('@aws-sdk/client-bedrock-agentcore', () => ({ + ...jest.requireActual('@aws-sdk/client-bedrock-agentcore'), + BedrockAgentCoreClient: jest.fn(() => ({ send: mockAgentcoreSend })), +})); + +import { ResumeMicrovmCommand, SuspendMicrovmCommand } from '@aws-sdk/client-lambda-microvms'; +import { resolveComputeStrategy, type SessionHandle } from '../../../src/handlers/shared/compute-strategy'; +import { MICROVM_LIFECYCLE_REQUEST_TIMEOUT_MS } from '../../../src/handlers/shared/strategies/lambda-microvm-strategy'; + +const handles: SessionHandle[] = [ + { strategyType: 'agentcore', sessionId: 'session-agentcore', runtimeArn: 'arn:runtime' }, + { strategyType: 'ecs', sessionId: 'arn:task', taskArn: 'arn:task', clusterArn: 'arn:cluster' }, + { strategyType: 'lambda-microvm', sessionId: 'mvm-one', microvmId: 'mvm-one', endpoint: 'https://unused.example' }, +]; +const microvm = handles[2] as Extract; + +beforeEach(() => { + jest.clearAllMocks(); + mockMicrovmSend.mockReset().mockResolvedValue({}); + mockEcsSend.mockReset().mockResolvedValue({}); + mockAgentcoreSend.mockReset().mockResolvedValue({}); +}); + +describe.each(handles.filter(handle => handle.strategyType !== 'lambda-microvm'))( + '$strategyType caller budgets', handle => { + test('an expired budget prevents poll/stop requests', async () => { + const controller = new AbortController(); + controller.abort(new Error('caller deadline')); + const options = { abortSignal: controller.signal }; + const strategy = resolveComputeStrategy({ compute_type: handle.strategyType, runtime_arn: '' }); + await expect(strategy.pollSession(handle, options)).rejects.toThrow('caller deadline'); + await expect(strategy.stopSession(handle, options)).resolves.toBeUndefined(); + expect(mockEcsSend).not.toHaveBeenCalled(); + expect(mockAgentcoreSend).not.toHaveBeenCalled(); + }); + test('stop propagates caller cancellation into its pending SDK request', async () => { + const controller = new AbortController(); + const send = handle.strategyType === 'ecs' ? mockEcsSend : mockAgentcoreSend; + send.mockImplementationOnce((_command, options) => new Promise((_resolve, reject) => { + options.abortSignal.addEventListener('abort', () => reject(new Error('caller deadline'))); + })); + const strategy = resolveComputeStrategy({ compute_type: handle.strategyType, runtime_arn: '' }); + const result = strategy.stopSession(handle, { abortSignal: controller.signal }); + controller.abort(); + await expect(result).resolves.toBeUndefined(); + expect(send).toHaveBeenCalledTimes(1); + }); + }, +); + +describe.each(['pollSession', 'suspendSession', 'resumeSession', 'stopSession'] as const)( + '%s composed budget', operation => { + test('does not send a control request after its caller deadline', async () => { + const controller = new AbortController(); + controller.abort(new Error('caller deadline')); + const strategy = resolveComputeStrategy({ compute_type: 'lambda-microvm', runtime_arn: '' }); + const result = strategy[operation](microvm, { abortSignal: controller.signal }); + if (operation === 'stopSession') await expect(result).resolves.toMatchObject({ outcome: 'unconfirmed' }); + else await expect(result).rejects.toThrow('caller deadline'); + expect(mockMicrovmSend).not.toHaveBeenCalled(); + }); + test('a caller can end the in-flight request before the default limit', async () => { + const controller = new AbortController(); + mockMicrovmSend.mockImplementationOnce((_command, options) => new Promise((_resolve, reject) => { + options.abortSignal.addEventListener('abort', () => reject(new Error('caller deadline'))); + })); + const strategy = resolveComputeStrategy({ compute_type: 'lambda-microvm', runtime_arn: '' }); + const result = strategy[operation](microvm, { abortSignal: controller.signal }); + const assertion = operation === 'stopSession' + ? expect(result).resolves.toMatchObject({ outcome: 'unconfirmed' }) + : expect(result).rejects.toThrow('caller deadline'); + controller.abort(); + await assertion; + expect(mockMicrovmSend).toHaveBeenCalledTimes(1); + expect(mockMicrovmSend.mock.calls[0][1].abortSignal.aborted).toBe(true); + }); + }, +); + +describe('MicroVM lifetime observations', () => { + const startedAt = new Date('2026-09-15T10:00:00Z'); + test.each(['RUNNING', 'SUSPENDING', 'SUSPENDED', 'TERMINATED', 'UNKNOWN'])( + 'retains the original service lifetime in %s as durable JSON data', async state => { + mockMicrovmSend.mockResolvedValue({ state, startedAt, maximumDurationInSeconds: 28_800 }); + const strategy = resolveComputeStrategy({ compute_type: 'lambda-microvm', runtime_arn: '' }); + const observation = await strategy.pollSession(microvm); + expect(JSON.parse(JSON.stringify(observation))).toMatchObject({ + microvmStartedAtMs: startedAt.getTime(), microvmMaximumDurationSeconds: 28_800, + }); + }, + ); + test.each([ + { startedAt: new Date('invalid'), maximumDurationInSeconds: 28_800 }, + { startedAt, maximumDurationInSeconds: 0 }, + { startedAt, maximumDurationInSeconds: -1 }, + { startedAt, maximumDurationInSeconds: 1.5 }, + { maximumDurationInSeconds: 28_800 }, + ])('does not invent a lifetime from incomplete service data: %j', async lifetime => { + mockMicrovmSend.mockResolvedValue({ state: 'RUNNING', ...lifetime }); + const strategy = resolveComputeStrategy({ compute_type: 'lambda-microvm', runtime_arn: '' }); + const observation = await strategy.pollSession(microvm); + expect(observation.microvmStartedAtMs).toBeUndefined(); + expect(observation.microvmMaximumDurationSeconds).toBeUndefined(); + }); +}); + +describe.each(['suspendSession', 'resumeSession'] as const)('%s contract', operation => { + test('records AWS acknowledgment identity without response or exception contents', async () => { + const strategy = resolveComputeStrategy({ compute_type: 'lambda-microvm', runtime_arn: '' }); + mockMicrovmSend.mockResolvedValueOnce({ + $metadata: { requestId: 'aws-control-123', headers: { Authorization: 'secret-response' } }, + payload: 'secret-payload', + }); + await expect(strategy[operation](microvm)).resolves.toEqual({ supported: true }); + expect(mockLifecycleLogger.info).toHaveBeenCalledWith( + 'MicroVM lifecycle request acknowledged', + expect.objectContaining({ microvm_id: 'mvm-one', aws_request_id: 'aws-control-123', elapsed_ms: expect.any(Number) }), + ); + mockMicrovmSend.mockRejectedValueOnce(Object.assign(new Error('secret-exception'), { + name: 'ConflictException', $metadata: { requestId: 'aws-control-456' }, + })); + await expect(strategy[operation](microvm)).rejects.toThrow(); + expect(mockLifecycleLogger.warn).toHaveBeenCalledWith( + 'MicroVM lifecycle request failed', + expect.objectContaining({ error_type: 'ConflictException', aws_request_id: 'aws-control-456' }), + ); + expect(JSON.stringify([mockLifecycleLogger.info.mock.calls, mockLifecycleLogger.warn.mock.calls])).not.toContain('secret-'); + }); + + test.each(handles.filter(h => h.strategyType !== 'lambda-microvm'))( + '$strategyType explicitly reports unsupported without calling AWS', async handle => { + const strategy = resolveComputeStrategy({ compute_type: handle.strategyType, runtime_arn: 'arn:runtime' }); + await expect(strategy[operation](handle)).resolves.toEqual({ supported: false }); + expect(mockMicrovmSend).not.toHaveBeenCalled(); + expect(mockEcsSend).not.toHaveBeenCalled(); + expect(mockAgentcoreSend).not.toHaveBeenCalled(); + }, + ); + + test.each(handles)('$strategyType rejects a handle from another backend', async handle => { + const strategy = resolveComputeStrategy({ compute_type: handle.strategyType, runtime_arn: 'arn:runtime' }); + for (const other of handles.filter(h => h.strategyType !== handle.strategyType)) { + await expect(strategy[operation](other)).rejects.toThrow(`${operation} called with non-${handle.strategyType} handle`); + } + expect(mockMicrovmSend).not.toHaveBeenCalled(); + expect(mockEcsSend).not.toHaveBeenCalled(); + expect(mockAgentcoreSend).not.toHaveBeenCalled(); + }); + + test('sends only the identifier and reports acknowledgement, without pretending to observe state', async () => { + const strategy = resolveComputeStrategy({ compute_type: 'lambda-microvm', runtime_arn: '' }); + await expect(strategy[operation](microvm)).resolves.toEqual({ supported: true }); + expect(mockMicrovmSend).toHaveBeenCalledTimes(1); + const [command, options] = mockMicrovmSend.mock.calls[0]; + expect(command).toBeInstanceOf(operation === 'suspendSession' ? SuspendMicrovmCommand : ResumeMicrovmCommand); + expect(command.input).toEqual({ microvmIdentifier: 'mvm-one' }); + expect(options.abortSignal).toBeInstanceOf(AbortSignal); + expect(options.abortSignal.aborted).toBe(false); + }); + + test.each(['', ' ', undefined])('rejects an empty runtime identifier (%s)', async microvmId => { + const strategy = resolveComputeStrategy({ compute_type: 'lambda-microvm', runtime_arn: '' }); + await expect(strategy[operation]({ ...microvm, microvmId } as SessionHandle)).rejects.toThrow('non-empty MicroVM identifier'); + expect(mockMicrovmSend).not.toHaveBeenCalled(); + }); + + test.each([ + 'AccessDeniedException', 'ConflictException', 'ResourceNotFoundException', + 'ThrottlingException', 'InternalServerException', 'ValidationException', + ])('preserves %s as an operational failure with a sanitized cause', async name => { + mockMicrovmSend.mockRejectedValueOnce(Object.assign(new Error( + 'failed https://example.test/task?X-Amz-Signature=BEARER-SECRET', + { cause: new Error('private request metadata: BEARER-SECRET') }, + ), { name })); + const strategy = resolveComputeStrategy({ compute_type: 'lambda-microvm', runtime_arn: '' }); + await strategy[operation](microvm).then( + () => { throw new Error('Expected lifecycle request to fail'); }, + (error: Error) => { + expect(error.message).toContain(name); + expect(error.cause).toMatchObject({ name }); + expect(inspect(error, { depth: null })).not.toContain('BEARER-SECRET'); + }, + ); + }); + + test('a repeated command cannot silently turn a simulated state conflict into success', async () => { + mockMicrovmSend.mockResolvedValueOnce({}).mockRejectedValueOnce(Object.assign(new Error('state changed'), { name: 'ConflictException' })); + const strategy = resolveComputeStrategy({ compute_type: 'lambda-microvm', runtime_arn: '' }); + await expect(strategy[operation](microvm)).resolves.toEqual({ supported: true }); + await expect(strategy[operation](microvm)).rejects.toThrow('ConflictException'); + expect(mockMicrovmSend.mock.calls[0][0].input).toEqual(mockMicrovmSend.mock.calls[1][0].input); + }); + + test('bounds the SDK request and surfaces abort for later reconciliation', async () => { + const controller = new AbortController(); + const timeout = jest.spyOn(AbortSignal, 'timeout').mockReturnValue(controller.signal); + try { + mockMicrovmSend.mockImplementationOnce((_command, options) => new Promise((_resolve, reject) => { + options.abortSignal.addEventListener('abort', () => reject(Object.assign(new Error('deadline'), { name: 'AbortError' }))); + })); + const strategy = resolveComputeStrategy({ compute_type: 'lambda-microvm', runtime_arn: '' }); + const result = expect(strategy[operation](microvm)).rejects.toThrow('AbortError'); + controller.abort(); + await result; + expect(timeout).toHaveBeenCalledWith(MICROVM_LIFECYCLE_REQUEST_TIMEOUT_MS); + } finally { + timeout.mockRestore(); + } + }); +}); diff --git a/cdk/test/handlers/shared/session-start-retry.test.ts b/cdk/test/handlers/shared/session-start-retry.test.ts index ee27474c1..75d9a122a 100644 --- a/cdk/test/handlers/shared/session-start-retry.test.ts +++ b/cdk/test/handlers/shared/session-start-retry.test.ts @@ -214,6 +214,17 @@ describe('startSessionWithRetry — Lambda MicroVM start failures (ADR-021)', () await expect(startSessionWithRetry({ startSession }, {} as never, d)).rejects.toBe(notFound); expect(startSession).toHaveBeenCalledTimes(1); }); + + it.each(['AccessDeniedException', 'UnauthorizedException', 'ValidationException', 'InvalidParameterValueException'])( + 'does NOT retry a MicroVM %s', async (name) => { + const error = new Error(`MicroVM RunMicrovm failed: ${name}: request rejected`); + const startSession = jest.fn().mockRejectedValueOnce(error); + const { d, emitReasons } = deps(); + await expect(startSessionWithRetry({ startSession }, {} as never, d)).rejects.toBe(error); + expect(startSession).toHaveBeenCalledTimes(1); + expect(emitReasons).toEqual([]); + }, + ); }); describe('startSessionWithRetry — other backends keep their pre-ADR-021 retry behaviour', () => { diff --git a/cdk/test/handlers/shared/strategies/ecs-strategy.test.ts b/cdk/test/handlers/shared/strategies/ecs-strategy.test.ts index db474ddd3..904e133cd 100644 --- a/cdk/test/handlers/shared/strategies/ecs-strategy.test.ts +++ b/cdk/test/handlers/shared/strategies/ecs-strategy.test.ts @@ -27,13 +27,9 @@ process.env.ECS_TASK_DEFINITION_ARN = TASK_DEF_ARN; process.env.ECS_SUBNETS = 'subnet-aaa,subnet-bbb'; process.env.ECS_SECURITY_GROUP = 'sg-12345'; process.env.ECS_CONTAINER_NAME = 'AgentContainer'; -// The top-of-file import's inline-fallback / no-op tests assume these OPTIONAL -// vars are ABSENT at load time. They are unset in a dev shell but the real ECS -// agent container HAS ECS_PAYLOAD_BUCKET set (#502) — so leaving this to ambient -// env made the build pass locally yet FAIL on ECS ("works local, dies on ECS"). -// The #502 / #299 describe blocks below set these via isolateModules; delete them -// here so the top-of-file import is hermetic regardless of the runner's env. -delete process.env.ECS_PAYLOAD_BUCKET; +// Pin the payload bucket and optional planning definition at import time so the +// test runner's own deployment environment cannot change these expectations. +process.env.ECS_PAYLOAD_BUCKET = 'payload-bucket'; delete process.env.ECS_PLANNING_TASK_DEFINITION_ARN; const mockSend = jest.fn(); @@ -44,17 +40,20 @@ jest.mock('@aws-sdk/client-ecs', () => ({ StopTaskCommand: jest.fn((input: unknown) => ({ _type: 'StopTask', input })), })); -const mockS3Send = jest.fn(); -jest.mock('@aws-sdk/client-s3', () => ({ - S3Client: jest.fn(() => ({ send: mockS3Send })), - PutObjectCommand: jest.fn((input: unknown) => ({ _type: 'PutObject', input })), - DeleteObjectCommand: jest.fn((input: unknown) => ({ _type: 'DeleteObject', input })), +const mockPrepare = jest.fn(); +const mockDelete = jest.fn(); +jest.mock('../../../../src/handlers/shared/payload-bootstrap', () => ({ + ...jest.requireActual('../../../../src/handlers/shared/payload-bootstrap'), + preparePayloadReference: (...args: unknown[]) => mockPrepare(...args), + deletePayloadReference: (...args: unknown[]) => mockDelete(...args), })); import { EcsComputeStrategy } from '../../../../src/handlers/shared/strategies/ecs-strategy'; beforeEach(() => { jest.clearAllMocks(); + mockPrepare.mockReset().mockImplementation(async ({ taskId }: { taskId: string }) => ({ version: 2, task_id: taskId, bootstrap_s3_uri: 's3://b/bootstrap/example.json', payload_url: 'https://signed.example/task', expires_at: Date.now()+900000 })); + mockDelete.mockResolvedValue(undefined); }); describe('EcsComputeStrategy', () => { @@ -103,21 +102,14 @@ describe('EcsComputeStrategy', () => { expect(envVars).toEqual(expect.arrayContaining([ { name: 'TASK_ID', value: 'TASK001' }, { name: 'REPO_URL', value: 'org/repo' }, - { name: 'TASK_DESCRIPTION', value: 'Fix the bug' }, { name: 'ISSUE_NUMBER', value: '42' }, { name: 'MAX_TURNS', value: '50' }, { name: 'CLAUDE_CODE_USE_BEDROCK', value: '1' }, ])); - // No ECS_PAYLOAD_BUCKET in this module's env → inline fallback (#502): the - // full payload rides in AGENT_PAYLOAD and nothing is written to S3. - const agentPayload = envVars.find((e: { name: string }) => e.name === 'AGENT_PAYLOAD'); - expect(agentPayload).toBeDefined(); - const parsed = JSON.parse(agentPayload.value); - expect(parsed.repo_url).toBe('org/repo'); - expect(parsed.prompt).toBe('Fix the bug'); - expect(envVars.find((e: { name: string }) => e.name === 'AGENT_PAYLOAD_S3_URI')).toBeUndefined(); - expect(mockS3Send).not.toHaveBeenCalled(); + const reference = envVars.find((e: { name: string }) => e.name === 'AGENT_PAYLOAD_REF'); + expect(JSON.parse(reference.value).task_id).toBe('TASK001'); + expect(mockPrepare).toHaveBeenCalledWith(expect.objectContaining({ payload: expect.objectContaining({ repo_url: 'org/repo', prompt: 'Fix the bug' }) })); // Container command override — runs Python directly instead of uvicorn expect(override.command).toBeDefined(); @@ -379,113 +371,47 @@ describe('EcsComputeStrategy', () => { }); }); -// #502: the S3-pointer path requires ECS_PAYLOAD_BUCKET to be set BEFORE the -// module is imported (it's a module-level constant). Re-import under -// jest.isolateModules with the env var set so these tests don't perturb the -// inline-fallback tests above. -describe('EcsComputeStrategy with ECS_PAYLOAD_BUCKET (S3-pointer path, #502)', () => { - const PAYLOAD_BUCKET = 'test-ecs-payload-bucket'; - - function loadStrategyWithBucket(): { - EcsComputeStrategy: typeof import('../../../../src/handlers/shared/strategies/ecs-strategy').EcsComputeStrategy; - deleteEcsPayload: typeof import('../../../../src/handlers/shared/strategies/ecs-strategy').deleteEcsPayload; - ecsPayloadKey: typeof import('../../../../src/handlers/shared/strategies/ecs-strategy').ecsPayloadKey; - } { - let mod!: ReturnType; - jest.isolateModules(() => { - process.env.ECS_PAYLOAD_BUCKET = PAYLOAD_BUCKET; - process.env.ECS_CLUSTER_ARN = CLUSTER_ARN; - process.env.ECS_TASK_DEFINITION_ARN = TASK_DEF_ARN; - process.env.ECS_SUBNETS = 'subnet-aaa,subnet-bbb'; - process.env.ECS_SECURITY_GROUP = 'sg-12345'; - process.env.ECS_CONTAINER_NAME = 'AgentContainer'; - // eslint-disable-next-line @typescript-eslint/no-require-imports - mod = require('../../../../src/handlers/shared/strategies/ecs-strategy'); - }); - return mod; - } - - afterEach(() => { - delete process.env.ECS_PAYLOAD_BUCKET; - }); - - test('writes payload to S3 and passes AGENT_PAYLOAD_S3_URI, not the inline blob', async () => { - mockS3Send.mockResolvedValueOnce({}); +describe('ECS v2 bootstrap integration', () => { + test('passes the prepared capability and consumes it before entering the pipeline', async () => { mockSend.mockResolvedValueOnce({ tasks: [{ taskArn: TASK_ARN }] }); + await new EcsComputeStrategy().startSession({ taskId: 'TASK001', userId: 'u1', payload: { task_id: 'TASK001', prompt: 'x'.repeat(10000) }, blueprintConfig: { compute_type: 'ecs', runtime_arn: '' } }); + expect(mockPrepare).toHaveBeenCalledWith(expect.objectContaining({ backend: 'ecs', taskId: 'TASK001', bucket: 'payload-bucket' })); + const override=mockSend.mock.calls[0][0].input.overrides.containerOverrides[0]; + expect(override.environment.some((e: { name: string }) => ['AGENT_PAYLOAD', 'AGENT_PAYLOAD_S3_URI', 'TASK_DESCRIPTION'].includes(e.name))).toBe(false); + const source=override.command[2]; + expect(source).toContain('load_ecs_payload()'); + expect(source.indexOf('load_ecs_payload()')).toBeLessThan(source.indexOf('from entrypoint')); + }); + test('delegates deletion of the payload and private launch capability', async () => { + const { deleteEcsPayload }=await import('../../../../src/handlers/shared/strategies/ecs-strategy'); + await deleteEcsPayload('TASK001'); + expect(mockDelete).toHaveBeenCalledWith('payload-bucket', 'TASK001'); + }); - const { EcsComputeStrategy: Strategy } = loadStrategyWithBucket(); - const strategy = new Strategy(); - await strategy.startSession({ + test('rejects oversized UTF-8 overrides before RunTask', async () => { + await expect(new EcsComputeStrategy().startSession({ taskId: 'TASK001', - userId: 'cognito-test', - payload: { repo_url: 'org/repo', prompt: 'Fix the bug', hydrated_context: { big: 'x'.repeat(10000) } }, - blueprintConfig: { compute_type: 'ecs', runtime_arn: '' }, - }); - - // PutObject to the payload bucket at /payload.json - expect(mockS3Send).toHaveBeenCalledTimes(1); - const put = mockS3Send.mock.calls[0][0]; - expect(put._type).toBe('PutObject'); - expect(put.input.Bucket).toBe(PAYLOAD_BUCKET); - expect(put.input.Key).toBe('TASK001/payload.json'); - expect(JSON.parse(put.input.Body).repo_url).toBe('org/repo'); - - // Override carries the URI pointer, NOT the inline payload - const envVars = mockSend.mock.calls[0][0].input.overrides.containerOverrides[0].environment; - const uri = envVars.find((e: { name: string }) => e.name === 'AGENT_PAYLOAD_S3_URI'); - expect(uri.value).toBe(`s3://${PAYLOAD_BUCKET}/TASK001/payload.json`); - expect(envVars.find((e: { name: string }) => e.name === 'AGENT_PAYLOAD')).toBeUndefined(); + userId: 'u1', + payload: { task_id: 'TASK001' }, + blueprintConfig: { compute_type: 'ecs', runtime_arn: '', system_prompt_overrides: '🌍'.repeat(2100) }, + })).rejects.toThrow('ECS container overrides exceed 8192 bytes'); + expect(mockSend).not.toHaveBeenCalled(); }); - test('boot command loads payload from S3 when the URI is set, else inline', async () => { - mockS3Send.mockResolvedValueOnce({}); - mockSend.mockResolvedValueOnce({ tasks: [{ taskArn: TASK_ARN }] }); - - const { EcsComputeStrategy: Strategy } = loadStrategyWithBucket(); - await new Strategy().startSession({ + test('redacts a capability even when preparation fails before RunTask', async () => { + mockPrepare.mockRejectedValueOnce(new Error('failed https://bucket.example/key?X-Amz-Signature=BEARER-SECRET')); + await expect(new EcsComputeStrategy().startSession({ taskId: 'TASK001', - userId: 'cognito-test', - payload: { repo_url: 'org/repo' }, + userId: 'u1', + payload: { task_id: 'TASK001' }, blueprintConfig: { compute_type: 'ecs', runtime_arn: '' }, - }); - - const cmd = mockSend.mock.calls[0][0].input.overrides.containerOverrides[0].command; - const src = cmd[2]; - // Reads the URI, fetches via boto3 S3 when set, falls back to inline env. - expect(src).toContain('AGENT_PAYLOAD_S3_URI'); - expect(src).toContain('get_object'); - expect(src).toContain('AGENT_PAYLOAD'); - // ABCA-487: the boot command maps the WHOLE payload via - // run_task_from_payload (not a hand-listed kwarg subset that dropped - // channel_source/channel_metadata → no Linear reactions on ECS). Assert we - // call the mapper and no longer hand-pick the old prompt/model_id kwargs. - expect(src).toContain('run_task_from_payload(p)'); - expect(src).not.toContain('task_description=p.get'); - expect(src).not.toContain('channel_source'); // never hand-listed; the mapper forwards it - }); - - test('deleteEcsPayload deletes the task payload object', async () => { - mockS3Send.mockResolvedValueOnce({}); - const { deleteEcsPayload, ecsPayloadKey } = loadStrategyWithBucket(); - await deleteEcsPayload('TASK001'); - expect(mockS3Send).toHaveBeenCalledTimes(1); - const del = mockS3Send.mock.calls[0][0]; - expect(del._type).toBe('DeleteObject'); - expect(del.input.Bucket).toBe(PAYLOAD_BUCKET); - expect(del.input.Key).toBe(ecsPayloadKey('TASK001')); - expect(ecsPayloadKey('TASK001')).toBe('TASK001/payload.json'); - }); - - test('deleteEcsPayload swallows S3 errors (best-effort — lifecycle is the backstop)', async () => { - mockS3Send.mockRejectedValueOnce(new Error('AccessDenied')); - const { deleteEcsPayload } = loadStrategyWithBucket(); - await expect(deleteEcsPayload('TASK001')).resolves.toBeUndefined(); + })).rejects.toThrow('failed [redacted payload URL]'); + expect(mockSend).not.toHaveBeenCalled(); }); }); // #299 ECS_RIGHTSIZED_PLANNING: the planning task def ARN is a module-level -// constant, so set it BEFORE import via isolateModules (mirrors the #502 bucket -// pattern above) — this keeps it out of the inline tests at the top. +// constant, so set it BEFORE import via isolateModules. describe('EcsComputeStrategy read-only planning-def selection (#299 ECS_RIGHTSIZED_PLANNING)', () => { const PLANNING_DEF_ARN = 'arn:aws:ecs:us-east-1:123456789012:task-definition/agent-planning:1'; @@ -548,12 +474,34 @@ describe('EcsComputeStrategy read-only planning-def selection (#299 ECS_RIGHTSIZ }); }); -describe('deleteEcsPayload without ECS_PAYLOAD_BUCKET', () => { +describe('ECS without ECS_PAYLOAD_BUCKET', () => { + let isolated: typeof import('../../../../src/handlers/shared/strategies/ecs-strategy'); + beforeEach(() => { + const bucket = process.env.ECS_PAYLOAD_BUCKET; + delete process.env.ECS_PAYLOAD_BUCKET; + try { + jest.isolateModules(() => { + // eslint-disable-next-line @typescript-eslint/no-require-imports + isolated = require('../../../../src/handlers/shared/strategies/ecs-strategy'); + }); + } finally { + process.env.ECS_PAYLOAD_BUCKET = bucket; + } + }); + test('no-ops when no payload bucket is configured', async () => { - // The top-of-file import has no ECS_PAYLOAD_BUCKET set. - // eslint-disable-next-line @typescript-eslint/no-require-imports - const { deleteEcsPayload } = require('../../../../src/handlers/shared/strategies/ecs-strategy'); - await expect(deleteEcsPayload('TASK001')).resolves.toBeUndefined(); - expect(mockS3Send).not.toHaveBeenCalled(); + await expect(isolated.deleteEcsPayload('TASK001')).resolves.toBeUndefined(); + expect(mockDelete).not.toHaveBeenCalled(); + }); + + test('rejects a launch before preparing or running anything', async () => { + await expect(new isolated.EcsComputeStrategy().startSession({ + taskId: 'TASK001', + userId: 'u1', + payload: { task_id: 'TASK001' }, + blueprintConfig: { compute_type: 'ecs', runtime_arn: '' }, + })).rejects.toThrow('ECS_PAYLOAD_BUCKET is required'); + expect(mockPrepare).not.toHaveBeenCalled(); + expect(mockSend).not.toHaveBeenCalled(); }); }); diff --git a/cdk/test/handlers/shared/strategies/lambda-microvm-strategy.test.ts b/cdk/test/handlers/shared/strategies/lambda-microvm-strategy.test.ts index 9d94f7cae..bcef69013 100644 --- a/cdk/test/handlers/shared/strategies/lambda-microvm-strategy.test.ts +++ b/cdk/test/handlers/shared/strategies/lambda-microvm-strategy.test.ts @@ -32,11 +32,9 @@ const MICROVM_ID = 'mvm-0123456789abcdef'; const ENDPOINT = 'https://mvm-0123456789abcdef.microvm.lambda.us-east-1.amazonaws.com'; // --- platform_config (ADR-021 P2) --- -// The FOUR required identifiers, and only those, are set for the main describes, -// so the default `platform_config` block is small and its exact serialized size is -// known — which the 4 KB boundary probes below depend on. The nine optional keys -// get their own describe (and are deleted here so a leaked env var from another -// suite cannot silently change the envelope's byte length). +// Set only the required platform fields for the main cases. Optional fields +// are tested separately so ambient environment cannot change payload size. +const APPROVAL_REQUESTS_API_URL = 'https://approval.execute-api.us-east-1.amazonaws.com/v1'; const TASK_TABLE_NAME = 'abca-task-table'; const TASK_EVENTS_TABLE_NAME = 'abca-task-events-table'; const GITHUB_TOKEN_SECRET_ARN = @@ -58,6 +56,7 @@ process.env.MICROVM_PAYLOAD_BUCKET = PAYLOAD_BUCKET; process.env.AWS_REGION = 'us-east-1'; delete process.env.MICROVM_INGRESS_CONNECTOR_ARNS; +process.env.APPROVAL_REQUESTS_API_URL = APPROVAL_REQUESTS_API_URL; process.env.TASK_TABLE_NAME = TASK_TABLE_NAME; process.env.TASK_EVENTS_TABLE_NAME = TASK_EVENTS_TABLE_NAME; process.env.GITHUB_TOKEN_SECRET_ARN = GITHUB_TOKEN_SECRET_ARN; @@ -78,10 +77,20 @@ for (const optional of [ } const mockSend = jest.fn(); +const mockClaimStart = jest.fn(); +const mockSaveHandle = jest.fn(); +const mockSaveCapability = jest.fn(); +jest.mock('../../../../src/handlers/shared/microvm-start', () => ({ + ...jest.requireActual('../../../../src/handlers/shared/microvm-start'), + claimMicrovmStart: (...args: unknown[]) => mockClaimStart(...args), + saveMicrovmStartHandle: (...args: unknown[]) => mockSaveHandle(...args), + saveMicrovmImageCapability: (...args: unknown[]) => mockSaveCapability(...args), +})); jest.mock('@aws-sdk/client-lambda-microvms', () => ({ LambdaMicrovmsClient: jest.fn(() => ({ send: mockSend })), RunMicrovmCommand: jest.fn((input: unknown) => ({ _type: 'RunMicrovm', input })), GetMicrovmCommand: jest.fn((input: unknown) => ({ _type: 'GetMicrovm', input })), + GetMicrovmImageVersionCommand: jest.fn((input: unknown) => ({ _type: 'GetMicrovmImageVersion', input })), TerminateMicrovmCommand: jest.fn((input: unknown) => ({ _type: 'TerminateMicrovm', input })), // Mirrors the real SDK's const-object enum so the strategy's switch keys on // the same literals the service returns. @@ -95,11 +104,12 @@ jest.mock('@aws-sdk/client-lambda-microvms', () => ({ }, })); -const mockS3Send = jest.fn(); -jest.mock('@aws-sdk/client-s3', () => ({ - S3Client: jest.fn(() => ({ send: mockS3Send })), - PutObjectCommand: jest.fn((input: unknown) => ({ _type: 'PutObject', input })), - DeleteObjectCommand: jest.fn((input: unknown) => ({ _type: 'DeleteObject', input })), +const mockPrepare = jest.fn(); +const mockDelete = jest.fn(); +jest.mock('../../../../src/handlers/shared/payload-bootstrap', () => ({ + ...jest.requireActual('../../../../src/handlers/shared/payload-bootstrap'), + preparePayloadReference: (...args: unknown[]) => mockPrepare(...args), + deletePayloadReference: (...args: unknown[]) => mockDelete(...args), })); // The real logger writes JSON to process.stdout/stderr, so level assertions need @@ -123,7 +133,6 @@ import { buildMicrovmPlatformConfig, deleteMicrovmPayload, microvmNoIngressConnectorArnForRegion, - microvmPayloadKey, } from '../../../../src/handlers/shared/strategies/lambda-microvm-strategy'; const BLUEPRINT: BlueprintConfig = { compute_type: 'lambda-microvm', runtime_arn: '' }; @@ -135,35 +144,11 @@ const BLUEPRINT: BlueprintConfig = { compute_type: 'lambda-microvm', runtime_arn const EXPECTED_PLATFORM_CONFIG = { task_table_name: TASK_TABLE_NAME, task_events_table_name: TASK_EVENTS_TABLE_NAME, + approval_requests_api_url: APPROVAL_REQUESTS_API_URL, github_token_secret_arn: GITHUB_TOKEN_SECRET_ARN, agent_session_role_arn: AGENT_SESSION_ROLE_ARN, }; -/** - * Build a payload whose serialized `{"agent_payload": …, "platform_config": …}` - * envelope is EXACTLY `targetBytes` long, so the 4 KB boundary can be probed on - * both sides. - * - * `platform_config` is part of the counted envelope (ADR-021 P2), so its bytes are - * subtracted from the payload's budget here. Asserts its own arithmetic — if the - * envelope shape or the platform block ever changes, this fails loudly rather than - * silently testing the wrong boundary. - */ -function payloadWithEnvelopeBytes(targetBytes: number): Record { - const overhead = Buffer.byteLength( - JSON.stringify({ agent_payload: { p: '' }, platform_config: EXPECTED_PLATFORM_CONFIG }), - 'utf8', - ); - const payload = { p: 'x'.repeat(targetBytes - overhead) }; - expect( - Buffer.byteLength( - JSON.stringify({ agent_payload: payload, platform_config: EXPECTED_PLATFORM_CONFIG }), - 'utf8', - ), - ).toBe(targetBytes); - return payload; -} - function runMicrovmOk() { mockSend.mockResolvedValueOnce({ microvmId: MICROVM_ID, @@ -251,22 +236,95 @@ async function withoutEnvAsync(keys: string[], body: () => Promise): Promi } } -/** {@link withoutEnvAsync}'s inverse: run `body` with extra env vars SET. */ -async function withEnvAsync(env: Record, body: () => Promise): Promise { - const saved = Object.fromEntries(Object.keys(env).map(key => [key, process.env[key]])); - try { - Object.assign(process.env, env); - await body(); - } finally { - for (const [key, value] of Object.entries(saved)) { - if (value === undefined) delete process.env[key]; - else process.env[key] = value; - } - } -} - beforeEach(() => { jest.clearAllMocks(); + mockSend.mockReset(); + mockPrepare.mockReset().mockImplementation(async ({ taskId }: { taskId: string }) => ({ version: 2, task_id: taskId, bootstrap_s3_uri: 's3://b/bootstrap/example.json', payload_url: 'https://signed.example/task', expires_at: Date.now()+900000 })); + mockDelete.mockResolvedValue(undefined); + mockClaimStart.mockReset().mockImplementation(async (taskId: string) => ({ clientToken: taskId, closed: false })); + mockSaveHandle.mockReset().mockResolvedValue(undefined); + mockSaveCapability.mockReset().mockResolvedValue(undefined); +}); + +describe('per-worker image capability', () => { + const input = { + taskId: 'CAP001', + userId: 'cognito-test', + payload: { task_id: 'CAP001' }, + blueprintConfig: BLUEPRINT, + }; + const identity = { imageArn: IMAGE_IDENTIFIER, imageVersion: 'actual-3.0' }; + const version = () => ({ + ...identity, + environmentVariables: { + [sharedConstants.microvm_lifecycle.image_protocol_env]: String(sharedConstants.microvm_lifecycle.protocol_version), + }, + hooks: { + port: sharedConstants.microvm_lifecycle.hook_port, + microvmImageHooks: { ready: 'ENABLED', validate: 'ENABLED' }, + microvmHooks: { + run: 'ENABLED', + terminate: 'ENABLED', + suspend: 'ENABLED', + resume: 'ENABLED', + suspendTimeoutInSeconds: sharedConstants.microvm_hook_budgets.lifecycle_hook_timeout_seconds, + resumeTimeoutInSeconds: sharedConstants.microvm_hook_budgets.lifecycle_hook_timeout_seconds, + }, + }, + }); + + test('saves the known worker before querying its actual image and conditionally enriching it', async () => { + mockSend.mockResolvedValueOnce({ ...makeHandle(), ...identity }) + .mockImplementationOnce(async (command, options) => { + expect(mockSaveHandle).toHaveBeenCalledWith(input.taskId, input.taskId, { ...makeHandle(), ...identity }); + expect(command).toEqual({ + _type: 'GetMicrovmImageVersion', + input: { imageIdentifier: IMAGE_IDENTIFIER, imageVersion: identity.imageVersion }, + }); + expect(options.abortSignal).toBeInstanceOf(AbortSignal); + return version(); + }); + const result = await new LambdaMicrovmComputeStrategy().startSession(input); + const expected = { ...makeHandle(), ...identity, lifecycleProtocol: '1' }; + expect(result).toEqual(expected); + expect(mockSaveCapability).toHaveBeenCalledWith(input.taskId, input.taskId, expected); + expect(mockSaveHandle.mock.invocationCallOrder[0]).toBeLessThan(mockSaveCapability.mock.invocationCallOrder[0]); + }); + + test.each(['lookup', 'marker', 'persistence'])('failed %s keeps the saved worker usable without enabling sleep', async failure => { + mockSend.mockResolvedValueOnce({ ...makeHandle(), ...identity, lifecycleProtocol: '1' }); + if (failure === 'lookup') { + mockSend.mockRejectedValueOnce(Object.assign(new Error('do-not-log-secret'), { name: 'AbortError' })); + } else { + const response = version(); + if (failure === 'marker') response.environmentVariables = {}; + mockSend.mockResolvedValueOnce(response); + if (failure === 'persistence') mockSaveCapability.mockRejectedValueOnce(new Error('do-not-log-secret')); + } + expect(await new LambdaMicrovmComputeStrategy().startSession(input)).toEqual({ ...makeHandle(), ...identity }); + expect(mockSaveHandle.mock.calls[0][2]).not.toHaveProperty('lifecycleProtocol'); + expect(mockSend.mock.calls.filter(([command]) => command._type === 'RunMicrovm')).toHaveLength(1); + expect(mockSend.mock.calls.filter(([command]) => command._type === 'TerminateMicrovm')).toHaveLength(0); + expect(JSON.stringify(mockLogger.warn.mock.calls)).not.toContain('do-not-log-secret'); + }); + + test.each([{}, { imageArn: IMAGE_IDENTIFIER }, { ...identity, imageArn: 'arn:other-image' }])( + 'missing or mismatched returned identity cannot borrow deployment capability: %j', async returned => { + mockSend.mockResolvedValueOnce({ ...makeHandle(), ...returned }); + const result = await new LambdaMicrovmComputeStrategy().startSession(input); + expect(result).not.toHaveProperty('lifecycleProtocol'); + expect(mockSend).toHaveBeenCalledTimes(1); + expect(mockSaveCapability).not.toHaveBeenCalled(); + }, + ); + + test.each([undefined, '1'])('replay preserves saved capability %s without querying the current deployment', async lifecycleProtocol => { + const saved = lifecycleProtocol ? { ...makeHandle(), ...identity, lifecycleProtocol } : makeHandle(); + mockClaimStart.mockResolvedValue({ clientToken: input.taskId, handle: saved, closed: false }); + expect(await new LambdaMicrovmComputeStrategy().startSession(input)).toEqual(saved); + expect(mockSend).not.toHaveBeenCalled(); + expect(mockSaveCapability).not.toHaveBeenCalled(); + }); }); describe('LambdaMicrovmComputeStrategy', () => { @@ -285,7 +343,7 @@ describe('LambdaMicrovmComputeStrategy', () => { blueprintConfig: BLUEPRINT, }); - expect(mockSend).toHaveBeenCalledTimes(1); + expect(mockSend).toHaveBeenCalledTimes(2); const call = mockSend.mock.calls[0][0]; expect(call._type).toBe('RunMicrovm'); expect(call.input.imageIdentifier).toBe(IMAGE_IDENTIFIER); @@ -298,6 +356,8 @@ describe('LambdaMicrovmComputeStrategy', () => { strategyType: 'lambda-microvm', microvmId: MICROVM_ID, endpoint: ENDPOINT, + imageArn: IMAGE_IDENTIFIER, + imageVersion: IMAGE_VERSION, }); }); @@ -414,142 +474,17 @@ describe('LambdaMicrovmComputeStrategy', () => { ]); }); - test('inlines a small payload in runHookPayload and never touches S3', async () => { - runMicrovmOk(); - - await new LambdaMicrovmComputeStrategy().startSession({ - taskId: 'TASK001', - userId: 'cognito-test', - payload: { repo_url: 'org/repo', prompt: 'Fix the bug', max_turns: 50 }, - blueprintConfig: BLUEPRINT, - }); - - expect(mockS3Send).not.toHaveBeenCalled(); - const envelope = JSON.parse(mockSend.mock.calls[0][0].input.runHookPayload); - expect(envelope.agent_payload).toEqual({ repo_url: 'org/repo', prompt: 'Fix the bug', max_turns: 50 }); - expect(envelope.agent_payload_s3_uri).toBeUndefined(); - // The MicroVM's substitute for the env block the other two backends get at - // deploy time — a snapshot must not bake it in (ADR-021 sub-decision 3), so - // it rides the envelope alongside the payload. - expect(envelope.platform_config).toEqual(EXPECTED_PLATFORM_CONFIG); - }); - - test('uploads an oversized payload to S3 and inlines only the pointer', async () => { - mockS3Send.mockResolvedValueOnce({}); - runMicrovmOk(); - - const big = { repo_url: 'org/repo', hydrated_context: { blob: 'x'.repeat(20_000) } }; - await new LambdaMicrovmComputeStrategy().startSession({ - taskId: 'TASK001', - userId: 'cognito-test', - payload: big, - blueprintConfig: BLUEPRINT, - }); - - // Same key shape as the ECS payload bucket: /payload.json - expect(mockS3Send).toHaveBeenCalledTimes(1); - const put = mockS3Send.mock.calls[0][0]; - expect(put._type).toBe('PutObject'); - expect(put.input.Bucket).toBe(PAYLOAD_BUCKET); - expect(put.input.Key).toBe('TASK001/payload.json'); - expect(put.input.ContentType).toBe('application/json'); - // The S3 object carries the payload with platform_config merged in at the top - // level, so an agent that resolves the pointer gets the config with it. - expect(JSON.parse(put.input.Body)).toEqual({ ...big, platform_config: EXPECTED_PLATFORM_CONFIG }); - - const runHookPayload = mockSend.mock.calls[0][0].input.runHookPayload; - const envelope = JSON.parse(runHookPayload); - expect(envelope.agent_payload_s3_uri).toBe(`s3://${PAYLOAD_BUCKET}/TASK001/payload.json`); - expect(envelope.agent_payload).toBeUndefined(); - // ...and ALSO inline on the pointer envelope, deliberately duplicated: the - // agent must be able to read its platform configuration whether it takes it - // off the hook body before fetching S3 or out of the fetched object. - expect(envelope.platform_config).toEqual(EXPECTED_PLATFORM_CONFIG); - // The whole point: the hook body must sit far under the 4 KB cap. - expect(Buffer.byteLength(runHookPayload, 'utf8')).toBeLessThan(MICROVM_RUN_HOOK_PAYLOAD_LIMIT_BYTES); - }); - - test('the inline/S3 branch point IS the 4096-byte service cap, with no headroom', () => { - // Live-measured, NOT read off the SDK docs (which say 16,384): the service - // rejects 4097 with "Member must have length less than or equal to 4096". - // Unlike ECS (whose 8192-byte cap is shared with env vars + command, so it - // needs a margin), runHookPayload is the entire counted string. - expect(MICROVM_RUN_HOOK_PAYLOAD_LIMIT_BYTES).toBe(4_096); - }); - - test('a payload whose envelope is EXACTLY 4096 bytes stays inline', async () => { - runMicrovmOk(); - - await new LambdaMicrovmComputeStrategy().startSession({ - taskId: 'TASK001', - userId: 'cognito-test', - payload: payloadWithEnvelopeBytes(4_096), - blueprintConfig: BLUEPRINT, - }); - - // Boundary is `<=`: 4096 passed the service's length validation live. - expect(mockS3Send).not.toHaveBeenCalled(); - const runHookPayload = mockSend.mock.calls[0][0].input.runHookPayload; - expect(Buffer.byteLength(runHookPayload, 'utf8')).toBe(4_096); - expect(JSON.parse(runHookPayload).agent_payload).toBeDefined(); - }); - - test('a payload whose envelope is 4097 bytes — one over — goes to S3', async () => { - mockS3Send.mockResolvedValueOnce({}); - runMicrovmOk(); - - await new LambdaMicrovmComputeStrategy().startSession({ - taskId: 'TASK001', - userId: 'cognito-test', - payload: payloadWithEnvelopeBytes(4_097), - blueprintConfig: BLUEPRINT, - }); - - expect(mockS3Send).toHaveBeenCalledTimes(1); - const envelope = JSON.parse(mockSend.mock.calls[0][0].input.runHookPayload); - expect(envelope.agent_payload_s3_uri).toBe(`s3://${PAYLOAD_BUCKET}/TASK001/payload.json`); - expect(envelope.agent_payload).toBeUndefined(); - }); - - test('a mid-sized envelope the SDK-documented 16KB cap would have inlined goes to S3', async () => { - mockS3Send.mockResolvedValueOnce({}); - runMicrovmOk(); - - // Regression guard for the live-verification fix: anything from 4,097 to - // 16,384 bytes used to be inlined and would be REJECTED by the service. - await new LambdaMicrovmComputeStrategy().startSession({ - taskId: 'TASK001', - userId: 'cognito-test', - payload: payloadWithEnvelopeBytes(13_000), - blueprintConfig: BLUEPRINT, - }); - - expect(mockS3Send).toHaveBeenCalledTimes(1); - const envelope = JSON.parse(mockSend.mock.calls[0][0].input.runHookPayload); - expect(envelope.agent_payload_s3_uri).toBeDefined(); + test.each(['small', 'x'.repeat(20000)])('delivers a persisted v2 reference for every payload size', async prompt => { + mockSend.mockResolvedValueOnce({ microvmId: MICROVM_ID, endpoint: ENDPOINT }); + await new LambdaMicrovmComputeStrategy().startSession({ taskId: 'TASK001', userId: 'u1', payload: { prompt }, blueprintConfig: BLUEPRINT }); + expect(mockPrepare).toHaveBeenCalledWith(expect.objectContaining({ bucket: PAYLOAD_BUCKET, backend: 'lambda-microvm', platformConfig: EXPECTED_PLATFORM_CONFIG, payload: { prompt } })); + const envelope=JSON.parse(mockSend.mock.calls[0][0].input.runHookPayload); + expect(envelope.version).toBe(2); + expect(envelope.task_id).toBe('TASK001'); + expect(envelope.platform_config).toBeUndefined(); expect(envelope.agent_payload).toBeUndefined(); }); - test('measures the serialized envelope in BYTES, so a multi-byte payload still goes to S3', async () => { - mockS3Send.mockResolvedValueOnce({}); - runMicrovmOk(); - - // 3-byte UTF-8 characters: 2000 chars is ~6 KB of bytes but only 2 KB of - // chars, so measuring String.length would have wrongly inlined this. - const payload = { prompt: '\u4f60'.repeat(2_000) }; - expect(JSON.stringify({ agent_payload: payload, platform_config: EXPECTED_PLATFORM_CONFIG }).length) - .toBeLessThan(MICROVM_RUN_HOOK_PAYLOAD_LIMIT_BYTES); - await new LambdaMicrovmComputeStrategy().startSession({ - taskId: 'TASK001', - userId: 'cognito-test', - payload, - blueprintConfig: BLUEPRINT, - }); - - expect(mockS3Send).toHaveBeenCalledTimes(1); - expect(JSON.parse(mockSend.mock.calls[0][0].input.runHookPayload).agent_payload_s3_uri).toBeDefined(); - }); - test('throws when RunMicrovm returns no microvmId', async () => { mockSend.mockResolvedValueOnce({ endpoint: ENDPOINT, state: 'PENDING' }); @@ -646,34 +581,6 @@ describe('LambdaMicrovmComputeStrategy', () => { expect(mockSend.mock.calls.filter(c => c[0]._type === 'TerminateMicrovm')).toHaveLength(0); }); - test('platform_config COUNTS toward the 4 KB boundary — an otherwise-inlineable payload goes to S3', async () => { - // The regression this locks: measuring `{agent_payload}` alone and then - // sending `{agent_payload, platform_config}` would inline an envelope the - // service rejects outright. So the branch decision has to be made on the - // FULL envelope. This payload is exactly 4 096 bytes WITHOUT the platform - // block — i.e. the old code would have inlined it — and must now upload. - mockS3Send.mockResolvedValueOnce({}); - runMicrovmOk(); - - const overheadWithoutConfig = Buffer.byteLength(JSON.stringify({ agent_payload: { p: '' } }), 'utf8'); - const payload = { p: 'x'.repeat(MICROVM_RUN_HOOK_PAYLOAD_LIMIT_BYTES - overheadWithoutConfig) }; - expect(Buffer.byteLength(JSON.stringify({ agent_payload: payload }), 'utf8')) - .toBe(MICROVM_RUN_HOOK_PAYLOAD_LIMIT_BYTES); - - await new LambdaMicrovmComputeStrategy().startSession({ - taskId: 'TASK001', - userId: 'cognito-test', - payload, - blueprintConfig: BLUEPRINT, - }); - - expect(mockS3Send).toHaveBeenCalledTimes(1); - const runHookPayload = mockSend.mock.calls[0][0].input.runHookPayload; - expect(JSON.parse(runHookPayload).agent_payload_s3_uri).toBeDefined(); - expect(Buffer.byteLength(runHookPayload, 'utf8')) - .toBeLessThanOrEqual(MICROVM_RUN_HOOK_PAYLOAD_LIMIT_BYTES); - }); - test.each(MICROVM_PLATFORM_CONFIG_REQUIRED_KEYS.map(key => [key]))( 'refuses to start — before any AWS call — when required platform config %s is missing', async (key) => { @@ -685,14 +592,14 @@ describe('LambdaMicrovmComputeStrategy', () => { userId: 'cognito-test', // Oversized on purpose: the guard must fire before the payload upload, // or a misconfiguration leaves orphan objects in the payload bucket. - payload: payloadWithEnvelopeBytes(20_000), + payload: { prompt: 'x'.repeat(20_000) }, blueprintConfig: BLUEPRINT, }); await expect(start).rejects.toThrow(new RegExp(`${key} <- ${envVar}`)); await expect(start).rejects.toThrow(/redeploy the stack/); expect(mockSend).not.toHaveBeenCalled(); - expect(mockS3Send).not.toHaveBeenCalled(); + expect(mockPrepare).not.toHaveBeenCalled(); }); }, ); @@ -715,72 +622,52 @@ describe('LambdaMicrovmComputeStrategy', () => { expect(JSON.stringify(started[1])).not.toContain(GITHUB_TOKEN_SECRET_ARN); }); - test('fails BEFORE the upload when even the POINTER envelope cannot fit (no orphan object)', async () => { - // The one shape with no smaller fallback: the payload has already been moved - // to S3, so if `{pointer + platform_config}` still exceeds 4 096 bytes there - // is nothing left to shed. Only a pathological identifier length can cause it - // — hence the check, and hence its placement BEFORE the PutObject so a - // misconfiguration cannot leave objects behind for the lifecycle rule to reap. - await withEnvAsync({ LOG_GROUP_NAME: `/aws/${'x'.repeat(5_000)}` }, async () => { - const start = new LambdaMicrovmComputeStrategy().startSession({ - taskId: 'TASK001', - userId: 'cognito-test', - payload: payloadWithEnvelopeBytes(20_000), - blueprintConfig: BLUEPRINT, - }); + test('rejects a saved reference exceeding the service cap before RunMicrovm', async () => { + mockPrepare.mockResolvedValueOnce({ version: 2, task_id: 'TASK001', payload_url: 'x'.repeat(MICROVM_RUN_HOOK_PAYLOAD_LIMIT_BYTES + 1) }); + await expect(new LambdaMicrovmComputeStrategy().startSession({ taskId: 'TASK001', userId: 'u1', payload: {}, blueprintConfig: BLUEPRINT })).rejects.toThrow('hook limit'); + expect(mockSend).not.toHaveBeenCalled(); + }); - await expect(start).rejects.toThrow(/pointer envelope is \d+ bytes/); - await expect(start).rejects.toThrow(/shorten the stack name/); - expect(mockS3Send).not.toHaveBeenCalled(); - expect(mockSend).not.toHaveBeenCalled(); + test.each(['bootstrap', 'run'])('redacts signed URLs throughout the %s error chain', async stage => { + const { inspect } = await import('node:util'); + const original = new Error('request failed https://bucket.example/key?X-Amz-Signature=BEARER-SECRET', { + cause: new Error('nested BEARER-SECRET'), }); + if (stage === 'bootstrap') mockPrepare.mockRejectedValueOnce(original); + else mockSend.mockRejectedValueOnce(original); + await new LambdaMicrovmComputeStrategy().startSession({ + taskId: 'TASK001', userId: 'u1', payload: {}, blueprintConfig: BLUEPRINT, + }).catch(error => { + expect(inspect(error, { depth: null })).not.toContain('BEARER-SECRET'); + }); + expect.assertions(1); }); - test('microvmPayloadKey matches the ECS payload key shape', () => { - expect(microvmPayloadKey('TASK001')).toBe('TASK001/payload.json'); + test('keeps the capability out of ordinary logs and start-receipt arguments', async () => { + mockPrepare.mockResolvedValueOnce({ version: 2, task_id: 'TASK001', payload_url: 'BEARER-SECRET' }); + runMicrovmOk(); + await new LambdaMicrovmComputeStrategy().startSession({ + taskId: 'TASK001', userId: 'u1', payload: {}, blueprintConfig: BLUEPRINT, + }); + expect(JSON.stringify([ + mockLogger.info.mock.calls, mockLogger.warn.mock.calls, mockLogger.error.mock.calls, + mockClaimStart.mock.calls, mockSaveHandle.mock.calls, + ])).not.toContain('BEARER-SECRET'); }); }); // --- finalize-time payload delete (review NB3 / ayushtr nit 2) --- // - // Not cosmetic parity with ECS. The execution role's payload-bucket grant is - // `grantRead` on the WHOLE bucket (the guest must read its object before any - // tenant identity exists) and keys are `/payload.json`, so a TTL-only - // reaper left every finished task's HYDRATED PROMPT readable by any concurrently - // running MicroVM — which runs untrusted repo code — for up to ~24 h. - describe('deleteMicrovmPayload — closes the cross-task payload-read window', () => { - test('deletes the task\'s own object from the payload bucket', async () => { + // Finalize revokes the single-object link and removes its private replay + // record; worker credentials already deny direct task-object reads. + describe('deleteMicrovmPayload', () => { + test('delegates cleanup of payload and private launch record', async () => { await deleteMicrovmPayload('TASK001'); - - expect(mockS3Send).toHaveBeenCalledTimes(1); - const call = mockS3Send.mock.calls[0][0]; - expect(call._type).toBe('DeleteObject'); - expect(call.input).toEqual({ - Bucket: PAYLOAD_BUCKET, - Key: microvmPayloadKey('TASK001'), - }); + expect(mockDelete).toHaveBeenCalledWith(PAYLOAD_BUCKET, 'TASK001', undefined); }); - - test('is best-effort — a failed delete never throws', async () => { - mockS3Send.mockRejectedValueOnce(new Error('AccessDenied')); - - await expect(deleteMicrovmPayload('TASK001')).resolves.toBeUndefined(); - // Not silent: the lifecycle rule is the backstop, but an operator must be - // able to see the delete is failing. - expect(mockLogger.warn).toHaveBeenCalledWith( - 'Failed to delete MicroVM payload object (non-fatal)', - expect.objectContaining({ task_id: 'TASK001', error: 'AccessDenied' }), - ); - }); - - test('deletes ONLY the given task\'s key — never a prefix or the bucket', async () => { - // The delete must not become a cleanup that can reach another task's object. - await deleteMicrovmPayload('TASK001'); - - const { input } = mockS3Send.mock.calls[0][0]; - expect(input.Key).toBe('TASK001/payload.json'); - expect(input.Key).not.toContain('*'); - expect(input).not.toHaveProperty('Prefix'); + test('scopes replacement cleanup to its own attempt', async () => { + await deleteMicrovmPayload('TASK001', 'attempt-two'); + expect(mockDelete).toHaveBeenCalledWith(PAYLOAD_BUCKET, 'TASK001', 'attempt-two'); }); }); @@ -796,7 +683,7 @@ describe('LambdaMicrovmComputeStrategy', () => { mockSend.mockResolvedValueOnce({ microvmId: MICROVM_ID, state }); const result = await new LambdaMicrovmComputeStrategy().pollSession(makeHandle()); - expect(result).toEqual({ status: expected }); + expect(result).toEqual({ status: expected, microvmState: state }); }); test('sends GetMicrovm keyed on microvmIdentifier', async () => { @@ -815,7 +702,7 @@ describe('LambdaMicrovmComputeStrategy', () => { // No task id, no DDB read — the strategy cannot see the task row at all, // which is exactly why the health rules live in the orchestrator. const result = await new LambdaMicrovmComputeStrategy().pollSession(makeHandle()); - expect(result).toEqual({ status: 'suspended' }); + expect(result).toEqual({ status: 'suspended', microvmState: 'SUSPENDED' }); }); test('treats ResourceNotFoundException as completed (a reaped MicroVM is gone, not broken)', async () => { @@ -824,7 +711,7 @@ describe('LambdaMicrovmComputeStrategy', () => { mockSend.mockRejectedValueOnce(err); const result = await new LambdaMicrovmComputeStrategy().pollSession(makeHandle()); - expect(result).toEqual({ status: 'completed' }); + expect(result).toEqual({ status: 'completed', microvmState: 'NOT_FOUND' }); }); test('rethrows non-NotFound errors so the caller can count poll failures', async () => { @@ -838,7 +725,7 @@ describe('LambdaMicrovmComputeStrategy', () => { mockSend.mockResolvedValueOnce({ microvmId: MICROVM_ID, state: 'HIBERNATING_SOMEDAY' }); const result = await new LambdaMicrovmComputeStrategy().pollSession(makeHandle()); - expect(result).toEqual({ status: 'running' }); + expect(result).toEqual({ status: 'running', microvmState: 'UNKNOWN' }); }); test('throws when the handle is not a lambda-microvm handle', async () => { @@ -865,7 +752,7 @@ describe('LambdaMicrovmComputeStrategy', () => { const result = await new LambdaMicrovmComputeStrategy().pollSession(makeHandle()); - expect(result).toEqual({ status: 'completed', reason }); + expect(result).toEqual({ status: 'completed', microvmState: 'TERMINATED', reason }); }); test('logs a WARNING when a terminal MicroVM carries a reason', async () => { @@ -889,7 +776,7 @@ describe('LambdaMicrovmComputeStrategy', () => { const result = await new LambdaMicrovmComputeStrategy().pollSession(makeHandle()); - expect(result).toEqual({ status: 'completed' }); + expect(result).toEqual({ status: 'completed', microvmState: 'TERMINATED' }); expect(mockLogger.warn).not.toHaveBeenCalled(); }); @@ -901,7 +788,7 @@ describe('LambdaMicrovmComputeStrategy', () => { const result = await new LambdaMicrovmComputeStrategy().pollSession(makeHandle()); - expect(result).toEqual({ status: 'running' }); + expect(result).toEqual({ status: 'running', microvmState: 'RUNNING' }); expect('reason' in result).toBe(false); }); @@ -914,7 +801,7 @@ describe('LambdaMicrovmComputeStrategy', () => { const result = await new LambdaMicrovmComputeStrategy().pollSession(makeHandle()); - expect(result).toEqual({ status: expected, reason: 'because' }); + expect(result).toEqual({ status: expected, microvmState: state === 'HIBERNATING_SOMEDAY' ? 'UNKNOWN' : state, reason: 'because' }); }); }); @@ -930,9 +817,24 @@ describe('LambdaMicrovmComputeStrategy', () => { expect(call.input).toEqual({ microvmIdentifier: MICROVM_ID }); }); + test.each(['TERMINATED', 'TERMINATING', 'RUNNING'])('confirms termination conflicts against %s', async state => { + mockSend.mockRejectedValueOnce(Object.assign(new Error('conflict'), { name: 'ConflictException' })) + .mockResolvedValueOnce({ state }); + const result = await new LambdaMicrovmComputeStrategy().stopSession(makeHandle()); + expect(result).toEqual(state === 'TERMINATED' ? { outcome: 'terminated' } + : { outcome: 'unconfirmed', error_type: 'ConflictException' }); + expect(mockSend.mock.calls[1][0]._type).toBe('GetMicrovm'); + }); + + test('confirms an absent worker after a termination conflict', async () => { + mockSend.mockRejectedValueOnce(Object.assign(new Error('conflict'), { name: 'ConflictException' })) + .mockRejectedValueOnce(Object.assign(new Error('gone'), { name: 'ResourceNotFoundException' })); + await expect(new LambdaMicrovmComputeStrategy().stopSession(makeHandle())).resolves.toEqual({ outcome: 'not-found' }); + }); + test.each([ ['ResourceNotFoundException', 'info'], - ['ConflictException', 'info'], + ['ConflictException', 'warn'], ['ThrottlingException', 'error'], ['AccessDeniedException', 'error'], ['InternalServerException', 'warn'], @@ -941,7 +843,10 @@ describe('LambdaMicrovmComputeStrategy', () => { err.name = errName; mockSend.mockRejectedValueOnce(err); - await expect(new LambdaMicrovmComputeStrategy().stopSession(makeHandle())).resolves.toBeUndefined(); + await expect(new LambdaMicrovmComputeStrategy().stopSession(makeHandle())).resolves.toEqual( + errName === 'ResourceNotFoundException' ? { outcome: 'not-found' } + : { outcome: 'unconfirmed', error_type: errName }, + ); const byLevel: Record = { info: mockLogger.info, @@ -995,7 +900,7 @@ describe('LambdaMicrovmComputeStrategy', () => { await expect(start).rejects.toThrow('MicroVM RunMicrovm failed: ThrottlingException: Rate exceeded'); }); - test('preserves the original error as `cause` so err.name stays inspectable', async () => { + test('preserves the AWS error name in a sanitized cause', async () => { const err = new Error('quota'); err.name = 'ServiceQuotaExceededException'; mockSend.mockRejectedValueOnce(err); @@ -1003,7 +908,7 @@ describe('LambdaMicrovmComputeStrategy', () => { await new LambdaMicrovmComputeStrategy() .startSession({ taskId: 'TASK001', userId: 'u', payload: {}, blueprintConfig: BLUEPRINT }) .catch((thrown: Error) => { - expect(thrown.cause).toBe(err); + expect(thrown.cause).not.toBe(err); expect((thrown.cause as Error).name).toBe('ServiceQuotaExceededException'); }); expect.assertions(2); @@ -1021,16 +926,16 @@ describe('LambdaMicrovmComputeStrategy', () => { test('marks a payload-upload failure so an S3 fault is attributed to this backend', async () => { const err = new Error('Access Denied'); err.name = 'AccessDenied'; - mockS3Send.mockRejectedValueOnce(err); + mockPrepare.mockRejectedValueOnce(err); const start = new LambdaMicrovmComputeStrategy().startSession({ taskId: 'TASK001', userId: 'cognito-test', - payload: payloadWithEnvelopeBytes(20_000), + payload: { prompt: 'x'.repeat(20_000) }, blueprintConfig: BLUEPRINT, }); - await expect(start).rejects.toThrow('MicroVM payload upload failed: AccessDenied: Access Denied'); + await expect(start).rejects.toThrow('MicroVM payload bootstrap failed: AccessDenied: Access Denied'); // Never reaches RunMicrovm — no half-started MicroVM on an upload fault. expect(mockSend).not.toHaveBeenCalled(); }); @@ -1098,7 +1003,7 @@ describe('LambdaMicrovmComputeStrategy without the MicroVM substrate deployed', await expect(start).rejects.toThrow(new RegExp(envVar)); // Fails BEFORE any AWS call — no half-started MicroVM, no orphan S3 object. expect(mockSend).not.toHaveBeenCalled(); - expect(mockS3Send).not.toHaveBeenCalled(); + expect(mockPrepare).not.toHaveBeenCalled(); }); test('MICROVM_IMAGE_VERSION is optional — the field is omitted so the service picks the default', async () => { @@ -1192,7 +1097,7 @@ describe('LambdaMicrovmComputeStrategy image-identifier validation', () => { userId: 'cognito-test', // Oversized on purpose: the guard must fire before the payload upload, or a // misconfiguration leaves orphan objects in the payload bucket. - payload: payloadWithEnvelopeBytes(20_000), + payload: { prompt: 'x'.repeat(20_000) }, blueprintConfig: BLUEPRINT, }); @@ -1202,7 +1107,7 @@ describe('LambdaMicrovmComputeStrategy image-identifier validation', () => { // ...and the remedy. await expect(start).rejects.toThrow(/--context compute_type=lambda-microvm/); expect(mockSend).not.toHaveBeenCalled(); - expect(mockS3Send).not.toHaveBeenCalled(); + expect(mockPrepare).not.toHaveBeenCalled(); }); test('accepts a full image ARN', async () => { @@ -1221,17 +1126,21 @@ describe('LambdaMicrovmComputeStrategy image-identifier validation', () => { }); describe('buildMicrovmPlatformConfig — the MicroVM substitute for a deploy-time env block', () => { - /** A fully-populated orchestrator environment: all thirteen keys present. */ + /** A fully-populated orchestrator environment. */ const FULL_ENV: NodeJS.ProcessEnv = { TASK_TABLE_NAME: 'tasks', TASK_EVENTS_TABLE_NAME: 'events', TASK_APPROVALS_TABLE_NAME: 'approvals', + APPROVAL_REQUESTS_API_URL: 'https://fixture.execute-api.us-east-1.amazonaws.com/v1/', NUDGES_TABLE_NAME: 'nudges', LOG_GROUP_NAME: '/aws/abca/application', ARTIFACTS_BUCKET_NAME: 'artifacts-bucket', TRACE_ARTIFACTS_BUCKET_NAME: 'trace-bucket', + CONTINUATION_BUCKET_NAME: 'continuation-bucket', GITHUB_TOKEN_SECRET_ARN: 'arn:aws:secretsmanager:us-east-1:123456789012:secret:gh-AbCdEf', LINEAR_OAUTH_SECRET_ARN: 'arn:aws:secretsmanager:us-east-1:123456789012:secret:bgagent-linear-oauth-acme-XyZ', + LINEAR_VAULT_ENABLED: 'true', + LINEAR_WORKLOAD_IDENTITY_NAME: 'abca_linear_oauth', JIRA_OAUTH_SECRET_ARN: 'arn:aws:secretsmanager:us-east-1:123456789012:secret:bgagent-jira-oauth-cloud1-XyZ', AGENT_SESSION_ROLE_ARN: 'arn:aws:iam::123456789012:role/SessionRole', AWS_SDK_UA_APP_ID: 'uksb-wt64nei4u6#backgroundagent-dev', @@ -1252,19 +1161,32 @@ describe('buildMicrovmPlatformConfig — the MicroVM substitute for a deploy-tim .toEqual(Object.keys(sharedConstants.microvm_platform_config.env_by_key)); expect([...MICROVM_PLATFORM_CONFIG_REQUIRED_KEYS]) .toEqual(sharedConstants.microvm_platform_config.required); + expect(Object.keys(sharedConstants.microvm_platform_config).sort()) + .toEqual(['account_anchor_key', 'arn_keys', 'env_by_key', 'required']); + expect(sharedConstants.microvm_platform_config.arn_keys).toEqual([ + 'github_token_secret_arn', + 'linear_oauth_secret_arn', + 'jira_oauth_secret_arn', + 'agent_session_role_arn', + ]); + expect(sharedConstants.microvm_platform_config.account_anchor_key).toBe('agent_session_role_arn'); - // Order is part of the contract: it is the serialization order, which the 4 KB - // inline/S3 branch decision is computed against. + // Keep the reviewed key inventory explicit. Payload/bootstrap hashing uses + // canonical JSON, so insertion order does not alter retry identity. expect([...MICROVM_PLATFORM_CONFIG_KEYS]).toEqual([ 'task_table_name', 'task_events_table_name', 'task_approvals_table_name', + 'approval_requests_api_url', 'nudges_table_name', 'log_group_name', 'artifacts_bucket_name', 'trace_artifacts_bucket_name', + 'continuation_bucket_name', 'github_token_secret_arn', 'linear_oauth_secret_arn', + 'linear_vault_enabled', + 'linear_workload_identity_name', 'jira_oauth_secret_arn', 'agent_session_role_arn', 'aws_sdk_ua_app_id', @@ -1277,17 +1199,17 @@ describe('buildMicrovmPlatformConfig — the MicroVM substitute for a deploy-tim // `global`, and wrong in a way that surfaces only as AccessDenied at turn 0. 'anthropic_model', ]); - expect(MICROVM_PLATFORM_CONFIG_KEYS).toHaveLength(14); // snake_case on the wire, matching every other key in the /run envelope. for (const key of MICROVM_PLATFORM_CONFIG_KEYS) { expect(key).toMatch(/^[a-z][a-z0-9_]*$/); } }); - test('pins the REQUIRED subset — these four are what a task cannot start without', () => { + test('pins the required platform fields including the trusted approval service', () => { expect([...MICROVM_PLATFORM_CONFIG_REQUIRED_KEYS]).toEqual([ 'task_table_name', 'task_events_table_name', + 'approval_requests_api_url', 'github_token_secret_arn', 'agent_session_role_arn', ]); @@ -1298,11 +1220,15 @@ describe('buildMicrovmPlatformConfig — the MicroVM substitute for a deploy-tim } }); - test('emits all fourteen keys, in declaration order, from a full environment', () => { + test('emits every configured key in declaration order from a full environment', () => { const config = buildMicrovmPlatformConfig(FULL_ENV); expect(Object.keys(config)).toEqual([...MICROVM_PLATFORM_CONFIG_KEYS]); expect(config.task_table_name).toBe('tasks'); + expect(config.approval_requests_api_url).toBe(FULL_ENV.APPROVAL_REQUESTS_API_URL); expect(config.nudges_table_name).toBe('nudges'); + expect(config.continuation_bucket_name).toBe('continuation-bucket'); + expect(config.linear_vault_enabled).toBe('true'); + expect(config.linear_workload_identity_name).toBe('abca_linear_oauth'); expect(config.agent_session_role_arn).toBe('arn:aws:iam::123456789012:role/SessionRole'); expect(config.anthropic_default_haiku_model).toBe('us.anthropic.claude-haiku-4-5-20251001-v1:0'); // The MAIN model, asserted BY VALUE rather than presence. This is the whole @@ -1316,6 +1242,7 @@ describe('buildMicrovmPlatformConfig — the MicroVM substitute for a deploy-tim const config = buildMicrovmPlatformConfig({ TASK_TABLE_NAME: 'tasks', TASK_EVENTS_TABLE_NAME: 'events', + APPROVAL_REQUESTS_API_URL, GITHUB_TOKEN_SECRET_ARN: 'arn:aws:secretsmanager:us-east-1:123456789012:secret:gh-AbCdEf', AGENT_SESSION_ROLE_ARN: 'arn:aws:iam::123456789012:role/SessionRole', }); @@ -1325,6 +1252,7 @@ describe('buildMicrovmPlatformConfig — the MicroVM substitute for a deploy-tim expect(Object.keys(config)).toEqual([ 'task_table_name', 'task_events_table_name', + 'approval_requests_api_url', 'github_token_secret_arn', 'agent_session_role_arn', ]); @@ -1359,6 +1287,7 @@ describe('buildMicrovmPlatformConfig — the MicroVM substitute for a deploy-tim }); test.each([ + ['APPROVAL_REQUESTS_API_URL', 'approval_requests_api_url'], ['TASK_TABLE_NAME', 'task_table_name'], ['TASK_EVENTS_TABLE_NAME', 'task_events_table_name'], ['GITHUB_TOKEN_SECRET_ARN', 'github_token_secret_arn'], @@ -1376,7 +1305,7 @@ describe('buildMicrovmPlatformConfig — the MicroVM substitute for a deploy-tim test('names EVERY missing required key at once, not just the first', () => { // One redeploy should fix all of them; reporting one per attempt turns a - // misconfiguration into four round-trips. + // misconfiguration into repeated round-trips. expect(() => buildMicrovmPlatformConfig({})).toThrow( /task_table_name.*task_events_table_name.*github_token_secret_arn.*agent_session_role_arn/s, ); @@ -1443,6 +1372,8 @@ describe('buildMicrovmPlatformConfig — the MicroVM substitute for a deploy-tim for (const optional of [ 'TASK_APPROVALS_TABLE_NAME', 'NUDGES_TABLE_NAME', 'LOG_GROUP_NAME', 'ARTIFACTS_BUCKET_NAME', 'TRACE_ARTIFACTS_BUCKET_NAME', 'LINEAR_OAUTH_SECRET_ARN', + 'LINEAR_VAULT_ENABLED', 'LINEAR_WORKLOAD_IDENTITY_NAME', + 'CONTINUATION_BUCKET_NAME', 'JIRA_OAUTH_SECRET_ARN', 'AWS_SDK_UA_APP_ID', 'ANTHROPIC_DEFAULT_HAIKU_MODEL', 'ANTHROPIC_MODEL', ]) { @@ -1459,18 +1390,6 @@ describe('buildMicrovmPlatformConfig — the MicroVM substitute for a deploy-tim const config = buildMicrovmPlatformConfig(); expect(config).toEqual(EXPECTED_PLATFORM_CONFIG); }); - - test('the full thirteen-key block fits the pointer envelope inside the 4 KB cap', () => { - // The one shape with no smaller fallback: if the pointer envelope itself - // exceeded 4 096 bytes there would be nothing left to move to S3. This asserts - // the design has real headroom rather than relying on the guard. - const pointerEnvelope = JSON.stringify({ - agent_payload_s3_uri: `s3://${PAYLOAD_BUCKET}/TASK001/payload.json`, - platform_config: buildMicrovmPlatformConfig(FULL_ENV), - }); - expect(Buffer.byteLength(pointerEnvelope, 'utf8')) - .toBeLessThan(MICROVM_RUN_HOOK_PAYLOAD_LIMIT_BYTES / 2); - }); }); describe('LambdaMicrovmComputeStrategy with the FULL platform_config environment', () => { @@ -1494,21 +1413,9 @@ describe('LambdaMicrovmComputeStrategy with the FULL platform_config environment for (const key of Object.keys(OPTIONAL_ENV)) delete process.env[key]; }); - test('delivers every configured identifier on the wire, in both envelope halves', async () => { - mockS3Send.mockResolvedValueOnce({}); - runMicrovmOk(); - - await new LambdaMicrovmComputeStrategy().startSession({ - taskId: 'TASK001', - userId: 'cognito-test', - payload: { repo_url: 'org/repo', hydrated_context: { blob: 'x'.repeat(10_000) } }, - blueprintConfig: BLUEPRINT, - }); - - const expected = { ...EXPECTED_PLATFORM_CONFIG, ...buildMicrovmPlatformConfig() }; - const envelope = JSON.parse(mockSend.mock.calls[0][0].input.runHookPayload); - expect(envelope.platform_config).toEqual(expected); - expect(Object.keys(envelope.platform_config)).toHaveLength(13); - expect(JSON.parse(mockS3Send.mock.calls[0][0].input.Body).platform_config).toEqual(expected); + test('passes every configured identifier to the authenticated manifest producer', async () => { + mockSend.mockResolvedValueOnce({ microvmId: MICROVM_ID, endpoint: ENDPOINT }); + await new LambdaMicrovmComputeStrategy().startSession({ taskId: 'TASK001', userId: 'u1', payload: {}, blueprintConfig: BLUEPRINT }); + expect(mockPrepare.mock.calls[0][0].platformConfig).toEqual(buildMicrovmPlatformConfig()); }); }); diff --git a/cdk/test/handlers/shared/task-cancellation.test.ts b/cdk/test/handlers/shared/task-cancellation.test.ts new file mode 100644 index 000000000..ff64d364c --- /dev/null +++ b/cdk/test/handlers/shared/task-cancellation.test.ts @@ -0,0 +1,144 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +// Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + +import { GetCommand, TransactWriteCommand, UpdateCommand } from '@aws-sdk/lib-dynamodb'; +import { TaskStatus } from '../../../src/constructs/task-status'; +import { cancelTaskState } from '../../../src/handlers/shared/task-cancellation'; +import type { TaskRecord } from '../../../src/handlers/shared/types'; + +const mockSend = jest.fn(); +jest.mock('../../../src/handlers/shared/ua', () => ({ + makeDocClient: () => ({ send: (...args: unknown[]) => mockSend(...args) }), +})); + +const task = { + task_id: 'task', + user_id: 'owner', + status: TaskStatus.AWAITING_APPROVAL, + awaiting_approval_request_id: 'gate', +} as TaskRecord; +const options = { + userId: 'owner', taskTable: 'Tasks', approvalsTable: 'Approvals', eventsTable: 'Events', retentionDays: 90, +}; +const pending = { status: 'PENDING', user_id: 'owner' }; +const conflict = { name: 'ConditionalCheckFailedException' }; +const transactionConflict = { + name: 'TransactionCanceledException', + CancellationReasons: [{ Code: 'None' }, { Code: 'ConditionalCheckFailed' }, { Code: 'None' }], +}; +beforeEach(() => mockSend.mockReset()); + +test('atomically closes the exact task, pending approval and audit event', async () => { + mockSend.mockResolvedValueOnce({ Item: pending }).mockResolvedValueOnce({}); + const result = await cancelTaskState(task, options); + expect(result.cancelledRequestId).toBe('gate'); + const command = mockSend.mock.calls[1][0] as TransactWriteCommand; + expect(command).toBeInstanceOf(TransactWriteCommand); + const [taskWrite, approvalWrite, eventWrite] = command.input.TransactItems!; + expect(taskWrite.Update).toMatchObject({ + Key: { task_id: 'task' }, + ConditionExpression: expect.stringContaining('awaiting_approval_request_id = :request'), + ExpressionAttributeValues: { ':observed': 'AWAITING_APPROVAL', ':request': 'gate', ':user': 'owner' }, + }); + expect(approvalWrite.Update).toMatchObject({ + Key: { task_id: 'task', request_id: 'gate' }, + ExpressionAttributeValues: { ':pending': 'PENDING', ':cancelled': 'CANCELLED', ':user': 'owner' }, + }); + expect(eventWrite.Put?.Item).toMatchObject({ + event_type: 'approval_cancelled', metadata: { request_id: 'gate', status: 'CANCELLED' }, + }); +}); + +test.each(['APPROVED', 'DENIED', 'TIMED_OUT', 'CANCELLED', 'STRANDED'])( + 'preserves an already %s approval', + async status => { + mockSend.mockResolvedValueOnce({ Item: { ...pending, status } }).mockResolvedValueOnce({}); + expect((await cancelTaskState(task, options)).cancelledRequestId).toBeUndefined(); + expect(mockSend.mock.calls[1][0]).toBeInstanceOf(UpdateCommand); + }, +); + +test('an approval winning the transaction race stays approved while its task is cancelled', async () => { + mockSend.mockResolvedValueOnce({ Item: pending }) + .mockRejectedValueOnce(transactionConflict) + .mockResolvedValueOnce({ Item: task }) + .mockResolvedValueOnce({ Item: { ...pending, status: 'APPROVED' } }) + .mockResolvedValueOnce({}); + await cancelTaskState(task, options); + expect(mockSend.mock.calls.map(([command]) => command.constructor)).toEqual([ + GetCommand, TransactWriteCommand, GetCommand, GetCommand, UpdateCommand, + ]); + expect(mockSend.mock.calls[2][0].input.ConsistentRead).toBe(true); +}); + +test('refreshes a newly created gate rather than cancelling only the old task snapshot', async () => { + const running = { ...task, status: TaskStatus.RUNNING, awaiting_approval_request_id: null }; + mockSend.mockRejectedValueOnce(conflict) + .mockResolvedValueOnce({ Item: task }) + .mockResolvedValueOnce({ Item: pending }) + .mockResolvedValueOnce({}); + const result = await cancelTaskState(running, options); + expect(result.cancelledRequestId).toBe('gate'); + expect(result.task).toEqual(task); + expect(mockSend.mock.calls[0][0].input.ConditionExpression).toContain('attribute_not_exists(awaiting_approval_request_id)'); + expect(mockSend.mock.calls[3][0]).toBeInstanceOf(TransactWriteCommand); +}); + +test.each([ + [undefined, 'missing'], + [{ ...task, status: TaskStatus.COMPLETED }, 'terminal'], + [{ ...task, user_id: 'other' }, 'forbidden'], +])('does not overwrite state after a conflicting cancellation (%s)', async (fresh, reason) => { + mockSend.mockResolvedValueOnce({ Item: pending }) + .mockRejectedValueOnce(transactionConflict) + .mockResolvedValueOnce({ Item: fresh }); + await expect(cancelTaskState(task, options)).rejects.toMatchObject({ reason }); + expect(mockSend).toHaveBeenCalledTimes(3); +}); + +test('does not fall back to a task-only cancellation after an unexpected transaction failure', async () => { + const error = new Error('DynamoDB unavailable'); + mockSend.mockResolvedValueOnce({ Item: pending }).mockRejectedValueOnce(error); + await expect(cancelTaskState(task, options)).rejects.toBe(error); + expect(mockSend).toHaveBeenCalledTimes(2); +}); + +test('bounded conflicts leave a retryable error, never an unconditional update', async () => { + for (let i = 0; i < 3; i++) { + mockSend.mockResolvedValueOnce({ Item: pending }) + .mockRejectedValueOnce(transactionConflict).mockResolvedValueOnce({ Item: task }); + } + await expect(cancelTaskState(task, options)).rejects.toMatchObject({ reason: 'conflict' }); + expect(mockSend).toHaveBeenCalledTimes(9); +}); + +test('a malformed foreign approval does not prevent cancellation or change that approval', async () => { + mockSend.mockResolvedValueOnce({ Item: { ...pending, user_id: 'other' } }).mockResolvedValueOnce({}); + await cancelTaskState(task, options); + expect(mockSend.mock.calls[1][0]).toBeInstanceOf(UpdateCommand); + expect(mockSend.mock.calls[1][0].input.TableName).toBe('Tasks'); +}); + +test('fails before mutating an approval wait when required tables are not wired', async () => { + await expect(cancelTaskState(task, { ...options, approvalsTable: undefined })) + .rejects.toThrow('requires approvals and events tables'); + expect(mockSend).not.toHaveBeenCalled(); +}); diff --git a/cdk/test/handlers/shared/task-concurrency-local.test.ts b/cdk/test/handlers/shared/task-concurrency-local.test.ts new file mode 100644 index 000000000..c5949ae57 --- /dev/null +++ b/cdk/test/handlers/shared/task-concurrency-local.test.ts @@ -0,0 +1,323 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +/** + * Opt-in real DynamoDB Local tests: + * ABCA_DDB_LOCAL_ENDPOINT=http://127.0.0.1: mise run testf -- task-concurrency-local + * The endpoint is restricted to loopback and every client uses dummy credentials. + */ +import { randomUUID } from 'node:crypto'; +import { CreateTableCommand, DeleteTableCommand, DynamoDBClient } from '@aws-sdk/client-dynamodb'; +import { DeleteCommand, DynamoDBDocumentClient, GetCommand, PutCommand, ScanCommand, UpdateCommand } from '@aws-sdk/lib-dynamodb'; + +const endpoint = process.env.ABCA_DDB_LOCAL_ENDPOINT; +if (process.env.CI === 'true' && !endpoint) { + throw new Error('CI requires ABCA_DDB_LOCAL_ENDPOINT; concurrency transaction tests must not skip'); +} +if (endpoint && (new URL(endpoint).hostname !== '127.0.0.1' || new URL(endpoint).protocol !== 'http:')) { + throw new Error('Capacity integration tests require an http://127.0.0.1 DynamoDB Local endpoint'); +} +const mockBeforeSend = jest.fn(); +const mockAfterSend = jest.fn(); +const mockClients: DynamoDBDocumentClient[] = []; +jest.mock('../../../src/handlers/shared/ua', () => { + const actual = jest.requireActual('../../../src/handlers/shared/ua'); + return { + ...actual, + makeDocClient: () => { + const client = actual.makeDocClient({ + endpoint: process.env.ABCA_DDB_LOCAL_ENDPOINT ?? 'http://127.0.0.1:1', + region: 'us-east-1', + credentials: { accessKeyId: 'local', secretAccessKey: 'local' }, + }); + const send = client.send.bind(client); + client.send = async (command: unknown) => { + await mockBeforeSend(command); + const result = await send(command); + await mockAfterSend(command, result); + return result; + }; + mockClients.push(client); + return client; + }, + }; +}); + +const suffix = randomUUID(); +const tasks = `capacity-tasks-${suffix}`; +const counters = `capacity-counters-${suffix}`; +const events = `capacity-events-${suffix}`; +Object.assign(process.env, { + TASK_TABLE_NAME: tasks, + USER_CONCURRENCY_TABLE_NAME: counters, + TASK_EVENTS_TABLE_NAME: events, + MEMORY_ID: '', +}); + +import { handler as reconcile } from '../../../src/handlers/reconcile-concurrency'; +import { failTask, finalizeTask } from '../../../src/handlers/shared/orchestrator'; +import { acquireTaskSlot, releaseTaskSlot } from '../../../src/handlers/shared/task-concurrency'; + +const raw = new DynamoDBClient({ + endpoint: endpoint ?? 'http://127.0.0.1:1', + region: 'us-east-1', + credentials: { accessKeyId: 'local', secretAccessKey: 'local' }, +}); +const admin = DynamoDBDocumentClient.from(raw); +const local = endpoint ? describe : describe.skip; +jest.setTimeout(30_000); + +local('task capacity against DynamoDB Local', () => { + beforeAll(async () => { + for (const [name, key] of [[tasks, 'task_id'], [counters, 'user_id'], [events, 'task_id']]) { + await raw.send(new CreateTableCommand({ + TableName: name, + BillingMode: 'PAY_PER_REQUEST', + AttributeDefinitions: [ + { AttributeName: key, AttributeType: 'S' }, + ...(name === events ? [{ AttributeName: 'event_id', AttributeType: 'S' as const }] : []), + ], + KeySchema: [ + { AttributeName: key, KeyType: 'HASH' }, + ...(name === events ? [{ AttributeName: 'event_id', KeyType: 'RANGE' as const }] : []), + ], + })); + } + }); + beforeEach(async () => { + mockBeforeSend.mockReset(); + mockAfterSend.mockReset(); + for (const [name, key] of [[tasks, 'task_id'], [counters, 'user_id'], [events, 'task_id']]) { + const page = await admin.send(new ScanCommand({ TableName: name })); + for (const item of page.Items ?? []) { + await admin.send(new DeleteCommand({ + TableName: name, + Key: { [key]: item[key], ...(name === events && { event_id: item.event_id }) }, + })); + } + } + }); + afterAll(async () => { + for (const name of [tasks, counters, events]) await raw.send(new DeleteTableCommand({ TableName: name })); + for (const client of mockClients) client.destroy(); + raw.destroy(); + }); + + async function seed(id: string, initialStatus = 'SUBMITTED', user = 'user') { + await admin.send(new PutCommand({ + TableName: tasks, + Item: { + task_id: id, user_id: user, status: initialStatus, memory_written: true, + }, + })); + } + async function status(id: string, value: string) { + await admin.send(new UpdateCommand({ + TableName: tasks, + Key: { task_id: id }, + UpdateExpression: 'SET #status = :status', + ExpressionAttributeNames: { '#status': 'status' }, + ExpressionAttributeValues: { ':status': value }, + })); + } + async function task(id: string) { + return (await admin.send(new GetCommand({ + TableName: tasks, Key: { task_id: id }, ConsistentRead: true, + }))).Item!; + } + async function count() { + return (await admin.send(new GetCommand({ + TableName: counters, Key: { user_id: 'user' }, ConsistentRead: true, + }))).Item?.active_count ?? 0; + } + async function twoSeats() { + await seed('one'); + await seed('two'); + await acquireTaskSlot('one', 'user', 3); + await acquireTaskSlot('two', 'user', 3); + } + + test('crash replay of real finalization preserves the other task reservation', async () => { + await twoSeats(); + await status('one', 'COMPLETED'); + await finalizeTask('one', { attempts: 10 }, 'user'); + await finalizeTask('one', { attempts: 10 }, 'user'); + expect(await count()).toBe(1); + expect((await task('one')).concurrency_slot.state).toBe('released'); + expect((await task('two')).concurrency_slot.state).toBe('held'); + }); + + test('simultaneous admission of one task creates exactly one reservation', async () => { + await seed('one'); + const results = await Promise.allSettled(Array.from({ length: 5 }, () => acquireTaskSlot('one', 'user', 3))); + expect(results.some(result => result.status === 'fulfilled' && result.value)).toBe(true); + expect(await acquireTaskSlot('one', 'user', 3)).toBe(true); + expect(await count()).toBe(1); + }); + + test('different tasks cannot exceed the admission cap', async () => { + const ids = ['a', 'b', 'c', 'd']; + for (const id of ids) await seed(id); + await Promise.allSettled(ids.map(id => acquireTaskSlot(id, 'user', 2))); + // Retry any transaction conflicts serially, as durable execution does. + for (const id of ids) await acquireTaskSlot(id, 'user', 2); + expect(await count()).toBe(2); + // Four fixed fixture reads; no input-dependent fan-out. + // eslint-disable-next-line @cdklabs/promiseall-no-unbounded-parallelism + const rows = await Promise.all([task('a'), task('b'), task('c'), task('d')]); + expect(rows.filter(row => row.concurrency_slot?.state === 'held')).toHaveLength(2); + }); + + test.each(['acquire', 'release'])('a lost committed %s response is recovered from the marker', async (operation) => { + await seed('one'); + if (operation === 'release') { + await acquireTaskSlot('one', 'user', 3); + await status('one', 'FAILED'); + } + mockAfterSend.mockImplementationOnce((command) => { + if (command.constructor.name === 'GetCommand') { + mockAfterSend.mockImplementationOnce(() => { throw new Error('transaction response lost'); }); + } + }); + if (operation === 'acquire') await acquireTaskSlot('one', 'user', 3); + else await releaseTaskSlot('one', 'user'); + expect(await count()).toBe(operation === 'acquire' ? 1 : 0); + expect((await task('one')).concurrency_slot.state).toBe(operation === 'acquire' ? 'held' : 'released'); + }); + + test('competing finalizers and direct cleanup only release once', async () => { + await twoSeats(); + await status('one', 'CANCELLED'); + await Promise.allSettled([ + finalizeTask('one', { attempts: 0 }, 'user'), + finalizeTask('one', { attempts: 0 }, 'user'), + releaseTaskSlot('one', 'user'), + ]); + await releaseTaskSlot('one', 'user'); + expect(await count()).toBe(1); + }); + + test('start failure and replay release a held slot once', async () => { + await twoSeats(); + await status('one', 'HYDRATING'); + await failTask('one', 'HYDRATING', 'start failed', 'user', true); + await failTask('one', 'HYDRATING', 'start failed', 'user', true); + expect(await count()).toBe(1); + expect((await task('one')).status).toBe('FAILED'); + }); + + test('approval waits retain capacity through reconciliation', async () => { + await twoSeats(); + await status('one', 'AWAITING_APPROVAL'); + expect(await releaseTaskSlot('one', 'user')).toBe(false); + await reconcile(); + expect(await count()).toBe(2); + }); + + test('queued and unadmitted terminal tasks cannot release another reservation', async () => { + await seed('running'); + await acquireTaskSlot('running', 'user', 3); + await seed('queued', 'QUEUED'); + await seed('cancelled', 'CANCELLED'); + expect(await acquireTaskSlot('queued', 'user', 3)).toBe(false); + await releaseTaskSlot('queued', 'user'); + await releaseTaskSlot('cancelled', 'user'); + expect(await count()).toBe(1); + }); + + test('owner mismatch changes neither the reservation nor counter', async () => { + await seed('one'); + await expect(acquireTaskSlot('one', 'someone-else', 3)).rejects.toThrow('owner'); + expect(await count()).toBe(0); + expect((await task('one')).concurrency_slot).toBeUndefined(); + }); + + test('cancellation between the read and admission transaction wins', async () => { + await seed('one'); + mockBeforeSend.mockImplementation(async (command) => { + if (command.constructor.name === 'TransactWriteCommand') { + mockBeforeSend.mockReset(); + await status('one', 'CANCELLED'); + } + }); + expect(await acquireTaskSlot('one', 'user', 3)).toBe(false); + expect(await count()).toBe(0); + }); + + test('an empty counter release preserves an admission that races into the empty-counter branch', async () => { + await seed('one'); + await acquireTaskSlot('one', 'user', 3); + await status('one', 'FAILED'); + await admin.send(new DeleteCommand({ TableName: counters, Key: { user_id: 'user' } })); + await seed('two'); + mockBeforeSend.mockImplementation(async (command) => { + if (command.constructor.name === 'TransactWriteCommand' + && command.input.TransactItems[1].Update.UpdateExpression.includes('if_not_exists')) { + mockBeforeSend.mockReset(); + await acquireTaskSlot('two', 'user', 3); + } + }); + await releaseTaskSlot('one', 'user'); + expect(await count()).toBe(1); + expect((await task('one')).concurrency_slot.state).toBe('released'); + expect((await task('two')).concurrency_slot.state).toBe('held'); + }); + + test('periodic cleanup recovers a crash between terminal status and release', async () => { + await twoSeats(); + await status('one', 'TIMED_OUT'); + await reconcile(); + await finalizeTask('one', { attempts: 10 }, 'user'); + expect(await count()).toBe(1); + }); + + test('an ambiguous legacy active task prevents guessing a replacement count', async () => { + await twoSeats(); + await seed('legacy', 'AWAITING_APPROVAL'); + await reconcile(); + expect(await count()).toBe(2); + expect((await task('legacy')).concurrency_slot).toBeUndefined(); + }); + + test('a revision change blocks stale repair even when the counter returns to its original number', async () => { + await seed('one'); + await acquireTaskSlot('one', 'user', 3); + await admin.send(new UpdateCommand({ + TableName: counters, + Key: { user_id: 'user' }, + UpdateExpression: 'SET active_count = :zero', + ExpressionAttributeValues: { ':zero': 0 }, + })); + mockAfterSend.mockImplementation(async (command) => { + if (command.constructor.name === 'ScanCommand' && command.input.TableName === tasks) { + mockAfterSend.mockReset(); + await seed('two'); + await acquireTaskSlot('two', 'user', 3); + await status('one', 'COMPLETED'); + await releaseTaskSlot('one', 'user'); + } + }); + await reconcile(); + // Admission then release brought the count back to zero, but changed its + // revision. The old scan must not write its answer into this new state. + expect(await count()).toBe(0); + await reconcile(); + expect(await count()).toBe(1); + }); +}); diff --git a/cdk/test/handlers/shared/task-concurrency.test.ts b/cdk/test/handlers/shared/task-concurrency.test.ts new file mode 100644 index 000000000..bca32da1e --- /dev/null +++ b/cdk/test/handlers/shared/task-concurrency.test.ts @@ -0,0 +1,246 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +const mockSend = jest.fn(); +jest.mock('node:timers/promises', () => ({ + setTimeout: jest.fn().mockResolvedValue(undefined), +})); +jest.mock('@aws-sdk/lib-dynamodb', () => ({ + GetCommand: jest.fn((input: unknown) => ({ kind: 'get', input })), + TransactWriteCommand: jest.fn((input: unknown) => ({ kind: 'transaction', input })), +})); +jest.mock('../../../src/handlers/shared/ua', () => ({ + makeDocClient: () => ({ send: mockSend }), +})); +process.env.TASK_TABLE_NAME = 'Tasks'; +process.env.USER_CONCURRENCY_TABLE_NAME = 'Counters'; + +import { setTimeout as delay } from 'node:timers/promises'; +import { acquireTaskSlot, releaseTaskSlot } from '../../../src/handlers/shared/task-concurrency'; + +const base = { task_id: 'task', user_id: 'user', status: 'SUBMITTED' }; +const held = { state: 'held', acquired_at: '2026-09-13T00:00:00Z' }; +function cancelled(index: number) { + return Object.assign(new Error('transaction cancelled'), { + name: 'TransactionCanceledException', + CancellationReasons: [0, 1].map(i => ({ Code: i === index ? 'ConditionalCheckFailed' : 'None' })), + }); +} +beforeEach(() => { + mockSend.mockReset(); + jest.mocked(delay).mockClear(); +}); + +function conflict(codes = ['None', 'TransactionConflict']) { + return Object.assign(new Error('transaction conflict'), { + name: 'TransactionCanceledException', + CancellationReasons: codes.map(Code => ({ Code })), + }); +} + +test.each(['acquire', 'release'])('%s retries counter contention with the identical guarded transaction', async operation => { + mockSend.mockResolvedValueOnce({ + Item: operation === 'acquire' ? base : { ...base, status: 'COMPLETED', concurrency_slot: held }, + }).mockRejectedValueOnce(conflict()).mockRejectedValueOnce(conflict()).mockResolvedValueOnce({}); + const result = operation === 'acquire' + ? await acquireTaskSlot('task', 'user', 10) + : await releaseTaskSlot('task', 'user'); + expect(result).toBe(true); + const transactions = mockSend.mock.calls.filter(([command]) => command.kind === 'transaction'); + expect(transactions).toHaveLength(3); + expect(transactions[1][0]).toBe(transactions[0][0]); + expect(transactions[2][0]).toBe(transactions[0][0]); + expect(delay).toHaveBeenCalledTimes(2); +}); + +test('persistent contention stays bounded and propagates the original error', async () => { + const error = conflict(); + mockSend.mockImplementation(async command => { + if (command.kind === 'get') return { Item: { ...base, status: 'COMPLETED', concurrency_slot: held } }; + throw error; + }); + await expect(releaseTaskSlot('task', 'user')).rejects.toBe(error); + expect(mockSend.mock.calls.filter(([command]) => command.kind === 'transaction')).toHaveLength(5); + expect(delay).toHaveBeenCalledTimes(4); + for (const [index, call] of jest.mocked(delay).mock.calls.entries()) { + expect(call[0]).toBeGreaterThanOrEqual(0); + expect(call[0]).toBeLessThan(50 * 2 ** index); + } +}); + +test('a mixed conflict and capacity condition failure follows normal admission without retrying', async () => { + mockSend.mockResolvedValueOnce({ Item: base }) + .mockRejectedValueOnce(conflict(['TransactionConflict', 'ConditionalCheckFailed'])) + .mockResolvedValueOnce({ Item: base }); + expect(await acquireTaskSlot('task', 'user', 10)).toBe(false); + expect(delay).not.toHaveBeenCalled(); +}); + +test('a worker lease changing during contention cannot release its reservation', async () => { + const task = { ...base, status: 'FAILED', concurrency_slot: held, continuation_launch: {}, microvm_start: { clientToken: 'attempt' } }; + const fenced = conflict(['None', 'None', 'ConditionalCheckFailed']); + mockSend.mockResolvedValueOnce({ Item: task }) + .mockResolvedValueOnce({ Item: { lease_user_id: 'user', lease_attempt_id: 'attempt', lease_state: 'CLOSED' } }) + .mockRejectedValueOnce(conflict(['None', 'TransactionConflict', 'None'])) + .mockRejectedValueOnce(fenced) + .mockResolvedValueOnce({ Item: task }); + await expect(releaseTaskSlot('task', 'user')).rejects.toBe(fenced); + const transactions = mockSend.mock.calls.filter(([command]) => command.kind === 'transaction'); + expect(transactions).toHaveLength(2); + expect(transactions[1][0]).toBe(transactions[0][0]); + expect(transactions[1][0].input.TransactItems[2].ConditionCheck.ConditionExpression).toContain('lease_state = :closed'); + expect(delay).toHaveBeenCalledTimes(1); +}); + +test.each([undefined, 'ACTIVE', 'FENCED', 'PARKED'])( + 'a terminal saved task keeps capacity until shutdown is confirmed (lease %s)', async leaseState => { + mockSend.mockResolvedValueOnce({ + Item: { ...base, status: 'FAILED', concurrency_slot: held, continuation_launch: {}, microvm_start: { clientToken: 'attempt' } }, + }).mockResolvedValueOnce({ + Item: { lease_user_id: 'user', lease_attempt_id: 'attempt', lease_state: leaseState }, + }); + expect(await releaseTaskSlot('task', 'user')).toBe(false); + expect(mockSend).toHaveBeenCalledTimes(2); + }, +); + +test('confirmed shutdown is checked atomically when returning a saved task’s seat', async () => { + mockSend.mockResolvedValueOnce({ + Item: { ...base, status: 'FAILED', concurrency_slot: held, continuation_launch: {}, microvm_start: { clientToken: 'attempt' } }, + }).mockResolvedValueOnce({ + Item: { lease_user_id: 'user', lease_attempt_id: 'attempt', lease_state: 'CLOSED' }, + }).mockResolvedValueOnce({}); + expect(await releaseTaskSlot('task', 'user')).toBe(true); + expect(mockSend.mock.calls[2][0].input.TransactItems[2].ConditionCheck).toMatchObject({ + Key: { task_id: 'worker-lease#task' }, + ExpressionAttributeValues: { ':user': 'user', ':attempt': 'attempt', ':closed': 'CLOSED' }, + }); +}); + +test('admission atomically ties a counter increment to an owned SUBMITTED task', async () => { + mockSend.mockResolvedValueOnce({ Item: base }).mockResolvedValueOnce({}); + expect(await acquireTaskSlot('task', 'user', 3)).toBe(true); + const transaction = mockSend.mock.calls[1][0].input; + expect(transaction.TransactItems).toHaveLength(2); + expect(transaction.TransactItems[0].Update).toMatchObject({ + TableName: 'Tasks', + Key: { task_id: 'task' }, + ConditionExpression: 'user_id = :user AND #status = :submitted AND attribute_not_exists(concurrency_slot)', + ExpressionAttributeValues: { ':slot': { state: 'held' }, ':user': 'user', ':submitted': 'SUBMITTED' }, + }); + expect(transaction.TransactItems[1].Update).toMatchObject({ + TableName: 'Counters', + Key: { user_id: 'user' }, + ConditionExpression: 'attribute_not_exists(active_count) OR active_count < :max', + ExpressionAttributeValues: { ':max': 3, ':version': transaction.ClientRequestToken }, + }); +}); + +test.each(['SUBMITTED', 'HYDRATING', 'RUNNING', 'AWAITING_APPROVAL', 'FINALIZING'])( + 'replaying admission in %s reuses the held reservation', async (status) => { + mockSend.mockResolvedValueOnce({ Item: { ...base, status, concurrency_slot: held } }); + expect(await acquireTaskSlot('task', 'user', 3)).toBe(true); + expect(mockSend).toHaveBeenCalledTimes(1); + }, +); + +test('a full counter declines admission without leaving a marker', async () => { + mockSend.mockResolvedValueOnce({ Item: base }).mockRejectedValueOnce(cancelled(1)) + .mockResolvedValueOnce({ Item: base }); + expect(await acquireTaskSlot('task', 'user', 3)).toBe(false); +}); + +test('a transaction outage is not misreported as a full counter', async () => { + mockSend.mockResolvedValueOnce({ Item: base }).mockRejectedValueOnce(new Error('throttled')) + .mockResolvedValueOnce({ Item: base }); + await expect(acquireTaskSlot('task', 'user', 3)).rejects.toThrow('throttled'); +}); + +test('a lost admission response recovers the committed held marker', async () => { + mockSend.mockResolvedValueOnce({ Item: base }).mockRejectedValueOnce(new Error('response lost')) + .mockResolvedValueOnce({ Item: { ...base, concurrency_slot: held } }); + expect(await acquireTaskSlot('task', 'user', 3)).toBe(true); +}); + +test('a cancelled task cannot acquire capacity', async () => { + mockSend.mockResolvedValueOnce({ Item: { ...base, status: 'CANCELLED' } }); + expect(await acquireTaskSlot('task', 'user', 3)).toBe(false); +}); + +test('a released task cannot reacquire even if its status is changed', async () => { + mockSend.mockResolvedValueOnce({ Item: { ...base, concurrency_slot: { ...held, state: 'released' } } }); + expect(await acquireTaskSlot('task', 'user', 3)).toBe(false); +}); + +test('owner mismatch never writes', async () => { + mockSend.mockResolvedValueOnce({ Item: base }); + await expect(acquireTaskSlot('task', 'other', 3)).rejects.toThrow('owner'); + expect(mockSend).toHaveBeenCalledTimes(1); +}); + +test.each(['COMPLETED', 'FAILED', 'CANCELLED', 'TIMED_OUT'])( + '%s releases the marker and count in one guarded transaction', async (status) => { + mockSend.mockResolvedValueOnce({ Item: { ...base, status, concurrency_slot: held } }) + .mockResolvedValueOnce({}); + expect(await releaseTaskSlot('task', 'user')).toBe(true); + const transaction = mockSend.mock.calls[1][0].input; + expect(transaction.TransactItems[0].Update).toMatchObject({ + ConditionExpression: 'user_id = :user AND concurrency_slot.#state = :held AND #status IN (:completed, :failed, :cancelled, :timedOut)', + ExpressionAttributeValues: { ':released': 'released', ':held': 'held' }, + }); + expect(transaction.TransactItems[1].Update).toMatchObject({ + UpdateExpression: 'SET active_count = active_count - :one, updated_at = :now, reservation_version = :version', + ConditionExpression: 'active_count > :zero', + }); + }, +); + +test('an active task retains its slot', async () => { + mockSend.mockResolvedValueOnce({ Item: { ...base, status: 'AWAITING_APPROVAL', concurrency_slot: held } }); + expect(await releaseTaskSlot('task', 'user')).toBe(false); + expect(mockSend).toHaveBeenCalledTimes(1); +}); + +test('unadmitted or legacy terminal tasks do not return another task seat', async () => { + mockSend.mockResolvedValueOnce({ Item: { ...base, status: 'FAILED' } }); + expect(await releaseTaskSlot('task', 'user')).toBe(false); +}); + +test('release recovers a competing commit or lost successful response', async () => { + mockSend.mockResolvedValueOnce({ Item: { ...base, status: 'FAILED', concurrency_slot: held } }) + .mockRejectedValueOnce(new Error('response lost')) + .mockResolvedValueOnce({ Item: { ...base, status: 'FAILED', concurrency_slot: { ...held, state: 'released' } } }); + expect(await releaseTaskSlot('task', 'user')).toBe(false); +}); + +test('an empty counter closes the marker without subtracting from later admissions', async () => { + const row = { Item: { ...base, status: 'FAILED', concurrency_slot: held } }; + mockSend.mockResolvedValueOnce(row).mockRejectedValueOnce(cancelled(1)) + .mockResolvedValueOnce(row).mockResolvedValueOnce({}); + expect(await releaseTaskSlot('task', 'user')).toBe(true); + const update = mockSend.mock.calls[3][0].input.TransactItems[1].Update; + expect(update.UpdateExpression).toBe('SET active_count = if_not_exists(active_count, :zero), updated_at = :now, reservation_version = :version'); + expect(update.ExpressionAttributeValues).not.toHaveProperty(':one'); +}); + +test('a release outage propagates with the held marker intact', async () => { + const row = { Item: { ...base, status: 'FAILED', concurrency_slot: held } }; + mockSend.mockResolvedValueOnce(row).mockRejectedValueOnce(new Error('unavailable')).mockResolvedValueOnce(row); + await expect(releaseTaskSlot('task', 'user')).rejects.toThrow('unavailable'); +}); diff --git a/cdk/test/handlers/slack-interactions.test.ts b/cdk/test/handlers/slack-interactions.test.ts index 9d1356022..eac51a432 100644 --- a/cdk/test/handlers/slack-interactions.test.ts +++ b/cdk/test/handlers/slack-interactions.test.ts @@ -26,6 +26,7 @@ jest.mock('@aws-sdk/lib-dynamodb', () => ({ DynamoDBDocumentClient: { from: jest.fn(() => ({ send: ddbSend })) }, GetCommand: jest.fn((input: unknown) => ({ _type: 'Get', input })), UpdateCommand: jest.fn((input: unknown) => ({ _type: 'Update', input })), + TransactWriteCommand: jest.fn((input: unknown) => ({ _type: 'TransactWrite', input })), })); const smSend = jest.fn(); @@ -39,6 +40,8 @@ const fetchMock = jest.fn(); process.env.SLACK_SIGNING_SECRET_ARN = 'arn:aws:secretsmanager:us-east-1:123:secret:bgagent/slack/signing-I'; process.env.TASK_TABLE_NAME = 'Tasks'; +process.env.TASK_APPROVALS_TABLE_NAME = 'Approvals'; +process.env.TASK_EVENTS_TABLE_NAME = 'Events'; process.env.SLACK_USER_MAPPING_TABLE_NAME = 'SlackMap'; import { invalidateSlackSecretCache } from '../../src/handlers/shared/slack-verify'; @@ -127,7 +130,7 @@ describe('slack-interactions handler', () => { // 1. user mapping lookup → platform user id ddbSend.mockResolvedValueOnce({ Item: { platform_user_id: 'user-42' } }); // 2. task lookup → same owner - ddbSend.mockResolvedValueOnce({ Item: { task_id: 'task-42', user_id: 'user-42', channel_metadata: {} } }); + ddbSend.mockResolvedValueOnce({ Item: { task_id: 'task-42', user_id: 'user-42', status: 'RUNNING', channel_metadata: {} } }); // 3. update → success ddbSend.mockResolvedValueOnce({}); @@ -159,22 +162,37 @@ describe('slack-interactions handler', () => { test('cancel_task on already-terminal task warns the user', async () => { ddbSend.mockResolvedValueOnce({ Item: { platform_user_id: 'user-42' } }); - ddbSend.mockResolvedValueOnce({ Item: { task_id: 'task-42', user_id: 'user-42' } }); - // ConditionalCheckFailedException => already in terminal state - const err = new Error('conditional failed'); - err.name = 'ConditionalCheckFailedException'; - ddbSend.mockRejectedValueOnce(err); + ddbSend.mockResolvedValueOnce({ Item: { task_id: 'task-42', user_id: 'user-42', status: 'COMPLETED' } }); const event = makeInteractionEvent(interactionPayload('cancel_task:task-42')); const result = await handler(event); expect(result.statusCode).toBe(200); const posted = fetchMock.mock.calls.find( ([url, opts]) => - isSlackHooksRequestUrl(url) && String((opts as { body: string }).body).includes('terminal state'), + isSlackHooksRequestUrl(url) && String((opts as { body: string }).body).includes('no longer available'), ); expect(posted).toBeTruthy(); }); + test('cancel_task closes the linked approval in the same transaction', async () => { + ddbSend.mockResolvedValueOnce({ Item: { platform_user_id: 'user-42' } }) + .mockResolvedValueOnce({ + Item: { + task_id: 'task-42', + user_id: 'user-42', + status: 'AWAITING_APPROVAL', + awaiting_approval_request_id: 'request-42', + }, + }) + .mockResolvedValueOnce({ Item: { user_id: 'user-42', status: 'PENDING' } }) + .mockResolvedValueOnce({}); + const response = await handler(makeInteractionEvent(interactionPayload('cancel_task:task-42'))); + expect(response.statusCode).toBe(200); + const transaction = ddbSend.mock.calls.find(([command]) => command._type === 'TransactWrite')![0].input; + expect(transaction.TransactItems[1].Update.Key.request_id).toBe('request-42'); + expect(transaction.TransactItems[2].Put.Item.event_type).toBe('approval_cancelled'); + }); + test('unknown action_id is ignored silently', async () => { const event = makeInteractionEvent(interactionPayload('other_action:xyz')); const result = await handler(event); diff --git a/cdk/test/handlers/slack-notify.test.ts b/cdk/test/handlers/slack-notify.test.ts index b7b10ccf3..e421dfe30 100644 --- a/cdk/test/handlers/slack-notify.test.ts +++ b/cdk/test/handlers/slack-notify.test.ts @@ -40,6 +40,7 @@ const fetchMock = jest.fn(); (global as unknown as { fetch: unknown }).fetch = fetchMock; process.env.TASK_TABLE_NAME = 'Tasks'; +process.env.TASK_APPROVALS_TABLE_NAME = 'Approvals'; import type { DynamoDBDocumentClient } from '@aws-sdk/lib-dynamodb'; import { dispatchSlackEvent, SlackApiError, type SlackDispatchEvent } from '../../src/handlers/slack-notify'; @@ -72,6 +73,44 @@ describe('dispatchSlackEvent', () => { }); }); + test.each([false, true])('approval delivery records success only after posting (retryable failure=%s)', async fails => { + ddbSend.mockResolvedValueOnce({ + Item: { + task_id: 't1', + user_id: 'u1', + status: 'AWAITING_APPROVAL', + awaiting_approval_request_id: 'g1', + channel_source: 'slack', + channel_metadata: { slack_team_id: 'T1', slack_channel_id: 'C1', slack_thread_ts: 'thread' }, + }, + }).mockResolvedValueOnce({ + Item: { + user_id: 'u1', + status: 'PENDING', + tool_name: 'Bash', + severity: 'high', + reason: 'Review this', + tool_input_preview: ' echo hello', + created_at: '2026-09-16T12:00:00Z', + timeout_s: 1800, + }, + }).mockResolvedValue({}); + if (fails) fetchMock.mockRejectedValueOnce(new Error('network unavailable')); + const dispatch = dispatchSlackEvent(mkEvent('t1', 'approval_requested', { request_id: 'g1' }), ddb); + if (fails) { + await expect(dispatch).rejects.toThrow('network unavailable'); + expect(ddbSend).toHaveBeenCalledTimes(2); + } else { + await dispatch; + const payload = JSON.parse(fetchMock.mock.calls[0][1].body); + expect(payload.thread_ts).toBe('thread'); + expect(payload.blocks[0].text.type).toBe('plain_text'); + expect(payload.blocks[0].text.text).toContain('bgagent approve t1 g1 --scope this_call'); + expect(ddbSend.mock.calls[2][0].input.ExpressionAttributeNames).toEqual({ '#marker': 'notified_slack_approval_requested' }); + expect(fetchMock.mock.invocationCallOrder[0]).toBeLessThan(ddbSend.mock.invocationCallOrder[2]); + } + }); + test('skips non-slack tasks without touching Slack', async () => { // A Slack-subscribed event on a non-Slack task must still short- // circuit cheaply — one DDB Get, no dedup write, no API call. diff --git a/cdk/test/handlers/start-session-composition.test.ts b/cdk/test/handlers/start-session-composition.test.ts index 1d2e816b9..4591dd4a0 100644 --- a/cdk/test/handlers/start-session-composition.test.ts +++ b/cdk/test/handlers/start-session-composition.test.ts @@ -20,8 +20,8 @@ /** * Integration-style tests for the start-session step composition: * resolveComputeStrategy → strategy.startSession → transitionTask → emitTaskEvent - * These verify that the orchestrate-task handler's step 4 logic correctly - * wires the strategy, state transitions, and event emission together. + * These compose the real helpers. The actual durable handler, including + * MicroVM recovery/finalization, is exercised in orchestrate-task-microvm.test.ts. */ const mockDdbSend = jest.fn(); @@ -45,6 +45,7 @@ jest.mock('@aws-sdk/client-lambda-microvms', () => ({ LambdaMicrovmsClient: jest.fn(() => ({ send: mockMicrovmSend })), RunMicrovmCommand: jest.fn((input: unknown) => ({ _type: 'RunMicrovm', input })), GetMicrovmCommand: jest.fn((input: unknown) => ({ _type: 'GetMicrovm', input })), + GetMicrovmImageVersionCommand: jest.fn((input: unknown) => ({ _type: 'GetMicrovmImageVersion', input })), TerminateMicrovmCommand: jest.fn((input: unknown) => ({ _type: 'TerminateMicrovm', input })), MicrovmState: { PENDING: 'PENDING', @@ -56,10 +57,17 @@ jest.mock('@aws-sdk/client-lambda-microvms', () => ({ }, })); -jest.mock('@aws-sdk/client-s3', () => ({ - S3Client: jest.fn(() => ({ send: jest.fn().mockResolvedValue({}) })), - PutObjectCommand: jest.fn((input: unknown) => ({ _type: 'PutObject', input })), - DeleteObjectCommand: jest.fn((input: unknown) => ({ _type: 'DeleteObject', input })), +// This suite checks strategy/metadata composition. Real bootstrap replay is +// covered by microvm-start-recovery and orchestrate-task-microvm. +jest.mock('../../src/handlers/shared/payload-bootstrap', () => ({ + ...jest.requireActual('../../src/handlers/shared/payload-bootstrap'), + preparePayloadReference: jest.fn(async ({ taskId }: { taskId: string }) => ({ + version: 2, + task_id: taskId, + bootstrap_s3_uri: 's3://bucket/bootstrap/example.json', + payload_url: 'https://signed.example/task', + expires_at: Date.now() + 900000, + })), })); jest.mock('../../src/handlers/shared/repo-config', () => ({ @@ -83,6 +91,7 @@ let ulidCounter = 0; jest.mock('ulid', () => ({ ulid: jest.fn(() => `ULID${ulidCounter++}`) })); process.env.TASK_TABLE_NAME = 'Tasks'; +process.env.APPROVAL_REQUESTS_API_URL = 'https://approval.execute-api.us-east-1.amazonaws.com/v1'; process.env.TASK_EVENTS_TABLE_NAME = 'TaskEvents'; process.env.USER_CONCURRENCY_TABLE_NAME = 'UserConcurrency'; process.env.RUNTIME_ARN = 'arn:aws:bedrock-agentcore:us-east-1:123456789012:runtime/test'; @@ -93,11 +102,7 @@ process.env.MICROVM_EXECUTION_ROLE_ARN = 'arn:aws:iam::123456789012:role/AbcaMic process.env.MICROVM_EGRESS_CONNECTOR_ARNS = 'arn:aws:lambda:us-east-1:123456789012:network-connector/egress-1'; process.env.MICROVM_PAYLOAD_BUCKET = 'test-microvm-payload-bucket'; -// platform_config (ADR-021 P2): the four REQUIRED identifiers the MicroVM -// strategy refuses to start a session without — they are the agent's only -// channel for them, since a snapshot must not bake configuration in. Read at -// call time by `buildMicrovmPlatformConfig`, but set here alongside the rest -// for clarity. +// Required platform configuration is supplied at launch, never baked into the image. process.env.GITHUB_TOKEN_SECRET_ARN = 'arn:aws:secretsmanager:us-east-1:123456789012:secret:abca/github-token-AbCdEf'; process.env.AGENT_SESSION_ROLE_ARN = 'arn:aws:iam::123456789012:role/AbcaAgentSessionRole'; @@ -199,7 +204,19 @@ describe('start-session step composition — lambda-microvm (ADR-021)', () => { const blueprintConfig: BlueprintConfig = { compute_type: 'lambda-microvm', runtime_arn: '' }; const payload = { repo_url: 'org/repo', task_id: taskId }; - test('startSession → buildComputeMetadata → transitionTask persists microvmId and endpoint', async () => { + beforeEach(() => { + let receipt: unknown; + mockDdbSend.mockReset().mockImplementation(async (command) => { + if (command._type === 'Get') { + return { Item: { user_id: 'cognito-test', status: TaskStatus.HYDRATING, microvm_start: receipt } }; + } + const saved = command.input.ExpressionAttributeValues?.[':receipt']; + if (saved) receipt = saved; + return {}; + }); + }); + + test('startSession → buildComputeMetadata → transitionTask persists the worker and actual image', async () => { mockMicrovmSend.mockResolvedValueOnce({ microvmId: MICROVM_ID, endpoint: ENDPOINT, @@ -207,8 +224,6 @@ describe('start-session step composition — lambda-microvm (ADR-021)', () => { imageArn: 'arn:image', imageVersion: '7', }); - mockDdbSend.mockResolvedValue({}); - const strategy = resolveComputeStrategy(blueprintConfig); const handle = await strategy.startSession({ taskId, userId: 'cognito-test', payload, blueprintConfig }); @@ -225,45 +240,53 @@ describe('start-session step composition — lambda-microvm (ADR-021)', () => { // The persisted attributes are what cancel-task (and P3's approve/deny // resume) read back, so assert them on the real UpdateCommand input. - const update = mockDdbSend.mock.calls.find(c => c[0]._type === 'Update')![0]; + const update = mockDdbSend.mock.calls.find(c => + c[0]._type === 'Update' && c[0].input.ExpressionAttributeValues[':toStatus'] === TaskStatus.RUNNING)![0]; const values = update.input.ExpressionAttributeValues as Record; expect(values[':attr_compute_type']).toBe('lambda-microvm'); - expect(values[':attr_compute_metadata']).toEqual({ microvmId: MICROVM_ID, endpoint: ENDPOINT }); + expect(values[':attr_compute_metadata']).toEqual({ + microvmId: MICROVM_ID, endpoint: ENDPOINT, imageArn: 'arn:image', imageVersion: '7', + }); expect(values[':attr_session_id']).toBe(MICROVM_ID); expect(values[':toStatus']).toBe(TaskStatus.RUNNING); }); - test('compute_metadata carries ONLY the two lifecycle keys (no image ARN)', async () => { + test('image identity alone does not claim verified lifecycle support', async () => { mockMicrovmSend.mockResolvedValueOnce({ microvmId: MICROVM_ID, endpoint: ENDPOINT, state: 'RUNNING', imageArn: 'arn:aws:lambda:us-east-1:123456789012:microvm-image/abca-agent', imageVersion: '7', + }).mockResolvedValueOnce({ + imageArn: 'arn:aws:lambda:us-east-1:123456789012:microvm-image/abca-agent', + imageVersion: '7', + hooks: {}, }); const strategy = resolveComputeStrategy(blueprintConfig); const handle = await strategy.startSession({ taskId, userId: 'cognito-test', payload, blueprintConfig }); const metadata = buildComputeMetadata(handle); - // ADR-021: the image ARN is deployment-time config, logged not persisted. - expect(Object.keys(metadata).sort()).toEqual(['endpoint', 'microvmId']); - expect(JSON.stringify(metadata)).not.toContain('microvm-image'); + expect(metadata).toEqual({ + microvmId: MICROVM_ID, + endpoint: ENDPOINT, + imageArn: 'arn:aws:lambda:us-east-1:123456789012:microvm-image/abca-agent', + imageVersion: '7', + }); + expect(metadata.lifecycleProtocol).toBeUndefined(); + expect(mockMicrovmSend.mock.calls[1][0]._type).toBe('GetMicrovmImageVersion'); }); - test('error path: a marked RunMicrovm failure flows into failTask', async () => { + test('a rejected RunMicrovm preserves its marked AWS exception for the caller', async () => { const err = new Error('Rate exceeded'); err.name = 'ThrottlingException'; mockMicrovmSend.mockRejectedValueOnce(err); - mockDdbSend.mockResolvedValue({}); const strategy = resolveComputeStrategy(blueprintConfig); await expect( strategy.startSession({ taskId, userId: 'cognito-test', payload, blueprintConfig }), ).rejects.toThrow('MicroVM RunMicrovm failed: ThrottlingException: Rate exceeded'); - - await failTask(taskId, TaskStatus.HYDRATING, 'Session start failed: boom', 'user-123', true); - expect(mockDdbSend).toHaveBeenCalled(); }); }); diff --git a/cdk/test/live/README.md b/cdk/test/live/README.md new file mode 100644 index 000000000..0b9fa8955 --- /dev/null +++ b/cdk/test/live/README.md @@ -0,0 +1,33 @@ +# MicroVM live probes + +These scripts create synthetic task records, S3 objects and short-lived workers +in an existing deployment. They are excluded from Jest. Use an owned, quiet test +deployment with `NO_INGRESS`, a working image version, AWS CLI v2, and credentials +authorized for its resources. Concurrent tasks can invalidate worker attribution. + +From `cdk/`, inspect cases and effects without calling AWS: + +```sh +mise exec -- node -r ts-node/register/transpile-only test/live/verify-microvm-start.live.ts +mise exec -- node -r ts-node/register/transpile-only test/live/verify-microvm-payload.live.ts +mise exec -- node -r ts-node/register/transpile-only test/live/verify-microvm-replay.live.ts +``` + +To execute, supply every target field and a private output directory: + +```sh +mise exec -- node -r ts-node/register/transpile-only test/live/verify-microvm-start.live.ts \ + --execute --account YOUR_ACCOUNT_ID --region YOUR_REGION \ + --stack YOUR_STACK --image-version YOUR_IMAGE_VERSION \ + --output /tmp/microvm-start-verification +``` + +Use the same arguments for the other two scripts, with separate output +directories. Start/payload probes also accept `--cases` with comma-separated +names from their inspection output. The replay probe includes waits beyond five +minutes. An assertion failure is a failed verification, even if some cases pass. +Inspect the output and worker inventory after interruption; do not assume cleanup +completed. Never publish raw output containing task data or payload URLs. + +The scripts test launch recovery, payload handling and service replay behavior. +They do not replace the [approval and replacement acceptance checks](../../../docs/verification/README.md). diff --git a/cdk/test/live/microvm-live-support.ts b/cdk/test/live/microvm-live-support.ts new file mode 100644 index 000000000..5f37db0ce --- /dev/null +++ b/cdk/test/live/microvm-live-support.ts @@ -0,0 +1,108 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import assert from 'node:assert/strict'; +import { execFile } from 'node:child_process'; +import { setTimeout as delay } from 'node:timers/promises'; +import { promisify } from 'node:util'; +import { GetFunctionConfigurationCommand, LambdaClient } from '@aws-sdk/client-lambda'; +import { + GetMicrovmCommand, GetMicrovmImageCommand, LambdaMicrovmsClient, + ListMicrovmsCommand, TerminateMicrovmCommand, +} from '@aws-sdk/client-lambda-microvms'; +import { GetCallerIdentityCommand, STSClient } from '@aws-sdk/client-sts'; +import { redactPayloadUrls } from '../../src/handlers/shared/payload-bootstrap'; +import { makeClient } from '../../src/handlers/shared/ua'; + +export interface LiveTarget { + account: string; + region: string; + stack: string; + imageVersion: string; +} + +const command = promisify(execFile); + +export function safeError(error: unknown): string { + return redactPayloadUrls(error instanceof Error ? `${error.name}: ${error.message}` : String(error)); +} + +export async function connectMicrovm(target: LiveTarget) { + process.env.AWS_REGION = target.region; + process.env.AWS_DEFAULT_REGION = target.region; + const aws = async (args: string[]): Promise => { + const response = await command('aws', [...args, '--region', target.region, '--output', 'json'], { + timeout: 20_000, maxBuffer: 2 * 1024 * 1024, + }); + return JSON.parse(response.stdout) as T; + }; + // A single SDK attempt exposes the actual response to each probe. + const cfg = { region: target.region, maxAttempts: 1 }; + const identity = await makeClient(STSClient, cfg).send(new GetCallerIdentityCommand({})); + assert.equal(identity.Account, target.account, 'Wrong AWS account'); + const stack = (await aws<{ + Stacks: { StackId: string; StackStatus: string; Outputs: { OutputKey: string; OutputValue: string }[] }[]; + }>(['cloudformation', 'describe-stacks', '--stack-name', target.stack])).Stacks[0]!; + assert(['CREATE_COMPLETE', 'UPDATE_COMPLETE'].includes(stack.StackStatus), 'Stack is not ready'); + const outputs = Object.fromEntries(stack.Outputs.map(o => [o.OutputKey, o.OutputValue])); + const resources = (await aws<{ + StackResourceSummaries: { ResourceType: string; LogicalResourceId: string; PhysicalResourceId: string }[]; + }>(['cloudformation', 'list-stack-resources', '--stack-name', target.stack])).StackResourceSummaries; + const fn = resources.find(r => r.ResourceType === 'AWS::Lambda::Function' + && r.LogicalResourceId.startsWith('TaskOrchestratorOrchestratorFn'))?.PhysicalResourceId; + assert(fn, 'Cannot identify coordinator'); + const configuration = await makeClient(LambdaClient, cfg).send(new GetFunctionConfigurationCommand({ FunctionName: fn })); + const env = configuration.Environment?.Variables ?? {}; + const image = env.MICROVM_IMAGE_IDENTIFIER; + const executionRole = env.MICROVM_EXECUTION_ROLE_ARN; + const ingress = env.MICROVM_INGRESS_CONNECTOR_ARNS?.split(',') ?? []; + const egress = env.MICROVM_EGRESS_CONNECTOR_ARNS?.split(',') ?? []; + const logGroup = outputs.MicrovmLogGroupName; + assert(image && executionRole && logGroup && egress.length, 'Missing MicroVM configuration'); + assert.equal(executionRole, outputs.MicrovmExecutionRoleArn); + assert.equal(ingress.length, 1); + assert(ingress[0]!.endsWith(':NO_INGRESS'), 'Live probes require NO_INGRESS'); + const mv = makeClient(LambdaMicrovmsClient, cfg); + const info = await mv.send(new GetMicrovmImageCommand({ imageIdentifier: image })); + assert.equal(info.latestActiveImageVersion, target.imageVersion, 'Unexpected active image'); + return { aws, mv, image, executionRole, ingress, egress, logGroup, env, stackId: stack.StackId }; +} + +export async function listWorkers(mv: LambdaMicrovmsClient, image: string) { + const workers = []; + let nextToken: string | undefined; + do { + const page = await mv.send(new ListMicrovmsCommand({ imageIdentifier: image, nextToken, maxResults: 50 })); + workers.push(...(page.items ?? [])); + nextToken = page.nextToken; + } while (nextToken); + return workers; +} + +export async function stopWorker(mv: LambdaMicrovmsClient, id: string): Promise { + const first = await mv.send(new GetMicrovmCommand({ microvmIdentifier: id })); + if (first.state === 'TERMINATED') return; + if (first.state !== 'TERMINATING') await mv.send(new TerminateMicrovmCommand({ microvmIdentifier: id })); + for (let attempt = 0; attempt < 30; attempt++) { + const state = await mv.send(new GetMicrovmCommand({ microvmIdentifier: id })); + if (state.state === 'TERMINATED') return; + await delay(2_000); + } + throw new Error(`Termination not confirmed for ${id}`); +} diff --git a/cdk/test/live/microvm-start-child.ts b/cdk/test/live/microvm-start-child.ts new file mode 100644 index 000000000..715bfa179 --- /dev/null +++ b/cdk/test/live/microvm-start-child.ts @@ -0,0 +1,128 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import assert from 'node:assert/strict'; +import { createHash } from 'node:crypto'; +import { appendFile } from 'node:fs/promises'; +import { LambdaMicrovmsClient, RunMicrovmCommand, type RunMicrovmCommandOutput } from '@aws-sdk/client-lambda-microvms'; +import { PutObjectCommand, S3Client } from '@aws-sdk/client-s3'; +import { DynamoDBDocumentClient, UpdateCommand } from '@aws-sdk/lib-dynamodb'; +import { safeError } from './microvm-live-support'; + +export interface ProbeContext { + taskId: string; + userId: string; + environment: Record; + manifestKey: string; + logGroup: string; +} +export interface Trace { + kind: string; + id?: string; + fingerprint?: string; + error?: string; + autoRetried?: boolean; + [key: string]: unknown; +} + +/** SDK fault injection is confined to a fresh local process, never the deployment. */ +export async function runStartChild(context: ProbeContext, fault: string, action: string, traceFile: string): Promise { + Object.assign(process.env, context.environment); + const trace = async (row: Trace) => appendFile(traceFile, `${JSON.stringify(row)}\n`, { mode: 0o600 }); + let injected = false; + const inject = async (point: string) => { + if (injected || ![`lost-${point}-reply`, `crash-after-${point}`].includes(fault)) return; + injected = true; + await trace({ kind: 'fault', point, fault }); + if (fault.startsWith('crash-')) process.exit(73); + throw Object.assign(new Error(`Synthetic lost ${point} response after AWS success`), { name: 'TimeoutError' }); + }; + // eslint-disable-next-line @typescript-eslint/unbound-method -- Reflect.apply supplies each actual client as this. + const originalVmSend = LambdaMicrovmsClient.prototype.send; + Object.defineProperty(LambdaMicrovmsClient.prototype, 'send', { + value: async function(this: LambdaMicrovmsClient, command: RunMicrovmCommand) { + if (!(command instanceof RunMicrovmCommand)) return Reflect.apply(originalVmSend, this, [command]); + // The deterministic test transport uses 180s instead of 8h and adds logs. + // Production hashing/receipt/payload/retry logic remains unchanged. + command.input.maximumDurationInSeconds = 180; + command.input.logging = { cloudWatch: { logGroup: context.logGroup } }; + const fingerprint = createHash('sha256').update(JSON.stringify(command.input)).digest('hex'); + await trace({ kind: 'run-request', fingerprint, clientToken: command.input.clientToken }); + const response = await Reflect.apply(originalVmSend, this, [command]) as RunMicrovmCommandOutput; + await trace({ kind: 'run-success', id: response.microvmId, fingerprint, requestId: response.$metadata.requestId }); + await inject('run'); + return response; + }, + }); + // eslint-disable-next-line @typescript-eslint/unbound-method -- Reflect.apply supplies each actual client as this. + const originalS3Send = S3Client.prototype.send; + Object.defineProperty(S3Client.prototype, 'send', { + value: async function(this: S3Client, command: PutObjectCommand) { + if (!(command instanceof PutObjectCommand)) return Reflect.apply(originalS3Send, this, [command]); + assert.equal(command.input.Bucket, context.environment.MICROVM_PAYLOAD_BUCKET); + const key = command.input.Key!; + assert([context.manifestKey, `${context.taskId}/payload.json`, `${context.taskId}/launch.json`].includes(key)); + const response = await Reflect.apply(originalS3Send, this, [command]); + await trace({ kind: 's3-committed', key }); + if (key.endsWith('/payload.json')) await inject('payload'); + if (key.endsWith('/launch.json')) await inject('launch'); + return response; + }, + }); + // eslint-disable-next-line @typescript-eslint/unbound-method -- Reflect.apply supplies each actual client as this. + const originalDdbSend = DynamoDBDocumentClient.prototype.send; + Object.defineProperty(DynamoDBDocumentClient.prototype, 'send', { + value: async function(this: DynamoDBDocumentClient, command: UpdateCommand) { + const handleWrite = command instanceof UpdateCommand && command.input.UpdateExpression?.includes('microvm_start.#handle'); + if (handleWrite && fault === 'handle-write-rejected') { + await trace({ kind: 'fault', point: 'before-handle-write', fault }); + throw Object.assign(new Error('Synthetic denied handle write'), { name: 'AccessDeniedException' }); + } + const response = await Reflect.apply(originalDdbSend, this, [command]); + if (handleWrite) { + await trace({ kind: 'handle-committed' }); + await inject('handle'); + } + return response; + }, + }); + // These production modules capture their environment at import time. + const { LambdaMicrovmComputeStrategy } = await import('../../src/handlers/shared/strategies/lambda-microvm-strategy.js'); + const { startSessionWithRetry } = await import('../../src/handlers/shared/session-start-retry.js'); + const input = { + taskId: context.taskId, + userId: context.userId, + payload: { task_id: context.taskId, description: action === 'changed' ? 'Changed synthetic instructions' : 'Synthetic recovery probe' }, + blueprintConfig: { compute_type: 'lambda-microvm' as const, runtime_arn: '' }, + }; + try { + const strategy = new LambdaMicrovmComputeStrategy(); + const result = action === 'auto-retry' + ? await startSessionWithRetry(strategy, input, { + taskId: context.taskId, + emitRetryEvent: async () => { await trace({ kind: 'retry-event' }); }, + logger: { warn: () => { /* Retry event and final result are captured above/below. */ } }, + }) + : { handle: await strategy.startSession(input), autoRetried: false }; + await trace({ kind: 'result', id: result.handle.sessionId, autoRetried: result.autoRetried }); + } catch (error) { + await trace({ kind: 'application-error', error: safeError(error) }); + process.exitCode = 1; + } +} diff --git a/cdk/test/live/verify-microvm-payload.live.ts b/cdk/test/live/verify-microvm-payload.live.ts new file mode 100644 index 000000000..c15754206 --- /dev/null +++ b/cdk/test/live/verify-microvm-payload.live.ts @@ -0,0 +1,429 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +// Runs against an ALREADY deployed image; no infrastructure/permission changes. +// Every synthetic payload deliberately fails before installing configuration or +// starting a pipeline. Signed URLs stay in memory. Never persist request bodies. +import assert from 'node:assert/strict'; +import { execFile } from 'node:child_process'; +import { createHash, randomUUID } from 'node:crypto'; +import { mkdir, writeFile } from 'node:fs/promises'; +import { join } from 'node:path'; +import { setTimeout as delay } from 'node:timers/promises'; +import { parseArgs, promisify } from 'node:util'; +import { GetFunctionConfigurationCommand, LambdaClient } from '@aws-sdk/client-lambda'; +import { + GetMicrovmCommand, GetMicrovmImageCommand, LambdaMicrovmsClient, + RunMicrovmCommand, TerminateMicrovmCommand, +} from '@aws-sdk/client-lambda-microvms'; +import { DeleteObjectCommand, GetObjectCommand, PutObjectCommand, S3Client } from '@aws-sdk/client-s3'; +import { GetCallerIdentityCommand, STSClient } from '@aws-sdk/client-sts'; +import { getSignedUrl } from '@aws-sdk/s3-request-presigner'; +import { + deletePayloadReference, PAYLOAD_BOOTSTRAP, type PayloadReference, preparePayloadReference, redactPayloadUrls, +} from '../../src/handlers/shared/payload-bootstrap'; +import { makeClient } from '../../src/handlers/shared/ua'; + +const CASES = [ + 'valid-transport', 'wrong-task', 'wrong-config', 'bad-json', + 'bad-signature', 'expired-url', 'revoked-url', 'foreign-manifest', + 'wrong-path', 'bad-manifest-digest', 'large-payload', +] as const; +const command = promisify(execFile); +type Case = typeof CASES[number]; +const EXPECTED: Record = { + 'valid-transport': 'platform_config is missing or blank for required key(s)', + 'wrong-task': 'downloaded payload does not belong to the referenced task', + 'wrong-config': 'payload configuration does not match the authenticated deployment manifest', + 'bad-json': 'task payload is not valid JSON', + 'bad-signature': 'task payload download returned HTTP 403', + 'expired-url': 'task payload download returned HTTP 403', + 'revoked-url': 'task payload download returned HTTP 404', + 'foreign-manifest': 'deployment manifest read failed (AccessDenied)', + 'wrong-path': 'payload reference must sign this task\'s exact S3 object', + 'bad-manifest-digest': 'deployment manifest digest does not match its key', + 'large-payload': 'platform_config is missing or blank for required key(s)', +}; + +function safeError(error: unknown): string { + return redactPayloadUrls(error instanceof Error ? `${error.name}: ${error.message}` : String(error)); +} + +async function main(): Promise { + const { values } = parseArgs({ + options: { + 'execute': { type: 'boolean', default: false }, + 'account': { type: 'string' }, + 'region': { type: 'string' }, + 'stack': { type: 'string' }, + 'image-version': { type: 'string' }, + 'output': { type: 'string' }, + 'cases': { type: 'string', default: CASES.join(',') }, + }, + }); + const cases = values.cases!.split(',') as Case[]; + assert(cases.length > 0 && cases.every(name => CASES.includes(name)), 'Unknown probe case'); + if (!values.execute) { + process.stdout.write(`${JSON.stringify({ cases, effects: 'Synthetic S3 objects and short-lived NO_INGRESS MicroVMs; exact cleanup' })}\n`); + return; + } + for (const key of ['account', 'region', 'stack', 'image-version', 'output'] as const) { + assert(values[key], `--${key} is required with --execute`); + } + const region = values.region!; + // The production producer resolves its client region from the environment. + process.env.AWS_REGION = region; + process.env.AWS_DEFAULT_REGION = region; + const aws = async (args: string[]): Promise => { + const response = await command('aws', [...args, '--region', region, '--output', 'json'], + { timeout: 15_000, maxBuffer: 2 * 1024 * 1024 }); + return JSON.parse(response.stdout) as T; + }; + const cfg = { region, maxAttempts: 2 }; + const s3 = makeClient(S3Client, cfg); + const mv = makeClient(LambdaMicrovmsClient, cfg); + const lambda = makeClient(LambdaClient, cfg); + const sts = makeClient(STSClient, cfg); + const identity = await sts.send(new GetCallerIdentityCommand({})); + assert.equal(identity.Account, values.account, 'Wrong AWS account'); + const stack = (await aws<{ + Stacks: { + StackId?: string; StackStatus?: string; Outputs?: { OutputKey: string; OutputValue: string }[]; + }[]; + }>(['cloudformation', 'describe-stacks', '--stack-name', values.stack!])).Stacks[0]; + assert(stack, 'Stack does not exist'); + assert(stack.StackStatus?.endsWith('_COMPLETE') && !stack.StackStatus.includes('ROLLBACK'), 'Stack is not ready'); + const outputs = Object.fromEntries((stack.Outputs ?? []).map(o => [o.OutputKey!, o.OutputValue!])); + const resources = (await aws<{ + StackResourceSummaries: { + ResourceType?: string; LogicalResourceId?: string; PhysicalResourceId?: string; + }[]; + }>(['cloudformation', 'list-stack-resources', '--stack-name', values.stack!])).StackResourceSummaries; + const fn = resources.find(r => r.ResourceType === 'AWS::Lambda::Function' + && r.LogicalResourceId?.startsWith('TaskOrchestratorOrchestratorFn'))?.PhysicalResourceId; + assert(fn, 'Cannot identify this stack’s coordinator'); + const env = (await lambda.send(new GetFunctionConfigurationCommand({ FunctionName: fn }))).Environment?.Variables ?? {}; + const bucket = env.MICROVM_PAYLOAD_BUCKET; + const image = env.MICROVM_IMAGE_IDENTIFIER; + const executionRole = env.MICROVM_EXECUTION_ROLE_ARN; + const ingress = env.MICROVM_INGRESS_CONNECTOR_ARNS?.split(',') ?? []; + const egress = env.MICROVM_EGRESS_CONNECTOR_ARNS?.split(',') ?? []; + const logGroup = outputs.MicrovmLogGroupName; + const foreignBucket = outputs.MicrovmArtifactBucketName; + assert(bucket && image && executionRole && logGroup && foreignBucket && egress.length, 'Missing MicroVM settings'); + assert.equal(executionRole, outputs.MicrovmExecutionRoleArn); + assert.equal(ingress.length, 1); + assert(ingress[0]!.endsWith(':NO_INGRESS'), 'Probe requires deployed NO_INGRESS'); + const imageInfo = await mv.send(new GetMicrovmImageCommand({ imageIdentifier: image })); + assert.equal(imageInfo.latestActiveImageVersion, values['image-version'], 'Unexpected active image'); + + const runId = `p2-bootstrap-${randomUUID()}`; + const directory = values.output!; + await mkdir(directory, { recursive: false, mode: 0o700 }); + const results: Record[] = []; + const owned = new Map(); + const active = new Set(); + const vmIds: string[] = []; + const taskIds: string[] = []; + const config = { log_group_name: runId }; // valid key, deliberately incomplete configuration + const manifestBody = JSON.stringify({ backend: 'lambda-microvm', platform_config: config, version: PAYLOAD_BOOTSTRAP.version }); + const expectedManifestKey = `${PAYLOAD_BOOTSTRAP.manifest_prefix}${createHash('sha256').update(manifestBody).digest('hex')}.json`; + const remember = (b: string, key: string) => owned.set(`${b}/${key}`, { bucket: b, key }); + const report = async (row: Record) => { + results.push(row); + process.stdout.write(`${JSON.stringify(row)}\n`); + await writeFile(join(directory, 'results.json'), JSON.stringify(results, null, 2), { mode: 0o600 }); + }; + const read = async (b: string, key: string) => { + const response = await s3.send(new GetObjectCommand({ Bucket: b, Key: key })); + assert(response.Body); + return { body: await response.Body.transformToString(), etag: response.ETag, requestId: response.$metadata.requestId }; + }; + const replace = async (key: string, body: string) => { + assert(owned.has(`${bucket}/${key}`), 'Refuse to change an unowned object'); + const current = await read(bucket, key); + await s3.send(new PutObjectCommand({ Bucket: bucket, Key: key, Body: body, IfMatch: current.etag })); + }; + const stop = async (id: string) => { + const initial = await mv.send(new GetMicrovmCommand({ microvmIdentifier: id })); + if (initial.state === 'TERMINATED') { + active.delete(id); + return; + } + if (initial.state !== 'TERMINATING') { + await mv.send(new TerminateMicrovmCommand({ microvmIdentifier: id })); + } + for (let i = 0; i < 30; i++) { + const state = await mv.send(new GetMicrovmCommand({ microvmIdentifier: id })); + if (state.state === 'TERMINATED') { + active.delete(id); + return; + } + await delay(2_000); + } + throw new Error(`Termination not confirmed for ${id}`); + }; + const awaitDiagnostic = async (id: string, started: number, expected: string, status: number) => { + const day = new Date(started).toISOString().slice(0, 10).replaceAll('-', '/'); + const prefix = `${day}[${values['image-version']}]${id}`; + for (let i = 0; i < 45; i++) { + // Reuse the installed AWS CLI for logs; this live test adds no runtime SDK dependency. + const page = await aws<{ events?: { message?: string }[] }>([ + 'logs', 'filter-log-events', '--log-group-name', logGroup, + '--log-stream-name-prefix', prefix, '--start-time', String(started), + '--limit', '1000', + ]); + const messages = (page.events ?? []).map(e => e.message ?? ''); + assert(!messages.some(m => /X-Amz-(?:Signature|Credential|Security-Token)=/.test(m)), 'Bearer URL leaked to worker logs'); + assert(!messages.some(m => m.includes('hook accepted task_id=') || m.includes('installed platform_config env')), + 'Probe unexpectedly installed configuration or started a pipeline'); + if (messages.some(m => m.includes(expected)) + && messages.some(m => m.includes(`/run HTTP/1.1" ${status}`))) { + await writeFile(join(directory, `${id}.log`), messages.map(redactPayloadUrls).join(''), { mode: 0o600 }); + return messages.filter(m => m.includes('/run hook') || m.includes('/run HTTP/1.1')) + .map(m => redactPayloadUrls(m.trim())); + } + await delay(2_000); + } + throw new Error(`Expected worker diagnostic not observed for ${id}`); + }; + + await writeFile(join(directory, 'context.json'), JSON.stringify({ + runId, + account: identity.Account, + caller: identity.Arn, + region, + stackId: stack.StackId, + image, + imageVersion: values['image-version'], + executionRole, + ingress, + egress, + bucket, + maximumDurationInSeconds: 180, + cases, + // Recovery coordinates contain no signed URLs. The service lifetime bounds + // workers if this process is killed; cleanup uses only these synthetic keys. + plannedCleanupObjects: [ + { bucket, key: expectedManifestKey }, + ...cases.flatMap(name => ['payload.json', 'launch.json'].map(filename => ({ + bucket, key: `${runId}-${name}/${filename}`, + }))), + ...(cases.includes('foreign-manifest') ? [{ bucket: foreignBucket, key: expectedManifestKey }] : []), + ], + scope: 'Production producer/S3 semantics with operator credentials; actual MicroVM consumer with unchanged worker role', + }, null, 2), { mode: 0o600 }); + + let failed = false; + try { + // This is the exact canonical manifest for our one-key synthetic config. + // Record ownership before preparing, so a lost S3 reply cannot leak it. + await assert.rejects(read(bucket, expectedManifestKey), (error: unknown) => + (error as { name?: string }).name === 'NoSuchKey'); + remember(bucket, expectedManifestKey); + for (const name of cases) { + const taskId = `${runId}-${name}`; + taskIds.push(taskId); + const key = `${taskId}/payload.json`; + remember(bucket, key); + remember(bucket, `${taskId}/launch.json`); + const input = { + bucket, + taskId, + backend: 'lambda-microvm' as const, + payload: { + task_id: taskId, + description: name === 'large-payload' ? 'x'.repeat(1024 * 1024) : `Synthetic ${runId}; never start a pipeline`, + }, + platformConfig: config, + }; + // A unique config makes this run's shared manifest unique too. Discover + // its exact key from the real producer, never delete a prefix/bucket. + // Wait for BOTH writers even if one fails, before any cleanup can run. + const prepared = await Promise.allSettled([preparePayloadReference(input), preparePayloadReference(input)]); + const refs = prepared.map(result => { + if (result.status === 'rejected') throw result.reason; + return result.value; + }); + assert.deepEqual(refs[0], refs[1], 'Competing preparations produced different capabilities'); + let reference: PayloadReference = refs[0]!; + const manifestKey = new URL(reference.bootstrap_s3_uri).pathname.slice(1); + assert.equal(manifestKey, expectedManifestKey, 'Producer manifest contract changed'); + assert.deepEqual(await preparePayloadReference(input), reference, 'Replay changed the saved capability'); + const original = await read(bucket, key); + await assert.rejects(preparePayloadReference({ + ...input, payload: { ...input.payload, description: 'Conflicting synthetic instructions' }, + }), /PAYLOAD_BOOTSTRAP_CONFLICT/); + assert.equal((await read(bucket, key)).body, original.body, 'Conflict overwrote instructions'); + // Positive control for every signed object before injecting a failure. + const download = await fetch(reference.payload_url, { redirect: 'manual', signal: AbortSignal.timeout(10_000) }); + assert.equal(download.status, 200, 'Signed URL positive control failed'); + assert.equal(await download.text(), original.body); + + switch (name) { + case 'wrong-task': { + const doc = JSON.parse(original.body); + doc.task_id = `${taskId}-different`; + await replace(key, JSON.stringify(doc)); + break; + } + case 'wrong-config': { + const doc = JSON.parse(original.body); + doc.platform_config = { ...config, github_token_secret_arn: `arn:aws:secretsmanager:${region}:${identity.Account}:secret:synthetic-other-workspace` }; + await replace(key, JSON.stringify(doc)); + break; + } + case 'bad-json': + await replace(key, '{ invalid synthetic JSON'); + break; + case 'bad-signature': { + const url = new URL(reference.payload_url); + const signature = url.searchParams.get('X-Amz-Signature')!; + url.searchParams.set('X-Amz-Signature', `${signature[0] === '0' ? '1' : '0'}${signature.slice(1)}`); + reference = { ...reference, payload_url: url.toString() }; + break; + } + case 'wrong-path': { + const url = new URL(reference.payload_url); + url.pathname = url.pathname.replace(/\/payload\.json$/, '/different.json'); + reference = { ...reference, payload_url: url.toString() }; + break; + } + case 'bad-manifest-digest': + // JSON meaning stays identical; the fingerprint must reject different bytes. + await replace(manifestKey, `${manifestBody} `); + break; + case 'expired-url': { + const url = await getSignedUrl(s3, new GetObjectCommand({ Bucket: bucket, Key: key }), { expiresIn: 1 }); + await delay(2_100); + const expired = await fetch(url, { redirect: 'manual', signal: AbortSignal.timeout(10_000) }); + assert.equal(expired.status, 403); + await expired.body?.cancel(); + reference = { ...reference, payload_url: url, expires_at: Date.now() - 1 }; + break; + } + case 'revoked-url': + await deletePayloadReference(bucket, taskId); + for (const filename of ['payload.json', 'launch.json']) { + await assert.rejects(read(bucket, `${taskId}/${filename}`), (error: unknown) => + (error as { name?: string }).name === 'NoSuchKey'); + } + break; + case 'foreign-manifest': { + const manifest = await read(bucket, manifestKey); + await assert.rejects(read(foreignBucket, manifestKey), (error: unknown) => + (error as { name?: string }).name === 'NoSuchKey'); + remember(foreignBucket, manifestKey); + await s3.send(new PutObjectCommand({ + Bucket: foreignBucket, Key: manifestKey, Body: manifest.body, IfNoneMatch: '*', + })); + assert.equal((await read(foreignBucket, manifestKey)).body, manifest.body); + reference = { ...reference, bootstrap_s3_uri: `s3://${foreignBucket}/${manifestKey}` }; + break; + } + case 'valid-transport': + case 'large-payload': + break; + } + const request = { + imageIdentifier: image, + imageVersion: values['image-version'], + executionRoleArn: executionRole, + ingressNetworkConnectors: ingress, + egressNetworkConnectors: egress, + logging: { cloudWatch: { logGroup } }, + maximumDurationInSeconds: 180, + clientToken: randomUUID(), + runHookPayload: JSON.stringify(reference), + }; + assert(Buffer.byteLength(request.runHookPayload) <= 4_096, 'Reference exceeds the verified hook limit'); + const started = Date.now(); + const startedVm = await mv.send(new RunMicrovmCommand(request)); + assert(startedVm.microvmId, 'Run returned no handle; bounded service lifetime is the backstop'); + const id = startedVm.microvmId; + active.add(id); + vmIds.push(id); + await report({ + case: name, + stage: 'started', + vmId: id, + taskId, + runRequestId: startedVm.$metadata.requestId, + getObjectRequestId: original.requestId, + signedDownloadRequestId: download.headers.get('x-amz-request-id'), + storedPayloadBytes: Buffer.byteLength(original.body), + hookReferenceBytes: Buffer.byteLength(request.runHookPayload), + requestFingerprint: createHash('sha256').update(JSON.stringify(request)).digest('hex'), + }); + try { + const replay = await mv.send(new RunMicrovmCommand(request)); + if (replay.microvmId && replay.microvmId !== id) { + active.add(replay.microvmId); + vmIds.push(replay.microvmId); + } + assert.equal(replay.microvmId, id, 'Same token/request created another MicroVM'); + const status = ['valid-transport', 'large-payload', 'wrong-task', 'wrong-config', 'wrong-path'].includes(name) ? 400 : 500; + const diagnostics = await awaitDiagnostic(id, started, EXPECTED[name], status); + const observed = await mv.send(new GetMicrovmCommand({ microvmIdentifier: id })); + assert.deepEqual(observed.ingressNetworkConnectors, ingress); + await report({ + case: name, + stage: 'passed', + vmId: id, + diagnostics, + observedState: observed.state, + repeatRunSameId: true, + preparationReplayAndConflict: true, + }); + } finally { + await stop(id); + } + } + } catch (error) { + failed = true; + await report({ stage: 'failed', error: safeError(error) }); + } finally { + for (const id of [...active]) { + try { await stop(id); } catch (error) { + failed = true; + await report({ stage: 'cleanup-failed', vmId: id, error: safeError(error) }); + } + } + // Call the real finalizer helper first (revoked-url separately asserts its + // effect); operator cleanup removes only this probe's exact synthetic keys. + for (const taskId of taskIds) await deletePayloadReference(bucket, taskId); + for (const object of owned.values()) { + await s3.send(new DeleteObjectCommand({ Bucket: object.bucket, Key: object.key })); + await assert.rejects(read(object.bucket, object.key), (error: unknown) => + (error as { name?: string }).name === 'NoSuchKey'); + } + await report({ + stage: active.size ? 'cleanup-incomplete' : 'cleanup-complete', + vmIds, + activeVmIds: [...active], + deletedObjects: owned.size, + }); + } + if (failed) process.exitCode = 1; +} + +main().catch(error => { + process.stderr.write(`${safeError(error)}\n`); + process.exitCode = 1; +}); diff --git a/cdk/test/live/verify-microvm-replay.live.ts b/cdk/test/live/verify-microvm-replay.live.ts new file mode 100644 index 000000000..f0d4f80f1 --- /dev/null +++ b/cdk/test/live/verify-microvm-replay.live.ts @@ -0,0 +1,172 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import assert from 'node:assert/strict'; +import { createHash, randomUUID } from 'node:crypto'; +import { mkdir, writeFile } from 'node:fs/promises'; +import { join } from 'node:path'; +import { setTimeout as delay } from 'node:timers/promises'; +import { parseArgs } from 'node:util'; +import { GetMicrovmCommand, RunMicrovmCommand, type RunMicrovmCommandInput } from '@aws-sdk/client-lambda-microvms'; +import { connectMicrovm, listWorkers, safeError, stopWorker } from './microvm-live-support'; + +async function main(): Promise { + const { values } = parseArgs({ + options: { + 'execute': { type: 'boolean', default: false }, + 'account': { type: 'string' }, + 'region': { type: 'string' }, + 'stack': { type: 'string' }, + 'image-version': { type: 'string' }, + 'output': { type: 'string' }, + }, + }); + if (!values.execute) { + process.stdout.write(`${JSON.stringify({ + cases: ['simultaneous-identical', 'changed-parameters', 'replay-after-termination', 'replay-at-30-130-305-seconds'], + effects: 'Disposable NO_INGRESS workers with rejected startup input; no S3/task rows; maximum lifetime 180 seconds', + })}\n`); + return; + } + for (const key of ['account', 'region', 'stack', 'image-version', 'output'] as const) assert(values[key], `--${key} is required`); + const target = { account: values.account!, region: values.region!, stack: values.stack!, imageVersion: values['image-version']! }; + const live = await connectMicrovm(target); + const directory = values.output!; + await mkdir(directory, { mode: 0o700 }); + const before = await listWorkers(live.mv, live.image); + const runId = `p2-replay-${randomUUID()}`; + const request: RunMicrovmCommandInput = { + imageIdentifier: live.image, + imageVersion: target.imageVersion, + executionRoleArn: live.executionRole, + ingressNetworkConnectors: live.ingress, + egressNetworkConnectors: live.egress, + logging: { cloudWatch: { logGroup: live.logGroup } }, + // Missing v2 reference: guaranteed rejection before S3/config/pipeline. + runHookPayload: JSON.stringify({ probe: runId }), + maximumDurationInSeconds: 180, + clientToken: runId, + }; + await writeFile(join(directory, 'context.json'), JSON.stringify({ + ...target, + runId, + stackId: live.stackId, + request, + beforeIds: before.map(w => w.microvmId), + scope: 'Direct service idempotency; no coordinator or application retry code', + }, null, 2), { mode: 0o600 }); + const results: Record[] = []; + const ids = new Set(); + const started = Date.now(); + const report = async (row: Record) => { + const timed = { ...row, elapsedMs: Date.now() - started }; + results.push(timed); + process.stdout.write(`${JSON.stringify(timed)}\n`); + await writeFile(join(directory, 'results.json'), JSON.stringify(results, null, 2), { mode: 0o600 }); + }; + const launch = async (label: string, input = request) => { + try { + const result = await live.mv.send(new RunMicrovmCommand(input)); + if (result.microvmId) ids.add(result.microvmId); + await report({ + label, + outcome: 'accepted', + id: result.microvmId, + state: result.state, + requestId: result.$metadata.requestId, + duration: result.maximumDurationInSeconds, + requestFingerprint: createHash('sha256').update(JSON.stringify(input)).digest('hex'), + }); + assert(result.microvmId, 'Run returned no handle'); + return result.microvmId; + } catch (error) { + await report({ + label, + outcome: 'error', + error: safeError(error), + requestId: (error as { $metadata?: { requestId?: string } }).$metadata?.requestId, + }); + throw error; + } + }; + try { + const simultaneous = await Promise.allSettled([launch('simultaneous-a'), launch('simultaneous-b')]); + const successes = simultaneous.filter(r => r.status === 'fulfilled'); + assert(successes.length, 'Neither simultaneous request succeeded'); + assert.equal(ids.size, 1, 'Identical requests created different workers'); + for (const rejected of simultaneous.filter(r => r.status === 'rejected')) { + assert.equal((rejected.reason as { name: string }).name, 'ConflictException', 'Unexpected simultaneous failure'); + } + const id = [...ids][0]!; + assert.equal(await launch('immediate-replay'), id); + // Observe whether AWS rejects changed parameters or returns the existing + // worker. Either way it must not create a second worker for this token. + try { + assert.equal(await launch('changed-duration', { ...request, maximumDurationInSeconds: 181 }), id); + await report({ label: 'changed-parameters', outcome: 'existing-worker-returned' }); + } catch (error) { + assert.equal((error as { name: string }).name, 'ValidationException', 'Unexpected changed-request failure'); + assert.match((error as Error).message, /clientToken was used with different request parameters/); + await report({ label: 'changed-parameters', outcome: 'conflict-rejected' }); + } + const state = await live.mv.send(new GetMicrovmCommand({ microvmIdentifier: id })); + assert.equal(state.maximumDurationInSeconds, 180, 'Changed request mutated the original worker'); + // Give the intentionally invalid startup hook time to reject on its own. + for (let i = 0; i < 30; i++) { + const current = await live.mv.send(new GetMicrovmCommand({ microvmIdentifier: id })); + if (current.state === 'TERMINATED') { + assert(current.stateReason?.includes('HTTP status 400'), 'Worker did not reject the invalid startup input'); + await report({ label: 'hook-rejected', id, state: current.state, reason: current.stateReason }); + break; + } + assert(i < 29, 'Worker did not terminate after startup rejection'); + await delay(2_000); + } + assert.equal(await launch('replay-after-termination'), id); + for (const seconds of [30, 130, 305]) { + await report({ label: 'waiting-for-replay', seconds }); + while (Date.now() - started < seconds * 1_000) { + await delay(Math.min(10_000, seconds * 1_000 - (Date.now() - started))); + } + assert.equal(await launch(`replay-at-${seconds}s`), id); + } + assert.equal(ids.size, 1); + await report({ stage: 'passed', boundedRetentionEvidenceSeconds: 305, id }); + } finally { + const cleanup = await Promise.allSettled([...ids].map(id => stopWorker(live.mv, id))); + const after = await listWorkers(live.mv, live.image); + const unaccounted = after.filter(w => !before.some(b => b.microvmId === w.microvmId) && !ids.has(w.microvmId!)); + await report({ + stage: 'cleanup', + knownIds: [...ids], + cleanup: cleanup.map((r, i) => ({ + id: [...ids][i], confirmed: r.status === 'fulfilled', ...(r.status === 'rejected' && { error: safeError(r.reason) }), + })), + unaccountedIds: unaccounted.map(w => w.microvmId), + }); + assert(cleanup.every(r => r.status === 'fulfilled'), 'Cleanup incomplete'); + // Unknown rows can belong to a concurrent caller: report, never delete them. + assert.equal(unaccounted.length, 0, 'Another worker appeared during this probe; attribution needs review'); + } +} + +main().catch(error => { + process.stderr.write(`${safeError(error)}\n`); + process.exitCode = 1; +}); diff --git a/cdk/test/live/verify-microvm-start.live.ts b/cdk/test/live/verify-microvm-start.live.ts new file mode 100644 index 000000000..74de7a7ae --- /dev/null +++ b/cdk/test/live/verify-microvm-start.live.ts @@ -0,0 +1,323 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import assert from 'node:assert/strict'; +import { spawn } from 'node:child_process'; +import { createHash, randomUUID } from 'node:crypto'; +import { mkdir, readFile, readdir, writeFile } from 'node:fs/promises'; +import { join } from 'node:path'; +import { setTimeout as delay } from 'node:timers/promises'; +import { parseArgs } from 'node:util'; +import { GetMicrovmCommand } from '@aws-sdk/client-lambda-microvms'; +import { DeleteObjectCommand, GetObjectCommand, S3Client } from '@aws-sdk/client-s3'; +import { DeleteCommand, GetCommand, PutCommand, UpdateCommand } from '@aws-sdk/lib-dynamodb'; +import { connectMicrovm, listWorkers, safeError, stopWorker } from './microvm-live-support'; +import { runStartChild, type ProbeContext, type Trace } from './microvm-start-child'; +import constants from '../../../contracts/constants.json'; +import { TaskStatus } from '../../src/constructs/task-status'; +import { deletePayloadReference, PAYLOAD_BOOTSTRAP, redactPayloadUrls } from '../../src/handlers/shared/payload-bootstrap'; +import { makeClient, makeDocClient } from '../../src/handlers/shared/ua'; + +const CASES = [ + 'control', 'lost-run-reply', 'crash-after-run', 'crash-after-payload', 'crash-after-launch', + 'lost-payload-reply', 'lost-launch-reply', 'lost-handle-reply', 'crash-after-handle', + 'handle-write-rejected', 'changed-input', 'canceled-handle', 'expired-receipt', +] as const; +type Case = typeof CASES[number]; + +async function main(): Promise { + const { values } = parseArgs({ + options: { + 'execute': { type: 'boolean', default: false }, + 'account': { type: 'string' }, + 'region': { type: 'string' }, + 'stack': { type: 'string' }, + 'image-version': { type: 'string' }, + 'output': { type: 'string' }, + 'cases': { type: 'string', default: CASES.join(',') }, + 'child-context': { type: 'string' }, + 'fault': { type: 'string', default: 'none' }, + 'action': { type: 'string', default: 'direct' }, + 'trace': { type: 'string' }, + }, + }); + if (!values.execute) { + process.stdout.write(`${JSON.stringify({ cases: CASES, effects: 'Synthetic S3 files/task rows and NO_INGRESS workers limited to 180 seconds; local process faults against AWS' })}\n`); + return; + } + if (values['child-context']) { + assert(values.trace); + await runStartChild(JSON.parse(await readFile(values['child-context'], 'utf8')) as ProbeContext, values.fault!, values.action!, values.trace); + return; + } + for (const key of ['account', 'region', 'stack', 'image-version', 'output'] as const) assert(values[key], `--${key} is required`); + const cases = values.cases!.split(',') as Case[]; + assert(cases.every(c => CASES.includes(c)) && new Set(cases).size === cases.length); + const target = { account: values.account!, region: values.region!, stack: values.stack!, imageVersion: values['image-version']! }; + const live = await connectMicrovm(target); + const bucket = live.env.MICROVM_PAYLOAD_BUCKET!; + const table = live.env.TASK_TABLE_NAME!; + assert(bucket && table); + const s3 = makeClient(S3Client, { region: target.region }); + const ddb = makeDocClient({ region: target.region }); + const directory = values.output!; + await mkdir(directory, { mode: 0o700 }); + const runId = `p2-start-${randomUUID()}`; + const environment: Record = { + ...Object.fromEntries(Object.values(constants.microvm_platform_config.env_by_key).map(name => [name, ''])), + AWS_REGION: target.region, + AWS_DEFAULT_REGION: target.region, + AWS_MAX_ATTEMPTS: '1', + TASK_TABLE_NAME: table, + TASK_EVENTS_TABLE_NAME: `${runId}-events`, + // Nonempty for the producer, deliberately malformed for guest validation. + // The guest rejects this before installing config, fetching secrets or work. + AGENT_SESSION_ROLE_ARN: 'synthetic-invalid-role', + GITHUB_TOKEN_SECRET_ARN: `arn:aws:secretsmanager:${target.region}:${target.account}:secret:synthetic`, + LOG_GROUP_NAME: runId, + MICROVM_IMAGE_IDENTIFIER: live.image, + MICROVM_IMAGE_VERSION: target.imageVersion, + MICROVM_PAYLOAD_BUCKET: bucket, + MICROVM_EXECUTION_ROLE_ARN: live.executionRole, + MICROVM_INGRESS_CONNECTOR_ARNS: live.ingress.join(','), + MICROVM_EGRESS_CONNECTOR_ARNS: live.egress.join(','), + }; + const config = Object.fromEntries(Object.entries(constants.microvm_platform_config.env_by_key) + .filter(([, name]) => environment[name]).map(([key, name]) => [key, environment[name]])); + const canonical = JSON.stringify({ version: PAYLOAD_BOOTSTRAP.version, backend: 'lambda-microvm', platform_config: config }, + (_key, value: unknown) => value && typeof value === 'object' && !Array.isArray(value) + ? Object.fromEntries(Object.entries(value).sort(([a], [b]) => a.localeCompare(b))) : value); + const manifestKey = `${PAYLOAD_BOOTSTRAP.manifest_prefix}${createHash('sha256').update(canonical).digest('hex')}.json`; + const plans = cases.map(name => ({ name, taskId: randomUUID() })); + const before = await listWorkers(live.mv, live.image); + const readObject = async (key: string) => { + const result = await s3.send(new GetObjectCommand({ Bucket: bucket, Key: key })); + assert(result.Body); + return result.Body.transformToString(); + }; + const absent = async (key: string) => assert.rejects(readObject(key), + (error: unknown) => (error as { name?: string }).name === 'NoSuchKey'); + await absent(manifestKey); + await writeFile(join(directory, 'context.json'), JSON.stringify({ + ...target, + runId, + environment, + manifestKey, + plans, + beforeIds: before.map(w => w.microvmId), + scope: 'Production strategy/storage under operator credentials; local child faults; 180s lifetime and logging transport overrides', + }, null, 2), { mode: 0o600 }); + const results: Record[] = []; + const report = async (row: Record) => { + results.push(row); + process.stdout.write(`${JSON.stringify(row)}\n`); + await writeFile(join(directory, 'results.json'), JSON.stringify(results, null, 2), { mode: 0o600 }); + }; + const allTraces = async (): Promise => { + const traces: Trace[] = []; + for (const path of (await readdir(directory)).filter(p => p.endsWith('.jsonl'))) { + traces.push(...(await readFile(join(directory, path), 'utf8')) + .trim().split('\n').filter(Boolean).map(line => JSON.parse(line) as Trace)); + } + return traces; + }; + const task = async (id: string) => (await ddb.send(new GetCommand({ + TableName: table, Key: { task_id: id }, ConsistentRead: true, + }))).Item; + const cleanupTask = async (id: string) => { + await deletePayloadReference(bucket, id); + await absent(`${id}/payload.json`); + await absent(`${id}/launch.json`); + if (await task(id)) { + await ddb.send(new DeleteCommand({ + TableName: table, + Key: { task_id: id }, + ConditionExpression: 'user_id = :owner', + ExpressionAttributeValues: { ':owner': runId }, + })); + } + assert.equal(await task(id), undefined); + }; + let attempt = 0; + const execute = async (context: ProbeContext, fault = 'none', action = 'direct') => { + const contextFile = join(directory, `${context.taskId}.json`); + await writeFile(contextFile, JSON.stringify(context, null, 2), { mode: 0o600 }); + const traceFile = join(directory, `${context.taskId}-${++attempt}.jsonl`); + await writeFile(traceFile, '', { mode: 0o600 }); + const outcome = await new Promise<{ code: number | null; signal: string | null; output: string }>((resolve, reject) => { + // Reuse the parent's resolved loader; npx may supply tsx from its cache. + const processChild = spawn(process.execPath, [...process.execArgv, __filename, '--execute', + '--child-context', contextFile, '--fault', fault, '--action', action, '--trace', traceFile], + { stdio: ['ignore', 'pipe', 'pipe'], timeout: 90_000 }); + let output = ''; + processChild.stdout.on('data', chunk => { output += String(chunk); }); + processChild.stderr.on('data', chunk => { output += String(chunk); }); + processChild.on('error', reject); + processChild.on('close', (code, signal) => resolve({ code, signal, output })); + }); + await writeFile(`${traceFile}.log`, redactPayloadUrls(outcome.output), { mode: 0o600 }); + assert.equal(outcome.signal, null, 'Child exceeded its time budget'); + const trace = (await readFile(traceFile, 'utf8')).trim().split('\n').filter(Boolean).map(line => JSON.parse(line) as Trace); + await report({ stage: 'child', taskId: context.taskId, fault, action, code: outcome.code, trace }); + return { code: outcome.code, trace }; + }; + let failed = false; + try { + for (const { name, taskId } of plans) { + await ddb.send(new PutCommand({ + TableName: table, + Item: { + task_id: taskId, + user_id: runId, + status: TaskStatus.HYDRATING, + created_at: new Date().toISOString(), + ttl: Math.floor(Date.now() / 1000) + 3_600, + }, + ConditionExpression: 'attribute_not_exists(task_id)', + })); + const context = { taskId, userId: runId, environment, manifestKey, logGroup: live.logGroup }; + const initialFault = ['changed-input', 'expired-receipt'].includes(name) ? 'crash-after-run' + : ['control', 'canceled-handle'].includes(name) ? 'none' : name; + const first = await execute(context, initialFault, name === 'lost-run-reply' ? 'auto-retry' : 'direct'); + const traces = [...first.trace]; + if (initialFault !== 'none') assert(first.trace.some(t => t.kind === 'fault'), 'Requested fault was not injected'); + if (initialFault.startsWith('crash-')) { + assert.equal(first.code, 73, 'Process did not reach the requested crash point'); + const beforeRetry = await task(taskId); + assert(beforeRetry?.microvm_start); + const storedPayload = await readObject(`${taskId}/payload.json`); + if (name === 'changed-input') { + const changed = await execute(context, 'none', 'changed'); + traces.push(...changed.trace); + assert.equal(changed.code, 1); + assert(changed.trace.some(t => t.error?.includes('MICROVM_START_INPUT_CHANGED'))); + assert(!changed.trace.some(t => t.kind === 'run-request')); + assert.equal(await readObject(`${taskId}/payload.json`), storedPayload); + } + if (name === 'expired-receipt') { + await report({ stage: 'waiting-for-receipt-expiry', taskId, expiresAt: beforeRetry.microvm_start.expiresAt }); + while (Date.now() <= beforeRetry.microvm_start.expiresAt) await delay(2_000); + const expired = await execute(context); + traces.push(...expired.trace); + assert.equal(expired.code, 1); + assert(expired.trace.some(t => t.error?.includes('MICROVM_START_OUTCOME_UNKNOWN'))); + assert(!expired.trace.some(t => t.kind === 'run-request')); + } else { + const recovered = await execute(context); + traces.push(...recovered.trace); + assert.equal(recovered.code, 0); + if (name === 'crash-after-handle') assert(!recovered.trace.some(t => t.kind === 'run-request')); + assert.equal(await readObject(`${taskId}/payload.json`), storedPayload); + } + } else if (name === 'handle-write-rejected') { + assert.equal(first.code, 1); + assert(first.trace.some(t => t.error?.includes('MICROVM_START_RECEIPT_SAVE_FAILED'))); + assert.equal((await task(taskId))?.microvm_start?.handle, undefined); + } else { + assert.equal(first.code, 0); + } + if (name === 'lost-run-reply') assert(first.trace.some(t => t.kind === 'result' && t.autoRetried)); + const launches = traces.filter(t => t.kind === 'run-success'); + const ids = new Set(launches.map(t => t.id)); + assert.equal(ids.size, 1, 'Application recovery created different workers'); + assert.equal(new Set(launches.map(t => t.fingerprint)).size, 1, 'Recovery changed the service request'); + const id = launches[0]!.id!; + if (!['expired-receipt', 'handle-write-rejected'].includes(name)) { + const saved = await task(taskId); + assert.equal(saved?.microvm_start?.clientToken, taskId); + assert.equal(saved?.microvm_start?.handle?.microvmId, id); + assert.equal(saved?.session_id, id); + } + if (name === 'canceled-handle') { + await ddb.send(new UpdateCommand({ + TableName: table, + Key: { task_id: taskId }, + UpdateExpression: 'SET #s = :c', + ExpressionAttributeNames: { '#s': 'status' }, + ExpressionAttributeValues: { ':c': TaskStatus.CANCELLED }, + })); + const closed = await execute(context); + assert.equal(closed.code, 1); + assert(closed.trace.some(t => t.error?.includes('MICROVM_START_TASK_CLOSED'))); + assert(!closed.trace.some(t => t.kind === 'run-request')); + } + for (let i = 0; i < 30; i++) { + const observed = await live.mv.send(new GetMicrovmCommand({ microvmIdentifier: id })); + if (observed.state === 'TERMINATED') break; + assert(i < 29, 'Synthetic worker did not terminate'); + await delay(2_000); + } + // Rejected writes/cancellation may stop the VM before its startup hook. + // Other cases must reach the malformed-anchor barrier, before installation. + let messages: string[] = []; + for (let i = 0; i < 30; i++) { + const worker = await live.mv.send(new GetMicrovmCommand({ microvmIdentifier: id })); + const day = worker.startedAt!.toISOString().slice(0, 10).replaceAll('-', '/'); + const logs = await live.aws<{ events?: { message?: string }[] }>([ + 'logs', 'filter-log-events', '--log-group-name', live.logGroup, + '--log-stream-name-prefix', `${day}[${target.imageVersion}]${id}`, + ]); + messages = (logs.events ?? []).map(e => e.message ?? ''); + assert(!messages.some(m => /X-Amz-(Signature|Credential|Security-Token)=/i.test(m)), 'Signed URL in logs'); + assert(!messages.some(m => /installed platform_config env|hook accepted task_id=/.test(m)), 'Pipeline unexpectedly started'); + if (['handle-write-rejected', 'canceled-handle'].includes(name)) break; + if (messages.some(m => m.includes('must be a well-formed IAM role ARN'))) break; + assert(i < 29, 'Expected guest configuration rejection not observed'); + await delay(2_000); + } + await writeFile(join(directory, `${id}.log`), messages.map(redactPayloadUrls).join(''), { mode: 0o600 }); + await cleanupTask(taskId); + await report({ + stage: 'passed', + case: name, + taskId, + id, + runRequests: launches.length, + requestFingerprint: launches[0]!.fingerprint, + taskAndPayloadDeleted: true, + }); + } + } catch (error) { + failed = true; + await report({ stage: 'failed', error: safeError(error) }); + } finally { + const ids = [...new Set((await allTraces()).filter(t => t.kind === 'run-success' && t.id).map(t => t.id!))]; + const workerCleanup = await Promise.allSettled(ids.map(id => stopWorker(live.mv, id))); + const taskCleanup = await Promise.allSettled(plans.map(p => cleanupTask(p.taskId))); + const manifestCleanup = await Promise.allSettled([s3.send(new DeleteObjectCommand({ Bucket: bucket, Key: manifestKey })).then(() => absent(manifestKey))]); + const cleanup = [...workerCleanup, ...taskCleanup, ...manifestCleanup]; + const after = await listWorkers(live.mv, live.image); + const unaccounted = after.filter(w => !before.some(b => b.microvmId === w.microvmId) && !ids.includes(w.microvmId!)); + await report({ + stage: 'cleanup', + ids, + tasks: plans.map(p => p.taskId), + confirmed: cleanup.every(r => r.status === 'fulfilled'), + unaccountedIds: unaccounted.map(w => w.microvmId), + errors: cleanup.filter(r => r.status === 'rejected').map(r => safeError(r.reason)), + }); + if (cleanup.some(r => r.status === 'rejected') || unaccounted.length) failed = true; + } + if (failed) process.exitCode = 1; +} + +main().catch(error => { + process.stderr.write(`${safeError(error)}\n`); + process.exitCode = 1; +}); diff --git a/cdk/test/scripts/check-constants-sync.test.ts b/cdk/test/scripts/check-constants-sync.test.ts index 8e251e1d5..17bbd2378 100644 --- a/cdk/test/scripts/check-constants-sync.test.ts +++ b/cdk/test/scripts/check-constants-sync.test.ts @@ -58,11 +58,17 @@ const SCRIPT_REL = 'scripts/check-constants-sync.ts'; const FIXTURE_FILES = [ SCRIPT_REL, 'contracts/constants.json', + 'agent/pyproject.toml', 'agent/src/policy.py', 'agent/src/jira_reactions.py', 'agent/src/server.py', 'agent/src/config.py', + 'agent/src/payload_bootstrap.py', + 'agent/src/microvm_http.py', + 'cdk/src/handlers/shared/payload-bootstrap.ts', 'cdk/src/constructs/lambda-microvm-compute.ts', + 'cdk/src/handlers/shared/microvm-image-capability.ts', + 'cdk/src/handlers/shared/strategies/lambda-microvm-strategy.ts', ]; interface RunResult { @@ -122,10 +128,73 @@ function patchContract(root: string, mutate: (json: Record) => void } describe('check-constants-sync', () => { + test('rejects an SDK upgrade without checkpoint compatibility verification', () => { + const result = runInMutatedRepo(root => { + write(root, 'agent/pyproject.toml', read(root, 'agent/pyproject.toml') + .replace(/claude-agent-sdk==[0-9.]+/, 'claude-agent-sdk==99.0.0')); + }); + expect(result.status).toBe(1); + expect(result.stderr).toContain('claude-agent-sdk pin must match'); + }); // Node's type-stripping runs the script from source; the suite is a handful of // subprocess spawns, so give it room on a cold cache. jest.setTimeout(60_000); + describe('MicroVM lifecycle image contract', () => { + test.each([ + ['protocol_version', 0], ['protocol_version', 1.5], ['protocol_version', '1'], + ['hook_port', 0], ['hook_port', 65536], ['hook_port', 8080.5], + ['maximum_duration_seconds', 0], ['maximum_duration_seconds', 28801], + ['maximum_duration_seconds', 28800.5], ['maximum_duration_seconds', '28800'], + ['image_protocol_env', 'AWS_ACCESS_KEY_ID'], ['image_protocol_env', 'ABCA_MICROVM_bad'], + ])('rejects invalid %s=%s', (key, value) => { + const result = runInMutatedRepo(root => patchContract(root, json => { + json.microvm_lifecycle[key] = value; + })); + expect(result.status).toBe(1); + expect(result.stderr).toContain('microvm_lifecycle'); + }); + test.each([ + ['cdk/src/constructs/lambda-microvm-compute.ts', 'AGENT_HOOK_PORT', '8080'], + ['cdk/src/handlers/shared/strategies/lambda-microvm-strategy.ts', 'MICROVM_MAX_DURATION_SECONDS', '28_800'], + ['cdk/src/constructs/lambda-microvm-compute.ts', 'LIFECYCLE_HOOK_TIMEOUT_SECONDS', '30'], + ['cdk/src/handlers/shared/microvm-image-capability.ts', 'MICROVM_LIFECYCLE_PROTOCOL', '"1"'], + ['cdk/src/handlers/shared/microvm-image-capability.ts', 'MICROVM_LIFECYCLE_PROTOCOL', 'String(1)'], + ['cdk/src/handlers/shared/microvm-image-capability.ts', 'MICROVM_IMAGE_PROTOCOL_ENV', '"ABCA_MICROVM_LIFECYCLE_PROTOCOL"'], + ])('rejects a literal %s/%s', (file, name, value) => { + const result = runInMutatedRepo(root => { + write(root, file, `${read(root, file)}\nexport const ${name} = ${value};\n`); + }); + expect(result.status).toBe(1); + expect(result.stderr).toContain(name); + expect(result.stderr).toContain('Cross-language constants drift detected'); + }); + }); + + describe('payload bootstrap contract', () => { + test.each([ + ['max_payload_bytes', 0], + ['minimum_url_lifetime_seconds', 901], + ['manifest_prefix', '../'], + ['launch_filename', 'payload.json'], + ])('rejects unsafe %s', (key, value) => { + const result = runInMutatedRepo(root => patchContract(root, json => { + json.payload_bootstrap[key] = value; + })); + expect(result.status).toBe(1); + expect(result.stderr).toContain('payload_bootstrap'); + }); + + test.each([ + ['agent/src/payload_bootstrap.py', 'CONTRACT = {}'], + ['cdk/src/handlers/shared/payload-bootstrap.ts', 'export const PAYLOAD_BOOTSTRAP = {};'], + ])('rejects a consumer with its own copy: %s', (file, source) => { + const result = runInMutatedRepo(root => write(root, file, source)); + expect(result.status).toBe(1); + expect(result.stderr).toContain('consumers must read the shared contract'); + }); + }); + describe('the real repository', () => { test('passes, and says what it actually checked', () => { const result = runInMutatedRepo(() => {}); @@ -367,6 +436,29 @@ describe('check-constants-sync', () => { }); describe('ARN-pinning contract (review B5)', () => { + test.each([ + ['new_resource_arn', 'NEW_RESOURCE'], + ['new_resource', 'NEW_RESOURCE_ARN'], + ])('rejects an unvalidated ARN field %s → %s', (key, envName) => { + const result = runInMutatedRepo((root) => { + patchContract(root, (json) => { + json.microvm_platform_config.env_by_key[key] = envName; + }); + }); + expect(result.status).toBe(1); + expect(result.stderr).toContain(`ARN-shaped key "${key}" is missing from arn_keys`); + }); + + test('accepts a new ARN field when its validation is also declared', () => { + const result = runInMutatedRepo((root) => { + patchContract(root, (json) => { + json.microvm_platform_config.env_by_key.new_resource_arn = 'NEW_RESOURCE_ARN'; + json.microvm_platform_config.arn_keys.push('new_resource_arn'); + }); + }); + expect(result.status).toBe(0); + }); + test('rejects an arn_keys entry that is not a wire key', () => { const result = runInMutatedRepo((root) => { patchContract(root, (json) => { @@ -417,6 +509,29 @@ describe('check-constants-sync', () => { }); describe('hook-budget invariant', () => { + test.each(['LIFECYCLE_HANDLER_BUDGET_S', 'LIFECYCLE_HOOK_TIMEOUT_S'])( + 'rejects a hardcoded lifecycle budget %s', + (name) => { + const result = runInMutatedRepo((root) => { + write(root, 'agent/src/microvm_http.py', + `${read(root, 'agent/src/microvm_http.py')}\n${name}: float = 20\n`); + }); + expect(result.status).toBe(1); + expect(result.stderr).toContain(name); + }, + ); + + test.each([0, 1])('rejects a lifecycle handler budget %s seconds beyond the service timeout', (offset) => { + const result = runInMutatedRepo((root) => { + patchContract(root, (json) => { + json.microvm_hook_budgets.lifecycle_handler_budget_seconds = + json.microvm_hook_budgets.lifecycle_hook_timeout_seconds + offset; + }); + }); + expect(result.status).toBe(1); + expect(result.stderr).toContain('lifecycle_handler_budget_seconds must be <'); + }); + test('rejects a warm-up budget that does not fit inside the hook timeout', () => { // The relationship the two-sided contract exists for: a warm-up that cannot // answer inside the service's hook budget turns a runtime fix into a build diff --git a/cdk/test/scripts/package-microvm-artifact.test.ts b/cdk/test/scripts/package-microvm-artifact.test.ts new file mode 100644 index 000000000..5e44862de --- /dev/null +++ b/cdk/test/scripts/package-microvm-artifact.test.ts @@ -0,0 +1,227 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { spawnSync } from 'child_process'; +import { createHash } from 'crypto'; +import * as fs from 'fs'; +import * as os from 'os'; +import * as path from 'path'; + +const REPO = path.resolve(__dirname, '../..'); +const BASE_KEY = 'microvm-images/agent-artifact.zip'; +const BUILD_INPUTS = [ + 'agent/pyproject.toml', 'agent/uv.lock', 'agent/src/runner.py', + 'agent/policies/test.cedar', 'agent/workflows/test.yaml', + 'agent/prepare-commit-msg.sh', 'agent/managed-settings.json', 'contracts/constants.json', +]; + +// An executable fake CLI exercises the real shell/Python packagers without +// network access. It validates PUT checksums and enforces If-None-Match against +// a local object store; unexpected AWS operations fail instead of falling back. +const AWS = `#!/usr/bin/env python3 +import base64, hashlib, json, os, pathlib, sys +args = sys.argv[1:] +store = pathlib.Path(os.environ["MICROVM_TEST_STORE"]) +with (store / "calls.jsonl").open("a") as f: + f.write(json.dumps(args) + "\\n") +def arg(name): + return args[args.index(name) + 1] +if args[:2] == ["cloudformation", "describe-stacks"]: + outputs = { + "MicrovmArtifactBucketName": "test-artifacts", + "MicrovmArtifactObjectKey": "microvm-images/agent-artifact.zip", + "MicrovmBuildRoleArn": "arn:aws:iam::123456789012:role/test", + "MicrovmEgressConnectorArns": "runtime-connector", + "MicrovmBuildEgressConnectorArns": "build-connector", + "MicrovmLogGroupName": "/aws/lambda-microvms/test", + } + if os.environ.get("MICROVM_TEST_BASE_KEY"): + outputs["MicrovmArtifactBaseObjectKey"] = os.environ["MICROVM_TEST_BASE_KEY"] + print(json.dumps({"Stacks": [{"Outputs": [ + {"OutputKey": k, "OutputValue": v} for k,v in outputs.items() + ]}]})) +elif args[:2] == ["s3api", "put-object"]: + if os.environ.get("MICROVM_TEST_PUT_ERROR"): + print("An error occurred (AccessDenied) when calling PutObject", file=sys.stderr) + sys.exit(254) + data = pathlib.Path(arg("--body")).read_bytes() + checksum = base64.b64encode(hashlib.sha256(data).digest()).decode() + assert arg("--checksum-sha256") == checksum + target = store / pathlib.PurePosixPath(arg("--key")).name + if "--if-none-match" in args: + assert arg("--if-none-match") == "*" + if target.exists(): + print("An error occurred (PreconditionFailed) when calling PutObject", file=sys.stderr) + sys.exit(254) + target.write_bytes(data) + print(json.dumps({"ChecksumSHA256": checksum})) +elif args[:2] == ["s3api", "head-object"]: + data = (store / pathlib.PurePosixPath(arg("--key")).name).read_bytes() + print("invalid-checksum" if os.environ.get("MICROVM_TEST_BAD_CHECKSUM") else + base64.b64encode(hashlib.sha256(data).digest()).decode()) +elif args[:2] == ["lambda-microvms", "create-microvm-image"]: + print(json.dumps({"imageArn": "arn:aws:lambda:us-west-2:123456789012:microvm-image:test", + "imageVersion": "1.0"})) +else: + print("Unexpected AWS operation: " + repr(args), file=sys.stderr) + sys.exit(2) +`; + +describe('MicroVM artifact packaging and immutable publication', () => { + let fixture: string; + let store: string; + + function write(relative: string, contents: string): void { + const filename = path.join(fixture, relative); + fs.mkdirSync(path.dirname(filename), { recursive: true }); + fs.writeFileSync(filename, contents); + } + + function packageArtifact(extraEnv: Record = {}, args: string[] = []) { + return spawnSync('bash', [ + path.join(fixture, 'cdk/scripts/package-microvm-artifact.sh'), ...args, + ], { + encoding: 'utf8', + timeout: 20_000, + env: { + ...process.env, + PATH: `${path.join(fixture, 'bin')}${path.delimiter}${process.env.PATH}`, + MICROVM_TEST_STORE: store, + ...extraEnv, + }, + }); + } + + function uploads(): string[][] { + return fs.readFileSync(path.join(store, 'calls.jsonl'), 'utf8').trim().split('\n') + .map(line => JSON.parse(line) as string[]) + .filter(args => args[0] === 's3api' && args[1] === 'put-object'); + } + + beforeEach(() => { + fixture = fs.mkdtempSync(path.join(os.tmpdir(), 'abca-microvm-package-test-')); + store = path.join(fixture, 'store'); + fs.mkdirSync(store); + for (const name of ['package-microvm-artifact.sh', 'build-microvm-artifact.py']) { + write(`cdk/scripts/${name}`, fs.readFileSync(path.join(REPO, 'scripts', name), 'utf8')); + } + write('agent/Dockerfile', 'FROM scratch\nCOPY agent/src/ /app/src/\n'); + for (const file of BUILD_INPUTS) write(file, `contents of ${file}\n`); + write('bin/aws', AWS); + fs.chmodSync(path.join(fixture, 'bin/aws'), 0o755); + fs.chmodSync(path.join(fixture, 'agent/prepare-commit-msg.sh'), 0o755); + }); + + afterEach(() => { + fs.rmSync(fixture, { recursive: true, force: true }); + }); + + test('same inputs reuse identical bytes despite dates/caches; changed source gets a new key', () => { + const first = packageArtifact(); + expect(first.status).toBe(0); + const firstArgs = uploads()[0]!; + const firstKey = firstArgs[firstArgs.indexOf('--key') + 1]!; + const bytes = fs.readFileSync(path.join(store, path.basename(firstKey))); + const hash = createHash('sha256').update(bytes).digest('hex'); + expect(firstKey).toBe(`microvm-images/agent-artifact-${hash}.zip`); + expect(first.stdout).toContain(`--context microvm_artifact_sha256=${hash}`); + + fs.utimesSync(path.join(fixture, 'agent/src/runner.py'), new Date(0), new Date(0)); + write('agent/src/__pycache__/runner.pyc', 'cache'); + write('agent/.coverage', 'not a build input'); + write('agent/tests/test_runner.py', 'not a build input'); + const repeated = packageArtifact(); + expect(repeated.status).toBe(0); + expect(repeated.stdout).toContain('Reusing checksum-verified existing artifact'); + expect(uploads()[1]).toContain(firstKey); + expect(fs.readFileSync(path.join(store, path.basename(firstKey)))).toEqual(bytes); + + write('agent/src/runner.py', 'changed runtime source'); + const changed = packageArtifact(); + expect(changed.status).toBe(0); + expect(uploads()[2]).not.toContain(firstKey); + + const archive = spawnSync('python3', ['-c', ` +import json, sys, zipfile +with zipfile.ZipFile(sys.argv[1]) as z: + print(json.dumps({i.filename: {"date": i.date_time, "mode": (i.external_attr >> 16) & 511} + for i in z.infolist()})) +`, path.join(store, path.basename(firstKey))], { encoding: 'utf8' }); + expect(archive.status).toBe(0); + const entries = JSON.parse(archive.stdout); + expect(Object.keys(entries).sort()).toEqual(['Dockerfile', ...BUILD_INPUTS].sort()); + expect(entries['agent/prepare-commit-msg.sh'].mode).toBe(0o755); + expect(entries['agent/src/runner.py']).toEqual({ date: [1980, 1, 1, 0, 0, 0], mode: 0o644 }); + }); + + test('rejects an existing object whose verified checksum differs', () => { + expect(packageArtifact().status).toBe(0); + const second = packageArtifact({ MICROVM_TEST_BAD_CHECKSUM: '1' }); + expect(second.status).not.toBe(0); + expect(second.stderr).toContain('refusing to overwrite'); + }); + + test('does not treat an authorization failure as an existing artifact', () => { + const result = packageArtifact({ MICROVM_TEST_PUT_ERROR: '1' }); + expect(result.status).not.toBe(0); + expect(result.stderr).toContain('AccessDenied'); + expect(fs.readFileSync(path.join(store, 'calls.jsonl'), 'utf8')).not.toContain('head-object'); + }); + + test('rejects symlinked inputs before any upload', () => { + fs.symlinkSync(path.join(fixture, 'agent/uv.lock'), path.join(fixture, 'agent/src/linked.py')); + const result = packageArtifact(); + expect(result.status).not.toBe(0); + expect(result.stderr).toContain('must not contain symlinks'); + expect(uploads()).toEqual([]); + }); + + test('uses the base-key output without nesting a previous digest into the next key', () => { + expect(packageArtifact({ MICROVM_TEST_BASE_KEY: 'custom/input.zip' }).status).toBe(0); + const args = uploads()[0]!; + expect(args[args.indexOf('--key') + 1]).toMatch(/^custom\/input-[a-f0-9]{64}\.zip$/); + }); + + test('keeps explicit out-of-band creation on the legacy fixed key', () => { + const result = packageArtifact({}, [ + '--create-image', '--base-image-arn', 'arn:aws:lambda:us-west-2:aws:microvm-image:al2023-1', + '--base-image-version', '1', + ]); + expect(result.status).toBe(0); + expect(uploads()[0]).toContain(BASE_KEY); + expect(uploads()[0]).not.toContain('--if-none-match'); + }); +}); + +test('the packager includes every local Dockerfile COPY source and no extra source tree', () => { + const dockerfile = fs.readFileSync(path.join(REPO, '../agent/Dockerfile'), 'utf8'); + const copied = dockerfile.split('\n') + .filter(line => line.startsWith('COPY ') && !line.includes('--from=')) + .flatMap(line => line.trim().split(/\s+/).slice(1, -1)) + .map(source => source.replace(/\/$/, '')); + const helper = spawnSync('python3', ['-c', ` +import importlib.util, json, sys +spec = importlib.util.spec_from_file_location("artifact", sys.argv[1]) +module = importlib.util.module_from_spec(spec) +spec.loader.exec_module(module) +print(json.dumps(module.INPUTS)) +`, path.join(REPO, 'scripts/build-microvm-artifact.py')], { encoding: 'utf8' }); + expect(helper.status).toBe(0); + expect(JSON.parse(helper.stdout).sort()).toEqual(copied.sort()); +}); diff --git a/cdk/test/stacks/agent.test.ts b/cdk/test/stacks/agent.test.ts index c20672955..f14825165 100644 --- a/cdk/test/stacks/agent.test.ts +++ b/cdk/test/stacks/agent.test.ts @@ -19,7 +19,7 @@ import * as fs from 'fs'; import * as path from 'path'; -import { App, AspectPriority, Aspects } from 'aws-cdk-lib'; +import { App, AspectPriority, Aspects, NestedStack } from 'aws-cdk-lib'; import { Match, Template } from 'aws-cdk-lib/assertions'; import { BEDROCK_GEO_REGION_CONTEXT_KEY, @@ -27,11 +27,13 @@ import { DEFAULT_BEDROCK_MODEL_IDS, } from '../../src/constructs/bedrock-models'; import * as lambdaMicrovmCompute from '../../src/constructs/lambda-microvm-compute'; +import { LambdaMicrovmStack } from '../../src/constructs/lambda-microvm-stack'; import { buildAppId, SolutionUaAspect } from '../../src/constructs/solution-ua-aspect'; import { AgentStack } from '../../src/stacks/agent'; describe('AgentStack', () => { let template: Template; + let concurrencyMaintenance: Template; beforeAll(() => { const app = new App(); @@ -39,12 +41,62 @@ describe('AgentStack', () => { env: { account: '123456789012', region: 'us-east-1' }, }); template = Template.fromStack(stack); + concurrencyMaintenance = Template.fromStack(stack.node.findChild('ConcurrencyMaintenance') as NestedStack); }); test('synthesizes without errors', () => { expect(template).toBeDefined(); }); + test('nests scheduled maintenance while retaining its data in the parent', () => { + const parentFunctions = Object.keys(template.findResources('AWS::Lambda::Function')); + for (const prefix of ['ConcurrencyReconciler', 'AdmissionQueuePickup', 'StrandedTaskReconciler', 'PendingUploadCleanup']) { + expect(parentFunctions.some(id => id.startsWith(prefix))).toBe(false); + expect(Object.keys(concurrencyMaintenance.findResources('AWS::Lambda::Function')) + .filter(id => id.startsWith(prefix))).toHaveLength(1); + } + concurrencyMaintenance.resourceCountIs('AWS::Lambda::Function', 4); + concurrencyMaintenance.resourceCountIs('AWS::Events::Rule', 4); + concurrencyMaintenance.resourceCountIs('AWS::S3::Bucket', 0); + concurrencyMaintenance.resourceCountIs('AWS::DynamoDB::Table', 0); + concurrencyMaintenance.hasResourceProperties('AWS::Events::Rule', { + ScheduleExpression: 'rate(15 minutes)', + }); + const fn = Object.entries(concurrencyMaintenance.findResources('AWS::Lambda::Function')) + .find(([id]) => id.startsWith('ConcurrencyReconciler'))![1]; + const variables = fn.Properties.Environment.Variables; + const nested = Object.entries(template.findResources('AWS::CloudFormation::Stack')) + .find(([id]) => id.startsWith('ConcurrencyMaintenanceNestedStack'))![1]; + expect(nested.Properties.Parameters[variables.TASK_TABLE_NAME.Ref]) + .toEqual({ Ref: expect.stringMatching(/^TaskTable/) }); + expect(nested.Properties.Parameters[variables.USER_CONCURRENCY_TABLE_NAME.Ref]) + .toEqual({ Ref: expect.stringMatching(/^UserConcurrencyTable/) }); + }); + + test('retains the guardrail version referenced by pinned durable environments', () => { + template.resourceCountIs('AWS::Bedrock::GuardrailVersion', 1); + template.hasResource('AWS::Bedrock::GuardrailVersion', { + DeletionPolicy: 'Retain', + UpdateReplacePolicy: 'Retain', + }); + }); + + test('AgentCore runtime has no direct DynamoDB grant, including capacity counters', () => { + const roles = Object.entries(template.findResources('AWS::IAM::Role')); + const runtimeRoleIds = roles.filter(([, role]) => + JSON.stringify(role.Properties.AssumeRolePolicyDocument).includes('bedrock-agentcore.amazonaws.com'), + ).map(([id]) => id); + expect(runtimeRoleIds.length).toBeGreaterThan(0); + const policies = Object.values(template.findResources('AWS::IAM::Policy')).filter( + (policy) => policy.Properties.Roles.some((role: { Ref?: string }) => + role.Ref && runtimeRoleIds.includes(role.Ref), + ), + ); + expect(policies.length).toBeGreaterThan(0); + expect(JSON.stringify(policies)).not.toContain('dynamodb:'); + expect(JSON.stringify(roles.filter(([id]) => runtimeRoleIds.includes(id)))).not.toContain('dynamodb:'); + }); + test('creates exactly 22 DynamoDB tables', () => { // task, task-events, repo, user-concurrency, budget, webhook, task-nudges, // task-approvals (Cedar HITL V2), @@ -101,6 +153,7 @@ describe('AgentStack', () => { for (const key of [ 'TASK_APPROVALS_TABLE_NAME', + 'APPROVAL_REQUESTS_API_URL', 'NUDGES_TABLE_NAME', 'LOG_GROUP_NAME', 'ARTIFACTS_BUCKET_NAME', @@ -114,8 +167,7 @@ describe('AgentStack', () => { ]) { expect(env[key]).toBeDefined(); } - // Plus the three the orchestrator already carried for its own work — together - // these cover all four identifiers the MicroVM strategy treats as required. + // These existing coordinator values complete the required platform configuration. expect(env.TASK_TABLE_NAME).toBeDefined(); expect(env.TASK_EVENTS_TABLE_NAME).toBeDefined(); expect(env.GITHUB_TOKEN_SECRET_ARN).toBeDefined(); @@ -135,6 +187,7 @@ describe('AgentStack', () => { for (const key of [ 'TASK_APPROVALS_TABLE_NAME', + 'APPROVAL_REQUESTS_API_URL', 'NUDGES_TABLE_NAME', 'LOG_GROUP_NAME', 'ARTIFACTS_BUCKET_NAME', @@ -1163,7 +1216,7 @@ describe('AgentStack with the ECS substrate gate (--context compute_type=ecs)', test('provisions an ECS cluster + both Fargate task definitions (build + planning)', () => { template.resourceCountIs('AWS::ECS::Cluster', 1); - // Two task defs — the 64 GB build def and the 8 GB read-only planning def + // Two task defs — the 16 GB build def and the 8 GB read-only planning def // (a read-only workflow runs on the smaller one). See // docs/design/ECS_RIGHTSIZED_PLANNING.md. template.resourceCountIs('AWS::ECS::TaskDefinition', 2); @@ -1173,6 +1226,20 @@ describe('AgentStack with the ECS substrate gate (--context compute_type=ecs)', template.hasOutput('ComputeSubstrate', { Value: 'ecs' }); }); + test('both ECS task definitions receive the same approval table as AgentCore', () => { + const runtime = Object.values(template.findResources('AWS::BedrockAgentCore::Runtime'))[0]; + const approvalTable = runtime.Properties.EnvironmentVariables.TASK_APPROVALS_TABLE_NAME; + expect(approvalTable).toBeDefined(); + const definitions = Object.values(template.findResources('AWS::ECS::TaskDefinition')); + expect(definitions).toHaveLength(2); + for (const definition of definitions) { + expect(definition.Properties.ContainerDefinitions[0].Environment).toContainEqual({ + Name: 'TASK_APPROVALS_TABLE_NAME', + Value: approvalTable, + }); + } + }); + test('the orchestrator gets the PLANNING task-def ARN, not just the build one', () => { // Without this env var the ECS strategy's `readOnly && // ECS_PLANNING_TASK_DEFINITION_ARN` guard is always falsy, so the planning @@ -1232,6 +1299,28 @@ describe('AgentStack with the Lambda MicroVMs substrate gate (--context compute_ const BASE_IMAGE_ARN = 'arn:aws:lambda:us-east-1:aws:microvm-image:al2023-1'; let template: Template; + let childTemplate: Template; + + function assertChildReference(value: unknown, resourceType: string, attribute: string, prefix = ''): void { + const [stackId, outputPath] = (value as { 'Fn::GetAtt': [string, string] })['Fn::GetAtt']; + expect(stackId).toMatch(/^MicrovmNestedStack/); + expect(outputPath).toMatch(/^Outputs\./); + const childOutput = childTemplate.toJSON().Outputs[outputPath.slice('Outputs.'.length)]; + const ids = Object.keys(childTemplate.findResources(resourceType)).filter(id => id.startsWith(prefix)); + expect(ids.length).toBeGreaterThan(0); + expect(childOutput.Value).toEqual({ 'Fn::GetAtt': [expect.stringMatching(new RegExp(`^(${ids.join('|')})$`)), attribute] }); + } + + function orchestratorEnvironment(): Record { + return Object.entries(template.findResources('AWS::Lambda::Function')) + .find(([id]) => id.includes('TaskOrchestratorOrchestratorFn'))![1] + .Properties.Environment.Variables; + } + + function imageResources(): unknown[] { + const arn = orchestratorEnvironment().MICROVM_IMAGE_IDENTIFIER; + return [arn, { 'Fn::Join': ['', [arn, ':*']] }]; + } beforeAll(() => { // Gate ON *and* an image configured — the steady state. The intermediate @@ -1241,21 +1330,27 @@ describe('AgentStack with the Lambda MicroVMs substrate gate (--context compute_ const app = new App({ context: { compute_type: 'lambda-microvm', + microvm_nested_stack: true, microvm_base_image_arn: BASE_IMAGE_ARN, microvm_base_image_version: '1', + microvm_artifact_sha256: 'a'.repeat(64), + microvm_managed_image_version: '7.0', }, }); const stack = new AgentStack(app, 'TestAgentStackMicrovm', { env: { account: '123456789012', region: 'us-east-1' }, }); template = Template.fromStack(stack); + childTemplate = Template.fromStack(stack.node.findChild('Microvm') as LambdaMicrovmStack); }); - test('provisions the MicroVM image + BOTH egress network connectors', () => { - template.resourceCountIs('AWS::Lambda::MicrovmImage', 1); + test('defaults to a child stack for the MicroVM image and both egress network connectors', () => { + template.resourceCountIs('AWS::Lambda::MicrovmImage', 0); + template.resourceCountIs('AWS::Lambda::NetworkConnector', 0); + childTemplate.resourceCountIs('AWS::Lambda::MicrovmImage', 1); // Runtime (443) + build-time (443 + 80, for apt-get) — see ADR-021's // build-time-egress security-table row. - template.resourceCountIs('AWS::Lambda::NetworkConnector', 2); + childTemplate.resourceCountIs('AWS::Lambda::NetworkConnector', 2); }); test('does NOT provision the ECS substrate (the gates are mutually exclusive)', () => { @@ -1271,6 +1366,7 @@ describe('AgentStack with the Lambda MicroVMs substrate gate (--context compute_ for (const output of [ 'MicrovmArtifactBucketName', 'MicrovmArtifactObjectKey', + 'MicrovmArtifactBaseObjectKey', 'MicrovmBuildRoleArn', 'MicrovmExecutionRoleArn', 'MicrovmEgressConnectorArns', @@ -1283,6 +1379,14 @@ describe('AgentStack with the Lambda MicroVMs substrate gate (--context compute_ } }); + test('forwards the uploaded artifact digest into the image URI and outputs', () => { + const key = `microvm-images/agent-artifact-${'a'.repeat(64)}.zip`; + template.hasOutput('MicrovmArtifactObjectKey', { Value: key }); + template.hasOutput('MicrovmArtifactBaseObjectKey', { Value: 'microvm-images/agent-artifact.zip' }); + const image = Object.values(childTemplate.findResources('AWS::Lambda::MicrovmImage'))[0]!; + expect(JSON.stringify(image.Properties.CodeArtifact.Uri)).toContain(key); + }); + test('the build and runtime egress connector outputs are DIFFERENT connectors', () => { const outputs = template.toJSON().Outputs as Record; expect(JSON.stringify(outputs.MicrovmEgressConnectorArns.Value)) @@ -1296,14 +1400,18 @@ describe('AgentStack with the Lambda MicroVMs substrate gate (--context compute_ const env = orchestrator.Properties.Environment.Variables as Record; expect(Object.keys(env).filter(k => k.startsWith('MICROVM_')).sort()).toEqual([ + 'MICROVM_APPROVAL_SUSPEND_ENABLED', + 'MICROVM_APPROVAL_SUSPEND_PARAMETER_NAME', 'MICROVM_EGRESS_CONNECTOR_ARNS', 'MICROVM_EXECUTION_ROLE_ARN', 'MICROVM_IMAGE_IDENTIFIER', + 'MICROVM_IMAGE_VERSION', 'MICROVM_INGRESS_CONNECTOR_ARNS', 'MICROVM_PAYLOAD_BUCKET', ]); - // Image version is deliberately unpinned. - expect(env.MICROVM_IMAGE_VERSION).toBeUndefined(); + // A runtime pin leaves the child-owned managed image in place. + expect(env.MICROVM_IMAGE_VERSION).toBe('7.0'); + expect(env.MICROVM_APPROVAL_SUSPEND_ENABLED).toBe('false'); // Ingress is NOT empty and NOT omitted: RunMicrovm attaches a PUBLIC // HTTP_INGRESS connector (with a public endpoint) when the field is absent, // so "no inbound" is an explicit control on every launch. @@ -1319,11 +1427,10 @@ describe('AgentStack with the Lambda MicroVMs substrate gate (--context compute_ const [, orchestrator] = Object.entries(fns) .find(([id]) => id.includes('TaskOrchestratorOrchestratorFn'))!; const env = orchestrator.Properties.Environment.Variables as Record; - expect(JSON.stringify(env.MICROVM_IMAGE_IDENTIFIER)) - .toMatch(/"Fn::GetAtt":\["LambdaMicrovmComputeImage[^"]*","ImageArn"\]/); + assertChildReference(env.MICROVM_IMAGE_IDENTIFIER, 'AWS::Lambda::MicrovmImage', 'ImageArn'); }); - test('grants the orchestrator exactly the P1 lifecycle actions, image-scoped', () => { + test('grants the orchestrator its launch, state, cleanup and image-capability actions, image-scoped', () => { const policies = Object.entries(template.findResources('AWS::IAM::Policy')) .filter(([id]) => id.includes('TaskOrchestrator')); const statements = policies.flatMap(([, p]) => p.Properties.PolicyDocument.Statement as Array<{ @@ -1336,34 +1443,35 @@ describe('AgentStack with the Lambda MicroVMs substrate gate (--context compute_ expect(lifecycle.Action).toEqual([ 'lambda:RunMicrovm', 'lambda:GetMicrovm', + 'lambda:GetMicrovmImageVersion', 'lambda:TerminateMicrovm', + 'lambda:SuspendMicrovm', + 'lambda:ResumeMicrovm', ]); // Every MicroVM lifecycle action authorizes against the *image* resource, // which is why "scoped to platform-created images" is achievable at all. - expect(JSON.stringify(lifecycle.Resource)).toMatch( - /"Fn::GetAtt":\["LambdaMicrovmComputeImage[^"]*","ImageArn"\]/, - ); + expect(lifecycle.Resource).toEqual(imageResources()); // PassNetworkConnector supports no resource-level permissions. const pass = statements.find(s => s.Sid === 'MicrovmPassNetworkConnector')!; expect(pass.Action).toBe('lambda:PassNetworkConnector'); expect(pass.Resource).toBe('*'); - // iam:PassRole for the execution role hand-off, service-conditioned. + // iam:PassRole for the execution role hand-off is scoped to the exact role. const passRole = statements.find(s => s.Sid === 'MicrovmPassExecutionRole')!; expect(passRole.Action).toBe('iam:PassRole'); }); - test('does NOT grant suspend/resume (P3) or auth-token minting (never)', () => { + test('grants supervisor sleep/wake without auth-token minting or shell access', () => { const rendered = JSON.stringify(template.toJSON()); - expect(rendered).not.toContain('lambda:SuspendMicrovm'); - expect(rendered).not.toContain('lambda:ResumeMicrovm'); + expect(rendered).toContain('lambda:SuspendMicrovm'); + expect(rendered).toContain('lambda:ResumeMicrovm'); expect(rendered).not.toContain('lambda:CreateMicrovmAuthToken'); expect(rendered).not.toContain('lambda:CreateMicrovmShellAuthToken'); expect(rendered).not.toContain('lambda:ConnectMicrovm'); }); - test('orchestrator may WRITE the payload bucket; nothing grants it delete', () => { + test('orchestrator may upload payloads and delete task payload objects at finalize', () => { const policies = Object.entries(template.findResources('AWS::IAM::Policy')) .filter(([id]) => id.includes('TaskOrchestrator')); const statements = policies.flatMap(([, p]) => p.Properties.PolicyDocument.Statement as Array<{ @@ -1371,13 +1479,23 @@ describe('AgentStack with the Lambda MicroVMs substrate gate (--context compute_ Resource: unknown; }>); const payloadStatements = statements.filter(s => - JSON.stringify(s.Resource).includes('LambdaMicrovmComputePayloadBucket')); + JSON.stringify(s.Resource).includes('ComputePayloadBucket')); const actions = payloadStatements.flatMap(s => Array.isArray(s.Action) ? s.Action : [s.Action]); expect(actions).toContain('s3:PutObject'); - // The bucket's lifecycle rule is the reaper on this backend — unlike the ECS - // path the orchestrator never deletes, so the grant must not exist. - expect(actions).not.toContain('s3:DeleteObject'); + expect(actions).toContain('s3:DeleteObject'); + expect(actions).toContain('s3:GetObject'); + expect(actions).toContain('s3:ListBucket'); + const deletion = payloadStatements.find(s => Array.isArray(s.Action) && s.Action.includes('s3:DeleteObject')); + const resources = deletion!.Resource as Array<{ 'Fn::Join': [string, unknown[]] }>; + const payloadArn = resources[0]['Fn::Join'][1][0]; + assertChildReference(payloadArn, 'AWS::S3::Bucket', 'Arn', 'ComputePayloadBucket'); + expect(deletion!.Resource).toEqual(['payload.json', 'launch.json'].map(filename => ({ + 'Fn::Join': ['', [ + payloadArn, + `/*/${filename}`, + ]], + }))); }); test('cancel Lambda may terminate a MicroVM (and only terminate), image-scoped', () => { @@ -1396,7 +1514,7 @@ describe('AgentStack with the Lambda MicroVMs substrate gate (--context compute_ // (TaskApi is built before the MicroVM construct), so the grant names ONE // image instead of an account/Region-wide `microvm-image:*`. const rendered = JSON.stringify(microvmStatements[0]!.Resource); - expect(rendered).toMatch(/LambdaMicrovmComputeImage[^"]*","ImageArn"/); + expect(microvmStatements[0]!.Resource).toEqual(imageResources()); expect(rendered).not.toContain('microvm-image:*'); }); @@ -1446,10 +1564,10 @@ describe('AgentStack with the Lambda MicroVMs substrate gate (--context compute_ }); test('MicroVM resources carry the backend cost-allocation tag', () => { - template.hasResourceProperties('AWS::Lambda::MicrovmImage', { + childTemplate.hasResourceProperties('AWS::Lambda::MicrovmImage', { Tags: Match.arrayWith([{ Key: 'abca:compute-backend', Value: 'lambda-microvm' }]), }); - template.hasResourceProperties('AWS::Lambda::NetworkConnector', { + childTemplate.hasResourceProperties('AWS::Lambda::NetworkConnector', { Tags: Match.arrayWith([{ Key: 'abca:compute-backend', Value: 'lambda-microvm' }]), }); }); @@ -1466,18 +1584,22 @@ describe('AgentStack with the Lambda MicroVMs substrate gate (--context compute_ const app = new App({ context: { compute_type: 'lambda-microvm', + microvm_nested_stack: true, microvm_region_override: true, microvm_base_image_arn: BASE_IMAGE_ARN, microvm_base_image_version: '1', + microvm_artifact_sha256: 'a'.repeat(64), }, }); - overriddenTemplate = Template.fromStack(new AgentStack(app, 'TestAgentStackMicrovmOverride', { + const stack = new AgentStack(app, 'TestAgentStackMicrovmOverride', { env: { account: '123456789012', region: 'eu-central-1' }, - })); + }); + Template.fromStack(stack); + overriddenTemplate = Template.fromStack(stack.node.findChild('Microvm') as LambdaMicrovmStack); }); test('fails synth when the stack Region has no Lambda MicroVMs', () => { - const app = new App({ context: { compute_type: 'lambda-microvm' } }); + const app = new App({ context: { compute_type: 'lambda-microvm', microvm_nested_stack: true } }); expect(() => new AgentStack(app, 'TestAgentStackMicrovmBadRegion', { env: { account: '123456789012', region: 'eu-central-1' }, })).toThrow(/AWS Lambda MicroVMs are not available in eu-central-1/); @@ -1525,21 +1647,25 @@ describe('AgentStack default (agentcore) deploy — MicroVM substrate absent', ( describe('AgentStack with the MicroVM gate on but no image configured (first deploy)', () => { let template: Template; + let childTemplate: Template; beforeAll(() => { // The bootstrap state: substrate provisioned so the artifact bucket exists, // but no image yet. Exercises the false branch of the shared // `isLambdaMicrovmImageConfigured` predicate that gates BOTH the // orchestrator's MICROVM_* wiring and the cancel Lambda's grant. - const app = new App({ context: { compute_type: 'lambda-microvm' } }); + const app = new App({ context: { compute_type: 'lambda-microvm', microvm_nested_stack: true } }); const stack = new AgentStack(app, 'TestAgentStackMicrovmNoImage', { env: { account: '123456789012', region: 'us-east-1' }, }); template = Template.fromStack(stack); + childTemplate = Template.fromStack(stack.node.findChild('Microvm') as LambdaMicrovmStack); }); test('provisions the substrate (buckets, roles, both connectors) but no image', () => { - template.resourceCountIs('AWS::Lambda::NetworkConnector', 2); + template.resourceCountIs('AWS::Lambda::NetworkConnector', 0); + childTemplate.resourceCountIs('AWS::Lambda::NetworkConnector', 2); + childTemplate.resourceCountIs('AWS::Lambda::MicrovmImage', 0); template.resourceCountIs('AWS::Lambda::MicrovmImage', 0); template.hasOutput('MicrovmArtifactBucketName', {}); // The build-time connector output is what the packaging script reads next, so @@ -1565,6 +1691,38 @@ describe('AgentStack with the MicroVM gate on but no image configured (first dep }); }); +describe.each([false, 'false'])('AgentStack MicroVM flat-layout escape hatch (%s)', microvmNested => { + let template: Template; + + beforeAll(() => { + const app = new App({ + context: { + compute_type: 'lambda-microvm', + microvm_nested_stack: microvmNested, + microvm_base_image_arn: 'arn:aws:lambda:us-east-1:aws:microvm-image:al2023-1', + microvm_base_image_version: '1', + microvm_artifact_sha256: 'a'.repeat(64), + }, + }); + template = Template.fromStack(new AgentStack(app, 'FlatMicrovmStack', { + env: { account: '123456789012', region: 'us-east-1' }, + })); + }); + + test('preserves flat resource identities when nesting is explicitly disabled', () => { + template.resourceCountIs('AWS::Lambda::MicrovmImage', 1); + template.resourceCountIs('AWS::Lambda::NetworkConnector', 2); + expect(Object.keys(template.findResources('AWS::Lambda::MicrovmImage'))[0]) + .toMatch(/^LambdaMicrovmComputeImage/); + expect(Object.keys(template.findResources('AWS::CloudFormation::Stack'))) + .not.toEqual(expect.arrayContaining([expect.stringMatching(/^MicrovmNestedStack/)])); + const roles = Object.entries(template.findResources('AWS::IAM::Role')) + .filter(([id]) => id.startsWith('LambdaMicrovmCompute')); + expect(roles).toHaveLength(3); + for (const [, role] of roles) expect(role.Properties.RoleName).toBeUndefined(); + }); +}); + describe('AgentStack MicroVM image ARN invariant', () => { let synthError: unknown; @@ -1575,7 +1733,7 @@ describe('AgentStack MicroVM image ARN invariant', () => { const configuredSpy = jest.spyOn(lambdaMicrovmCompute, 'isLambdaMicrovmImageConfigured') .mockReturnValue(true); try { - const app = new App({ context: { compute_type: 'lambda-microvm' } }); + const app = new App({ context: { compute_type: 'lambda-microvm', microvm_nested_stack: true } }); const stack = new AgentStack(app, 'TestAgentStackMicrovmInvariant', { env: { account: '123456789012', region: 'us-east-1' }, }); @@ -1597,7 +1755,7 @@ describe('AgentStack MicroVM image ARN invariant', () => { }); describe('AgentStack solution attribution (#319): AWS_SDK_UA_APP_ID via stack-level aspect', () => { - let template: Template; + let templates: Template[]; beforeAll(() => { const app = new App(); @@ -1611,7 +1769,11 @@ describe('AgentStack solution attribution (#319): AWS_SDK_UA_APP_ID via stack-le Aspects.of(stack).add(new SolutionUaAspect(buildAppId('UaAgentStack')), { priority: AspectPriority.MUTATING, }); - template = Template.fromStack(stack); + templates = [ + Template.fromStack(stack), + ...stack.node.findAll().filter((child): child is NestedStack => NestedStack.isNestedStack(child)) + .map(child => Template.fromStack(child)), + ]; }); // CDK synthesizes its own framework-owned Lambdas that are NOT part of the @@ -1621,22 +1783,24 @@ describe('AgentStack solution attribution (#319): AWS_SDK_UA_APP_ID via stack-le // `AWS679f53fac002430cb0da5b7982bd2287…` `cr.AwsCustomResource` singleton // (which CDK happens to give the env var today, but whose attribution we do // not want to depend on across CDK upgrades). Every framework-owned id is - // enumerated explicitly so the coverage assertion below cannot silently + // enumerated explicitly, including the registry Provider's framework handlers, + // so the coverage assertion below cannot silently // stop covering an ABCA Lambda by relabelling it as "framework". const FRAMEWORK_LAMBDA_ID = - /^(CustomResourceProviderHandler|CustomS3AutoDeleteObjects|CustomVpcRestrictDefaultSG|AWS679f53fac002430cb0da5b7982bd2287)/; + /^(CustomResourceProviderHandler|CustomS3AutoDeleteObjects|CustomVpcRestrictDefaultSG|AWS679f53fac002430cb0da5b7982bd2287|AgentRegistryProviderframework)/; test('every solution Lambda carries AWS_SDK_UA_APP_ID (traverses nested scope)', () => { - const functions = template.findResources('AWS::Lambda::Function'); - const abcaLambdas = Object.entries(functions).filter( + const functions = templates.flatMap(template => Object.entries(template.findResources('AWS::Lambda::Function'))); + const abcaLambdas = functions.filter( ([id]) => !FRAMEWORK_LAMBDA_ID.test(id), ); // exact count — update when adding/removing a Lambda construct (#319). // A loose `toBeGreaterThan` let a whole integration construct disappear // unnoticed; the exact count fails if a Lambda is dropped OR if a new one // is added without being attributed below. - // 47 = 46 on main + RemoveWorkspaceFn (DELETE /v1/linear/workspaces/{slug}). - expect(abcaLambdas.length).toBe(47); + // 47 existing handlers (including concurrency repair) + 2 registry + // provisioning handlers + 4 registry API handlers in nested stacks. + expect(abcaLambdas.length).toBe(54); // Every ABCA-authored Lambda must carry the canonical `#` app-id. Collect // any offenders so a failure names the exact logical id(s) that are naked. const unattributed = abcaLambdas @@ -1653,8 +1817,8 @@ describe('AgentStack solution attribution (#319): AWS_SDK_UA_APP_ID via stack-le // The trap: these functions live inside integration constructs several // scopes below the stack. The env-var still resolves the canonical `#` // form (not the mangled `-` variant). - const functions = template.findResources('AWS::Lambda::Function'); - const nested = Object.entries(functions).filter(([id]) => + const functions = templates.flatMap(template => Object.entries(template.findResources('AWS::Lambda::Function'))); + const nested = functions.filter(([id]) => /Jira|Slack|Linear/.test(id), ); expect(nested.length).toBeGreaterThan(0); @@ -1910,30 +2074,105 @@ describe('AgentStack Linear identity vault gate (#809)', () => { expect([...workloadNames]).toEqual(['abca_linear_oauth_LinearVaultWritersStack']); }); - test('MicroVM + vault is REFUSED by name, not left to the resource counter', () => { - // Pinning a limitation, not a behaviour. The vault IS wired for the MicroVM substrate - // — platform_config carries the workload name and the guest's execution role gets the - // mint grant — but the two cannot be enabled together today: 505 resources against a - // HARD limit of 500 (microvm alone 496, the vault alone 488). Claiming MicroVM support - // without saying so would be false. - // - // The stack refuses the combination itself rather than letting the counter throw, - // because the counter's message is a per-type census that never mentions either flag — - // the operator cannot tell from it what to change. - // - // Reclaiming room means nesting a subsystem. MicroVM (+19 resources) is the cheapest - // candidate and currently deployed nowhere, but nesting it needs the session-role trust - // wiring to stop referencing a child resource (it creates a parent↔child cycle today). - // - // When the room is found, this test should be replaced by a real parity assertion. - const app = new App({ - context: { enableLinearIdentityVault: true, compute_type: 'lambda-microvm' }, - }); - expect(() => Template.fromStack( - new AgentStack(app, 'LinearVaultMicrovmStack', { + describe('MicroVM vault configuration and execution-role grants (#857)', () => { + let template: Template; + let childTemplates: Template[]; + const workloadName = 'abca_linear_oauth_LinearVaultMicrovmStack'; + + beforeAll(() => { + const app = new App({ + context: { + enableLinearIdentityVault: true, + compute_type: 'lambda-microvm', + microvm_nested_stack: true, + microvm_base_image_arn: 'arn:aws:lambda:us-east-1:aws:microvm-image:al2023-1', + microvm_base_image_version: '1', + microvm_artifact_sha256: 'a'.repeat(64), + }, + }); + const stack = new AgentStack(app, 'LinearVaultMicrovmStack', { env: { account: '123456789012', region: 'us-east-1' }, - }), - )).toThrow(/enableLinearIdentityVault cannot be combined with compute_type=lambda-microvm/); + }); + template = Template.fromStack(stack); + childTemplates = stack.node.findAll() + .filter((child): child is NestedStack => NestedStack.isNestedStack(child)) + .map(child => Template.fromStack(child)); + }); + + test('delivers the same workload identity through the coordinator and runtime configuration', () => { + template.hasResourceProperties('Custom::LinearWorkloadIdentity', { + WorkloadName: workloadName, + }); + template.hasOutput('LinearVaultWorkloadName', { Value: workloadName }); + const orchestrator = Object.entries(template.findResources('AWS::Lambda::Function')) + .find(([id]) => id.includes('TaskOrchestratorOrchestratorFn'))!; + expect(orchestrator[1].Properties.Environment.Variables).toMatchObject({ + LINEAR_VAULT_ENABLED: 'true', + LINEAR_WORKLOAD_IDENTITY_NAME: workloadName, + }); + template.hasResourceProperties('AWS::BedrockAgentCore::Runtime', { + EnvironmentVariables: Match.objectLike({ + LINEAR_VAULT_ENABLED: 'true', + LINEAR_WORKLOAD_IDENTITY_NAME: workloadName, + }), + }); + }); + + test('grants the MicroVM role both token operations on the Linear vault resources', () => { + const roles = template.findResources('AWS::IAM::Role'); + const executionRole = Object.keys(roles) + .find(id => id.startsWith('LambdaMicrovmComputeExecutionRole'))!; + expect(executionRole).toBeDefined(); + const policies = Object.values(template.findResources('AWS::IAM::Policy')) + .filter(policy => policy.Properties.Roles?.some( + (role: { Ref?: string }) => role.Ref === executionRole, + )); + const statements = policies.flatMap(policy => policy.Properties.PolicyDocument.Statement); + const arn = (suffix: string) => ({ + 'Fn::Join': ['', [ + 'arn:', { Ref: 'AWS::Partition' }, `:bedrock-agentcore:us-east-1:123456789012:${suffix}`, + ]], + }); + const directory = arn('workload-identity-directory/default'); + const identity = arn(`workload-identity-directory/default/workload-identity/${workloadName}`); + for (const action of [ + 'bedrock-agentcore:GetWorkloadAccessTokenForUserId', + 'bedrock-agentcore:GetResourceOauth2Token', + ]) { + const grants = statements.filter(statement => + [statement.Action].flat().includes(action)); + expect(grants).toHaveLength(1); + expect(grants[0].Effect).toBe('Allow'); + expect(grants[0].Resource).toEqual(expect.arrayContaining([directory, identity])); + expect([grants[0].Resource].flat()).not.toContain('*'); + } + const oauthGrant = statements.find(statement => + [statement.Action].flat().includes('bedrock-agentcore:GetResourceOauth2Token')); + expect(oauthGrant.Resource).toEqual(expect.arrayContaining([ + arn('token-vault/default'), + arn('token-vault/default/oauth2credentialprovider/bgagent-linear-oauth-*'), + ])); + const policyJson = JSON.stringify(policies); + expect(policyJson).toContain('bedrock-agentcore-identity!default/oauth2/bgagent-linear-oauth-*'); + expect(policyJson).not.toContain('bedrock-agentcore:CreateWorkloadIdentity'); + expect(policyJson).not.toContain('bedrock-agentcore:DeleteWorkloadIdentity'); + }); + + test('keeps task vault configuration and credentials out of the reusable image', () => { + const images = childTemplates.flatMap(child => + Object.values(child.findResources('AWS::Lambda::MicrovmImage'))); + expect(images).toHaveLength(1); + const imageJson = JSON.stringify(images); + for (const field of [ + 'LINEAR_VAULT_ENABLED', + 'LINEAR_WORKLOAD_IDENTITY_NAME', + 'access_token', + 'refresh_token', + workloadName, + ]) { + expect(imageJson).not.toContain(field); + } + }); }); test('the source graph names no Linear-minting handler that is unwired', () => { @@ -2008,29 +2247,77 @@ describe('AgentStack CloudFormation resource budget 500 with cushion', () => { // `AWS::CDK::Metadata`. Budget the synthesized number, so add that resource back. const SYNTH_ONLY_RESOURCES = 1; - const COMPUTE_TYPES = ['agentcore', 'ecs', 'lambda-microvm']; - const CELLS = COMPUTE_TYPES.flatMap(computeType => - [false, true].map(enableToolGateway => ({ computeType, enableToolGateway })), - ); + const MICROVM_CONFIGURATIONS = [ + { name: 'microvm-bootstrap', context: { compute_type: 'lambda-microvm' } }, + { + name: 'microvm-imported', + context: { + compute_type: 'lambda-microvm', + microvm_image_identifier: 'arn:aws:lambda:us-east-1:123456789012:microvm-image:existing-agent', + microvm_image_version: '6.0', + }, + }, + { + name: 'microvm-managed', + context: { + compute_type: 'lambda-microvm', + microvm_base_image_arn: 'arn:aws:lambda:us-east-1:aws:microvm-image:al2023-1', + microvm_base_image_version: '1', + microvm_artifact_sha256: 'a'.repeat(64), + }, + }, + ]; + const CONFIGURATIONS = [ + { name: 'agentcore', context: { compute_type: 'agentcore' } }, + { name: 'ecs', context: { compute_type: 'ecs' } }, + ...MICROVM_CONFIGURATIONS.flatMap(configuration => [ + { name: `${configuration.name}-nested`, context: { ...configuration.context, microvm_nested_stack: true } }, + { + name: `${configuration.name}-flat`, + context: { ...configuration.context, microvm_nested_stack: false }, + }, + ]), + ]; + const CELLS = CONFIGURATIONS.flatMap(configuration => + [false, true].flatMap(enableToolGateway => + [false, true].map(enableLinearIdentityVault => ({ + ...configuration, enableToolGateway, enableLinearIdentityVault, + })))); describe.each(CELLS)( - 'compute_type=$computeType enableToolGateway=$enableToolGateway', - ({ computeType, enableToolGateway }) => { + '$name enableToolGateway=$enableToolGateway enableLinearIdentityVault=$enableLinearIdentityVault', + ({ context, enableToolGateway, enableLinearIdentityVault }) => { let template: Template; + let templates: Template[]; beforeAll(() => { - const app = new App({ context: { compute_type: computeType, enableToolGateway } }); + const app = new App({ context: { ...context, enableToolGateway, enableLinearIdentityVault } }); const stack = new AgentStack(app, 'BudgetStack', { env: { account: '123456789012', region: 'us-east-1' }, }); // Throws `TooManyResourcesInStack` if this cell is over the hard quota, so // reaching the assertions below is itself part of the guard. template = Template.fromStack(stack); + templates = [ + template, + ...stack.node.findAll().filter((child): child is NestedStack => NestedStack.isNestedStack(child)) + .map(child => Template.fromStack(child)), + ]; }); test('stays inside the resource budget', () => { - const resourceCount = Object.keys(template.toJSON().Resources ?? {}).length; - expect(resourceCount + SYNTH_ONLY_RESOURCES).toBeLessThanOrEqual(RESOURCE_BUDGET); + for (const item of templates) { + const rendered = item.toJSON(); + const resourceCount = Object.keys(rendered.Resources ?? {}).length; + expect(resourceCount + SYNTH_ONLY_RESOURCES).toBeLessThanOrEqual(RESOURCE_BUDGET); + expect(Buffer.byteLength(JSON.stringify(rendered))).toBeLessThanOrEqual(1024 * 1024); + expect(Object.keys(rendered.Parameters ?? {}).length).toBeLessThanOrEqual(200); + expect(Object.keys(rendered.Outputs ?? {}).length).toBeLessThanOrEqual(200); + } + // Even changing every resource must fit the nested-operation quota. + const total = templates.reduce((count, item) => + count + Object.keys(item.toJSON().Resources ?? {}).length + SYNTH_ONLY_RESOURCES, 0); + expect(total).toBeLessThanOrEqual(2500); }); test('emits no Lambda permission for the API Gateway console test-invoke stage', () => { diff --git a/cdk/test/stacks/microvm-layout-context.test.ts b/cdk/test/stacks/microvm-layout-context.test.ts new file mode 100644 index 000000000..b3a31312f --- /dev/null +++ b/cdk/test/stacks/microvm-layout-context.test.ts @@ -0,0 +1,93 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { App } from 'aws-cdk-lib'; +import { Template } from 'aws-cdk-lib/assertions'; +import { AgentStack } from '../../src/stacks/agent'; + +const env = { account: '123456789012', region: 'us-east-1' }; +const imageContext = { + microvm_base_image_arn: 'arn:aws:lambda:us-east-1:aws:microvm-image:al2023-1', + microvm_base_image_version: '1', + microvm_artifact_sha256: 'a'.repeat(64), +}; + +test.each([{}, imageContext])('rejects an omitted layout before synthesizing MicroVM resources: %j', context => { + const app = new App({ context: { compute_type: 'lambda-microvm', ...context } }); + expect(() => new AgentStack(app, 'ImplicitLayout', { env })) + .toThrow('microvm_nested_stack must be explicitly selected'); +}); + +describe.each([ + ...[true, 'true', false, 'false'].map((value, index) => ({ + name: `ExplicitLayout${index}`, + context: { compute_type: 'lambda-microvm', microvm_nested_stack: value }, + })), + { name: 'Agentcore', context: { compute_type: 'agentcore' } }, + { name: 'Ecs', context: { compute_type: 'ecs' } }, + ...['agentcore', 'ecs'].map(compute => ({ + name: `${compute}WithUnusedPrefix`, + context: { compute_type: compute, microvm_resource_name_prefix: 'retained-config' }, + })), +])('MicroVM layout selection: $name', ({ name, context }) => { + let template: Template; + beforeAll(() => { + template = Template.fromStack(new AgentStack(new App({ context }), name, { env })); + }); + test('accepts explicit MicroVM layouts and leaves other backends unaffected', () => { + expect(template.toJSON().Resources).toBeDefined(); + }); +}); + +describe('MicroVM resource prefix validation', () => { + test.each([false, 'false'])('explains an explicitly disabled nested layout (%s)', value => { + const app = new App({ + context: { + compute_type: 'lambda-microvm', + microvm_nested_stack: value, + microvm_resource_name_prefix: 'migration', + }, + }); + expect(() => new AgentStack(app, 'FlatPrefix', { env })) + .toThrow('microvm_resource_name_prefix cannot be used with microvm_nested_stack=false'); + }); + + test.each([42, true, null])('identifies non-string prefix %s in nested mode', value => { + const app = new App({ + context: { + compute_type: 'lambda-microvm', + microvm_nested_stack: true, + microvm_resource_name_prefix: value, + }, + }); + expect(() => new AgentStack(app, 'InvalidPrefix', { env })) + .toThrow('microvm_resource_name_prefix must be a string'); + }); + + test('accepts a valid prefix with an explicit nested layout', () => { + const app = new App({ + context: { + compute_type: 'lambda-microvm', + microvm_nested_stack: true, + microvm_resource_name_prefix: 'migration', + }, + }); + expect(() => new AgentStack(app, 'DefaultNestedPrefix', { env })).not.toThrow(); + }); +}); diff --git a/cdk/test/stacks/microvm-managed-image-nag.test.ts b/cdk/test/stacks/microvm-managed-image-nag.test.ts new file mode 100644 index 000000000..791889f58 --- /dev/null +++ b/cdk/test/stacks/microvm-managed-image-nag.test.ts @@ -0,0 +1,99 @@ +/** + * MIT No Attribution + * + * Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. + * + * Permission is hereby granted, free of charge, to any person obtaining a copy of + * the Software without restriction, including without limitation the rights to + * use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of + * the Software, and to permit persons to whom the Software is furnished to do so. + * + * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR + * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, + * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE + * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER + * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, + * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE + * SOFTWARE. + */ + +import { ManagedPolicy, PolicyStatement, Role } from 'aws-cdk-lib/aws-iam'; +import { AGENTCORE_AZS_CONTEXT_KEY } from '../../src/constructs/agentcore-azs'; +import { OrchestrationReconciler } from '../../src/constructs/orchestration-reconciler'; +import { TaskOrchestrator } from '../../src/constructs/task-orchestrator'; +import { buildApp } from '../../src/main'; + +const configurations = [false, true].flatMap(enableLinearIdentityVault => + [false, true].flatMap(extraWildcard => + [false, true].map(microvmNested => ({ enableLinearIdentityVault, extraWildcard, microvmNested })))); + +describe.each(configurations)('managed MicroVM security checks (vault=$enableLinearIdentityVault, extra wildcard=$extraWildcard, nested=$microvmNested)', ({ enableLinearIdentityVault, extraWildcard, microvmNested }) => { + let errors: string[]; + + beforeAll(async () => { + const app = await buildApp({ + account: '123456789012', + region: 'us-west-2', + // A concrete region verifies even explicit AZ overrides. Keep this + // template-security test independent of credentials and live EC2. + describeAzs: async () => [ + { zoneName: 'us-west-2a', zoneId: 'usw2-az1' }, + { zoneName: 'us-west-2b', zoneId: 'usw2-az2' }, + ], + appProps: { + context: { + compute_type: 'lambda-microvm', + microvm_nested_stack: microvmNested, + enableLinearIdentityVault, + enableToolGateway: true, + microvm_base_image_arn: 'arn:aws:lambda:us-west-2:aws:microvm-image:al2023-1', + microvm_base_image_version: '1', + microvm_artifact_sha256: 'a'.repeat(64), + [AGENTCORE_AZS_CONTEXT_KEY]: ['us-west-2a', 'us-west-2b'], + }, + }, + }); + const stack = app.node.findChild('backgroundagent-dev'); + const orchestrator = stack.node.findChild('TaskOrchestrator') as TaskOrchestrator; + if (enableLinearIdentityVault) { + const reconciler = stack.node.findChild('OrchestrationReconciler') as OrchestrationReconciler; + // Bundled policies can overflow at different boundaries from unit-test + // policies. Exercise the documented exceptions on both owning roles. + for (const role of [orchestrator.fn.role, reconciler.fn.role]) { + new ManagedPolicy(role as Role, 'OverflowPolicyVaultProbe', { + statements: [ + new PolicyStatement({ + actions: ['bedrock-agentcore:GetResourceOauth2Token'], + resources: ['arn:aws:bedrock-agentcore:us-west-2:123456789012:token-vault/default/oauth2credentialprovider/bgagent-linear-oauth-*'], + }), + new PolicyStatement({ + actions: ['secretsmanager:GetSecretValue'], + resources: ['arn:aws:secretsmanager:us-west-2:123456789012:secret:bedrock-agentcore-identity!default/oauth2/bgagent-linear-oauth-*'], + }), + ], + }); + } + } + if (extraWildcard) { + // Model a future overflow document without relying on the policy + // splitter placing an extra statement in a particular document. + new ManagedPolicy(orchestrator.fn.role as Role, 'OverflowPolicy999', { + statements: [new PolicyStatement({ + actions: ['ssm:GetParameter'], + resources: ['arn:aws:ssm:us-west-2:123456789012:parameter/unrelated-*'], + })], + }); + } + // The real entry point installs cdk-nag. The CLI fails on these assembly + // annotations even though in-process synth itself does not throw. + const artifact = app.synth().getStackByName('backgroundagent-dev'); + errors = artifact.messages.filter(message => message.level === 'error') + .map(message => String(message.entry.data)); + }, 60_000); + + test('accepts documented Jira access while still reporting unrelated wildcard grants', () => { + expect(errors).toEqual(extraWildcard + ? [expect.stringMatching(/AwsSolutions-IAM5.*parameter\/unrelated-\*/)] + : []); + }); +}); diff --git a/cli/src/commands/pending.ts b/cli/src/commands/pending.ts index 2f8266fed..dce05ea87 100644 --- a/cli/src/commands/pending.ts +++ b/cli/src/commands/pending.ts @@ -65,7 +65,7 @@ function renderText(pending: readonly PendingApprovalSummary[]): void { } console.log(` preview: ${p.tool_input_preview}`); console.log(` created: ${p.created_at}`); - console.log(` expires: ${p.expires_at} (timeout_s=${p.timeout_s})`); + console.log(` expires: ${p.expires_at ?? 'no automatic expiry'}`); console.log( ` approve: bgagent approve ${p.task_id} ${p.request_id}`, ); diff --git a/cli/src/commands/submit.ts b/cli/src/commands/submit.ts index 8fd3feec1..0c41f1710 100644 --- a/cli/src/commands/submit.ts +++ b/cli/src/commands/submit.ts @@ -36,6 +36,9 @@ import { INITIAL_APPROVALS_MAX_ENTRY_LENGTH, MAX_BUDGET_USD_MAX, MAX_BUDGET_USD_MIN, + MICROVM_SLEEP_AFTER_S_DEFAULT, + MICROVM_SLEEP_AFTER_S_MAX, + MICROVM_SLEEP_AFTER_S_MIN, } from '../types'; import { exitCodeForStatus, waitForTask } from '../wait'; @@ -85,9 +88,14 @@ export function makeSubmitCommand(): Command { .option( '--approval-timeout ', `Cedar HITL per-task default approval timeout (${APPROVAL_TIMEOUT_S_MIN}-${APPROVAL_TIMEOUT_S_MAX}s). ` - + 'Overrides the platform default of 300s. Per-rule @approval_timeout_s still min-wins at gate-firing.', + + 'Default: no automatic expiry (0). Explicit per-rule timeouts still apply.', parseInt, ) + .option( + '--microvm-sleep-after ', + `Sleep while waiting for approval after this many seconds (default ${MICROVM_SLEEP_AFTER_S_DEFAULT}; ` + + 'off keeps the worker awake). MicroVM only; requires platform sleep support and does not extend approval deadlines.', + ) .option( '--pre-approve ', 'Cedar HITL pre-approval scope to seed at task start (repeatable). ' @@ -141,15 +149,25 @@ export function makeSubmitCommand(): Command { if ( isNaN(opts.approvalTimeout) || !Number.isInteger(opts.approvalTimeout) - || opts.approvalTimeout < APPROVAL_TIMEOUT_S_MIN + || (opts.approvalTimeout !== 0 && opts.approvalTimeout < APPROVAL_TIMEOUT_S_MIN) || opts.approvalTimeout > APPROVAL_TIMEOUT_S_MAX ) { throw new CliError( - `--approval-timeout must be an integer between ${APPROVAL_TIMEOUT_S_MIN} ` + `--approval-timeout must be 0 (no expiry), or an integer between ${APPROVAL_TIMEOUT_S_MIN} ` + `and ${APPROVAL_TIMEOUT_S_MAX} seconds.`, ); } } + let microvmSleepAfterS: number | undefined; + if (opts.microvmSleepAfter !== undefined) { + const raw = opts.microvmSleepAfter as string; + microvmSleepAfterS = raw === 'off' ? 0 : /^\d+$/.test(raw) ? Number(raw) : NaN; + if (!Number.isSafeInteger(microvmSleepAfterS) + || microvmSleepAfterS < MICROVM_SLEEP_AFTER_S_MIN || microvmSleepAfterS > MICROVM_SLEEP_AFTER_S_MAX) { + throw new CliError('--microvm-sleep-after must be off or an integer between ' + + `${MICROVM_SLEEP_AFTER_S_MIN} and ${MICROVM_SLEEP_AFTER_S_MAX} seconds.`); + } + } const preApproveRaw = (opts.preApprove ?? []) as readonly string[]; let initialApprovals: readonly ApprovalScope[] | undefined; if (preApproveRaw.length > 0) { @@ -226,6 +244,7 @@ export function makeSubmitCommand(): Command { ...(prNumber !== undefined && { pr_number: prNumber }), ...(opts.trace && { trace: true }), ...(opts.approvalTimeout !== undefined && { approval_timeout_s: opts.approvalTimeout }), + ...(microvmSleepAfterS !== undefined && { microvm_sleep_after_s: microvmSleepAfterS }), ...(initialApprovals !== undefined && { initial_approvals: initialApprovals }), ...(attachments.length > 0 && { attachments }), }; diff --git a/cli/src/commands/watch.ts b/cli/src/commands/watch.ts index 6cefbfab8..caaaa4c09 100644 --- a/cli/src/commands/watch.ts +++ b/cli/src/commands/watch.ts @@ -113,6 +113,8 @@ const PROGRESS_EVENT_TYPES = new Set([ 'agent_cost_update', 'agent_error', 'agent_blocked', + 'approval_decision_recorded', + 'approval_cancelled', ]); /** Format an event timestamp to a short local time string. */ @@ -154,7 +156,9 @@ function renderMilestoneSuffix(meta: Record): string { if (meta.severity != null) parts.push(`[sev=${String(meta.severity)}]`); if (meta.request_id != null) parts.push(`request_id=${String(meta.request_id)}`); if (meta.scope != null) parts.push(`scope=${String(meta.scope)}`); - if (meta.timeout_s != null) parts.push(`timeout=${String(meta.timeout_s)}s`); + if (meta.status != null) parts.push(`status=${String(meta.status)}`); + if (meta.timeout_s != null) parts.push(meta.timeout_s === 0 ? 'no automatic expiry' : `timeout=${String(meta.timeout_s)}s`); + if (meta.reason != null) parts.push(`reason=${String(meta.reason)}`); const ruleIds = meta.matching_rule_ids; if (Array.isArray(ruleIds) && ruleIds.length > 0) { parts.push(`rules=${ruleIds.map(String).join(',')}`); @@ -177,7 +181,7 @@ function renderMilestoneSuffix(meta: Record): string { } /** Render a single progress event as a human-readable line. */ -export function renderEvent(event: TaskEvent): string { +export function renderEvent(event: TaskEvent, taskId?: string): string { const time = formatTime(event.timestamp); const meta = event.metadata; @@ -208,8 +212,21 @@ export function renderEvent(event: TaskEvent): string { } case 'agent_milestone': { const milestone = String(meta.milestone ?? ''); - return `[${time}] ★ ${milestone}${renderMilestoneSuffix(meta)}`; + let line = `[${time}] ★ ${milestone}${renderMilestoneSuffix(meta)}`; + if (milestone === 'approval_requested') { + if (meta.input_preview) line += `\n Action: ${String(meta.input_preview)}`; + const requestId = String(meta.request_id ?? ''); + if (taskId && [taskId, requestId].every(id => /^[A-Za-z0-9_-]{1,128}$/.test(id))) { + line += `\n bgagent approve ${taskId} ${requestId} --scope this_call`; + line += `\n bgagent deny ${taskId} ${requestId}`; + } + } + return line; } + case 'approval_decision_recorded': + return `[${time}] Decision saved${renderMilestoneSuffix(meta)}`; + case 'approval_cancelled': + return `[${time}] Approval request closed${renderMilestoneSuffix(meta)}`; case 'agent_cost_update': { const cost = meta.cost_usd != null ? `$${Number(meta.cost_usd).toFixed(COST_USD_DECIMALS)}` : '$?'; const input = meta.input_tokens ?? 0; @@ -302,7 +319,7 @@ interface Formatter { emit(ev: TaskEvent): void; } -export function makeFormatter(isJson: boolean): Formatter { +export function makeFormatter(isJson: boolean, taskId?: string): Formatter { return { emit(ev: TaskEvent): void { if (isJson) { @@ -310,7 +327,7 @@ export function makeFormatter(isJson: boolean): Formatter { return; } if (PROGRESS_EVENT_TYPES.has(ev.event_type)) { - console.log(renderEvent(ev)); + console.log(renderEvent(ev, taskId)); } }, }; @@ -561,7 +578,7 @@ export function makeWatchCommand(): Command { throw e; } - const formatter = makeFormatter(isJson); + const formatter = makeFormatter(isJson, taskId); // Task already terminated — print the snapshot tail and exit. if ((TERMINAL_STATUSES as readonly string[]).includes(snapshot.taskStatus)) { diff --git a/cli/src/types.ts b/cli/src/types.ts index e1b6b681a..c3d455f7e 100644 --- a/cli/src/types.ts +++ b/cli/src/types.ts @@ -167,6 +167,8 @@ export interface ErrorClassification { /** Task detail returned by GET /v1/tasks/{task_id}. */ export interface TaskDetail { + /** Configured MicroVM approval-wait delay; 0 disables sleep. Absent on legacy records. */ + readonly microvm_sleep_after_s?: number; readonly task_id: string; readonly status: TaskStatusType; /** ``null`` for a repo-less workflow (#248 Phase 3). */ @@ -451,6 +453,8 @@ export interface CreateTaskResponse extends TaskDetail { /** Create task request body for POST /v1/tasks. */ export interface CreateTaskRequest { + /** MicroVM approval-wait seconds before sleep (0 = off, omitted = 600). Does not extend approval deadlines. */ + readonly microvm_sleep_after_s?: number; /** Optional since #248 Phase 3: repo-less workflows submit without it. */ readonly repo?: string; readonly issue_number?: number; @@ -471,7 +475,8 @@ export interface CreateTaskRequest { */ readonly trace?: boolean; /** Cedar HITL per-task default approval timeout (design §7.3 step 5). - * Valid range ``[APPROVAL_TIMEOUT_S_MIN, APPROVAL_TIMEOUT_S_MAX]``. */ + * Zero retains unanswered requests; positive values use + * ``[APPROVAL_TIMEOUT_S_MIN, APPROVAL_TIMEOUT_S_MAX]``. */ readonly approval_timeout_s?: number; /** Cedar HITL pre-approval allowlist seeded at task start (§7.3 step 4). * Each entry must be a valid ``ApprovalScope``. */ @@ -742,6 +747,7 @@ export type ApprovalStatus = | 'PENDING' | 'APPROVED' | 'DENIED' + | 'CANCELLED' | 'TIMED_OUT' | 'STRANDED'; @@ -793,7 +799,7 @@ export interface PendingApprovalSummary { readonly reason: string; readonly created_at: string; readonly timeout_s: number; - readonly expires_at: string; + readonly expires_at: string | null; /** Cedar rule ids that matched this request — shown by * ``bgagent pending`` so users can see which rule fired without * spelunking TaskEventsTable. */ @@ -839,7 +845,12 @@ export const APPROVAL_TIMEOUT_S_MIN = 30; export const APPROVAL_TIMEOUT_S_MAX = 3600; /** Default approval_timeout_s when the submit payload omits it. */ -export const APPROVAL_TIMEOUT_S_DEFAULT = 300; +export const APPROVAL_TIMEOUT_S_DEFAULT = 0; + +/** Per-task MicroVM sleep delay bounds; zero disables automatic sleep. */ +export const MICROVM_SLEEP_AFTER_S_MIN = 0; +export const MICROVM_SLEEP_AFTER_S_MAX = 3600; +export const MICROVM_SLEEP_AFTER_S_DEFAULT = 600; /** Minimum allowed max_budget_usd (1 cent). * Sourced from ``contracts/constants.json`` via cdk types.ts (#258). */ diff --git a/cli/test/commands/submit.test.ts b/cli/test/commands/submit.test.ts index 0d6d825c4..d9cc70eec 100644 --- a/cli/test/commands/submit.test.ts +++ b/cli/test/commands/submit.test.ts @@ -422,6 +422,28 @@ describe('submit command', () => { }); describe('Cedar HITL extensions', () => { + test.each([['off', 0], ['0', 0], ['30', 30], ['600', 600], ['3600', 3600]])( + 'forwards MicroVM sleep delay %s without changing approval timeout', async (value, expected) => { + mockCreateTask.mockResolvedValue({ task_id: 't-sleep', status: 'SUBMITTED' }); + await makeSubmitCommand().parseAsync([ + 'node', 'test', '--repo', 'owner/repo', '--task', 'ok', '--microvm-sleep-after', String(value), + ]); + const [body] = mockCreateTask.mock.calls[0]; + expect(body.microvm_sleep_after_s).toBe(expected); + expect(body).not.toHaveProperty('approval_timeout_s'); + }, + ); + test.each(['-1', '3601', '1.5', '30seconds', '1e2', '', ' ', 'NaN'])('rejects malformed MicroVM delay %j', async value => { + await expect(makeSubmitCommand().parseAsync([ + 'node', 'test', '--repo', 'owner/repo', '--task', 'ok', '--microvm-sleep-after', value, + ])).rejects.toThrow(/--microvm-sleep-after must be off or an integer/); + expect(mockCreateTask).not.toHaveBeenCalled(); + }); + test('omits the sleep override so the server captures its default', async () => { + mockCreateTask.mockResolvedValue({ task_id: 't-sleep', status: 'SUBMITTED' }); + await makeSubmitCommand().parseAsync(['node', 'test', '--repo', 'owner/repo', '--task', 'ok']); + expect(mockCreateTask.mock.calls[0][0]).not.toHaveProperty('microvm_sleep_after_s'); + }); // --approval-timeout ------------------------------------------------------ test('forwards --approval-timeout as approval_timeout_s', async () => { diff --git a/cli/test/commands/watch.test.ts b/cli/test/commands/watch.test.ts index 2cc65a32f..972e4cd1c 100644 --- a/cli/test/commands/watch.test.ts +++ b/cli/test/commands/watch.test.ts @@ -133,6 +133,8 @@ describe('renderEvent', () => { timeout_s: 300, matching_rule_ids: ['rule-1', 'rule-2'], scope: 'this_call', + reason: 'Needs your permission', + input_preview: 'git push --force', }, }); const output = renderEvent(event); @@ -142,6 +144,19 @@ describe('renderEvent', () => { expect(output).toContain('timeout=300s'); expect(output).toContain('rules=rule-1,rule-2'); expect(output).toContain('scope=this_call'); + expect(output).toContain('Needs your permission'); + expect(output).toContain('Action: git push --force'); + expect(renderEvent(event, 'task-1')).toContain('bgagent approve task-1 req-abc --scope this_call'); + expect(renderEvent(event, 'task-1')).toContain('bgagent deny task-1 req-abc'); + }); + + test('renders a closed approval with its cancellation reason', () => { + const output = renderEvent(makeEvent({ + event_type: 'approval_cancelled', + metadata: { request_id: 'gate', status: 'CANCELLED', reason: 'Task cancelled by its owner' }, + })); + expect(output).toContain('Approval request closed'); + expect(output).toContain('Task cancelled by its owner'); }); test('renders policy_decision milestone metadata', () => { diff --git a/contracts/constants.json b/contracts/constants.json index 07a624ba3..f6ce84503 100644 --- a/contracts/constants.json +++ b/contracts/constants.json @@ -7,7 +7,12 @@ "approval_timeout_s": { "min": 30, "max": 3600, - "default": 300 + "default": 0 + }, + "microvm_sleep_after_s": { + "min": 0, + "max": 3600, + "default": 600 }, "max_budget_usd": { "min": 0.01, @@ -36,12 +41,16 @@ "task_table_name": "TASK_TABLE_NAME", "task_events_table_name": "TASK_EVENTS_TABLE_NAME", "task_approvals_table_name": "TASK_APPROVALS_TABLE_NAME", + "approval_requests_api_url": "APPROVAL_REQUESTS_API_URL", "nudges_table_name": "NUDGES_TABLE_NAME", "log_group_name": "LOG_GROUP_NAME", "artifacts_bucket_name": "ARTIFACTS_BUCKET_NAME", "trace_artifacts_bucket_name": "TRACE_ARTIFACTS_BUCKET_NAME", + "continuation_bucket_name": "CONTINUATION_BUCKET_NAME", "github_token_secret_arn": "GITHUB_TOKEN_SECRET_ARN", "linear_oauth_secret_arn": "LINEAR_OAUTH_SECRET_ARN", + "linear_vault_enabled": "LINEAR_VAULT_ENABLED", + "linear_workload_identity_name": "LINEAR_WORKLOAD_IDENTITY_NAME", "jira_oauth_secret_arn": "JIRA_OAUTH_SECRET_ARN", "agent_session_role_arn": "AGENT_SESSION_ROLE_ARN", "aws_sdk_ua_app_id": "AWS_SDK_UA_APP_ID", @@ -51,6 +60,7 @@ "required": [ "task_table_name", "task_events_table_name", + "approval_requests_api_url", "github_token_secret_arn", "agent_session_role_arn" ], @@ -62,10 +72,38 @@ ], "account_anchor_key": "agent_session_role_arn" }, + "payload_bootstrap": { + "version": 2, + "manifest_prefix": "bootstrap/", + "launch_filename": "launch.json", + "max_manifest_bytes": 16384, + "max_payload_bytes": 8388608, + "url_ttl_seconds": 900, + "minimum_url_lifetime_seconds": 300 + }, "microvm_hook_budgets": { "ready_hook_timeout_seconds": 300, "warmup_total_budget_seconds": 240, - "warmup_required_timeout_seconds": 120 + "warmup_required_timeout_seconds": 120, + "lifecycle_hook_timeout_seconds": 30, + "lifecycle_handler_budget_seconds": 20 + }, + "microvm_lifecycle": { + "protocol_version": 1, + "image_protocol_env": "ABCA_MICROVM_LIFECYCLE_PROTOCOL", + "hook_port": 8080, + "maximum_duration_seconds": 28800 + }, + "microvm_continuation": { + "version": 1, + "lease_key_prefix": "worker-lease#", + "object_key_prefix": "continuations/", + "max_manifest_bytes": 2097152, + "max_workspace_bytes": 1073741824, + "max_conversation_bytes": 16777216, + "park_after_seconds": 3600, + "retirement_margin_seconds": 300, + "verified_sdk_version": "0.2.110" }, "linear_vault": { "_why": "The vault caches a grant keyed by the WHOLE token request, customParameters included. A one-token divergence between any two of these copies makes every resolve a cache miss, which post-#812 is reported as consent-required and can latch a healthy workspace revoked.", diff --git a/contracts/constants.md b/contracts/constants.md index c9224ce3d..fb211fc02 100644 --- a/contracts/constants.md +++ b/contracts/constants.md @@ -21,10 +21,15 @@ the contract. This is the neutral location both runtimes read. |---|---|---| | `agent/src/shared_constants.py` | `/app/contracts/constants.json` | import-time | | `agent/src/policy.py`, `agent/src/jira_reactions.py` | `SHARED_CONSTANTS` | import-time | -| `agent/src/server.py` | `SHARED_CONSTANTS["microvm_platform_config"]`, `SHARED_CONSTANTS["microvm_hook_budgets"]` | import-time | +| `agent/src/payload_bootstrap.py` | `SHARED_CONSTANTS["payload_bootstrap"]` | import-time | +| `cdk/src/handlers/shared/payload-bootstrap.ts`, `cdk/src/constructs/payload-bootstrap-permissions.ts` | `payload_bootstrap` | import-time | +| `agent/src/server.py` | `SHARED_CONSTANTS["microvm_platform_config"]`, `SHARED_CONSTANTS["microvm_hook_budgets"]`, `SHARED_CONSTANTS["microvm_lifecycle"]` | import-time | +| `agent/src/microvm_http.py` | `SHARED_CONSTANTS["microvm_hook_budgets"]` | import-time | +| `agent/src/hooks.py` | `microvm_lifecycle.maximum_duration_seconds`, `approval_timeout_s.max` | SDK matcher construction | | `cdk/src/handlers/shared/types.ts`, `jira-app-actor.ts` | `../../../../contracts/constants.json` | synth-time `import` | | `cdk/src/handlers/shared/strategies/lambda-microvm-strategy.ts` | `microvm_platform_config` | synth-time `import`, read per session start | -| `cdk/src/constructs/lambda-microvm-compute.ts` | `microvm_hook_budgets` | synth-time `import` | +| `cdk/src/constructs/lambda-microvm-compute.ts` | `microvm_hook_budgets`, `microvm_lifecycle` | synth-time `import` | +| `cdk/src/handlers/shared/microvm-image-capability.ts` | `microvm_hook_budgets`, `microvm_lifecycle` | runtime `import` | | `cdk/src/constructs/blueprint.ts` | re-exports from `types.ts` | synth-time | | `cli/test/constants-parity.test.ts` | package-safe literal parity | test-time | @@ -45,7 +50,7 @@ JSON at TypeScript compile time via `resolveJsonModule`. "approval_timeout_s": { "min": 30, "max": 3600, - "default": 300 + "default": 0 }, "max_budget_usd": { "min": 0.01, @@ -77,10 +82,27 @@ JSON at TypeScript compile time via `resolveJsonModule`. "jira_oauth_secret_arn", "agent_session_role_arn"], "account_anchor_key": "agent_session_role_arn" }, + "payload_bootstrap": { + "version": 2, + "manifest_prefix": "bootstrap/", + "launch_filename": "launch.json", + "max_manifest_bytes": 16384, + "max_payload_bytes": 8388608, + "url_ttl_seconds": 900, + "minimum_url_lifetime_seconds": 300 + }, "microvm_hook_budgets": { "ready_hook_timeout_seconds": 300, "warmup_total_budget_seconds": 240, - "warmup_required_timeout_seconds": 120 + "warmup_required_timeout_seconds": 120, + "lifecycle_hook_timeout_seconds": 30, + "lifecycle_handler_budget_seconds": 20 + }, + "microvm_lifecycle": { + "protocol_version": 1, + "image_protocol_env": "ABCA_MICROVM_LIFECYCLE_PROTOCOL", + "hook_port": 8080, + "maximum_duration_seconds": 28800 } } ``` @@ -94,14 +116,25 @@ JSON at TypeScript compile time via `resolveJsonModule`. - **`approval_gate_cap.default`** — value applied when a blueprint omits the field. 50 is the design-decision default (see `docs/design/CEDAR_HITL_GATES.md` decision #13). -- **`approval_timeout_s.min`** — floor for `approval_timeout_s` (§6 - decision #6). 30 seconds — below this, humans cannot realistically - respond to an approval prompt. -- **`approval_timeout_s.max`** — absolute ceiling for `approval_timeout_s` - before the `maxLifetime - 300` clip is applied (§7.3). 3600 seconds - (1 hour). +- **`approval_timeout_s.min`** — minimum positive explicit timeout: 30 seconds. + Zero is separately accepted and means no automatic deadline. +- **`approval_timeout_s.max`** — maximum positive explicit timeout: 3600 seconds + (1 hour). MicroVM continuation keeps this human deadline separate from a + worker's service lifetime. - **`approval_timeout_s.default`** — value applied when the submit payload - omits `approval_timeout_s`. 300 seconds (5 minutes) per §6 decision #6. + omits `approval_timeout_s`: zero, or no automatic expiry. Pending rows have + no DynamoDB TTL; task closure starts retention cleanup. Positive rule + annotations still apply, with the shortest positive deadline winning. +- **`microvm_continuation`** — the versioned checkpoint contract. It pins object + prefixes, maximum archive sizes and the SDK version verified for conversation + and exact budget recovery. A saved approval can release its worker after one + hour when sleep is enabled, or five minutes before the worker's eight-hour + limit when sleep is disabled. The request remains open in both cases. +- **`microvm_sleep_after_s`** — per-task delay before sleeping during a + pending human approval: whole seconds from 0 to 3600, default 600 + (10 minutes). Zero disables sleep. Task creation persists the resolved + preference; only the MicroVM supervisor consumes it. It does not extend + approval deadlines or override the deployment's suspension switch. - **`max_budget_usd.min`** — floor for a task's `max_budget_usd` (1 cent). Validated server-side (`validation.ts`) and pre-validated by `bgagent submit --max-budget` (#258). @@ -121,11 +154,11 @@ JSON at TypeScript compile time via `resolveJsonModule`. variable the agent installs it as (UPPER_SNAKE). This block is unlike the others — it is a **security allowlist**, not a tuning bound. The MicroVM image is a snapshot whose env is frozen at build time, so the agent's non-secret - platform env arrives in the `/run` hook payload instead; the values land in + platform env arrives through an authenticated v2 manifest and task payload instead; the values land in `os.environ`, which makes an unrecognised key an env-injection attempt. The consumer (`agent/src/server.py`) therefore **rejects** any `platform_config` - carrying a key that is not in this map. Values are non-secret identifiers - (table/bucket names, secret ARNs, role ARNs) only. + carrying a key that is not in this map. Values are non-secret configuration: + resource identifiers and the Linear vault enabled flag/workload name. - **`microvm_platform_config.required`** — the subset without which a task cannot run (task + event tables, GitHub secret ARN, session role ARN). A `/run` hook whose `platform_config` misses or blanks any of these is rejected with HTTP 400. @@ -138,15 +171,15 @@ JSON at TypeScript compile time via `resolveJsonModule`. that disagrees with the anchor below on partition or account, is rejected (HTTP 400 `MICROVM_RUN_PLATFORM_CONFIG_INVALID`). This is **fail-fast plus defence in depth, not an ownership proof** — the account-scoped IAM grants are - what actually deny a foreign read, and an in-account redirect is deliberately - *not* covered (see `MICROVM_PLATFORM_CONFIG_ARN_KEYS` in - `agent/src/server.py` for the full statement of what this does and does not buy). + what actually deny a foreign read, and this ARN check alone does not cover an in-account redirect. The v2 bootstrap + authenticates a deployment manifest with IAM and requires the downloaded config + to match it; that separate check rejects same-account workspace substitutions. - **`microvm_platform_config.account_anchor_key`** — which `arn_keys` entry supplies the expected partition + account. Must be `agent_session_role_arn`-shaped: a payload key rather than `os.environ` or an `sts:GetCallerIdentity`, because the environment is empty by construction on this backend (the snapshot bakes nothing) - and this check runs on the path that must make zero AWS calls beyond the S3 - payload fetch. Two invariants follow and both are enforced: the anchor must be in + and the ARN consistency check itself makes no AWS calls. The v2 bootstrap + manifest read and signed task download precede it. Two invariants follow and both are enforced: the anchor must be in `arn_keys`, **and** it must be in `required` — an optional anchor would let a payload disarm the whole check by simply omitting it. @@ -159,6 +192,18 @@ duplication is deliberate: a malformed entry here would silently **widen** what agent accepts from a network payload, so neither side is trusted to be the only gate. +- **`payload_bootstrap`** — the shared ECS/MicroVM v2 transport. `version` is + required on references, manifests and task documents. `manifest_prefix` and + `launch_filename` define the authenticated settings directory and private retry + record. Byte limits bound manifest and task downloads; URL lifetime is at most + `url_ttl_seconds`, shortened by known signer credential expiry, with initial + creation requiring `minimum_url_lifetime_seconds`. The constants checker + rejects nonpositive/noninteger bounds, unsafe/colliding paths, a minimum above + its maximum, and consumers that stop reading the shared block. Changing the + protocol requires matching coordinator, worker images and policies; old unsigned + transports are rejected. See repository runbook + `docs/verification/645-payload-bootstrap.md` for rollout and live checks. + - **`microvm_hook_budgets.ready_hook_timeout_seconds`** — the `/ready` build-hook budget the CDK construct declares to `CreateMicrovmImage` (`READY_HOOK_TIMEOUT_SECONDS` in `cdk/src/constructs/lambda-microvm-compute.ts`). @@ -173,7 +218,7 @@ gate. purpose: a cold 225 MiB `exec` has no predictable duration, which is the lesson of P2-F5. -Unlike every other block here, these three are not independent tuning bounds — they +These three warm-up values are not independent tuning bounds — they are a **relationship**: `warmup_required < warmup_total < ready_hook`. The warm-up must finish inside the budget the service holds the hook to, or a fix for a runtime failure turns into a build failure. A relationship cannot be enforced from one side, @@ -184,6 +229,32 @@ literal re-declaration on **either** side — the Python constants *and* `agent/src/server.py` re-checks the same ordering at import time, so a bad contract fails the drift check *and* the image build. +The runtime lifecycle pair has its own relationship: +`lifecycle_handler_budget_seconds < lifecycle_hook_timeout_seconds`. The 20-second +handler limit covers reading the body, draining activity and checkpoint/refresh +work together. The declared 30-second service hook timeout leaves response headroom. +`microvm_http.py` checks the ordering at import time; the drift script checks +positive integer values, ordering and hardcoded Python/TypeScript redeclarations. +The image declares suspend/resume using this service timeout. These values do not +enable automatic suspension. + +`microvm_lifecycle` owns protocol version `1`, marker name +`ABCA_MICROVM_LIFECYCLE_PROTOCOL`, hook port `8080`, and the 28,800-second +maximum VM lifetime. The strategy sends that lifetime to `RunMicrovm`; the +agent sizes its PreToolUse SDK callback timeout to outlive the same bound by +120 seconds. This callback timeout lets an expired approval finish reconciliation +after a delayed wake; the original approval deadline still controls permission. +Other backends use the maximum approval interval plus the same margin. + +The marker is baked into +the immutable image; it is neither a credential nor task/deployment configuration. +`/validate` rejects a supplied unsupported marker. The coordinator checks the exact +image ARN/version returned by Run, including all six enabled hooks and lifecycle +budgets, before persisting `lifecycleProtocol` alongside that worker's `imageArn` +and `imageVersion`. Legacy or unverified workers cannot start a new suspension. +The drift gate validates this shape and rejects literal copies in its TypeScript +consumers; the artifact-script parity test checks the shell hook/environment JSON. + The published CLI package contains only `lib/`, so it cannot load the repository contract at runtime. It mirrors these values as literals and `cli/test/constants-parity.test.ts` makes drift a CI failure. The standalone diff --git a/docs/AGENTS.md b/docs/AGENTS.md index 896e0b166..3af7434b4 100644 --- a/docs/AGENTS.md +++ b/docs/AGENTS.md @@ -17,6 +17,7 @@ Pre-commit hook `docs-sync` runs sync automatically when prek hooks are installe ## Testing - **Sync + build:** `mise //docs:build` (required before PR if you touched guides, design, ADRs, or `CONTRIBUTING.md`) +- **Rendered links:** `mise //docs:build` checks all internal page and anchor links in `dist/`; external URLs are not checked offline. - **Sync only:** `mise //docs:sync` then `git diff docs/src/content/docs/` — commit mirror changes alongside sources - **Astro check:** `mise //docs:check` (or `cd docs && npm run docs:check`) @@ -29,6 +30,7 @@ CI **"Fail build on mutation"** rejects PRs where committed Starlight mirrors do | `docs/guides/` | WRITE | User and developer guides | | `docs/design/` | WRITE | Architecture and design docs | | `docs/decisions/` | WRITE | ADRs | +| `docs/verification/` | WRITE | Operator runbooks mirrored to `/verification/` | | `docs/imgs/` | WRITE | Static images | | `CONTRIBUTING.md` (repo root) | WRITE | Mirrored to Starlight | | `docs/src/content/docs/` | READ only | Generated — never edit by hand | diff --git a/docs/decisions/ADR-004-tabula-rasa-documentation.md b/docs/decisions/ADR-004-tabula-rasa-documentation.md index 579a476f3..8a6689448 100644 --- a/docs/decisions/ADR-004-tabula-rasa-documentation.md +++ b/docs/decisions/ADR-004-tabula-rasa-documentation.md @@ -44,7 +44,7 @@ Never force a novice to read expert material to proceed. Never force an expert t ### Self-contained references When referencing another document: -- State what the reader gets from it: "See [Deployment Guide](link) for AWS account setup (required before this step)" +- State what the reader gets from it: "See [Deployment Guide](../guides/DEPLOYMENT_GUIDE.md) for AWS account setup (required before this step)" - Never assume the reader has read it - Never use "as mentioned above" — each section must stand alone after context compaction diff --git a/docs/decisions/ADR-016-pluggable-identity-and-auth.md b/docs/decisions/ADR-016-pluggable-identity-and-auth.md index 757f2bcc6..5bf28b7d4 100644 --- a/docs/decisions/ADR-016-pluggable-identity-and-auth.md +++ b/docs/decisions/ADR-016-pluggable-identity-and-auth.md @@ -160,7 +160,7 @@ Per-user `McpCredential` selection requires the Gateway to know *which task-user | P6 | **Trusted task-user identity propagation** (prerequisite for per-user MCP on the general plane, P5): specify + validate a user-scoped inbound identity the Gateway authorizer trusts, replacing the M2M JWT for per-user credential selection. | **Blocks per-user `McpCredential`.** Until done, MCP credentials are workspace-scoped at best. | | P7 | Jira + Slack `ChannelCredential` (same shape as P1); GitHub `GithubOauth2` behind a flag, retire the shared PAT; OBO `act`-claim delegation feeding #237. | Flag-gated; per-surface. | -**Substrate exception — `lambda-microvm` (2026-09-02):** P1's vault cannot be enabled on the MicroVM substrate. The two together synthesize 505 resources against CloudFormation's hard 500-resource limit (MicroVM alone 496, the vault alone 488), so `AgentStack` refuses the combination at synth, naming both context flags. The MicroVM wiring itself is complete — `platform_config` carries the workload name and the guest execution role holds the mint grant — so this is a capacity limit, not a design gap, and it lifts as soon as a subsystem moves into a nested stack. +**Lambda MicroVMs:** The vault uses the guest's compute execution role, with its workload name delivered in authenticated `platform_config`. The former resource-count refusal was removed under #857 after stack reductions and nesting; backend/image-mode tests now include vault and gateway combinations, and check each template's resource and size budgets. See the [Linear setup guide](../guides/LINEAR_SETUP_GUIDE.md#using-the-vault-with-lambda-microvms). **Substrate independence (verified 2026-07-21, both proven live):** the vault path works on any compute. AgentCore Runtime injects the Workload Access Token as the `WorkloadAccessToken` header; ECS/Fargate/Lambda bootstrap it via `GetWorkloadAccessTokenForJWT(workloadName, userToken=)` against a **standalone** (non-service-linked) workload identity, then call `GetResourceOauth2Token`. Runtime-managed (service-linked) workload identities cannot self-vend, so the ECS path needs a manually-created workload identity. The runtime execution role today has `GetWorkloadAccessToken*` but **not** `GetResourceOauth2Token` — P1 adds it, plus `GetSecretValue` scoped to that surface's providers (`bedrock-agentcore-identity!default/oauth2/*`; see the implementation notes below for why this is narrower than the wildcard first anticipated here). diff --git a/docs/decisions/ADR-021-lambda-microvms-compute-backend.md b/docs/decisions/ADR-021-lambda-microvms-compute-backend.md index 49bf51482..2ad7ec195 100644 --- a/docs/decisions/ADR-021-lambda-microvms-compute-backend.md +++ b/docs/decisions/ADR-021-lambda-microvms-compute-backend.md @@ -1,432 +1,220 @@ # ADR-021: AWS Lambda MicroVMs as a third ComputeStrategy backend -> Number: candidate ADR-021 (ADR-020 is the highest accepted on `main`; ADR-018 is claimed by open PR [#548](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/548), ADR-019 by open PR [#663](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/663)). Numbers are never reused. If a lower number frees before merge, renumber and coordinate with those PRs. +> **Implementation status (2026-09-18): P1 and P2 are merged; P3 is in draft review.** Approval sleep/wake, retained requests, conversation/workspace recovery and nested infrastructure have live acceptance evidence. Reusable migration and final integration checks remain open; see [verification status](../verification/README.md). This ADR defines P1–P3, not an official P4. **Status:** proposed **Date:** 2026-07-29 ## Context -ABCA selects a per-repo compute backend through the Blueprint's `compute_type` field (`cdk/src/handlers/shared/repo-config.ts`). Two backends exist today, resolved by `resolveComputeStrategy` (`cdk/src/handlers/shared/compute-strategy.ts`) behind a uniform `ComputeStrategy` interface (`startSession` / `pollSession` / `stopSession`): +ABCA selects compute per repository through Blueprint `compute_type`. AgentCore Runtime is the default; ECS Fargate supports larger workloads. [#645](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/645) adds Lambda MicroVMs for explicit lifecycle control and reduced compute usage during human approval waits. -- **AgentCore Runtime** (`agentcore`, default) — managed Firecracker MicroVM per session. Invoked via `InvokeAgentRuntime`; liveness is inferred from agent heartbeats in DynamoDB plus the FastAPI `/ping` endpoint (`agent/src/server.py`) — the strategy's `pollSession` is a stub that always reports `running`. Constraints: 2 GB image limit, no substrate-level suspend API exposed to the orchestrator. -- **ECS on Fargate** (`ecs`) — always-on Fargate task (16 vCPU / 120 GB, ARM64) for repos that exceed AgentCore's limits ([#596](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/596)). Invoked via `RunTask` in batch mode (bypassing the HTTP server); liveness via `DescribeTasks`. No suspend — an idle task blocked on an approval wait burns full compute the whole time. - -**AWS Lambda MicroVMs** (launched 2026-06-22) is a serverless Firecracker sandbox primitive that AWS positions explicitly for AI coding agents: VM-level isolation, snapshot-based near-instant launch, **suspend/resume with full memory + disk state preserved** (compute charges stop while suspended), a dedicated JWE-authenticated HTTPS endpoint per instance, lifecycle hooks (`/run`, `/suspend`, `/resume`, `/terminate`), and up to 8 hours per session. It is **not** classic Lambda: the 15-minute function cap does not apply, and [COMPUTE.md](../design/COMPUTE.md)'s "Lambda: poor fit" verdict refers to functions, not MicroVMs. - -[#645](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/645) proposes adding Lambda MicroVMs as a third `ComputeStrategy`. The fit is strong but not free — pre-implementation review of the service's lifecycle model surfaced real design tensions this ADR resolves: +Lambda MicroVMs are managed Firecracker virtual machines, separate from ordinary Lambda functions. A MicroVM can preserve memory and disk while suspended and can live for eight hours, including suspended time. The function service’s fifteen-minute limit does not apply. ### Capability comparison (delta rows only — full matrix in COMPUTE.md) -| | AgentCore Runtime | ECS on Fargate | **Lambda MicroVMs** | +| Capability | AgentCore Runtime | ECS Fargate | Lambda MicroVMs | |---|---|---|---| -| Isolation | MicroVM (managed) | Task-level (Firecracker) | MicroVM (Firecracker) | -| Max duration | 8 h | No cap | 8 h (running + suspended) — **verified** (`L-B430C318 = 8 hours`) | -| Suspend/resume | No orchestrator-visible API | No | **Yes** — explicit API + idle policy, state preserved, no compute charge while suspended. **Verified**: suspend reaches `SUSPENDED` in ~1 s with no `idlePolicy`; resume restores `RUNNING` in ~1 s with `microvmId` **and** `endpoint` byte-identical | -| Resources | AgentCore-managed | 16 vCPU / 120 GB / 20–200 GB disk | **Baseline 8 GiB RAM / 4 vCPU, auto-scaling to a 32 GiB / 16 vCPU peak**; 32 GB disk. `minimumMemoryInMiB` configures the BASELINE (max 8,192 MiB); the service scales vertically on demand — capacity is baseline-priced with 4× burst headroom | -| Packaging | ECR image ≤ 2 GB | ECR image, no hard cap | **Zip + Dockerfile in S3 → service-built snapshot image** (versioned, storage billed) | -| Invocation | `InvokeAgentRuntime` (SigV4) | `RunTask` + container overrides | `RunMicrovm` (image **ARN** required — a bare name is rejected) → dedicated HTTPS endpoint + JWE token (`CreateMicrovmAuthToken`, ≤ 60 min TTL) | -| Liveness | Agent heartbeat + `/ping` | `DescribeTasks` | MicroVM state (RUNNING / SUSPENDED / TERMINATED) via control-plane API **and** agent heartbeat (see sub-decision 1) | -| Session storage | `/mnt/workspace` FUSE (no `flock()`) | Ephemeral disk | Native disk in snapshot — **survives suspend/resume, `flock()` works** | -| Architecture | ARM64 | ARM64 | ARM64 (Graviton) | -| Regions (launch) | Broad | Broad | 5 (us-east-1/2, us-west-2, eu-west-1, ap-northeast-1) | - -> Rows marked **verified** were discharged empirically on 2026-07-31 (us-east-1); see `docs/verification/645-p1-lambda-microvm-runbook.md`. Two originally documented constants were **refuted** by that run and are corrected throughout this ADR: the `runHookPayload` cap (16 KB → 4 KB) and the claim that `RunMicrovm` accepts a bare image name (it does not). -> -> **Source hierarchy for service facts.** Where sources disagree, the higher tier wins, and the tier is recorded next to the claim: -> -> 1. **Live boundary probe** — a request the service actually accepted or rejected in this account/Region. Strongest, and the only tier that can refute the others. (Both corrections above came from here.) -> 2. **Modeled constraints** — the CLI/SDK request schema and enumerated allowed values (`--generate-cli-skeleton`, `ARM_64`, `ENABLED|DISABLED`). Authoritative about request *shape*; silent about runtime behaviour. -> 3. **Service developer guide** — including its sizing/scaling tables. Authoritative about *semantics* a probe cannot see, which is exactly how the memory row was fixed: a probe can only show that 32,768 MiB is rejected as a `minimumMemoryInMiB` value; only the guide explains that the field is a BASELINE and that the service scales to a 32 GiB peak on its own. A boundary probe is the strongest evidence about a boundary and says nothing about what the boundary *means*. -> 4. **SDK docstrings** — generated, and demonstrably stale here: `runHookPayload` documents "Maximum: 16,384 bytes" against an enforced 4,096. -> 5. **Launch blogs / skills / toolkit material** — orientation only; never load-bearing on its own. -> -> **Omitted API fields mean "service default", never "none".** Two live findings drove this rule: leaving `ingressNetworkConnectors` unset attaches a PUBLIC `HTTP_INGRESS` connector, and leaving `/ready` out of `hooks` makes the image un-creatable once any lifecycle hook is enabled. So for any security-relevant field, the desired posture must be **requested explicitly** and the test must assert the **outcome** (the ARN present in the request, the hook enabled) rather than the omission (`expect(field).toBeUndefined()`) — an omission assertion passes just as happily when the service is silently choosing something wider. -> -> On the **memory** row specifically: `CreateMicrovmImage` enumerates the baseline sizes a base image supports — `[512, 1024, 2048, 4096, 8192]` MiB for `al2023-1` — and rejects anything else, which is why the construct validates against that list at synth. The 32 GiB / 16 vCPU peak is reached by the service's own vertical scaling, not by asking for it. The account memory quota (`L-CD1C0CC4`, 1024 GB, "burst up to 4×") is an aggregate across MicroVMs, not a per-VM limit; note that concurrency arithmetic should be done against the PEAK, not the baseline, since that is what a busy fleet can actually consume. +| Packaging | ECR image, 2 GB limit | ECR image | ZIP + Dockerfile in S3 → versioned snapshot | +| Duration | Eight hours | No task duration cap | Eight hours, running + suspended | +| Explicit suspend/resume in ABCA | Unsupported | Unsupported | Control-plane APIs; memory and disk retained | +| Storage | Ephemeral disk plus preview persistent FUSE mount | Configurable ephemeral disk | 32 GB native disk; supports `flock()` | +| Sizing | Service-managed | Configurable; larger sustained workloads | ABCA baseline 8,192 MiB; service guide lists up to 32 GiB / 16 vCPU | +| Invocation/liveness | Invoke API + agent heartbeat | RunTask/DescribeTasks | RunMicrovm/GetMicrovm + agent heartbeat | + +See [COMPUTE.md](../design/COMPUTE.md) for the full comparison and costs. Suspending stops compute charges, but snapshot storage and save/restore charges remain. A shorter sleep delay does not guarantee lower total cost. + +Recorded P1 probes established a 4,096-byte `runHookPayload` limit, an image-ARN requirement and accepted baseline values of 512, 1,024, 2,048, 4,096 and 8,192 MiB. Those observations override conflicting generated SDK descriptions for the tested account/Region. They do not measure guest-visible launch memory, vertical-scaling latency or sustained workload fit. The service guide’s capacity figures and live observations must remain distinguishable. ### Design tensions the strategy must resolve -1. **Idle detection is inbound-traffic-based; the ABCA agent is outbound-only.** MicroVM idle policies suspend when no traffic arrives at the *endpoint*. A busy agent running a 40-minute build receives no inbound traffic and would be suspended mid-work by a naive idle policy. Conversely, "no inbound traffic" is the agent's *normal* state. -2. **No self-suspend.** The agent cannot suspend its own MicroVM from inside; only an external `SuspendMicrovm` call can. Suspend decisions must be owned by the orchestrator — which aligns with the unified liveness model proposed in [#491](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/491). -3. **Snapshots bake state.** The image snapshot is captured once at build time; every MicroVM resumes from it. Secrets, tokens, and per-task identity must arrive at run time (`runHookPayload`, ≤ 4 KB, or fetched in the `/run` hook), never at image build. CSPRNG reseeding is a consequence of the same property and is scoped to **P3** — see the amended note under sub-decision 2 and the risk bullet, which record why the exposure is negligible today. -4. **Auth tokens are short-lived.** JWE tokens max out at 60 minutes; any orchestrator→agent HTTP interaction over the endpoint needs token refresh, unlike AgentCore's SigV4 invoke or ECS's no-endpoint model. -5. **Identity delta — narrower than it looks.** Most AgentCore services ABCA uses are standalone and substrate-portable: Memory is already consumed from ECS via an IAM grant plus `MEMORY_ID` (`EcsAgentCluster`), and Gateway (ADR-019) is portable by design (SigV4 inbound). The genuinely Runtime-coupled piece is the workload-access-token **delivery mechanism** (`runtimeUserId` → `WorkloadAccessToken` request header → `BedrockAgentCoreContext`, used by `resolve_linear_api_token()`), which has no MicroVM equivalent. The ECS backend already lives with this delta (env-var token delivery); MicroVMs inherit the same posture until the pluggable identity work ([#249](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/249), ADR-016) redesigns the seam. -6. **The service's defaults are not our posture.** Two of them, both discovered live: `RunMicrovm` attaches a **public** `HTTP_INGRESS` connector (and mints a public `*.lambda-microvm..on.aws` endpoint) when `ingressNetworkConnectors` is omitted, and `CreateMicrovmImage` **requires** the `/ready` hook whenever any lifecycle hook is enabled. Neither posture can be reached by leaving a field out — each needs an explicit control (see sub-decisions 3 and 4). +1. Traffic-based idle policies observe inbound endpoint traffic. A busy coding agent mostly sends outbound requests, so absence of inbound traffic cannot establish that it is idle. +2. Suspension needs an external controller. The agent can prepare a safe checkpoint, but the coordinator owns the service call. +3. Image snapshots share build-time process state. Credentials, task identity and deployment configuration must arrive after launch. +4. A worker’s eight-hour lifetime is shorter than an unanswered approval may remain useful. Durable task state must outlive the worker. +5. New packaging, IAM and lifecycle behavior need live checks; a successful CDK synth cannot validate service semantics. ## Decision -Adopt **AWS Lambda MicroVMs as a third, opt-in `ComputeStrategy` backend** named `lambda-microvm`, selected per repo via Blueprint `compute_type`. AgentCore remains the default. Five sub-decisions: +Add `lambda-microvm` as an opt-in `ComputeStrategy`. AgentCore remains the default. ### 1. Strategy shape: extend the interface with mandatory suspend/resume -`ComputeType` widens to `'agentcore' | 'ecs' | 'lambda-microvm'` (mirrored in `cli/src/types.ts` and the CLI's inline unions). `SessionHandle` gains a `{ strategyType: 'lambda-microvm', microvmId, endpoint }` variant — `microvmId` because every lifecycle API (`suspend-microvm`, `resume-microvm`, `terminate-microvm`, `get-microvm`, and `create-microvm-auth-token` — the latter not called in P1–P3, see sub-decision 3) takes only the MicroVM identifier, and `endpoint` because it is per-session (minted by `RunMicrovm`) and required for any orchestrator→agent HTTP interaction. Note the naming seam: the handle field is `microvmId` (matching `RunMicrovmResponse.microvmId`), while the request key on every lifecycle command is `microvmIdentifier` — the strategy is the only place that translates between the two. The image ARN is deliberately **not** in the handle: like the ECS task definition ARN, it is deployment-time configuration consumed by `startSession` (from construct-injected environment) and recorded in the session-start log entry for diagnostics, not per-session lifecycle state. - -**A second, sharper naming seam: `imageIdentifier` must be an ARN.** The name suggests a bare image name is acceptable — `create-microvm-image --name` takes one, and this ADR originally assumed `run-microvm --image-identifier` would too. It does not: a bare name is rejected with `ValidationException: Malformed ARN - doesn't start with 'arn:'`, and so is `list-microvm-image-builds --image-identifier ` (`Invalid ARN format`). The construct therefore resolves an operator-supplied name to its exact `arn:${Partition}:lambda:${Region}:${Account}:microvm-image:${Name}` ARN **once** — the same value it scopes the lifecycle IAM grant to — and injects THAT as `MICROVM_IMAGE_IDENTIFIER`. One derivation, two consumers, so a request field and an IAM resource can never disagree. The strategy validates the invariant and fails fast with the remedy, because the service's own error names neither the env var nor the fix. - -The `ComputeStrategy` interface gains **mandatory** `suspendSession(handle)` / `resumeSession(handle)` methods returning a typed result (`{ supported: false } | { supported: true }`-shaped, exact type at implementation time); the widening lands in P3 (see sub-decision 5), in one commit across all three strategies. Mandatory-with-explicit-stub is the codebase idiom, not optional methods: no behavioral interface in the codebase has an optional method, `AgentCoreComputeStrategy.pollSession` is already a mandatory explicit stub rather than an optional member, and the exhaustive-`never` switch culture means a fourth backend must make a compile-checked decision about its suspend semantics instead of silently falling through a `strategy.suspendSession?.()` feature-detection. The agentcore and ecs strategies return `unsupported` (not a silent success — a suspend that silently no-ops would let the orchestrator believe compute billing stopped when it did not); the orchestrator gates its suspend policy on the typed response, consistent with how `pollTaskStatus` already branches explicitly on `computeType`. - -**Poll semantics — the strategy reports, the orchestrator interprets.** `pollSession(handle)` receives only the session handle and cannot see task state, so the health rules must live where the DynamoDB status lives. `SessionStatus` gains a `'suspended'` variant; the strategy maps `GetMicrovm` state mechanically and the **orchestrator** cross-references against the task row — the same division of labor `finalPollState` already uses for ECS (substrate stopped + non-terminal DynamoDB status → failed) and `pollTaskStatus` uses for agentcore heartbeats: substrate `suspended` + task `AWAITING_APPROVAL` is healthy (orchestrator-intended suspend); `suspended` with any other task status is an anomaly to surface, not fail-fast; substrate terminal + non-terminal task status → classify failed. +All strategies implement `startSession`, `pollSession`, `stopSession`, `suspendSession` and `resumeSession`. The lifecycle methods return an explicit supported/unsupported result. AgentCore and ECS return unsupported without making suspension API calls; MicroVM bounds each control request. An accepted API request does not prove the transition completed. -**Liveness on this backend is substrate state AND agent heartbeat.** The substrate cross-check above answers one question — "is the VM still there?" — and P2 established that it is not sufficient on its own. `GetMicrovm` catches a MicroVM that *died*; it cannot catch a MicroVM that is alive and reporting `RUNNING` while the pipeline **inside the guest** is hung, deadlocked, or was OOM-killed. That state is not hypothetical or self-correcting on this substrate, and the P2 live run narrowed *why* without weakening the conclusion. The service does reap a VM whose run hook FAILS (a 4xx makes it terminate within ~12 s — see sub-decision 2), so the P1 evidence for this paragraph (a hook-less image sitting in `RUNNING` indefinitely with no `stateReason`) no longer describes an ABCA image. What the service reaps is a hook *result*; it has no view into the guest afterwards. So the surviving — and more realistic — hang case is a `/run` hook that returned **200** and a pipeline that then hung, deadlocked or was OOM-killed behind it: the substrate stays `RUNNING`, the service is satisfied, and nothing else notices. Left to the substrate check alone, such a task would burn the orchestrator's full ~8.5 h poll window — billing an 8-hour reservation — before the safety net fired. +The MicroVM handle records `microvmId`, `endpoint`, the actual launched `imageArn`/`imageVersion`, and verified `lifecycleProtocol` when available. Lifecycle request fields use `microvmIdentifier`. Save the known handle before optional image discovery so a failed capability check cannot lose cleanup ownership. Current deployment settings cannot establish an older worker’s capabilities. -The in-guest half is already being written: the agent updates `agent_heartbeat_at` on the task row unconditionally, with no backend awareness, so the timestamp exists on every substrate. Only the orchestrator's *reaction* to it was backend-scoped — `pollTaskStatus` evaluated staleness for `agentcore` alone — which is the gap P2 closes by extending it to `lambda-microvm`. The grace and stale thresholds are the SAME on both: the timestamp is written by the same pipeline code at the same cadence, so a backend-specific window would encode a difference that does not exist. The two signals stay complementary rather than redundant — the substrate check is the crash detector, the heartbeat is the hang detector — and the check remains scoped to task status `RUNNING`, which is what keeps a deliberately suspended VM during an approval wait (sub-decision 2, P3) from being read as a dead one. `ecs` is deliberately left out: `DescribeTasks` reports a real container exit *with an exit code* (OOM-kill included) and the ECS poll block already interprets it with its own patience counters, so adding the heartbeat there would give one backend two independently-tuned kill paths for the same failure. - -The service's `MicrovmState` enum has **six** members, not three, so the mapping is stated exhaustively (one line of rationale each, mirrored in the strategy's doc comment): - -| `MicrovmState` | `SessionStatus` | Why | +| Service state | Strategy result | Coordinator responsibility | |---|---|---| -| `PENDING` | `running` | Still booting; the same way ECS's `PENDING`/`PROVISIONING` map to `running`. | -| `RUNNING` | `running` | — | -| `SUSPENDING` | `suspended` | Already on its way to frozen; reporting `running` would tell the orchestrator compute is still progressing when it is not. Both suspend states land on a report the orchestrator treats as benign-or-anomalous depending on task status, never as failure. **Never observable in practice** — suspend reached `SUSPENDED` in under 1 s live — so it is mapped for completeness and nothing may *wait* for it. | -| `SUSPENDED` | `suspended` | — | -| `TERMINATING` | `completed` | Terminal-bound and carries no exit code, so "the substrate is gone" is all the strategy can honestly say. | -| `TERMINATED` | `completed` | Success vs failure is the orchestrator's call — it cross-references the DynamoDB status. **This is the load-bearing terminal signal**, not `NotFound` (see below). | -| *unrecognized* | `running` | A future service enum addition must never fail a healthy task; the strategy warns and keeps polling. | - -**`GetMicrovm` `ResourceNotFoundException` → `completed`, but as a LATE fallback.** This deliberately diverges from `ecs-strategy`, where `DescribeTasks` returning no task maps to `failed`. ECS keeps stopped tasks describable for roughly an hour, so a missing task there really is anomalous; a MicroVM is eventually reaped from the control plane **by design**, so `failed` would fail tasks that finished cleanly. - -What the live run corrected is the *timing*, not the mapping. `NotFound` is **not the near-term terminal signal**: a terminated MicroVM reported `TERMINATING` at +1 s, `TERMINATED` at +3 s, and was **still `TERMINATED` ~10 minutes later** and at every subsequent checkpoint — `ResourceNotFoundException` was never observed in that window. So the branch that actually fires in practice is `TERMINATED → completed` in the table above; the `NotFound` rule covers only a VM reaped after a long gap (a poll resumed after a crash, a stranded-task reconciler sweep). Both are required and neither substitutes for the other: without the `TERMINATED` row the orchestrator would poll a finished VM until its safety net fired, and without the `NotFound` rule a late sweep would classify a cleanly-finished task as a poll error. - -The divergence is safe because it does not weaken detection: the orchestrator still fails the task when a terminal report lands while the DynamoDB status is non-terminal, so a genuine mid-run disappearance is caught — it simply receives the substrate-failure classification instead of a misleading poll error. Because that cross-check acts on a status read earlier in the same poll cycle, the orchestrator **re-reads the task row before failing** (the normal shutdown order is "agent writes terminal status → agent exits → VM terminates", which a stale read would otherwise turn into a spurious failure); ECS buys the same protection with a five-consecutive-poll patience counter instead. - -Neither the mapping nor the NotFound rule is a health decision: both are mechanical restatements of substrate state, which is what keeps the "strategy reports, orchestrator interprets" split intact. - -Normative requirements (EARS, per [ADR-020](./ADR-020-ears-requirements-syntax.md)): - -- When a task's Blueprint sets `compute_type: 'lambda-microvm'`, the orchestrator shall resolve the `LambdaMicrovmComputeStrategy` via `resolveComputeStrategy`. -- When `startSession` is invoked, the strategy shall call `RunMicrovm` with `maximumDurationInSeconds` set to 28 800 (the service maximum, matching AgentCore's 8-hour session cap and sitting inside the orchestrator's ~8.5 h safety-net poll window). -- When `startSession` is invoked, the strategy shall pass a fully-qualified MicroVM image ARN as `imageIdentifier`. -- If the configured image identifier is not an ARN, then the strategy shall fail the session start with an error naming the environment variable and the redeploy remedy, before performing any AWS call. -- When `startSession` returns, the orchestrator shall persist the MicroVM handle (`microvmId`, `endpoint`) in the task row's `compute_metadata` (the field `cancel-task.ts` already reads ECS handles from). -- The strategy shall omit `idlePolicy` on every `RunMicrovm` call, in every phase. -- The orchestrator shall be the sole initiator of suspension, via `suspendSession`. -- When `pollSession` observes MicroVM state `SUSPENDED` or `SUSPENDING`, the strategy shall report `suspended` without interpreting task state. -- When `pollSession` observes MicroVM state `TERMINATED` or `TERMINATING`, the strategy shall report `completed` (the observable terminal state persists for at least ~10 minutes, so `completed` shall not depend on the MicroVM being reaped). -- When `pollSession` observes a MicroVM state it does not recognize, the strategy shall report `running`. -- If `GetMicrovm` reports that the MicroVM does not exist, then the strategy shall report `completed`. -- If the strategy reports a terminal substrate state while the task's DynamoDB status is non-terminal, then the orchestrator shall re-read the task row and, if it is still non-terminal, classify the task as failed with a substrate-failure remedy. -- If the strategy reports `suspended` while the task's DynamoDB status is not `AWAITING_APPROVAL`, then the orchestrator shall surface an anomaly event and shall not fail-fast the task. -- While a `lambda-microvm` task's DynamoDB status is `RUNNING`, if the task's `agent_heartbeat_at` is stale (or absent past the grace window) by the same thresholds the orchestrator applies to `agentcore`, then the orchestrator shall treat the session as unhealthy and stop polling — the substrate `GetMicrovm` check shall remain the crash detector, and the heartbeat shall be the in-guest hang detector. -- The task-detail API response shall include `agent_heartbeat_at`, and the CLI shall surface it while the task is non-terminal (P2r2-F11: the field drove the orchestrator's hang detector but was never projected, so no operator could observe the signal — and its invisibility produced a wrong verification conclusion). -- The task-**summary** API response (`GET /v1/tasks`) shall also include `agent_heartbeat_at`, and `bgagent list` shall render it as an age column. Extending the field to the list response is a deliberate widening of the requirement above rather than an incidental one: the detail-only projection makes liveness a per-task question, and an operator checking a fleet of tasks one `bgagent status` at a time is exactly how the hung task P2r2-F11 describes went unnoticed. Same suppression rule as the detail view (terminal tasks and never-beaten tasks render a placeholder) so the two views cannot disagree. -- If `suspendSession` or `resumeSession` is invoked on a strategy that does not support suspension, then the strategy shall return an explicit unsupported result. -- When the agent process reaches a terminal state, the agent shall exit. -- When the orchestrator finalizes a `lambda-microvm` task, the orchestrator shall call `terminate-microvm` (termination shall not rely on any substrate timeout, and shall not rely on the MicroVM self-terminating — it does not). - -*On the omitted `idlePolicy`:* if the block is present all three fields are required, so omission is the unambiguous disabled state the invariant test asserts. This deliberately forgoes `suspendedDurationSeconds` — it lives *inside* `idlePolicy` and cannot be set without re-enabling the traffic-idle machinery — so the suspended-state bound is `maximumDurationInSeconds` plus orchestrator termination and the stranded-approval reconciler (see sub-decision 2). A tighter substrate-level suspended-TTL remains available later as an additive `idlePolicy` change if operators want it. *On the fixed `maximumDurationInSeconds`:* no wall-clock task budget exists in the platform (budgets are `max_turns` / `max_budget_usd`), so the value is parity with AgentCore's 8 h cap rather than derived policy; a Blueprint override can be added later if a real need appears. - -### 2. Lifecycle: suspend/resume reconciled with the agent-owned approval poll - -The headline economic win is suspend during **HITL approval waits** (Cedar approval gates, [CEDAR_HITL_GATES.md](../design/CEDAR_HITL_GATES.md)): while a task waits on a human decision, the MicroVM is suspended (compute charges stop; memory/disk state — cloned repo, warm build caches — is preserved) and resumed when the decision lands. Under Cedar decision #6 the approval window is bounded (default 300 s, ceiling 1 h, timeout → deny), so the saving per gate is bounded at ~1 h of compute — real at 16 vCPU, and it makes any future extension of gate ceilings (the off-hours posture §14.8 deliberately defers) cheap on this backend. - -The handshake must respect the existing approval mechanics: the agent **discovers decisions itself** by polling DynamoDB (`_poll_for_decision`, monotonic timeout), the approve/deny Lambda writes only the decision rows, and `AWAITING_APPROVAL` holds the concurrency slot (Cedar decision #7). Nothing "delivers" an approval to the agent, and suspension freezes the agent's monotonic clock — so the design is: - -- **Suspend — orchestrator-owned.** The orchestrator's durable poll observes `AWAITING_APPROVAL` on a `lambda-microvm` task and calls `suspendSession` after a grace period, and only when the gate's remaining window exceeds grace + resume overhead (suspending a 30 s gate is pure loss). Suspend is a policy decision on a poll observation, not a user action. -- **Resume — inline in the approve/deny Lambdas, orchestrator poll as backstop.** After the transactional decision write commits, `ApproveTaskFn`/`DenyTaskFn` load the MicroVM handle from the task row's `compute_metadata` (persisted at session start — the same field `cancel-task.ts` reads ECS handles from) via a post-commit strongly-consistent `GetItem`, then call `resumeSession` **best-effort**: on failure they log a warning and write a resume-orphan task event; the decision response never fails on a compute error (the decision row is already durable). The orchestrator poll reconciles: decision row present + MicroVM still `SUSPENDED` → retry resume (idempotent). +| RUNNING or starting | `running` | Read task state and applicable heartbeat | +| SUSPENDING / SUSPENDED | `suspended` | Reconcile the approval and lifecycle intent | +| TERMINATING / TERMINATED / not found | `completed` | Re-read task state; distinguish completion, recoverable checkpoint and failure | +| Unknown future state | `running`, with warning | Continue bounded observation | - *Why inline rather than poll-only — codebase precedent:* resume-on-approve is structurally identical to task cancellation — a user-initiated, latency-sensitive action whose purpose is an immediate compute-lifecycle side effect. `cancel-task.ts` already resolves this exact tension: the API-plane handler invokes ECS `StopTask` / AgentCore `StopRuntimeSession` inline, best-effort (a failed stop logs a warning and the state transition stands; a `task_cancel_compute_orphan` event is written when no stoppable compute handle exists, `reason: missing_runtime_handle`) — with the conditional IAM wired in `task-api.ts`. The resume path goes one step further than the precedent by also writing the orphan event on *failed* resume calls, because a failed resume strands a suspended VM awaiting a decision — a stronger liveness consequence than a failed stop of an already-cancelled task. The alternative (orchestrator-poll-only resume) preserves single-owner lifecycle purity but pays up to a full poll interval (~30 s) of latency on every approval, and the purity argument was already litigated and declined for cancel. `approve-task.ts` is deliberately minimal today (security-critical ownership comparison, Cedar finding #6); the resume call is therefore added *after* the transaction commits, cannot alter the decision outcome, and carries one conditional `lambda:ResumeMicrovm` grant — the same blast-radius trade the cancel handler accepted in review. -- **Timeout under freeze — the agent re-bases on the wall clock it already owns.** The agent's monotonic gate timer freezes while suspended, so resuming near the deadline is not enough: the frozen timer would still hold its remaining budget and fire the deny minutes *after* the user-visible window — colliding with the approval row's TTL (`created_at + timeout_s + 120s`) and triggering the "row reaped → stranded" fallback on a healthy gate. Instead, the gate expires at **`min(monotonic budget, created_at + timeout_s)`**, evaluated on each poll iteration and on `/resume`. This is not a new principle: Cedar decision #6 is already "min wins" for timeouts, the wall-clock deadline is already durable in the approval row the agent itself writes (`created_at` is in the agent's own clock domain — no skew), and §13.12's late-approval race fix already establishes that the durable row is authoritative over the agent's local timer. Deny authority stays agent-side (the conditional `TIMED_OUT` write + ConsistentRead re-read race protection is untouched); the orchestrator's resume at `deadline − margin` is purely the wake-up mechanism, with no correctness role. -- **Backstops, not mechanisms.** `maximumDurationInSeconds` (mandatory on every `RunMicrovm`, pinned at 28 800 s — see sub-decision 1) is the substrate kill switch bounding running **and** suspended time; the orchestrator's finalization `terminate-microvm` is the active cleanup path; the stranded-approval reconciler retains its role for orphaned waits. No `idlePolicy`-based bound is used in any phase — see sub-decision 1's omit-`idlePolicy` invariant. +A strategy result is not a task outcome. A terminal substrate can leave a recoverable pending-approval checkpoint; otherwise a non-terminal task needs failure classification. Capacity release separately requires confirmed physical shutdown, not merely the strategy’s `completed` result. - **The active terminate is still mandatory on the SUCCESS path, and P2 sharpened why.** P1 concluded flatly that "nothing self-terminates": a hook-less MicroVM reached `RUNNING` in 12 s and stayed there with no `stateReason` through every checkpoint. P2 refuted that *for the failure path only* — with `run: ENABLED`, a run hook that answers 4xx makes the **service** terminate the VM within ~12 s, `stateReason: "Run lifecycle hook returned HTTP status 400. Please check your hook endpoint and application logs for more details."`, after which `suspend-microvm` correctly refuses it. That is a real improvement in cost posture and a direct benefit of declaring hooks (see also the failure-path row in the phasing table, sub-decision 3). - - It does **not** relieve the orchestrator of anything, because the two cases are disjoint. The service reaps a hook *result* it did not like; it has no view of the guest once the hook returned 200. So a task that starts normally — the overwhelming majority — has no service-side reaper at all, and a VM whose pipeline finished, crashed after `/run`, or hung is reaped by nobody but `TerminateMicrovm`. A leaked handle therefore remains a cost incident that bills until the 8 h cap; only the "the guest rejected its own payload" corner now cleans itself up. -- **Concurrency slot stays held** during suspend. Cedar decision #7's rationale ("container alive, consuming memory") weakens under suspend, and the harder replacement rationale — "AWS counts `SUSPENDED` MicroVMs toward the account memory quota, so releasing ABCA's slot would not free real capacity" — is **undischarged**: the suspended VM stayed in `list-microvms` at every checkpoint, but that only proves *listed*. `L-CD1C0CC4` (1024 GB, account-scoped) exposes no `UsageMetric`, `AWS/Usage` carries only `CallCount` per API, and no MicroVM memory metric exists in any namespace, so consumption is **not observable safely** — proving it would need a large concurrent fleet. The conclusion (hold the slot) stands as the conservative choice, not as a verified fact. Size the arithmetic against the 32 GiB **peak** rather than the 8 GiB baseline: a busy fleet scales up, so peak is what actually competes for the account quota. - -The agent's `/suspend` hook flushes progress events (durable writes before returning 200, within the 60 s hook budget); `/resume` reseeds CSPRNGs and refreshes cached credentials. - -**Amended (P2 review): CSPRNG reseeding is P3 scope, and the P2 exposure is negligible.** The original EARS requirement below implied a `/run`-time reseed had to land with P2; it did not, and shipping P2 with the requirement unmet-but-asserted was itself the defect. Measured exposure, which is what changed the scoping: - -- The only consumer of a non-cryptographic PRNG in `agent/src` is `progress_writer.py`'s `getrandbits(80)`, used for the random half of a **ULID**. `os.urandom` / `secrets` are not seeded from the snapshot at all, and nothing in the agent derives a key, token, or nonce from `random`. -- That ULID is a DynamoDB **sort key under a `task_id` partition**. A collision therefore needs two events in the *same task* at the *same millisecond* with the same 80 random bits — and identical PRNG state across two MicroVMs restored from one snapshot does not produce that, because the events are in different partitions. The worst case is a duplicate progress event within one task, not a security boundary. -- There is no credential exposure from the snapshot's PRNG state: build-role credentials are kept out of the snapshot structurally (the build hooks make zero AWS calls, so `boto3.DEFAULT_SESSION` is never populated — see `_aws_silent_log`), and per-task credentials arrive via `platform_config.agent_session_role_arn` at `/run`. - -So the reseed moves to P3 alongside `/suspend` + `/resume`, where a *resumed* VM — which really does continue with the exact PRNG state it was frozen with, repeatedly — makes it load-bearing rather than theoretical. P3 must reseed on both `/run` and `/resume`, and must not treat the ULID as the only consumer: any future use of `random` for anything security-relevant needs the reseed in place first. +AgentCore and MicroVM start a heartbeat writer every 45 seconds. Their RUNNING tasks use the same 120-second startup grace and 240-second stale threshold. ECS’s batch entrypoint does not start that writer, so applying this check to ECS would fail healthy tasks. Heartbeats establish writer liveness, not progress of every coding thread. Returning from approval to RUNNING refreshes the heartbeat atomically. Task detail/list APIs and CLI views expose the timestamp. Normative requirements (EARS): -- While a `lambda-microvm` task is in `AWAITING_APPROVAL` and the gate's remaining window exceeds the configured grace period plus resume overhead, the orchestrator shall call `suspendSession` after the grace period. -- When the approve or deny Lambda commits a decision for a `lambda-microvm` task, the Lambda shall load the MicroVM handle from `compute_metadata` and call `resumeSession` best-effort. -- If the inline resume fails, then the Lambda shall record a resume-orphan task event and shall still return the decision outcome. -- While a decision row exists and the MicroVM remains `SUSPENDED`, the orchestrator shall retry `resumeSession`. -- While a `lambda-microvm` task waits on an approval gate, the agent shall evaluate gate expiry as the earlier of its monotonic budget and the row's wall-clock deadline (`created_at + timeout_s`), on each poll iteration and on `/resume`. -- If gate expiry is reached without a decision, then the agent shall deny. -- If no decision arrives by the gate's wall-clock deadline minus the resume margin, then the orchestrator shall resume the MicroVM so the agent can evaluate expiry and fire the deny agent-side. - -### 3. Packaging: same agent image source, new build path - -The existing agent container (`agent/` Dockerfile, already ARM64) is repackaged as a zip + Dockerfile artifact in S3 and built into a versioned `MicrovmImage` via `CreateMicrovmImage`. The agent runs its existing FastAPI server (`agent/src/server.py`) — the MicroVM path uses the HTTP entrypoint like AgentCore, not ECS's batch bypass — plus the runtime lifecycle hooks (`/run`, `/suspend`, `/resume`, `/terminate`) and the `/ready` + `/validate` build hooks, all on the same port the server already listens on (8080, declared as the image's `hooks.port`). Runtime hooks are fast-notification only (1–60 s): `/run` validates the payload and starts the pipeline **asynchronously**, mirroring how the agent loop already runs in a background thread behind `/ping` on AgentCore. - -**Hook phasing — corrected: `/ready` + `/run` are both P1.** The original plan split *declaring* a hook from *serving* it, putting `/run`'s declaration in P1 and its implementation in P2. Live verification proved that split is **not a reachable service state**, on two independent counts: - -- `CreateMicrovmImage` rejects an image that enables any lifecycle hook without `/ready`: *"The ready (/ready) MicroVM image hook must be enabled when any MicroVM lifecycle hook (run, resume, suspend, or terminate) is enabled."* So a P1 image declaring only `/run` is **not creatable**. -- With `/ready` added but unserved, both chipset builds fail: *"Ready hook check failed: the application returned a client error (HTTP 4xx) response."* So a declared hook must be served in the same phase. -- And an image with **no** hooks at all — the only other creatable shape — cannot receive a payload: *"The run hook must be enabled in the MicroVM image to pass the run hook payload."* So deferring hooks entirely also defers the whole payload-delivery channel. - -The phasing is therefore: +- When a Blueprint selects `lambda-microvm`, the orchestrator shall use the MicroVM strategy and persist its handle in `compute_metadata`. +- Every launch shall use a full image ARN, `maximumDurationInSeconds=28800`, explicit `NO_INGRESS` and no `idlePolicy`. Invalid image configuration shall fail before launch. +- Before allowing suspension, the coordinator shall verify the actual launched image version and lifecycle protocol. +- When a suspended state conflicts with task state, the coordinator shall record and reconcile the anomaly within bounded recovery windows. +- When the task ends, the coordinator shall actively terminate its MicroVM. The service lifetime is a backstop, not the normal cleanup mechanism. +- Uncertain starts, lost replies and supervisor replay shall preserve the original worker identity, ownership and lifetime rather than launch duplicate workers. -| Hook | Declared by | Served by the agent | Notes | -|---|---|---|---| -| `/ready` | **P1** (construct enables `hooks.microvmImageHooks.ready`) | **P1** | MANDATORY, not a quality nicety — see above. A 200 proves uvicorn is bound and `server` imported cleanly (pulling in `pipeline` → `runner` → the policy engine), so a missing policy file fails the BUILD instead of the first task. **Since P2-F5 it also WARMS the snapshot** — the hook's 200 is what the service waits for before capturing the snapshot, making this the only place a warm page can be created, and the 225 MiB `claude` binary was cold in it (see the P2-F5 correction below). A required warm-up failure answers 503, so a snapshot that cannot exec the agent's own CLI fails the image build instead of every task. Still makes ZERO AWS calls, logging included (a `--version` exec is neither an AWS call nor a network call). | -| `/run` | **P1** (construct sets `hooks.microvmHooks.run`) | **P1** | The payload-delivery channel. Must be served in P1 because `/ready` forces hooks to exist at all, and a hook-less image cannot accept `runHookPayload`. Since P2 it is also the **platform-configuration** channel (see "Platform configuration delivery" below). | -| `/validate` | **P2** (construct sets `hooks.microvmImageHooks.validate`) | **P2** | An **image** (build-time) hook, and a **shallow self-check only**: server alive, hook routes registered, interpreter + contract sanity. It runs under the BUILD role, which deliberately holds no Bedrock / Secrets Manager / DynamoDB grants, so it must make **zero AWS API calls** and must not touch credential resolution — the "deeper warm-up assertions (Bedrock reachability, Memory access, tool availability)" this ADR originally assigned here are **not implementable**: every one of them would `AccessDenied` and fail every image build. They belong to the first task's own error handling. 200 when the checks pass, 503 while still initialising. | -| `/terminate` | **P2** (construct sets `hooks.microvmHooks.terminate`) | **P2** | Best-effort final flush: a final structured log line, then 200 — always, inside the hook budget, even with nothing running. It must **not** write terminal task status (the orchestrator finalizes the task and *then* calls `TerminateMicrovm`, so a terminate hook that wrote a status would race that finalization and could clobber the real outcome) and must not join the pipeline thread. There is nothing buffered to flush: `_ProgressWriter` does a synchronous `put_item` per event, so durability is per-write. "Always 200" also covers the BODY: the handler reads the raw request rather than a typed model, because a typed body is validated before the handler runs and would answer 422 to malformed JSON — a reported hook failure on a successful teardown. Safe to declare because `TerminateMicrovm` removes the VM with or without in-guest cooperation. **Correction (P2-F8):** the service sends `microvmId: ""` on this hook, unlike `/run` where it is populated, so an empty id is expected-normal and this hook cannot join the guest record to the control-plane one — `/run`'s accepted line carries that correlation instead. | -| `/suspend`, `/resume` | **P3** | **P3** | Declaring a runtime hook the agent does not serve fails the corresponding lifecycle transition, so each is declared only in the phase that implements it. P1 termination is the orchestrator's `TerminateMicrovm`, which needs no in-guest cooperation. | +### 2. Lifecycle: suspend/resume reconciled with the agent-owned approval poll -Consequence to state plainly, replacing the original "a P1-built MicroVM image is not runnable end to end": **a P1 image is creatable, launchable and payload-deliverable, but carries no smoke-parity guarantee.** P1 delivers the strategy, the construct, the roles/buckets/connectors, the image resource, the packaging script, and the `/ready` + `/run` endpoints — so a `lambda-microvm` task can start a MicroVM and hand it a payload. What P1 has **not** established is anything P2 owns: AgentCore Memory grants and `MEMORY_ID` delivery, the agent's non-secret env parity inside the snapshot, egress specifics from a running MicroVM, and heartbeat/progress behaviour end to end. No clone → change → PR run has happened on this substrate. P2 ("smoke parity") is the phase that closes that gap. The construct and the packaging script both surface exactly this at synth/run time (`abca:microvm-image-p1-smoke-unverified`) so an operator cannot mistake a launchable substrate for a verified one. +Unanswered approvals have no deadline by default (`approval_timeout_s=0`). Explicit task deadlines range from 30 to 3,600 seconds; a positive policy-rule deadline can also apply. The independent sleep preference defaults to 600 seconds per approval wait. `microvm_sleep_after_s=0` keeps the task awake. -**The `AWS::Lambda::MicrovmImage` L1 enforces the API's enums (P2-F2, live 2026-08-06).** This closes the one item P1 left explicitly open, and it closes it against the construct's own stated reasoning. CloudFormation's generated types make `cpuConfigurations[].architecture` and all four `hooks.*` fields plain strings and document no allowed values, from which P1 concluded that the CloudFormation surface takes a *hook path* while the API takes an `ENABLED`/`DISABLED` flag, and that both were correct for their own surface. CloudFormation refused the change set at **early validation** — the stack was never touched, so there was no rollback and no runtime symptom to trace back — on five values: +Automatic suspension also requires the deployment’s `microvm_approval_suspend_enabled` opt-in, which defaults false for new deployments. A live Parameter Store switch lets existing durable executions stop initiating new suspensions without changing their pinned Lambda environment. The verified normal deployment has this opt-in enabled. Turning it off does not abandon already-suspended workers. -``` -/aws/lambda-microvms/runtime/v1/run is not a valid enum value. Supported values: [DISABLED, ENABLED] - (at /Resources/…/Properties/Hooks/MicrovmHooks/Run) … and the same for Terminate, Ready, Validate -arm64 is not a valid enum value. Supported values: [ARM_64] - (at /Resources/…/Properties/CpuConfigurations/0/Architecture) -``` +The coordinator observes the exact pending gate and waits for the sleep delay. It skips sleep when a timed gate has too little time left before its wake margin. The guest holds a coding barrier, drains acknowledged progress and commits a checkpoint before accepting `/suspend`. Lifecycle HTTP responses explicitly close their connections before freeze to avoid reuse of a stale pooled connection; see the [transport evidence summary](../verification/README.md#recorded-acceptance). -Three consequences. First, the CloudFormation surface is **identical** to the API surface, and the packaging script (`--cpu-configurations '[{"architecture":"ARM_64"}]'`, `--hooks '{"microvmHooks":{"run":"ENABLED",…}}'`) had it right all along. Second, the "CDK-managed (recommended)" bootstrap path was **non-functional** for the whole of P1 and P2 — the out-of-band `--create-image` script was the only working path — and no unit test, `cdk synth` or cdk-nag rule could see it, because the types accept any string. Third, **hook paths are not configurable on either surface**: the service calls fixed well-known routes (proved by the build and run logs, which POST to exactly the `/aws/lambda-microvms/runtime/v1/*` paths the agent serves), so the route constants in the construct are an agent-side cross-package contract ONLY and must never be sent as property values again. Also discharged in passing: the `microvmImageHooks` property name and nesting are correct — CloudFormation resolved `…/Hooks/MicrovmImageHooks/Ready` and objected only to its value. +Approval and denial handlers commit the decision first. They then read the current handle consistently, persist wake intent and request resume best-effort. A wake failure records diagnostics and does not undo the accepted decision. The durable supervisor retries and observes both service state and guest consumption of the decision; RUNNING alone does not prove the tool was released. -**A snapshot is only as warm as the pages touched before it was captured (P2-F5, live 2026-08-07).** This is the defect that stopped the P2 smoke run one step short of a pull request, and it is a property of the substrate rather than a bug in any one file. Every task failed at turn 0, reproducibly: +A sleeping worker retains its concurrency reservation. For a longer wait, ABCA verifies a complete, version-pinned conversation/workspace checkpoint, fences the attempt, confirms shutdown and releases the reservation. A later decision can admit one replacement through the original published coordinator. It restores Git state, required workspace files, the actual SDK conversation, exact pending tool inputs and cumulative usage with fresh scoped credentials. See the [continuation protocol](../design/ORCHESTRATOR.md#retained-microvm-approvals). -``` -TimeoutExpired: Command '['claude', '--version']' timed out after 10 seconds -``` +Normative requirements (EARS): -The binary was fine — in the identical image, locally, `claude --version` answers `2.1.191 (Claude Code)` in under a second. It is a **225 MiB (236,305,136-byte) statically-linked ELF** that nothing had exec'd before the snapshot was taken, so on a guest restored ~50 s earlier the first `exec` had to fault all of those pages in from lazily-restored storage, and 10 s was not enough. `/ready` existed precisely so "the snapshot is taken with a warm server", and the snapshot was warm for uvicorn and stone cold for the binary that does all the work. +- While all sleep gates and guest safety checks hold, the coordinator shall suspend after the configured delay; it shall remain the sole initiator of suspension. +- When a decision is committed, the API shall preserve that outcome even if optional wake or replacement dispatch fails. +- A timed request shall retain its original wall-clock deadline and, within the original process, its monotonic cap. The earlier limit wins; wake/replacement shall not restart either decision window. +- Before an unanswered timed gate reaches its deadline, the coordinator shall wake the worker with the configured margin. For a retired attempt, it may resolve expiry atomically while admitting a replacement. +- Before releasing a worker reservation, the coordinator shall verify checkpoint integrity, fence ownership and confirm physical shutdown. An uncertain control request shall not release capacity. +- A replacement shall preserve the task/request/tool identity and usage totals, obtain fresh scoped credentials and use a new authenticated launch reference. +- Pending requests shall have no storage TTL. Cancellation or terminal cleanup shall close unanswered requests and preserve recorded decisions without extending an existing retention TTL. +- The guest shall reseed application PRNG state from OS entropy on `/run` and `/resume`; security-sensitive values shall continue to use cryptographic randomness. -Both halves of the fix are kept, because they answer different questions. `/ready` now **exec's the heavyweight binaries before returning 200** (`claude` required, `git`/`node` best-effort), which is the only mechanism that can make the shipped snapshot warm — and its own budget rises to 300 s, well inside the 3600 s build-hook window, because it now does work whose duration is a cold `exec`. Two structural rules keep that honest, because **per-command timeouts do not compose**: the required command runs FIRST with its own budget so no best-effort warm-up can starve the one that decides whether the snapshot is usable, and the best-effort ones then SHARE the remainder of a total warm-up ceiling that sits inside the hook budget with margin (240 s against 300 s). Without them, three commands at 120 s each would be 360 s — a fix for a runtime failure that produces a build failure instead — and a single hung optional command could hold up a 200 that the required warm-up had already earned. Separately, the version probe's timeout goes from 10 s to 60 s: a probe that exists to print a version string into a log line gains nothing from a tight bound and loses the whole task when it trips. The general rule this generalises to, and the reason it belongs in the ADR rather than only in a comment: **on this backend, a first-touch cost that other substrates pay during container start is deferred to the first task instead**, so anything large and lazily-loaded is a turn-0 hazard unless it is touched in `/ready`. +The platform verifies ownership and the exact approved action. It does not implement a separate semantic relevance checker; the agent decides whether the proposed work still makes sense. -**Payload delivery** reuses the ECS strategy's S3-pointer pattern, adapted to `runHookPayload` (**≤ 4 KB** — measured, see below): payloads that fit ride inline; the rest are uploaded by the strategy to a platform payload bucket (the ECS payload bucket pattern in `ecs-agent-cluster.ts`: orchestrator write access, compute-role read-only scoped to the bucket, lifecycle expiry on objects) with the S3 URI in `runHookPayload` in place of the payload itself — the MicroVM **execution role** holds the read grant, exactly as the ECS task role does today. Since P2 the hook body also carries `platform_config` in both branches, so `runHookPayload` is never *only* the URI (see the canonical shapes below). +### 3. Packaging: same agent image source, new build path -The cap is **4 096 bytes**, not the 16 384 the SDK documents. Measured exactly: 4 096 passes, 4 097 is rejected with *"Value at 'runHookPayload' failed to satisfy constraint: Member must have length less than or equal to 4096"*. Two consequences follow. First, the original threshold would have inlined every envelope between 4 097 and 16 384 bytes and had the service reject all of them. Second, and more structurally: **the S3-pointer path is now the dominant one, and inline is the exception.** A hydrated task payload (prompt + issue thread + repo context) essentially always exceeds 4 KB, so "small payloads ride inline" describes tiny repo-less prompts rather than the common case. The payload bucket is therefore not a rarely-exercised overflow valve but a required part of every normal task, which raises its lifecycle rule (`MICROVM_PAYLOAD_TTL_DAYS`) and the execution role's read grant from edge-case plumbing to load-bearing. +Package the existing ARM64 `agent/Dockerfile` and its local inputs into a deterministic ZIP. Managed builds use `microvm-images/agent-artifact-.zip`; deploy the digest with the managed base-image ARN/version. A changed object URI triggers CloudFormation to build a new image version. The first deployment may create only infrastructure so the artifact bucket exists before upload. An external image identifier supports out-of-band builds. -**Canonical wire shapes.** Three, and the producer (`lambda-microvm-strategy.ts`) emits exactly these: +All six hooks share the FastAPI listener on port 8080. AWS hook properties accept `ENABLED`/`DISABLED`, not route paths; the architecture enum is `ARM_64`. -| Where | Exact shape | +| Hook | Contract | |---|---| -| `runHookPayload`, inline branch | `{"agent_payload": {…}, "platform_config": {…}}` | -| `runHookPayload`, pointer branch | `{"agent_payload_s3_uri": "s3://…", "platform_config": {…}}` | -| the object at that S3 URI | `{…agent_payload fields…, "platform_config": {…}}` — the payload's own fields at the TOP level, with the config merged in beside them | - -Two asymmetries are deliberate and must not be "tidied" without changing both sides. First, **the S3 object is not the envelope**: the payload's fields sit at the top level (that is what P1 uploaded, before `platform_config` existed) rather than nested under an `agent_payload` key. Second, **`platform_config` is duplicated** on the pointer path — once beside the pointer, once inside the uploaded object. It costs a few hundred bytes and buys the property that the config is reachable whichever end of the fetch a reader looks at, which matters because it is the agent's only substitute for an env block. - -The agent's reader is deliberately more permissive than this contract: it also accepts an S3 object shaped like the envelope (`agent_payload` nested), and `platform_config` present in only one of the two places (the fetched object wins, the hook body is the fallback). Those are **defensive compatibility** for the independent deploy cadences of a snapshot image and the orchestrator Lambda — a tolerant reader, not an alternative contract. A producer must emit the three shapes above. - -**Platform configuration delivery (P2): payload-sourced, allowlisted, fail-closed.** The other two backends hand the agent its non-secret platform env at launch — AgentCore Runtime env vars, ECS container overrides — and there is no equivalent on this substrate: a MicroVM starts from a **snapshot**, so its process environment is whatever was frozen at *image build* time and is then replayed by every MicroVM launched from that image version. Baking the deployment's identifiers into the snapshot would make them **version-frozen**: a redeploy that renames a table, adds a bucket or rotates the session role would leave every existing image version describing a deployment that no longer exists, and the drift would surface as a task-time `ResourceNotFound` rather than a deploy-time error. So the values travel with the task instead: `platform_config` is a SIBLING of the payload — beside `agent_payload` in the inline branch, beside `agent_payload_s3_uri` in the pointer branch, and merged in beside the payload's own fields inside the S3 object (the canonical shapes above give each one exactly) — whose snake_case keys the agent installs into `os.environ` as their UPPER_SNAKE equivalents. A payload value therefore **wins** over any pre-existing/image value — the orchestrator is describing the live deployment, the snapshot is describing a past one. `platform_config` carries **non-secret identifiers only** (table and bucket names, secret ARNs, the session-role ARN); secrets are still fetched at `/run` time from Secrets Manager using those ARNs, so the snapshot-must-stay-secret-free requirement above is untouched. Per-task fields — `memory_id` and friends — stay inside `agent_payload`: `platform_config` configures the *process*, `agent_payload` describes the *task*. - -Two rules make it safe. First, **the allowlist fails closed**: the agent installs a fixed set of keys and *rejects the entire run* (HTTP 400, nothing spawned, not one key installed) if the block carries anything else. These values become environment variables of the process that spawns the agent's tool subprocesses, so an unrecognised key is an attempt to set an arbitrary variable in the agent (`AWS_ENDPOINT_URL`, `LD_PRELOAD`, `PATH`, …) — an injection attempt, not a forward-compatibility gap, which is why unknown keys are refused rather than filtered out. Second, **installation happens before any credential or pipeline initialisation** on the hook path: the very next step reads `GITHUB_TOKEN_SECRET_ARN` to resolve the GitHub token and `AGENT_SESSION_ROLE_ARN` to scope the task's credentials, so installing later would silently resolve the whole task against the snapshot's frozen env. The one call that must precede installation is the S3 payload fetch (the config is *inside* the fetched object), which therefore runs on the ambient compute role via the attributed platform client — and it is the ONLY one: the same rule covers **logging**, so every `/run` log line before the install is stdout-only. The CloudWatch writer would otherwise resolve credentials and pin a boto3 default session (region included) off whatever a snapshot happened to bake, which is the build-hook defect one phase later. Nothing is lost — in the intended deployment there is no baked `LOG_GROUP_NAME`, so those lines would have gone to stdout anyway, and the reason for every pre-install rejection also travels in the structured 4xx/5xx body the service surfaces. A **required subset** (task table, task-events table, GitHub token secret ARN, session-role ARN) is rejected as `…_INCOMPLETE` when missing or blank — a distinct wire code from the `…_INVALID` allowlist rejection, because the remedies differ (deployment wiring vs. producer bug). A `/run` envelope with *no* `platform_config` at all is still accepted, loudly warned: the image snapshot and the orchestrator Lambda deploy on independent cadences, and a new image must not require a same-instant orchestrator. The key set is a cross-package contract in `contracts/constants.json` (`microvm_platform_config`), consumed by the agent's `/run` hook and produced by the orchestrator, with shape and required-subset invariants enforced by `scripts/check-constants-sync.ts`. - -**No orchestrator→agent HTTP path exists in P1–P3**: payload arrives through the `/run` hook, all agent work is outbound, and therefore **no JWE auth tokens are minted at all** — token minting (and its ≤ 60 min TTL refresh problem) is deferred until a real consumer exists (e.g. operator shell access, [#391](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/391)). The `endpoint` stays in the `SessionHandle` because it is genuinely per-session state that becomes load-bearing the day such a consumer appears. But note the service does not agree by default: omitting `ingressNetworkConnectors` on `RunMicrovm` attaches a **public** `HTTP_INGRESS` connector, so the strategy passes the Lambda-managed `NO_INGRESS` connector explicitly on every launch (see sub-decision 4's security table). - -**Constraint accepted:** the configured **baseline is 8 GiB RAM / 4 vCPU** and the service scales vertically to a **32 GiB / 16 vCPU peak** on its own, with 32 GB of disk. So capacity is baseline-priced with 4× burst headroom — good for the bursty compile-and-test shape of an agent task — but the SUSTAINED ceiling is still 32 GiB, so repos that motivated the 120 GB ECS sizing stay on `ecs`. What the construct configures (and validates) is the baseline; the peak is not something a deployment asks for. - -Normative requirements (EARS). **Each requirement's own `(Pn)` tag is authoritative**; there is no blanket phase for the list. The tags are per-requirement because the original list *was* split P1/P2 on the assumption that a hook could be declared in one phase and served in a later one — which the service does not permit (see the phasing table above), so the hook-serving requirements collapsed into P1 while the P2 items below arrived with the P2 hooks and `platform_config`: - -- (P1) The image build shall not embed secrets, tokens, or per-task identity in the snapshot. -- (P1) If the task payload exceeds the 4 KB `runHookPayload` limit, then the strategy shall upload the payload to the platform payload bucket and pass its S3 URI in `runHookPayload` in place of the payload (from P2, alongside `platform_config`). -- (P1) The MicroVM execution role shall hold read-only access to the payload bucket, scoped to that bucket. -- (P1) Where no ingress is configured for a deployment, the strategy shall pass the Lambda-managed `NO_INGRESS` network connector on every `RunMicrovm` call (the field shall not be omitted). -- (P1) Where the image enables any MicroVM lifecycle hook, the image shall also enable the `/ready` hook and the agent shall serve it. -- (P1) When the `/run` hook receives the task payload, the agent shall validate it, start the pipeline asynchronously, and return HTTP 200 within the hook budget. -- (P1) If the `/run` hook payload cannot be resolved to a task payload, then the agent shall reject the hook with a client error and shall not start a pipeline. -- (P1) The agent shall not execute the clone→verify→PR pipeline on the hook path. -- (P1) The agent shall resolve credentials at `/run` time. -- (P1) Where a deployment configures a MicroVM image before smoke parity is verified, the platform shall warn that the backend has no smoke-parity guarantee. -- (P2) Where the image declares the `/validate` build hook, the agent shall serve it. -- (P2) The `/validate` hook shall make no AWS API calls. -- (P2) When the `/run` hook receives `platform_config`, the agent shall install only allowlisted keys into the environment before pipeline initialization. -- (P2) If `platform_config` carries a key that is not on the allowlist, then the agent shall reject the run with a 400 and shall install none of the block's keys. -- (P2) If a required `platform_config` key is missing, then the agent shall reject the run with a 400. -- (P2) Where a `platform_config` value and an image-baked environment value disagree, the agent shall use the `platform_config` value. -- (P2) Until `platform_config` is installed, the `/run` hook shall make no AWS API call other than the payload fetch, and shall log to stdout only. -- (P2) Where the image declares the `/terminate` hook, the agent shall return 200 within the hook budget for any request body — including a malformed, empty or absent one — and shall not write terminal task status. -- (P2) When the `/ready` hook runs, the agent shall exec the agent CLI binary before returning 200, so that its pages are resident when the snapshot is captured. -- (P2) If a required `/ready` warm-up does not complete successfully, then the agent shall report not-ready (HTTP 503) rather than allow the snapshot to be taken. -- (P2) The `/ready` hook shall make no AWS API call, warm-up included. - -### 4. Infra and IAM: conditional resources behind bootstrap `ComputeTypes` - -Mirroring the ECS pattern: a `compute-lambda-microvm` bootstrap policy (`cdk/src/bootstrap/policies/`) gated on the `ComputeTypes` CFN parameter; a CDK construct provisioning the build role, execution role (admitted to the per-session role via `AgentSessionRole.admitComputeRole`, which was designed for exactly this), the S3 artifact bucket wiring, and image build automation. Egress uses the platform VPC via egress network connectors so the DNS Firewall / security-group / flow-log stack in [COMPUTE.md](../design/COMPUTE.md) applies unchanged; ingress is suppressed with the Lambda-managed `NO_INGRESS` connector (no `SHELL_INGRESS` — it is noted as a candidate for [#391](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/391), operator session access, as a separate decision). - -Two networking facts the construct has to encode, both established live: - -- **A `VPC_EGRESS` connector requires an operator role.** CloudFormation's generated L1 types `operatorRole` as optional and this ADR originally assumed Lambda would manage the ENIs with its own service-linked role. It does not: the connector fails to create with *"NetworkConnectorOperatorRole is required for VPC_EGRESS connector type"*. The construct creates one role — trusting the bare `lambda.amazonaws.com` service principal (see the trust-policy fact below), carrying `AWSLambdaVPCAccessExecutionRole` plus the ENI / tag / private-IP actions that policy omits — and shares it across both connectors, since it manages interfaces rather than traffic. -- **The MicroVM-facing roles cannot carry a confused-deputy source condition.** All three (build, execution, connector operator) trust the bare `lambda.amazonaws.com` service principal with **no** `aws:SourceAccount` / `aws:SourceArn`, and that is a forced choice, not an oversight: the Lambda MicroVMs service presents no source key when it assumes them, so a trust policy carrying one is unassumable. Two symptoms of the one cause, both live 2026-08-06/07 and both blocking: - - - Both `AWS::Lambda::NetworkConnector` resources `CREATE_FAILED` **deterministically** (on a freshly deleted stack, so not propagation lag — which matters, because *"The service is unable to assume the provided NetworkConnectorOperatorRole. Please verify the trust policy on the role."* is also the classic propagation symptom and a re-run is the obvious wrong guess). Removing the condition → both created within a second. - - `RunMicrovm` failed with a **misleading `iam:PassRole` AccessDenied on the caller**, with the orchestrator's grant present, `simulate-principal-policy` returning `allowed`, no permissions boundary, and a temporary *unconditioned* `iam:PassRole` **also** denied. The real cause was the execution role's trust; removing its conditions made the next submission reach `RUNNING` in 6 s. So the service reports a role it cannot pass-and-assume as an identity-policy denial on the principal passing it. - - Recorded plainly because the fix looks like a regression to anyone applying the standard service-principal pattern — and because it *was* a regression in the other direction: P1's standalone-validated operator-role probe had no conditions and worked, and the P1 F2 fix then added them "to mirror the build/execution roles". `sts:TagSession` stays: the service needs both actions and it was never implicated. - -- **Neither can the `iam:PassRole` grants carry an `iam:PassedToService` condition — same root cause, identity side (P2r2-F9 + P2r2-F10, live 2026-08-07 run 2).** An earlier revision of this ADR recorded the opposite, that the identity-side condition "was exonerated" by run 1's elimination. **That was a false negative**, and its cause is worth recording because it is a general trap: run 1 tested the conditioned grant by *adding* a temporary unconditioned `iam:PassRole` and watching the task still fail — but the temporary grant remained attached through the later submissions that succeeded, so the conditioned grant was never once tested against a working trust policy. A contaminated control. - - Run 2 ran the clean experiment — same exact-ARN resource, same ~5-minute IAM settle, one variable. It removed the run-1 workaround **first** (submission 4: denied) and only then added the unconditioned grant back on the same resource (submission 5: `RUNNING`), which is the ordering run 1 got wrong: - - | Orchestrator `iam:PassRole` on the execution role | Result | - |---|---| - | exact ARN **+ `iam:PassedToService: lambda.amazonaws.com`** | **DENIED** (two independent submissions) | - | exact ARN, **no condition** | **`RUNNING` in 9 s** | - - The denial lands on the **caller**, which is what makes it so misleading — the statement names that exact ARN and `simulate-principal-policy` answers `allowed`: - - ``` - User: arn:aws:sts:::assumed-role/backgroundagent-dev-TaskOrchestratorOrchestratorFn-… - is not authorized to perform: iam:PassRole on resource: - arn:aws:iam:::role/backgroundagent-dev-LambdaMicrovmComputeExecutionRo-… - because no identity-based policy allows the iam:PassRole action - ``` - - And the same key blocks the *other* PassRole path, which run 1 never reached because the enum defect (P2-F2) stopped it earlier: CloudFormation could not pass the **build role** at `CreateMicrovmImage` under the bootstrap `infrastructure` policy's allowlisted `IAMPassRole`. Verbatim, so the diagnosis does not have to be taken on trust: +| `/ready` | Required when runtime hooks are enabled; execute required binary warm-up before snapshot capture | +| `/validate` | Check local readiness, routes and configuration contracts without AWS calls | +| `/run` | Authenticate/install launch configuration and start the pipeline asynchronously | +| `/terminate` | Close the local coding barrier, log and acknowledge any request body; do not join the pipeline or write terminal task status | +| `/suspend` | Drain acknowledged progress and commit the current safe checkpoint within the hook budget | +| `/resume` | Renew credentials and reconcile the original gate before releasing coding | - ``` - LambdaMicrovmComputeImage… CREATE_FAILED - User: arn:aws:sts:::assumed-role/cdk-hnb659fds-cfn-exec-role--us-east-1/AWSCloudFormation - is not authorized to perform: iam:PassRole on resource: - arn:aws:iam:::role/backgroundagent-dev-LambdaMicrovmComputeBuildRoleF0-… - because no identity-based policy allows the iam:PassRole action - (Service: LambdaMicrovms, Status Code: 403) - ``` +Warm-up budgets come from `contracts/constants.json`; the total guest budget must remain below the image hook timeout. Cold-binary startup exceeded the hook budget in an earlier image; warm-up moved this work into image preparation. Historical sizes and timings are not sizing guarantees for later builds. - Three pieces of evidence pin that to the *condition* rather than to a stale bootstrap or a wrong resource pattern: +**Authenticated v2 payload transport.** Every ECS and MicroVM task uses S3. The coordinator publishes a deployment manifest, conditionally creates the task payload and sends a short-lived signed URL for that one object. The serialized MicroVM reference must fit 4,096 bytes; payloads are bounded at 8 MiB and manifests at 16 KiB. - 1. the live `IaCRole-ABCA-Infrastructure` policy was byte-identical to this branch's `cdk/bootstrap/policies/infrastructure.json`, so `cdk bootstrap --force` would have changed nothing; - 2. `aws iam simulate-principal-policy --policy-source-arn --action-names iam:PassRole --resource-arns ` returned `allowed` **with** `--context-entries ContextKeyName=iam:PassedToService,ContextKeyValues=lambda.amazonaws.com,ContextKeyType=string` and `implicitDeny` with no context entry — so the resource pattern matches and the condition key is the only remaining variable; - 3. the **control**: the out-of-band `create-microvm-image` call passed the *same build role* to the *same service* successfully, using operator credentials that carry no such condition. The role's trust is therefore fine and the denial is genuinely caller-side. +| Location | Shape | +|---|---| +| Hook reference / ECS `AGENT_PAYLOAD_REF` | `{version:2, task_id, bootstrap_s3_uri, payload_url, expires_at}` | +| `bootstrap/.json` | `{version:2, backend, platform_config}` | +| `/payload.json` | `{version:2, task_id, agent_payload, platform_config}` | +| Private `/launch.json` | `{fingerprint, reference}` for coordinator replay | - So: **the Lambda MicroVMs service presents no usable value for `iam:PassedToService` on either PassRole path** (CloudFormation → build role at `CreateMicrovmImage`; orchestrator → execution role at `RunMicrovm`), exactly as it presents no `aws:SourceAccount` on the assume-role path. One root cause, two more symptoms. Both statements therefore drop the condition, and the fix is deliberately asymmetric so it stays contained: +The worker authenticates only its deployment’s `bootstrap/*` using ambient credentials. Other payload-bucket object reads and listing are explicitly denied; signed task downloads carry coordinator authorization. The task ID and configuration must exactly match the reference and authenticated manifest before installation. MicroVM `platform_config` contains allowlisted non-secret identifiers, including role/secret ARNs; it does not contain credentials. ECS already receives its deployment configuration through task settings, so its manifest config is empty. - - `task-orchestrator.ts` sid `MicrovmPassExecutionRole` — condition removed; the **exact execution-role ARN** is now the whole of the scoping, which is why that resource must never be relaxed to a prefix or `*`. - - a new sid `MicrovmPassRoles` in the **conditional** `compute-lambda-microvm` bootstrap policy — unconditioned `iam:PassRole` on the build- and connector-operator role **name prefixes only** (not the execution role, which CloudFormation never passes). The shared `infrastructure` `IAMPassRole` keeps its allowlist, so no other role in the stack loses that constraint, and an agentcore-only bootstrap never gains an unconditioned pass at all. **Operators must re-bootstrap** (bundle ≥ 1.6.0) for the CDK-managed image path to work. +The coordinator saves the exact reference for idempotent replay. Signed URLs last at most 900 seconds, bounded by known credential expiry, and initial creation requires at least 300 seconds. An expired saved reference fails without re-signing the same launch request. Finalization deletes task payload/launch objects; one-day asynchronous S3 expiry is the backstop. - If AWS documents the value the service does present, adding it to both statements restores the condition. `microvms.lambda.amazonaws.com`, `lambda-microvms.amazonaws.com` and `microvms.amazonaws.com` were all `implicitDeny` against the conditioned policy, so any one of them would serve as the allowlist entry if it turns out to be right. Note that CloudTrail carries **no `lambda-microvms` management events at all** today, so the value cannot be read out of a log — only confirmed by AWS or found by a bounded sweep. +Normative requirements (EARS): - Compensating controls, enumerated per role. The deployment-role's shared `IAMPassRole` grant is **name-prefix-scoped**, so it technically covers all three roles; in practice only the orchestrator actively invokes `iam:PassRole` on the execution role (CloudFormation never requests it). The other two roles are passed **to** the deployment role (not to themselves). +- Image builds shall not embed credentials, task identity or deployment-specific configuration. Build hooks shall avoid AWS clients and credential caches. +- The worker shall reject legacy envelopes, oversized/malformed data, mismatched provenance and unknown configuration keys before starting a pipeline or installing configuration. +- Before configuration installation, the worker shall make only the bootstrap manifest/payload reads and shall log diagnostics to stdout. +- Signed URLs shall not appear in ordinary logs, agent-readable task rows or repository subprocess environments. Downloads shall use the exact regional S3 HTTPS object without redirects/proxies and with bounded response sizes. +- Producers, images and IAM shall be upgraded together; incompatible workers must be drained before switching transport. - | Role | Who can pass it, and how that grant is scoped | - |---|---| - | **Execution role** | The **orchestrator Lambda only**, at `RunMicrovm` — `iam:PassRole` scoped to this role's **exact ARN**, no condition (`constructs/task-orchestrator.ts`, sid `MicrovmPassExecutionRole`). The referenced comment contains the authoritative two-arm experiment evidence that the condition is the true blocker (not a permissions gap or stale bootstrap). | - | **Build role** | The **CloudFormation deployment role**, at `CreateMicrovmImage` (the L1's `buildRoleArn`) — via the new `MicrovmPassRoles` statement, scoped to `role/backgroundagent-dev-LambdaMicrovmComputeBuild*`, no condition. Also whoever runs `package-microvm-artifact.sh --create-image` out of band, using their own credentials. | - | **Connector operator role** | The **CloudFormation deployment role**, at `AWS::Lambda::NetworkConnector` create/update (`operatorRole`) — the same statement, scoped to `role/backgroundagent-dev-LambdaMicrovmComputeConnector*`. | +The [payload contract](../verification/645-payload-bootstrap.md) and [live checks](../verification/README.md) record the implementation and validation. This transport does not establish complete hostile-worker isolation: other platform grants remain, and a stolen signed URL is usable until expiry or revocation. - The rest of the posture: every resource these roles can reach is account-scoped by ARN **except two deliberate `Resource: '*'` statements** — `ec2:DescribeAvailabilityZones` on the execution role (EC2 describe actions have no resource-level scoping; read-only, no mutation, no data access, needed so a CDK target repo's `cdk synth` build gate can resolve AZ context on a fresh clone) and the connector operator role's ENI/tag/private-IP statement (`CreateNetworkInterface` is authorized before the ENI exists and the `Describe*` calls take no resource, which is why the AWS-managed VPC-access policy uses `*` too). Both are justified in the construct's cdk-nag `AwsSolutions-IAM5` suppressions, which is where a reviewer should check them rather than here. The Logs grants are prefix-scoped (`/aws/lambda-microvms/*` plus one named log group), i.e. wildcards inside a namespace, not `*`. Separately, the **orchestrator's** `lambda:PassNetworkConnector` is also `Resource: '*'` and unavoidably so — the AWS-managed connectors live in the `aws` account, outside any ARN we could enumerate (justified in `task-orchestrator.ts`, sid `MicrovmPassNetworkConnector`). Finally: none of the three roles holds `iam:*`, none has cross-account trust, and the only `sts:AssumeRole` any of them has is the execution role's, scoped to the per-task SessionRole. +No ABCA endpoint consumer exists in P1–P3. The platform grants no `CreateMicrovmAuthToken` permission and mints no JWE tokens. `NO_INGRESS` can still return an endpoint URL; an unauthenticated 403 verifies the authentication boundary, not valid-token reachability. - If AWS later populates a source key on this path, adding it to the shared principal fixes all three roles and both `sts` actions at once. +### 4. Infra and IAM: conditional resources behind bootstrap `ComputeTypes` -- **Build-time egress needs port 80; runtime does not.** `agent/Dockerfile` installs Debian packages and `apt-get` fetches over plain HTTP, so a 443-only egress path fails every snapshot build (`Could not connect to deb.debian.org:80 … exit code: 100`). Rather than widen the runtime posture, the construct provisions a **second, build-only** connector on the same private subnets with a 443 + 80 security group, referenced solely by the image resource and the packaging script. The agent at run time still has 443-only egress. +The backend adds build/runtime VPC connectors, build artifacts, launch payloads, logs, roles and a managed or external image. Its bootstrap policy is conditional on `ComputeTypes` including `lambda-microvm`. A VPC egress connector requires an operator role. Build egress permits ports 80/443 for package installation; runtime egress permits 443 through the platform VPC. -- Where the bootstrap `ComputeTypes` parameter includes `lambda-microvm`, the generated template shall attach the `IaCRole-ABCA-Compute-LambdaMicrovms` policy to the CloudFormation execution role. -- The orchestrator role shall receive only the MicroVM lifecycle actions it calls (`lambda:RunMicrovm`, `lambda:SuspendMicrovm`, `lambda:ResumeMicrovm`, `lambda:TerminateMicrovm`, `lambda:GetMicrovm` for `pollSession`, and `lambda:PassNetworkConnector`, which is required even for the default connectors), scoped to platform-created images. -- Where the `lambda-microvm` backend is enabled, the approve and deny Lambdas shall receive `lambda:ResumeMicrovm` and `lambda:GetMicrovm` — conditionally, mirroring the cancel handler's conditional `RUNTIME_ARN` wiring in `task-api.ts`. -- The trust policy of every MicroVM-facing role shall name `lambda.amazonaws.com` and shall carry no source-condition key (the service presents none; see the trust-policy fact above). -- The `iam:PassRole` grant the orchestrator uses for the MicroVM execution role shall carry no `iam:PassedToService` condition and shall be scoped to that role's exact ARN. -- Where the bootstrap `ComputeTypes` parameter includes `lambda-microvm`, the `IaCRole-ABCA-Compute-LambdaMicrovms` policy shall grant `iam:PassRole` without an `iam:PassedToService` condition, scoped to the MicroVM build- and connector-operator role name prefixes, and shall not extend that grant to the MicroVM execution role. -- The shared `IaCRole-ABCA-Infrastructure` `iam:PassRole` statement shall retain its `iam:PassedToService` allowlist. -- The MicroVM execution role shall hold `logs:CreateLogStream` and `logs:PutLogEvents` on the application log group whose name is delivered in `platform_config`, scoped to that log group. +**Trust and PassRole limitation.** Recorded live checks rejected `aws:SourceAccount`/`aws:SourceArn` conditions on the MicroVM-facing roles and `iam:PassedToService` on the MicroVM PassRole paths. The working roles trust `lambda.amazonaws.com` without those conditions; build/execution roles also allow `sts:TagSession`. IAM simulation with caller-supplied condition values did not prove that the service supplied those values. Reintroduce a condition only after verifying service support. -`lambda:CreateMicrovmAuthToken` is granted to no role in P1–P3 (no JWE consumer exists; see sub-decision 3). +| Role/action | Scope and responsibility | +|---|---| +| Coordinator lifecycle APIs | Configured image ARN and its version-qualified sibling; includes actual-version capability lookup | +| Coordinator `iam:PassRole` | Exact execution-role ARN, without `iam:PassedToService` | +| Deployment `iam:PassRole` | Backend-specific build/operator role name patterns; shared infrastructure allowlist remains intact | +| Approval/denial handlers | Observe/resume the configured image after committing a decision; dispatch parked continuations | +| Build role | Selected immutable artifact plus manual-build key; MicroVM log writes | +| Execution role | Bootstrap manifests, startup secrets, allowlisted models, Memory and logs; tenant data through the per-task SessionRole | +| Connector operator | Tested ENI/tag/private-IP permissions plus AWSLambdaVPCAccessExecutionRole | -**Cost attribution.** `cdk/src/main.ts` currently tags the whole stack with a single `compute_type` context value (default `agentcore`) — already imprecise with two backends, wrong with three. P1 must add backend-identifying cost-allocation tags on the MicroVM-specific resources (images, payload/artifact bucket wiring, log groups) and revisit the stack-level tag semantics (e.g. a `compute_types` list), keeping attribution consistent with [#645](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/645)'s cost/attribution acceptance criterion. +`lambda:PassNetworkConnector` has no resource-level authorization support and therefore uses `Resource: *`. The operator role also has wildcard ENI permissions; some mutations can be scoped in IAM, so their current tested wildcard is not evidence that narrower permissions are impossible. DescribeAvailabilityZones needs a wildcard for fresh CDK repository lookups. These exceptions and namespace wildcards are documented in the construct’s cdk-nag suppressions. -- Where a deployment enables the `lambda-microvm` backend, MicroVM-specific resources shall carry backend-identifying cost-allocation tags. +**Nested infrastructure.** `microvm_nested_stack` must be explicitly selected until flat-to-nested migration is verified. New and already-nested deployments should use `true`, putting MicroVM resources in a nested stack. Omission fails synthesis. The shared execution role stays in the parent to avoid a role-trust dependency cycle. Bootstrap 1.9.0 covers nested deployment roles. Before upgrading an existing flat installation, set and retain `microvm_nested_stack=false` until completing the reviewed overlap/drain migration, explicit image/coordinator pins and rollback checks described in the [migration prerequisites](../verification/645-p3-nested-stack.md). Reusable migration commands are still being completed. Moving construct paths alone is not a safe migration. Preserve `microvm_resource_name_prefix` after a migrated deployment. #### Security bar vs existing backends ([#645](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/645) acceptance criterion) -| Control | AgentCore | ECS | Lambda MicroVMs | Delta | -|---|---|---|---|---| -| Egress, runtime (DNS Firewall, TCP 443 SG, flow logs) | Platform VPC | Platform VPC | Platform VPC via egress network connector | None | -| Egress, image build | ECR build outside the platform VPC | ECR build outside the platform VPC | Platform VPC via a **separate build-only connector, TCP 443 + 80** (`apt-get` is plain HTTP) | New surface: build-time egress is wider than runtime egress by one port, on a connector no running MicroVM can use | -| Tenant-data scoping | Per-session role (`admitComputeRole`) | Per-session role | Per-session role, execution role admitted identically | None | -| Secrets delivery | Runtime env + Identity injection | Task env vars | Fetched at `/run`; never in snapshot | New surface: snapshot must stay secret-free (EARS req., sub-decision 3) | -| Non-secret platform config (table/bucket names, secret + role ARNs) | Runtime env vars | Task env vars | `platform_config` in the `/run` payload, installed into the process env | New surface: the values are attacker-relevant *as env vars* (`LD_PRELOAD`, `AWS_ENDPOINT_URL`), so the agent installs a fixed **allowlist** and rejects the whole run on any other key (EARS req., sub-decision 3) | -| Inbound exposure | None (SigV4 invoke only) | None (no endpoint) | **None — but only because the strategy passes `NO_INGRESS` explicitly.** The service default is a PUBLIC `HTTP_INGRESS` connector plus a public `*.lambda-microvm..on.aws` endpoint; no tokens are minted in P1–P3 either way | New surface **and** a new failure mode: "no inbound" is an active control, not an absence. Drop the `NO_INGRESS` argument and every agent MicroVM gets a public endpoint (EARS req., sub-decision 3) | -| IAM condition keys on the compute-role trust **and** on the `iam:PassRole` grants that hand it over | Trust pinned with `aws:SourceAccount`; `PassRole` under the allowlisted bootstrap statement | Trust pinned per-service; `PassRole` under the allowlisted bootstrap statement | **Neither is possible.** All three MicroVM-facing roles trust the bare `lambda.amazonaws.com` with no `aws:SourceAccount`/`aws:SourceArn`, **and** both `iam:PassRole` grants (orchestrator → execution role at `RunMicrovm`; CloudFormation → build role at `CreateMicrovmImage`) carry no `iam:PassedToService` — the service presents no usable value for any of those keys, and each condition is a hard blocker while present (live-verified, blocking, four times across two runs) | **Real, evidenced gap that does not close from our side, and it is wider than the trust policy alone.** `lambda.amazonaws.com` is shared with every other Lambda feature, so neither the account pin nor the passed-to-service pin is available on this path. Compensated per role (table in sub-decision 4): the **execution** role is passable by the **orchestrator only** (at `RunMicrovm`), restricted to its **exact ARN**; the **build** and **connector-operator** roles are passable by the CloudFormation deployment role under a new **conditional, per-backend, name-prefix-scoped** statement (`MicrovmPassRoles`, bootstrap ≥ 1.6.0) that deliberately excludes the execution role. The shared allowlisted `IAMPassRole` (`role/backgroundagent-dev-*`) is left intact to avoid widening the grant for ~30 other roles, so while it technically matches the execution role, only the orchestrator actively reaches for it. Resources are account-scoped by ARN apart from two justified `Resource: \'*\'` statements (`ec2:DescribeAvailabilityZones`; the operator role\'s ENI management — both carry cdk-nag IAM5 suppressions). No `iam:*`, no cross-account trust. Revisit if AWS ever documents the values the service presents; CloudTrail records no `lambda-microvms` events, so they cannot be read from logs | -| Per-task observability writes | Runtime writes to the vended APPLICATION_LOGS group | Task role writes to the task log group | Execution role writes to the SAME APPLICATION_LOGS group, granted against the group `platform_config` names (P2-F4) | None — but only after P2-F4: the name was delivered a phase before the grant, so the agent attempted the write and every per-task line (and `METRICS_REPORT`) was `AccessDenied`, degrading silently to guest stdout | -| Session isolation | MicroVM | Task-level | MicroVM (Firecracker) | None (≥ ECS) | -| State reuse | None | None | Snapshot shared across MicroVMs | New surface: CSPRNG reseed + credential refresh on `/run`/`/resume` — **P3 scope** (P2 exposure measured as negligible: sole `random` consumer is a ULID sort key under a `task_id` partition; `os.urandom`/`secrets` unaffected; no credential derives from `random`). Credential refresh IS in P2: per-task credentials arrive via `platform_config` at `/run` | -| Workload-token injection | Yes (Runtime-coupled) | No (env-var posture) | No (env-var posture) | Shared with ECS; deferred to [#249](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/249)/ADR-016 | -| Operator shell access | No | No | Not enabled (`SHELL_INGRESS` omitted; candidate for [#391](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/391)) | None by default | -| Auth-token minting | n/a | n/a | `CreateMicrovmAuthToken` granted to no role in any phase | Verified static-only: any principal holding the action *can* mint a working JWE, including against a `SUSPENDED` MicroVM, so the posture rests entirely on the grant being absent | +| Control | MicroVM posture | Difference to account for | +|---|---|---| +| Runtime egress | Platform VPC, DNS Firewall, HTTPS security group, flow logs | Separate build connector additionally allows HTTP | +| Tenant access | Task-scoped SessionRole | Shared compute-role permissions remain outside that boundary | +| Configuration/secrets | Authenticated runtime identifiers; credentials resolved after launch | Shared snapshot must not capture credentials or task identity | +| Ingress | Explicit NO_INGRESS; no platform token minting | Service defaults would select HTTP_INGRESS if the field were omitted | +| Trust/PassRole | Exact role/image scopes where supported | Source-condition limitations described above | +| Logs | Service image group plus platform APPLICATION_LOGS | Both namespaces need explicit grants | +| Retained work | Versioned, bounded, checksum-verified checkpoint | Protect stored conversation/files and clean terminal state | +| Workload identity | Runtime credentials and task-role refresh | Linear vault identifiers arrive through `platform_config`; the compute execution role mints tokens ([setup](../guides/LINEAR_SETUP_GUIDE.md#using-the-vault-with-lambda-microvms)) | + +MicroVM-specific resources carry `abca:compute-backend=lambda-microvm` cost tags. A stack-wide compute tag cannot accurately attribute a mixed-backend deployment by itself. #### Regional availability enforcement -Lambda MicroVMs launched in 5 regions (us-east-1/2, us-west-2, eu-west-1, ap-northeast-1) and will expand. ABCA is a single-region deployment, so the constraint is binary per stack: either the stack region supports the backend or the backend does not exist there. Enforcement is layered — one static check where offline determinism is required, live probes everywhere else so the platform self-heals as AWS adds regions (`list-managed-microvm-images` is the documented read-only availability probe): +The launch-region list is us-east-1, us-east-2, us-west-2, eu-west-1 and ap-northeast-1. New Regions can be supported before the static list is updated: -| Stage | Mechanism | Check | -|---|---|---| -| CDK synth/deploy | Static region constant (single exported list, documented update path) | Synth fails when `ComputeTypes` includes `lambda-microvm` in an unlisted region; context-flag escape hatch for newly launched regions ahead of the constant update | -| Repo onboarding | Live probe from the CLI | `bgagent repo onboard --compute-type lambda-microvm` calls `list-managed-microvm-images` in the stack region and rejects with a remedy (supported-region list + suggest `agentcore`/`ecs`) | -| `bgagent platform doctor` | Live probe (precedent: `checkBedrockModel`) | Reports backend availability for the stack region whenever any active blueprint selects `lambda-microvm` | -| Orchestration (defense in depth) | Error classification | `startSession` failures from a missing regional endpoint classify to a typed remedy in `error-classifier.ts`, never a cryptic SDK error on the task | +- Synth rejects a concrete unlisted Region unless `microvm_region_override` is set. An unresolved Region defers to live checks. +- CLI onboarding and platform doctor probe `list-managed-microvm-images`. +- Runtime regional failures receive a configuration remedy instead of an opaque SDK error. -- If an operator onboards a repo with `compute_type: 'lambda-microvm'` and the availability probe fails for the stack region, then the CLI shall reject the onboarding with the supported-region list and alternative backends as the remedy. -- If `startSession` fails because the MicroVM service is unavailable in the stack region, then the orchestrator shall classify the failure with a configuration remedy and shall not retry. -- When the platform doctor runs in a deployment where any active blueprint selects `lambda-microvm`, the doctor shall probe MicroVM availability in the stack region and report the result. +### 5. Rollout: phased, AgentCore remains the default backend -### 5. Rollout: phased, default unchanged +| Phase | Delivered behavior | +|---|---| +| P1 | Strategy, infrastructure, bootstrap/types, minimal `/ready` + `/run` serving; merged in [#689](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/689) | +| P2 | Clone → change → PR, progress/logs/Memory, runtime configuration, `/validate` + `/terminate`; merged in [#733](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/733), with takeover follow-up validation | +| P3 | Approval-aware sleep/wake, credential renewal, original deadlines, retained requests, complete checkpoint/replacement recovery, nested deployment and live acceptance | + +Activate sleep only after verifying the deployed image and coordinator together. Keep a compatible published coordinator and explicit image pin for rollback. Normal acceptance includes the 600-second default, explicit expiry, new and existing task off-switch behavior, replacement, cleanup and preservation of unrelated infrastructure. -- **P1 — strategy + infra + minimal hook serving:** `LambdaMicrovmComputeStrategy` (start/poll/stop), CDK construct, bootstrap policy, types sync, unit + CDK assertion tests, and the agent's `/ready` + `/run` endpoints. No suspend yet. The image IS creatable and launchable and the payload DOES reach the agent — but there is **no smoke-parity guarantee** (sub-decision 3's phasing table). -- **P2 — smoke parity:** the agent serves `/terminate` + `/validate` and installs its platform env from the `/run` payload (see sub-decision 3's "Platform configuration delivery"); agent completes clone → change → PR on the backend with progress visible to `bgagent watch`; failure classification entries in `error-classifier.ts`; **AgentCore Memory parity** (IAM grant + `MEMORY_ID` delivery, following the `EcsAgentCluster` pattern — Memory is a standalone service already consumed cross-substrate, and omitting the grant silently no-ops cross-session learning); the agent's remaining non-secret env parity inside the snapshot. -- **P3 — suspend/resume:** the interface widening from sub-decision 1 (mandatory methods, all three strategies in one commit), HITL-wait suspend policy, inline resume in the approve/deny Lambda with orchestrator-poll reconciliation (sub-decision 2), timeout-under-freeze wall-clock handling; coordinate with [#491](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/491)'s unified liveness model and update Cedar decision #7's rationale note. -- **Out of scope:** replacing AgentCore as default; classic Lambda functions as a runtime; GPU; the Runtime-coupled workload-access-token injection path (delivery mechanism exists only on AgentCore Runtime; MicroVMs adopt the ECS env-var posture until [#249](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/249)/ADR-016 redesign the seam). Gateway integration is orthogonal: ADR-019/[#641](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/641) is substrate-portable by design and applies to this backend when it lands. +Approval responses are supported through the authenticated CLI and owner-authored `approve`/`deny` replies to Linear approval comments. Changing the default backend, GPU support, native Slack approval buttons and operator shell access remain outside this ADR. Future work needs its own scope; “P4” is not an approved phase here. ## Consequences -- (+) **Suspend/resume economics.** Tasks idling on approval waits stop billing compute while preserving full state — bounded at ~1 h per gate under the current Cedar ceiling (decision #6), and the enabler for cheap off-hours gate-ceiling extensions later (§14.8). Verified end to end at the substrate level: suspend and resume each complete in ~1 s, and `microvmId` **and** `endpoint` survive the cycle byte-identical, so a stored `SessionHandle` remains valid. -- (+) **VM-level isolation without cluster ops.** Firecracker isolation with no ECS cluster, task definition, or capacity management; one-session-per-MicroVM maps 1:1 onto ABCA's task model. -- (+) **Escapes AgentCore's 2 GB image limit and FUSE `flock()` workaround** — native disk in the snapshot supports `uv`/`mise` without the split-storage scheme. Note the size comparison must say WHICH measure it means: the same agent tree is 1.799 GB as an OCI image (629.7 MB compressed, i.e. under AgentCore's limit) but reports `codeInstallSizeInBytes` of 2.17 GiB as a MicroVM snapshot (i.e. over it). The two straddle the limit and are not interchangeable; memory/disk snapshot sizes are a third thing again and must not be summed into the comparison. -- (+) **Liveness becomes explicit.** Unlike AgentCore's stub `pollSession`, the strategy can report real substrate state, strengthening the [#491](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/491) unification. -- (−) **32 GiB sustained ceiling (8 GiB baseline + automatic 4× vertical scaling), 32 GB disk.** Not a successor to the ECS backend for heavy CI-parity builds; the platform now maintains three backends. -- (−) **Capacity is baseline-priced with burst headroom, which is a narrower promise than "32 GiB".** The deployment configures an 8 GiB / 4 vCPU baseline and the service scales to 32 GiB / 16 vCPU on demand — well matched to an agent task, which is idle-ish while waiting on the model and spiky during builds, and cheaper than reserving the peak. But it is *burst*, not a reservation: a workload that needs 32 GiB **sustained** is relying on scaling behaviour this ADR has not measured, and against ECS's 120 GB the gap for sustained-memory workloads is unchanged. So the value proposition remains the suspend economics, the observable control-plane state machine, and the absence of cluster ops — with capacity now a fair-to-good fit rather than a hard blocker. Repos with genuinely sustained heavy builds still belong on `ecs`. -- (−) **New packaging pipeline.** Zip + Dockerfile + service-side image builds with versioned snapshots (storage billed per version) alongside the existing ECR flow; image versions need lifecycle cleanup — including versions left behind by FAILED builds, and noting the last version of an image cannot be deleted individually (delete the image, which reaps it). -- (−) **The payload bucket is on the hot path, not the overflow path.** With a 4 KB `runHookPayload` cap, virtually every real task delivers its payload via S3, so the bucket, its TTL rule and the execution role's read grant are load-bearing for normal operation rather than an edge case (sub-decision 3). -- (−) **8-hour hard cap includes suspended time**, and with `idlePolicy` omitted there is no tighter substrate-level suspended-TTL — the suspended-state bound is `maximumDurationInSeconds` plus orchestrator termination and the stranded reconciler. A manually suspended VM was observed alive at 1 h with no TTL in sight (observation truncated there), so nothing contradicts this bound, but nothing narrows it either. Under today's 1 h gate ceiling it is comfortably sufficient; any future extension of gate ceilings must revisit the bound (an additive `idlePolicy` change) and give the orchestrator a checkpoint-and-restart path (push branch, new session) beyond the cap. -- (!) **Idle-policy foot-gun.** Traffic-based auto-suspend would freeze a busy outbound-only agent; the decision to disable auto-suspend must be enforced in code and covered by tests, not left to configuration discipline. -- (!) **Service defaults are not the desired posture.** Two live-caught cases (public `HTTP_INGRESS` by default; `/ready` mandatory) mean an omitted field on this backend does not mean "off" — it can mean "the service picks, and it picks wider than we want". Every new `RunMicrovm` / `CreateMicrovmImage` field should be assumed to have an opinionated default until checked. -- (!) **Nothing self-terminates on the paths that matter** — superseding P1's unqualified version of this bullet. With `run: ENABLED` the service DOES reap a VM whose run hook returns 4xx (~12 s, `stateReason: "Run lifecycle hook returned HTTP status 400."`, live-verified), so a guest that rejects its own payload cleans itself up. That is the only self-cleaning case: the service reaps a hook *result*, and once `/run` has answered 200 it has no view of the guest. A VM whose task finished, crashed after `/run`, or hung stays `RUNNING` and billing until the 8 h cap, so the orchestrator's `TerminateMicrovm` on finalize remains the only cleanup for normal operation and a leaked handle is still a cost incident. -- (!) **Snapshot uniqueness.** Shared memory snapshots mean every MicroVM restored from one image starts with identical PRNG state. **Re-scoped to P3 in the P2 review, with the exposure measured rather than assumed** (see the amendment under sub-decision 2): the sole `random` consumer in `agent/src` is a ULID sort key under a `task_id` partition, `os.urandom`/`secrets` are unaffected, and no credential or token derives from `random` — so the P2 exposure is a possible duplicate progress event, not a security boundary. It becomes load-bearing at P3, where a *resumed* VM continues from frozen state repeatedly; the reseed must land on `/run` and `/resume` together with those hooks. Asserting the requirement while leaving it unimplemented was the real defect, and this amendment is the fix. -- (−) **The agent stack template is at 98.6 % of CloudFormation's 1 MB limit** (985,886 bytes) and 486 of 500 resources with a MicroVM image configured — ~14 KB of headroom, i.e. roughly one more construct, and down from 98.4 % / ~16 KB one run earlier. Not caused by this backend (the MicroVM construct is ~6 KB of it) but reached by it, and it will block deploys for reasons that have nothing to do with MicroVMs. Tracked in [#735](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/735); the candidate remedies are `suppressTemplateIndentation` and a stack split. -- (!) **Snapshot WARMTH is a first-class property, not an optimisation.** A snapshot inherits only the pages something touched before it was captured, so a large lazily-loaded artifact — the 225 MiB `claude` binary, and anything similar added later — pays its first-touch cost on the *first task* instead of at container start. That cost failed every task at turn 0 in the P2 smoke run (P2-F5). Anything heavyweight added to the image must be exec'd in `/ready`, and any timeout guarding a first touch must be sized for a cold page fault rather than for the work itself. -- (!) **Regional availability (5 regions at launch, expanding)** — enforced in layers (synth-time static check, onboarding + doctor live probes, orchestration-time classification; see sub-decision 4). The static CDK constant is the one piece that rots as AWS expands; its update path and context-flag escape hatch are deliberate. -- (!) **Workload-token injection delta persists** (shared with the ECS backend) until [#249](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/249)/ADR-016 land; document it in the security bar comparison rather than blocking on it. Memory and Gateway are explicitly *not* deltas — both are standalone services consumed via IAM from any substrate. +- Approval waits can stop consuming compute without discarding the question or the saved work. Snapshot and checkpoint costs still apply. +- Native disk supports build-tool locking, and the service exposes explicit worker state without a cluster to operate. +- ABCA now maintains three backends and an additional artifact/snapshot lifecycle. Failed or unused image versions also need cleanup. +- Eight hours remains a per-worker limit. Retained approvals rely on verified retirement and replacement rather than extending a worker indefinitely. +- Suspended workers retain ABCA capacity until confirmed retirement. The service’s account-quota treatment of suspended memory was not established by the recorded probes. +- Memory baseline validation and published peak capacity are not workload benchmarks. Sustained heavy builds require measured sizing and may fit ECS better. +- A healthy heartbeat does not prove coding progress. Hook failures may cause service termination, but successful hook acceptance does not remove ABCA’s cleanup responsibility. +- Service error wording can hide transport errors. The pooled-hook mitigation has local and live evidence; historical internal dispatch traces remain unavailable. ## Testing -**P1 (start/poll/stop — no suspend):** - -- Unit tests for the strategy: start/poll/stop mapping (including `SessionStatus` `'suspended'` reported mechanically, without task-state interpretation), payload-size branching (inline vs S3 pointer, at the exact 4 096/4 097-byte boundary), the image-identifier-must-be-an-ARN guard, the explicit `NO_INGRESS` argument (including the blank-env-var fallback, which must never omit the field), error classification (`ServiceQuotaExceededException`, `ThrottlingException`, `ResourceNotFoundException`, regional-unavailability), the omit-`idlePolicy` invariant, and `maximumDurationInSeconds` fixed at 28 800. -- Agent tests: `/ready` returns 200 once the server is up and starts nothing; `/run` accepts both envelope shapes (inline and S3 pointer), starts the pipeline asynchronously through the same mapper `/invocations` uses, returns before the pipeline finishes, and rejects every unusable envelope with a named code before spawning; `/suspend` and `/resume` are NOT served (`/validate` and `/terminate` joined the served set in P2). -- Orchestrator tests: substrate-terminal + non-terminal task status → failed classification; `suspended` + non-`AWAITING_APPROVAL` status → anomaly event, no fail-fast; `compute_metadata` persisted with `microvmId`/`endpoint` after `startSession`. -- CDK assertions: MicroVM resources present only when `ComputeTypes` includes the backend; synth failure for unsupported regions (plus the context-flag escape hatch); memory size validated against the accepted list at synth; the connector operator role and its trust; two connectors with the build-only one carrying port 80 and the runtime one not; the `/ready` + `/run` hook declaration and the absence of the others; IAM actions scoped as specified (orchestrator lifecycle set; no `CreateMicrovmAuthToken` anywhere); payload-bucket grants (execution role read-only); backend cost-allocation tags; types-sync check covers the widened `ComputeType`. -- CLI tests: onboarding rejection with remedy when the availability probe fails; doctor check present when a blueprint selects the backend. -- P1 verification items (external service facts) — **executed 2026-07-31, us-east-1**; see `docs/verification/645-p1-lambda-microvm-runbook.md` for the full evidence. Discharged: `runHookPayload` limit (**4 096**, not 16 KB), the accepted baseline memory sizes (`[512…8192]` MiB — note the developer guide, not the probe, is what establishes that this is a BASELINE with a 32 GiB peak), image-identifier ARN requirement, IAM action names and the observed image-ARN shape, region probe behaviour, manual suspend/resume without `idlePolicy`, terminate timing and the `TERMINATED`-persists-≥10-min finding, the default public `HTTP_INGRESS`, and the `/ready` requirement. **Not** discharged: account-quota treatment of `SUSPENDED` MicroVMs (not observable safely), suspended TTL beyond 1 h (truncated), the vertical-scaling behaviour itself (no workload here approached the baseline, so the 4× peak is documented rather than observed), and the `AWS::Lambda::MicrovmImage` CloudFormation value shapes (never exercised — the run used the out-of-band script path; **discharged, and REFUTED, by the P2 run — see P2-F2 in sub-decision 3**). Record the closed answers in COMPUTE.md. - -**P2 (smoke parity):** - -- Agent tests: `/validate` returns 200 with its individual check results, 503 while initialising, reports a missing hook route / unsupported interpreter, starts nothing, and makes **zero AWS calls even with `LOG_GROUP_NAME` set** (asserted by poisoning the boto3 and CloudWatch-writer seams — the same assertion covers `/ready`); `/terminate` returns 200 with no body at all, with a malformed / non-object / wrong-content-type / whitespace-only body, when the body read itself fails, with a pipeline still running (without joining it), and when its own best-effort step raises — and never calls `task_state.write_terminal`; a structural assertion that the route carries no typed body param keeps the 422 from being reintroduced. -- `platform_config` tests: the allowlist and required subset are read from `contracts/constants.json` (the wire key set is additionally asserted literally, as the agent-side tripwire on a contract edit); an unknown key rejects the whole block with nothing installed; a non-object block and a non-string value are rejected; a blank/`null` optional value is skipped without clobbering an image value while a blank required value is rejected; a payload value beats a pre-existing env value; installation is observed to happen before the GitHub-token resolver runs and before any pipeline thread exists; the config is picked up from the inline envelope, from beside the S3 pointer, and from inside the fetched object (inner wins); an envelope with no `platform_config` is still accepted with a warning. -- Snapshot credential hygiene: a subprocess probe asserts that importing `server` and serving `/ready` + `/validate` imports neither `boto3` nor `botocore`, caches no `aws_session` session, and spawns no CloudWatch writer thread — the property that keeps a build-role credential chain and the build-time region out of the snapshot. -- `/run` pre-install silence: with a **baked `LOG_GROUP_NAME`** (the hostile case — without it the assertions pass vacuously) every AWS/credential seam (`boto3.client`/`Session`, the `aws_session` factories, `_debug_cw`/`_warn_cw`) is armed to raise until the install succeeds. Asserted on the accepted path, on all three rejection paths (bad envelope, `platform_config` invalid, `platform_config` incomplete) and on the failed-fetch 500 — where the seams stay armed for the whole request, because a rejected run installed nothing and so earns no AWS call. The permitted exception is asserted POSITIVELY: exactly one client is built pre-install, for `s3`, through the attributed factory. -- `/ready` warm-up tests (P2-F5): the hook exec's each configured binary exactly once with a generous timeout; `claude` is the only REQUIRED entry; a timeout, a missing binary, a non-zero exit and an unexpected `OSError` each produce **503 with the reason logged to stdout** rather than a 200 or a 500; a best-effort failure still reports ready; the warm-up makes zero AWS calls with `LOG_GROUP_NAME` baked. Plus the backstop half: the `claude --version` probe's bound is asserted to be ≥ 60 s and to be applied to the *exec* rather than to the PATH lookup, and a missing CLI warns instead of raising. -- CDK assertions (P2-F1/F2/F4): no source-condition key on any of the three MicroVM-facing role trusts, and no `aws:SourceAccount`/`aws:SourceArn` string anywhere in them; hook properties are `ENABLED` and the architecture is `ARM_64`, with a negative assertion that **no** hook route string appears anywhere in the rendered image resource; the agent hook routes are asserted against their own dedicated constant (the template no longer carries a path to compare); the execution role holds `logs:CreateLogStream`/`PutLogEvents` on the application log group and the two logs grants stay separate; the stack wires the SAME log group it delivers as `platform_config.log_group_name`. -- Smoke (gated like the ECS backend): clone → change → PR with `bgagent watch` progress; Memory write parity (no AccessDenied no-op). **Run 1 (2026-08-06) FAILED at `implement`, turn 0 — no PR. Run 2 (2026-08-07) PASSED: two tasks clone → change → commit → push → PR, `COMPLETED`, 12 turns / $0.279 / 153 s** (`docs/verification/645-p2-smoke-runbook.md`), which also discharged P2-F1, P2-F2, P2-F4, P2-F5 and the dual-signal-liveness item (45 s heartbeat cadence observed across a 181 s `RUNNING` window). **The row is not yet fully closed:** run 2 needed one live IAM workaround, and establishing why produced P2r2-F10 (the identity-side `iam:PassedToService`) and P2r2-F9 (its CloudFormation twin). Both are fixed in source above and neither has been re-exercised live, so what remains is a re-run on a re-bootstrapped account with no workarounds. - -**P3 (suspend/resume):** +The [verification summary](../verification/README.md) distinguishes recorded live acceptance from open PR checks. Required coverage includes: -- Unit tests: suspend/resume mapping; agentcore/ecs `unsupported` stubs; approve/deny inline resume with handle loaded from `compute_metadata`. -- HITL lifecycle tests: inline resume failure leaves the decision outcome intact and records the orphan event; orchestrator backstop retries resume; gate expiry fires at `min(monotonic budget, created_at + timeout_s)` — including the suspend/resume case where the monotonic budget exceeds the wall-clock remainder — without disturbing the §13.12 late-approval race protection. -- Smoke: suspend/resume across a simulated approval wait preserving workspace state. +- Strategy state mapping, explicit unsupported results, ARN/Region validation, NO_INGRESS, omitted idlePolicy and bounded uncertain-start recovery. +- Hook readiness/warm-up, AWS-silent build hooks, authenticated payload installation, arbitrary terminate bodies and lifecycle connection closure. +- Decision/timeout/cancellation races, image capability, coding/progress barriers, original deadlines, durable replay and exact-attempt capacity ownership. +- Real S3 version/integrity/access checks, real DynamoDB transactions and actual SDK conversation/Git/workspace recovery after process and disk loss. +- Cloud approve/deny/expiry/cancellation, repeated sleep/wake, AWS credential renewal after actual expiry, retirement/replacement and final resource cleanup. +- Nested fresh deployment and overlapping migration, compatible image/coordinator rollback, normal CLI feedback and live off-switch acceptance. -**All phases:** docs sync for [COMPUTE.md](../design/COMPUTE.md) (new column distinguishing MicroVMs from classic Lambda) and [ORCHESTRATOR.md](../design/ORCHESTRATOR.md) (liveness + suspend lifecycle). +Live evidence must distinguish service acknowledgment from completed guest recovery and synthetic handler checks from actual external-channel submissions. ## References -- Issue [#645](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/645) — originating RFC proposal -- Issue [#491](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/491) — unified liveness decision model (soft dependency, P3) -- Issue [#641](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/641) / ADR-019 (PR [#663](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/663)) — substrate-portable tool plane -- PR [#596](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/596) — ECS Fargate backend (pattern source for conditional wiring) -- [AWS Lambda MicroVMs](https://docs.aws.amazon.com/lambda/latest/dg/lambda-microvms-guide.html) — developer guide; [Running and using MicroVMs](https://docs.aws.amazon.com/lambda/latest/dg/microvms-launching.html) — lifecycle APIs and hooks -- [Agent Toolkit for AWS — aws-lambda-microvms skill](https://github.com/aws/agent-toolkit-for-aws/blob/main/skills/specialized-skills/serverless-skills/aws-lambda-microvms/SKILL.md) — operational constraints (no self-suspend, idle-policy semantics, snapshot uniqueness, size limits) -- [ADR-020](./ADR-020-ears-requirements-syntax.md) — EARS syntax used for the normative requirements above -- [CEDAR_HITL_GATES.md](../design/CEDAR_HITL_GATES.md) — approval-gate mechanics (decisions #6, #7) the suspend/resume handshake preserves; `cancel-task.ts` / `task-api.ts` — the inline best-effort + reconciler-backstop pattern the resume path mirrors -- [COMPUTE.md](../design/COMPUTE.md), [ORCHESTRATOR.md](../design/ORCHESTRATOR.md) — design docs to be updated by the implementing PRs +- [Issue #645](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/645) — originating proposal +- [Issue #491](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/491) — unified liveness model +- [Issue #641](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/641) — substrate-portable tool plane +- [AWS Lambda MicroVMs guide](https://docs.aws.amazon.com/lambda/latest/dg/lambda-microvms-guide.html) and [lifecycle APIs/hooks](https://docs.aws.amazon.com/lambda/latest/dg/microvms-launching.html) +- [ADR-020](./ADR-020-ears-requirements-syntax.md) — requirement syntax +- [Compute](../design/COMPUTE.md), [orchestrator](../design/ORCHESTRATOR.md), [approval gates](../design/CEDAR_HITL_GATES.md) diff --git a/docs/decisions/ADR-023-trusted-approval-writer.md b/docs/decisions/ADR-023-trusted-approval-writer.md new file mode 100644 index 000000000..73daa07e4 --- /dev/null +++ b/docs/decisions/ADR-023-trusted-approval-writer.md @@ -0,0 +1,68 @@ +# ADR-023: Trusted approval writer and retained human decisions + +**Status:** proposed +**Date:** 2026-09-21 + +## Context + +Workers previously had direct write access to approval records. Restricting a +worker to its own task did not prevent it from replacing a pending request with +an approved record. MicroVM continuation makes the distinction between worker +authority and human consent especially important, but the same defect affects +ECS and AgentCore. + +A short approval deadline also couples human response time to compute lifetime. +Someone taking time to consider a request should not lose it merely because the +worker should stop consuming resources. + +## Decision + +Use a small IAM-authenticated API and Lambda as the worker-facing approval +writer. Its interface creates a pending request or closes one after a worker +timeout/failure. It cannot record human approval or denial, notification markers, +or retention TTLs. Human decisions remain in the owner-authenticated decision +handlers. Workers retain approval reads and transaction condition checks. + +The session role can invoke only the task path named by its `task_id` tag. +Creation atomically writes the request and transitions its task from running to +awaiting approval. The transaction checks the task owner/state and, for MicroVM, +the coordinator-owned worker lease. A stale worker cannot use its old lease to +create or close a request. + +Unanswered requests have no decision deadline by default. Explicit positive +timeouts remain available. Task cancellation, terminal failure and invalidated +worker execution can close requests independently. Compute lifetime is separate: +MicroVM can checkpoint and release a worker while retaining a request; ECS and +AgentCore currently cannot restore that waiting execution into a replacement. +Their task execution limits still apply. + +This implementation is included in the P3 review because retaining approvals +without protecting their decision records would preserve the security defect. +The ADR remains proposed for maintainer review. + +## Consequences + +- Existing deployments must pause submissions, drain old workers and deploy + matching infrastructure and images. Old images still attempt direct writes, + which the new IAM policy denies. See the + [upgrade procedure](../guides/DEPLOYMENT_GUIDE.md#upgrading-approval-permissions). +- The service adds one signed request per creation/closure and becomes an + availability dependency. Failure denies permission to proceed; it never + becomes human consent. +- The service validates record shape, not the truth of worker-supplied policy + descriptions. A preview can be truncated and is not a proof of the full tool + input. `TIMED_OUT` currently also represents a worker polling failure; it is + not proof that a human deadline elapsed. Human `DENIED` remains distinct. +- The compute role still chooses session tags when assuming the session role. + This API protects the decision-writing boundary but does not provide full + isolation from a compromised worker retaining ambient compute credentials. +- Linear consent is read back from Linear and must come from the mapped owner, + excluding bots and the saved OAuth token's own identity. Other webhook paths + still need a separate review of the legacy shared OAuth/signing-secret bundle. + +## References + +- [MicroVM backend, issue #645](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/645) +- [Implementation and security review, PR #904](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/904) +- [Approval trust boundaries](../design/CEDAR_HITL_GATES.md#121-trust-boundaries) +- [ADR-021: Lambda MicroVM compute backend](./ADR-021-lambda-microvms-compute-backend.md) diff --git a/docs/design/API_CONTRACT.md b/docs/design/API_CONTRACT.md index a62e30bd3..4d9f64e1b 100644 --- a/docs/design/API_CONTRACT.md +++ b/docs/design/API_CONTRACT.md @@ -254,7 +254,7 @@ Returns full details of a task. Users can only access their own tasks. } ``` -`agent_heartbeat_at` is the agent's last in-guest liveness beat, or `null`. The agent writes it every 45 s on the `agentcore` and `lambda-microvm` backends; on `ecs` it is written once at start, because that backend runs the pipeline directly instead of serving HTTP. The orchestrator reads the same field to detect a hung agent inside a healthy compute environment ([ORCHESTRATOR.md](./ORCHESTRATOR.md#dynamodb-heartbeat-agentcore-and-lambda-microvms)), so a value that is minutes old on a `RUNNING` task is the signal, not the timestamp itself. `null` on records written before the field existed. +`agent_heartbeat_at` is the agent's last in-guest liveness beat, or `null`. The agent writes it every 45 s on the `agentcore` and `lambda-microvm` backends; on `ecs` it is written once at start, because that backend runs the pipeline directly instead of serving HTTP. The orchestrator reads the same field to detect a hung agent inside a healthy compute environment ([ORCHESTRATOR.md](./ORCHESTRATOR.md#liveness-monitoring)), so a value that is minutes old on a `RUNNING` task is the signal, not the timestamp itself. `null` on records written before the field existed. `error_classification` is a derived field computed at response time from `error_message`. When `error_message` is `null`, `error_classification` is `null`. When present, it contains: @@ -382,7 +382,7 @@ When a task pauses in `AWAITING_APPROVAL` (Cedar soft-deny gate), the owner appr { "data": { "task_id": "01HYX...", "request_id": "...", "status": "DENIED", "decided_at": "2025-03-15T10:35:00Z" } } ``` -**Errors:** `400 VALIDATION_ERROR`, `401 UNAUTHORIZED`, `404 REQUEST_NOT_FOUND` (collapses "row missing" and "wrong caller"), `409 REQUEST_ALREADY_DECIDED`, `409 TASK_NOT_AWAITING_APPROVAL`. +**Errors:** `400 VALIDATION_ERROR`, `401 UNAUTHORIZED`, `404 REQUEST_NOT_FOUND` (collapses missing, inaccessible, closed or expired approval rows), `409 TASK_NOT_AWAITING_APPROVAL` (task-only state conflict). ### List pending approvals @@ -608,7 +608,7 @@ There is no per-user request-rate or "tasks-per-hour" limiter on task creation. | `RATE_LIMIT_EXCEEDED` | 429 | Rate/concurrency gate exceeded — per-task nudge limit, the application rate limiter on approval endpoints, or the user concurrency limit on confirm-uploads | | `BUDGET_EXCEEDED` | 429 | A configured user or Cognito-team monthly budget reached 100% with hard stop enabled | | `REQUEST_NOT_FOUND` | 404 | Cedar HITL approval request not found (also returned when the caller does not own it) | -| `REQUEST_ALREADY_DECIDED` | 409 | Cedar HITL approval request was already approved or denied | +| `REQUEST_ALREADY_DECIDED` | 409 | Legacy error-code enum; current approve/deny handlers return `404 REQUEST_NOT_FOUND` for closed or inaccessible approval rows | | `TASK_NOT_AWAITING_APPROVAL` | 409 | Task is not in `AWAITING_APPROVAL`, so the approval decision does not apply | | `INTERNAL_ERROR` | 500 | Unexpected server error | | `SERVICE_UNAVAILABLE` | 503 | Downstream dependency unavailable (retry with backoff) | diff --git a/docs/design/ATTACHMENTS.md b/docs/design/ATTACHMENTS.md index 7f5932100..da5963e9c 100644 --- a/docs/design/ATTACHMENTS.md +++ b/docs/design/ATTACHMENTS.md @@ -1290,13 +1290,13 @@ The task record is preserved with status `CANCELLED` — `bgagent status **Status:** Core implemented; this document remains the authoritative design reference. +> **Status:** Core implemented. The September update below supersedes the original bounded-wait assumptions in the historical design and examples. > **Companion:** [`INTERACTIVE_AGENTS.md`](./INTERACTIVE_AGENTS.md) §9.3 (pointing here), §7 (state machine). > **Design locked:** 2026-04-23 (Sam ↔ assistant discussion). > **Rev:** 5 (2026-05-06 — fold in parallel adversarial + advocate review of the timeout design: late-approval re-read on TIMED_OUT ConditionCheckFailed; user-visible timeout-cap milestones; ceiling-shrink milestone; Runtime JWT bound verified as auto-refreshed IAM; three new tuning metrics; explicit off-hours trade-off section; notification-delivery-failure boundary. IMPL-24 through IMPL-28 added.). > **Implementation:** Core shipped. The 3-outcome engine (`agent/src/policy.py`), default policy sets (`agent/policies/hard_deny.cedar`, `agent/policies/soft_deny.cedar`), approval Lambdas (`cdk/src/handlers/{approve-task,deny-task,get-pending,get-policies}.ts`) wired into `cdk/src/constructs/task-api.ts` (routes `/tasks/{id}/approve`, `/deny`, `/pending`, `/repos/{repo_id}/policies`), the cross-engine parity fixtures (`contracts/cedar-parity/`), and the exact engine pins are all on `main`. §15's task list is preserved as a historical implementation record; see the note at the top of §15 for what (if anything) remains unbuilt. +> +> **Current source behavior (2026-09-22):** the task default is `approval_timeout_s=0`, +> meaning no decision deadline. Workers create/close requests through the IAM-authenticated +> [trusted approval writer](../decisions/ADR-023-trusted-approval-writer.md); they cannot +> write approval rows directly. Explicit task settings are 30–3,600 seconds; +> positive policy-rule deadlines still apply. Pending rows have no DynamoDB TTL, +> and `expires_at` is nullable. Task closure cancels unanswered requests and adds +> retention TTL without changing already-recorded decisions. Closed approval-row +> conditions return `404 REQUEST_NOT_FOUND`; task-only conflicts return 409. +> MicroVM checkpoint/retirement/replacement separates human waiting from worker +> lifetime and capacity; other backends retain their existing runtime limits. +> Linear accepts an owner’s `approve` or `deny` reply to the approval comment; +> the handler verifies the actual comment through Linear’s API. Slack uses CLI +> response instructions. See the current +> [user guide](../guides/USER_GUIDE.md#approval-gates-cedar-hitl) and +> [continuation protocol](./ORCHESTRATOR.md#retained-microvm-approvals). +> The normal deployment passed retained-request, ten-minute sleep, explicit-expiry +> and sleep-off/rollback acceptance; see the +> [deployment record](../verification/README.md). --- @@ -12,7 +31,7 @@ 1. [What we are building, in one paragraph](#1-what-we-are-building-in-one-paragraph) 2. [The three-outcome model and why Cedar alone can't give it](#2-the-three-outcome-model) -3. [Design decisions (locked)](#3-design-decisions-locked) +3. [Design decisions](#3-design-decisions) 4. [End-to-end request flow](#4-end-to-end-request-flow) 5. [Cedar policy authoring guide](#5-cedar-policy-authoring-guide) 6. [Engine implementation](#6-engine-implementation) @@ -109,19 +128,19 @@ The winning property: **policy authors can put on their "security-review-approve --- -## 3. Design decisions (locked) +## 3. Design decisions -Settled during the 2026-04-23 design discussion and extended after the 2026-04-24 and 2026-05-06 reviews. Each has detailed rationale in those conversations; summary here for implementers. **23 decisions**, all locked unless an adversarial review finding explicitly reopened a concern. +Settled during the 2026-04-23 design discussion and extended after the 2026-04-24 and 2026-05-06 reviews. Each has detailed rationale in those conversations; summary here for implementers. The September retained-request update also amends deadline, recovery and capacity behavior. | # | Decision | Summary | |---|---|---| | 1 | **Cedar encoding: two policy sets** | Physical hard-deny vs soft-deny split, validated via `@tier(...)` annotation. | | 2 | **Hook point: extend `PreToolUse`, not `can_use_tool`** | PreToolUse is already async-compatible, already wired to Cedar, and already owns the tool-governance boundary. | | 3 | **Wait mechanism: DDB strongly-consistent polling, 2s → 5s backoff** | Initial 2s cadence for the first 30s, then 5s. `ConsistentRead=True` so the agent never misses an approval that already landed. | -| 4 | **Scope allowlist: in-process, seeded from persisted `initial_approvals`** | Runtime escalation lives in the `PolicyEngine` instance. Submit-time `--pre-approve` flags persist on TaskTable and seed the allowlist at container startup. Lost on restart (rare; reconciler fails stranded tasks). | +| 4 | **Scope allowlist: in-process, seeded from persisted `initial_approvals`** | Runtime grants live in the `PolicyEngine` instance. Submit-time grants seed it at startup; a verified MicroVM checkpoint also saves session grant scopes and denial-cache entries for replacement. | | 5 | **CLI UX: standalone `bgagent approve/deny` + `--pre-approve ` + `bgagent policies list` + `bgagent pending`** | No inline interactive prompt in the streaming CLI for v1. Discovery + listing commands solve the request_id/rule_id copy problem. | -| 6 | **Timeouts: per-task default + per-rule Cedar annotation override, min wins, bounded floor + ceiling, fail-closed** | Per-task default: **300s** (5 min), overridable via `--approval-timeout` on submit and bounded by `[30, min(3600, maxLifetime - 300)]`. Floor: 30s (engine-enforced on both task default and rule annotations). Ceiling: `min(1h, maxLifetime_remaining - cleanup_margin)` — sized so the TTL on the approval row always covers the decision window. On timeout → deny (never auto-approve). See §14.8 for the off-hours trade-off this posture deliberately accepts. | -| 7 | **Concurrency slots: AWAITING_APPROVAL holds the slot** | Matches PAUSED semantics. Container is alive, consuming memory. | +| 6 | **Timeouts: per-task default + per-rule Cedar annotation override, min wins, bounded floor + ceiling, fail-closed** | Per-task default: **0**, meaning no deadline. Explicit task deadlines are 30–3,600 seconds; positive matching rule deadlines can shorten them. Rule annotations still require at least 30 seconds. Workers without checkpoint/replacement support retain their existing lifetime limit. Pending rows have no storage TTL. Explicit timeout means deny, never auto-approve. | +| 7 | **Concurrency slots: AWAITING_APPROVAL holds the slot** | Bounds unfinished sessions and their eventual resume demand, including a suspended MicroVM. This is an ABCA admission policy; suspended AWS memory-quota consumption remains unverified. | | 8 | **Hard-deny is absolute** | No `--pre-approve` scope, and no blueprint `disable:` directive, can bypass it. CreateTaskFn validates and rejects `rule:`; blueprint loader rejects `disable:` entries that name built-in hard-deny rules. | | 9 | **Submit-time scope cap: 20 entries, ≤128 chars each** | Keeps audit trail legible, bounds allowlist check cost, limits abuse-vector damage. | | 10 | **Cedar annotations (verified working)** | `@rule_id(...)`, `@tier(...)`, `@approval_timeout_s(...)`, `@severity(...)`, `@category(...)`. Recoverable via `cedarpy.policies_to_json_str()` → JSON. Multi-match merging: min timeout wins (clamped by floor), max severity wins. | @@ -137,13 +156,13 @@ Settled during the 2026-04-23 design discussion and extended after the 2026-04-2 | 20 | **`write_path:` scope** | Added so users can pre-approve file writes under specific path patterns (e.g., `write_path:docs/**`) without needing to grant all Writes. Validation uses Python `fnmatch` at runtime; glob semantics are a Cedar-`like` superset (§6.4, §5.5). | | 21 | **`tool_group:file_write` convenience scope** | Resolves to `{Write, Edit}`. Prevents the surprise of pre-approving `Write` and still getting gated on `Edit`. | | 22 | **Pre-implementation spike: cedarpy annotation round-trip** | Day 1 of implementation validates that `policies_to_json_str()` returns annotations in the expected shape. If the API has changed, fall back to policy-ID prefix conventions. | -| 23 | **Cedar engine parity contract (Python `cedarpy` ↔ JS `cedar-wasm`)** | Both engines are pinned in `mise.toml`. A golden-file parity test runs in CI: for each `(policy, input)` fixture the test asserts Python and WASM return the same `decision` and the same set of matching rule IDs. Policy authors who upgrade either engine must refresh the golden file; drift fails the build. See §15.6 and Appendix B. | +| 23 | **Cedar engine parity contract (Python `cedarpy` ↔ JS `cedar-wasm`)** | Both engines are pinned in their package manifests. A golden-file parity test runs in CI: for each `(policy, input)` fixture the test asserts Python and WASM return the same `decision` and the same set of matching rule IDs. Policy authors who upgrade either engine must refresh the golden file; drift fails the build. See §15.6 and Appendix B. | --- ## 4. End-to-end request flow -Narrative walk-through of the happy path. Sequence diagrams in the round-trip Mermaid below. +Narrative walk-through with an explicit 600-second task deadline and a custom `force_push_any` policy annotated with `@approval_timeout_s("300")`. Built-in starter rules do not set deadlines. Sequence diagrams are below. ### Setup (task start) @@ -159,7 +178,7 @@ Narrative walk-through of the happy path. Sequence diagrams in the round-trip Me - rejects blueprint whose combined `cedar_policies` text exceeds the 64 KB cap (§12.4) regardless of origin - resolves `approval_gate_cap` = `Blueprint.security.approvalGateCap ?? 50`; rejects if outside `[1, 500]` (decision #13) 5. Task persists. `approval_timeout_s`, `approval_gate_cap`, and `initial_approvals` become DDB attributes on the task row (cap is captured at submit time so mid-task blueprint edits do not shift the cap beneath a running task). -6. Container spawns on Runtime-JWT. `PolicyEngine.__init__` loads: +6. Container starts with IAM runtime credentials. `PolicyEngine.__init__` loads: - `HARD_DENY_POLICIES` (built-in + repo blueprint's `security.cedarPolicies.hard`; blueprint `disable:` may suppress non-built-in rules only, §5.1, §15.4) - `SOFT_DENY_POLICIES` (built-in + repo blueprint's `security.cedarPolicies.soft`; blueprint `disable:` may suppress soft-deny rules freely) - Annotation lookup table: `{policy_id: {annotation: value}}` built from `cedarpy.policies_to_json_str()` once, cached for the task lifetime @@ -193,7 +212,7 @@ Narrative walk-through of the happy path. Sequence diagrams in the round-trip Me ) → effective = 300s ``` - If `maxLifetime_remaining_s - CLEANUP_MARGIN_120S < FLOOR_30S`, hook returns DENY immediately with reason `"insufficient lifetime for approval"` (§13.7). + This lifetime ceiling applies to workers without continuation support. A continuation-capable MicroVM omits it because the approval can outlive the worker. With no positive task/rule deadline, the effective timeout is zero (no decision deadline). 12. Hook checks per-task approval-gate cap (default 50, configurable per blueprint via `security.approvalGateCap`; §5.1) and per-minute rate limit (20/task, per-container). If either exceeded → DENY with reason `"approval-gate cap exceeded"` (fail-closed). 13. Hook mints `request_id = _ulid()` (26-char ULID). @@ -211,15 +230,17 @@ Narrative walk-through of the happy path. Sequence diagrams in the round-trip Me "status": "PENDING", "created_at": "2026-04-23T14:00:00Z", "timeout_s": 300, - "ttl": 1734567890, # created_at + timeout_s + CLEANUP_MARGIN_120S; always covers the decision window + "expires_at": "2026-04-23T14:05:00Z", # explicit decision deadline; no retention TTL "user_id": "...", "repo": "my-org/my-app" } ``` -15. **Atomic transition** — hook issues `TransactWriteItems` with two operations: +15. **Atomic transition** — the hook sends an IAM-signed request to the approval + service, which issues `TransactWriteItems` with these operations (and a + worker-lease condition for MicroVM): - Put on `TaskApprovalsTable` (new row with status=PENDING) - ConditionalUpdate on `TaskTable`: `status = :awaiting, awaiting_approval_request_id = :rid WHERE status = :running` - Both succeed or both fail. On `TransactionCanceledException` (most likely the TaskTable condition fails because another process moved the status), the hook emits `approval_write_failed` and returns DENY. + Both succeed or both fail. If the service rejects a task/lease conflict or the request fails, the hook emits `approval_write_failed` and returns DENY. 16. Hook emits `agent_milestone("approval_requested", {...})` to both `ProgressWriter` (DDB audit) and `sse_adapter` (live stream). Best-effort emission — transactional write has already committed; milestone failure is observability degradation, not state degradation. 17. Terminal A stream renders: ``` @@ -230,36 +251,12 @@ Narrative walk-through of the happy path. Sequence diagrams in the round-trip Me timeout 300s ``` Severity colors the line (respecting `NO_COLOR` env var). -18. Hook enters poll loop with strongly-consistent reads: - ```python - async def _poll_for_decision(task_id, request_id, timeout_s): - start = time.monotonic() - interval = 2 - consecutive_failures = 0 - while True: - elapsed = time.monotonic() - start - if elapsed >= timeout_s: - return TimedOut() - if elapsed > 30: - interval = 5 # backoff - try: - row = await _ddb_get_approval(task_id, request_id, ConsistentRead=True) - consecutive_failures = 0 - if row is None: - # Row disappeared between write and poll — treat as stranded - return TimedOut(reason="approval row missing; fail-closed") - if row["status"] != "PENDING": - return Decided(row) - except Exception as exc: - consecutive_failures += 1 - if consecutive_failures == 3: - log("WARN", f"approval poll degraded for {request_id}: {exc}") - emit_milestone("approval_poll_degraded", {...}) - if consecutive_failures >= 10: - return TimedOut(reason="approval poll consecutive failures") - await asyncio.sleep(interval) - ``` -19. The approval CAP and local-timeout paths ALWAYS attempt to write the row to TIMED_OUT (best-effort conditional update `status = :pending`) before returning. This prevents orphan PENDING rows when the agent bails internally. +18. Hook enters the poll loop with strongly-consistent reads. For an explicit positive timeout, the deadline is captured with the original approval row, before database writes/notifications: UTC expiry is `created_at + timeout_s`, capped by the original monotonic remaining duration. `deadline.remaining_s()` takes the smaller remainder and clamps at zero, so a frozen guest clock or backward UTC correction cannot restart the window. + An untimed request has no expiry. The executable poll loop in + [hooks.py](../../agent/src/hooks.py) also handles cancelled/closed requests, + read failures and continuation barriers; it must not be replaced by a loop + that checks only APPROVED/DENIED or starts a fresh timer after wake. +19. The local-timeout path asks the trusted approval service to conditionally mark the pending row TIMED_OUT before returning. If that write loses or fails, the hook rereads consistently to honor an already-committed decision. The approval-cap check runs before row creation and has no row to update. ### User responds @@ -272,7 +269,7 @@ Narrative walk-through of the happy path. Sequence diagrams in the round-trip Me - ConditionalUpdate on `TaskApprovalsTable`: `#status = :pending AND user_id = :caller AND task_id = :task_id` → flip to APPROVED - ConditionalUpdate on `TaskTable`: `#status = :awaiting AND awaiting_approval_request_id = :rid` → (no-op update, pure state guard; keeps status AWAITING_APPROVAL until the agent's resume transaction flips it RUNNING) Both conditions must hold or the entire transaction is cancelled. No TOCTOU window, no "approved a cancelled task" 202 surprise. - - On `TransactionCanceledException` with per-item `CancellationReasons`: distinguishes between (a) approvals row missing (404 `REQUEST_NOT_FOUND`), (b) approvals row wrong user (404 `REQUEST_NOT_FOUND` — don't leak existence), (c) approvals row wrong status (409 `REQUEST_ALREADY_DECIDED`), (d) task no longer AWAITING_APPROVAL (409 `TASK_NOT_AWAITING_APPROVAL`). + - On `TransactionCanceledException` with per-item `CancellationReasons`: returns 404 `REQUEST_NOT_FOUND` for any approval-row condition failure (missing, foreign-owned or already closed), or 409 `TASK_NOT_AWAITING_APPROVAL` for a task-only condition failure. - Records audit event to TaskEventsTable directly (`approval_decision_recorded`) so the 90-day audit trail is owned by the Lambda, not dependent on agent milestones. - Returns 202 `{task_id, request_id, status: "APPROVED", scope, decided_at}` or error. 24. Agent's poll reads the `APPROVED` row on next tick (within 2-5s). @@ -297,6 +294,7 @@ sequenceDiagram participant Engine as PolicyEngine participant Events as TaskEventsTable participant Approvals as TaskApprovalsTable + participant Requests as Approval request service participant CLI participant User participant Lambda as ApproveTaskFn @@ -305,8 +303,9 @@ sequenceDiagram Agent->>Hook: tool call (Bash git push --force) Hook->>Engine: evaluate_tool_use Engine-->>Hook: REQUIRE_APPROVAL (soft-deny force_push_any) - Hook->>Approvals: TransactWriteItems - Note right of Hook: Put approval row PENDING
plus TaskTable status
to AWAITING_APPROVAL + Hook->>Requests: IAM-signed create for this task + Requests->>Approvals: TransactWriteItems + Note right of Requests: Put approval row PENDING
plus TaskTable status
to AWAITING_APPROVAL Hook->>Events: approval_requested milestone Events-->>CLI: live stream with approval_requested CLI-->>User: bgagent approve TASK REQ @@ -318,7 +317,8 @@ sequenceDiagram Lambda-->>CLI: 202 APPROVED Hook->>Approvals: poll with ConsistentRead Approvals-->>Hook: status APPROVED - Hook->>Approvals: TransactWriteItems, TaskTable to RUNNING + Hook->>Approvals: ConditionCheck on recorded decision + Note right of Hook: Same transaction updates TaskTable to RUNNING;
worker does not modify the approval row Hook->>Engine: allowlist.add(scope) if scope is not this_call Hook-->>Agent: permissionDecision allow Note over Stop: (not used on approval path) @@ -420,7 +420,7 @@ Fail-on-error is the right posture for blueprint misconfiguration — silent-fal |---|---|---|---| | `@rule_id("...")` | **Yes on soft-deny**, recommended on hard-deny | Kebab-case or snake_case identifier, unique across both tiers | Stable ID for `--pre-approve rule:X`, for audit trail, and for the `bgagent policies` discovery endpoint. `PolicyEngine.__init__` raises on duplicates. | | `@tier("hard"\|"soft")` | **Yes** | Exactly one of "hard" or "soft" | Validates policy is in the correct file/section. Engine rejects mismatch at load time. | -| `@approval_timeout_s("N")` | No | Integer seconds ≥ 30 | Per-rule timeout. If absent, uses the task default (**300s** by default, overridable via submit-time `--approval-timeout`; see decision #6). Has no effect on hard-deny rules. Values below the floor are rejected at load time. Values below **120s** emit a blueprint-load WARN but are accepted down to the 30s floor — almost no human responds to an approval request in under 2 minutes, so sub-120s is usually a policy-authoring mistake (see IMPL-25). Loader policy: STRICT at the floor (30s, reject) and ADVISORY below 120s (warn, accept). | +| `@approval_timeout_s("N")` | No | Integer seconds ≥ 30 | Per-rule timeout. If absent, uses the task setting (**0/no deadline** by default, configurable via submit-time `--approval-timeout`; see decision #6). Has no effect on hard-deny rules. Values below the floor are rejected at load time. Values below **120s** emit a blueprint-load WARN but are accepted down to the 30s floor — almost no human responds to an approval request in under 2 minutes, so sub-120s is usually a policy-authoring mistake (see IMPL-25). Loader policy: STRICT at the floor (30s, reject) and ADVISORY below 120s (warn, accept). | | `@severity("low"\|"medium"\|"high")` | No | One of the three | Shown in CLI approval prompt, colored by severity. Default: "medium". | | `@category("...")` | No | "destructive", "network", "filesystem", "auth", or free-form | UX grouping. CLI could filter approvals by category. Not enforced. | @@ -449,7 +449,7 @@ forbid (principal, action == Agent::Action::"execute_bash", resource) when { context.command like "*DROP TABLE*" }; ``` -**Gate destructive git ops** (soft-deny — part of the built-in starter set): +**Gate destructive git ops** (custom timed variants of the built-in starter rules): ```cedar @tier("soft") @rule_id("force_push_any") @@ -484,7 +484,7 @@ forbid (principal, action == Agent::Action::"execute_bash", resource) A force-push to any branch needs approval in 300s. A force-push to `main` or `prod` gives the user 600s with elevated severity. A non-force push to a protected branch (`main`/`prod`/`master`/`release/*`) also gates — catches the case where an agent directly pushes rather than opening a PR. If a command matches both `force_push_any` and `force_push_main`, multi-match merging picks `min(300, 600) = 300s` and `max(medium, high) = high`. -**Protect sensitive file paths** (soft-deny — part of the built-in starter set): +**Protect sensitive file paths** (custom timed variants of the built-in starter rules): ```cedar @tier("soft") @rule_id("write_env_files") @@ -671,35 +671,12 @@ The recent-decision cache is a simple `dict[(tool_name, input_sha), (decision, r When multiple soft-deny rules match a single tool call: -```python -def _merge_annotations(self, policy_ids: list[str]) -> dict: - rule_ids, timeouts, severities = [], [], [] - for pid in policy_ids: - ann = self._annotations[pid] - rule_ids.append(ann.get("rule_id", pid)) - if "approval_timeout_s" in ann: - try: - t = int(ann["approval_timeout_s"]) - if t >= FLOOR_30S: - timeouts.append(t) - except ValueError: - log("WARN", f"malformed @approval_timeout_s on {ann.get('rule_id', pid)}") - severities.append(ann.get("severity", "medium")) - - # Task default always eligible - timeouts.append(self._task_default_timeout_s) - - raw_min_timeout = min(timeouts) - return { - "rule_ids": rule_ids, - "timeout_s": max(FLOOR_30S, raw_min_timeout), # floor enforcement - "severity": _max_severity(severities), # "high" > "medium" > "low" - } -``` +The implementation is `_merge_annotations` in [policy.py](../../agent/src/policy.py). +It takes the shortest **positive** matching-rule or task timeout; zero means no +deadline and is excluded from that minimum. If no positive timeout exists, the +result is zero. Positive results retain the 30-second floor. The highest matching +severity governs the displayed severity, and matching rule IDs are preserved. -**Rationale for min/max choices**: -- **Timeout → min (above floor)**: multiple rules matching means multiple concerns. Users should have *less* time to decide when stakes are higher. Floor prevents unusable 5s windows. -- **Severity → max**: the most severe concern governs the UX coloring. ### 6.4 Allowlist data structure @@ -766,161 +743,30 @@ class ApprovalAllowlist: PreToolUse hook (compressed for doc; implementation will be richer): -```python -async def pre_tool_use_hook(hook_input, tool_use_id, ctx, *, - engine, task_id, user_id, progress, sse_adapter, - task_default_timeout_s): - tool_name, tool_input = _extract(hook_input) - decision = engine.evaluate_tool_use(tool_name, tool_input) - - if decision.outcome == Outcome.ALLOW: - return _allow() - if decision.outcome == Outcome.DENY: - return _deny(decision.reason) - - # REQUIRE_APPROVAL path. - # Cap + rate-limit check. Per-minute rate limit is per-container; on - # container restart the counter resets. The per-task approvalGateCap - # (blueprint-configurable, default 50) is persisted and bounds cumulative - # damage across restarts (§13.6). - if engine.approval_gate_count >= engine.approval_gate_cap: - return _deny(f"approval-gate cap exceeded ({engine.approval_gate_cap}/task)") - if engine.approvals_in_last_minute >= APPROVAL_RATE_LIMIT: - return _deny("approval-gate rate limit exceeded (20/min)") - - # Compute effective timeout with floor/ceiling. - remaining = _remaining_maxlifetime_s() - effective_timeout = max( - FLOOR_30S, - min(decision.timeout_s or task_default_timeout_s, - task_default_timeout_s, - remaining - CLEANUP_MARGIN_120S), - ) - if remaining - CLEANUP_MARGIN_120S < FLOOR_30S: - return _deny(f"insufficient maxLifetime remaining ({remaining}s) for approval") - - request_id = _ulid() - engine.approval_gate_count += 1 - - row = { - "task_id": task_id, "request_id": request_id, - "tool_name": tool_name, - "tool_input_preview": _strip_ansi(_preview(tool_input))[:256], - "tool_input_sha256": _sha256(_serialize(tool_input)), - "reason": decision.reason, "severity": decision.severity, - "matching_rule_ids": list(decision.matching_rule_ids), - "status": "PENDING", - "created_at": _iso_now(), - "timeout_s": effective_timeout, - "ttl": int(time.time()) + effective_timeout + CLEANUP_MARGIN_120S, - "user_id": user_id, "repo": engine.repo, - } - - # ATOMIC: put approval row + transition TaskTable status in one transaction. - try: - await _transact_write_approval_request(task_id, request_id, row) - except TransactionCanceledException as exc: - # Either the task was concurrently cancelled, or status wasn't RUNNING. - _emit("approval_write_failed", {"request_id": request_id, "reason": str(exc)}) - return _deny("approval system unavailable") - - _emit("approval_requested", { - "request_id": request_id, "tool_name": tool_name, - "input_preview": row["tool_input_preview"], - "reason": decision.reason, "severity": decision.severity, - "timeout_s": effective_timeout, - "matching_rule_ids": list(decision.matching_rule_ids), - }) - - outcome = await _poll_for_decision(task_id, request_id, effective_timeout) - - # On TIMED_OUT, attempt to write the row to TIMED_OUT so future reads see - # a terminal state (not orphaned PENDING). The conditional write is guarded - # by `status = :pending` — if the user's APPROVE landed between our last - # poll and this write, the condition fails. In that case we MUST re-read - # the row and honor whatever terminal state won the race; otherwise local - # `outcome.status = "TIMED_OUT"` is stale and we would deny a call the user - # just approved ("I approved it" → agent denies). See §13.12 and the - # scenario below. - if outcome.status == "TIMED_OUT": - wrote_timeout = await _best_effort_update_status( - task_id, request_id, "TIMED_OUT", - reason=outcome.reason, - # Returns True on successful write, False on ConditionCheckFailed. - ) - if not wrote_timeout: - # Re-read the row with ConsistentRead — user's decision beat us. - row = await _ddb_get_approval(task_id, request_id, ConsistentRead=True) - if row is not None and row["status"] == "APPROVED": - # Late-approve wins. Honor it. Rebuild the outcome so the - # downstream allow flow (scope propagation, milestone emission, - # resume transaction) runs identically to the normal approve - # path. - outcome = Decided( - status="APPROVED", - scope=row.get("scope"), - decided_by=row.get("user_id"), - decided_at=row.get("decided_at"), - ) - _emit("approval_late_win", { - "request_id": request_id, - "outcome": "APPROVED", - "reason": "user decision landed during TIMED_OUT write", - }) - elif row is not None and row["status"] == "DENIED": - outcome = Decided( - status="DENIED", - reason=row.get("deny_reason") or "denied", - decided_at=row.get("decided_at"), - ) - # If status is still PENDING (rare — concurrent reaper race) or - # the row is gone (TTL reaped before we could read it), fall - # through with the original TIMED_OUT outcome; fail-closed deny. - - # ATOMIC: resume TaskTable status RUNNING, conditional on awaiting_approval_request_id matching. - try: - await _transact_resume(task_id, request_id) - except TransactionCanceledException: - # User cancelled (or some other path) during poll; abandon gracefully. - _emit("approval_resume_failed", {"request_id": request_id}) - return _deny("task no longer awaiting approval") - - if outcome.status == "APPROVED": - if outcome.scope and outcome.scope != "this_call": - engine._allowlist.add(outcome.scope) - _emit("approval_granted", {"request_id": request_id, - "scope": outcome.scope or "this_call", - "decided_at": outcome.decided_at}) - return _allow() - - # DENIED or TIMED_OUT — cache for 60s + queue denial injection. - engine._recent_decisions.record( - tool_name, _sha256(_serialize(tool_input)), - decision="DENIED" if outcome.status == "DENIED" else "TIMED_OUT", - reason=outcome.reason, - ) - # Truncated reason for guaranteed-surface permissionDecisionReason. - # Best-effort richer injection via _denial_between_turns_hook; may be - # pre-empted by _cancel_between_turns_hook on a concurrently-cancelled - # task. See §4 "Denial with steering text" scenario. - permission_decision_reason = _truncate( - outcome.reason or f"User {outcome.status.lower()}", max_len=500 - ) - if outcome.status == "DENIED": - # Queue steering injection via Stop hook's between_turns_hooks. - engine._queue_denial_injection( - request_id=request_id, - reason=outcome.reason, # already sanitized by DenyTaskFn - decided_at=outcome.decided_at, - ) - _emit("approval_denied" if outcome.status == "DENIED" else "approval_timed_out", - {"request_id": request_id, "reason": outcome.reason}) - return _deny(permission_decision_reason) -``` - -`engine._queue_denial_injection` appends to a list consumed by `_denial_between_turns_hook` — registered **after** `_nudge_between_turns_hook` in the `between_turns_hooks` list (which itself runs after `_cancel_between_turns_hook`). At the next Stop hook fire, the denial is emitted as `…` XML (sanitized via `_xml_escape` from the shared utility introduced with Phase 2). If a `bgagent cancel` has landed between the deny and the next Stop seam, `_cancel_between_turns_hook` short-circuits the dispatcher and the denial text is NOT injected — in which case the guaranteed surface is `permissionDecisionReason` on the hook return. See finding #2 scenario in §4 for the cancel-vs-deny race reasoning. - -**Scenario (§13.12 VM-throttle + late-approval race).** User Alice hits a soft-deny gate at t=0 with `timeout_s=300`. The AgentCore VM is evicted from its warm CPU share around t=285 due to noisy-neighbor pressure on the host; poll ticks stretch by ~400ms. Alice, seeing the approval prompt in Terminal A, types `bgagent approve 01KPW... 01KPR...` at t=294. The approve-transaction lands in DDB at t=294.7 (APPROVED). The agent's next poll-tick was due at t=290 but the VM throttle delayed it to t=295.1. The monotonic wall-clock already shows elapsed >300 (actual since-start ~300.3s), so `_poll_for_decision` returns `TimedOut()`. The hook runs `_best_effort_update_status("TIMED_OUT", ... WHERE status = :pending)` — the conditional fails because the row is APPROVED. **Without the re-read**, the hook would proceed with stale local `outcome.status = "TIMED_OUT"`, queue a denial injection, and return `{"permissionDecision": "deny"}` — Alice sees "I approved it" on Terminal B but the agent denies the tool call anyway. **With the re-read** (the `wrote_timeout` branch in the pseudocode above): the hook fetches the row with ConsistentRead, sees `status = APPROVED`, rebuilds `outcome` from the row (preserving `scope`, `decided_by`, `decided_at`), emits an `approval_late_win` milestone, runs the normal resume transaction + allow flow, and returns `{"permissionDecision": "allow"}`. Alice's tool runs. The cost is one extra strongly-consistent GetItem on the race path; the benefit is that user intent is authoritative. Without this fix, a timer design that is otherwise sound would produce a confounding and unrecoverable UX. See IMPL-24, §13.12, and §15.2 task #43 for the race test. +The executable flow lives in [hooks.py](../../agent/src/hooks.py), with signed +request creation/closure in [approval_requests.py](../../agent/src/approval_requests.py) +and task-state transactions in [task_state.py](../../agent/src/task_state.py). + +1. Evaluate policy, existing grants, gate cap and creation-rate limit. +2. Resolve the shortest positive task/rule deadline. Zero remains untimed. + For a positive deadline, apply a remaining-worker-lifetime ceiling when one + is supplied. A continuation runtime separates this deadline from worker life. +3. Ask the trusted service to atomically create the pending request and move the + task to `AWAITING_APPROVAL`; MicroVM requests also check the active worker lease. + Pending rows have no storage TTL. A failed write denies the action. +4. Emit the notification milestone and poll with strongly consistent reads, + preserving the original UTC/monotonic deadline. No deadline means no timer expiry. +5. If polling times out or fails, ask the trusted service to close the pending + request. If a human decision won the race, reread and preserve that decision. +6. Resume only the matching active task/gate and worker lease. Approval allows the + action and applicable scope; denial is returned to the tool hook and queued as + best-effort steering for the next Stop hook. Cancellation prevents resumption. + +The worker's `TIMED_OUT` outcome currently also covers polling failures; it is +not proof that an explicit human deadline elapsed. This limitation is recorded in +[ADR-023](../decisions/ADR-023-trusted-approval-writer.md). + +**Scenario (§13.12 VM-throttle + late-approval race).** Alice hits a gate with a 300-second window and commits APPROVED at t=294.7s. Scheduling delays prevent the next agent poll until t=300.2s. That poll sees the original deadline has passed and returns TIMED_OUT before reading the row. The conditional TIMED_OUT write then loses because APPROVED is already stored. The hook rereads with `ConsistentRead`, preserves Alice's scope and decision metadata, emits `approval_late_win`, and proceeds through the guarded resume transaction and allow flow. Without that reread it would deny an already-approved call. See IMPL-24, §13.12, and §15.2 task #43 for the race test. --- @@ -948,8 +794,7 @@ Content-Type: application/json | 202 | — | Success | `{task_id, request_id, status: "APPROVED", scope, decided_at}` | | 400 | `VALIDATION_ERROR` | Bad scope format, missing fields | `{error, message, field}` | | 401 | `UNAUTHORIZED` | Missing/invalid JWT | — | -| 404 | `REQUEST_NOT_FOUND` | Row missing OR wrong user (both surfaces 404 to prevent enumeration) | — | -| 409 | `REQUEST_ALREADY_DECIDED` | Approvals row status != PENDING | `{error, message, current_status}` | +| 404 | `REQUEST_NOT_FOUND` | Approval-row condition failed: missing, foreign-owned or already closed | — | | 409 | `TASK_NOT_AWAITING_APPROVAL` | Task's current status is not AWAITING_APPROVAL | `{error, message, current_status}` | | 429 | `RATE_LIMIT_EXCEEDED` | Per-user > 30 approve/min | — | | 503 | `SERVICE_UNAVAILABLE` | DDB throttled or upstream failure | — | @@ -998,21 +843,32 @@ await ddb.transactWriteItems({ }); ``` -On `TransactionCanceledException`, `ApproveTaskFn` inspects the per-item `CancellationReasons` to distinguish cases: -- ApprovalsTable condition failed with `OldImage` absent → 404 `REQUEST_NOT_FOUND` -- ApprovalsTable condition failed with `OldImage.user_id != caller` → 404 (same code, prevent existence oracle) -- ApprovalsTable condition failed with `OldImage.status != "PENDING"` → 409 `REQUEST_ALREADY_DECIDED` -- TaskTable condition failed (status changed) → 409 `TASK_NOT_AWAITING_APPROVAL` +On `TransactionCanceledException`, `ApproveTaskFn` inspects per-item +`CancellationReasons`. It does not request or classify old item images: + +- Any approval-row condition failure → 404 `REQUEST_NOT_FOUND`, including + cancellation, timeout and an already-recorded decision. +- A task-only condition failure → 409 `TASK_NOT_AWAITING_APPROVAL`. This is symmetric with the agent-side `TransactWriteItems` pattern (§4 step 25a) used for the resume transition — Lambdas and agent speak the same atomic-update contract. -**Ownership**: `user_id` stored on TaskApprovalsTable and compared against `caller_user_id` in the ConditionExpression is the Cognito `sub` claim **verbatim**. The Lambda extracts `sub` from the validated JWT and uses it as-is: no prefix stripping, no tenant mapping, no format normalization. If we ever introduce per-tenant user ID namespacing, that transformation MUST happen at the **write** path (i.e. before the agent writes the row in §4 step 14) rather than at compare time, so the ConditionExpression always compares identical-shape identifiers. See finding #6 scenario below. +**Ownership**: `user_id` stored on TaskApprovalsTable and compared against `caller_user_id` in the ConditionExpression is the Cognito `sub` claim **verbatim**. The Lambda extracts `sub` from the validated JWT and uses it as-is: no prefix stripping, no tenant mapping, no format normalization. If we ever introduce per-tenant user ID namespacing, that transformation MUST happen at the **write** path (i.e. before the approval service persists the row prepared in §4 step 14) rather than at compare time, so the ConditionExpression always compares identical-shape identifiers. See finding #6 scenario below. -After successful transaction, `ApproveTaskFn` writes an audit event to `TaskEventsTable` (`approval_decision_recorded` event_type), ensuring the 90-day audit trail is owned by the Lambda path — not dependent on the agent's milestone emission. +After a successful transaction, the decision handler attempts the authoritative `approval_decision_recorded` audit event independently of the agent milestone. Audit-delivery failure does not undo the saved decision. **Scenario (finding #6):** Three months from now, a platform engineer adds a multi-tenant mode where Cognito `sub` becomes `tenant-abc:01JXZ...`. They update the agent's row-write path to prefix-strip: `user_id = sub.split(":", 1)[1]`, storing `01JXZ...` on TaskApprovalsTable. They forget to update `ApproveTaskFn`. Now the Lambda reads `sub = "tenant-abc:01JXZ..."` from the JWT and compares it against the stored `01JXZ...` — condition fails, 404 on every approve, all tasks stranded. The fix as written: "the Cognito sub is compared verbatim; any transformation must happen at write time, not at compare time" — if the agent writes the full `sub`, the Lambda compares the full `sub`; if either side transforms, both sides must. The CI assertion is a unit test that extracts `user_id` from a sample row and asserts it matches the `sub` claim of a sample JWT byte-for-byte. This test would fail on the prefix-strip refactor above and force the engineer to update both sides. Without this hard rule, ownership-in-condition silently breaks under any future identity refactor. -**Scenario (finding #7):** A user submits a risky task at 10:00 AM. At 10:05 AM the agent hits a soft-deny gate. At 10:05:30 AM the user on Terminal B runs `bgagent cancel 01KPW...`, which lands as CancelTaskFn writes `status=CANCELLING`. At 10:05:31 AM the user — forgetting they just cancelled, or running from a different terminal where they didn't see the cancel — runs `bgagent approve 01KPW... 01KPR...`. Without the cross-table transaction, the Lambda's GetItem on TaskTable (separate call) might read the stale RUNNING state, then UpdateItem on TaskApprovalsTable succeeds because the approvals row is still PENDING → 202 returned. The user sees "approved!" but the task is dying. With the TransactWriteItems pattern, both conditions must hold: the TaskTable guard `status = AWAITING_APPROVAL` fails (because it's now CANCELLING), the entire transaction rolls back, the Lambda returns 409 `TASK_NOT_AWAITING_APPROVAL` with `current_status: CANCELLING`. The user sees "cannot approve: task is already cancelling" and correctly understands state. The cost is one extra table in the transaction (two instead of one) — still within DDB's 100-item limit and nowhere near the 4 MB request size. Symmetric with the agent's resume transaction, which already does the cross-table guard. +**Scenario (finding #7):** A user cancels a task and then approves its old request +from another terminal. Cancellation writes `CANCELLED` directly; there is no +`CANCELLING` task state. The approval transaction cannot commit because the task +is no longer `AWAITING_APPROVAL`. The P3 cancellation update also atomically closes +the linked `PENDING` approval as `CANCELLED` and writes an `approval_cancelled` +event. A later decision is rejected: the existing API returns +`404 REQUEST_NOT_FOUND` for missing, foreign or already-decided approval rows, +including a cancelled row. If approval committed first, cancellation preserves +the recorded decision while cancelling the task. See the +[P3 approval verification record](../verification/README.md) +for source versus deployment status. ### 7.2 `POST /v1/tasks/{task_id}/deny` @@ -1058,7 +914,7 @@ New field reference: | Field | Type | Required? | Default | Description | |---|---|---|---|---| -| `approval_timeout_s` | integer seconds | No | **300** | Per-task default approval timeout. Bounded by `[30, min(3600, maxLifetime - 300)]`. Per-rule `@approval_timeout_s` annotations may clip this further (min-wins; see decision #6 and §6.3). Default matches the §10.2 `TaskTable.approval_timeout_s` default. | +| `approval_timeout_s` | integer seconds | No | **0** | Zero retains unanswered requests without a deadline. Positive values must be 30–3,600. The shortest positive task/rule deadline wins. | | `initial_approvals` | list of scope strings | No | `[]` | Pre-approval allowlist scopes (≤20 entries, ≤128 chars each). Validated per §7.3 rules below. | `CreateTaskFn` validations: @@ -1072,7 +928,7 @@ New field reference: - `write_path:X` — same rules as bash_pattern - `rule:X` — X must exist in the (built-in + target repo's blueprint) soft-deny policy set per the shared policy-parsing library; hard-deny rule IDs rejected - `all_session` — rejected if `Blueprint.security.maxPreApprovalScope` forbids -5. `approval_timeout_s` within `[30, min(3600, maxLifetime - 300)]` — cap at 1 hour OR (maxLifetime - 5min), whichever is smaller. Prevents multi-hour slot-exhaustion attacks and keeps approval windows within the TTL budget. +5. `approval_timeout_s` is zero or an integer from 30 through 3,600. MicroVM worker lifetime and capacity are bounded separately through checkpointed retirement; requests are not deleted by a pending-row TTL. 6. Combined `hard + soft + disable + custom` Cedar text size ≤ 64 KB (§12.4); reject on overflow. ### 7.4 Degenerate-pattern detection @@ -1131,7 +987,11 @@ Rate-limited 30/min/user; cached 5min per repo in-Lambda. ### 7.7 `GET /v1/pending` — list pending approvals across user's active tasks -Returns all approvals with `status=PENDING` owned by the caller. Backing index: `user_id-status-index` GSI on `TaskApprovalsTable` (see §10.1). +Returns up to 100 caller-owned pending approvals whose consistently read tasks +are still awaiting that exact request. The `user_id-status-index` GSI supplies +candidates; pagination continues past cancelled/orphaned rows within a shared +five-second read budget. A read failure returns an error rather than a misleading +partial list. **Request**: `GET /v1/pending` with Cognito auth. @@ -1199,7 +1059,7 @@ bgagent submit --task "..." --pre-approve all_session --yes `--pre-approve-file` reads a YAML/JSON array of scope strings — supports the 20-entry cap without command-line bloat. -`--approval-timeout` default (CLI and server): **300 seconds** (5 min), matching decision #6, §7.3, and the `TaskTable.approval_timeout_s` default in §10.2. Accepted range `[30, min(3600, maxLifetime - 300)]` — CLI validates client-side and the server re-validates. `bgagent submit --help` surfaces the default explicitly. +`--approval-timeout` default (CLI and server): **0**, meaning no automatic decision deadline. A positive value must be 30–3,600 seconds. The CLI validates it and the server re-validates it. Matching positive rule deadlines still apply. ### 8.3 Streaming UX @@ -1279,21 +1139,25 @@ TaskApprovalsTable rows are **terminal on first decision** — a row never re-op ```mermaid stateDiagram-v2 - [*] --> PENDING: Agent writes row
(TransactWriteItems with
TaskTable → AWAITING_APPROVAL) + [*] --> PENDING: Approval service creates row
(transaction with
TaskTable → AWAITING_APPROVAL) PENDING --> APPROVED: ApproveTaskFn
(cross-table transaction) PENDING --> DENIED: DenyTaskFn
(cross-table transaction) - PENDING --> TIMED_OUT: Agent poll timeout
(best-effort update) - PENDING --> STRANDED: Reconciler detects
orphan (age > 2×timeout_s) + PENDING --> TIMED_OUT: Approval service records
worker timeout + PENDING --> CANCELLED: Task owner cancels
(cross-table transaction) APPROVED --> [*]: terminal DENIED --> [*]: terminal TIMED_OUT --> [*]: terminal - STRANDED --> [*]: terminal + CANCELLED --> [*]: terminal note right of APPROVED - TTL = created_at + timeout_s + 120s - DDB reaps row after TTL + Terminal retention TTL + never expires a pending request end note ``` +`STRANDED` remains a recognized row status in the type contract. The current +reconciler does not write it: it fails the owning task and closes pending requests +as `CANCELLED`, as described below. + ### 9.3 Orchestrator impact - `waitStrategy` adds `AWAITING_APPROVAL` as non-terminal. @@ -1305,7 +1169,7 @@ stateDiagram-v2 **AWAITING_APPROVAL holds the user's concurrency slot.** -Rationale: the Docker container is alive. Memory allocated. The AgentCore microVM pool is committed. Releasing the slot while the resource is still held lies to accounting and opens a resource-exhaustion vector. +Rationale: the task still owns an unfinished compute session and may resume work. Retaining its ABCA reservation prevents an unbounded collection of parked tasks from bypassing admission control. The rule also applies to P3 Lambda MicroVM suspension; suspended AWS memory-quota consumption remains unverified and is not the basis for claiming quota usage. Resume and terminal cleanup use the existing task-owned reservation protocol. Concrete behavior: @@ -1322,18 +1186,28 @@ t=45m: Task #1 completes. count → 9. Bob can submit task #11. AgentCore Runtime's `maxLifetime = 28800s` (8h) is an absolute timer from session start. It does NOT pause during `AWAITING_APPROVAL`. -This has a concrete implication: the hook computes an `effective_timeout` bounded by `maxLifetime - remaining - CLEANUP_MARGIN_120S`. If the task has been running 7h55m and hits a soft-deny gate, the effective timeout might be clamped to a much shorter value than the task default. Below the 30s floor → immediate DENY with reason `"insufficient lifetime"`. +When the worker supplies a remaining-lifetime estimate, the hook refuses to open a new gate if fewer than 30 seconds remain after the 120-second cleanup margin. This check applies to timed and untimed approvals and reports `"insufficient maxLifetime remaining (s) for approval"`. A positive approval timeout is also capped by that remaining budget. + +The default approval timeout is `0`: no decision deadline. Worker lifetime does not turn silence into a human denial. ECS and AgentCore do not currently restore a waiting agent into a replacement worker; when their task execution limit is reached, the task closes and its pending approval is cancelled. MicroVM can retain a verified checkpoint and continue on a replacement worker. Setting its sleep delay to `0` disables early sleep/retirement, but the coordinator still attempts retirement before the service lifetime ends. This option trades idle cost for faster replies. ### 9.6 Stranded-approval reconciliation -`reconcile-stranded-tasks.ts` gains an AWAITING_APPROVAL-aware branch: +`reconcile-stranded-tasks.ts` has an AWAITING_APPROVAL-aware branch: -- Detects tasks in AWAITING_APPROVAL with `age > 2 * timeout_s` -- Best-effort conditional-updates TaskApprovalsTable row → `STRANDED` status -- Transitions TaskTable → `FAILED` with reason `"approval stranded (container eviction)"` -- Emits `approval_stranded` event to TaskEventsTable +- Uses `APPROVAL_STRANDED_TIMEOUT_SECONDS`, default 30,600 seconds (8.5 hours), measured from + entry into the current status. It does not calculate twice each row's timeout. +- Conditionally changes the task to `FAILED` if it is still awaiting approval, + recording the elapsed wait and a recovery suggestion. +- Emits `task_stranded`, `task_failed` and a wrapped `approval_stranded` milestone. + The legacy milestone has no request ID. +- Closes pending approval rows as `CANCELLED`. Late approval is also rejected by + the task-state guard. A saved MicroVM continuation is handled by its dedicated + coordinator instead of this timeout. -This closes the container-eviction gap. Without this, a container restart mid-approval would leave the task hanging until the user manually cancelled. +The notification helper can recover that legacy milestone's request identity +from the consistently read failed task and its saved stranded cause. It verifies +approval ownership before rendering feedback and does not change the row. +The timer is a backstop, not proof of a particular container failure. `reconcile-concurrency.ts` (scheduled every 5 min) already scans for orphaned concurrency counters; with `AWAITING_APPROVAL` added to `ACTIVE_STATUSES` it correctly counts awaiting tasks as active. @@ -1347,31 +1221,11 @@ The design assumes a human is watching. For truly unattended tasks (scheduled au ### 10.1 New DynamoDB table: `TaskApprovalsTable` -```typescript -new dynamodb.Table(this, 'Table', { - partitionKey: { name: 'task_id', type: dynamodb.AttributeType.STRING }, - sortKey: { name: 'request_id', type: dynamodb.AttributeType.STRING }, // ULID - billingMode: dynamodb.BillingMode.PAY_PER_REQUEST, - pointInTimeRecovery: true, - timeToLiveAttribute: 'ttl', - stream: dynamodb.StreamViewType.NEW_AND_OLD_IMAGES, // (evaluated — may drop; see §11) - removalPolicy: RemovalPolicy.RETAIN, -}); - -// v1 GSI — backs `GET /v1/pending` and `bgagent pending`. -// Required at v1 ship, not deferred — see finding #8 scenario in §7.7. -table.addGlobalSecondaryIndex({ - indexName: 'user_id-status-index', - partitionKey: { name: 'user_id', type: dynamodb.AttributeType.STRING }, - sortKey: { name: 'status', type: dynamodb.AttributeType.STRING }, - projectionType: dynamodb.ProjectionType.INCLUDE, - nonKeyAttributes: [ - 'task_id', 'request_id', 'tool_name', 'tool_input_preview', - 'severity', 'reason', 'created_at', 'timeout_s', - 'matching_rule_ids', - ], -}); -``` +[TaskApprovalsTable](../../cdk/src/constructs/task-approvals-table.ts) defines the +`task_id` partition key, `request_id` sort key, optional retention TTL attribute +`ttl`, and `user_id-status-index` GSI. Streams are disabled; TaskEventsTable carries +the audit/fan-out stream. The GSI discovers candidates; the pending endpoint then +strongly reads the approval and owning task before returning a request. **Projection is fixed at design time.** DynamoDB rejects in-place updates to a GSI's `nonKeyAttributes` (the CloudFormation error @@ -1399,17 +1253,18 @@ Attributes: | `reason` | S | Yes | Cedar matching rule description | | `severity` | S | Yes | "low" \| "medium" \| "high" | | `matching_rule_ids` | L | Yes | List (not Set — can be empty) of soft-deny rule IDs | -| `status` | S | Yes | PENDING \| APPROVED \| DENIED \| TIMED_OUT \| STRANDED | +| `status` | S | Yes | PENDING \| APPROVED \| DENIED \| TIMED_OUT \| STRANDED \| CANCELLED | | `created_at` | S | Yes | ISO8601 | -| `decided_at` | S | No | Set when status != PENDING | +| `decided_at` | S | No | Set when the trusted service or platform closes a request | | `scope` | S | No | Set on APPROVED | | `deny_reason` | S | No | Set on DENIED; sanitized user text | | `timeout_s` | N | Yes | Resolved timeout for audit | -| `ttl` | N | Yes | `created_at_epoch + timeout_s + CLEANUP_MARGIN_120S` — always covers the decision window | +| `expires_at` | S/null | Yes | Original decision deadline; null when `timeout_s=0` | +| `ttl` | N | No | Retention cleanup set on task closure; absent while pending | | `user_id` | S | Yes | Cognito `sub` **verbatim**; used in ownership check `ConditionExpression` (§7.1 finding #6) | | `repo` | S | Yes | Denormalized for fan-out | -**TTL sizing**: the TTL is always `timeout_s + 120s`, so a 300s approval window has a 420s TTL, a 3600s window has a 3720s TTL. The row never expires during the decision window. After the decision + a short grace period, DDB's eventual-consistency TTL reaper cleans up. +**TTL and decision deadlines are separate.** Pending approval rows have no TTL, including explicitly timed requests. Task closure cancels unanswered requests and assigns retention TTL to its approval records. Already-recorded decisions are preserved, and retries do not extend an existing retention deadline. **Why a list, not a StringSet, for `matching_rule_ids`**: DDB string sets cannot be empty. Pathological no-match soft-deny hits would fail to persist. Lists handle empty gracefully. @@ -1421,7 +1276,7 @@ Five new attributes on the existing task row: | Name | Type | Required | Description | |---|---|---|---| -| `approval_timeout_s` | N | No | Default timeout for soft-deny gates. Default 300. | +| `approval_timeout_s` | N | No | Task setting for soft-deny gates. Default 0/no deadline; positive values are 30–3,600 seconds. | | `initial_approvals` | L | No | List of scope strings from submit time | | `awaiting_approval_request_id` | S | No | Set when status = AWAITING_APPROVAL; cleared on transition back (via joint `UpdateExpression`) | | `approval_gate_count` | N | No | Running counter of approval gates fired on this task; used to enforce `approval_gate_cap` (decision #13) | @@ -1479,16 +1334,39 @@ Emitted to both `ProgressWriter` (DDB, 90d) and `sse_adapter` (live stream). Plu ### 11.2 Fan-out plane interaction — Slack button → Cognito mapping -Approval events flow to the fan-out Lambda via TaskEventsTable Streams (the existing Phase 1b path). They are dispatched to Slack / GitHub / Email stubs. - -**TaskApprovalsTable Streams are not consumed by the fan-out Lambda**. The approval row is working state; the audit trail is in TaskEventsTable. Enabling Streams on TaskApprovalsTable would be redundant and add noise. Final design: TaskApprovalsTable DOES NOT have Streams enabled. (Retains the `stream` attribute commented out for future use if needed.) - -Fan-out dispatch rules (extending Phase 1b stubs): -- Slack: on `approval_requested` OR `approval_stranded` — "Agent @task_id requests approval for Bash: `git push --force`" -- Email: on `approval_requested` with `severity: high` -- GitHub: none - -**Rate-limited per-user**: 10 approval-related fan-out messages per user per minute. Prevents notification-spam from malicious users driving up approval-gate count. +Approval events flow to the fan-out Lambda via TaskEventsTable Streams. The P3 +implementation adds Slack and Linear notifications with the saved action, +reason, decision deadline and exact CLI approve/deny commands. Recorded decisions, +cancellations, timeouts and stranded waits also produce messages. The response +path supports the CLI and native Linear thread replies. Slack approval buttons +and the Slack OAuth/button design below remain proposed. Email remains a log-only stub and GitHub does not +receive approval messages. Deployment status is recorded in the +[P3 verification record](../verification/README.md). + +**TaskApprovalsTable Streams are not consumed by the fan-out Lambda**. The approval row is working state; the audit trail is in TaskEventsTable. Enabling Streams on TaskApprovalsTable would be redundant and add noise. Final design: TaskApprovalsTable DOES NOT have Streams enabled. + +Slack and Linear route `approval_requested`, `approval_decision_recorded`, +`approval_timed_out`, `approval_cancelled` and `approval_stranded`. The dispatcher +reads the current approval row and owning task before displaying a pending request, +and records successful delivery per request/channel. Delivery failure does not mark +the message delivered. A post that succeeds just before receipt persistence fails +can still produce a duplicate on retry, except for Linear approval prompts and +reply acknowledgements, which use deterministic comment IDs. + +For Linear, the platform saves a thread binding to the exact workspace, issue, +task, request and owner before posting the prompt. A verified Comment/create +webhook containing only `approve` or `deny` in that thread resolves the commenter's +linked platform identity. The owner then uses the same atomic decision function, +rate limit, deadline and current-request guards as the API. Approval grants +`this_call`. A trusted source comment ID saved in the decision transaction makes +webhook retries recognize the original result. Top-level comments, edits and +unbound threads never select a pending request. Pending bindings have no TTL; +closure starts 90-day retention. MicroVM decisions use the existing wake or +continuation path; other backends keep their existing polling behavior. + +**Proposed notification rate limit:** 10 approval-related messages per user per +minute. This dispatcher limit is not implemented. Existing gate-creation caps and +API rate limits remain, but they are not a substitute for notification throttling. **Notification plane is observability, not state (see §13.14).** Notification delivery failures do NOT pause the approval timer — coupling the two creates a bypass where an adversary who takes down the webhook gets an unbounded approval window. The timer runs on the agent's local clock keyed to `created_at`; `bgagent pending` is the recovery path for users who suspect notifications are broken (backed by `user_id-status-index` GSI, §7.7). For the off-hours / unattended trade-off that this posture implies, see §14.8. @@ -1614,11 +1492,33 @@ These alarms transition to `ALARM` state in CloudWatch and appear in the console ### 12.1 Trust boundaries -- **Agent container ↔ TaskApprovalsTable**: IAM role on the runtime has `GetItem` / `PutItem` / conditional `UpdateItem` on the table. Agent writes pending, reads decisions, writes TIMED_OUT on internal timeout. +- **Agent container ↔ TaskApprovalsTable**: workers have task-scoped reads and + transaction condition checks, with no direct item writes or deletion. +- **Agent container ↔ approval request service**: IAM restricts signed + `POST /v1/tasks/{task_id}` to the session's `task_id` tag. The service creates + `PENDING` rows or conditionally records a non-human `TIMED_OUT`; it rejects + human decisions, notification markers, retention TTL and other extra fields. + IAM binds the caller to a task path; the transaction prevents concurrent task + ownership/state changes and, for MicroVM, checks the active worker lease. + The tool preview, its hash, severity, reason and matching rules are worker + assertions, not independently evaluated policy results. The preview can be + truncated, so its bytes cannot be used to verify the full-input hash. This protects approval records; it does not sandbox code running + inside the agent or bind ambient compute credentials to one task. - **User CLI ↔ API Gateway**: Cognito JWT (same authorizer as `/tasks/*`). Cognito `sub` is the canonical caller identity, used **verbatim** in DDB `ConditionExpression` (§7.1, finding #6). - **ApproveTaskFn/DenyTaskFn ↔ TaskApprovalsTable + TaskTable**: Lambda IAM policy allows `UpdateItem` on both tables under `TransactWriteItems`. Authorization is in the ConditionExpression (ownership AND state), not in a separate IAM boundary. - **Blueprint origin**: blueprints are CDK-deployed constructs (see `cdk/src/constructs/blueprint.ts`). Platform operators deploy them. Users cannot upload arbitrary blueprint.yaml from the target repo. This property is load-bearing for the security model — if blueprint origin ever becomes user-uploaded, the blueprint-injection section (§12.4) must be re-evaluated. The 64 KB text cap (§5.1, finding #12) and `disable:` hard-deny rejection (finding #9) are applied regardless of origin as defense in depth. -- **Slack → ApproveTaskFn**: mediated by the fan-out Lambda + `SlackUserMappingTable` (§11.2). Slack admin cannot forge mappings; Slack approvals capped at `severity: low|medium` (finding #4). +- **Linear → decision handlers**: the mapped task owner must author the actual + reply returned by Linear's API. Bot comments and comments authored as the + saved OAuth token identity are rejected. For approvals, install with the + default `actor=app`; diagnostic user-mode credentials cannot approve on + behalf of their authorizing user. Use the authenticated CLI in that case. + A webhook signature alone is insufficient: legacy workers can read a bundle + containing OAuth credentials and the webhook signing key. The authoritative + readback protects this approval path, not every other webhook action. + The Slack button/proxy design in §11.2 remains future work. + +See [ADR-023](../decisions/ADR-023-trusted-approval-writer.md) for the decision and limits, and [upgrading approval permissions](../guides/DEPLOYMENT_GUIDE.md#upgrading-approval-permissions) +before updating an existing deployment. ### 12.2 Ownership encoded in ConditionExpression @@ -1627,9 +1527,12 @@ No TOCTOU window. The `TransactWriteItems` (§7.1) encodes across two tables: - TaskApprovalsTable: `#status = :pending AND user_id = :caller` - TaskTable: `#status = :awaiting AND awaiting_approval_request_id = :rid` -Authorization + approvals-state + task-state transition all atomic. A compromised internal caller (Lambda with raw DDB access) or a logic bug in a future refactor that forgets the ownership check still can't flip rows without matching the `user_id`. The task-state guard additionally prevents the "approve succeeds on a cancelled task" race (finding #7). +The handlers check authorization, approval state and task state atomically. +These conditions prevent races; they do not constrain a compromised Lambda with +raw table-write permission, which could omit them. Only trusted control-plane +handlers receive that permission. -`user_id` comparison is against Cognito `sub` **verbatim** — byte-for-byte equality. Any future identity transformation (per-tenant prefixing, namespacing) must apply to BOTH the write path (agent-side row write) AND the compare path (Lambda ConditionExpression) simultaneously, or the comparison silently fails under the new format. A unit test (§15.3) enforces this: given a sample JWT, extract `sub`, write a row, then assert the stored `user_id` equals `sub` byte-for-byte. +`user_id` comparison is against Cognito `sub` **verbatim** — byte-for-byte equality. Any future identity transformation (per-tenant prefixing, namespacing) must apply to BOTH the write path (trusted approval service) AND the compare path (Lambda ConditionExpression) simultaneously, or the comparison silently fails under the new format. A unit test (§15.3) enforces this: given a sample JWT, extract `sub`, write a row, then assert the stored `user_id` equals `sub` byte-for-byte. ### 12.3 Race prevention @@ -1638,11 +1541,11 @@ Authorization + approvals-state + task-state transition all atomic. A compromise - User's CLI writes `APPROVED WHERE status = :pending` (via TransactWriteItems) - One wins atomically - The loser: - - If TIMED_OUT wins: user gets 409 `REQUEST_ALREADY_DECIDED`. User sees "approval expired". + - If TIMED_OUT wins: a later decision gets 404 `REQUEST_NOT_FOUND`; the saved row is already closed. - If APPROVED wins: agent's poll reads APPROVED on next tick. Agent proceeds. **Race 2 — double-approve**: -- Two concurrent CLI invocations. Second gets 409 `REQUEST_ALREADY_DECIDED`. Idempotent. +- Two concurrent CLI invocations. Only one decision commits; the second gets 404 `REQUEST_NOT_FOUND`. **Race 3 — cancel during AWAITING_APPROVAL**: - Agent writes `RUNNING WHERE status = :awaiting AND awaiting_approval_request_id = :rid` @@ -1742,7 +1645,7 @@ Tracked as IMPL-22. Without these telemetry-driven re-evaluations, 50 will ossif ### 12.10 JWT replay -Cognito JWT with signature + expiry validation on API Gateway. Approval row conditional-update prevents replay from mutating state. Slack button replays similarly mediated by `SlackUserMappingTable` (§11.2) + Slack's own request signing. +Cognito JWT with signature + expiry validation on API Gateway. Approval row conditional-update prevents replay from mutating state. Native Slack approval buttons are not implemented; §11.2 describes the proposed flow. --- @@ -1784,7 +1687,7 @@ The per-task `approvalGateCap` (decision #13; default 50, configurable) is **per ### 13.7 Insufficient lifetime remaining for approval -If `remaining_maxLifetime - CLEANUP_MARGIN_120S < FLOOR_30S`, hook immediately returns DENY with reason `"insufficient maxLifetime for approval"`. Task continues without a gate — or, if the gate was load-bearing, fails gracefully in RUNNING state. +Without a continuation runtime, if `remaining_maxLifetime - CLEANUP_MARGIN_120S < FLOOR_30S`, the hook immediately denies the tool with reason `"insufficient maxLifetime remaining ({n}s) for approval"`, where `{n}` is the remaining lifetime in seconds. No approval request is created and the tool does not run. The agent receives the denial and may choose another action. ### 13.8 PreToolUse hook itself crashes @@ -1810,34 +1713,38 @@ Addressed by the parity contract (decision #23, §15.6). Golden-file CI test run ### 13.12 VM-throttle + late-approval race -The agent's poll loop computes a local timeout wall-clock (`timeout_s` worth of elapsed monotonic time). If the VM is throttled by the hypervisor — either an AgentCore noisy-neighbor eviction window, or a CPU-throttle under memory pressure — poll ticks can stretch past their nominal cadence. In the worst case, the user's APPROVE transaction lands in DDB a few hundred milliseconds before the agent's local clock trips past `timeout_s` and the agent attempts to write `status = TIMED_OUT WHERE status = :pending`. The ConditionCheckFailed path fires (APPROVED already won), but without a re-read the agent's local state is stale: `outcome.status == "TIMED_OUT"` locally while DDB holds APPROVED. The agent would return DENY, the user sees "I approved it" and the agent still blocks — a confounding experience that also violates the design principle that user-observed state is authoritative. +The agent's poll loop retains the original approval row's UTC expiry and a monotonic cap, using whichever expires first. Slow writes/notifications count toward that same window. This also covers a MicroVM whose monotonic clock stops while suspended; the `/resume` hook reuses this deadline and wakes the decision loop. If CPU throttling or suspension delays polling beyond expiry, a user's APPROVE transaction may already have committed. The agent asks the trusted service to conditionally record `TIMED_OUT` while the row remains `PENDING`. The condition fails because APPROVED already won. Without a re-read, the agent's local state would remain TIMED_OUT while DDB holds APPROVED, and it would incorrectly deny the approved call. -**Mitigation**: the §6.5 pseudocode re-reads the approval row with `ConsistentRead=True` whenever `_best_effort_update_status("TIMED_OUT", ...)` returns ConditionCheckFailed, and honors whatever terminal state the row carries: +**Mitigation**: the hook rereads the approval row with `ConsistentRead=True` when the service reports that another decision won the timeout race, and honors the recorded decision: - If `status == "APPROVED"`: rebuild the local `outcome` to reflect APPROVED, preserving `scope`, `decided_by`, `decided_at`, and proceed through the normal allow flow (scope-propagation, `approval_granted` milestone, resume transaction, return `{"permissionDecision": "allow"}`). Emit a `approval_late_win` milestone so operator telemetry can count races. - If `status == "DENIED"`: honor the denial text the user submitted. Agent returns DENY with the user's sanitized reason as `permissionDecisionReason` (same surface as normal deny). -- If `status` is still PENDING (rare — concurrent reaper race) or the row is gone (TTL reaped): fall through with the original TIMED_OUT outcome; fail-closed deny. +- If `status` is still PENDING (rare — concurrent reaper race) or the row is missing: fall through with the original TIMED_OUT outcome; fail-closed deny. The scenario is bounded by the polling cadence (2-5s ticks) and DDB's strongly-consistent read latency (tens of ms), so the re-read adds at most one extra GetItem to the racing path — acceptable cost for honoring user intent. See IMPL-24 and §15.2 task #43 (race tests). ### 13.13 Runtime JWT expiry during approval wait -**Context in this codebase (verified 2026-05-06).** The AgentCore Runtime container authenticates outbound AWS API calls (DynamoDB, Secrets Manager, etc.) via the container's IAM role, which the SDK resolves through the instance-metadata-service equivalent and auto-refreshes transparently. There is no user-presented JWT with a short rolling expiry consumed by the container's own API calls — `grep -rn -iE 'runtime.jwt|jwt.refresh|token_expiry' agent/src/` returns nothing (only `token_usage` for LLM billing and `GITHUB_TOKEN` for git operations). AgentCore Runtime invocation on the Lambda side uses sigv4 via `InvokeAgentRuntimeCommand` (see `cdk/src/handlers/shared/strategies/agentcore-strategy.ts`) — also auto-refreshed AWS credentials, not a user JWT. The "Runtime-JWT" label in §4 step 6 and the sequence diagrams below refers to the **caller-facing SSE auth** (Terminal A's Cognito ID token presented to API Gateway to stream task events) — it does not authenticate the container's own DDB writes. - -**Therefore, for v1: no separate Runtime JWT expiry term is required in the ceiling computation.** The `maxLifetime` term (AgentCore's hard lifetime of 8h) is the only upper bound we control; IAM credentials refresh automatically within that window. The ceiling definition in decision #6 stands as `min(1h, maxLifetime_remaining - cleanup_margin)`. +Workers authenticate AWS calls with IAM role credentials, not the user's Cognito +JWT. The CLI's Cognito token protects its platform API calls; an expired CLI login +requires reauthentication but does not itself delete the saved approval request. +MicroVM resume refreshes runtime and task-role credentials before releasing coding. -**If the auth model changes** (e.g. a future design introduces a container-held user JWT to authenticate `permissionDecisionReason` attribution, or to carry the caller's Cognito `sub` end-to-end for per-user DDB conditions), the ceiling MUST be extended to `min(1h, maxLifetime_remaining - 120s, runtime_jwt_expiry - 120s)` and this section updated. Tracked as IMPL-27 so the contract is reviewed whenever the auth shape changes. The failure signature if this bound is missed: the container's IAM calls succeed but some JWT-gated channel (e.g. Terminal A's SSE stream) quietly 403s mid-approval-wait; the user's decision lands in DDB but the agent's poll fails to deliver `approval_granted` to the live stream. Today that channel is best-effort observability, not state — but a future state-bearing channel would need the ceiling term. +Credential renewal does not extend a worker's service lifetime. Apply the worker +and explicit-deadline rules in §9.5; do not impose a new approval deadline based on +the user's API token expiry. A future worker-held user-token design would need a +separate review of that boundary. ### 13.14 Notification delivery failure -Fan-out delivery failures (Slack down, email bounce, webhook 5xx) do **NOT** pause the approval timer. The timer runs on the agent's local clock, keyed to the `created_at` timestamp on the DDB row — it is independent of whether any notification channel succeeded in alerting the human. +A failed notification does not change the recorded decision deadline. Timed +requests retain their original UTC/monotonic deadline; untimed requests remain +available until task closure or another valid resolution. Neither case permits +the action without approval. -**Rationale (security):** coupling the timer to notification-plane availability creates a bypass. An adversary who takes down the webhook (or poisons the Slack rate limit) would get an unbounded approval window; worse, a compromised tenant could deliberately suppress their own notifications to escape gates. Fail-closed on timer expiry is invariant; delivery is best-effort observability. +Users can discover unanswered requests with `bgagent pending`, which reads through +the authenticated API independently of notification delivery. A late-discovered +explicitly timed request does not receive a fresh decision window. -**Recovery path for the user:** `bgagent pending` queries `TaskApprovalsTable` directly via the `user_id-status-index` GSI (§7.7, §10.1) — it does not depend on notification delivery. A user who suspects notifications are broken can poll `bgagent pending` at any time to see all live approvals. If the notification never landed and the user finds a gate via `bgagent pending`, they can `bgagent approve/deny` normally; the timer is still running against the original `created_at`, not against when the user found it. - -**Operational signal:** `approval_timed_out` events carry `timeout_s` and the `created_at`/`decided_at` delta. A rising `approval_timed_out` rate with flat `approval_requested` rate (measured via `ApprovalTimeoutClipRate` and `ApprovalDecisionLatency` in §11.3) is the telemetry that indicates notification breakage, not an unresponsive user. - -**The fail-closed posture on timer expiry remains unchanged.** Delivery-availability-aware scheduling is the notification plane's job (see §14.8 and INTERACTIVE_AGENTS.md notification-plane design); the timer does not reason about it. ### 13.15 Fail-closed summary @@ -1898,7 +1805,7 @@ BLOCKED[]: (resource: ) # when a resource is named ### 14.1 Scenario A: force-push with per-rule timeout -Setup: repo `my-org/my-app` blueprint extends soft-deny with `force_push_main` (@approval_timeout_s=600). Task default is 300s. +Setup: repo `my-org/my-app` blueprint extends soft-deny with `force_push_main` (@approval_timeout_s=600). This example explicitly sets the task timeout to 300s. ```bash $ bgagent run --repo my-org/my-app \ @@ -2016,7 +1923,7 @@ Each phase has explicit scope. Matches real-world review workflows. Visible in a ### 14.6 Scenario F: VM-throttle + late-approval race (trace) -Setup: task default 300s; force-push gate fires. User Alice approves at the very edge of the timeout window while the VM is throttled. +Setup: task explicitly configured with a 300-second timeout; force-push gate fires. User Alice approves at the very edge of the timeout window while the VM is throttled. ``` t=0.00s PreToolUse hook fires: Bash "git push --force origin feature-x" @@ -2025,22 +1932,18 @@ t=0.02s TransactWriteItems: approval row PENDING + TaskTable AWAITING_APPROVA t=0.03s agent_milestone: approval_requested → Terminal A stream t=0.04s _poll_for_decision begins; interval=2s for first 30s, then 5s -... (poll ticks every 5s from t=30 to t=295) ... +... (poll ticks every 5s from t=30 to t=285) ... t=285.0s host hypervisor evicts VM from warm CPU share (noisy neighbor). - Next scheduled poll was t=290.0s; actual scheduling delay ~5.1s. + Next scheduled poll was t=290.0s; actual scheduling delay ~10.2s. t=294.7s Alice's bgagent approve lands at API Gateway. ApproveTaskFn TransactWriteItems: ApprovalsTable: PENDING → APPROVED (user_id matches, status was PENDING) TaskTable: state guard holds (still AWAITING_APPROVAL, rid matches) → 202 returned to CLI; Alice sees "approved!" in Terminal B -t=295.1s agent's delayed poll tick fires. Elapsed wall-clock = 295.1s. - Monotonic elapsed is 295.1 > timeout_s=300? NO — but the poll - function computes `elapsed >= timeout_s` and on the NEXT tick - (t=300.2s) it will exceed. -t=300.2s next tick: elapsed=300.2 ≥ timeout_s=300 → TimedOut() returned. - (Alice's APPROVED write at t=294.7s was MISSED — the previous - poll was due at t=295.0 but the VM throttle stretched it past.) +t=300.2s delayed tick: original deadline has passed → TimedOut() returned + before reading the approval row. Alice's APPROVED write at + t=294.7s has not yet been observed by this agent. t=300.3s _best_effort_update_status("TIMED_OUT", ... WHERE status = :pending) → ConditionCheckFailed (row is APPROVED, not PENDING) → wrote_timeout = False @@ -2077,16 +1980,20 @@ Setup: Bob submits with `--approval-timeout 600`. The blueprint has a `write_cre Bob sees both the pre-submit warning (`approval_timeout_capped_at_submit`) and the per-gate cap event (`approval_timeout_capped`) so he understands why his 600s didn't apply. Without these milestones, the user sees only `timeout: 300s` in the approval banner and may think the CLI dropped their setting. Both events are captured in the event stream and surface via `bgagent watch`. See §11.1, §11.3 (`ApprovalTimeoutClipRate`), and Fix 4 / IMPL-26 in §16. -### 14.8 Off-hours and unattended tasks (known trade-off) +### 14.8 Off-hours and unattended tasks -**Known trade-off: off-hours failure.** Because timeouts are fail-closed (decision #6), a task running overnight with pending approvals will fail if no approver responds in time. This is deliberate — auto-approve on timeout would make "wait the reviewer out" the attacker's winning strategy; see decision #6 and §13.15 fail-closed summary. +Unanswered requests now remain available by default. On MicroVMs, a verified +checkpoint and worker retirement bound resource use while the person is away; +their later answer can start a replacement worker. Other compute substrates keep +their existing runtime limits. This does not approve any action automatically. -**For overnight / unattended runs, choose one of:** -- `--pre-approve all_session --yes` to bypass gates entirely for that task (accept the broader trust grant; see §7.3). -- Configure escalation on the `approval_requested` event via the notification plane (see `docs/design/INTERACTIVE_AGENTS.md` for channel configuration). Route to whoever is on-call; escalation schedule is the tenant's responsibility, not the timeout engine's. -- Schedule the task during business hours. +For a time-sensitive request, set an explicit positive decision deadline. Its +expiry remains fail-closed: the tool is denied, and the agent decides what to do +next. Notification routing can still notify an on-call reviewer; scheduling that +escalation belongs to the notification layer. -The `approvalGateCap` (decision #13) will force-fail the task after approximately `cap × task_default_timeout_s` of unanswered gates — default worst case ~4h at cap=50 / timeout=300s. Plan accordingly. +`approvalGateCap` limits the number of gates reached by a task, not elapsed human +waiting time. It does not impose a deadline on one unanswered request. **Why the timer itself is timezone-unaware:** Business-hours logic belongs in the notification plane, not the authorization engine. Baking calendars or on-call rotations into the timer couples the security boundary to a scheduling system it doesn't own. Same rule evaluated at 9am and 3am because the security property (adversary cannot wait out review) is time-invariant. Delivery-availability-aware scheduling is the notification plane's job via subscribed-channel health and escalation policies. @@ -2124,7 +2031,7 @@ See §17.18 for the off-hours escalation future-work primitive, and §13.14 for | # | Package | File | Change | |---|---|---|---| | 1 | agent | Spike | Validate cedarpy.policies_to_json_str() returns annotations. Confirm `diagnostics.reasons` shape for multi-match. If API diverges, update §6 before proceeding. | -| 2 | mise + agent + cdk | `mise.toml`, `agent/pyproject.toml`, `cdk/package.json` | Pin `cedarpy==4.8.0` (agent) and `@cedar-policy/cedar-wasm==4.10.0` (cdk). The two bindings are intentionally on different version lines — verified compatible via the parity fixtures, not required to be equal. Both pinned exactly, not `^` or `~` — decision #23 / finding #1. | +| 2 | mise + agent + cdk | `mise.toml`, `agent/pyproject.toml`, `cdk/package.json` | Pin both Cedar bindings exactly in their package manifests and verify the pair through the shared parity fixtures; see §15.6. | | 3 | agent + cdk | `contracts/cedar-parity/*.json` (shared fixture dir; follows precedent set by `contracts/memory-hash-vectors.json`) | Golden-file parity fixtures: `(policy_set, input) → {decision, matching_rule_ids}`. Agent side loads via `cedarpy`; Lambda side via `cedar-wasm`. Divergence fails CI. | | 4 | agent | `src/policy.py` | Extend `PolicyDecision` (outcome/timeout_s/severity/matching_rule_ids/allowed-property). Split `_DEFAULT_POLICIES` into hard + soft. Add annotation parsing. Implement `ApprovalAllowlist` + `RecentDecisionCache` (50-entry LRU cap, independent of `approvalGateCap`). Load-time validation (rule_id uniqueness, tier mismatch, annotation floor, 64 KB cap, disable-list hard-deny rejection, `approvalGateCap` bounds check `1 ≤ N ≤ 500`). `PolicyEngine.__init__` accepts `approval_gate_cap` sourced from blueprint (default 50). | | 5 | agent | `policies/hard_deny.cedar` (new) | Migrate current hard-deny rules + add DROP TABLE. Annotations. | @@ -2189,7 +2096,7 @@ See §17.18 for the off-hours escalation future-work primitive, and §13.14 for - Cancel during AWAITING_APPROVAL (agent-side resume race) - Cancel during approve Lambda (cross-table transaction catches it — finding #7) - Cancel during deny with queued denial injection (between-turns hook pre-empted; `permissionDecisionReason` still delivered — finding #2) - - Late approval after TIMED_OUT (expect 409) + - Late approval after TIMED_OUT (expect 404 `REQUEST_NOT_FOUND`) - **VM-throttle + late-approve race (IMPL-24, §13.12)**: user's APPROVE lands in DDB before `_best_effort_update_status("TIMED_OUT")` can claim the row; agent must re-read with ConsistentRead and honor APPROVED (not return stale TIMED_OUT). Includes the DENIED variant and the "still PENDING" fall-through where neither side wins. - **Chaos tests**: - Container restart mid-approval (simulated via kill + reconciler) @@ -2219,11 +2126,13 @@ Built-in policies shipped with the agent: **Hard-deny (absolute, cannot be disabled by blueprint)**: `rm_slash`, `write_git_internals`, `write_git_internals_nested`, `drop_table`. Absolute; no scope bypasses them; blueprint `disable:` cannot remove them (§5.1, finding #9). **Soft-deny starter set (require approval by default, may be disabled by blueprint)**: -- `force_push_any` — `like "*git push --force*"` — medium, 300s -- `push_to_protected_branch` — pushes to `main`/`master`/`prod`/`release/*` (non-force) — medium, 300s -- `force_push_main` — force-push specifically to `main`/`prod` — high, 600s -- `write_env_files` — `like "*.env"` — high, 600s -- `write_credentials` — `like "*credentials*"` — high, 300s +- `force_push_any` — `like "*git push --force*"` — medium +- `push_to_protected_branch` — pushes to `main`/`master`/`prod`/`release/*` (non-force) — medium +- `force_push_main` — force-push specifically to `main`/`prod` — high +- `write_env_files` — `like "*.env"` — high +- `write_credentials` — `like "*credentials*"` — high + +These built-in rules inherit the task deadline: zero/no deadline by default. Custom rule annotations may impose a positive deadline. Users who want fully autonomous execution (no approval gates) pass `--pre-approve all_session --yes` at submit. Repos that want additional gates add them via `Blueprint.security.cedarPolicies.soft`. Repos that want a different policy set can override specific built-in **soft-deny** rules by `@rule_id` via the blueprint's `security.cedarPolicies.disable` list. The `disable:` mechanism is restricted: it may NOT include any built-in hard-deny rule_id, and the blueprint loader rejects such configurations at task start. @@ -2258,19 +2167,22 @@ Rollout steps: ### 15.5 Backward compatibility -- Existing tasks without `initial_approvals` → empty list → no pre-approvals, default `approval_timeout_s = 300` +- Tasks without `initial_approvals` receive an empty list and no pre-approvals. + New tasks default to `approval_timeout_s = 0`. An already-persisted approval + retains its original timeout; an upgrade or replacement does not extend it. - Existing policies without `@rule_id` / `@tier` → engine fails to start (fail-closed). Blueprint authors must add annotations explicitly during migration. - `PolicyDecision.allowed` property provides backward compat for existing `if not decision.allowed` callers - Hook return shape unchanged — Phase 1a/1b tests continue to pass ### 15.6 Shared Cedar parsing — cross-engine parity contract -The agent runtime uses Python [`cedarpy@4.8.0`](https://pypi.org/project/cedarpy/); the Lambda side (`CreateTaskFn`, `ApproveTaskFn`, `DenyTaskFn`, `GetPoliciesFn`) uses [`@cedar-policy/cedar-wasm@4.10.0`](https://www.npmjs.com/package/@cedar-policy/cedar-wasm) — AWS's official WASM-compiled Cedar engine. Same Rust core, two bindings. Because these engines evolve independently, we ship a **parity contract** (decision #23, finding #1) to catch drift before deploy. - -**Version pinning.** Both engines are pinned exactly (not `^` or `~`) in the monorepo's canonical manifest files. The two bindings are deliberately on **different version lines** — they are NOT required to be equal. `cedarpy` and `cedar-wasm` follow independent release cadences over the shared Cedar Rust core, and the currently-shipped pins (`cedarpy==4.8.0` ↔ `@cedar-policy/cedar-wasm==4.10.0`) are an intentional, tested-compatible skew: the parity fixtures in `contracts/cedar-parity/` are what certify that this specific pair produces identical `(decision, matching_rule_ids)` on every fixture. The rule is "move together and re-verify parity when you bump either side," not "keep the version strings equal." -- `agent/pyproject.toml`: `cedarpy==4.8.0` -- `cdk/package.json`: `"@cedar-policy/cedar-wasm": "4.10.0"` -- `mise.toml` documents the pinned versions in a comment for operator visibility +The agent uses Python `cedarpy`; the policy Lambdas use +`@cedar-policy/cedar-wasm`. Exact current pins live in +[agent/pyproject.toml](../../agent/pyproject.toml) and +[cdk/package.json](../../cdk/package.json). Binding version strings need not be +identical. Upgrade them as a tested pair and run the shared +[parity fixtures](../../contracts/cedar-parity/README.md), which check decisions +and matching rule IDs across both engines. **Lambda layer packaging** (finding #5). The cedar-wasm package is 4.1 MB unzipped. Shipping it in the deployment bundle of each of the 4 policy Lambdas would consume ~16 MB of unzipped bundle size — manageable on its own but leaves little room for AWS SDK + other deps as the codebase grows, and threatens the Lambda 250 MB unzipped limit under realistic growth. Solution: package cedar-wasm as a **Lambda layer** (`cedar-wasm-layer.ts`, task #10 in §15.2), attached to each policy Lambda. This reduces each Lambda's deployment bundle to just the handler code + thin wrapper around the layer import. Policy Lambdas are configured with ≥ 512 MB memory to accommodate WASM module instantiation under concurrent invocation (measured under 100-concurrent bursts in §15.3 Lambda memory tests). @@ -2317,7 +2229,7 @@ flowchart LR When policy authors upgrade either engine, the parity fixture must be re-generated (a small helper script dumps decisions from both engines; the human confirms the change is intentional). -**Scenario (finding #1, illustrative):** This example uses hypothetical versions (e.g. cedarpy `4.10.1` → `4.11.0`) to show the *class* of bug the parity contract catches; it does not describe the real shipped pins (which are the intentional `cedarpy==4.8.0` ↔ `cedar-wasm==4.10.0` skew documented above). The point is that closeness of version strings — even within the same minor line — is no guarantee of behavioral parity, which is exactly why the golden fixtures, not the version numbers, are the source of truth. A platform engineer runs `mise run deps:update` which bumps cedarpy from 4.10.1 to 4.11.0. They notice cedar-wasm is still 4.10.0 but assume it's fine because both say "4.x". Between these versions, cedarpy added support for a new `context has` operator that cedar-wasm doesn't yet have. A new blueprint soft-deny rule uses `context has "approved_context"`. On deploy: +**Scenario (finding #1, illustrative):** This example uses hypothetical versions (e.g. cedarpy `4.10.1` → `4.11.0`) to show the *class* of bug the parity contract catches; it does not describe the real shipped pins (read the package manifests for the actual pins). The point is that closeness of version strings — even within the same minor line — is no guarantee of behavioral parity, which is exactly why the golden fixtures, not the version numbers, are the source of truth. A platform engineer runs `mise run deps:update` which bumps cedarpy from 4.10.1 to 4.11.0. They notice cedar-wasm is still 4.10.0 but assume it's fine because both say "4.x". Between these versions, cedarpy added support for a new `context has` operator that cedar-wasm doesn't yet have. A new blueprint soft-deny rule uses `context has "approved_context"`. On deploy: - Agent-side `PolicyEngine.__init__` parses the rule successfully; engine loads normally. - `CreateTaskFn` on the Lambda side calls cedar-wasm `policyToJson()` — it throws: `ParseError: unknown operator 'has' at line 3`. - User submits a task against that repo. `CreateTaskFn` crashes mid-validation. Error message: "500 Internal Server Error" (because the Lambda didn't handle the upstream parse error gracefully). @@ -2412,7 +2324,7 @@ Items from the design reviews not captured above as design changes — to be add **IMPL-9** (functional P1-3): Runtime allowlist revocation. Not shipped in v1. Placeholder: `bgagent revoke-approval ` noted in §17. -**IMPL-10** (functional P1-12): `approval_timeout_s` default 300 documented consistently in §3 #6, §7.3 table, §10.2 attribute description. +**IMPL-10** (historical functional P1-12): the original 300-second default was superseded by the September retained-request default of `0` (no decision deadline). **IMPL-11** (functional P2-8): CLI `run.ts` command exists from Phase 1b. `submit.ts` also exists. `--pre-approve` / `--approval-timeout` flags added to both. @@ -2602,7 +2514,7 @@ See §15.2. Net new files: ~15. Net modified files: ~15. Total LOC estimate: ~40 - [ ] Backward compat: Phase 1a/1b tests pass without modification - [ ] ULID length references are 26 chars throughout CLI + docs - [ ] **Re-read approval row on TIMED_OUT ConditionCheckFailed (IMPL-24)**: `_best_effort_update_status("TIMED_OUT")` failure path re-reads with ConsistentRead and honors APPROVED/DENIED if the user's decision beat the agent's timer; emits `approval_late_win` milestone. See §6.5 pseudocode, §13.12 VM-throttle race, §14.6 trace, §15.2 task #43. -- [ ] **Default `--approval-timeout` is 300s** documented consistently in decision #6, §5.2, §7.3 field table, §8.2 CLI flags, and §10.2 TaskTable schema. +- [ ] **Default `--approval-timeout` is 0 (no deadline)** documented consistently in decision #6, §5.2, §7.3 field table, §8.2 CLI flags, and §10.2 TaskTable schema. - [ ] **Sub-120s `@approval_timeout_s` emits WARN (IMPL-25)** at blueprint load; sub-30s still rejected. `bgagent lint-policies` (§17.14) surfaces the same WARN pre-submit. - [ ] **User-visible timeout milestones (IMPL-26)**: `approval_timeout_capped` (per-gate, on SSE stream), `approval_timeout_capped_at_submit` (on `POST /v1/tasks` response), `approval_ceiling_shrinking` (once per task at lifetime threshold). All carry `{requested_timeout_s, effective_timeout_s, reason}`. - [ ] **Runtime JWT ceiling (IMPL-27)**: no separate JWT expiry term required in v1 — container uses auto-refreshed IAM credentials (verified by grep of `agent/src/`). Ceiling stays `min(1h, maxLifetime_remaining - cleanup_margin)`. Review if container auth shape changes (see §13.13). diff --git a/docs/design/COMPUTE.md b/docs/design/COMPUTE.md index 48670a6df..cba02c388 100644 --- a/docs/design/COMPUTE.md +++ b/docs/design/COMPUTE.md @@ -18,7 +18,7 @@ The default runtime is **Amazon Bedrock AgentCore Runtime**, which runs each ses | **Startup** | Service-managed | Slim images help | Snapshot resume | Warm ASGs + pre-pull | Karpenter + pre-pull | Backend-dependent | Provisioned concurrency | Snapshot pools (DIY) | | **GPU** | No | No | No | Yes | Yes | Yes (EC2/EKS backend) | No | Yes (with passthrough) | | **Ops burden** | Low (managed) | Low | Low (managed) | Medium | High | Low-Medium | Low | **Very high** | -| **Cost model** | vCPU-hrs + GB-hrs | vCPU + mem/sec | Baseline-priced (8 GiB / 4 vCPU) with 4× vertical burst (32 GiB / 16 vCPU peak); suspended time is storage-only | EC2 + EBS | EKS control + EC2 | Underlying compute | Request + duration | EC2 metal + your ops | +| **Cost model** | vCPU-hrs + GB-hrs | vCPU + mem/sec | Baseline compute (8 GiB / 4 vCPU) plus additional burst usage (up to 32 GiB / 16 vCPU); no suspended compute charge, but snapshot storage and read/write charges remain | EC2 + EBS | EKS control + EC2 | Underlying compute | Request + duration | EC2 metal + your ops | | **Fit** | **Default choice** | Repos > 2 GB image | Suspend/resume economics; approval-wait-heavy workloads; default-sized repos. Heavy sustained-memory builds stay on ECS | GPU, heavy toolchains | Max flexibility | Queued batch jobs | **Poor** (15 min cap) | Best potential, highest cost | > **Lambda MicroVMs are not Lambda functions.** They are a different compute primitive, so the functions column's 15-minute cap and poor-fit verdict do not apply. See [ADR-021](../decisions/ADR-021-lambda-microvms-compute-backend.md). @@ -77,11 +77,32 @@ See [ORCHESTRATOR.md](./ORCHESTRATOR.md) for how the orchestrator handles these ## Lambda MicroVMs backend -Lambda MicroVMs are an opt-in third backend, selected per repository with `compute_type: lambda-microvm`; AgentCore remains the default. Image configuration has three states: a managed base-image ARN and version creates the snapshot image in CDK; an external image identifier uses a snapshot built out of band; and supplying neither provisions only the roles, buckets, and connectors needed for the bootstrap deploy. `cdk/scripts/package-microvm-artifact.sh` packages the agent as zip + Dockerfile, uploads it to the artifact bucket, and can create the external image. Lambda MicroVMs are available in five launch regions (us-east-1, us-east-2, us-west-2, eu-west-1, ap-northeast-1) and will expand; the platform enforces regional availability in layers via a synth-time constant, onboarding live probes, and orchestration-time classification. +Lambda MicroVMs are an opt-in third backend, selected per repository with `compute_type: lambda-microvm`; AgentCore remains the default. Image configuration has three states: a managed base-image ARN, version and artifact digest create the snapshot image in CDK; an external image identifier uses a snapshot built out of band; and supplying neither image provisions only the roles, buckets, and connectors needed for the bootstrap deploy. Lambda MicroVMs are available in five launch regions (us-east-1, us-east-2, us-west-2, eu-west-1, ap-northeast-1) and will expand; the platform enforces regional availability in layers via a synth-time constant, onboarding live probes, and orchestration-time classification. -Because a snapshot freezes its build-time environment, deployment-specific, non-secret identifiers travel in the `/run` hook's `platform_config` block instead. The strategy sends the canonical inline envelope or, when that envelope exceeds the verified 4,096-byte `runHookPayload` limit, an S3-pointer envelope with the configuration also merged into the uploaded payload. The agent accepts only allowlisted keys and installs them before pipeline initialization; [ADR-021 §3](../decisions/ADR-021-lambda-microvms-compute-backend.md#3-packaging-same-agent-image-source-new-build-path) defines the exact wire shapes and validation rules. +New MicroVM installations should select `microvm_nested_stack=true` and require bootstrap bundle 1.9.0. +The layout setting is temporarily required; omission fails synthesis until migration is verified. +Before upgrading an existing flat installation, set `microvm_nested_stack=false` +and retain it until completing the [resource migration](../verification/645-p3-nested-stack.md). -Networking separates image build from execution: the build-only connector permits TCP 80 and 443 because the Dockerfile uses `apt-get`, while running MicroVMs retain 443-only egress through the platform VPC. Every launch explicitly passes the Lambda-managed `NO_INGRESS` connector; omission would select the service's public-ingress default. The P2 image declares and serves `/ready` and `/validate` at build time and `/run` and `/terminate` at runtime. `/suspend` and `/resume` remain disabled until their P3 implementation. +For managed images, run `cdk/scripts/package-microvm-artifact.sh --stack-name ` after each agent change. It packages the Dockerfile's local inputs into a deterministic ZIP and uploads it under `microvm-images/agent-artifact-.zip`. The checksum is a fingerprint of the uploaded bytes: identical inputs reuse the same verified object, and changed inputs produce a new filename. Deploy with the printed `--context microvm_artifact_sha256=` alongside `microvm_base_image_arn` and `microvm_base_image_version`, and retain these inputs for later deployments. The changed S3 URI tells CloudFormation to update the existing image; overwriting the old fixed filename alone does not. Missing or malformed digests fail synthesis. An initial deployment without an image must create the bucket first. The script's explicit `--create-image` alternative retains the fixed base key and calls the image API directly. + +Because a snapshot freezes its build-time environment, current deployment identifiers arrive through the v2 payload bootstrap. The coordinator publishes a non-secret deployment manifest and sends a single-object signed download URL for the task. The worker reads only its deployment's `bootstrap/*` with ambient credentials; other object reads and payload-bucket listing are explicitly denied. The downloaded task identity and configuration must match the authenticated manifest before configuration installation. The serialized reference fits the verified 4,096-byte hook limit; all payload sizes use S3. ECS shares this transport through `AGENT_PAYLOAD_REF`. [ADR-021 §3](../decisions/ADR-021-lambda-microvms-compute-backend.md#3-packaging-same-agent-image-source-new-build-path) defines the wire format and compatibility requirements; [live payload checks](../verification/README.md) record validation. + +Networking separates image build from execution: the build-only connector permits TCP 80 and 443 because the Dockerfile uses `apt-get`, while running MicroVMs retain 443-only egress through the platform VPC. Every launch explicitly passes the Lambda-managed `NO_INGRESS` connector; omission would select the service's public-ingress default. Images declare and serve `/ready` and `/validate` at build time and `/run`, `/terminate`, `/suspend` and `/resume` at runtime. Automatic suspension is a separate deployment opt-in. Registry HTTP/SSE tools therefore need reachable HTTPS/443 endpoints; remote non-443 tools are unsupported under the default policies of AgentCore and ECS as well. Local `stdio` tools can run, with their outbound traffic subject to the same restriction. Asset resolution/loading does not probe connectivity; see [registry network support](./REGISTRY.md#2-asset-kinds-for-mvp). + +P3 connects the guest checkpoint/credential hooks, actual image-version capability, durable supervisor and post-commit approval wake. The `microvm_approval_suspend_enabled` deployment context defaults false and controls both a static opt-in and a live Parameter Store switch. Existing durable executions reread the live switch before new suspension because their original Lambda environment is pinned. Recovery and the original service lifetime survive supervisor replay; the API preserves accepted decisions when optional wake fails. AgentCore/ECS retain explicit unsupported pause/wake results. Explicit approval deadlines use the original UTC/monotonic deadline. See the [lifecycle diagnostics](../verification/645-p3-lifecycle-diagnostics.md) for failure investigation, and the [acceptance status](../verification/README.md) for deployment evidence. + +Tasks can set `microvm_sleep_after_s` (CLI: `--microvm-sleep-after `). +The default is 600 seconds of waiting for each approval; zero disables sleep. +Creation persists the resolved preference, while legacy rows use the default. +The global suspension switch still takes precedence. Unanswered approvals have +no deadline by default (`approval_timeout_s=0`); an explicit finite deadline +remains available. A short timed request stays awake when there is too little +useful sleep time before its wake margin. Neither suspension nor replacement +extends that original deadline. For a longer wait, a verified conversation and +workspace checkpoint lets ABCA retire the worker, release capacity and start a +replacement after the answer. Snapshot storage and save/restore fees mean the +sleep delay is a user preference, not a guarantee of savings for every pause. ## ECS Fargate task sizing (build vs. planning) @@ -104,7 +125,7 @@ The agent harness is the layer around the LLM that manages the execution loop: c The platform uses the [Claude Agent SDK](https://github.com/anthropics/claude-agent-sdk-python) as the harness. It provides the agent loop, built-in tools (filesystem, shell), and streaming message reception for per-turn trajectory capture (token usage, cost, tool calls). -**Execution model:** Tasks are fully unattended and one-shot. The agent loop runs in a background thread so the FastAPI `/ping` endpoint stays responsive on the main thread. The agent thread uses `asyncio.run()` with the stdlib event loop (uvicorn is configured with `--loop asyncio` to avoid uvloop conflicts with subprocess SIGCHLD handling). +**Execution model:** The agent runs automatically until it completes, needs a human decision or is stopped. The HTTP entrypoint runs its loop in a background thread so FastAPI `/ping` stays responsive; ECS uses the batch entrypoint. The agent thread uses `asyncio.run()` with the stdlib event loop (uvicorn is configured with `--loop asyncio` to avoid uvloop conflicts with subprocess SIGCHLD handling). **System prompt:** Selected by workflow from a shared base template (`agent/src/prompts/base.py`) with per-workflow sections (`coding/new-task-v1`, `coding/pr-iteration-v1`, `coding/pr-review-v1`). The platform defines what the agent should do; the harness executes it. @@ -119,7 +140,7 @@ The platform uses the [Claude Agent SDK](https://github.com/anthropics/claude-ag | GitHub | AgentCore Gateway + Identity | Clone, push, PR, issues | | Web search | AgentCore Gateway | Documentation lookups | -Plugins, skills, and MCP servers are out of scope for MVP. Additional tools can be added via Gateway integration. +Agents can use configured skills and MCP tools through the [registry](./REGISTRY.md). Their runtime network access remains subject to the compute backend’s egress rules. ### Policy enforcement diff --git a/docs/design/DEPLOYMENT_ROLES.md b/docs/design/DEPLOYMENT_ROLES.md index 9f10f1967..a3b4c464d 100644 --- a/docs/design/DEPLOYMENT_ROLES.md +++ b/docs/design/DEPLOYMENT_ROLES.md @@ -30,7 +30,7 @@ The policies are split into six IAM managed policies (each under the 6,144-chara > **Placeholder substitution**: Replace `ACCOUNT_ID` with your 12-digit AWS account ID and `REGION` with your deployment region (e.g., `us-east-1`) throughout this document. -These policies are not created or attached manually. The repository generates them — and a custom bootstrap template that wires all six into the CloudFormation execution role — from the TypeScript sources, then bootstraps with that template: +These policies are not created or attached manually. The repository generates them and a custom bootstrap template that attaches the selected policies to the CloudFormation execution role: ```bash # Regenerate artifacts (policies JSON + template YAML) and bootstrap. @@ -48,14 +48,29 @@ aws cloudformation update-stack --stack-name CDKToolkit --use-previous-template aws cloudformation describe-stacks --stack-name CDKToolkit --query 'Stacks[0].Parameters' ``` -Under the hood, `mise //cdk:bootstrap` runs `npx cdk bootstrap --template bootstrap/bootstrap-template.yaml` (see `cdk/mise.toml`). The generated template defines six inline `AWS::IAM::ManagedPolicy` resources that **replace** the default `AdministratorAccess` on the CloudFormation execution role; the `IaCRole-ABCA-Compute-ECS` and `IaCRole-ABCA-Compute-LambdaMicrovms` policies are conditional on the `ComputeTypes` parameter including their respective backend. The policy sources are `cdk/src/bootstrap/policies/{infrastructure,application,observability,compute-agentcore,compute-ecs,compute-lambda-microvm}.ts`, compiled to `cdk/bootstrap/policies/*.json` by `cdk/scripts/generate-bootstrap-artifacts.ts`. +Under the hood, `mise //cdk:bootstrap` runs `npx cdk bootstrap --template bootstrap/bootstrap-template.yaml` (see `cdk/mise.toml`). The generated template defines six `AWS::IAM::ManagedPolicy` resources that **replace** the default `AdministratorAccess` on the CloudFormation execution role; the `IaCRole-ABCA-Compute-ECS` and `IaCRole-ABCA-Compute-LambdaMicrovms` policies are conditional on the `ComputeTypes` parameter including their respective backend. The policy sources are `cdk/src/bootstrap/policies/{infrastructure,application,observability,compute-agentcore,compute-ecs,compute-lambda-microvm}.ts`, compiled to `cdk/bootstrap/policies/*.json` by `cdk/scripts/generate-bootstrap-artifacts.ts`. + +**Re-bootstrap to bundle 1.7.0 or later before deploying nested stacks.** The +execution role also needs the generated inline policy +`PassExecutionRoleToCloudFormation`, from +`cdk/src/bootstrap/nested-stack-policy.ts`. It grants `iam:PassRole` on that exact +execution role, with `iam:PassedToService=cloudformation.amazonaws.com`. The ARN +uses the bootstrap partition, account, Region and qualifier; it grants no access +to pass other roles. Without it, a fresh 1.6.0 deployment fails change-set +validation when CloudFormation tries to pass its role to the registry's nested +stacks. Preserve all existing bootstrap parameters when updating the template. + +Bundle 1.7.0 also corrects `BootstrapPolicyHash`: it includes nested policy fields +and the generated inline policy. Earlier hashes could remain unchanged after +Action, Resource or Condition changes. A matching old hash is insufficient proof +that deployed permissions match source. > **CloudFormation inline-template limit — 51,200 characters**: This is a second, independent size ceiling, distinct from the per-policy IAM 6,144-character limit above. `cdk bootstrap --template` sends the template inline as `TemplateBody`; above 51,200 characters the CLI has to stage it in S3 instead, which it **cannot** do while bootstrapping a fresh account, because that bucket is one of the resources bootstrap creates. The result is a hard `BootstrapStackRequired` failure with no way through `cdk bootstrap`, `--force` included ([#864](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/864)). > > Two consequences for anyone editing the policies: > > - **The gated size is not the file's size on disk.** The CLI parses the file, discards its formatting, and re-serialises the parsed object before measuring. Reformatting `bootstrap-template.yaml` therefore changes nothing; only the *content* moves the number. Check it with `npx cdk bootstrap --show-template --template bootstrap/bootstrap-template.yaml | wc -c`. -> - **Each `PolicyDocument` is emitted as a minified JSON string**, not a nested YAML mapping. Both are valid for this `Json`-typed property and IAM stores the string parsed, but a string scalar survives the CLI's re-serialisation on one line — which is what keeps the body under the ceiling (45,743 characters, versus 53,369 as mappings). +> - **Each managed-policy `PolicyDocument` is emitted as a minified JSON string**, not a nested YAML mapping. Both are valid for this `Json`-typed property and IAM stores the string parsed, but a string scalar survives the CLI's re-serialisation on one line. The original fix reduced the body to 45,743 characters from 53,369; subsequent policy additions remain subject to the budget. The small inline self-role policy uses a mapping to resolve its ARN with `Fn::Sub`. > > `cdk/scripts/generate-bootstrap-template.ts` fails the build when the body exceeds the budget in `cdk/src/bootstrap/template-size.ts`, so adding statements surfaces the problem at generation time rather than against somebody's fresh account. Both the guard and its regression test obtain the size by invoking `cdk bootstrap --show-template` on the committed artifact — the CLI is the component that makes the inline-vs-S3 decision, so asking it directly cannot drift the way a local copy of its serialiser would. No AWS credentials are required. @@ -92,7 +107,7 @@ Under the hood, `mise //cdk:bootstrap` runs `npx cdk bootstrap --template bootst For deploying the `backgroundagent-dev` stack. This single stack contains all platform resources including the AgentCore runtime, ECS compute (when enabled), API Gateway, Cognito, DynamoDB tables, VPC, DNS Firewall, and observability infrastructure. -> **IAM managed policy size limit**: A single managed policy cannot exceed 6,144 characters. The permissions below are split into six policies to stay under this limit (three always-applied, plus three compute-variant policies). They are wired into the CloudFormation execution role by the generated bootstrap template; see [Using these policies](#using-these-policies). +> **IAM managed policy size limit**: A single managed policy cannot exceed 6,144 characters. The permissions below are split into six policies to stay under this limit (four always applied, including AgentCore, plus optional ECS and MicroVM policies). The separate inline self-role grant is described above. They are wired into the CloudFormation execution role by the generated bootstrap template; see [Using these policies](#using-these-policies). ### IaCRole-ABCA-Infrastructure @@ -286,6 +301,12 @@ CloudFormation stack operations, IAM roles/policies, VPC networking, and Route 5 DynamoDB tables, Lambda functions, API Gateway, Cognito, WAFv2, EventBridge, SQS, CloudFront, and Secrets Manager. When ECS Fargate compute is enabled, add the ECS statement below to this policy. +Agent Registry provisioning also uses a Step Functions workflow to wait for +asynchronous creation and deletion. Its construct explicitly names the workflow +with the parent stack's `backgroundagent-dev-` prefix so it fits this policy, +including when deployed in a nested stack. Adopting this name in an existing +deployment replaces the provider's waiter state machine. + ```json { "Version": "2012-10-17", @@ -823,6 +844,31 @@ The second statement, `MicrovmPassRoles`, is the one exception to the rule that > **Operators must re-bootstrap for this.** The statement ships in bootstrap policy bundle **1.6.0**; a CDKToolkit stack bootstrapped at 1.5.0 or earlier will fail the CDK-managed MicroVM image deploy with a caller-side `iam:PassRole` AccessDenied on the build role. Check `CDKToolkit`'s `BootstrapPolicyVersion` output, and re-run `mise //cdk:bootstrap` (with `ComputeTypes` including `lambda-microvm`) if it is behind. +P3 additionally requires **bundle 1.8.0** for `MicrovmSuspendConfiguration`. + +The nested MicroVM layout requires **bundle 1.9.0**. Its child stack uses the +explicit parent-derived names `backgroundagent-dev-MicrovmBuildRole` and +`backgroundagent-dev-MicrovmConnectorRole`; `MicrovmPassRoles` admits those two +exact names in addition to the legacy flat-layout prefixes. The execution role +stays in the parent and is still excluded. Re-bootstrap before deploying the +child stack. Before upgrading an existing flat deployment, set and retain +`microvm_nested_stack=false` until its resource migration is complete; changing ownership is not an ordinary +in-place update. See the [nested-stack runbook](../verification/645-p3-nested-stack.md). + +For a reviewed migration that keeps old and new resources side by side, +`microvm_resource_name_prefix` gives the nested image, network connectors and log +group distinct names. It requires nested mode and a concrete 1–40 character +letter/digit/hyphen prefix. Build/operator IAM role names remain derived from the +parent deployment so the bootstrap's existing `PassRole` scope still applies. +Keep the selected prefix stable in later deployments. This option alone does +not preserve the old resources or their runtime permissions; those remain part +of the migration procedure. +This statement lets CloudFormation manage and tag the live suspension setting +at `//microvm-approval-suspend-enabled`. The +coordinator gets only `GetParameter` on its exact parameter. Existing durable +executions retain their Lambda version and reread this setting before new +suspension, so disable can reach executions already running. + ```json { "Statement": [ @@ -857,9 +903,24 @@ The second statement, `MicrovmPassRoles`, is the one exception to the rule that "Effect": "Allow", "Resource": [ "arn:aws:iam::*:role/backgroundagent-dev-LambdaMicrovmComputeBuild*", - "arn:aws:iam::*:role/backgroundagent-dev-LambdaMicrovmComputeConnector*" + "arn:aws:iam::*:role/backgroundagent-dev-LambdaMicrovmComputeConnector*", + "arn:aws:iam::*:role/backgroundagent-dev-MicrovmBuildRole", + "arn:aws:iam::*:role/backgroundagent-dev-MicrovmConnectorRole" ], "Sid": "MicrovmPassRoles" + }, + { + "Action": [ + "ssm:GetParameters", + "ssm:PutParameter", + "ssm:DeleteParameter", + "ssm:AddTagsToResource", + "ssm:RemoveTagsFromResource", + "ssm:ListTagsForResource" + ], + "Effect": "Allow", + "Resource": "arn:aws:ssm:*:*:parameter/backgroundagent-*/microvm-approval-suspend-enabled", + "Sid": "MicrovmSuspendConfiguration" } ], "Version": "2012-10-17" diff --git a/docs/design/INTERACTIVE_AGENTS.md b/docs/design/INTERACTIVE_AGENTS.md index c17955d97..fa6d98678 100644 --- a/docs/design/INTERACTIVE_AGENTS.md +++ b/docs/design/INTERACTIVE_AGENTS.md @@ -1,6 +1,8 @@ # Interactive Agents: Async Interaction Design -> **Status:** Active design +> **Status:** Historical interaction design with current approval-path corrections (2026-09-22). +> Proposed dispatcher services and Slack buttons below are not all implemented; use the +> [API contract](./API_CONTRACT.md) and [approval guide](../guides/USER_GUIDE.md#approval-gates-cedar-hitl) for supported interfaces. > **Branch:** `feature/interactive-background-agents` > **Last updated:** 2026-04-29 (rev 6) @@ -19,7 +21,7 @@ This document describes the interactivity surfaces layered on top of that model 3. **Watch** — `bgagent watch ` polls `TaskEventsTable` with an adaptive interval (500 ms when events are arriving, back-off to 5 s when idle). Same endpoint used under the hood for foreground-block UX on `ask` and for HITL approval waits. 4. **Nudge** — `bgagent nudge ""` writes a row into `TaskNudgesTable`. The agent reads pending nudges between turns, acknowledges with a `nudge_acknowledged` milestone event, and integrates the nudge on its next turn. 5. **Ask** — `bgagent ask ""` (Phase 2) writes a question row. The agent answers at the next between-turns boundary; the answer surfaces as a `status_response` event. CLI default is foreground block-and-poll with a spinner; task and answer are both durable if the CLI disconnects. -6. **Approval gates** — Phase 3 Cedar-driven hard gates. Agent emits `approval_requested`, waits for a decision from `bgagent approve` / `bgagent deny` or a Slack button-press. Detailed design in [`CEDAR_HITL_GATES.md`](./CEDAR_HITL_GATES.md). +6. **Approval gates** — Phase 3 Cedar-driven hard gates. Agent emits `approval_requested`, waits for a decision from `bgagent approve` / `bgagent deny` or an owner-authored Linear thread reply. Slack notifications provide CLI instructions; buttons remain proposed. Detailed design in [`CEDAR_HITL_GATES.md`](./CEDAR_HITL_GATES.md). ### Core architectural choices @@ -27,7 +29,7 @@ This document describes the interactivity surfaces layered on top of that model - **Durable event table (`TaskEventsTable`)** is the one source of truth for agent progress. Every reader — CLI, Slack/GitHub/email dispatchers, status Lambda — reads from this table, never from the live agent. - **Polling-only CLI.** No SSE, no WebSockets. DDB eventually-consistent reads with an `event_id` cursor are cheap, reliable, and compute-agnostic. - **Notification plane as first-class.** A FanOutConsumer Lambda subscribes to `TaskEventsTable` DDB Streams and routes per-event-type to per-channel dispatcher Lambdas (Slack, email, GitHub comment). Per-channel defaults ship in v1. -- **Agent interaction via the hook mechanism the Claude Agent SDK provides.** Nudges, asks, and approvals all use `Stop` / between-turns hooks; no mechanism outside the SDK's contract is required. +- **Agent interaction via the hook mechanism the Claude Agent SDK provides.** Nudges and asks use `Stop` / between-turns hooks; approval gates pause in `PreToolUse`, with denial steering delivered at a later Stop hook; no mechanism outside the SDK's contract is required. --- @@ -111,7 +113,7 @@ This document describes the interactivity surfaces layered on top of that model │ DELETE /tasks/{id} cancel │ │ POST /tasks/{id}/nudge nudge │ │ POST /tasks/{id}/asks ask (P2) │ - │ POST /tasks/{id}/approvals approve P3 │ + │ POST /tasks/{id}/approve approve │ │ POST /webhooks/tasks GH webhook │ └───────────┬──────────────────────────────────┘ │ @@ -223,17 +225,16 @@ Consumer: agent between-turns hook reads pending nudges, emits `nudge_acknowledg ### 3.7 TaskApprovalsTable (Phase 3) Phase 3 approval-request spine. Detailed schema in [`CEDAR_HITL_GATES.md`](./CEDAR_HITL_GATES.md). Semantics summary: -- Agent writes an approval row with the request context. -- Agent transitions `RUNNING → AWAITING_APPROVAL` and enters a poll loop. -- User responds via REST (`POST /tasks/{id}/approvals/{request_id}`) or via a Slack button dispatched by the notification plane. +- The worker calls the trusted approval service, which atomically creates the pending row and transitions the task `RUNNING → AWAITING_APPROVAL`. The worker then polls for a decision. +- The owner responds through `POST /tasks/{id}/approve` or `/deny` with `request_id` in the body, the equivalent CLI commands, or an `approve`/`deny` reply to the Linear approval comment. - On decision, agent transitions back to `RUNNING`; denial reasons are injected as Stop-hook steering on the next turn. ### 3.8 FanOutConsumer (router) Lambda subscribed to `TaskEventsTable` DDB Streams (relying on the DynamoDB Streams **default** `ParallelizationFactor` of 1, which preserves per-`task_id` ordering by shard — not set explicitly in `fanout-consumer.ts`; see §6.1). Reads per-task notification config (from `TaskTable` metadata or `RepoTable` defaults), filters events by channel subscription, and invokes per-channel dispatcher Lambdas. -- **SlackDispatchFn** — posts to configured channel / DM. Includes action buttons for `approval_required` events. -- **EmailDispatchFn** — SES. +- **SlackDispatchFn** — posts to configured channel / DM. Approval notifications currently contain CLI response instructions; buttons remain proposed. +- **EmailDispatchFn** — proposed SES delivery; current email dispatch is a log-only stub. - **GitHubDispatchFn** — edits a single GitHub issue comment in place via `PATCH /repos/{o}/{r}/issues/comments/{id}`. On 404 (comment deleted upstream) falls back to POSTing a fresh comment. Per-task ordering is guaranteed upstream by the DDB Streams default `ParallelizationFactor` of 1 (see §6.1), so no conditional-request header is needed (and GitHub's REST API does not accept `If-Match` on this endpoint — see §6.4). Detailed routing and default filters in §6. @@ -296,8 +297,8 @@ Authentication: Cognito User Pool ID token in `Authorization` header for all RES | `agent_milestone` | Agent code (pipeline, hooks) | Named checkpoint (`repo_cloned`, `pr_opened`, `nudge_acknowledged`, ...) | | `agent_cost_update` | Runner | Cumulative token + dollar cost | | `agent_error` | Runner | Handled exception | -| `approval_required` (P3) | PreToolUse Cedar hook | Cedar policy requires user decision | -| `approval_decided` (P3) | Approve/Deny Lambda | User responded | +| `approval_requested` (P3) | PreToolUse Cedar hook | Cedar policy requires user decision | +| `approval_decision_recorded` (P3) | Approve/Deny Lambda | User responded | | `status_response` (P2) | Between-turns hook | Agent answered an `ask` | | `nudge_acknowledged` | Between-turns hook | Agent saw a nudge before incorporating it | | `pr_created` | Pipeline | PR opened for the task | @@ -320,7 +321,7 @@ Consumers page `TaskEventsTable` using `event_id` as a cursor: `KeyConditionExpr ### 5.1 `bgagent submit` ``` -$ bgagent submit --repo org/repo "fix the auth timeout bug" +$ bgagent submit --repo org/repo --task "fix the auth timeout bug" task submitted: abc123 ``` @@ -411,9 +412,9 @@ Flags: HITL approval commands. All flows are REST + DDB; no streaming. Detailed design in [`CEDAR_HITL_GATES.md`](./CEDAR_HITL_GATES.md). Summary: -- Agent emits `approval_required` with the tool context. -- Notification plane dispatches the event (Slack with action buttons, email, GitHub). -- User responds via `bgagent approve `, `bgagent deny --reason "…"`, or Slack button click. +- Agent emits `approval_requested` with the tool context. +- The notification plane sends approval messages to Slack and Linear. Email is a stub; GitHub does not receive approval messages. +- The owner responds via `bgagent approve `, `bgagent deny --reason "…"`, or an `approve`/`deny` reply to the Linear approval comment. - Agent's poll loop sees the decision and proceeds or deny-steers. ### 5.7 `bgagent cancel` @@ -444,19 +445,22 @@ TaskEventsTable ──DDB Stream──▶ FanOutConsumer - Router reads per-task notification config (channel enablement + event-type filters), then invokes the relevant dispatcher Lambda(s) per event. - Dispatchers are separate Lambdas so a GitHub API outage doesn't block Slack notifications. -### 6.2 Per-channel defaults (v1) +### 6.2 Original proposed per-channel defaults (v1) | Channel | Default subscribed events | Opt-in via `--verbose` | |---|---|---| -| **Slack** | `task_completed`, `task_failed`, `task_cancelled`, `pr_created`, `agent_error`, `approval_required`, `status_response` | adds `agent_milestone` | -| **Email** | `task_completed`, `task_failed`, `approval_required` | — | +| **Slack** | `task_completed`, `task_failed`, `task_cancelled`, `pr_created`, `agent_error`, `approval_requested`, `status_response` | adds `agent_milestone` | +| **Email** | `task_completed`, `task_failed`, `approval_requested` | — | | **GitHub issue comment** | `pr_created`, terminal status (single edit-in-place comment) | — already minimal | Rationale: if Slack pings on every milestone, users mute the bot within days. Default to the minimal set that surfaces decision-requiring events and completion; power users opt into verbose streams. ### 6.3 Slack approval buttons -`approval_required` events delivered to Slack include `Approve` / `Deny` action buttons. On click, Slack invokes an interaction callback Lambda which writes to `TaskApprovalsTable` via the same `POST /approvals` path the CLI uses. This gives the common case (reviewer in Slack, not at a terminal) a one-click response path. +**Proposed, not implemented.** Current Slack approval messages contain CLI +commands. A future button callback would need to map the Slack user to the task +owner and use the authenticated decision path; a valid Slack signature alone +would not authorize approval. Linear thread replies already support owner decisions. ### 6.4 GitHub issue comment — edit-in-place @@ -482,7 +486,7 @@ Submitted with the task (optional) or resolved from repo defaults: { "notifications": { "slack": { "enabled": true, "channel": "#coding-agents", "events": ["default"] }, - "email": { "enabled": true, "events": ["approval_required", "task_failed"] }, + "email": { "enabled": true, "events": ["approval_requested", "task_failed"] }, "github": { "enabled": true, "events": ["default"] } } } diff --git a/docs/design/ORCHESTRATOR.md b/docs/design/ORCHESTRATOR.md index 188eea2bd..0ca3d0ad0 100644 --- a/docs/design/ORCHESTRATOR.md +++ b/docs/design/ORCHESTRATOR.md @@ -19,7 +19,7 @@ The orchestrator sits between the API layer and the agent runtime. Changes to ta ## Responsibilities -The orchestrator is deliberately scoped. It handles coordination and bookkeeping but never touches agent logic, compute infrastructure, or memory storage. This clear boundary means a crashed agent does not leave orphaned state, and platform invariants (concurrency limits, event audit, cancellation) cannot be bypassed by agent code. +The orchestrator handles coordination, compute lifecycle calls and finalization bookkeeping. The agent runs the coding workflow. Recovery still needs explicit guards around external effects, including saved start receipts and task-owned capacity reservations. ### What the orchestrator owns @@ -32,7 +32,7 @@ The orchestrator is deliberately scoped. It handles coordination and bookkeeping | Result inference | Determine success or failure from agent response, DynamoDB record, and GitHub state | | Finalization | Update status, emit events, release concurrency, persist audit records | | Cancellation | Stop the session and drive the task to CANCELLED at any point | -| Concurrency | Track per-user and system-wide running task counts with atomic counters | +| Concurrency | Track per-user capacity with task-owned reservations and an atomic counter | ### What the orchestrator does NOT own @@ -96,7 +96,7 @@ stateDiagram-v2 AWAITING_APPROVAL --> RUNNING : Approved or denied (resume) AWAITING_APPROVAL --> CANCELLED : User cancels mid-approval - AWAITING_APPROVAL --> FAILED : Stranded-approval reconciler + AWAITING_APPROVAL --> FAILED : Infrastructure loss or stranded wait FINALIZING --> COMPLETED : PR or commits found FINALIZING --> FAILED : No useful work @@ -121,11 +121,11 @@ stateDiagram-v2 | `HYDRATING` | `FAILED` | Hydration error | GitHub API failure, guardrail blocks content, Bedrock unavailable | | `RUNNING` | `AWAITING_APPROVAL` | Cedar soft-deny gate fires | Tool call triggers a soft-deny policy rule during execution | | `RUNNING` | `FINALIZING` | Session ends | Response received or session terminated | -| `RUNNING` | `TIMED_OUT` | Max duration exceeded | AgentCore and Lambda MicroVMs have an 8h substrate cap; the orchestrator's own safety-net poll window is `MAX_POLL_ATTEMPTS` (1020) × 30s ≈ 8.5h, after which a still-`RUNNING` task is driven to `TIMED_OUT` | +| `RUNNING` | `TIMED_OUT` | Max duration exceeded | AgentCore and Lambda MicroVMs have an 8h substrate cap; MicroVM supervision retains the original service deadline across replay; other backends retain the 1,020-attempt safety window (about 8.5h at 30s) | | `RUNNING` | `FAILED` | Session crash | Heartbeat or substrate liveness lost (see Liveness monitoring) | | `AWAITING_APPROVAL` | `RUNNING` | Approved or denied | Human decision received; agent resumes | | `AWAITING_APPROVAL` | `CANCELLED` | User cancels | Explicit cancel while awaiting approval | -| `AWAITING_APPROVAL` | `FAILED` | Stranded reconciler | Approval request orphaned (agent died mid-wait) | +| `AWAITING_APPROVAL` | `FAILED` | Infrastructure failure or stranded wait | Lost compute, exhausted supervisor recovery/window, or an orphaned approval; the approval decision is not rewritten | | `FINALIZING` | `COMPLETED` | Success inferred | PR exists or commits on branch | | `FINALIZING` | `FAILED` | Failure inferred | No commits, no PR, or agent reported error | @@ -152,7 +152,7 @@ Multiple timeout mechanisms work together to prevent runaway tasks. Substrate ti | Type | Default | Effect | |---|---|---| -| Max session duration | 8 hours | AgentCore caps a session at 8h; Lambda MicroVMs use `maximumDurationInSeconds: 28,800`, including suspended time. The orchestrator's safety-net poll loop runs up to `MAX_POLL_ATTEMPTS` (1020) × 30s ≈ 8.5h; a task still `RUNNING` when that window is exhausted is driven to `TIMED_OUT`. | +| Max session duration | 8 hours | AgentCore caps a session at 8h; Lambda MicroVMs use `maximumDurationInSeconds: 28,800`, including suspended time. MicroVM uses the saved absolute service deadline, so fast transition polling cannot shorten the session. Other backends retain the 1,020-attempt safety window. An exhausted approval wait uses FAILED, its allowed infrastructure-failure transition. | | Idle timeout | Backend-specific | AgentCore has an idle timeout. Lambda MicroVMs omit `idlePolicy` because inbound-traffic idleness would suspend an outbound-only agent while it is working. See Liveness monitoring. | | Max turns | 100 (range 1-500) | Agent stops after N model invocations. Configurable per task or per repo. | | Max cost budget | $0.01-$100 | Agent stops when budget is reached. Per-task or per-repo via Blueprint. | @@ -178,8 +178,8 @@ The orchestrator (`orchestrate-task.ts`) runs these as distinct durable-executio Validates the task before any compute is consumed. Checks run in order: 1. **Repo onboarding** - `GetItem` on `RepoTable`. If not found or inactive, reject with `REPO_NOT_ONBOARDED`. This runs at the API handler level (`createTaskCore`) for fast rejection. -2. **User concurrency** - Atomic check-and-increment on `UserConcurrency` counter. If at limit (default 10), the task is **queued, not failed** (#441): it transitions `SUBMITTED → QUEUED` and a scheduled admission-queue pickup Lambda re-attempts admission in FIFO order (by `created_at`) as slots free up, flipping `QUEUED → SUBMITTED` and re-invoking the orchestrator. The pickup Lambda does a read-only capacity pre-check; the orchestrator's atomic increment remains the single writer of the counter, so a pickup that loses the race harmlessly re-queues without losing FIFO position. `GET /tasks/{id}` surfaces `queue_position` and `estimated_wait_s` while queued. -3. **System concurrency** - Compare total running + hydrating tasks to the configured system limit and selected-backend quotas. +2. **User concurrency** - One transaction creates an internal `concurrency_slot` reservation on the task and increments `UserConcurrency.active_count`, subject to the configured cap (default 3). A retry reuses a held reservation. At the cap, the task transitions `SUBMITTED → QUEUED` and a scheduled pickup retries in FIFO order by `created_at`. Pickup and upload confirmation only inspect capacity; the orchestrator owns reservation acquisition. A task that acquired a reservation concurrently cannot be put back in the queue. `GET /tasks/{id}` surfaces `queue_position` and `estimated_wait_s` while queued. +3. **Backend capacity** - AWS also enforces the selected backend's service quotas when compute starts. 4. **Rate limiting** - Sliding window counter (10 tasks/hour per user). Rate-limit rejections happen at submit time and are rejected, not queued (unlike the concurrency cap, which queues). 5. **Idempotency** - If the request includes an idempotency key and a task with that key exists, return the existing task. @@ -205,7 +205,13 @@ The orchestrator resolves the repository's `ComputeStrategy` and calls `startSes AgentCore's session ID is pre-generated and reused on retry. ECS and Lambda MicroVMs use their substrate identifiers as session IDs. -If `RunMicrovm` succeeds but persisting the session handle or emitting the start event fails, the start step terminates the MicroVM best-effort using its in-memory handle before propagating the original error. This orphan reap is required because no later poll or finalization step can recover an unpersisted handle. +MicroVM starts first save an internal `microvm_start` receipt on the task: its stable client token (the task ID), a fingerprint of the request and full payload, creation time, and a local replay deadline. This happens before payload upload or `RunMicrovm`. Retries must match the saved fingerprint; a changed request cannot overwrite the earlier task's input. A returned handle is saved in the receipt and the normal task metadata before registration finishes. A replay can recover that handle without another start call. Registration uses strongly consistent reads to observe cancellation and already-committed writes. + +If a handle-save or registration response is lost, the code checks the committed task before terminating a known computer. If registration failed, cleanup remains best-effort and the receipt retains any saved handle for diagnosis. Failure to emit `session_started` alone does not terminate a registered MicroVM. MicroVM start failures persist the task outcome and reach `finalize-before-session`; they do not also release concurrency in the start step. Finalization begins with a strongly consistent read so a recently saved failure or cancellation is not reported using an older active state. + +An unanswered service request may already have created a MicroVM. Recovery reuses the same token and request within a **120-second local window**; after the window, it refuses another `RunMicrovm` call. This window is an application guard, not a verified AWS token-retention promise. An unrecovered request is reported as `MICROVM_START_OUTCOME_UNKNOWN` and requires inspection before submitting another task. Cancellation after an unanswered request records a `microvm_start_outcome_unknown` event. Without an ID, immediate termination cannot be guaranteed; the eight-hour service lifetime bound still applies. + +The MicroVM `start-session` step disables automatic durable **error** retries after its own recovery attempt. Crash replay is still possible and uses the saved receipt. Live AWS token-retention, changed-request and concurrent-conflict behavior remain verification gates. ### Step 5: Await completion @@ -217,7 +223,7 @@ The orchestrator polls for completion using `waitForCondition` from the Durable | ECS | `DescribeTasks`, including container exit status and exit code | | Lambda MicroVMs | `GetMicrovm` state plus agent heartbeat | -While waiting between polls, the durable orchestrator suspends without compute charges. If the session is terminated externally (crash, timeout, cancellation), the poll detects it and the orchestrator proceeds to finalization using GitHub-based result inference as fallback. +While waiting between polls, the durable orchestrator suspends without compute charges. If the session is terminated externally (crash, timeout, cancellation), the poll detects it and the orchestrator proceeds to finalization after a strongly consistent task read; it preserves an already committed terminal result. ### Step 6: Finalization @@ -246,9 +252,9 @@ After the session ends, the orchestrator determines the outcome from multiple si ### Step execution contract -Every step in the pipeline satisfies these properties: +Configured workflow steps target the following contract. The top-level durable orchestrator still needs explicit guards around external effects, as described under recovery below. -- **Idempotent** - Safe to retry after crashes. Context hydration produces the same prompt for the same inputs; session-start retry semantics are implemented by each backend strategy. +- **Replay-aware** - A retry must preserve task intent and avoid repeating external effects. Session-start recovery is implemented by each backend strategy; a checkpoint alone is not an idempotency guarantee. - **Timeout-bounded** - Each step has a configurable timeout to prevent blocking the pipeline. - **Failure-aware** - Returns `success` or `failed`. Infrastructure failures (throttle, transient errors) trigger exponential backoff retries (default: 2 retries, base 1s, max 10s). Explicit failures transition to `FAILED` without retry. - **Least-privilege input** - Each step receives only the `blueprintConfig` fields it needs. Custom Lambda steps get credential ARNs stripped. @@ -276,7 +282,9 @@ Liveness detection varies by compute backend. AgentCore sessions use DynamoDB he - **Grace period** (120s) - After entering `RUNNING`, the orchestrator waits before expecting heartbeats (covers container startup). - **Stale threshold** (240s) - If the heartbeat exists but is older than this, the session is treated as lost. -- **Early crash** - If no heartbeat is ever set after the combined window (360s), the agent died before the pipeline started. +- **Early crash** - If no heartbeat is ever set after the combined window (360s), the session is treated as lost; a process failure or failed DynamoDB writes can cause this. + +Approval waits suppress heartbeat writes. When the agent consumes a decision and restores `RUNNING`, its conditional transaction also refreshes `agent_heartbeat_at`, so the first poll after a long wait does not mistake the old timestamp for a crash. When the session is unhealthy, the task transitions to `FAILED` with "Agent session lost: no recent heartbeat." @@ -284,14 +292,50 @@ When the session is unhealthy, the task transitions to `FAILED` with "Agent sess **Lambda MicroVM state polling.** Liveness is a dual signal. The strategy maps `GetMicrovm` mechanically: `PENDING`/`RUNNING` report `running`, `SUSPENDING`/`SUSPENDED` report `suspended`, and `TERMINATING`/`TERMINATED` report terminal completion. The orchestrator supplies the health interpretation: -- `suspended` is healthy only while the task is `AWAITING_APPROVAL`; in any other task state it emits an anomaly and keeps polling rather than failing recoverable work. -- A terminal substrate report paired with a non-terminal task is a failure, but the orchestrator first re-reads the task row to confirm the agent did not write a terminal result between the original read and VM termination. -- Substrate state detects a dead VM; heartbeat staleness detects a hung, deadlocked, or OOM-killed pipeline inside a VM that still reports `RUNNING`. +- Intentional suspension requires the matching pending gate and saved suspend intent. Unexpected suspension emits one anomaly per episode and starts bounded wake recovery, preserving recoverable work. +- A terminal substrate report paired with a non-terminal task first checks for a complete, acknowledged approval checkpoint. Such a checkpoint can retire the old attempt and retain the task for a replacement. Without one, finalization strongly re-reads the task row before classifying a substrate failure. +- Substrate state detects a dead VM; heartbeat staleness detects loss of the heartbeat writer inside a VM that still reports `RUNNING`. The independent heartbeat thread can continue during a pipeline hang, so a fresh timestamp is not proof of progress. + +The P3 supervisor saves intent before control calls and rechecks the gate before and after them. Its durable state retains an absolute service lifetime, consecutive failures, recovery start time and next delay. Three failed cycles or 120 seconds of unconfirmed wake cannot become an indefinite wait. AWS RUNNING does not end recovery while the guest remains stuck on a decided/expired approval; fresh guest liveness is required. API approve/deny commit first, then attempt a bounded wake without changing the decision response. Automatic suspension defaults off via `microvm_approval_suspend_enabled`; disabling new sleep preserves wake and cleanup. See the [lifecycle diagnostics](../verification/645-p3-lifecycle-diagnostics.md). + +Unanswered approvals have no deadline by default. A checkpointed MicroVM wait can +retire after an hour, or before that worker's lifetime ends. Retirement and +replacement follow the [retained approval protocol](#retained-microvm-approvals). `TERMINATED` is the normal terminal signal and remains observable for at least 10 minutes. `ResourceNotFoundException` maps to completion only as a late fallback after the control-plane record is eventually reaped; polling does not wait for `NotFound`. **`/ping` health endpoint (AgentCore only).** The agent's FastAPI server responds to AgentCore's `/ping` calls while the coding task runs in a separate thread. AgentCore sees `HealthyBusy` and keeps the session alive. +### Retained MicroVM approvals + +A pending approval has no deletion timer by default. Closing its task cancels +unanswered requests and retains the decision history for 90 days. An explicit +positive approval deadline still applies; waking or replacing a worker does not +restart it. The agent decides whether the approved action remains relevant. + +At the approval barrier, the worker saves the conversation, exact pending tool +inputs, workflow context, cumulative usage and Git/workspace archive. The +coordinator verifies the checksummed S3 object versions before fencing the old +attempt through its coordinator-owned `worker-lease#` record. Workers may +read and condition-check that lease but cannot modify it. Only confirmed shutdown +allows the task to become `PARKED` and release its concurrency reservation. + +An answer admits one replacement when capacity permits. Admission, the new lease +and capacity reservation are one DynamoDB transaction; a deterministic Durable +execution name deduplicates invocation. The replacement uses the original +published coordinator and exact image version, restores the files/conversation, +and consumes the recorded answer. Remaining cost and turns are the original +allowance minus accumulated usage. Repository-free tasks preserve scratch files +using a private directory and local Git baseline. + +The scheduled continuation manager retries unfinished retirement, missed +dispatches and terminal cleanup, saving its scan cursor between invocations. +It cannot launch, suspend or resume workers; launching stays with the pinned +coordinator. A failed start response does not prove no worker exists. Capacity +remains reserved until shutdown is confirmed, or the full service lifetime has +elapsed for an unknown handle. The exact-attempt lease becomes `CLOSED` before +atomic release; `TERMINATING` alone is insufficient. + ### The idle timeout problem AgentCore terminates sessions after 15 minutes of inactivity. Since coding tasks may have long pauses between tool calls (builds, complex reasoning), the agent uses `add_async_task` to register background work. The SDK reports `HealthyBusy` via `/ping` while any async task is active, preventing idle termination. @@ -314,7 +358,7 @@ Long-running distributed systems fail. The orchestrator is designed so that ever | Hydration | Guardrail API unavailable | Fail the task (fail-closed: unscreened content never reaches agent) | | Session start | Selected compute service throttled | Exponential backoff. Fail after retries exhausted. | | Session start | Session crashes immediately | AgentCore: heartbeat never set, detected after 360s grace window. ECS: `DescribeTasks` reports failure. Lambda MicroVMs: `GetMicrovm` reports terminal state or the heartbeat never appears. | -| Running | Agent crashes mid-task | AgentCore: heartbeat goes stale. ECS: `DescribeTasks` reports stopped task. Lambda MicroVMs: `GetMicrovm` detects VM death and heartbeat staleness detects an in-guest hang. Finalization inspects GitHub for partial work. | +| Running | Agent crashes mid-task | AgentCore: heartbeat goes stale. ECS: `DescribeTasks` reports stopped task. Lambda MicroVMs: `GetMicrovm` detects VM death and heartbeat staleness detects loss of the in-guest writer. Finalization preserves committed task results and records a specific failure for an active lost session. | | Running | Agent hits turn or budget limit | Session ends normally. Finalize based on what was produced. | | Running | Idle for 15 min | AgentCore kills session. Task transitions to `TIMED_OUT`. | | Finalization | GitHub API down | Retry 3x. If still failing, mark `FAILED` with infrastructure reason. | @@ -323,14 +367,14 @@ Long-running distributed systems fail. The orchestrator is designed so that ever ### Recovery mechanisms 1. **Durable execution** - Lambda Durable Functions checkpoints at each state transition and replays after crashes. -2. **Idempotent operations** - All steps are safe to retry. +2. **Replay guards** - Operations need their own idempotency controls; checkpointing alone does not make external effects exactly-once. MicroVM starts use saved receipts. Capacity acquisition and release use task-owned markers updated atomically with the user counter. This guards the seat count; terminal audit events are still allowed to repeat on replay. 3. **Stuck-task scanner** - Periodic Lambda detects tasks stuck beyond expected durations and either resumes or fails them. -4. **Counter reconciliation** - Lambda runs every 15 minutes, compares counters to actual running task counts, corrects drift. Emits `counter_drift_corrected` CloudWatch metric. +4. **Counter reconciliation** - Every 15 minutes, the Lambda strongly scans counter records followed by task reservations. It repairs a count only if the saved counter revision is unchanged, then releases terminal held reservations. Structured logs record repairs, failures, empty counters and ambiguous legacy ownership; this handler does not publish a `counter_drift_corrected` metric. 5. **Dead-letter queue** - Tasks that exhaust retries go to DLQ for investigation. ## Concurrency and scaling -Each task runs in its own isolated compute session with no shared mutable state at the compute layer. The orchestrator manages concurrency purely at the coordination layer: atomic counters track how many tasks are active per user and system-wide, and admission control enforces limits before resources are consumed. +Each task runs in an isolated compute session. The orchestrator reserves capacity per user before starting compute; AWS separately enforces backend quotas. Approval waits keep their reservation, including when P3 suspends the MicroVM. This bounds unfinished sessions and their eventual resume demand; AWS memory-quota use while suspended remains unverified. ### Capacity limits @@ -338,15 +382,24 @@ Each task runs in its own isolated compute session with no shared mutable state |---|---|---| | `invoke_agent_runtime` TPS | 25 per agent/account | AgentCore quota (adjustable) | | Concurrent sessions | Account-level limit | AgentCore quota | -| Per-user concurrency | Configurable (default 3-5) | Platform config | -| System-wide max tasks | Configurable | Bounded by selected-backend quotas | +| Per-user concurrency | Configurable (default 3) | `MAX_CONCURRENT_TASKS_PER_USER` | ### Counter management -- **UserConcurrency** - DynamoDB item per user with `active_count`. Incremented atomically (`active_count < max`) at admission, decremented at finalization. -- **SystemConcurrency** - Single DynamoDB item, same pattern. +A reservation is a saved seat for one task. `task-concurrency.ts` owns both operations: + +- **Acquire:** require a matching owner, `SUBMITTED` status and no prior reservation; set `concurrency_slot.state = held` and increment the counter in the same transaction. A held reservation is reused on replay. A released task cannot reserve again; a new attempt gets a new task ID. +- **Release:** require a terminal task and a held reservation; mark it `released` and decrement the counter in the same transaction. Normal finalization, early failure and stranded cleanup use this helper. A missing or already-released marker does not decrement. An active task keeps its seat. + +Finalization attempts release even if an audit event fails. A failed status write that leaves the task active is propagated for retry. Cancellation before admission/pre-flight completion also checks for a held terminal reservation. If a crash separates the terminal write from release, the scheduled counter reconciler completes release later. + +Every reservation change writes a fresh `reservation_version` on the counter. Reconciliation uses strongly consistent **base-table** scans, once for counters and once for tasks, then compares the saved revision before repairing a count. A count-only comparison cannot detect an increment followed by a decrement. Held terminal reservations are included until their release commits. A partial scan never installs a partial count. These scans consume read capacity across retained rows; check scan duration and capacity at deployment scale. + +If a release discovers an empty/missing counter, it closes the marker without subtracting from seats reserved meanwhile. A concurrent repair can leave a conservative overcount until the next sweep. `CONCURRENCY_EMPTY_COUNTER` records that condition. + +Older active records without reservation markers are ambiguous: status alone does not prove admission. The helper never guesses that they own a seat; reconciliation skips count repair for that user and logs `CONCURRENCY_RESERVATION_UNKNOWN`. Pause new submissions and drain old executions before deploying this protocol across all counter writers. After old tasks settle, reconcile and resume admissions. Rollback also requires draining tasks using the newer protocol; mixing old direct decrements with new markers does not provide this guarantee. -Concurrency is always released in `finalizeTask` (step 6), never inside the poll loop. ECS poll failure paths call `failTask` with `releaseConcurrency: false` to transition the task to `FAILED` without decrementing — `finalizeTask` handles the single decrement after re-reading the task state. The heartbeat-detected crash path also guards against double-decrement by only releasing the counter after a successful state transition. If the transition fails (task already terminal), it re-reads and acts accordingly. +This protocol protects against replay and competing **cooperative** writers. The agent role can currently update/replace its own task row, so internal fields are not yet protected against a compromised agent. Coordinator-only metadata storage or constrained agent writes remain a separate security prerequisite. ## Implementation @@ -398,7 +451,7 @@ At 500 concurrent tasks, peak TPS is ~16.7 - well within the 25 TPS AgentCore qu ## Data model -Three DynamoDB tables back the orchestrator: one for task state, one for the audit log, and one for concurrency counters. The Tasks table is the source of truth for every task; the orchestrator reads and writes it at every state transition. TaskEvents is append-only and powers the `GET /v1/tasks/{id}/events` API. UserConcurrency is a lightweight counter table used only during admission and finalization. +Three DynamoDB tables back the orchestrator: one for task state, one for the audit log, and one for concurrency counters. The Tasks table is the source of truth for every task; the orchestrator reads and writes it at every state transition. TaskEvents is append-only and powers the `GET /v1/tasks/{id}/events` API. UserConcurrency stores reservation counts and revision tokens used by admission, cleanup and reconciliation. ### Tasks table (DynamoDB) @@ -416,7 +469,8 @@ Three DynamoDB tables back the orchestrator: one for task state, one for the aud | `branch_name` | String | `bgagent/{task_id}/{slug}` for new tasks; PR's `head_ref` for PR tasks | | `session_id` | String? | Backend session identifier (AgentCore session ID, ECS task ARN, or MicroVM ID) | | `compute_type` | String? | Selected backend: `agentcore`, `ecs`, or `lambda-microvm` | -| `compute_metadata` | Map? | Backend lifecycle handle; Lambda MicroVMs persist `microvmId` and `endpoint` | +| `compute_metadata` | Map? | Backend lifecycle handle; Lambda MicroVMs persist `microvmId`, `endpoint`, actual image identity and verified lifecycle protocol when available | +| `concurrency_slot` | Map? | Internal reservation `{state, acquired_at, released_at?}`; excluded from public task responses | | `execution_id` | String? | Durable execution ID | | `pr_url` | String? | PR URL (set during finalization) | | `error_message` | String? | Error reason if FAILED | @@ -452,7 +506,7 @@ Append-only audit log. See [OBSERVABILITY.md](./OBSERVABILITY.md). | Field | Type | Description | |---|---|---| | `user_id` (PK) | String | User ID | -| `active_count` | Number | Running task count | +| `active_count` | Number | Held reservation count, including approval waits and pending terminal cleanup | +| `reservation_version` | String? | Fresh revision token on every reservation mutation or count repair | -Increment: `SET active_count = active_count + 1` with `ConditionExpression: active_count < :max`. -Decrement: `SET active_count = active_count - 1` with `ConditionExpression: active_count > 0`. +Counter changes belong to the reservation transactions described above. Do not add a standalone increment or decrement: it bypasses per-task replay protection. diff --git a/docs/design/REGISTRY.md b/docs/design/REGISTRY.md index a74c1e222..8136e424f 100644 --- a/docs/design/REGISTRY.md +++ b/docs/design/REGISTRY.md @@ -37,6 +37,12 @@ A **registry asset** is a versioned, immutable-per-version runtime artifact that **Cedar parity.** Registry Cedar text reaches the agent through the **same** `cedar_policies` payload field as inline blueprint policies, so it is byte-identical from the `PolicyEngine`'s view. **Skills** are prompt text only: a skill cannot invoke tools; its `tool_hints` are advisory prose referencing tools an MCP server separately provides (no transitive dependency — the operator attaches both). +**Runtime network support (#818).** An MCP server is a program the coding agent calls to use extra tools. Remote `http`/`sse` assets need an **HTTPS endpoint reachable on TCP port 443** under the shipped **AgentCore, ECS and Lambda MicroVM** network policies. Remote ports such as 80, 8080 or 8443 are unsupported on all three defaults; this is not a MicroVM-only limitation. A `stdio` asset starts a local program and uses its input/output pipes, but that program's outbound network requests still face the runtime policy. The MicroVM image builder's separate 80/443 permission does not apply to task execution. + +Registry resolution validates/pins the asset and the loader writes its configuration; neither probes network reachability or rejects an asset merely because its URL uses a different port. This is a documented support constraint, not a new synth/onboarding validator. Use a reachable HTTPS/443 service or proxy and verify DNS, routing, TLS and authentication on the deployed backend. Port 443 alone does not prove reachability. + +**Task delivery.** `resolved_assets` stays in the shared agent payload. ECS and MicroVM now deliver every payload through v2 authenticated bootstrap and a single-object signed URL. Registry content may exceed 4,096 bytes without expanding the MicroVM hook reference; the 8 MiB task-document cap still applies. Local integration coverage resolves a large MCP asset, runs the real payload assembly, verifies S3 bytes and the small hook reference, then separately exercises the Python hook mapper and local `.mcp.json` loader. That proves delivery/configuration, not a live connection to the remote tool. See [COMPUTE.md](./COMPUTE.md) for runtime behavior. + ## 3. Substrate mapping (the core design) Agent Registry answers *"what servers/skills exist, find me one"* (discovery metadata + semantic search). ABCA needs *"give me the exact runtime config to load this pinned asset."* These are two different objects, and the service validates the discovery object against the official schemas (an MCP record's body must be a valid MCP `server.json`, not our `.mcp.json`). So **every record carries BOTH a discovery descriptor AND ABCA's runtime payload.** diff --git a/docs/design/SECURITY.md b/docs/design/SECURITY.md index 3935d020e..00dee4cd3 100644 --- a/docs/design/SECURITY.md +++ b/docs/design/SECURITY.md @@ -41,12 +41,18 @@ Three authentication mechanisms protect the platform, matching its input channel **Per-session IAM scoping** - The agent does not use its long-lived compute role (the AgentCore Runtime `ExecutionRole`, ECS Fargate task role, or Lambda MicroVMs execution role) for tenant data. Instead, at task startup it assumes a per-task **SessionRole** via `sts:AssumeRole` with session tags `{user_id, repo, task_id}`, and uses the resulting short-lived credentials for all DynamoDB and S3 tenant-data access. The SessionRole's policies self-constrain on those tags: -- **DynamoDB**: item access on the four `task_id`-partitioned tables (task, events, approvals, nudges) is gated by a `dynamodb:LeadingKeys` condition equal to `${aws:PrincipalTag/task_id}`, so a session can read or write only its own task's rows. `Scan` is not granted (it ignores leading-keys). `task_id` is the isolation boundary because it is the base-table partition key — `LeadingKeys` cannot bind to a GSI partition key such as `user_id`. +- **DynamoDB**: item access on the four `task_id`-partitioned tables (task, events, approvals, nudges) is gated by a `dynamodb:LeadingKeys` condition equal to `${aws:PrincipalTag/task_id}` and requires that context key to be present. `Scan` is not granted. The main task table permits reads and only attribute-scoped `UpdateItem` writes: the agent can report status, heartbeat, results and approval details, but cannot replace/delete the row or edit coordinator-owned start receipts, capacity reservations, owner identity or compute handles. The reviewed write list is `cdk/src/constructs/agent-task-write-attributes.json`; Python tests exercise the current writers against it. Events and nudges retain task-scoped item writes. Approval records permit reads and transaction condition checks only. `LeadingKeys` binds the base-table partition key, not a GSI key such as `user_id`. +- **Approval requests**: workers call an IAM-signed API path restricted to their task tag. Its trusted handler can create `PENDING` or conditionally record `TIMED_OUT`; human decisions, notification markers and retention fields are unavailable to the worker. IAM binds the caller to the tagged task path. The transaction guards against concurrent task ownership/state changes and checks MicroVM worker leases. See the [approval trust boundary](./CEDAR_HITL_GATES.md#121-trust-boundaries) and [upgrade procedure](../guides/DEPLOYMENT_GUIDE.md#upgrading-approval-permissions). - **S3**: trace writes and attachment reads are scoped to the `/${aws:PrincipalTag/user_id}/` object prefix. -The compute role retains only non-tenant access (Bedrock model invocation — already ARN-scoped; CloudWatch Logs; the GitHub PAT secret, read once before the SessionRole is assumed; AgentCore Memory) plus `sts:AssumeRole`/`sts:TagSession` on the SessionRole. Because the agent runs under credentials that are themselves an assumed role, its `AssumeRole` is *role chaining* — capped at one hour regardless of the role's max session duration — so the agent uses a **refreshable** credential provider that re-assumes before expiry (tasks can run up to the 8-hour `maxLifetime`). The design is backend-agnostic: the same SessionRole and agent code serve all three backends (AgentCore Runtime, ECS Fargate, Lambda MicroVMs). A compromised agent session is therefore confined to its own task's data, enforced at the IAM layer rather than by application-code conventions. The policy structure (the `dynamodb:LeadingKeys` condition on `${aws:PrincipalTag/task_id}`, per-user S3 prefixes, and `Scan` exclusion) is asserted by CDK template tests, and the refreshable-credential and session-tag flow by agent unit tests; the matching-tag → allow / mismatched-or-absent-tag → deny behaviour was additionally confirmed once via the IAM policy simulator during development. +The compute role retains only non-tenant access (Bedrock model invocation — already ARN-scoped; CloudWatch Logs; the GitHub PAT secret, read once before the SessionRole is assumed; AgentCore Memory) plus `sts:AssumeRole`/`sts:TagSession` on the SessionRole. Because the agent runs under credentials that are themselves an assumed role, its `AssumeRole` is *role chaining* — capped at one hour regardless of the role's max session duration — so the agent uses a **refreshable** credential provider that re-assumes before expiry (tasks can run up to the 8-hour `maxLifetime`). The design is backend-agnostic: the same SessionRole and agent code serve all three backends (AgentCore Runtime, ECS Fargate, Lambda MicroVMs). Existing scoped credentials are constrained to their tags, but the compute role chooses those tags when assuming the role. The trust policy does not independently bind a worker to one task; this is not complete isolation against a compromised worker with ambient credentials. No worker needs direct access to the shared capacity counter. ECS's configuration without a session role retains unscoped task reads/reporting updates, with the same attribute restriction, but cannot enable approval gates; the construct rejects approval wiring without the session role and service URL. The policy structure and Python writer compatibility are tested locally. A historical IAM simulator check covered the original tag conditions; it does not verify the new attribute restriction or deployed transaction authorization. The [acceptance checklist](../verification/README.md#live-acceptance-for-an-installation) includes effective-role authorization checks. + +**Lambda MicroVMs compute-role delta** — The compute role additionally reads the per-workspace channel-OAuth secrets (`bgagent-linear-oauth-*`, `bgagent-jira-oauth-*`) and its payload bucket's `bootstrap/*` manifests. Ambient task-object reads and payload-bucket listing are explicitly denied. It also has read-only `ec2:DescribeAvailabilityZones` for repository CDK synthesis; that API has no resource-level scope. + +Recorded live calls rejected source-conditioned MicroVM role trust and service-conditioned PassRole grants. The integration therefore omits those conditions on the affected paths, while restricting which roles can be passed and which resources they may access. Reintroduce conditions only after verifying service support. [ADR-021 §4](../decisions/ADR-021-lambda-microvms-compute-backend.md#4-infra-and-iam-conditional-resources-behind-bootstrap-computetypes) records the role responsibilities and limits. + +**ECS/MicroVM task bootstrap (v2)** — The coordinator publishes non-secret deployment manifests and supplies a short-lived signed URL for exactly one task's payload. The worker reads the manifest with its ambient role, whose explicit deny outside its own `bootstrap/*` prevents a foreign public bucket from authorizing fake configuration. It verifies downloaded task identity and exact manifest/config agreement before installing MicroVM settings. The coordinator privately persists the URL outside TaskTable for retry and deletes both task objects at finalization. Worker roles cannot read/list task objects using their ambient credentials. A signed URL is a bearer capability: redact it in errors and keep it out of task rows and repository subprocess environments. Existing session-tag trust and other platform grants remain separate limitations. The [payload guide](../verification/645-payload-bootstrap.md) describes coordinated image/coordinator/policy upgrades; the [acceptance summary](../verification/README.md) distinguishes recorded AWS checks from remaining validation. -**Lambda MicroVMs compute-role delta** - On this backend the compute role additionally holds prefix-scoped `secretsmanager:GetSecretValue` on the per-workspace channel-OAuth secrets (`bgagent-linear-oauth-*`, `bgagent-jira-oauth-*`), read access to the `/run` payload bucket (the task payload arrives as an S3 object, not as environment), and `ec2:DescribeAvailabilityZones` (`Resource: *`, read-only — EC2 describe actions have no resource-level scoping) for a CDK repo's own synth gate. It is also **the only compute role in the platform whose trust policy carries no confused-deputy condition**: the Lambda MicroVMs service populates no `aws:SourceAccount` or `aws:SourceArn` when it assumes the role, so a trust policy carrying one is unassumable — verified live, not assumed (connector creation failed deterministically, and `RunMicrovm` surfaced the same root cause as a misleading caller-side `iam:PassRole` denial). The compensating controls are that each of the three MicroVM roles can be passed to `lambda.amazonaws.com` only by a named principal — the orchestrator, via an `iam:PassRole` scoped to the execution role's exact ARN, and the CloudFormation deployment role for the build and connector-operator roles — that every other resource they reach is account-scoped by ARN apart from two justified `Resource: *` read/create-time statements, and that none of them holds `iam:*`, cross-account trust, or any `sts:AssumeRole` beyond the execution role's scoped hop to the per-task SessionRole. Full evidence, the two-arm PassRole experiment, and the alternatives considered are in [ADR-021 §4](../decisions/ADR-021-lambda-microvms-compute-backend.md#4-infra-and-iam-conditional-resources-behind-bootstrap-computetypes). > Out of scope for this control and tracked separately as GitHub issues: replacing the shared GitHub PAT (GitHub App / Token Vault), binding credentials to the MicroVM via attestation, and scoping AgentCore Memory (namespace isolation by `actorId`/`sessionId` remains its boundary). diff --git a/docs/guides/CEDAR_POLICY_GUIDE.md b/docs/guides/CEDAR_POLICY_GUIDE.md index 15fd92322..0ad43f54b 100644 --- a/docs/guides/CEDAR_POLICY_GUIDE.md +++ b/docs/guides/CEDAR_POLICY_GUIDE.md @@ -63,11 +63,14 @@ Every tool call the agent makes is evaluated as a Cedar `(principal, action, res |---|---|---|---| | `@rule_id("...")` | **Yes on soft-deny** (recommended on hard-deny) | Unique kebab/snake-case identifier | Stable ID for `--pre-approve rule:X`, for audit events, and for `bgagent policies show --rule X`. Engine rejects duplicates at task start. | | `@tier("hard"\|"soft")` | **Yes** | Exactly one of `"hard"` or `"soft"` | Must match the file section. Mismatches fail task start. | -| `@approval_timeout_s("N")` | No | Integer seconds ≥ 30 | Per-rule timeout. Defaults to 300 s (overridable per-task via `--approval-timeout`). When multiple soft rules match, the engine picks the minimum. Values < 120 s emit a load-time warning; values < 30 s are rejected. Ignored on hard-deny. | +| `@approval_timeout_s("N")` | No | Integer seconds ≥ 30 | Optional per-rule decision deadline. If absent, uses the task setting, whose default `0` means no deadline. The shortest positive task/rule deadline wins. Values < 120 s emit a load-time warning; values < 30 s are rejected. Ignored on hard-deny. | | `@severity("low"\|"medium"\|"high")` | No | One of three | Displayed in the approval prompt. Default: `medium`. | | `@category("...")` | No | `destructive`, `network`, `filesystem`, `auth`, or free-form | Optional UX grouping. Not enforced. | -**Rule of thumb:** every soft-deny rule must have `@rule_id` and should set `@severity` + `@approval_timeout_s` explicitly. Users scanning `bgagent pending` lean on these fields to triage quickly. +**Rule of thumb:** every soft-deny rule must have `@rule_id` and should set +`@severity`. Add `@approval_timeout_s` only when the workflow needs a decision +deadline. Omit the annotation to let the task choose; unlike the task's `0` +setting, a zero-valued rule annotation is invalid. ## Common patterns diff --git a/docs/guides/DEPLOYMENT_GUIDE.md b/docs/guides/DEPLOYMENT_GUIDE.md index d38efe314..ab8696654 100644 --- a/docs/guides/DEPLOYMENT_GUIDE.md +++ b/docs/guides/DEPLOYMENT_GUIDE.md @@ -4,31 +4,34 @@ This guide covers deploying ABCA into an AWS account, including compute backend ## Architecture overview -ABCA deploys as a **single CDK stack** (`backgroundagent-dev`) containing all platform resources. The stack uses a `ComputeStrategy` interface to support three compute backends within the same stack: +ABCA deploys through a **root CDK stack** (`backgroundagent-dev`) and nested stacks for parts of the platform, including registry infrastructure. A `ComputeStrategy` interface supports three compute backends: | Aspect | AgentCore (default) | ECS Fargate (opt-in) | Lambda MicroVMs (experimental) | |--------|--------------------|--------------------|--------------------| | **Compute** | Bedrock AgentCore Runtime (Firecracker MicroVMs) | ECS Fargate containers | AWS Lambda MicroVMs | -| **Resources** | 2 vCPU, 8 GB RAM, 2 GB max image size | 2 vCPU, 4 GB RAM | 8 GB baseline / 32 GB peak memory | +| **Resources** | 2 vCPU, 8 GB RAM, 2 GB max image size | Build: 4 vCPU / 16 GiB; planning: 2 vCPU / 8 GiB (configurable) | 8 GiB baseline / 32 GiB peak memory | | **Orchestration** | Durable Lambda (checkpoint/replay) | Same durable Lambda via `ComputeStrategy` | Same durable Lambda via `ComputeStrategy` | | **Agent mode** | FastAPI server (HTTP invocation) | Batch (run-to-completion) | FastAPI server (lifecycle hooks) | | **Startup** | ~10s (warm MicroVM) | ~60-180s (Fargate cold start) | ~6s to `RUNNING` (live-measured) | | **Max duration** | 8 hours (AgentCore service limit) | 9 hours (orchestrator `executionTimeout`) | 8 hours (`maximumDurationInSeconds`) | -All backends are orchestrated by the same durable Lambda function. The `ComputeStrategy` interface abstracts `startSession()`, `pollSession()`, and `stopSession()` -- the ECS strategy calls `ecs:RunTask` / `ecs:DescribeTasks` / `ecs:StopTask` directly from the Lambda. No Step Functions are used. +All backends are orchestrated by the same durable Lambda function. The `ComputeStrategy` interface abstracts `startSession()`, `pollSession()`, and `stopSession()` -- the ECS strategy calls `ecs:RunTask` / `ecs:DescribeTasks` / `ecs:StopTask` directly from the Lambda. Task orchestration does not use Step Functions; other platform features may use them. -ECS Fargate is currently **opt-in** -- the `EcsAgentCluster` construct is present in the stack code but commented out. To enable it, uncomment the ECS blocks in `cdk/src/stacks/agent.ts`. +ECS Fargate is **opt-in**. Deploy with `--context compute_type=ecs`; the stack enables `EcsAgentCluster` from that context flag. ### Lambda MicroVMs backend (experimental) -> **Not for production.** `lambda-microvm` carries no smoke-parity guarantee for an unattended deployment. Keep production repositories on `agentcore` or `ecs`. Synth emits an unsuppressible warning to this effect whenever the backend is selected. Design detail: [COMPUTE.md](../design/COMPUTE.md) and [ADR-021](../decisions/ADR-021-lambda-microvms-compute-backend.md). +> **Not for production.** `lambda-microvm` carries no smoke-parity guarantee for an unattended deployment. Keep production repositories on `agentcore` or `ecs`. Synth emits a verification warning whenever a MicroVM image is configured; selecting the backend without an image emits a separate setup warning. Design detail: [COMPUTE.md](../design/COMPUTE.md) and [ADR-021](../decisions/ADR-021-lambda-microvms-compute-backend.md). -Selecting it is a synth-time context flag: +For a new installation, select the backend and nested layout: ```bash -mise //cdk:deploy -- --context compute_type=lambda-microvm +mise //cdk:deploy -- --context compute_type=lambda-microvm --context microvm_nested_stack=true ``` +Existing flat installations must retain `microvm_nested_stack=false` until +completing the [resource migration](../verification/645-p3-nested-stack.md). + **You must re-bootstrap first.** This is the single most common way this backend fails, and the failure does not look like a configuration problem: 1. Check the bootstrap policy bundle already deployed in the account: @@ -53,9 +56,11 @@ mise //cdk:deploy -- --context compute_type=lambda-microvm Operational notes specific to this backend: -- **Nothing self-terminates.** A MicroVM whose task finished, crashed, or hung stays `RUNNING` and billing until the 8-hour cap. The orchestrator calls `TerminateMicrovm` on finalize, and the heartbeat-staleness check catches a hung guest inside a healthy VM -- but a leaked handle is a cost incident. The one exception: the service reaps a VM whose `/run` hook returns 4xx (~12s). +- **Nothing self-terminates.** A MicroVM whose task finished, crashed, or hung stays `RUNNING` and billing until the 8-hour cap. The orchestrator calls `TerminateMicrovm` on finalize, and the heartbeat-staleness check detects loss of the in-guest heartbeat writer (a pipeline hang can leave that writer running) -- but a leaked handle is a cost incident. The one exception: the service reaps a VM whose `/run` hook returns 4xx (~12s). - **Logs** land in `/aws/lambda-microvms/`. Guest stdout goes there too, which is the fallback path when the agent cannot reach the application log group. -- **Deployment identifiers are not baked into the image.** The snapshot carries no configuration; table names, secret ARNs, and the per-task session-role ARN arrive in the `/run` payload as a `platform_config` block. A version-skewed orchestrator that does not send it is refused rather than run with tenant scoping disabled. +- **Deployment identifiers are not baked into the image.** Current table names, secret ARNs and session-role ARN arrive through the v2 IAM-authenticated manifest and signed task document. Old unsigned envelopes are refused. Deploy matching coordinator code, worker images and IAM with admissions paused and old tasks drained; the repository runbook `docs/verification/645-payload-bootstrap.md` records the procedure and pending live checks. +- **Registry tools share the runtime network restriction.** Remote HTTP/SSE MCP assets need reachable HTTPS/443 endpoints. AgentCore and ECS defaults also block remote non-443 ports. `stdio` programs run locally but their outbound calls remain restricted; resolution does not test connectivity. The image builder's 80/443 access does not widen runtime egress. See [REGISTRY.md](../design/REGISTRY.md). +- **Logging failures have a fallback record.** Debug/warn CloudWatch failures emit `cloudwatch_write_failed` to stdout with writer/task/error class, without another AWS call or sensitive log text. This is not a configured metric/alarm; verify collection in the guest log stream while the VM is running. ### Optional Agent Registry @@ -241,6 +246,66 @@ Triggers via `workflow_run` when `build.yml` completes successfully. The pipelin ## Known deployment issues +### Upgrading approval permissions + +Worker approval creation and timeout writes now use an IAM-authenticated +service. Deploy the matching agent image and CDK together: old workers write +directly to DynamoDB and cannot create new gates after those permissions are removed. + +1. Pause submissions from the CLI, integrations and schedules during the upgrade. + Let existing tasks finish, or have their owners cancel them. Include tasks + awaiting approval, suspended MicroVMs and retained continuations; an empty + running-container list does not prove the deployment has drained. +2. Build the agent from the same revision as the CDK. For an externally managed + MicroVM image, publish that build and select its new version before resuming + submissions. A suspended VM keeps its old code. +3. Deploy the stack. Check that the SessionRole has approval-table reads and + condition checks only, plus `execute-api:Invoke` restricted to its task tag. + CDK supplies `APPROVAL_REQUESTS_API_URL` to all three compute backends. + Custom ECS constructs must provide both the SessionRole and service URL; + approval wiring without them is rejected before deployment. A MicroVM + manifest missing the URL is rejected before the worker starts. + If an AgentCore environment was edited to remove the URL, the worker reports + `APPROVAL_REQUESTS_API_URL is required for cloud approval requests` at its + next gate rather than attempting a direct DynamoDB write. Redeploy the + matching stack and image. +4. Submit a test task that triggers a known approval rule on each enabled backend. + Verify that the request appears, an owner decision resumes it, and an explicit + deadline records `TIMED_OUT` without overwriting a human decision. Then resume + normal submissions. + +If an old worker survives the upgrade, its next approval write fails closed. +Existing rows remain readable; do not restore direct writes to work around a stale +image. Roll forward with the matching image. Rolling back IAM restores the original +approval-record vulnerability and requires a deliberate operator decision. + +### Scheduled maintenance stack + +Concurrency repair, admission-queue pickup, stranded-task repair and pending-upload +cleanup run in the `ConcurrencyMaintenance` nested stack. MicroVM continuation +recovery also runs there when an image is configured. An upgrade recreates the +stateless functions, roles and schedules; their task tables and storage stay in +the parent stack. This applies to every compute backend, including AgentCore, +and replaces roughly twenty resources, depending on enabled features. +CloudFormation creates the new schedules before deleting the old ones, so both +can fire during the update. Task mutations and the continuation scan cursor use +conditional writes to tolerate that overlap. The drain above is required for +the approval permission/image upgrade, not a scheduling gap. + +Review the replacements in the change set. Existing Lambda log groups remain +under their old generated names; use the new function's log group for post-upgrade +invocations and retain the old groups when investigating earlier runs. +Both flat and nested MicroVM layouts support the optional +tool gateway and Linear Identity vault without exceeding the template budget. + +New MicroVM installations must explicitly select `microvm_nested_stack=true` +(bootstrap bundle 1.9.0). +Before upgrading an existing flat MicroVM deployment, save +`"microvm_nested_stack": false` in its CDK context or pass +`--context microvm_nested_stack=false` on every deploy. Keep this escape hatch +until completing the [resource migration](../verification/645-p3-nested-stack.md). +Omitting the setting fails synthesis. Selecting `true` does not migrate existing flat resources. + ### AgentCore unsupported Availability Zones **Affects:** Fresh deploys in accounts whose default Availability Zones don't line up with the zones AgentCore supports for the region. diff --git a/docs/guides/LINEAR_SETUP_GUIDE.md b/docs/guides/LINEAR_SETUP_GUIDE.md index e4b6685b9..0ce9f5ae3 100644 --- a/docs/guides/LINEAR_SETUP_GUIDE.md +++ b/docs/guides/LINEAR_SETUP_GUIDE.md @@ -11,7 +11,7 @@ Set up the ABCA Linear integration so that applying a label to a Linear issue tr ## How it works -You create a Linear OAuth app and authorize it on your workspace. When someone adds the trigger label to an issue in a mapped project, Linear fires a webhook at ABCA; the receiver verifies the HMAC signature, looks up the workspace, resolves a Linear API token, and creates a task. The agent clones the repo, makes the change, opens a PR, and comments back on the issue. +You create a Linear OAuth app and authorize it on your workspace. When someone adds the trigger label to an issue in a mapped project, Linear sends a webhook to ABCA. The receiver verifies its signature and invokes a processor, which resolves the workspace token and creates a task. The agent clones the repo, makes the change, opens a PR, and reports back on the issue. The app is installed with `actor=app`, so everything ABCA writes is attributed to the app rather than to whoever clicked Authorize. @@ -21,10 +21,10 @@ One of two places, chosen automatically at setup time: | | When it's used | What's stored | |---|---|---| -| **AgentCore Identity vault** | The stack was deployed with `--context enableLinearIdentityVault=true` | Nothing long-lived. AgentCore holds the refresh token and mints short-lived access tokens on demand. | +| **AgentCore Identity vault** | The stack was deployed with `--context enableLinearIdentityVault=true` | AgentCore manages the OAuth grant and token refresh. ABCA retains the OAuth client credentials and workspace/webhook metadata in Secrets Manager. | | **Secrets Manager** | Otherwise — including regions where AgentCore Identity isn't available | An OAuth token bundle in `bgagent-linear-oauth-`, refreshed and rotated by ABCA. | -The vault is unavailable on the `lambda-microvm` substrate — see [Not available with `compute_type=lambda-microvm`](#not-available-with-compute_typelambda-microvm) below. +The vault can be enabled with AgentCore, ECS or Lambda MicroVM compute. `bgagent linear setup` picks whichever the deployment supports and tells you which one it used. There is no flag. If the vault isn't available it prints one line and continues on Secrets Manager: @@ -34,11 +34,16 @@ AgentCore Identity not available in us-east-1 — using Secrets Manager. A workspace that started on Secrets Manager and later moves to the vault **keeps** its Secrets Manager token as a fallback. A workspace onboarded straight onto the vault has no such token by design — it needs the vault to be reachable. -When a workspace's authorization dies, ABCA records it on the registry row and publishes to the stack's operational alert topic. That topic has **no subscribers unless you deployed with `alertEmail`**, so set it if you want to hear about a dead workspace rather than discover it from `bgagent platform doctor`. +When a workspace's authorization dies, ABCA records it on the registry row and publishes to the stack's operational alert topic. A new topic starts without subscribers; configure `alertEmail` or subscribe another destination to receive those alerts. A revoked legacy Secrets Manager fallback does not by itself mean the active vault grant is revoked. -#### Not available with `compute_type=lambda-microvm` +#### Using the vault with Lambda MicroVMs -The vault and the Lambda MicroVMs substrate cannot be enabled on the same stack. Together they synthesize 505 CloudFormation resources against a hard limit of 500 (MicroVM alone is 496, the vault alone 488), so `cdk deploy` refuses the combination by name at synth rather than failing partway through. Use the vault on the `agentcore` or `ecs` substrate; a MicroVM stack stays on Secrets Manager until the stack reclaims room. +Deploy with `compute_type=lambda-microvm`, `enableLinearIdentityVault=true` and an explicit `microvm_nested_stack` value (`true` for new/already-nested installations; retain `false` for existing flat installations until migration). The coordinator sends the workload identity name through authenticated `platform_config`; the guest uses its compute execution role to obtain a Linear token. Credentials are not baked into the MicroVM image. + +When upgrading an existing MicroVM deployment, rebuild the guest image too: +the coordinator and guest must both support the vault configuration fields. + +The runtime resolves the vault in its AWS Region. A grant in another Region or under another workload identity does not automatically carry over. The old resource-count guard ([#857](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/857)) has been replaced with configuration, permission and deployment-budget checks. #### One workload identity per stack diff --git a/docs/guides/QUICK_START.mdx b/docs/guides/QUICK_START.mdx index 6322efc73..8b01d97a8 100644 --- a/docs/guides/QUICK_START.mdx +++ b/docs/guides/QUICK_START.mdx @@ -480,7 +480,7 @@ node lib/bin/bgagent.js deny \ --reason "Don't force-push shared branches; open a revert PR instead" ``` -The task transitions back to `RUNNING` immediately on a decision. The denial reason is injected into the agent's context so it can adapt rather than retry the same tool call. If no decision arrives within the rule's timeout (300 s by default), the gate is treated as a denial with `timed_out` as the reason. +A decision lets the task continue; a sleeping or retired MicroVM first wakes or restores its saved work. The denial reason is passed to the agent so it can adapt. Requests have no decision deadline by default. If an explicit task or rule deadline expires, the gate is treated as a denial with `timed_out` as the reason. Worker lifetime limits remain separate. If you want a task to run without interactive gates (e.g. an unattended overnight job), pre-approve the scopes you trust up-front: diff --git a/docs/guides/USER_GUIDE.md b/docs/guides/USER_GUIDE.md index 7d55b39e5..4128b7339 100644 --- a/docs/guides/USER_GUIDE.md +++ b/docs/guides/USER_GUIDE.md @@ -638,7 +638,8 @@ Created: 2026-04-01T00:39:51.271Z | `--max-budget` | Maximum cost budget in USD (0.01–100). Overrides per-repo Blueprint default. No default limit. | | `--idempotency-key` | Idempotency key for deduplication. | | `--trace` | Enable detailed tracing: raises progress preview cap to 4 KB and uploads full NDJSON trajectory to S3 on completion. Download with `bgagent trace download`. | -| `--approval-timeout` | Cedar HITL per-task approval timeout in seconds (default 300). A matching rule with its own `@approval_timeout_s` annotation still takes the minimum. See [Approval gates](#approval-gates-cedar-hitl). | +| `--approval-timeout` | Cedar HITL decision window: `0` (default) keeps unanswered requests available; a positive value sets a 30–3,600 second deadline. A matching rule can require a shorter positive deadline. See [Approval gates](#approval-gates-cedar-hitl). | +| `--microvm-sleep-after` | Seconds to wait for approval before putting a Lambda MicroVM to sleep (default 600 = 10 minutes; 0–3600 accepted). Use `off` to keep it awake. Requires the deployment's automatic-sleep feature to be enabled; does not change approval deadlines or affect other compute backends. | | `--pre-approve` | Cedar HITL scope to approve up-front (repeatable). Same scope forms as `bgagent approve --scope`. Hard-deny rules are always enforced. | | `--wait` | Poll until the task reaches a terminal status. | | `--output` | Output format: `text` (default) or `json`. | @@ -841,10 +842,27 @@ When a rule marked `@tier("soft")` matches a tool call: 1. The agent stops before invoking the tool. 2. A row is atomically written to the approvals table and the task status flips to `AWAITING_APPROVAL`. 3. A progress event (`approval_requested`) is emitted so `bgagent watch` shows the gate in real time. -4. The task waits for your decision up to the rule's timeout (default 300 s, configurable per-rule and per-task). +4. The task waits for your decision without an automatic deadline by default. An explicit per-task or policy-rule timeout can limit that window. 5. On approval, the agent proceeds; on denial, the deny reason is best-effort injected back into the agent's context so it can adapt; on timeout, the gate is treated as a denial with `timed_out` as the reason. -A decision is recorded at most once per request. Replaying approve/deny on the same `(task_id, request_id)` is idempotent. +A decision is recorded at most once per request. A repeated decision cannot change an already closed request. + +### Responding in Linear + +For a Linear task, the bot posts an **Approval needed** comment with the action +and reason. Reply **approve** or **deny** in that comment's thread. You do not +need a task ID, request ID, or bot mention. Your Linear account must be linked to +the ABCA account that submitted the task. + +`approve` allows the displayed action once. The bot acknowledges the saved +decision. An old thread cannot approve a newer request; replies to closed requests +explain that no new decision was recorded. Use the CLI for broader approval scopes +or a denial reason. Editing an existing comment does not submit a decision. + +This works on every compute backend. A sleeping MicroVM wakes to receive the +decision; AgentCore and ECS receive it through their existing approval wait. +Sleep does not set the approval deadline. Requests have no automatic deadline by +default, and an explicitly configured deadline still applies. ### Listing pending approvals @@ -852,7 +870,7 @@ A decision is recorded at most once per request. Replaying approve/deny on the s node lib/bin/bgagent.js pending ``` -Lists every approval across your tasks that is currently awaiting your decision. The default text output gives you the `request_id`, tool, severity, the reason the rule matched, the tool-input preview, the expiry time, and ready-to-run `approve` / `deny` command lines. Pipe through `--output json` for scripting. +Lists every approval across your tasks that is currently awaiting your decision. The default text output gives you the `request_id`, tool, severity, the reason the rule matched, the tool-input preview, the deadline or “no automatic expiry,” and ready-to-run `approve` / `deny` command lines. The JSON `expires_at` is `null` when there is no deadline. Cancelled or completed tasks no longer have answerable requests. ```text 1 pending approval(s): @@ -864,7 +882,7 @@ Lists every approval across your tasks that is currently awaiting your decision. rules: force_push_any preview: git push --force origin feature/xyz created: 2026-05-13T12:04:12Z - expires: 2026-05-13T12:09:12Z (timeout_s=300) + expires: no automatic expiry approve: bgagent approve 01KN37PZ77P1W19D71DTZ15X6X 01R... deny: bgagent deny 01KN37PZ77P1W19D71DTZ15X6X 01R... --reason "..." ``` @@ -924,13 +942,33 @@ node lib/bin/bgagent.js submit --repo owner/repo --issue 42 \ --pre-approve tool_type:Bash \ --pre-approve write_path:tests/** -# Per-task timeout override (platform default is 300s) +# Optional ten-minute decision deadline (the default has no deadline) node lib/bin/bgagent.js submit --repo owner/repo --issue 42 --approval-timeout 600 ``` `--pre-approve` can be repeated up to the platform limit (see `bgagent submit --help` for the current cap). Valid scope forms are the same as the `approve --scope` table above. Hard-deny rules are still enforced — `--pre-approve` only short-circuits soft-deny rules. -`--approval-timeout` sets the task-wide default; a rule with its own `@approval_timeout_s` annotation still takes the minimum of the two. +`--approval-timeout 0` keeps unanswered requests available. A positive setting +limits the decision window; the shortest positive deadline from the task and +matching policy rules wins. Zero does not disable a policy rule's explicit +deadline. Cancelling the task closes its requests. + +For Lambda MicroVM tasks, `--microvm-sleep-after 600` selects the default +10-minute delay; `--microvm-sleep-after 120` selects two minutes and +`--microvm-sleep-after off` keeps the worker awake. The delay starts when each +approval request is created. Waking for approval, denial, or an approaching +deadline remains automatic. Sleep never starts a new approval timer. + +An unanswered request can outlive its MicroVM. After a longer wait, ABCA saves +the workspace and conversation, stops the worker, and releases its capacity. +Your answer can then start a replacement when capacity is available. A request +remaining open does not mean its old computer must stay alive. Sleeping saves +compute charges but adds snapshot save/restore charges and wake-up time; short +pauses can cost more than staying awake. The API equivalent is +`microvm_sleep_after_s` (zero means off); task +details return the saved setting. Automatic suspension is disabled by default for +new deployments. An operator enables it after [verifying the deployed image and +coordinator](../verification/README.md#live-acceptance-for-an-installation). ## Webhook integration diff --git a/docs/mise.toml b/docs/mise.toml index 191fab83e..2a6639dc5 100644 --- a/docs/mise.toml +++ b/docs/mise.toml @@ -8,13 +8,13 @@ description = "Install from monorepo root (Yarn workspaces)" run = "cd .. && yarn install --check-files" [tasks.sync] -description = "Sync docs content from guides/design/CONTRIBUTING" +description = "Sync docs content from guides/design/decisions/verification/CONTRIBUTING" run = "node scripts/sync-starlight.mjs" [tasks.build] description = "Starlight docs site build" depends = [":sync"] -run = "./node_modules/.bin/astro build" +run = "./node_modules/.bin/astro build && node scripts/check-site-links.mjs" [tasks.dev] description = "Starlight dev server" @@ -26,7 +26,7 @@ depends = [":sync"] run = "./node_modules/.bin/astro sync && ./node_modules/.bin/astro check" [tasks.link-check] -# Scope: docs/guides/, docs/design/, docs/decisions/, and root *.md files. +# Scope: docs/guides/, docs/design/, docs/decisions/, docs/verification/, and root *.md files. # Not scanned: docs/README.md, agent/, cli/, contracts/, docs/abca-plugin/. description = "Check for broken links in Markdown sources" run = "bash scripts/link-check.sh" diff --git a/docs/package.json b/docs/package.json index 2e24f0e8d..64e13c7ef 100644 --- a/docs/package.json +++ b/docs/package.json @@ -5,7 +5,7 @@ "dev": "astro dev", "start": "astro dev", "docs:check": "npm run sync && astro check", - "docs:build": "npm run sync && astro build", + "docs:build": "npm run sync && astro build && node scripts/check-site-links.mjs", "build": "npm run docs:build", "preview": "astro preview", "security:retire": "retire --path . --severity high", diff --git a/docs/scripts/check-site-links.mjs b/docs/scripts/check-site-links.mjs new file mode 100644 index 000000000..6d1ea7b11 --- /dev/null +++ b/docs/scripts/check-site-links.mjs @@ -0,0 +1,50 @@ +import fs from 'node:fs'; +import path from 'node:path'; + +const root = path.resolve(import.meta.dirname, '..', 'dist'); +const repoRoot = path.resolve(import.meta.dirname, '..', '..'); +const repoBlobPrefix = 'https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/'; +const base = '/sample-autonomous-cloud-coding-agents/'; +const pages = fs.readdirSync(root, { recursive: true }).filter(file => file.endsWith('.html')); +if (pages.length < 10) throw new Error('Expected a built documentation site before checking links'); + +const errors = new Set(); +const contents = new Map(); +function read(file) { + if (!contents.has(file)) contents.set(file, fs.readFileSync(file, 'utf8')); + return contents.get(file); +} + +for (const file of pages) { + const origin = `https://local${base}${file.replace(/index\.html$/, '')}`; + for (const match of read(path.join(root, file)).matchAll(/]*\bhref="([^"]*)"/g)) { + const href = match[1].replaceAll('&', '&'); + if (href.startsWith(repoBlobPrefix)) { + const relative = decodeURIComponent(href.slice(repoBlobPrefix.length).split(/[?#]/)[0]); + const target = path.resolve(repoRoot, relative); + if (!target.startsWith(`${repoRoot}${path.sep}`) || !fs.existsSync(target) || !fs.statSync(target).isFile()) { + errors.add(`${file}: missing repository file: ${href}`); + } + continue; + } + // This check is offline. External availability is outside its scope. + if (/^(?:[a-z]+:|\/\/)/i.test(href)) continue; + const url = new URL(href, origin); + if (!url.pathname.startsWith(base)) { + errors.add(`${file}: link escapes the site's base: ${href}`); + continue; + } + let target = path.join(root, decodeURIComponent(url.pathname.slice(base.length))); + if (fs.existsSync(target) && fs.statSync(target).isDirectory()) { + target = path.join(target, 'index.html'); + } + if (!fs.existsSync(target)) { + errors.add(`${file}: missing page: ${href}`); + } else if (url.hash && !read(target).includes(`id="${decodeURIComponent(url.hash.slice(1))}"`)) { + errors.add(`${file}: missing anchor: ${href}`); + } + } +} + +if (errors.size) throw new Error(`Broken documentation links:\n${[...errors].join('\n')}`); +console.log(`Checked internal page and anchor links in ${pages.length} rendered pages.`); diff --git a/docs/scripts/link-check.sh b/docs/scripts/link-check.sh index 23f3632a8..0ca170cef 100644 --- a/docs/scripts/link-check.sh +++ b/docs/scripts/link-check.sh @@ -4,7 +4,7 @@ set -euo pipefail filelist=$(mktemp) trap 'rm -f "$filelist"' EXIT -{ find guides design decisions -name '*.md' -print0; find .. -maxdepth 1 -name '*.md' -print0; } > "$filelist" +{ find guides design decisions verification -name '*.md' -print0; find .. -maxdepth 1 -name '*.md' -print0; } > "$filelist" count=$(tr '\0' '\n' < "$filelist" | grep -c .) if [ "$count" -lt 10 ]; then diff --git a/docs/scripts/sync-starlight.mjs b/docs/scripts/sync-starlight.mjs index 037180aee..dd895329f 100644 --- a/docs/scripts/sync-starlight.mjs +++ b/docs/scripts/sync-starlight.mjs @@ -8,7 +8,7 @@ const docsBase = '/sample-autonomous-cloud-coding-agents'; function normalizeFileStem(input) { const cleaned = input - .replace(/\.md$/i, '') + .replace(/\.mdx?$/i, '') .replace(/[^a-zA-Z0-9]+/g, '-') .replace(/^-+|-+$/g, '') .toLowerCase(); @@ -18,7 +18,15 @@ function normalizeFileStem(input) { return `${cleaned.charAt(0).toUpperCase()}${cleaned.slice(1)}`; } -function rewriteDocsLinkTarget(target) { +function rewriteDocsLinkTarget(target, sourcePath) { + if (target?.startsWith('#') && path.basename(sourcePath) === 'USER_GUIDE.md') { + const route = { + '#joining-an-existing-deployment': '/using/authentication#joining-an-existing-deployment', + '#get-stack-outputs': '/using/authentication#get-stack-outputs', + '#approval-gates-cedar-hitl': '/using/approval-gates-cedar-hitl', + }[target]; + if (route) return route; + } if (!target || target.startsWith('#') || target.startsWith('/')) { return undefined; } @@ -27,12 +35,28 @@ function rewriteDocsLinkTarget(target) { } const [pathPart, anchor] = target.split('#'); - if (!pathPart.toLowerCase().endsWith('.md')) { - return undefined; - } const normalizedPath = pathPart.replaceAll('\\', '/'); - const stem = path.basename(normalizedPath, '.md'); + // Resolve relative links in the source tree before assigning a site route. + // In particular, ./README.md inside verification is not architecture/readme. + const sourceTarget = path.resolve(path.dirname(sourcePath), normalizedPath); + if (sourceTarget === path.join(targetRoot, 'index.md')) return '/'; + if (sourceTarget === path.join(docsRoot, 'decisions')) return '/decisions/readme'; + const relativeSource = path.relative(repoRoot, sourceTarget).replaceAll('\\', '/'); + if (!relativeSource.startsWith('../') && fs.existsSync(sourceTarget) + && (!sourceTarget.startsWith(docsRoot + path.sep) || sourceTarget === path.join(docsRoot, 'README.md') + || sourceTarget.startsWith(path.join(docsRoot, 'abca-plugin') + path.sep))) { + const kind = fs.statSync(sourceTarget).isDirectory() ? 'tree' : 'blob'; + return `https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/${kind}/main/${relativeSource}${anchor ? `#${anchor}` : ''}`; + } + if (!/\.mdx?$/i.test(pathPart)) return undefined; + for (const directory of ['verification', 'decisions']) { + if (path.dirname(sourceTarget) === path.join(docsRoot, directory)) { + const slug = normalizeFileStem(path.basename(sourceTarget)).toLowerCase(); + return `/${directory}/${slug}${anchor ? `#${anchor}` : ''}`; + } + } + const stem = path.basename(normalizedPath).replace(/\.mdx?$/i, ''); const slug = normalizeFileStem(stem).toLowerCase(); const anchorSuffix = anchor ? `#${anchor}` : ''; @@ -71,6 +95,9 @@ function rewriteDocsLinkTarget(target) { const userGuideAnchorRoutes = { overview: '/using/overview', authentication: '/using/authentication', + 'joining-an-existing-deployment': '/using/authentication#joining-an-existing-deployment', + 'get-stack-outputs': '/using/authentication#get-stack-outputs', + 'operator-commands-stack-admin': '/using/using-the-cli#operator-commands-stack-admin', 'repository-onboarding': '/customizing/repository-onboarding', 'per-repo-overrides': '/customizing/per-repo-overrides', 'monthly-user-and-team-budgets': '/customizing/per-repo-overrides#monthly-user-and-team-budgets', @@ -80,6 +107,7 @@ function rewriteDocsLinkTarget(target) { 'webhook-integration': '/using/webhook-integration', 'task-lifecycle': '/using/task-lifecycle', 'what-the-agent-does': '/using/what-the-agent-does', + 'approval-gates-cedar-hitl': '/using/approval-gates-cedar-hitl', 'tips-for-being-a-good-citizen': '/using/tips-for-being-a-good-citizen', }; if (stem === 'USER_GUIDE' && anchor) { @@ -99,11 +127,11 @@ function rewriteDocsLinkTarget(target) { return `/architecture/${slug}${anchorSuffix}`; } -function ensureFrontmatter(content, title) { +function ensureFrontmatter(content, title, sourcePath) { const normalized = content .replaceAll('../imgs/', `${docsBase}/imgs/`) .replace(/\[([^\]]+)\]\(([^)]+)\)/g, (match, label, target) => { - const rewritten = rewriteDocsLinkTarget(target); + const rewritten = rewriteDocsLinkTarget(target, sourcePath); if (!rewritten) { return match; } @@ -111,12 +139,8 @@ function ensureFrontmatter(content, title) { // must carry that prefix — otherwise they resolve to the domain root // and 404. Starlight prefixes its own nav links automatically, but our // rewritten body links are raw markdown and need it added explicitly - // (same reason the image rewrites above include docsBase). Every - // non-undefined return from rewriteDocsLinkTarget is a `/…` route (bare - // `#…` anchors and external links return undefined and keep their - // original text above), so the prefix always applies. - // (Fixes the broken in-body design-doc links.) - return `[${label}](${docsBase}${rewritten})`; + // Repository-only documents instead retain their absolute GitHub URL. + return `[${label}](${rewritten.startsWith('/') ? docsBase : ''}${rewritten})`; }); const trimmed = normalized.trimStart(); @@ -152,7 +176,7 @@ function mirrorMarkdownFile(sourcePath, targetRelativePath) { const ext = path.extname(sourcePath); const stem = path.basename(sourcePath, ext); const fallbackTitle = normalizeFileStem(stem).replace(/-/g, ' '); - const out = ensureFrontmatter(raw, fallbackTitle); + const out = ensureFrontmatter(raw, fallbackTitle, sourcePath); writeFile(path.join(docsRoot, targetRelativePath), out); } @@ -168,7 +192,7 @@ function mirrorDirectory(sourceDir, targetDirRelative) { const sourcePath = path.join(sourceDir, file); const raw = fs.readFileSync(sourcePath, 'utf8'); const fallbackTitle = normalizeFileStem(file).replace(/-/g, ' '); - const out = ensureFrontmatter(raw, fallbackTitle); + const out = ensureFrontmatter(raw, fallbackTitle, sourcePath); const normalizedName = `${normalizeFileStem(file)}.md`; writeFile(path.join(docsRoot, targetDirRelative, normalizedName), out); } @@ -200,7 +224,7 @@ function splitGuide(sourcePath, targetDirRelative, introTitle) { const raw = fs.readFileSync(sourcePath, 'utf8'); const parts = raw.split(/\n##\s+/g); const intro = parts.shift() ?? ''; - const introOut = ensureFrontmatter(intro.trim(), introTitle); + const introOut = ensureFrontmatter(intro.trim(), introTitle, sourcePath); writeFile(path.join(docsRoot, targetDirRelative, 'Introduction.md'), introOut); for (const part of parts) { @@ -208,7 +232,7 @@ function splitGuide(sourcePath, targetDirRelative, introTitle) { const heading = (firstNewline === -1 ? part : part.slice(0, firstNewline)).trim(); const body = firstNewline === -1 ? '' : part.slice(firstNewline + 1).trim(); const filename = `${normalizeFileStem(heading)}.md`; - const out = ensureFrontmatter(body, heading); + const out = ensureFrontmatter(body, heading, sourcePath); writeFile(path.join(docsRoot, targetDirRelative, filename), out); } } @@ -316,6 +340,9 @@ mirrorDirectory(path.join(docsRoot, 'design'), path.join('src', 'content', 'docs // --- Decision records (ADRs): mirror to decisions/ --- mirrorDirectory(path.join(docsRoot, 'decisions'), path.join('src', 'content', 'docs', 'decisions')); +// Verification runbooks ship with the same revision as their referring pages. +mirrorDirectory(path.join(docsRoot, 'verification'), path.join('src', 'content', 'docs', 'verification')); + // --- Static assets: copy source image dir into the site's public/ --- // Guides reference images as `../imgs/foo.png`; ensureFrontmatter() turns // those into absolute `//imgs/foo.png` URLs, which Astro serves from diff --git a/docs/src/content/docs/architecture/Api-contract.md b/docs/src/content/docs/architecture/Api-contract.md index b8adba636..dc3e71aa4 100644 --- a/docs/src/content/docs/architecture/Api-contract.md +++ b/docs/src/content/docs/architecture/Api-contract.md @@ -258,7 +258,7 @@ Returns full details of a task. Users can only access their own tasks. } ``` -`agent_heartbeat_at` is the agent's last in-guest liveness beat, or `null`. The agent writes it every 45 s on the `agentcore` and `lambda-microvm` backends; on `ecs` it is written once at start, because that backend runs the pipeline directly instead of serving HTTP. The orchestrator reads the same field to detect a hung agent inside a healthy compute environment ([ORCHESTRATOR.md](/sample-autonomous-cloud-coding-agents/architecture/orchestrator#dynamodb-heartbeat-agentcore-and-lambda-microvms)), so a value that is minutes old on a `RUNNING` task is the signal, not the timestamp itself. `null` on records written before the field existed. +`agent_heartbeat_at` is the agent's last in-guest liveness beat, or `null`. The agent writes it every 45 s on the `agentcore` and `lambda-microvm` backends; on `ecs` it is written once at start, because that backend runs the pipeline directly instead of serving HTTP. The orchestrator reads the same field to detect a hung agent inside a healthy compute environment ([ORCHESTRATOR.md](/sample-autonomous-cloud-coding-agents/architecture/orchestrator#liveness-monitoring)), so a value that is minutes old on a `RUNNING` task is the signal, not the timestamp itself. `null` on records written before the field existed. `error_classification` is a derived field computed at response time from `error_message`. When `error_message` is `null`, `error_classification` is `null`. When present, it contains: @@ -386,7 +386,7 @@ When a task pauses in `AWAITING_APPROVAL` (Cedar soft-deny gate), the owner appr { "data": { "task_id": "01HYX...", "request_id": "...", "status": "DENIED", "decided_at": "2025-03-15T10:35:00Z" } } ``` -**Errors:** `400 VALIDATION_ERROR`, `401 UNAUTHORIZED`, `404 REQUEST_NOT_FOUND` (collapses "row missing" and "wrong caller"), `409 REQUEST_ALREADY_DECIDED`, `409 TASK_NOT_AWAITING_APPROVAL`. +**Errors:** `400 VALIDATION_ERROR`, `401 UNAUTHORIZED`, `404 REQUEST_NOT_FOUND` (collapses missing, inaccessible, closed or expired approval rows), `409 TASK_NOT_AWAITING_APPROVAL` (task-only state conflict). ### List pending approvals @@ -612,7 +612,7 @@ There is no per-user request-rate or "tasks-per-hour" limiter on task creation. | `RATE_LIMIT_EXCEEDED` | 429 | Rate/concurrency gate exceeded — per-task nudge limit, the application rate limiter on approval endpoints, or the user concurrency limit on confirm-uploads | | `BUDGET_EXCEEDED` | 429 | A configured user or Cognito-team monthly budget reached 100% with hard stop enabled | | `REQUEST_NOT_FOUND` | 404 | Cedar HITL approval request not found (also returned when the caller does not own it) | -| `REQUEST_ALREADY_DECIDED` | 409 | Cedar HITL approval request was already approved or denied | +| `REQUEST_ALREADY_DECIDED` | 409 | Legacy error-code enum; current approve/deny handlers return `404 REQUEST_NOT_FOUND` for closed or inaccessible approval rows | | `TASK_NOT_AWAITING_APPROVAL` | 409 | Task is not in `AWAITING_APPROVAL`, so the approval decision does not apply | | `INTERNAL_ERROR` | 500 | Unexpected server error | | `SERVICE_UNAVAILABLE` | 503 | Downstream dependency unavailable (retry with backoff) | diff --git a/docs/src/content/docs/architecture/Attachments.md b/docs/src/content/docs/architecture/Attachments.md index e92b64771..730ec1d12 100644 --- a/docs/src/content/docs/architecture/Attachments.md +++ b/docs/src/content/docs/architecture/Attachments.md @@ -1294,13 +1294,13 @@ The task record is preserved with status `CANCELLED` — `bgagent status **Status:** Core implemented; this document remains the authoritative design reference. +> **Status:** Core implemented. The September update below supersedes the original bounded-wait assumptions in the historical design and examples. > **Companion:** [`INTERACTIVE_AGENTS.md`](/sample-autonomous-cloud-coding-agents/architecture/interactive-agents) §9.3 (pointing here), §7 (state machine). > **Design locked:** 2026-04-23 (Sam ↔ assistant discussion). > **Rev:** 5 (2026-05-06 — fold in parallel adversarial + advocate review of the timeout design: late-approval re-read on TIMED_OUT ConditionCheckFailed; user-visible timeout-cap milestones; ceiling-shrink milestone; Runtime JWT bound verified as auto-refreshed IAM; three new tuning metrics; explicit off-hours trade-off section; notification-delivery-failure boundary. IMPL-24 through IMPL-28 added.). > **Implementation:** Core shipped. The 3-outcome engine (`agent/src/policy.py`), default policy sets (`agent/policies/hard_deny.cedar`, `agent/policies/soft_deny.cedar`), approval Lambdas (`cdk/src/handlers/{approve-task,deny-task,get-pending,get-policies}.ts`) wired into `cdk/src/constructs/task-api.ts` (routes `/tasks/{id}/approve`, `/deny`, `/pending`, `/repos/{repo_id}/policies`), the cross-engine parity fixtures (`contracts/cedar-parity/`), and the exact engine pins are all on `main`. §15's task list is preserved as a historical implementation record; see the note at the top of §15 for what (if anything) remains unbuilt. +> +> **Current source behavior (2026-09-22):** the task default is `approval_timeout_s=0`, +> meaning no decision deadline. Workers create/close requests through the IAM-authenticated +> [trusted approval writer](/sample-autonomous-cloud-coding-agents/decisions/adr-023-trusted-approval-writer); they cannot +> write approval rows directly. Explicit task settings are 30–3,600 seconds; +> positive policy-rule deadlines still apply. Pending rows have no DynamoDB TTL, +> and `expires_at` is nullable. Task closure cancels unanswered requests and adds +> retention TTL without changing already-recorded decisions. Closed approval-row +> conditions return `404 REQUEST_NOT_FOUND`; task-only conflicts return 409. +> MicroVM checkpoint/retirement/replacement separates human waiting from worker +> lifetime and capacity; other backends retain their existing runtime limits. +> Linear accepts an owner’s `approve` or `deny` reply to the approval comment; +> the handler verifies the actual comment through Linear’s API. Slack uses CLI +> response instructions. See the current +> [user guide](/sample-autonomous-cloud-coding-agents/using/approval-gates-cedar-hitl) and +> [continuation protocol](/sample-autonomous-cloud-coding-agents/architecture/orchestrator#retained-microvm-approvals). +> The normal deployment passed retained-request, ten-minute sleep, explicit-expiry +> and sleep-off/rollback acceptance; see the +> [deployment record](/sample-autonomous-cloud-coding-agents/verification/readme). --- @@ -16,7 +35,7 @@ title: Cedar hitl gates 1. [What we are building, in one paragraph](#1-what-we-are-building-in-one-paragraph) 2. [The three-outcome model and why Cedar alone can't give it](#2-the-three-outcome-model) -3. [Design decisions (locked)](#3-design-decisions-locked) +3. [Design decisions](#3-design-decisions) 4. [End-to-end request flow](#4-end-to-end-request-flow) 5. [Cedar policy authoring guide](#5-cedar-policy-authoring-guide) 6. [Engine implementation](#6-engine-implementation) @@ -113,19 +132,19 @@ The winning property: **policy authors can put on their "security-review-approve --- -## 3. Design decisions (locked) +## 3. Design decisions -Settled during the 2026-04-23 design discussion and extended after the 2026-04-24 and 2026-05-06 reviews. Each has detailed rationale in those conversations; summary here for implementers. **23 decisions**, all locked unless an adversarial review finding explicitly reopened a concern. +Settled during the 2026-04-23 design discussion and extended after the 2026-04-24 and 2026-05-06 reviews. Each has detailed rationale in those conversations; summary here for implementers. The September retained-request update also amends deadline, recovery and capacity behavior. | # | Decision | Summary | |---|---|---| | 1 | **Cedar encoding: two policy sets** | Physical hard-deny vs soft-deny split, validated via `@tier(...)` annotation. | | 2 | **Hook point: extend `PreToolUse`, not `can_use_tool`** | PreToolUse is already async-compatible, already wired to Cedar, and already owns the tool-governance boundary. | | 3 | **Wait mechanism: DDB strongly-consistent polling, 2s → 5s backoff** | Initial 2s cadence for the first 30s, then 5s. `ConsistentRead=True` so the agent never misses an approval that already landed. | -| 4 | **Scope allowlist: in-process, seeded from persisted `initial_approvals`** | Runtime escalation lives in the `PolicyEngine` instance. Submit-time `--pre-approve` flags persist on TaskTable and seed the allowlist at container startup. Lost on restart (rare; reconciler fails stranded tasks). | +| 4 | **Scope allowlist: in-process, seeded from persisted `initial_approvals`** | Runtime grants live in the `PolicyEngine` instance. Submit-time grants seed it at startup; a verified MicroVM checkpoint also saves session grant scopes and denial-cache entries for replacement. | | 5 | **CLI UX: standalone `bgagent approve/deny` + `--pre-approve ` + `bgagent policies list` + `bgagent pending`** | No inline interactive prompt in the streaming CLI for v1. Discovery + listing commands solve the request_id/rule_id copy problem. | -| 6 | **Timeouts: per-task default + per-rule Cedar annotation override, min wins, bounded floor + ceiling, fail-closed** | Per-task default: **300s** (5 min), overridable via `--approval-timeout` on submit and bounded by `[30, min(3600, maxLifetime - 300)]`. Floor: 30s (engine-enforced on both task default and rule annotations). Ceiling: `min(1h, maxLifetime_remaining - cleanup_margin)` — sized so the TTL on the approval row always covers the decision window. On timeout → deny (never auto-approve). See §14.8 for the off-hours trade-off this posture deliberately accepts. | -| 7 | **Concurrency slots: AWAITING_APPROVAL holds the slot** | Matches PAUSED semantics. Container is alive, consuming memory. | +| 6 | **Timeouts: per-task default + per-rule Cedar annotation override, min wins, bounded floor + ceiling, fail-closed** | Per-task default: **0**, meaning no deadline. Explicit task deadlines are 30–3,600 seconds; positive matching rule deadlines can shorten them. Rule annotations still require at least 30 seconds. Workers without checkpoint/replacement support retain their existing lifetime limit. Pending rows have no storage TTL. Explicit timeout means deny, never auto-approve. | +| 7 | **Concurrency slots: AWAITING_APPROVAL holds the slot** | Bounds unfinished sessions and their eventual resume demand, including a suspended MicroVM. This is an ABCA admission policy; suspended AWS memory-quota consumption remains unverified. | | 8 | **Hard-deny is absolute** | No `--pre-approve` scope, and no blueprint `disable:` directive, can bypass it. CreateTaskFn validates and rejects `rule:`; blueprint loader rejects `disable:` entries that name built-in hard-deny rules. | | 9 | **Submit-time scope cap: 20 entries, ≤128 chars each** | Keeps audit trail legible, bounds allowlist check cost, limits abuse-vector damage. | | 10 | **Cedar annotations (verified working)** | `@rule_id(...)`, `@tier(...)`, `@approval_timeout_s(...)`, `@severity(...)`, `@category(...)`. Recoverable via `cedarpy.policies_to_json_str()` → JSON. Multi-match merging: min timeout wins (clamped by floor), max severity wins. | @@ -141,13 +160,13 @@ Settled during the 2026-04-23 design discussion and extended after the 2026-04-2 | 20 | **`write_path:` scope** | Added so users can pre-approve file writes under specific path patterns (e.g., `write_path:docs/**`) without needing to grant all Writes. Validation uses Python `fnmatch` at runtime; glob semantics are a Cedar-`like` superset (§6.4, §5.5). | | 21 | **`tool_group:file_write` convenience scope** | Resolves to `{Write, Edit}`. Prevents the surprise of pre-approving `Write` and still getting gated on `Edit`. | | 22 | **Pre-implementation spike: cedarpy annotation round-trip** | Day 1 of implementation validates that `policies_to_json_str()` returns annotations in the expected shape. If the API has changed, fall back to policy-ID prefix conventions. | -| 23 | **Cedar engine parity contract (Python `cedarpy` ↔ JS `cedar-wasm`)** | Both engines are pinned in `mise.toml`. A golden-file parity test runs in CI: for each `(policy, input)` fixture the test asserts Python and WASM return the same `decision` and the same set of matching rule IDs. Policy authors who upgrade either engine must refresh the golden file; drift fails the build. See §15.6 and Appendix B. | +| 23 | **Cedar engine parity contract (Python `cedarpy` ↔ JS `cedar-wasm`)** | Both engines are pinned in their package manifests. A golden-file parity test runs in CI: for each `(policy, input)` fixture the test asserts Python and WASM return the same `decision` and the same set of matching rule IDs. Policy authors who upgrade either engine must refresh the golden file; drift fails the build. See §15.6 and Appendix B. | --- ## 4. End-to-end request flow -Narrative walk-through of the happy path. Sequence diagrams in the round-trip Mermaid below. +Narrative walk-through with an explicit 600-second task deadline and a custom `force_push_any` policy annotated with `@approval_timeout_s("300")`. Built-in starter rules do not set deadlines. Sequence diagrams are below. ### Setup (task start) @@ -163,7 +182,7 @@ Narrative walk-through of the happy path. Sequence diagrams in the round-trip Me - rejects blueprint whose combined `cedar_policies` text exceeds the 64 KB cap (§12.4) regardless of origin - resolves `approval_gate_cap` = `Blueprint.security.approvalGateCap ?? 50`; rejects if outside `[1, 500]` (decision #13) 5. Task persists. `approval_timeout_s`, `approval_gate_cap`, and `initial_approvals` become DDB attributes on the task row (cap is captured at submit time so mid-task blueprint edits do not shift the cap beneath a running task). -6. Container spawns on Runtime-JWT. `PolicyEngine.__init__` loads: +6. Container starts with IAM runtime credentials. `PolicyEngine.__init__` loads: - `HARD_DENY_POLICIES` (built-in + repo blueprint's `security.cedarPolicies.hard`; blueprint `disable:` may suppress non-built-in rules only, §5.1, §15.4) - `SOFT_DENY_POLICIES` (built-in + repo blueprint's `security.cedarPolicies.soft`; blueprint `disable:` may suppress soft-deny rules freely) - Annotation lookup table: `{policy_id: {annotation: value}}` built from `cedarpy.policies_to_json_str()` once, cached for the task lifetime @@ -197,7 +216,7 @@ Narrative walk-through of the happy path. Sequence diagrams in the round-trip Me ) → effective = 300s ``` - If `maxLifetime_remaining_s - CLEANUP_MARGIN_120S < FLOOR_30S`, hook returns DENY immediately with reason `"insufficient lifetime for approval"` (§13.7). + This lifetime ceiling applies to workers without continuation support. A continuation-capable MicroVM omits it because the approval can outlive the worker. With no positive task/rule deadline, the effective timeout is zero (no decision deadline). 12. Hook checks per-task approval-gate cap (default 50, configurable per blueprint via `security.approvalGateCap`; §5.1) and per-minute rate limit (20/task, per-container). If either exceeded → DENY with reason `"approval-gate cap exceeded"` (fail-closed). 13. Hook mints `request_id = _ulid()` (26-char ULID). @@ -215,15 +234,17 @@ Narrative walk-through of the happy path. Sequence diagrams in the round-trip Me "status": "PENDING", "created_at": "2026-04-23T14:00:00Z", "timeout_s": 300, - "ttl": 1734567890, # created_at + timeout_s + CLEANUP_MARGIN_120S; always covers the decision window + "expires_at": "2026-04-23T14:05:00Z", # explicit decision deadline; no retention TTL "user_id": "...", "repo": "my-org/my-app" } ``` -15. **Atomic transition** — hook issues `TransactWriteItems` with two operations: +15. **Atomic transition** — the hook sends an IAM-signed request to the approval + service, which issues `TransactWriteItems` with these operations (and a + worker-lease condition for MicroVM): - Put on `TaskApprovalsTable` (new row with status=PENDING) - ConditionalUpdate on `TaskTable`: `status = :awaiting, awaiting_approval_request_id = :rid WHERE status = :running` - Both succeed or both fail. On `TransactionCanceledException` (most likely the TaskTable condition fails because another process moved the status), the hook emits `approval_write_failed` and returns DENY. + Both succeed or both fail. If the service rejects a task/lease conflict or the request fails, the hook emits `approval_write_failed` and returns DENY. 16. Hook emits `agent_milestone("approval_requested", {...})` to both `ProgressWriter` (DDB audit) and `sse_adapter` (live stream). Best-effort emission — transactional write has already committed; milestone failure is observability degradation, not state degradation. 17. Terminal A stream renders: ``` @@ -234,36 +255,12 @@ Narrative walk-through of the happy path. Sequence diagrams in the round-trip Me timeout 300s ``` Severity colors the line (respecting `NO_COLOR` env var). -18. Hook enters poll loop with strongly-consistent reads: - ```python - async def _poll_for_decision(task_id, request_id, timeout_s): - start = time.monotonic() - interval = 2 - consecutive_failures = 0 - while True: - elapsed = time.monotonic() - start - if elapsed >= timeout_s: - return TimedOut() - if elapsed > 30: - interval = 5 # backoff - try: - row = await _ddb_get_approval(task_id, request_id, ConsistentRead=True) - consecutive_failures = 0 - if row is None: - # Row disappeared between write and poll — treat as stranded - return TimedOut(reason="approval row missing; fail-closed") - if row["status"] != "PENDING": - return Decided(row) - except Exception as exc: - consecutive_failures += 1 - if consecutive_failures == 3: - log("WARN", f"approval poll degraded for {request_id}: {exc}") - emit_milestone("approval_poll_degraded", {...}) - if consecutive_failures >= 10: - return TimedOut(reason="approval poll consecutive failures") - await asyncio.sleep(interval) - ``` -19. The approval CAP and local-timeout paths ALWAYS attempt to write the row to TIMED_OUT (best-effort conditional update `status = :pending`) before returning. This prevents orphan PENDING rows when the agent bails internally. +18. Hook enters the poll loop with strongly-consistent reads. For an explicit positive timeout, the deadline is captured with the original approval row, before database writes/notifications: UTC expiry is `created_at + timeout_s`, capped by the original monotonic remaining duration. `deadline.remaining_s()` takes the smaller remainder and clamps at zero, so a frozen guest clock or backward UTC correction cannot restart the window. + An untimed request has no expiry. The executable poll loop in + [hooks.py](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/agent/src/hooks.py) also handles cancelled/closed requests, + read failures and continuation barriers; it must not be replaced by a loop + that checks only APPROVED/DENIED or starts a fresh timer after wake. +19. The local-timeout path asks the trusted approval service to conditionally mark the pending row TIMED_OUT before returning. If that write loses or fails, the hook rereads consistently to honor an already-committed decision. The approval-cap check runs before row creation and has no row to update. ### User responds @@ -276,7 +273,7 @@ Narrative walk-through of the happy path. Sequence diagrams in the round-trip Me - ConditionalUpdate on `TaskApprovalsTable`: `#status = :pending AND user_id = :caller AND task_id = :task_id` → flip to APPROVED - ConditionalUpdate on `TaskTable`: `#status = :awaiting AND awaiting_approval_request_id = :rid` → (no-op update, pure state guard; keeps status AWAITING_APPROVAL until the agent's resume transaction flips it RUNNING) Both conditions must hold or the entire transaction is cancelled. No TOCTOU window, no "approved a cancelled task" 202 surprise. - - On `TransactionCanceledException` with per-item `CancellationReasons`: distinguishes between (a) approvals row missing (404 `REQUEST_NOT_FOUND`), (b) approvals row wrong user (404 `REQUEST_NOT_FOUND` — don't leak existence), (c) approvals row wrong status (409 `REQUEST_ALREADY_DECIDED`), (d) task no longer AWAITING_APPROVAL (409 `TASK_NOT_AWAITING_APPROVAL`). + - On `TransactionCanceledException` with per-item `CancellationReasons`: returns 404 `REQUEST_NOT_FOUND` for any approval-row condition failure (missing, foreign-owned or already closed), or 409 `TASK_NOT_AWAITING_APPROVAL` for a task-only condition failure. - Records audit event to TaskEventsTable directly (`approval_decision_recorded`) so the 90-day audit trail is owned by the Lambda, not dependent on agent milestones. - Returns 202 `{task_id, request_id, status: "APPROVED", scope, decided_at}` or error. 24. Agent's poll reads the `APPROVED` row on next tick (within 2-5s). @@ -301,6 +298,7 @@ sequenceDiagram participant Engine as PolicyEngine participant Events as TaskEventsTable participant Approvals as TaskApprovalsTable + participant Requests as Approval request service participant CLI participant User participant Lambda as ApproveTaskFn @@ -309,8 +307,9 @@ sequenceDiagram Agent->>Hook: tool call (Bash git push --force) Hook->>Engine: evaluate_tool_use Engine-->>Hook: REQUIRE_APPROVAL (soft-deny force_push_any) - Hook->>Approvals: TransactWriteItems - Note right of Hook: Put approval row PENDING
plus TaskTable status
to AWAITING_APPROVAL + Hook->>Requests: IAM-signed create for this task + Requests->>Approvals: TransactWriteItems + Note right of Requests: Put approval row PENDING
plus TaskTable status
to AWAITING_APPROVAL Hook->>Events: approval_requested milestone Events-->>CLI: live stream with approval_requested CLI-->>User: bgagent approve TASK REQ @@ -322,7 +321,8 @@ sequenceDiagram Lambda-->>CLI: 202 APPROVED Hook->>Approvals: poll with ConsistentRead Approvals-->>Hook: status APPROVED - Hook->>Approvals: TransactWriteItems, TaskTable to RUNNING + Hook->>Approvals: ConditionCheck on recorded decision + Note right of Hook: Same transaction updates TaskTable to RUNNING;
worker does not modify the approval row Hook->>Engine: allowlist.add(scope) if scope is not this_call Hook-->>Agent: permissionDecision allow Note over Stop: (not used on approval path) @@ -424,7 +424,7 @@ Fail-on-error is the right posture for blueprint misconfiguration — silent-fal |---|---|---|---| | `@rule_id("...")` | **Yes on soft-deny**, recommended on hard-deny | Kebab-case or snake_case identifier, unique across both tiers | Stable ID for `--pre-approve rule:X`, for audit trail, and for the `bgagent policies` discovery endpoint. `PolicyEngine.__init__` raises on duplicates. | | `@tier("hard"\|"soft")` | **Yes** | Exactly one of "hard" or "soft" | Validates policy is in the correct file/section. Engine rejects mismatch at load time. | -| `@approval_timeout_s("N")` | No | Integer seconds ≥ 30 | Per-rule timeout. If absent, uses the task default (**300s** by default, overridable via submit-time `--approval-timeout`; see decision #6). Has no effect on hard-deny rules. Values below the floor are rejected at load time. Values below **120s** emit a blueprint-load WARN but are accepted down to the 30s floor — almost no human responds to an approval request in under 2 minutes, so sub-120s is usually a policy-authoring mistake (see IMPL-25). Loader policy: STRICT at the floor (30s, reject) and ADVISORY below 120s (warn, accept). | +| `@approval_timeout_s("N")` | No | Integer seconds ≥ 30 | Per-rule timeout. If absent, uses the task setting (**0/no deadline** by default, configurable via submit-time `--approval-timeout`; see decision #6). Has no effect on hard-deny rules. Values below the floor are rejected at load time. Values below **120s** emit a blueprint-load WARN but are accepted down to the 30s floor — almost no human responds to an approval request in under 2 minutes, so sub-120s is usually a policy-authoring mistake (see IMPL-25). Loader policy: STRICT at the floor (30s, reject) and ADVISORY below 120s (warn, accept). | | `@severity("low"\|"medium"\|"high")` | No | One of the three | Shown in CLI approval prompt, colored by severity. Default: "medium". | | `@category("...")` | No | "destructive", "network", "filesystem", "auth", or free-form | UX grouping. CLI could filter approvals by category. Not enforced. | @@ -453,7 +453,7 @@ forbid (principal, action == Agent::Action::"execute_bash", resource) when { context.command like "*DROP TABLE*" }; ``` -**Gate destructive git ops** (soft-deny — part of the built-in starter set): +**Gate destructive git ops** (custom timed variants of the built-in starter rules): ```cedar @tier("soft") @rule_id("force_push_any") @@ -488,7 +488,7 @@ forbid (principal, action == Agent::Action::"execute_bash", resource) A force-push to any branch needs approval in 300s. A force-push to `main` or `prod` gives the user 600s with elevated severity. A non-force push to a protected branch (`main`/`prod`/`master`/`release/*`) also gates — catches the case where an agent directly pushes rather than opening a PR. If a command matches both `force_push_any` and `force_push_main`, multi-match merging picks `min(300, 600) = 300s` and `max(medium, high) = high`. -**Protect sensitive file paths** (soft-deny — part of the built-in starter set): +**Protect sensitive file paths** (custom timed variants of the built-in starter rules): ```cedar @tier("soft") @rule_id("write_env_files") @@ -675,35 +675,12 @@ The recent-decision cache is a simple `dict[(tool_name, input_sha), (decision, r When multiple soft-deny rules match a single tool call: -```python -def _merge_annotations(self, policy_ids: list[str]) -> dict: - rule_ids, timeouts, severities = [], [], [] - for pid in policy_ids: - ann = self._annotations[pid] - rule_ids.append(ann.get("rule_id", pid)) - if "approval_timeout_s" in ann: - try: - t = int(ann["approval_timeout_s"]) - if t >= FLOOR_30S: - timeouts.append(t) - except ValueError: - log("WARN", f"malformed @approval_timeout_s on {ann.get('rule_id', pid)}") - severities.append(ann.get("severity", "medium")) - - # Task default always eligible - timeouts.append(self._task_default_timeout_s) - - raw_min_timeout = min(timeouts) - return { - "rule_ids": rule_ids, - "timeout_s": max(FLOOR_30S, raw_min_timeout), # floor enforcement - "severity": _max_severity(severities), # "high" > "medium" > "low" - } -``` +The implementation is `_merge_annotations` in [policy.py](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/agent/src/policy.py). +It takes the shortest **positive** matching-rule or task timeout; zero means no +deadline and is excluded from that minimum. If no positive timeout exists, the +result is zero. Positive results retain the 30-second floor. The highest matching +severity governs the displayed severity, and matching rule IDs are preserved. -**Rationale for min/max choices**: -- **Timeout → min (above floor)**: multiple rules matching means multiple concerns. Users should have *less* time to decide when stakes are higher. Floor prevents unusable 5s windows. -- **Severity → max**: the most severe concern governs the UX coloring. ### 6.4 Allowlist data structure @@ -770,161 +747,30 @@ class ApprovalAllowlist: PreToolUse hook (compressed for doc; implementation will be richer): -```python -async def pre_tool_use_hook(hook_input, tool_use_id, ctx, *, - engine, task_id, user_id, progress, sse_adapter, - task_default_timeout_s): - tool_name, tool_input = _extract(hook_input) - decision = engine.evaluate_tool_use(tool_name, tool_input) - - if decision.outcome == Outcome.ALLOW: - return _allow() - if decision.outcome == Outcome.DENY: - return _deny(decision.reason) - - # REQUIRE_APPROVAL path. - # Cap + rate-limit check. Per-minute rate limit is per-container; on - # container restart the counter resets. The per-task approvalGateCap - # (blueprint-configurable, default 50) is persisted and bounds cumulative - # damage across restarts (§13.6). - if engine.approval_gate_count >= engine.approval_gate_cap: - return _deny(f"approval-gate cap exceeded ({engine.approval_gate_cap}/task)") - if engine.approvals_in_last_minute >= APPROVAL_RATE_LIMIT: - return _deny("approval-gate rate limit exceeded (20/min)") - - # Compute effective timeout with floor/ceiling. - remaining = _remaining_maxlifetime_s() - effective_timeout = max( - FLOOR_30S, - min(decision.timeout_s or task_default_timeout_s, - task_default_timeout_s, - remaining - CLEANUP_MARGIN_120S), - ) - if remaining - CLEANUP_MARGIN_120S < FLOOR_30S: - return _deny(f"insufficient maxLifetime remaining ({remaining}s) for approval") - - request_id = _ulid() - engine.approval_gate_count += 1 - - row = { - "task_id": task_id, "request_id": request_id, - "tool_name": tool_name, - "tool_input_preview": _strip_ansi(_preview(tool_input))[:256], - "tool_input_sha256": _sha256(_serialize(tool_input)), - "reason": decision.reason, "severity": decision.severity, - "matching_rule_ids": list(decision.matching_rule_ids), - "status": "PENDING", - "created_at": _iso_now(), - "timeout_s": effective_timeout, - "ttl": int(time.time()) + effective_timeout + CLEANUP_MARGIN_120S, - "user_id": user_id, "repo": engine.repo, - } - - # ATOMIC: put approval row + transition TaskTable status in one transaction. - try: - await _transact_write_approval_request(task_id, request_id, row) - except TransactionCanceledException as exc: - # Either the task was concurrently cancelled, or status wasn't RUNNING. - _emit("approval_write_failed", {"request_id": request_id, "reason": str(exc)}) - return _deny("approval system unavailable") - - _emit("approval_requested", { - "request_id": request_id, "tool_name": tool_name, - "input_preview": row["tool_input_preview"], - "reason": decision.reason, "severity": decision.severity, - "timeout_s": effective_timeout, - "matching_rule_ids": list(decision.matching_rule_ids), - }) - - outcome = await _poll_for_decision(task_id, request_id, effective_timeout) - - # On TIMED_OUT, attempt to write the row to TIMED_OUT so future reads see - # a terminal state (not orphaned PENDING). The conditional write is guarded - # by `status = :pending` — if the user's APPROVE landed between our last - # poll and this write, the condition fails. In that case we MUST re-read - # the row and honor whatever terminal state won the race; otherwise local - # `outcome.status = "TIMED_OUT"` is stale and we would deny a call the user - # just approved ("I approved it" → agent denies). See §13.12 and the - # scenario below. - if outcome.status == "TIMED_OUT": - wrote_timeout = await _best_effort_update_status( - task_id, request_id, "TIMED_OUT", - reason=outcome.reason, - # Returns True on successful write, False on ConditionCheckFailed. - ) - if not wrote_timeout: - # Re-read the row with ConsistentRead — user's decision beat us. - row = await _ddb_get_approval(task_id, request_id, ConsistentRead=True) - if row is not None and row["status"] == "APPROVED": - # Late-approve wins. Honor it. Rebuild the outcome so the - # downstream allow flow (scope propagation, milestone emission, - # resume transaction) runs identically to the normal approve - # path. - outcome = Decided( - status="APPROVED", - scope=row.get("scope"), - decided_by=row.get("user_id"), - decided_at=row.get("decided_at"), - ) - _emit("approval_late_win", { - "request_id": request_id, - "outcome": "APPROVED", - "reason": "user decision landed during TIMED_OUT write", - }) - elif row is not None and row["status"] == "DENIED": - outcome = Decided( - status="DENIED", - reason=row.get("deny_reason") or "denied", - decided_at=row.get("decided_at"), - ) - # If status is still PENDING (rare — concurrent reaper race) or - # the row is gone (TTL reaped before we could read it), fall - # through with the original TIMED_OUT outcome; fail-closed deny. - - # ATOMIC: resume TaskTable status RUNNING, conditional on awaiting_approval_request_id matching. - try: - await _transact_resume(task_id, request_id) - except TransactionCanceledException: - # User cancelled (or some other path) during poll; abandon gracefully. - _emit("approval_resume_failed", {"request_id": request_id}) - return _deny("task no longer awaiting approval") - - if outcome.status == "APPROVED": - if outcome.scope and outcome.scope != "this_call": - engine._allowlist.add(outcome.scope) - _emit("approval_granted", {"request_id": request_id, - "scope": outcome.scope or "this_call", - "decided_at": outcome.decided_at}) - return _allow() - - # DENIED or TIMED_OUT — cache for 60s + queue denial injection. - engine._recent_decisions.record( - tool_name, _sha256(_serialize(tool_input)), - decision="DENIED" if outcome.status == "DENIED" else "TIMED_OUT", - reason=outcome.reason, - ) - # Truncated reason for guaranteed-surface permissionDecisionReason. - # Best-effort richer injection via _denial_between_turns_hook; may be - # pre-empted by _cancel_between_turns_hook on a concurrently-cancelled - # task. See §4 "Denial with steering text" scenario. - permission_decision_reason = _truncate( - outcome.reason or f"User {outcome.status.lower()}", max_len=500 - ) - if outcome.status == "DENIED": - # Queue steering injection via Stop hook's between_turns_hooks. - engine._queue_denial_injection( - request_id=request_id, - reason=outcome.reason, # already sanitized by DenyTaskFn - decided_at=outcome.decided_at, - ) - _emit("approval_denied" if outcome.status == "DENIED" else "approval_timed_out", - {"request_id": request_id, "reason": outcome.reason}) - return _deny(permission_decision_reason) -``` - -`engine._queue_denial_injection` appends to a list consumed by `_denial_between_turns_hook` — registered **after** `_nudge_between_turns_hook` in the `between_turns_hooks` list (which itself runs after `_cancel_between_turns_hook`). At the next Stop hook fire, the denial is emitted as `…` XML (sanitized via `_xml_escape` from the shared utility introduced with Phase 2). If a `bgagent cancel` has landed between the deny and the next Stop seam, `_cancel_between_turns_hook` short-circuits the dispatcher and the denial text is NOT injected — in which case the guaranteed surface is `permissionDecisionReason` on the hook return. See finding #2 scenario in §4 for the cancel-vs-deny race reasoning. - -**Scenario (§13.12 VM-throttle + late-approval race).** User Alice hits a soft-deny gate at t=0 with `timeout_s=300`. The AgentCore VM is evicted from its warm CPU share around t=285 due to noisy-neighbor pressure on the host; poll ticks stretch by ~400ms. Alice, seeing the approval prompt in Terminal A, types `bgagent approve 01KPW... 01KPR...` at t=294. The approve-transaction lands in DDB at t=294.7 (APPROVED). The agent's next poll-tick was due at t=290 but the VM throttle delayed it to t=295.1. The monotonic wall-clock already shows elapsed >300 (actual since-start ~300.3s), so `_poll_for_decision` returns `TimedOut()`. The hook runs `_best_effort_update_status("TIMED_OUT", ... WHERE status = :pending)` — the conditional fails because the row is APPROVED. **Without the re-read**, the hook would proceed with stale local `outcome.status = "TIMED_OUT"`, queue a denial injection, and return `{"permissionDecision": "deny"}` — Alice sees "I approved it" on Terminal B but the agent denies the tool call anyway. **With the re-read** (the `wrote_timeout` branch in the pseudocode above): the hook fetches the row with ConsistentRead, sees `status = APPROVED`, rebuilds `outcome` from the row (preserving `scope`, `decided_by`, `decided_at`), emits an `approval_late_win` milestone, runs the normal resume transaction + allow flow, and returns `{"permissionDecision": "allow"}`. Alice's tool runs. The cost is one extra strongly-consistent GetItem on the race path; the benefit is that user intent is authoritative. Without this fix, a timer design that is otherwise sound would produce a confounding and unrecoverable UX. See IMPL-24, §13.12, and §15.2 task #43 for the race test. +The executable flow lives in [hooks.py](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/agent/src/hooks.py), with signed +request creation/closure in [approval_requests.py](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/agent/src/approval_requests.py) +and task-state transactions in [task_state.py](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/agent/src/task_state.py). + +1. Evaluate policy, existing grants, gate cap and creation-rate limit. +2. Resolve the shortest positive task/rule deadline. Zero remains untimed. + For a positive deadline, apply a remaining-worker-lifetime ceiling when one + is supplied. A continuation runtime separates this deadline from worker life. +3. Ask the trusted service to atomically create the pending request and move the + task to `AWAITING_APPROVAL`; MicroVM requests also check the active worker lease. + Pending rows have no storage TTL. A failed write denies the action. +4. Emit the notification milestone and poll with strongly consistent reads, + preserving the original UTC/monotonic deadline. No deadline means no timer expiry. +5. If polling times out or fails, ask the trusted service to close the pending + request. If a human decision won the race, reread and preserve that decision. +6. Resume only the matching active task/gate and worker lease. Approval allows the + action and applicable scope; denial is returned to the tool hook and queued as + best-effort steering for the next Stop hook. Cancellation prevents resumption. + +The worker's `TIMED_OUT` outcome currently also covers polling failures; it is +not proof that an explicit human deadline elapsed. This limitation is recorded in +[ADR-023](/sample-autonomous-cloud-coding-agents/decisions/adr-023-trusted-approval-writer). + +**Scenario (§13.12 VM-throttle + late-approval race).** Alice hits a gate with a 300-second window and commits APPROVED at t=294.7s. Scheduling delays prevent the next agent poll until t=300.2s. That poll sees the original deadline has passed and returns TIMED_OUT before reading the row. The conditional TIMED_OUT write then loses because APPROVED is already stored. The hook rereads with `ConsistentRead`, preserves Alice's scope and decision metadata, emits `approval_late_win`, and proceeds through the guarded resume transaction and allow flow. Without that reread it would deny an already-approved call. See IMPL-24, §13.12, and §15.2 task #43 for the race test. --- @@ -952,8 +798,7 @@ Content-Type: application/json | 202 | — | Success | `{task_id, request_id, status: "APPROVED", scope, decided_at}` | | 400 | `VALIDATION_ERROR` | Bad scope format, missing fields | `{error, message, field}` | | 401 | `UNAUTHORIZED` | Missing/invalid JWT | — | -| 404 | `REQUEST_NOT_FOUND` | Row missing OR wrong user (both surfaces 404 to prevent enumeration) | — | -| 409 | `REQUEST_ALREADY_DECIDED` | Approvals row status != PENDING | `{error, message, current_status}` | +| 404 | `REQUEST_NOT_FOUND` | Approval-row condition failed: missing, foreign-owned or already closed | — | | 409 | `TASK_NOT_AWAITING_APPROVAL` | Task's current status is not AWAITING_APPROVAL | `{error, message, current_status}` | | 429 | `RATE_LIMIT_EXCEEDED` | Per-user > 30 approve/min | — | | 503 | `SERVICE_UNAVAILABLE` | DDB throttled or upstream failure | — | @@ -1002,21 +847,32 @@ await ddb.transactWriteItems({ }); ``` -On `TransactionCanceledException`, `ApproveTaskFn` inspects the per-item `CancellationReasons` to distinguish cases: -- ApprovalsTable condition failed with `OldImage` absent → 404 `REQUEST_NOT_FOUND` -- ApprovalsTable condition failed with `OldImage.user_id != caller` → 404 (same code, prevent existence oracle) -- ApprovalsTable condition failed with `OldImage.status != "PENDING"` → 409 `REQUEST_ALREADY_DECIDED` -- TaskTable condition failed (status changed) → 409 `TASK_NOT_AWAITING_APPROVAL` +On `TransactionCanceledException`, `ApproveTaskFn` inspects per-item +`CancellationReasons`. It does not request or classify old item images: + +- Any approval-row condition failure → 404 `REQUEST_NOT_FOUND`, including + cancellation, timeout and an already-recorded decision. +- A task-only condition failure → 409 `TASK_NOT_AWAITING_APPROVAL`. This is symmetric with the agent-side `TransactWriteItems` pattern (§4 step 25a) used for the resume transition — Lambdas and agent speak the same atomic-update contract. -**Ownership**: `user_id` stored on TaskApprovalsTable and compared against `caller_user_id` in the ConditionExpression is the Cognito `sub` claim **verbatim**. The Lambda extracts `sub` from the validated JWT and uses it as-is: no prefix stripping, no tenant mapping, no format normalization. If we ever introduce per-tenant user ID namespacing, that transformation MUST happen at the **write** path (i.e. before the agent writes the row in §4 step 14) rather than at compare time, so the ConditionExpression always compares identical-shape identifiers. See finding #6 scenario below. +**Ownership**: `user_id` stored on TaskApprovalsTable and compared against `caller_user_id` in the ConditionExpression is the Cognito `sub` claim **verbatim**. The Lambda extracts `sub` from the validated JWT and uses it as-is: no prefix stripping, no tenant mapping, no format normalization. If we ever introduce per-tenant user ID namespacing, that transformation MUST happen at the **write** path (i.e. before the approval service persists the row prepared in §4 step 14) rather than at compare time, so the ConditionExpression always compares identical-shape identifiers. See finding #6 scenario below. -After successful transaction, `ApproveTaskFn` writes an audit event to `TaskEventsTable` (`approval_decision_recorded` event_type), ensuring the 90-day audit trail is owned by the Lambda path — not dependent on the agent's milestone emission. +After a successful transaction, the decision handler attempts the authoritative `approval_decision_recorded` audit event independently of the agent milestone. Audit-delivery failure does not undo the saved decision. **Scenario (finding #6):** Three months from now, a platform engineer adds a multi-tenant mode where Cognito `sub` becomes `tenant-abc:01JXZ...`. They update the agent's row-write path to prefix-strip: `user_id = sub.split(":", 1)[1]`, storing `01JXZ...` on TaskApprovalsTable. They forget to update `ApproveTaskFn`. Now the Lambda reads `sub = "tenant-abc:01JXZ..."` from the JWT and compares it against the stored `01JXZ...` — condition fails, 404 on every approve, all tasks stranded. The fix as written: "the Cognito sub is compared verbatim; any transformation must happen at write time, not at compare time" — if the agent writes the full `sub`, the Lambda compares the full `sub`; if either side transforms, both sides must. The CI assertion is a unit test that extracts `user_id` from a sample row and asserts it matches the `sub` claim of a sample JWT byte-for-byte. This test would fail on the prefix-strip refactor above and force the engineer to update both sides. Without this hard rule, ownership-in-condition silently breaks under any future identity refactor. -**Scenario (finding #7):** A user submits a risky task at 10:00 AM. At 10:05 AM the agent hits a soft-deny gate. At 10:05:30 AM the user on Terminal B runs `bgagent cancel 01KPW...`, which lands as CancelTaskFn writes `status=CANCELLING`. At 10:05:31 AM the user — forgetting they just cancelled, or running from a different terminal where they didn't see the cancel — runs `bgagent approve 01KPW... 01KPR...`. Without the cross-table transaction, the Lambda's GetItem on TaskTable (separate call) might read the stale RUNNING state, then UpdateItem on TaskApprovalsTable succeeds because the approvals row is still PENDING → 202 returned. The user sees "approved!" but the task is dying. With the TransactWriteItems pattern, both conditions must hold: the TaskTable guard `status = AWAITING_APPROVAL` fails (because it's now CANCELLING), the entire transaction rolls back, the Lambda returns 409 `TASK_NOT_AWAITING_APPROVAL` with `current_status: CANCELLING`. The user sees "cannot approve: task is already cancelling" and correctly understands state. The cost is one extra table in the transaction (two instead of one) — still within DDB's 100-item limit and nowhere near the 4 MB request size. Symmetric with the agent's resume transaction, which already does the cross-table guard. +**Scenario (finding #7):** A user cancels a task and then approves its old request +from another terminal. Cancellation writes `CANCELLED` directly; there is no +`CANCELLING` task state. The approval transaction cannot commit because the task +is no longer `AWAITING_APPROVAL`. The P3 cancellation update also atomically closes +the linked `PENDING` approval as `CANCELLED` and writes an `approval_cancelled` +event. A later decision is rejected: the existing API returns +`404 REQUEST_NOT_FOUND` for missing, foreign or already-decided approval rows, +including a cancelled row. If approval committed first, cancellation preserves +the recorded decision while cancelling the task. See the +[P3 approval verification record](/sample-autonomous-cloud-coding-agents/verification/readme) +for source versus deployment status. ### 7.2 `POST /v1/tasks/{task_id}/deny` @@ -1062,7 +918,7 @@ New field reference: | Field | Type | Required? | Default | Description | |---|---|---|---|---| -| `approval_timeout_s` | integer seconds | No | **300** | Per-task default approval timeout. Bounded by `[30, min(3600, maxLifetime - 300)]`. Per-rule `@approval_timeout_s` annotations may clip this further (min-wins; see decision #6 and §6.3). Default matches the §10.2 `TaskTable.approval_timeout_s` default. | +| `approval_timeout_s` | integer seconds | No | **0** | Zero retains unanswered requests without a deadline. Positive values must be 30–3,600. The shortest positive task/rule deadline wins. | | `initial_approvals` | list of scope strings | No | `[]` | Pre-approval allowlist scopes (≤20 entries, ≤128 chars each). Validated per §7.3 rules below. | `CreateTaskFn` validations: @@ -1076,7 +932,7 @@ New field reference: - `write_path:X` — same rules as bash_pattern - `rule:X` — X must exist in the (built-in + target repo's blueprint) soft-deny policy set per the shared policy-parsing library; hard-deny rule IDs rejected - `all_session` — rejected if `Blueprint.security.maxPreApprovalScope` forbids -5. `approval_timeout_s` within `[30, min(3600, maxLifetime - 300)]` — cap at 1 hour OR (maxLifetime - 5min), whichever is smaller. Prevents multi-hour slot-exhaustion attacks and keeps approval windows within the TTL budget. +5. `approval_timeout_s` is zero or an integer from 30 through 3,600. MicroVM worker lifetime and capacity are bounded separately through checkpointed retirement; requests are not deleted by a pending-row TTL. 6. Combined `hard + soft + disable + custom` Cedar text size ≤ 64 KB (§12.4); reject on overflow. ### 7.4 Degenerate-pattern detection @@ -1135,7 +991,11 @@ Rate-limited 30/min/user; cached 5min per repo in-Lambda. ### 7.7 `GET /v1/pending` — list pending approvals across user's active tasks -Returns all approvals with `status=PENDING` owned by the caller. Backing index: `user_id-status-index` GSI on `TaskApprovalsTable` (see §10.1). +Returns up to 100 caller-owned pending approvals whose consistently read tasks +are still awaiting that exact request. The `user_id-status-index` GSI supplies +candidates; pagination continues past cancelled/orphaned rows within a shared +five-second read budget. A read failure returns an error rather than a misleading +partial list. **Request**: `GET /v1/pending` with Cognito auth. @@ -1203,7 +1063,7 @@ bgagent submit --task "..." --pre-approve all_session --yes `--pre-approve-file` reads a YAML/JSON array of scope strings — supports the 20-entry cap without command-line bloat. -`--approval-timeout` default (CLI and server): **300 seconds** (5 min), matching decision #6, §7.3, and the `TaskTable.approval_timeout_s` default in §10.2. Accepted range `[30, min(3600, maxLifetime - 300)]` — CLI validates client-side and the server re-validates. `bgagent submit --help` surfaces the default explicitly. +`--approval-timeout` default (CLI and server): **0**, meaning no automatic decision deadline. A positive value must be 30–3,600 seconds. The CLI validates it and the server re-validates it. Matching positive rule deadlines still apply. ### 8.3 Streaming UX @@ -1283,21 +1143,25 @@ TaskApprovalsTable rows are **terminal on first decision** — a row never re-op ```mermaid stateDiagram-v2 - [*] --> PENDING: Agent writes row
(TransactWriteItems with
TaskTable → AWAITING_APPROVAL) + [*] --> PENDING: Approval service creates row
(transaction with
TaskTable → AWAITING_APPROVAL) PENDING --> APPROVED: ApproveTaskFn
(cross-table transaction) PENDING --> DENIED: DenyTaskFn
(cross-table transaction) - PENDING --> TIMED_OUT: Agent poll timeout
(best-effort update) - PENDING --> STRANDED: Reconciler detects
orphan (age > 2×timeout_s) + PENDING --> TIMED_OUT: Approval service records
worker timeout + PENDING --> CANCELLED: Task owner cancels
(cross-table transaction) APPROVED --> [*]: terminal DENIED --> [*]: terminal TIMED_OUT --> [*]: terminal - STRANDED --> [*]: terminal + CANCELLED --> [*]: terminal note right of APPROVED - TTL = created_at + timeout_s + 120s - DDB reaps row after TTL + Terminal retention TTL + never expires a pending request end note ``` +`STRANDED` remains a recognized row status in the type contract. The current +reconciler does not write it: it fails the owning task and closes pending requests +as `CANCELLED`, as described below. + ### 9.3 Orchestrator impact - `waitStrategy` adds `AWAITING_APPROVAL` as non-terminal. @@ -1309,7 +1173,7 @@ stateDiagram-v2 **AWAITING_APPROVAL holds the user's concurrency slot.** -Rationale: the Docker container is alive. Memory allocated. The AgentCore microVM pool is committed. Releasing the slot while the resource is still held lies to accounting and opens a resource-exhaustion vector. +Rationale: the task still owns an unfinished compute session and may resume work. Retaining its ABCA reservation prevents an unbounded collection of parked tasks from bypassing admission control. The rule also applies to P3 Lambda MicroVM suspension; suspended AWS memory-quota consumption remains unverified and is not the basis for claiming quota usage. Resume and terminal cleanup use the existing task-owned reservation protocol. Concrete behavior: @@ -1326,18 +1190,28 @@ t=45m: Task #1 completes. count → 9. Bob can submit task #11. AgentCore Runtime's `maxLifetime = 28800s` (8h) is an absolute timer from session start. It does NOT pause during `AWAITING_APPROVAL`. -This has a concrete implication: the hook computes an `effective_timeout` bounded by `maxLifetime - remaining - CLEANUP_MARGIN_120S`. If the task has been running 7h55m and hits a soft-deny gate, the effective timeout might be clamped to a much shorter value than the task default. Below the 30s floor → immediate DENY with reason `"insufficient lifetime"`. +When the worker supplies a remaining-lifetime estimate, the hook refuses to open a new gate if fewer than 30 seconds remain after the 120-second cleanup margin. This check applies to timed and untimed approvals and reports `"insufficient maxLifetime remaining (s) for approval"`. A positive approval timeout is also capped by that remaining budget. + +The default approval timeout is `0`: no decision deadline. Worker lifetime does not turn silence into a human denial. ECS and AgentCore do not currently restore a waiting agent into a replacement worker; when their task execution limit is reached, the task closes and its pending approval is cancelled. MicroVM can retain a verified checkpoint and continue on a replacement worker. Setting its sleep delay to `0` disables early sleep/retirement, but the coordinator still attempts retirement before the service lifetime ends. This option trades idle cost for faster replies. ### 9.6 Stranded-approval reconciliation -`reconcile-stranded-tasks.ts` gains an AWAITING_APPROVAL-aware branch: +`reconcile-stranded-tasks.ts` has an AWAITING_APPROVAL-aware branch: -- Detects tasks in AWAITING_APPROVAL with `age > 2 * timeout_s` -- Best-effort conditional-updates TaskApprovalsTable row → `STRANDED` status -- Transitions TaskTable → `FAILED` with reason `"approval stranded (container eviction)"` -- Emits `approval_stranded` event to TaskEventsTable +- Uses `APPROVAL_STRANDED_TIMEOUT_SECONDS`, default 30,600 seconds (8.5 hours), measured from + entry into the current status. It does not calculate twice each row's timeout. +- Conditionally changes the task to `FAILED` if it is still awaiting approval, + recording the elapsed wait and a recovery suggestion. +- Emits `task_stranded`, `task_failed` and a wrapped `approval_stranded` milestone. + The legacy milestone has no request ID. +- Closes pending approval rows as `CANCELLED`. Late approval is also rejected by + the task-state guard. A saved MicroVM continuation is handled by its dedicated + coordinator instead of this timeout. -This closes the container-eviction gap. Without this, a container restart mid-approval would leave the task hanging until the user manually cancelled. +The notification helper can recover that legacy milestone's request identity +from the consistently read failed task and its saved stranded cause. It verifies +approval ownership before rendering feedback and does not change the row. +The timer is a backstop, not proof of a particular container failure. `reconcile-concurrency.ts` (scheduled every 5 min) already scans for orphaned concurrency counters; with `AWAITING_APPROVAL` added to `ACTIVE_STATUSES` it correctly counts awaiting tasks as active. @@ -1351,31 +1225,11 @@ The design assumes a human is watching. For truly unattended tasks (scheduled au ### 10.1 New DynamoDB table: `TaskApprovalsTable` -```typescript -new dynamodb.Table(this, 'Table', { - partitionKey: { name: 'task_id', type: dynamodb.AttributeType.STRING }, - sortKey: { name: 'request_id', type: dynamodb.AttributeType.STRING }, // ULID - billingMode: dynamodb.BillingMode.PAY_PER_REQUEST, - pointInTimeRecovery: true, - timeToLiveAttribute: 'ttl', - stream: dynamodb.StreamViewType.NEW_AND_OLD_IMAGES, // (evaluated — may drop; see §11) - removalPolicy: RemovalPolicy.RETAIN, -}); - -// v1 GSI — backs `GET /v1/pending` and `bgagent pending`. -// Required at v1 ship, not deferred — see finding #8 scenario in §7.7. -table.addGlobalSecondaryIndex({ - indexName: 'user_id-status-index', - partitionKey: { name: 'user_id', type: dynamodb.AttributeType.STRING }, - sortKey: { name: 'status', type: dynamodb.AttributeType.STRING }, - projectionType: dynamodb.ProjectionType.INCLUDE, - nonKeyAttributes: [ - 'task_id', 'request_id', 'tool_name', 'tool_input_preview', - 'severity', 'reason', 'created_at', 'timeout_s', - 'matching_rule_ids', - ], -}); -``` +[TaskApprovalsTable](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/cdk/src/constructs/task-approvals-table.ts) defines the +`task_id` partition key, `request_id` sort key, optional retention TTL attribute +`ttl`, and `user_id-status-index` GSI. Streams are disabled; TaskEventsTable carries +the audit/fan-out stream. The GSI discovers candidates; the pending endpoint then +strongly reads the approval and owning task before returning a request. **Projection is fixed at design time.** DynamoDB rejects in-place updates to a GSI's `nonKeyAttributes` (the CloudFormation error @@ -1403,17 +1257,18 @@ Attributes: | `reason` | S | Yes | Cedar matching rule description | | `severity` | S | Yes | "low" \| "medium" \| "high" | | `matching_rule_ids` | L | Yes | List (not Set — can be empty) of soft-deny rule IDs | -| `status` | S | Yes | PENDING \| APPROVED \| DENIED \| TIMED_OUT \| STRANDED | +| `status` | S | Yes | PENDING \| APPROVED \| DENIED \| TIMED_OUT \| STRANDED \| CANCELLED | | `created_at` | S | Yes | ISO8601 | -| `decided_at` | S | No | Set when status != PENDING | +| `decided_at` | S | No | Set when the trusted service or platform closes a request | | `scope` | S | No | Set on APPROVED | | `deny_reason` | S | No | Set on DENIED; sanitized user text | | `timeout_s` | N | Yes | Resolved timeout for audit | -| `ttl` | N | Yes | `created_at_epoch + timeout_s + CLEANUP_MARGIN_120S` — always covers the decision window | +| `expires_at` | S/null | Yes | Original decision deadline; null when `timeout_s=0` | +| `ttl` | N | No | Retention cleanup set on task closure; absent while pending | | `user_id` | S | Yes | Cognito `sub` **verbatim**; used in ownership check `ConditionExpression` (§7.1 finding #6) | | `repo` | S | Yes | Denormalized for fan-out | -**TTL sizing**: the TTL is always `timeout_s + 120s`, so a 300s approval window has a 420s TTL, a 3600s window has a 3720s TTL. The row never expires during the decision window. After the decision + a short grace period, DDB's eventual-consistency TTL reaper cleans up. +**TTL and decision deadlines are separate.** Pending approval rows have no TTL, including explicitly timed requests. Task closure cancels unanswered requests and assigns retention TTL to its approval records. Already-recorded decisions are preserved, and retries do not extend an existing retention deadline. **Why a list, not a StringSet, for `matching_rule_ids`**: DDB string sets cannot be empty. Pathological no-match soft-deny hits would fail to persist. Lists handle empty gracefully. @@ -1425,7 +1280,7 @@ Five new attributes on the existing task row: | Name | Type | Required | Description | |---|---|---|---| -| `approval_timeout_s` | N | No | Default timeout for soft-deny gates. Default 300. | +| `approval_timeout_s` | N | No | Task setting for soft-deny gates. Default 0/no deadline; positive values are 30–3,600 seconds. | | `initial_approvals` | L | No | List of scope strings from submit time | | `awaiting_approval_request_id` | S | No | Set when status = AWAITING_APPROVAL; cleared on transition back (via joint `UpdateExpression`) | | `approval_gate_count` | N | No | Running counter of approval gates fired on this task; used to enforce `approval_gate_cap` (decision #13) | @@ -1483,16 +1338,39 @@ Emitted to both `ProgressWriter` (DDB, 90d) and `sse_adapter` (live stream). Plu ### 11.2 Fan-out plane interaction — Slack button → Cognito mapping -Approval events flow to the fan-out Lambda via TaskEventsTable Streams (the existing Phase 1b path). They are dispatched to Slack / GitHub / Email stubs. - -**TaskApprovalsTable Streams are not consumed by the fan-out Lambda**. The approval row is working state; the audit trail is in TaskEventsTable. Enabling Streams on TaskApprovalsTable would be redundant and add noise. Final design: TaskApprovalsTable DOES NOT have Streams enabled. (Retains the `stream` attribute commented out for future use if needed.) - -Fan-out dispatch rules (extending Phase 1b stubs): -- Slack: on `approval_requested` OR `approval_stranded` — "Agent @task_id requests approval for Bash: `git push --force`" -- Email: on `approval_requested` with `severity: high` -- GitHub: none - -**Rate-limited per-user**: 10 approval-related fan-out messages per user per minute. Prevents notification-spam from malicious users driving up approval-gate count. +Approval events flow to the fan-out Lambda via TaskEventsTable Streams. The P3 +implementation adds Slack and Linear notifications with the saved action, +reason, decision deadline and exact CLI approve/deny commands. Recorded decisions, +cancellations, timeouts and stranded waits also produce messages. The response +path supports the CLI and native Linear thread replies. Slack approval buttons +and the Slack OAuth/button design below remain proposed. Email remains a log-only stub and GitHub does not +receive approval messages. Deployment status is recorded in the +[P3 verification record](/sample-autonomous-cloud-coding-agents/verification/readme). + +**TaskApprovalsTable Streams are not consumed by the fan-out Lambda**. The approval row is working state; the audit trail is in TaskEventsTable. Enabling Streams on TaskApprovalsTable would be redundant and add noise. Final design: TaskApprovalsTable DOES NOT have Streams enabled. + +Slack and Linear route `approval_requested`, `approval_decision_recorded`, +`approval_timed_out`, `approval_cancelled` and `approval_stranded`. The dispatcher +reads the current approval row and owning task before displaying a pending request, +and records successful delivery per request/channel. Delivery failure does not mark +the message delivered. A post that succeeds just before receipt persistence fails +can still produce a duplicate on retry, except for Linear approval prompts and +reply acknowledgements, which use deterministic comment IDs. + +For Linear, the platform saves a thread binding to the exact workspace, issue, +task, request and owner before posting the prompt. A verified Comment/create +webhook containing only `approve` or `deny` in that thread resolves the commenter's +linked platform identity. The owner then uses the same atomic decision function, +rate limit, deadline and current-request guards as the API. Approval grants +`this_call`. A trusted source comment ID saved in the decision transaction makes +webhook retries recognize the original result. Top-level comments, edits and +unbound threads never select a pending request. Pending bindings have no TTL; +closure starts 90-day retention. MicroVM decisions use the existing wake or +continuation path; other backends keep their existing polling behavior. + +**Proposed notification rate limit:** 10 approval-related messages per user per +minute. This dispatcher limit is not implemented. Existing gate-creation caps and +API rate limits remain, but they are not a substitute for notification throttling. **Notification plane is observability, not state (see §13.14).** Notification delivery failures do NOT pause the approval timer — coupling the two creates a bypass where an adversary who takes down the webhook gets an unbounded approval window. The timer runs on the agent's local clock keyed to `created_at`; `bgagent pending` is the recovery path for users who suspect notifications are broken (backed by `user_id-status-index` GSI, §7.7). For the off-hours / unattended trade-off that this posture implies, see §14.8. @@ -1618,11 +1496,33 @@ These alarms transition to `ALARM` state in CloudWatch and appear in the console ### 12.1 Trust boundaries -- **Agent container ↔ TaskApprovalsTable**: IAM role on the runtime has `GetItem` / `PutItem` / conditional `UpdateItem` on the table. Agent writes pending, reads decisions, writes TIMED_OUT on internal timeout. +- **Agent container ↔ TaskApprovalsTable**: workers have task-scoped reads and + transaction condition checks, with no direct item writes or deletion. +- **Agent container ↔ approval request service**: IAM restricts signed + `POST /v1/tasks/{task_id}` to the session's `task_id` tag. The service creates + `PENDING` rows or conditionally records a non-human `TIMED_OUT`; it rejects + human decisions, notification markers, retention TTL and other extra fields. + IAM binds the caller to a task path; the transaction prevents concurrent task + ownership/state changes and, for MicroVM, checks the active worker lease. + The tool preview, its hash, severity, reason and matching rules are worker + assertions, not independently evaluated policy results. The preview can be + truncated, so its bytes cannot be used to verify the full-input hash. This protects approval records; it does not sandbox code running + inside the agent or bind ambient compute credentials to one task. - **User CLI ↔ API Gateway**: Cognito JWT (same authorizer as `/tasks/*`). Cognito `sub` is the canonical caller identity, used **verbatim** in DDB `ConditionExpression` (§7.1, finding #6). - **ApproveTaskFn/DenyTaskFn ↔ TaskApprovalsTable + TaskTable**: Lambda IAM policy allows `UpdateItem` on both tables under `TransactWriteItems`. Authorization is in the ConditionExpression (ownership AND state), not in a separate IAM boundary. - **Blueprint origin**: blueprints are CDK-deployed constructs (see `cdk/src/constructs/blueprint.ts`). Platform operators deploy them. Users cannot upload arbitrary blueprint.yaml from the target repo. This property is load-bearing for the security model — if blueprint origin ever becomes user-uploaded, the blueprint-injection section (§12.4) must be re-evaluated. The 64 KB text cap (§5.1, finding #12) and `disable:` hard-deny rejection (finding #9) are applied regardless of origin as defense in depth. -- **Slack → ApproveTaskFn**: mediated by the fan-out Lambda + `SlackUserMappingTable` (§11.2). Slack admin cannot forge mappings; Slack approvals capped at `severity: low|medium` (finding #4). +- **Linear → decision handlers**: the mapped task owner must author the actual + reply returned by Linear's API. Bot comments and comments authored as the + saved OAuth token identity are rejected. For approvals, install with the + default `actor=app`; diagnostic user-mode credentials cannot approve on + behalf of their authorizing user. Use the authenticated CLI in that case. + A webhook signature alone is insufficient: legacy workers can read a bundle + containing OAuth credentials and the webhook signing key. The authoritative + readback protects this approval path, not every other webhook action. + The Slack button/proxy design in §11.2 remains future work. + +See [ADR-023](/sample-autonomous-cloud-coding-agents/decisions/adr-023-trusted-approval-writer) for the decision and limits, and [upgrading approval permissions](/sample-autonomous-cloud-coding-agents/getting-started/deployment-guide#upgrading-approval-permissions) +before updating an existing deployment. ### 12.2 Ownership encoded in ConditionExpression @@ -1631,9 +1531,12 @@ No TOCTOU window. The `TransactWriteItems` (§7.1) encodes across two tables: - TaskApprovalsTable: `#status = :pending AND user_id = :caller` - TaskTable: `#status = :awaiting AND awaiting_approval_request_id = :rid` -Authorization + approvals-state + task-state transition all atomic. A compromised internal caller (Lambda with raw DDB access) or a logic bug in a future refactor that forgets the ownership check still can't flip rows without matching the `user_id`. The task-state guard additionally prevents the "approve succeeds on a cancelled task" race (finding #7). +The handlers check authorization, approval state and task state atomically. +These conditions prevent races; they do not constrain a compromised Lambda with +raw table-write permission, which could omit them. Only trusted control-plane +handlers receive that permission. -`user_id` comparison is against Cognito `sub` **verbatim** — byte-for-byte equality. Any future identity transformation (per-tenant prefixing, namespacing) must apply to BOTH the write path (agent-side row write) AND the compare path (Lambda ConditionExpression) simultaneously, or the comparison silently fails under the new format. A unit test (§15.3) enforces this: given a sample JWT, extract `sub`, write a row, then assert the stored `user_id` equals `sub` byte-for-byte. +`user_id` comparison is against Cognito `sub` **verbatim** — byte-for-byte equality. Any future identity transformation (per-tenant prefixing, namespacing) must apply to BOTH the write path (trusted approval service) AND the compare path (Lambda ConditionExpression) simultaneously, or the comparison silently fails under the new format. A unit test (§15.3) enforces this: given a sample JWT, extract `sub`, write a row, then assert the stored `user_id` equals `sub` byte-for-byte. ### 12.3 Race prevention @@ -1642,11 +1545,11 @@ Authorization + approvals-state + task-state transition all atomic. A compromise - User's CLI writes `APPROVED WHERE status = :pending` (via TransactWriteItems) - One wins atomically - The loser: - - If TIMED_OUT wins: user gets 409 `REQUEST_ALREADY_DECIDED`. User sees "approval expired". + - If TIMED_OUT wins: a later decision gets 404 `REQUEST_NOT_FOUND`; the saved row is already closed. - If APPROVED wins: agent's poll reads APPROVED on next tick. Agent proceeds. **Race 2 — double-approve**: -- Two concurrent CLI invocations. Second gets 409 `REQUEST_ALREADY_DECIDED`. Idempotent. +- Two concurrent CLI invocations. Only one decision commits; the second gets 404 `REQUEST_NOT_FOUND`. **Race 3 — cancel during AWAITING_APPROVAL**: - Agent writes `RUNNING WHERE status = :awaiting AND awaiting_approval_request_id = :rid` @@ -1746,7 +1649,7 @@ Tracked as IMPL-22. Without these telemetry-driven re-evaluations, 50 will ossif ### 12.10 JWT replay -Cognito JWT with signature + expiry validation on API Gateway. Approval row conditional-update prevents replay from mutating state. Slack button replays similarly mediated by `SlackUserMappingTable` (§11.2) + Slack's own request signing. +Cognito JWT with signature + expiry validation on API Gateway. Approval row conditional-update prevents replay from mutating state. Native Slack approval buttons are not implemented; §11.2 describes the proposed flow. --- @@ -1788,7 +1691,7 @@ The per-task `approvalGateCap` (decision #13; default 50, configurable) is **per ### 13.7 Insufficient lifetime remaining for approval -If `remaining_maxLifetime - CLEANUP_MARGIN_120S < FLOOR_30S`, hook immediately returns DENY with reason `"insufficient maxLifetime for approval"`. Task continues without a gate — or, if the gate was load-bearing, fails gracefully in RUNNING state. +Without a continuation runtime, if `remaining_maxLifetime - CLEANUP_MARGIN_120S < FLOOR_30S`, the hook immediately denies the tool with reason `"insufficient maxLifetime remaining ({n}s) for approval"`, where `{n}` is the remaining lifetime in seconds. No approval request is created and the tool does not run. The agent receives the denial and may choose another action. ### 13.8 PreToolUse hook itself crashes @@ -1814,34 +1717,38 @@ Addressed by the parity contract (decision #23, §15.6). Golden-file CI test run ### 13.12 VM-throttle + late-approval race -The agent's poll loop computes a local timeout wall-clock (`timeout_s` worth of elapsed monotonic time). If the VM is throttled by the hypervisor — either an AgentCore noisy-neighbor eviction window, or a CPU-throttle under memory pressure — poll ticks can stretch past their nominal cadence. In the worst case, the user's APPROVE transaction lands in DDB a few hundred milliseconds before the agent's local clock trips past `timeout_s` and the agent attempts to write `status = TIMED_OUT WHERE status = :pending`. The ConditionCheckFailed path fires (APPROVED already won), but without a re-read the agent's local state is stale: `outcome.status == "TIMED_OUT"` locally while DDB holds APPROVED. The agent would return DENY, the user sees "I approved it" and the agent still blocks — a confounding experience that also violates the design principle that user-observed state is authoritative. +The agent's poll loop retains the original approval row's UTC expiry and a monotonic cap, using whichever expires first. Slow writes/notifications count toward that same window. This also covers a MicroVM whose monotonic clock stops while suspended; the `/resume` hook reuses this deadline and wakes the decision loop. If CPU throttling or suspension delays polling beyond expiry, a user's APPROVE transaction may already have committed. The agent asks the trusted service to conditionally record `TIMED_OUT` while the row remains `PENDING`. The condition fails because APPROVED already won. Without a re-read, the agent's local state would remain TIMED_OUT while DDB holds APPROVED, and it would incorrectly deny the approved call. -**Mitigation**: the §6.5 pseudocode re-reads the approval row with `ConsistentRead=True` whenever `_best_effort_update_status("TIMED_OUT", ...)` returns ConditionCheckFailed, and honors whatever terminal state the row carries: +**Mitigation**: the hook rereads the approval row with `ConsistentRead=True` when the service reports that another decision won the timeout race, and honors the recorded decision: - If `status == "APPROVED"`: rebuild the local `outcome` to reflect APPROVED, preserving `scope`, `decided_by`, `decided_at`, and proceed through the normal allow flow (scope-propagation, `approval_granted` milestone, resume transaction, return `{"permissionDecision": "allow"}`). Emit a `approval_late_win` milestone so operator telemetry can count races. - If `status == "DENIED"`: honor the denial text the user submitted. Agent returns DENY with the user's sanitized reason as `permissionDecisionReason` (same surface as normal deny). -- If `status` is still PENDING (rare — concurrent reaper race) or the row is gone (TTL reaped): fall through with the original TIMED_OUT outcome; fail-closed deny. +- If `status` is still PENDING (rare — concurrent reaper race) or the row is missing: fall through with the original TIMED_OUT outcome; fail-closed deny. The scenario is bounded by the polling cadence (2-5s ticks) and DDB's strongly-consistent read latency (tens of ms), so the re-read adds at most one extra GetItem to the racing path — acceptable cost for honoring user intent. See IMPL-24 and §15.2 task #43 (race tests). ### 13.13 Runtime JWT expiry during approval wait -**Context in this codebase (verified 2026-05-06).** The AgentCore Runtime container authenticates outbound AWS API calls (DynamoDB, Secrets Manager, etc.) via the container's IAM role, which the SDK resolves through the instance-metadata-service equivalent and auto-refreshes transparently. There is no user-presented JWT with a short rolling expiry consumed by the container's own API calls — `grep -rn -iE 'runtime.jwt|jwt.refresh|token_expiry' agent/src/` returns nothing (only `token_usage` for LLM billing and `GITHUB_TOKEN` for git operations). AgentCore Runtime invocation on the Lambda side uses sigv4 via `InvokeAgentRuntimeCommand` (see `cdk/src/handlers/shared/strategies/agentcore-strategy.ts`) — also auto-refreshed AWS credentials, not a user JWT. The "Runtime-JWT" label in §4 step 6 and the sequence diagrams below refers to the **caller-facing SSE auth** (Terminal A's Cognito ID token presented to API Gateway to stream task events) — it does not authenticate the container's own DDB writes. - -**Therefore, for v1: no separate Runtime JWT expiry term is required in the ceiling computation.** The `maxLifetime` term (AgentCore's hard lifetime of 8h) is the only upper bound we control; IAM credentials refresh automatically within that window. The ceiling definition in decision #6 stands as `min(1h, maxLifetime_remaining - cleanup_margin)`. +Workers authenticate AWS calls with IAM role credentials, not the user's Cognito +JWT. The CLI's Cognito token protects its platform API calls; an expired CLI login +requires reauthentication but does not itself delete the saved approval request. +MicroVM resume refreshes runtime and task-role credentials before releasing coding. -**If the auth model changes** (e.g. a future design introduces a container-held user JWT to authenticate `permissionDecisionReason` attribution, or to carry the caller's Cognito `sub` end-to-end for per-user DDB conditions), the ceiling MUST be extended to `min(1h, maxLifetime_remaining - 120s, runtime_jwt_expiry - 120s)` and this section updated. Tracked as IMPL-27 so the contract is reviewed whenever the auth shape changes. The failure signature if this bound is missed: the container's IAM calls succeed but some JWT-gated channel (e.g. Terminal A's SSE stream) quietly 403s mid-approval-wait; the user's decision lands in DDB but the agent's poll fails to deliver `approval_granted` to the live stream. Today that channel is best-effort observability, not state — but a future state-bearing channel would need the ceiling term. +Credential renewal does not extend a worker's service lifetime. Apply the worker +and explicit-deadline rules in §9.5; do not impose a new approval deadline based on +the user's API token expiry. A future worker-held user-token design would need a +separate review of that boundary. ### 13.14 Notification delivery failure -Fan-out delivery failures (Slack down, email bounce, webhook 5xx) do **NOT** pause the approval timer. The timer runs on the agent's local clock, keyed to the `created_at` timestamp on the DDB row — it is independent of whether any notification channel succeeded in alerting the human. +A failed notification does not change the recorded decision deadline. Timed +requests retain their original UTC/monotonic deadline; untimed requests remain +available until task closure or another valid resolution. Neither case permits +the action without approval. -**Rationale (security):** coupling the timer to notification-plane availability creates a bypass. An adversary who takes down the webhook (or poisons the Slack rate limit) would get an unbounded approval window; worse, a compromised tenant could deliberately suppress their own notifications to escape gates. Fail-closed on timer expiry is invariant; delivery is best-effort observability. +Users can discover unanswered requests with `bgagent pending`, which reads through +the authenticated API independently of notification delivery. A late-discovered +explicitly timed request does not receive a fresh decision window. -**Recovery path for the user:** `bgagent pending` queries `TaskApprovalsTable` directly via the `user_id-status-index` GSI (§7.7, §10.1) — it does not depend on notification delivery. A user who suspects notifications are broken can poll `bgagent pending` at any time to see all live approvals. If the notification never landed and the user finds a gate via `bgagent pending`, they can `bgagent approve/deny` normally; the timer is still running against the original `created_at`, not against when the user found it. - -**Operational signal:** `approval_timed_out` events carry `timeout_s` and the `created_at`/`decided_at` delta. A rising `approval_timed_out` rate with flat `approval_requested` rate (measured via `ApprovalTimeoutClipRate` and `ApprovalDecisionLatency` in §11.3) is the telemetry that indicates notification breakage, not an unresponsive user. - -**The fail-closed posture on timer expiry remains unchanged.** Delivery-availability-aware scheduling is the notification plane's job (see §14.8 and INTERACTIVE_AGENTS.md notification-plane design); the timer does not reason about it. ### 13.15 Fail-closed summary @@ -1902,7 +1809,7 @@ BLOCKED[]: (resource: ) # when a resource is named ### 14.1 Scenario A: force-push with per-rule timeout -Setup: repo `my-org/my-app` blueprint extends soft-deny with `force_push_main` (@approval_timeout_s=600). Task default is 300s. +Setup: repo `my-org/my-app` blueprint extends soft-deny with `force_push_main` (@approval_timeout_s=600). This example explicitly sets the task timeout to 300s. ```bash $ bgagent run --repo my-org/my-app \ @@ -2020,7 +1927,7 @@ Each phase has explicit scope. Matches real-world review workflows. Visible in a ### 14.6 Scenario F: VM-throttle + late-approval race (trace) -Setup: task default 300s; force-push gate fires. User Alice approves at the very edge of the timeout window while the VM is throttled. +Setup: task explicitly configured with a 300-second timeout; force-push gate fires. User Alice approves at the very edge of the timeout window while the VM is throttled. ``` t=0.00s PreToolUse hook fires: Bash "git push --force origin feature-x" @@ -2029,22 +1936,18 @@ t=0.02s TransactWriteItems: approval row PENDING + TaskTable AWAITING_APPROVA t=0.03s agent_milestone: approval_requested → Terminal A stream t=0.04s _poll_for_decision begins; interval=2s for first 30s, then 5s -... (poll ticks every 5s from t=30 to t=295) ... +... (poll ticks every 5s from t=30 to t=285) ... t=285.0s host hypervisor evicts VM from warm CPU share (noisy neighbor). - Next scheduled poll was t=290.0s; actual scheduling delay ~5.1s. + Next scheduled poll was t=290.0s; actual scheduling delay ~10.2s. t=294.7s Alice's bgagent approve lands at API Gateway. ApproveTaskFn TransactWriteItems: ApprovalsTable: PENDING → APPROVED (user_id matches, status was PENDING) TaskTable: state guard holds (still AWAITING_APPROVAL, rid matches) → 202 returned to CLI; Alice sees "approved!" in Terminal B -t=295.1s agent's delayed poll tick fires. Elapsed wall-clock = 295.1s. - Monotonic elapsed is 295.1 > timeout_s=300? NO — but the poll - function computes `elapsed >= timeout_s` and on the NEXT tick - (t=300.2s) it will exceed. -t=300.2s next tick: elapsed=300.2 ≥ timeout_s=300 → TimedOut() returned. - (Alice's APPROVED write at t=294.7s was MISSED — the previous - poll was due at t=295.0 but the VM throttle stretched it past.) +t=300.2s delayed tick: original deadline has passed → TimedOut() returned + before reading the approval row. Alice's APPROVED write at + t=294.7s has not yet been observed by this agent. t=300.3s _best_effort_update_status("TIMED_OUT", ... WHERE status = :pending) → ConditionCheckFailed (row is APPROVED, not PENDING) → wrote_timeout = False @@ -2081,16 +1984,20 @@ Setup: Bob submits with `--approval-timeout 600`. The blueprint has a `write_cre Bob sees both the pre-submit warning (`approval_timeout_capped_at_submit`) and the per-gate cap event (`approval_timeout_capped`) so he understands why his 600s didn't apply. Without these milestones, the user sees only `timeout: 300s` in the approval banner and may think the CLI dropped their setting. Both events are captured in the event stream and surface via `bgagent watch`. See §11.1, §11.3 (`ApprovalTimeoutClipRate`), and Fix 4 / IMPL-26 in §16. -### 14.8 Off-hours and unattended tasks (known trade-off) +### 14.8 Off-hours and unattended tasks -**Known trade-off: off-hours failure.** Because timeouts are fail-closed (decision #6), a task running overnight with pending approvals will fail if no approver responds in time. This is deliberate — auto-approve on timeout would make "wait the reviewer out" the attacker's winning strategy; see decision #6 and §13.15 fail-closed summary. +Unanswered requests now remain available by default. On MicroVMs, a verified +checkpoint and worker retirement bound resource use while the person is away; +their later answer can start a replacement worker. Other compute substrates keep +their existing runtime limits. This does not approve any action automatically. -**For overnight / unattended runs, choose one of:** -- `--pre-approve all_session --yes` to bypass gates entirely for that task (accept the broader trust grant; see §7.3). -- Configure escalation on the `approval_requested` event via the notification plane (see `docs/design/INTERACTIVE_AGENTS.md` for channel configuration). Route to whoever is on-call; escalation schedule is the tenant's responsibility, not the timeout engine's. -- Schedule the task during business hours. +For a time-sensitive request, set an explicit positive decision deadline. Its +expiry remains fail-closed: the tool is denied, and the agent decides what to do +next. Notification routing can still notify an on-call reviewer; scheduling that +escalation belongs to the notification layer. -The `approvalGateCap` (decision #13) will force-fail the task after approximately `cap × task_default_timeout_s` of unanswered gates — default worst case ~4h at cap=50 / timeout=300s. Plan accordingly. +`approvalGateCap` limits the number of gates reached by a task, not elapsed human +waiting time. It does not impose a deadline on one unanswered request. **Why the timer itself is timezone-unaware:** Business-hours logic belongs in the notification plane, not the authorization engine. Baking calendars or on-call rotations into the timer couples the security boundary to a scheduling system it doesn't own. Same rule evaluated at 9am and 3am because the security property (adversary cannot wait out review) is time-invariant. Delivery-availability-aware scheduling is the notification plane's job via subscribed-channel health and escalation policies. @@ -2128,7 +2035,7 @@ See §17.18 for the off-hours escalation future-work primitive, and §13.14 for | # | Package | File | Change | |---|---|---|---| | 1 | agent | Spike | Validate cedarpy.policies_to_json_str() returns annotations. Confirm `diagnostics.reasons` shape for multi-match. If API diverges, update §6 before proceeding. | -| 2 | mise + agent + cdk | `mise.toml`, `agent/pyproject.toml`, `cdk/package.json` | Pin `cedarpy==4.8.0` (agent) and `@cedar-policy/cedar-wasm==4.10.0` (cdk). The two bindings are intentionally on different version lines — verified compatible via the parity fixtures, not required to be equal. Both pinned exactly, not `^` or `~` — decision #23 / finding #1. | +| 2 | mise + agent + cdk | `mise.toml`, `agent/pyproject.toml`, `cdk/package.json` | Pin both Cedar bindings exactly in their package manifests and verify the pair through the shared parity fixtures; see §15.6. | | 3 | agent + cdk | `contracts/cedar-parity/*.json` (shared fixture dir; follows precedent set by `contracts/memory-hash-vectors.json`) | Golden-file parity fixtures: `(policy_set, input) → {decision, matching_rule_ids}`. Agent side loads via `cedarpy`; Lambda side via `cedar-wasm`. Divergence fails CI. | | 4 | agent | `src/policy.py` | Extend `PolicyDecision` (outcome/timeout_s/severity/matching_rule_ids/allowed-property). Split `_DEFAULT_POLICIES` into hard + soft. Add annotation parsing. Implement `ApprovalAllowlist` + `RecentDecisionCache` (50-entry LRU cap, independent of `approvalGateCap`). Load-time validation (rule_id uniqueness, tier mismatch, annotation floor, 64 KB cap, disable-list hard-deny rejection, `approvalGateCap` bounds check `1 ≤ N ≤ 500`). `PolicyEngine.__init__` accepts `approval_gate_cap` sourced from blueprint (default 50). | | 5 | agent | `policies/hard_deny.cedar` (new) | Migrate current hard-deny rules + add DROP TABLE. Annotations. | @@ -2193,7 +2100,7 @@ See §17.18 for the off-hours escalation future-work primitive, and §13.14 for - Cancel during AWAITING_APPROVAL (agent-side resume race) - Cancel during approve Lambda (cross-table transaction catches it — finding #7) - Cancel during deny with queued denial injection (between-turns hook pre-empted; `permissionDecisionReason` still delivered — finding #2) - - Late approval after TIMED_OUT (expect 409) + - Late approval after TIMED_OUT (expect 404 `REQUEST_NOT_FOUND`) - **VM-throttle + late-approve race (IMPL-24, §13.12)**: user's APPROVE lands in DDB before `_best_effort_update_status("TIMED_OUT")` can claim the row; agent must re-read with ConsistentRead and honor APPROVED (not return stale TIMED_OUT). Includes the DENIED variant and the "still PENDING" fall-through where neither side wins. - **Chaos tests**: - Container restart mid-approval (simulated via kill + reconciler) @@ -2223,11 +2130,13 @@ Built-in policies shipped with the agent: **Hard-deny (absolute, cannot be disabled by blueprint)**: `rm_slash`, `write_git_internals`, `write_git_internals_nested`, `drop_table`. Absolute; no scope bypasses them; blueprint `disable:` cannot remove them (§5.1, finding #9). **Soft-deny starter set (require approval by default, may be disabled by blueprint)**: -- `force_push_any` — `like "*git push --force*"` — medium, 300s -- `push_to_protected_branch` — pushes to `main`/`master`/`prod`/`release/*` (non-force) — medium, 300s -- `force_push_main` — force-push specifically to `main`/`prod` — high, 600s -- `write_env_files` — `like "*.env"` — high, 600s -- `write_credentials` — `like "*credentials*"` — high, 300s +- `force_push_any` — `like "*git push --force*"` — medium +- `push_to_protected_branch` — pushes to `main`/`master`/`prod`/`release/*` (non-force) — medium +- `force_push_main` — force-push specifically to `main`/`prod` — high +- `write_env_files` — `like "*.env"` — high +- `write_credentials` — `like "*credentials*"` — high + +These built-in rules inherit the task deadline: zero/no deadline by default. Custom rule annotations may impose a positive deadline. Users who want fully autonomous execution (no approval gates) pass `--pre-approve all_session --yes` at submit. Repos that want additional gates add them via `Blueprint.security.cedarPolicies.soft`. Repos that want a different policy set can override specific built-in **soft-deny** rules by `@rule_id` via the blueprint's `security.cedarPolicies.disable` list. The `disable:` mechanism is restricted: it may NOT include any built-in hard-deny rule_id, and the blueprint loader rejects such configurations at task start. @@ -2262,19 +2171,22 @@ Rollout steps: ### 15.5 Backward compatibility -- Existing tasks without `initial_approvals` → empty list → no pre-approvals, default `approval_timeout_s = 300` +- Tasks without `initial_approvals` receive an empty list and no pre-approvals. + New tasks default to `approval_timeout_s = 0`. An already-persisted approval + retains its original timeout; an upgrade or replacement does not extend it. - Existing policies without `@rule_id` / `@tier` → engine fails to start (fail-closed). Blueprint authors must add annotations explicitly during migration. - `PolicyDecision.allowed` property provides backward compat for existing `if not decision.allowed` callers - Hook return shape unchanged — Phase 1a/1b tests continue to pass ### 15.6 Shared Cedar parsing — cross-engine parity contract -The agent runtime uses Python [`cedarpy@4.8.0`](https://pypi.org/project/cedarpy/); the Lambda side (`CreateTaskFn`, `ApproveTaskFn`, `DenyTaskFn`, `GetPoliciesFn`) uses [`@cedar-policy/cedar-wasm@4.10.0`](https://www.npmjs.com/package/@cedar-policy/cedar-wasm) — AWS's official WASM-compiled Cedar engine. Same Rust core, two bindings. Because these engines evolve independently, we ship a **parity contract** (decision #23, finding #1) to catch drift before deploy. - -**Version pinning.** Both engines are pinned exactly (not `^` or `~`) in the monorepo's canonical manifest files. The two bindings are deliberately on **different version lines** — they are NOT required to be equal. `cedarpy` and `cedar-wasm` follow independent release cadences over the shared Cedar Rust core, and the currently-shipped pins (`cedarpy==4.8.0` ↔ `@cedar-policy/cedar-wasm==4.10.0`) are an intentional, tested-compatible skew: the parity fixtures in `contracts/cedar-parity/` are what certify that this specific pair produces identical `(decision, matching_rule_ids)` on every fixture. The rule is "move together and re-verify parity when you bump either side," not "keep the version strings equal." -- `agent/pyproject.toml`: `cedarpy==4.8.0` -- `cdk/package.json`: `"@cedar-policy/cedar-wasm": "4.10.0"` -- `mise.toml` documents the pinned versions in a comment for operator visibility +The agent uses Python `cedarpy`; the policy Lambdas use +`@cedar-policy/cedar-wasm`. Exact current pins live in +[agent/pyproject.toml](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/agent/pyproject.toml) and +[cdk/package.json](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/cdk/package.json). Binding version strings need not be +identical. Upgrade them as a tested pair and run the shared +[parity fixtures](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/contracts/cedar-parity/README.md), which check decisions +and matching rule IDs across both engines. **Lambda layer packaging** (finding #5). The cedar-wasm package is 4.1 MB unzipped. Shipping it in the deployment bundle of each of the 4 policy Lambdas would consume ~16 MB of unzipped bundle size — manageable on its own but leaves little room for AWS SDK + other deps as the codebase grows, and threatens the Lambda 250 MB unzipped limit under realistic growth. Solution: package cedar-wasm as a **Lambda layer** (`cedar-wasm-layer.ts`, task #10 in §15.2), attached to each policy Lambda. This reduces each Lambda's deployment bundle to just the handler code + thin wrapper around the layer import. Policy Lambdas are configured with ≥ 512 MB memory to accommodate WASM module instantiation under concurrent invocation (measured under 100-concurrent bursts in §15.3 Lambda memory tests). @@ -2321,7 +2233,7 @@ flowchart LR When policy authors upgrade either engine, the parity fixture must be re-generated (a small helper script dumps decisions from both engines; the human confirms the change is intentional). -**Scenario (finding #1, illustrative):** This example uses hypothetical versions (e.g. cedarpy `4.10.1` → `4.11.0`) to show the *class* of bug the parity contract catches; it does not describe the real shipped pins (which are the intentional `cedarpy==4.8.0` ↔ `cedar-wasm==4.10.0` skew documented above). The point is that closeness of version strings — even within the same minor line — is no guarantee of behavioral parity, which is exactly why the golden fixtures, not the version numbers, are the source of truth. A platform engineer runs `mise run deps:update` which bumps cedarpy from 4.10.1 to 4.11.0. They notice cedar-wasm is still 4.10.0 but assume it's fine because both say "4.x". Between these versions, cedarpy added support for a new `context has` operator that cedar-wasm doesn't yet have. A new blueprint soft-deny rule uses `context has "approved_context"`. On deploy: +**Scenario (finding #1, illustrative):** This example uses hypothetical versions (e.g. cedarpy `4.10.1` → `4.11.0`) to show the *class* of bug the parity contract catches; it does not describe the real shipped pins (read the package manifests for the actual pins). The point is that closeness of version strings — even within the same minor line — is no guarantee of behavioral parity, which is exactly why the golden fixtures, not the version numbers, are the source of truth. A platform engineer runs `mise run deps:update` which bumps cedarpy from 4.10.1 to 4.11.0. They notice cedar-wasm is still 4.10.0 but assume it's fine because both say "4.x". Between these versions, cedarpy added support for a new `context has` operator that cedar-wasm doesn't yet have. A new blueprint soft-deny rule uses `context has "approved_context"`. On deploy: - Agent-side `PolicyEngine.__init__` parses the rule successfully; engine loads normally. - `CreateTaskFn` on the Lambda side calls cedar-wasm `policyToJson()` — it throws: `ParseError: unknown operator 'has' at line 3`. - User submits a task against that repo. `CreateTaskFn` crashes mid-validation. Error message: "500 Internal Server Error" (because the Lambda didn't handle the upstream parse error gracefully). @@ -2416,7 +2328,7 @@ Items from the design reviews not captured above as design changes — to be add **IMPL-9** (functional P1-3): Runtime allowlist revocation. Not shipped in v1. Placeholder: `bgagent revoke-approval ` noted in §17. -**IMPL-10** (functional P1-12): `approval_timeout_s` default 300 documented consistently in §3 #6, §7.3 table, §10.2 attribute description. +**IMPL-10** (historical functional P1-12): the original 300-second default was superseded by the September retained-request default of `0` (no decision deadline). **IMPL-11** (functional P2-8): CLI `run.ts` command exists from Phase 1b. `submit.ts` also exists. `--pre-approve` / `--approval-timeout` flags added to both. @@ -2606,7 +2518,7 @@ See §15.2. Net new files: ~15. Net modified files: ~15. Total LOC estimate: ~40 - [ ] Backward compat: Phase 1a/1b tests pass without modification - [ ] ULID length references are 26 chars throughout CLI + docs - [ ] **Re-read approval row on TIMED_OUT ConditionCheckFailed (IMPL-24)**: `_best_effort_update_status("TIMED_OUT")` failure path re-reads with ConsistentRead and honors APPROVED/DENIED if the user's decision beat the agent's timer; emits `approval_late_win` milestone. See §6.5 pseudocode, §13.12 VM-throttle race, §14.6 trace, §15.2 task #43. -- [ ] **Default `--approval-timeout` is 300s** documented consistently in decision #6, §5.2, §7.3 field table, §8.2 CLI flags, and §10.2 TaskTable schema. +- [ ] **Default `--approval-timeout` is 0 (no deadline)** documented consistently in decision #6, §5.2, §7.3 field table, §8.2 CLI flags, and §10.2 TaskTable schema. - [ ] **Sub-120s `@approval_timeout_s` emits WARN (IMPL-25)** at blueprint load; sub-30s still rejected. `bgagent lint-policies` (§17.14) surfaces the same WARN pre-submit. - [ ] **User-visible timeout milestones (IMPL-26)**: `approval_timeout_capped` (per-gate, on SSE stream), `approval_timeout_capped_at_submit` (on `POST /v1/tasks` response), `approval_ceiling_shrinking` (once per task at lifetime threshold). All carry `{requested_timeout_s, effective_timeout_s, reason}`. - [ ] **Runtime JWT ceiling (IMPL-27)**: no separate JWT expiry term required in v1 — container uses auto-refreshed IAM credentials (verified by grep of `agent/src/`). Ceiling stays `min(1h, maxLifetime_remaining - cleanup_margin)`. Review if container auth shape changes (see §13.13). diff --git a/docs/src/content/docs/architecture/Compute.md b/docs/src/content/docs/architecture/Compute.md index c574d6350..41de3b5f9 100644 --- a/docs/src/content/docs/architecture/Compute.md +++ b/docs/src/content/docs/architecture/Compute.md @@ -22,10 +22,10 @@ The default runtime is **Amazon Bedrock AgentCore Runtime**, which runs each ses | **Startup** | Service-managed | Slim images help | Snapshot resume | Warm ASGs + pre-pull | Karpenter + pre-pull | Backend-dependent | Provisioned concurrency | Snapshot pools (DIY) | | **GPU** | No | No | No | Yes | Yes | Yes (EC2/EKS backend) | No | Yes (with passthrough) | | **Ops burden** | Low (managed) | Low | Low (managed) | Medium | High | Low-Medium | Low | **Very high** | -| **Cost model** | vCPU-hrs + GB-hrs | vCPU + mem/sec | Baseline-priced (8 GiB / 4 vCPU) with 4× vertical burst (32 GiB / 16 vCPU peak); suspended time is storage-only | EC2 + EBS | EKS control + EC2 | Underlying compute | Request + duration | EC2 metal + your ops | +| **Cost model** | vCPU-hrs + GB-hrs | vCPU + mem/sec | Baseline compute (8 GiB / 4 vCPU) plus additional burst usage (up to 32 GiB / 16 vCPU); no suspended compute charge, but snapshot storage and read/write charges remain | EC2 + EBS | EKS control + EC2 | Underlying compute | Request + duration | EC2 metal + your ops | | **Fit** | **Default choice** | Repos > 2 GB image | Suspend/resume economics; approval-wait-heavy workloads; default-sized repos. Heavy sustained-memory builds stay on ECS | GPU, heavy toolchains | Max flexibility | Queued batch jobs | **Poor** (15 min cap) | Best potential, highest cost | -> **Lambda MicroVMs are not Lambda functions.** They are a different compute primitive, so the functions column's 15-minute cap and poor-fit verdict do not apply. See [ADR-021](/sample-autonomous-cloud-coding-agents/architecture/adr-021-lambda-microvms-compute-backend). +> **Lambda MicroVMs are not Lambda functions.** They are a different compute primitive, so the functions column's 15-minute cap and poor-fit verdict do not apply. See [ADR-021](/sample-autonomous-cloud-coding-agents/decisions/adr-021-lambda-microvms-compute-backend). The backend is selected per repo via `compute_type` in the Blueprint config. The orchestrator resolves the strategy and delegates session start, polling, and termination to the strategy implementation. See [REPO_ONBOARDING.md](/sample-autonomous-cloud-coding-agents/architecture/repo-onboarding) for the `ComputeStrategy` interface. @@ -81,11 +81,32 @@ See [ORCHESTRATOR.md](/sample-autonomous-cloud-coding-agents/architecture/orches ## Lambda MicroVMs backend -Lambda MicroVMs are an opt-in third backend, selected per repository with `compute_type: lambda-microvm`; AgentCore remains the default. Image configuration has three states: a managed base-image ARN and version creates the snapshot image in CDK; an external image identifier uses a snapshot built out of band; and supplying neither provisions only the roles, buckets, and connectors needed for the bootstrap deploy. `cdk/scripts/package-microvm-artifact.sh` packages the agent as zip + Dockerfile, uploads it to the artifact bucket, and can create the external image. Lambda MicroVMs are available in five launch regions (us-east-1, us-east-2, us-west-2, eu-west-1, ap-northeast-1) and will expand; the platform enforces regional availability in layers via a synth-time constant, onboarding live probes, and orchestration-time classification. +Lambda MicroVMs are an opt-in third backend, selected per repository with `compute_type: lambda-microvm`; AgentCore remains the default. Image configuration has three states: a managed base-image ARN, version and artifact digest create the snapshot image in CDK; an external image identifier uses a snapshot built out of band; and supplying neither image provisions only the roles, buckets, and connectors needed for the bootstrap deploy. Lambda MicroVMs are available in five launch regions (us-east-1, us-east-2, us-west-2, eu-west-1, ap-northeast-1) and will expand; the platform enforces regional availability in layers via a synth-time constant, onboarding live probes, and orchestration-time classification. -Because a snapshot freezes its build-time environment, deployment-specific, non-secret identifiers travel in the `/run` hook's `platform_config` block instead. The strategy sends the canonical inline envelope or, when that envelope exceeds the verified 4,096-byte `runHookPayload` limit, an S3-pointer envelope with the configuration also merged into the uploaded payload. The agent accepts only allowlisted keys and installs them before pipeline initialization; [ADR-021 §3](/sample-autonomous-cloud-coding-agents/architecture/adr-021-lambda-microvms-compute-backend#3-packaging-same-agent-image-source-new-build-path) defines the exact wire shapes and validation rules. +New MicroVM installations should select `microvm_nested_stack=true` and require bootstrap bundle 1.9.0. +The layout setting is temporarily required; omission fails synthesis until migration is verified. +Before upgrading an existing flat installation, set `microvm_nested_stack=false` +and retain it until completing the [resource migration](/sample-autonomous-cloud-coding-agents/verification/645-p3-nested-stack). -Networking separates image build from execution: the build-only connector permits TCP 80 and 443 because the Dockerfile uses `apt-get`, while running MicroVMs retain 443-only egress through the platform VPC. Every launch explicitly passes the Lambda-managed `NO_INGRESS` connector; omission would select the service's public-ingress default. The P2 image declares and serves `/ready` and `/validate` at build time and `/run` and `/terminate` at runtime. `/suspend` and `/resume` remain disabled until their P3 implementation. +For managed images, run `cdk/scripts/package-microvm-artifact.sh --stack-name ` after each agent change. It packages the Dockerfile's local inputs into a deterministic ZIP and uploads it under `microvm-images/agent-artifact-.zip`. The checksum is a fingerprint of the uploaded bytes: identical inputs reuse the same verified object, and changed inputs produce a new filename. Deploy with the printed `--context microvm_artifact_sha256=` alongside `microvm_base_image_arn` and `microvm_base_image_version`, and retain these inputs for later deployments. The changed S3 URI tells CloudFormation to update the existing image; overwriting the old fixed filename alone does not. Missing or malformed digests fail synthesis. An initial deployment without an image must create the bucket first. The script's explicit `--create-image` alternative retains the fixed base key and calls the image API directly. + +Because a snapshot freezes its build-time environment, current deployment identifiers arrive through the v2 payload bootstrap. The coordinator publishes a non-secret deployment manifest and sends a single-object signed download URL for the task. The worker reads only its deployment's `bootstrap/*` with ambient credentials; other object reads and payload-bucket listing are explicitly denied. The downloaded task identity and configuration must match the authenticated manifest before configuration installation. The serialized reference fits the verified 4,096-byte hook limit; all payload sizes use S3. ECS shares this transport through `AGENT_PAYLOAD_REF`. [ADR-021 §3](/sample-autonomous-cloud-coding-agents/decisions/adr-021-lambda-microvms-compute-backend#3-packaging-same-agent-image-source-new-build-path) defines the wire format and compatibility requirements; [live payload checks](/sample-autonomous-cloud-coding-agents/verification/readme) record validation. + +Networking separates image build from execution: the build-only connector permits TCP 80 and 443 because the Dockerfile uses `apt-get`, while running MicroVMs retain 443-only egress through the platform VPC. Every launch explicitly passes the Lambda-managed `NO_INGRESS` connector; omission would select the service's public-ingress default. Images declare and serve `/ready` and `/validate` at build time and `/run`, `/terminate`, `/suspend` and `/resume` at runtime. Automatic suspension is a separate deployment opt-in. Registry HTTP/SSE tools therefore need reachable HTTPS/443 endpoints; remote non-443 tools are unsupported under the default policies of AgentCore and ECS as well. Local `stdio` tools can run, with their outbound traffic subject to the same restriction. Asset resolution/loading does not probe connectivity; see [registry network support](/sample-autonomous-cloud-coding-agents/architecture/registry#2-asset-kinds-for-mvp). + +P3 connects the guest checkpoint/credential hooks, actual image-version capability, durable supervisor and post-commit approval wake. The `microvm_approval_suspend_enabled` deployment context defaults false and controls both a static opt-in and a live Parameter Store switch. Existing durable executions reread the live switch before new suspension because their original Lambda environment is pinned. Recovery and the original service lifetime survive supervisor replay; the API preserves accepted decisions when optional wake fails. AgentCore/ECS retain explicit unsupported pause/wake results. Explicit approval deadlines use the original UTC/monotonic deadline. See the [lifecycle diagnostics](/sample-autonomous-cloud-coding-agents/verification/645-p3-lifecycle-diagnostics) for failure investigation, and the [acceptance status](/sample-autonomous-cloud-coding-agents/verification/readme) for deployment evidence. + +Tasks can set `microvm_sleep_after_s` (CLI: `--microvm-sleep-after `). +The default is 600 seconds of waiting for each approval; zero disables sleep. +Creation persists the resolved preference, while legacy rows use the default. +The global suspension switch still takes precedence. Unanswered approvals have +no deadline by default (`approval_timeout_s=0`); an explicit finite deadline +remains available. A short timed request stays awake when there is too little +useful sleep time before its wake margin. Neither suspension nor replacement +extends that original deadline. For a longer wait, a verified conversation and +workspace checkpoint lets ABCA retire the worker, release capacity and start a +replacement after the answer. Snapshot storage and save/restore fees mean the +sleep delay is a user preference, not a guarantee of savings for every pause. ## ECS Fargate task sizing (build vs. planning) @@ -108,7 +129,7 @@ The agent harness is the layer around the LLM that manages the execution loop: c The platform uses the [Claude Agent SDK](https://github.com/anthropics/claude-agent-sdk-python) as the harness. It provides the agent loop, built-in tools (filesystem, shell), and streaming message reception for per-turn trajectory capture (token usage, cost, tool calls). -**Execution model:** Tasks are fully unattended and one-shot. The agent loop runs in a background thread so the FastAPI `/ping` endpoint stays responsive on the main thread. The agent thread uses `asyncio.run()` with the stdlib event loop (uvicorn is configured with `--loop asyncio` to avoid uvloop conflicts with subprocess SIGCHLD handling). +**Execution model:** The agent runs automatically until it completes, needs a human decision or is stopped. The HTTP entrypoint runs its loop in a background thread so FastAPI `/ping` stays responsive; ECS uses the batch entrypoint. The agent thread uses `asyncio.run()` with the stdlib event loop (uvicorn is configured with `--loop asyncio` to avoid uvloop conflicts with subprocess SIGCHLD handling). **System prompt:** Selected by workflow from a shared base template (`agent/src/prompts/base.py`) with per-workflow sections (`coding/new-task-v1`, `coding/pr-iteration-v1`, `coding/pr-review-v1`). The platform defines what the agent should do; the harness executes it. @@ -123,7 +144,7 @@ The platform uses the [Claude Agent SDK](https://github.com/anthropics/claude-ag | GitHub | AgentCore Gateway + Identity | Clone, push, PR, issues | | Web search | AgentCore Gateway | Documentation lookups | -Plugins, skills, and MCP servers are out of scope for MVP. Additional tools can be added via Gateway integration. +Agents can use configured skills and MCP tools through the [registry](/sample-autonomous-cloud-coding-agents/architecture/registry). Their runtime network access remains subject to the compute backend’s egress rules. ### Policy enforcement diff --git a/docs/src/content/docs/architecture/Deployment-roles.md b/docs/src/content/docs/architecture/Deployment-roles.md index 9603bf56d..3e09a9c42 100644 --- a/docs/src/content/docs/architecture/Deployment-roles.md +++ b/docs/src/content/docs/architecture/Deployment-roles.md @@ -34,7 +34,7 @@ The policies are split into six IAM managed policies (each under the 6,144-chara > **Placeholder substitution**: Replace `ACCOUNT_ID` with your 12-digit AWS account ID and `REGION` with your deployment region (e.g., `us-east-1`) throughout this document. -These policies are not created or attached manually. The repository generates them — and a custom bootstrap template that wires all six into the CloudFormation execution role — from the TypeScript sources, then bootstraps with that template: +These policies are not created or attached manually. The repository generates them and a custom bootstrap template that attaches the selected policies to the CloudFormation execution role: ```bash # Regenerate artifacts (policies JSON + template YAML) and bootstrap. @@ -52,14 +52,29 @@ aws cloudformation update-stack --stack-name CDKToolkit --use-previous-template aws cloudformation describe-stacks --stack-name CDKToolkit --query 'Stacks[0].Parameters' ``` -Under the hood, `mise //cdk:bootstrap` runs `npx cdk bootstrap --template bootstrap/bootstrap-template.yaml` (see `cdk/mise.toml`). The generated template defines six inline `AWS::IAM::ManagedPolicy` resources that **replace** the default `AdministratorAccess` on the CloudFormation execution role; the `IaCRole-ABCA-Compute-ECS` and `IaCRole-ABCA-Compute-LambdaMicrovms` policies are conditional on the `ComputeTypes` parameter including their respective backend. The policy sources are `cdk/src/bootstrap/policies/{infrastructure,application,observability,compute-agentcore,compute-ecs,compute-lambda-microvm}.ts`, compiled to `cdk/bootstrap/policies/*.json` by `cdk/scripts/generate-bootstrap-artifacts.ts`. +Under the hood, `mise //cdk:bootstrap` runs `npx cdk bootstrap --template bootstrap/bootstrap-template.yaml` (see `cdk/mise.toml`). The generated template defines six `AWS::IAM::ManagedPolicy` resources that **replace** the default `AdministratorAccess` on the CloudFormation execution role; the `IaCRole-ABCA-Compute-ECS` and `IaCRole-ABCA-Compute-LambdaMicrovms` policies are conditional on the `ComputeTypes` parameter including their respective backend. The policy sources are `cdk/src/bootstrap/policies/{infrastructure,application,observability,compute-agentcore,compute-ecs,compute-lambda-microvm}.ts`, compiled to `cdk/bootstrap/policies/*.json` by `cdk/scripts/generate-bootstrap-artifacts.ts`. + +**Re-bootstrap to bundle 1.7.0 or later before deploying nested stacks.** The +execution role also needs the generated inline policy +`PassExecutionRoleToCloudFormation`, from +`cdk/src/bootstrap/nested-stack-policy.ts`. It grants `iam:PassRole` on that exact +execution role, with `iam:PassedToService=cloudformation.amazonaws.com`. The ARN +uses the bootstrap partition, account, Region and qualifier; it grants no access +to pass other roles. Without it, a fresh 1.6.0 deployment fails change-set +validation when CloudFormation tries to pass its role to the registry's nested +stacks. Preserve all existing bootstrap parameters when updating the template. + +Bundle 1.7.0 also corrects `BootstrapPolicyHash`: it includes nested policy fields +and the generated inline policy. Earlier hashes could remain unchanged after +Action, Resource or Condition changes. A matching old hash is insufficient proof +that deployed permissions match source. > **CloudFormation inline-template limit — 51,200 characters**: This is a second, independent size ceiling, distinct from the per-policy IAM 6,144-character limit above. `cdk bootstrap --template` sends the template inline as `TemplateBody`; above 51,200 characters the CLI has to stage it in S3 instead, which it **cannot** do while bootstrapping a fresh account, because that bucket is one of the resources bootstrap creates. The result is a hard `BootstrapStackRequired` failure with no way through `cdk bootstrap`, `--force` included ([#864](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/864)). > > Two consequences for anyone editing the policies: > > - **The gated size is not the file's size on disk.** The CLI parses the file, discards its formatting, and re-serialises the parsed object before measuring. Reformatting `bootstrap-template.yaml` therefore changes nothing; only the *content* moves the number. Check it with `npx cdk bootstrap --show-template --template bootstrap/bootstrap-template.yaml | wc -c`. -> - **Each `PolicyDocument` is emitted as a minified JSON string**, not a nested YAML mapping. Both are valid for this `Json`-typed property and IAM stores the string parsed, but a string scalar survives the CLI's re-serialisation on one line — which is what keeps the body under the ceiling (45,743 characters, versus 53,369 as mappings). +> - **Each managed-policy `PolicyDocument` is emitted as a minified JSON string**, not a nested YAML mapping. Both are valid for this `Json`-typed property and IAM stores the string parsed, but a string scalar survives the CLI's re-serialisation on one line. The original fix reduced the body to 45,743 characters from 53,369; subsequent policy additions remain subject to the budget. The small inline self-role policy uses a mapping to resolve its ARN with `Fn::Sub`. > > `cdk/scripts/generate-bootstrap-template.ts` fails the build when the body exceeds the budget in `cdk/src/bootstrap/template-size.ts`, so adding statements surfaces the problem at generation time rather than against somebody's fresh account. Both the guard and its regression test obtain the size by invoking `cdk bootstrap --show-template` on the committed artifact — the CLI is the component that makes the inline-vs-S3 decision, so asking it directly cannot drift the way a local copy of its serialiser would. No AWS credentials are required. @@ -96,7 +111,7 @@ Under the hood, `mise //cdk:bootstrap` runs `npx cdk bootstrap --template bootst For deploying the `backgroundagent-dev` stack. This single stack contains all platform resources including the AgentCore runtime, ECS compute (when enabled), API Gateway, Cognito, DynamoDB tables, VPC, DNS Firewall, and observability infrastructure. -> **IAM managed policy size limit**: A single managed policy cannot exceed 6,144 characters. The permissions below are split into six policies to stay under this limit (three always-applied, plus three compute-variant policies). They are wired into the CloudFormation execution role by the generated bootstrap template; see [Using these policies](#using-these-policies). +> **IAM managed policy size limit**: A single managed policy cannot exceed 6,144 characters. The permissions below are split into six policies to stay under this limit (four always applied, including AgentCore, plus optional ECS and MicroVM policies). The separate inline self-role grant is described above. They are wired into the CloudFormation execution role by the generated bootstrap template; see [Using these policies](#using-these-policies). ### IaCRole-ABCA-Infrastructure @@ -290,6 +305,12 @@ CloudFormation stack operations, IAM roles/policies, VPC networking, and Route 5 DynamoDB tables, Lambda functions, API Gateway, Cognito, WAFv2, EventBridge, SQS, CloudFront, and Secrets Manager. When ECS Fargate compute is enabled, add the ECS statement below to this policy. +Agent Registry provisioning also uses a Step Functions workflow to wait for +asynchronous creation and deletion. Its construct explicitly names the workflow +with the parent stack's `backgroundagent-dev-` prefix so it fits this policy, +including when deployed in a nested stack. Adopting this name in an existing +deployment replaces the provider's waiter state machine. + ```json { "Version": "2012-10-17", @@ -827,6 +848,31 @@ The second statement, `MicrovmPassRoles`, is the one exception to the rule that > **Operators must re-bootstrap for this.** The statement ships in bootstrap policy bundle **1.6.0**; a CDKToolkit stack bootstrapped at 1.5.0 or earlier will fail the CDK-managed MicroVM image deploy with a caller-side `iam:PassRole` AccessDenied on the build role. Check `CDKToolkit`'s `BootstrapPolicyVersion` output, and re-run `mise //cdk:bootstrap` (with `ComputeTypes` including `lambda-microvm`) if it is behind. +P3 additionally requires **bundle 1.8.0** for `MicrovmSuspendConfiguration`. + +The nested MicroVM layout requires **bundle 1.9.0**. Its child stack uses the +explicit parent-derived names `backgroundagent-dev-MicrovmBuildRole` and +`backgroundagent-dev-MicrovmConnectorRole`; `MicrovmPassRoles` admits those two +exact names in addition to the legacy flat-layout prefixes. The execution role +stays in the parent and is still excluded. Re-bootstrap before deploying the +child stack. Before upgrading an existing flat deployment, set and retain +`microvm_nested_stack=false` until its resource migration is complete; changing ownership is not an ordinary +in-place update. See the [nested-stack runbook](/sample-autonomous-cloud-coding-agents/verification/645-p3-nested-stack). + +For a reviewed migration that keeps old and new resources side by side, +`microvm_resource_name_prefix` gives the nested image, network connectors and log +group distinct names. It requires nested mode and a concrete 1–40 character +letter/digit/hyphen prefix. Build/operator IAM role names remain derived from the +parent deployment so the bootstrap's existing `PassRole` scope still applies. +Keep the selected prefix stable in later deployments. This option alone does +not preserve the old resources or their runtime permissions; those remain part +of the migration procedure. +This statement lets CloudFormation manage and tag the live suspension setting +at `//microvm-approval-suspend-enabled`. The +coordinator gets only `GetParameter` on its exact parameter. Existing durable +executions retain their Lambda version and reread this setting before new +suspension, so disable can reach executions already running. + ```json { "Statement": [ @@ -861,9 +907,24 @@ The second statement, `MicrovmPassRoles`, is the one exception to the rule that "Effect": "Allow", "Resource": [ "arn:aws:iam::*:role/backgroundagent-dev-LambdaMicrovmComputeBuild*", - "arn:aws:iam::*:role/backgroundagent-dev-LambdaMicrovmComputeConnector*" + "arn:aws:iam::*:role/backgroundagent-dev-LambdaMicrovmComputeConnector*", + "arn:aws:iam::*:role/backgroundagent-dev-MicrovmBuildRole", + "arn:aws:iam::*:role/backgroundagent-dev-MicrovmConnectorRole" ], "Sid": "MicrovmPassRoles" + }, + { + "Action": [ + "ssm:GetParameters", + "ssm:PutParameter", + "ssm:DeleteParameter", + "ssm:AddTagsToResource", + "ssm:RemoveTagsFromResource", + "ssm:ListTagsForResource" + ], + "Effect": "Allow", + "Resource": "arn:aws:ssm:*:*:parameter/backgroundagent-*/microvm-approval-suspend-enabled", + "Sid": "MicrovmSuspendConfiguration" } ], "Version": "2012-10-17" diff --git a/docs/src/content/docs/architecture/Identity-and-auth.md b/docs/src/content/docs/architecture/Identity-and-auth.md index c91ede25f..eef62f87f 100644 --- a/docs/src/content/docs/architecture/Identity-and-auth.md +++ b/docs/src/content/docs/architecture/Identity-and-auth.md @@ -4,10 +4,10 @@ title: Identity and auth # Identity and authentication — worked examples -ABCA today carries its own inbound identity (Amazon Cognito for the CLI and REST API, HMAC-SHA256 for webhooks) and resolves outbound credentials per integration through hand-rolled Secrets Manager resolvers (`resolve_github_token`, `resolve_linear_api_token`, and the Jira OAuth/Forge resolver). This doc maps each of those integration shapes onto the [Amazon Bedrock AgentCore Identity](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/identity.html) primitives: workload identity, the token vault, and the three outbound flows. A contributor adding a new integration can then see which flow to pick and what the wire looks like at each hop. Every code block below is grounded in the [AgentCore developer guide](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/identity.html) and the [CreateOauth2CredentialProvider API reference](https://docs.aws.amazon.com/bedrock-agentcore-control/latest/APIReference/API_CreateOauth2CredentialProvider.html); the binding decision for ABCA is captured in [ADR-016](/sample-autonomous-cloud-coding-agents/architecture/adr-016-pluggable-identity-and-auth). +ABCA today carries its own inbound identity (Amazon Cognito for the CLI and REST API, HMAC-SHA256 for webhooks) and resolves outbound credentials per integration through hand-rolled Secrets Manager resolvers (`resolve_github_token`, `resolve_linear_api_token`, and the Jira OAuth/Forge resolver). This doc maps each of those integration shapes onto the [Amazon Bedrock AgentCore Identity](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/identity.html) primitives: workload identity, the token vault, and the three outbound flows. A contributor adding a new integration can then see which flow to pick and what the wire looks like at each hop. Every code block below is grounded in the [AgentCore developer guide](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/identity.html) and the [CreateOauth2CredentialProvider API reference](https://docs.aws.amazon.com/bedrock-agentcore-control/latest/APIReference/API_CreateOauth2CredentialProvider.html); the binding decision for ABCA is captured in [ADR-016](/sample-autonomous-cloud-coding-agents/decisions/adr-016-pluggable-identity-and-auth). - **Use this doc for:** picking the right outbound flow (USER_FEDERATION / M2M / OBO) for a new integration, reading the principal at each hop of a webhook-triggered async agent, and seeing the Linear before/after the token vault replaces. -- **Related docs:** [SECURITY.md](/sample-autonomous-cloud-coding-agents/architecture/security) for the security boundaries and the shared-PAT limitation, [ADR-016](/sample-autonomous-cloud-coding-agents/architecture/adr-016-pluggable-identity-and-auth) for the two-seam pluggable-auth decision, and the Authentication page (`/using/authentication`) for ABCA's current inbound auth. +- **Related docs:** [SECURITY.md](/sample-autonomous-cloud-coding-agents/architecture/security) for the security boundaries and the shared-PAT limitation, [ADR-016](/sample-autonomous-cloud-coding-agents/decisions/adr-016-pluggable-identity-and-auth) for the two-seam pluggable-auth decision, and the Authentication page (`/using/authentication`) for ABCA's current inbound auth. ## Design principle @@ -188,7 +188,7 @@ The config field is `grantType`, singular. The nested actor field is `actorToken A human creates a ticket and adds a label (`agent:triage` on Jira; the same shape covers the GitHub-issue-label trigger the ABCA team is discussing in #abca). Automation fires a webhook, the agent triages, comments, and maybe closes, acting on the user's behalf where write-back matters. This is the realistic trigger pattern: no interactive browser, identity arrives as webhook payload, and consent happens later when the agent calls back. -> **Illustrative target, not the shipped Jira write path.** The flow below shows a future user-bound token-vault design. Current ABCA Jira uses 3LO only for inbound reads and human lookup. Outbound writes go to an HMAC-signed Forge web trigger, which calls Jira with `api.asApp()` so the actor is the installed `bgagent` app. The human Jira account remains task-attribution data, not the outbound credential. See [ADR-015](/sample-autonomous-cloud-coding-agents/architecture/adr-015-jira-integration). +> **Illustrative target, not the shipped Jira write path.** The flow below shows a future user-bound token-vault design. Current ABCA Jira uses 3LO only for inbound reads and human lookup. Outbound writes go to an HMAC-signed Forge web trigger, which calls Jira with `api.asApp()` so the actor is the installed `bgagent` app. The human Jira account remains task-attribution data, not the outbound credential. See [ADR-015](/sample-autonomous-cloud-coding-agents/decisions/adr-015-jira-integration). The architecture chain: diff --git a/docs/src/content/docs/architecture/Interactive-agents.md b/docs/src/content/docs/architecture/Interactive-agents.md index f968bc830..fa90f8f8c 100644 --- a/docs/src/content/docs/architecture/Interactive-agents.md +++ b/docs/src/content/docs/architecture/Interactive-agents.md @@ -4,7 +4,9 @@ title: Interactive agents # Interactive Agents: Async Interaction Design -> **Status:** Active design +> **Status:** Historical interaction design with current approval-path corrections (2026-09-22). +> Proposed dispatcher services and Slack buttons below are not all implemented; use the +> [API contract](/sample-autonomous-cloud-coding-agents/architecture/api-contract) and [approval guide](/sample-autonomous-cloud-coding-agents/using/approval-gates-cedar-hitl) for supported interfaces. > **Branch:** `feature/interactive-background-agents` > **Last updated:** 2026-04-29 (rev 6) @@ -23,7 +25,7 @@ This document describes the interactivity surfaces layered on top of that model 3. **Watch** — `bgagent watch ` polls `TaskEventsTable` with an adaptive interval (500 ms when events are arriving, back-off to 5 s when idle). Same endpoint used under the hood for foreground-block UX on `ask` and for HITL approval waits. 4. **Nudge** — `bgagent nudge ""` writes a row into `TaskNudgesTable`. The agent reads pending nudges between turns, acknowledges with a `nudge_acknowledged` milestone event, and integrates the nudge on its next turn. 5. **Ask** — `bgagent ask ""` (Phase 2) writes a question row. The agent answers at the next between-turns boundary; the answer surfaces as a `status_response` event. CLI default is foreground block-and-poll with a spinner; task and answer are both durable if the CLI disconnects. -6. **Approval gates** — Phase 3 Cedar-driven hard gates. Agent emits `approval_requested`, waits for a decision from `bgagent approve` / `bgagent deny` or a Slack button-press. Detailed design in [`CEDAR_HITL_GATES.md`](/sample-autonomous-cloud-coding-agents/architecture/cedar-hitl-gates). +6. **Approval gates** — Phase 3 Cedar-driven hard gates. Agent emits `approval_requested`, waits for a decision from `bgagent approve` / `bgagent deny` or an owner-authored Linear thread reply. Slack notifications provide CLI instructions; buttons remain proposed. Detailed design in [`CEDAR_HITL_GATES.md`](/sample-autonomous-cloud-coding-agents/architecture/cedar-hitl-gates). ### Core architectural choices @@ -31,7 +33,7 @@ This document describes the interactivity surfaces layered on top of that model - **Durable event table (`TaskEventsTable`)** is the one source of truth for agent progress. Every reader — CLI, Slack/GitHub/email dispatchers, status Lambda — reads from this table, never from the live agent. - **Polling-only CLI.** No SSE, no WebSockets. DDB eventually-consistent reads with an `event_id` cursor are cheap, reliable, and compute-agnostic. - **Notification plane as first-class.** A FanOutConsumer Lambda subscribes to `TaskEventsTable` DDB Streams and routes per-event-type to per-channel dispatcher Lambdas (Slack, email, GitHub comment). Per-channel defaults ship in v1. -- **Agent interaction via the hook mechanism the Claude Agent SDK provides.** Nudges, asks, and approvals all use `Stop` / between-turns hooks; no mechanism outside the SDK's contract is required. +- **Agent interaction via the hook mechanism the Claude Agent SDK provides.** Nudges and asks use `Stop` / between-turns hooks; approval gates pause in `PreToolUse`, with denial steering delivered at a later Stop hook; no mechanism outside the SDK's contract is required. --- @@ -115,7 +117,7 @@ This document describes the interactivity surfaces layered on top of that model │ DELETE /tasks/{id} cancel │ │ POST /tasks/{id}/nudge nudge │ │ POST /tasks/{id}/asks ask (P2) │ - │ POST /tasks/{id}/approvals approve P3 │ + │ POST /tasks/{id}/approve approve │ │ POST /webhooks/tasks GH webhook │ └───────────┬──────────────────────────────────┘ │ @@ -227,17 +229,16 @@ Consumer: agent between-turns hook reads pending nudges, emits `nudge_acknowledg ### 3.7 TaskApprovalsTable (Phase 3) Phase 3 approval-request spine. Detailed schema in [`CEDAR_HITL_GATES.md`](/sample-autonomous-cloud-coding-agents/architecture/cedar-hitl-gates). Semantics summary: -- Agent writes an approval row with the request context. -- Agent transitions `RUNNING → AWAITING_APPROVAL` and enters a poll loop. -- User responds via REST (`POST /tasks/{id}/approvals/{request_id}`) or via a Slack button dispatched by the notification plane. +- The worker calls the trusted approval service, which atomically creates the pending row and transitions the task `RUNNING → AWAITING_APPROVAL`. The worker then polls for a decision. +- The owner responds through `POST /tasks/{id}/approve` or `/deny` with `request_id` in the body, the equivalent CLI commands, or an `approve`/`deny` reply to the Linear approval comment. - On decision, agent transitions back to `RUNNING`; denial reasons are injected as Stop-hook steering on the next turn. ### 3.8 FanOutConsumer (router) Lambda subscribed to `TaskEventsTable` DDB Streams (relying on the DynamoDB Streams **default** `ParallelizationFactor` of 1, which preserves per-`task_id` ordering by shard — not set explicitly in `fanout-consumer.ts`; see §6.1). Reads per-task notification config (from `TaskTable` metadata or `RepoTable` defaults), filters events by channel subscription, and invokes per-channel dispatcher Lambdas. -- **SlackDispatchFn** — posts to configured channel / DM. Includes action buttons for `approval_required` events. -- **EmailDispatchFn** — SES. +- **SlackDispatchFn** — posts to configured channel / DM. Approval notifications currently contain CLI response instructions; buttons remain proposed. +- **EmailDispatchFn** — proposed SES delivery; current email dispatch is a log-only stub. - **GitHubDispatchFn** — edits a single GitHub issue comment in place via `PATCH /repos/{o}/{r}/issues/comments/{id}`. On 404 (comment deleted upstream) falls back to POSTing a fresh comment. Per-task ordering is guaranteed upstream by the DDB Streams default `ParallelizationFactor` of 1 (see §6.1), so no conditional-request header is needed (and GitHub's REST API does not accept `If-Match` on this endpoint — see §6.4). Detailed routing and default filters in §6. @@ -300,8 +301,8 @@ Authentication: Cognito User Pool ID token in `Authorization` header for all RES | `agent_milestone` | Agent code (pipeline, hooks) | Named checkpoint (`repo_cloned`, `pr_opened`, `nudge_acknowledged`, ...) | | `agent_cost_update` | Runner | Cumulative token + dollar cost | | `agent_error` | Runner | Handled exception | -| `approval_required` (P3) | PreToolUse Cedar hook | Cedar policy requires user decision | -| `approval_decided` (P3) | Approve/Deny Lambda | User responded | +| `approval_requested` (P3) | PreToolUse Cedar hook | Cedar policy requires user decision | +| `approval_decision_recorded` (P3) | Approve/Deny Lambda | User responded | | `status_response` (P2) | Between-turns hook | Agent answered an `ask` | | `nudge_acknowledged` | Between-turns hook | Agent saw a nudge before incorporating it | | `pr_created` | Pipeline | PR opened for the task | @@ -324,7 +325,7 @@ Consumers page `TaskEventsTable` using `event_id` as a cursor: `KeyConditionExpr ### 5.1 `bgagent submit` ``` -$ bgagent submit --repo org/repo "fix the auth timeout bug" +$ bgagent submit --repo org/repo --task "fix the auth timeout bug" task submitted: abc123 ``` @@ -415,9 +416,9 @@ Flags: HITL approval commands. All flows are REST + DDB; no streaming. Detailed design in [`CEDAR_HITL_GATES.md`](/sample-autonomous-cloud-coding-agents/architecture/cedar-hitl-gates). Summary: -- Agent emits `approval_required` with the tool context. -- Notification plane dispatches the event (Slack with action buttons, email, GitHub). -- User responds via `bgagent approve `, `bgagent deny --reason "…"`, or Slack button click. +- Agent emits `approval_requested` with the tool context. +- The notification plane sends approval messages to Slack and Linear. Email is a stub; GitHub does not receive approval messages. +- The owner responds via `bgagent approve `, `bgagent deny --reason "…"`, or an `approve`/`deny` reply to the Linear approval comment. - Agent's poll loop sees the decision and proceeds or deny-steers. ### 5.7 `bgagent cancel` @@ -448,19 +449,22 @@ TaskEventsTable ──DDB Stream──▶ FanOutConsumer - Router reads per-task notification config (channel enablement + event-type filters), then invokes the relevant dispatcher Lambda(s) per event. - Dispatchers are separate Lambdas so a GitHub API outage doesn't block Slack notifications. -### 6.2 Per-channel defaults (v1) +### 6.2 Original proposed per-channel defaults (v1) | Channel | Default subscribed events | Opt-in via `--verbose` | |---|---|---| -| **Slack** | `task_completed`, `task_failed`, `task_cancelled`, `pr_created`, `agent_error`, `approval_required`, `status_response` | adds `agent_milestone` | -| **Email** | `task_completed`, `task_failed`, `approval_required` | — | +| **Slack** | `task_completed`, `task_failed`, `task_cancelled`, `pr_created`, `agent_error`, `approval_requested`, `status_response` | adds `agent_milestone` | +| **Email** | `task_completed`, `task_failed`, `approval_requested` | — | | **GitHub issue comment** | `pr_created`, terminal status (single edit-in-place comment) | — already minimal | Rationale: if Slack pings on every milestone, users mute the bot within days. Default to the minimal set that surfaces decision-requiring events and completion; power users opt into verbose streams. ### 6.3 Slack approval buttons -`approval_required` events delivered to Slack include `Approve` / `Deny` action buttons. On click, Slack invokes an interaction callback Lambda which writes to `TaskApprovalsTable` via the same `POST /approvals` path the CLI uses. This gives the common case (reviewer in Slack, not at a terminal) a one-click response path. +**Proposed, not implemented.** Current Slack approval messages contain CLI +commands. A future button callback would need to map the Slack user to the task +owner and use the authenticated decision path; a valid Slack signature alone +would not authorize approval. Linear thread replies already support owner decisions. ### 6.4 GitHub issue comment — edit-in-place @@ -486,7 +490,7 @@ Submitted with the task (optional) or resolved from repo defaults: { "notifications": { "slack": { "enabled": true, "channel": "#coding-agents", "events": ["default"] }, - "email": { "enabled": true, "events": ["approval_required", "task_failed"] }, + "email": { "enabled": true, "events": ["approval_requested", "task_failed"] }, "github": { "enabled": true, "events": ["default"] } } } diff --git a/docs/src/content/docs/architecture/Orchestrator.md b/docs/src/content/docs/architecture/Orchestrator.md index f1bfc6dd9..bd55896b7 100644 --- a/docs/src/content/docs/architecture/Orchestrator.md +++ b/docs/src/content/docs/architecture/Orchestrator.md @@ -23,7 +23,7 @@ The orchestrator sits between the API layer and the agent runtime. Changes to ta ## Responsibilities -The orchestrator is deliberately scoped. It handles coordination and bookkeeping but never touches agent logic, compute infrastructure, or memory storage. This clear boundary means a crashed agent does not leave orphaned state, and platform invariants (concurrency limits, event audit, cancellation) cannot be bypassed by agent code. +The orchestrator handles coordination, compute lifecycle calls and finalization bookkeeping. The agent runs the coding workflow. Recovery still needs explicit guards around external effects, including saved start receipts and task-owned capacity reservations. ### What the orchestrator owns @@ -36,7 +36,7 @@ The orchestrator is deliberately scoped. It handles coordination and bookkeeping | Result inference | Determine success or failure from agent response, DynamoDB record, and GitHub state | | Finalization | Update status, emit events, release concurrency, persist audit records | | Cancellation | Stop the session and drive the task to CANCELLED at any point | -| Concurrency | Track per-user and system-wide running task counts with atomic counters | +| Concurrency | Track per-user capacity with task-owned reservations and an atomic counter | ### What the orchestrator does NOT own @@ -100,7 +100,7 @@ stateDiagram-v2 AWAITING_APPROVAL --> RUNNING : Approved or denied (resume) AWAITING_APPROVAL --> CANCELLED : User cancels mid-approval - AWAITING_APPROVAL --> FAILED : Stranded-approval reconciler + AWAITING_APPROVAL --> FAILED : Infrastructure loss or stranded wait FINALIZING --> COMPLETED : PR or commits found FINALIZING --> FAILED : No useful work @@ -125,11 +125,11 @@ stateDiagram-v2 | `HYDRATING` | `FAILED` | Hydration error | GitHub API failure, guardrail blocks content, Bedrock unavailable | | `RUNNING` | `AWAITING_APPROVAL` | Cedar soft-deny gate fires | Tool call triggers a soft-deny policy rule during execution | | `RUNNING` | `FINALIZING` | Session ends | Response received or session terminated | -| `RUNNING` | `TIMED_OUT` | Max duration exceeded | AgentCore and Lambda MicroVMs have an 8h substrate cap; the orchestrator's own safety-net poll window is `MAX_POLL_ATTEMPTS` (1020) × 30s ≈ 8.5h, after which a still-`RUNNING` task is driven to `TIMED_OUT` | +| `RUNNING` | `TIMED_OUT` | Max duration exceeded | AgentCore and Lambda MicroVMs have an 8h substrate cap; MicroVM supervision retains the original service deadline across replay; other backends retain the 1,020-attempt safety window (about 8.5h at 30s) | | `RUNNING` | `FAILED` | Session crash | Heartbeat or substrate liveness lost (see Liveness monitoring) | | `AWAITING_APPROVAL` | `RUNNING` | Approved or denied | Human decision received; agent resumes | | `AWAITING_APPROVAL` | `CANCELLED` | User cancels | Explicit cancel while awaiting approval | -| `AWAITING_APPROVAL` | `FAILED` | Stranded reconciler | Approval request orphaned (agent died mid-wait) | +| `AWAITING_APPROVAL` | `FAILED` | Infrastructure failure or stranded wait | Lost compute, exhausted supervisor recovery/window, or an orphaned approval; the approval decision is not rewritten | | `FINALIZING` | `COMPLETED` | Success inferred | PR exists or commits on branch | | `FINALIZING` | `FAILED` | Failure inferred | No commits, no PR, or agent reported error | @@ -156,7 +156,7 @@ Multiple timeout mechanisms work together to prevent runaway tasks. Substrate ti | Type | Default | Effect | |---|---|---| -| Max session duration | 8 hours | AgentCore caps a session at 8h; Lambda MicroVMs use `maximumDurationInSeconds: 28,800`, including suspended time. The orchestrator's safety-net poll loop runs up to `MAX_POLL_ATTEMPTS` (1020) × 30s ≈ 8.5h; a task still `RUNNING` when that window is exhausted is driven to `TIMED_OUT`. | +| Max session duration | 8 hours | AgentCore caps a session at 8h; Lambda MicroVMs use `maximumDurationInSeconds: 28,800`, including suspended time. MicroVM uses the saved absolute service deadline, so fast transition polling cannot shorten the session. Other backends retain the 1,020-attempt safety window. An exhausted approval wait uses FAILED, its allowed infrastructure-failure transition. | | Idle timeout | Backend-specific | AgentCore has an idle timeout. Lambda MicroVMs omit `idlePolicy` because inbound-traffic idleness would suspend an outbound-only agent while it is working. See Liveness monitoring. | | Max turns | 100 (range 1-500) | Agent stops after N model invocations. Configurable per task or per repo. | | Max cost budget | $0.01-$100 | Agent stops when budget is reached. Per-task or per-repo via Blueprint. | @@ -182,8 +182,8 @@ The orchestrator (`orchestrate-task.ts`) runs these as distinct durable-executio Validates the task before any compute is consumed. Checks run in order: 1. **Repo onboarding** - `GetItem` on `RepoTable`. If not found or inactive, reject with `REPO_NOT_ONBOARDED`. This runs at the API handler level (`createTaskCore`) for fast rejection. -2. **User concurrency** - Atomic check-and-increment on `UserConcurrency` counter. If at limit (default 10), the task is **queued, not failed** (#441): it transitions `SUBMITTED → QUEUED` and a scheduled admission-queue pickup Lambda re-attempts admission in FIFO order (by `created_at`) as slots free up, flipping `QUEUED → SUBMITTED` and re-invoking the orchestrator. The pickup Lambda does a read-only capacity pre-check; the orchestrator's atomic increment remains the single writer of the counter, so a pickup that loses the race harmlessly re-queues without losing FIFO position. `GET /tasks/{id}` surfaces `queue_position` and `estimated_wait_s` while queued. -3. **System concurrency** - Compare total running + hydrating tasks to the configured system limit and selected-backend quotas. +2. **User concurrency** - One transaction creates an internal `concurrency_slot` reservation on the task and increments `UserConcurrency.active_count`, subject to the configured cap (default 3). A retry reuses a held reservation. At the cap, the task transitions `SUBMITTED → QUEUED` and a scheduled pickup retries in FIFO order by `created_at`. Pickup and upload confirmation only inspect capacity; the orchestrator owns reservation acquisition. A task that acquired a reservation concurrently cannot be put back in the queue. `GET /tasks/{id}` surfaces `queue_position` and `estimated_wait_s` while queued. +3. **Backend capacity** - AWS also enforces the selected backend's service quotas when compute starts. 4. **Rate limiting** - Sliding window counter (10 tasks/hour per user). Rate-limit rejections happen at submit time and are rejected, not queued (unlike the concurrency cap, which queues). 5. **Idempotency** - If the request includes an idempotency key and a task with that key exists, return the existing task. @@ -209,7 +209,13 @@ The orchestrator resolves the repository's `ComputeStrategy` and calls `startSes AgentCore's session ID is pre-generated and reused on retry. ECS and Lambda MicroVMs use their substrate identifiers as session IDs. -If `RunMicrovm` succeeds but persisting the session handle or emitting the start event fails, the start step terminates the MicroVM best-effort using its in-memory handle before propagating the original error. This orphan reap is required because no later poll or finalization step can recover an unpersisted handle. +MicroVM starts first save an internal `microvm_start` receipt on the task: its stable client token (the task ID), a fingerprint of the request and full payload, creation time, and a local replay deadline. This happens before payload upload or `RunMicrovm`. Retries must match the saved fingerprint; a changed request cannot overwrite the earlier task's input. A returned handle is saved in the receipt and the normal task metadata before registration finishes. A replay can recover that handle without another start call. Registration uses strongly consistent reads to observe cancellation and already-committed writes. + +If a handle-save or registration response is lost, the code checks the committed task before terminating a known computer. If registration failed, cleanup remains best-effort and the receipt retains any saved handle for diagnosis. Failure to emit `session_started` alone does not terminate a registered MicroVM. MicroVM start failures persist the task outcome and reach `finalize-before-session`; they do not also release concurrency in the start step. Finalization begins with a strongly consistent read so a recently saved failure or cancellation is not reported using an older active state. + +An unanswered service request may already have created a MicroVM. Recovery reuses the same token and request within a **120-second local window**; after the window, it refuses another `RunMicrovm` call. This window is an application guard, not a verified AWS token-retention promise. An unrecovered request is reported as `MICROVM_START_OUTCOME_UNKNOWN` and requires inspection before submitting another task. Cancellation after an unanswered request records a `microvm_start_outcome_unknown` event. Without an ID, immediate termination cannot be guaranteed; the eight-hour service lifetime bound still applies. + +The MicroVM `start-session` step disables automatic durable **error** retries after its own recovery attempt. Crash replay is still possible and uses the saved receipt. Live AWS token-retention, changed-request and concurrent-conflict behavior remain verification gates. ### Step 5: Await completion @@ -221,7 +227,7 @@ The orchestrator polls for completion using `waitForCondition` from the Durable | ECS | `DescribeTasks`, including container exit status and exit code | | Lambda MicroVMs | `GetMicrovm` state plus agent heartbeat | -While waiting between polls, the durable orchestrator suspends without compute charges. If the session is terminated externally (crash, timeout, cancellation), the poll detects it and the orchestrator proceeds to finalization using GitHub-based result inference as fallback. +While waiting between polls, the durable orchestrator suspends without compute charges. If the session is terminated externally (crash, timeout, cancellation), the poll detects it and the orchestrator proceeds to finalization after a strongly consistent task read; it preserves an already committed terminal result. ### Step 6: Finalization @@ -250,9 +256,9 @@ After the session ends, the orchestrator determines the outcome from multiple si ### Step execution contract -Every step in the pipeline satisfies these properties: +Configured workflow steps target the following contract. The top-level durable orchestrator still needs explicit guards around external effects, as described under recovery below. -- **Idempotent** - Safe to retry after crashes. Context hydration produces the same prompt for the same inputs; session-start retry semantics are implemented by each backend strategy. +- **Replay-aware** - A retry must preserve task intent and avoid repeating external effects. Session-start recovery is implemented by each backend strategy; a checkpoint alone is not an idempotency guarantee. - **Timeout-bounded** - Each step has a configurable timeout to prevent blocking the pipeline. - **Failure-aware** - Returns `success` or `failed`. Infrastructure failures (throttle, transient errors) trigger exponential backoff retries (default: 2 retries, base 1s, max 10s). Explicit failures transition to `FAILED` without retry. - **Least-privilege input** - Each step receives only the `blueprintConfig` fields it needs. Custom Lambda steps get credential ARNs stripped. @@ -280,7 +286,9 @@ Liveness detection varies by compute backend. AgentCore sessions use DynamoDB he - **Grace period** (120s) - After entering `RUNNING`, the orchestrator waits before expecting heartbeats (covers container startup). - **Stale threshold** (240s) - If the heartbeat exists but is older than this, the session is treated as lost. -- **Early crash** - If no heartbeat is ever set after the combined window (360s), the agent died before the pipeline started. +- **Early crash** - If no heartbeat is ever set after the combined window (360s), the session is treated as lost; a process failure or failed DynamoDB writes can cause this. + +Approval waits suppress heartbeat writes. When the agent consumes a decision and restores `RUNNING`, its conditional transaction also refreshes `agent_heartbeat_at`, so the first poll after a long wait does not mistake the old timestamp for a crash. When the session is unhealthy, the task transitions to `FAILED` with "Agent session lost: no recent heartbeat." @@ -288,14 +296,50 @@ When the session is unhealthy, the task transitions to `FAILED` with "Agent sess **Lambda MicroVM state polling.** Liveness is a dual signal. The strategy maps `GetMicrovm` mechanically: `PENDING`/`RUNNING` report `running`, `SUSPENDING`/`SUSPENDED` report `suspended`, and `TERMINATING`/`TERMINATED` report terminal completion. The orchestrator supplies the health interpretation: -- `suspended` is healthy only while the task is `AWAITING_APPROVAL`; in any other task state it emits an anomaly and keeps polling rather than failing recoverable work. -- A terminal substrate report paired with a non-terminal task is a failure, but the orchestrator first re-reads the task row to confirm the agent did not write a terminal result between the original read and VM termination. -- Substrate state detects a dead VM; heartbeat staleness detects a hung, deadlocked, or OOM-killed pipeline inside a VM that still reports `RUNNING`. +- Intentional suspension requires the matching pending gate and saved suspend intent. Unexpected suspension emits one anomaly per episode and starts bounded wake recovery, preserving recoverable work. +- A terminal substrate report paired with a non-terminal task first checks for a complete, acknowledged approval checkpoint. Such a checkpoint can retire the old attempt and retain the task for a replacement. Without one, finalization strongly re-reads the task row before classifying a substrate failure. +- Substrate state detects a dead VM; heartbeat staleness detects loss of the heartbeat writer inside a VM that still reports `RUNNING`. The independent heartbeat thread can continue during a pipeline hang, so a fresh timestamp is not proof of progress. + +The P3 supervisor saves intent before control calls and rechecks the gate before and after them. Its durable state retains an absolute service lifetime, consecutive failures, recovery start time and next delay. Three failed cycles or 120 seconds of unconfirmed wake cannot become an indefinite wait. AWS RUNNING does not end recovery while the guest remains stuck on a decided/expired approval; fresh guest liveness is required. API approve/deny commit first, then attempt a bounded wake without changing the decision response. Automatic suspension defaults off via `microvm_approval_suspend_enabled`; disabling new sleep preserves wake and cleanup. See the [lifecycle diagnostics](/sample-autonomous-cloud-coding-agents/verification/645-p3-lifecycle-diagnostics). + +Unanswered approvals have no deadline by default. A checkpointed MicroVM wait can +retire after an hour, or before that worker's lifetime ends. Retirement and +replacement follow the [retained approval protocol](#retained-microvm-approvals). `TERMINATED` is the normal terminal signal and remains observable for at least 10 minutes. `ResourceNotFoundException` maps to completion only as a late fallback after the control-plane record is eventually reaped; polling does not wait for `NotFound`. **`/ping` health endpoint (AgentCore only).** The agent's FastAPI server responds to AgentCore's `/ping` calls while the coding task runs in a separate thread. AgentCore sees `HealthyBusy` and keeps the session alive. +### Retained MicroVM approvals + +A pending approval has no deletion timer by default. Closing its task cancels +unanswered requests and retains the decision history for 90 days. An explicit +positive approval deadline still applies; waking or replacing a worker does not +restart it. The agent decides whether the approved action remains relevant. + +At the approval barrier, the worker saves the conversation, exact pending tool +inputs, workflow context, cumulative usage and Git/workspace archive. The +coordinator verifies the checksummed S3 object versions before fencing the old +attempt through its coordinator-owned `worker-lease#` record. Workers may +read and condition-check that lease but cannot modify it. Only confirmed shutdown +allows the task to become `PARKED` and release its concurrency reservation. + +An answer admits one replacement when capacity permits. Admission, the new lease +and capacity reservation are one DynamoDB transaction; a deterministic Durable +execution name deduplicates invocation. The replacement uses the original +published coordinator and exact image version, restores the files/conversation, +and consumes the recorded answer. Remaining cost and turns are the original +allowance minus accumulated usage. Repository-free tasks preserve scratch files +using a private directory and local Git baseline. + +The scheduled continuation manager retries unfinished retirement, missed +dispatches and terminal cleanup, saving its scan cursor between invocations. +It cannot launch, suspend or resume workers; launching stays with the pinned +coordinator. A failed start response does not prove no worker exists. Capacity +remains reserved until shutdown is confirmed, or the full service lifetime has +elapsed for an unknown handle. The exact-attempt lease becomes `CLOSED` before +atomic release; `TERMINATING` alone is insufficient. + ### The idle timeout problem AgentCore terminates sessions after 15 minutes of inactivity. Since coding tasks may have long pauses between tool calls (builds, complex reasoning), the agent uses `add_async_task` to register background work. The SDK reports `HealthyBusy` via `/ping` while any async task is active, preventing idle termination. @@ -318,7 +362,7 @@ Long-running distributed systems fail. The orchestrator is designed so that ever | Hydration | Guardrail API unavailable | Fail the task (fail-closed: unscreened content never reaches agent) | | Session start | Selected compute service throttled | Exponential backoff. Fail after retries exhausted. | | Session start | Session crashes immediately | AgentCore: heartbeat never set, detected after 360s grace window. ECS: `DescribeTasks` reports failure. Lambda MicroVMs: `GetMicrovm` reports terminal state or the heartbeat never appears. | -| Running | Agent crashes mid-task | AgentCore: heartbeat goes stale. ECS: `DescribeTasks` reports stopped task. Lambda MicroVMs: `GetMicrovm` detects VM death and heartbeat staleness detects an in-guest hang. Finalization inspects GitHub for partial work. | +| Running | Agent crashes mid-task | AgentCore: heartbeat goes stale. ECS: `DescribeTasks` reports stopped task. Lambda MicroVMs: `GetMicrovm` detects VM death and heartbeat staleness detects loss of the in-guest writer. Finalization preserves committed task results and records a specific failure for an active lost session. | | Running | Agent hits turn or budget limit | Session ends normally. Finalize based on what was produced. | | Running | Idle for 15 min | AgentCore kills session. Task transitions to `TIMED_OUT`. | | Finalization | GitHub API down | Retry 3x. If still failing, mark `FAILED` with infrastructure reason. | @@ -327,14 +371,14 @@ Long-running distributed systems fail. The orchestrator is designed so that ever ### Recovery mechanisms 1. **Durable execution** - Lambda Durable Functions checkpoints at each state transition and replays after crashes. -2. **Idempotent operations** - All steps are safe to retry. +2. **Replay guards** - Operations need their own idempotency controls; checkpointing alone does not make external effects exactly-once. MicroVM starts use saved receipts. Capacity acquisition and release use task-owned markers updated atomically with the user counter. This guards the seat count; terminal audit events are still allowed to repeat on replay. 3. **Stuck-task scanner** - Periodic Lambda detects tasks stuck beyond expected durations and either resumes or fails them. -4. **Counter reconciliation** - Lambda runs every 15 minutes, compares counters to actual running task counts, corrects drift. Emits `counter_drift_corrected` CloudWatch metric. +4. **Counter reconciliation** - Every 15 minutes, the Lambda strongly scans counter records followed by task reservations. It repairs a count only if the saved counter revision is unchanged, then releases terminal held reservations. Structured logs record repairs, failures, empty counters and ambiguous legacy ownership; this handler does not publish a `counter_drift_corrected` metric. 5. **Dead-letter queue** - Tasks that exhaust retries go to DLQ for investigation. ## Concurrency and scaling -Each task runs in its own isolated compute session with no shared mutable state at the compute layer. The orchestrator manages concurrency purely at the coordination layer: atomic counters track how many tasks are active per user and system-wide, and admission control enforces limits before resources are consumed. +Each task runs in an isolated compute session. The orchestrator reserves capacity per user before starting compute; AWS separately enforces backend quotas. Approval waits keep their reservation, including when P3 suspends the MicroVM. This bounds unfinished sessions and their eventual resume demand; AWS memory-quota use while suspended remains unverified. ### Capacity limits @@ -342,15 +386,24 @@ Each task runs in its own isolated compute session with no shared mutable state |---|---|---| | `invoke_agent_runtime` TPS | 25 per agent/account | AgentCore quota (adjustable) | | Concurrent sessions | Account-level limit | AgentCore quota | -| Per-user concurrency | Configurable (default 3-5) | Platform config | -| System-wide max tasks | Configurable | Bounded by selected-backend quotas | +| Per-user concurrency | Configurable (default 3) | `MAX_CONCURRENT_TASKS_PER_USER` | ### Counter management -- **UserConcurrency** - DynamoDB item per user with `active_count`. Incremented atomically (`active_count < max`) at admission, decremented at finalization. -- **SystemConcurrency** - Single DynamoDB item, same pattern. +A reservation is a saved seat for one task. `task-concurrency.ts` owns both operations: + +- **Acquire:** require a matching owner, `SUBMITTED` status and no prior reservation; set `concurrency_slot.state = held` and increment the counter in the same transaction. A held reservation is reused on replay. A released task cannot reserve again; a new attempt gets a new task ID. +- **Release:** require a terminal task and a held reservation; mark it `released` and decrement the counter in the same transaction. Normal finalization, early failure and stranded cleanup use this helper. A missing or already-released marker does not decrement. An active task keeps its seat. + +Finalization attempts release even if an audit event fails. A failed status write that leaves the task active is propagated for retry. Cancellation before admission/pre-flight completion also checks for a held terminal reservation. If a crash separates the terminal write from release, the scheduled counter reconciler completes release later. + +Every reservation change writes a fresh `reservation_version` on the counter. Reconciliation uses strongly consistent **base-table** scans, once for counters and once for tasks, then compares the saved revision before repairing a count. A count-only comparison cannot detect an increment followed by a decrement. Held terminal reservations are included until their release commits. A partial scan never installs a partial count. These scans consume read capacity across retained rows; check scan duration and capacity at deployment scale. + +If a release discovers an empty/missing counter, it closes the marker without subtracting from seats reserved meanwhile. A concurrent repair can leave a conservative overcount until the next sweep. `CONCURRENCY_EMPTY_COUNTER` records that condition. + +Older active records without reservation markers are ambiguous: status alone does not prove admission. The helper never guesses that they own a seat; reconciliation skips count repair for that user and logs `CONCURRENCY_RESERVATION_UNKNOWN`. Pause new submissions and drain old executions before deploying this protocol across all counter writers. After old tasks settle, reconcile and resume admissions. Rollback also requires draining tasks using the newer protocol; mixing old direct decrements with new markers does not provide this guarantee. -Concurrency is always released in `finalizeTask` (step 6), never inside the poll loop. ECS poll failure paths call `failTask` with `releaseConcurrency: false` to transition the task to `FAILED` without decrementing — `finalizeTask` handles the single decrement after re-reading the task state. The heartbeat-detected crash path also guards against double-decrement by only releasing the counter after a successful state transition. If the transition fails (task already terminal), it re-reads and acts accordingly. +This protocol protects against replay and competing **cooperative** writers. The agent role can currently update/replace its own task row, so internal fields are not yet protected against a compromised agent. Coordinator-only metadata storage or constrained agent writes remain a separate security prerequisite. ## Implementation @@ -402,7 +455,7 @@ At 500 concurrent tasks, peak TPS is ~16.7 - well within the 25 TPS AgentCore qu ## Data model -Three DynamoDB tables back the orchestrator: one for task state, one for the audit log, and one for concurrency counters. The Tasks table is the source of truth for every task; the orchestrator reads and writes it at every state transition. TaskEvents is append-only and powers the `GET /v1/tasks/{id}/events` API. UserConcurrency is a lightweight counter table used only during admission and finalization. +Three DynamoDB tables back the orchestrator: one for task state, one for the audit log, and one for concurrency counters. The Tasks table is the source of truth for every task; the orchestrator reads and writes it at every state transition. TaskEvents is append-only and powers the `GET /v1/tasks/{id}/events` API. UserConcurrency stores reservation counts and revision tokens used by admission, cleanup and reconciliation. ### Tasks table (DynamoDB) @@ -420,7 +473,8 @@ Three DynamoDB tables back the orchestrator: one for task state, one for the aud | `branch_name` | String | `bgagent/{task_id}/{slug}` for new tasks; PR's `head_ref` for PR tasks | | `session_id` | String? | Backend session identifier (AgentCore session ID, ECS task ARN, or MicroVM ID) | | `compute_type` | String? | Selected backend: `agentcore`, `ecs`, or `lambda-microvm` | -| `compute_metadata` | Map? | Backend lifecycle handle; Lambda MicroVMs persist `microvmId` and `endpoint` | +| `compute_metadata` | Map? | Backend lifecycle handle; Lambda MicroVMs persist `microvmId`, `endpoint`, actual image identity and verified lifecycle protocol when available | +| `concurrency_slot` | Map? | Internal reservation `{state, acquired_at, released_at?}`; excluded from public task responses | | `execution_id` | String? | Durable execution ID | | `pr_url` | String? | PR URL (set during finalization) | | `error_message` | String? | Error reason if FAILED | @@ -456,7 +510,7 @@ Append-only audit log. See [OBSERVABILITY.md](/sample-autonomous-cloud-coding-ag | Field | Type | Description | |---|---|---| | `user_id` (PK) | String | User ID | -| `active_count` | Number | Running task count | +| `active_count` | Number | Held reservation count, including approval waits and pending terminal cleanup | +| `reservation_version` | String? | Fresh revision token on every reservation mutation or count repair | -Increment: `SET active_count = active_count + 1` with `ConditionExpression: active_count < :max`. -Decrement: `SET active_count = active_count - 1` with `ConditionExpression: active_count > 0`. +Counter changes belong to the reservation transactions described above. Do not add a standalone increment or decrement: it bypasses per-task replay protection. diff --git a/docs/src/content/docs/architecture/Registry.md b/docs/src/content/docs/architecture/Registry.md index e6e2457ac..1d0ac94fd 100644 --- a/docs/src/content/docs/architecture/Registry.md +++ b/docs/src/content/docs/architecture/Registry.md @@ -41,6 +41,12 @@ A **registry asset** is a versioned, immutable-per-version runtime artifact that **Cedar parity.** Registry Cedar text reaches the agent through the **same** `cedar_policies` payload field as inline blueprint policies, so it is byte-identical from the `PolicyEngine`'s view. **Skills** are prompt text only: a skill cannot invoke tools; its `tool_hints` are advisory prose referencing tools an MCP server separately provides (no transitive dependency — the operator attaches both). +**Runtime network support (#818).** An MCP server is a program the coding agent calls to use extra tools. Remote `http`/`sse` assets need an **HTTPS endpoint reachable on TCP port 443** under the shipped **AgentCore, ECS and Lambda MicroVM** network policies. Remote ports such as 80, 8080 or 8443 are unsupported on all three defaults; this is not a MicroVM-only limitation. A `stdio` asset starts a local program and uses its input/output pipes, but that program's outbound network requests still face the runtime policy. The MicroVM image builder's separate 80/443 permission does not apply to task execution. + +Registry resolution validates/pins the asset and the loader writes its configuration; neither probes network reachability or rejects an asset merely because its URL uses a different port. This is a documented support constraint, not a new synth/onboarding validator. Use a reachable HTTPS/443 service or proxy and verify DNS, routing, TLS and authentication on the deployed backend. Port 443 alone does not prove reachability. + +**Task delivery.** `resolved_assets` stays in the shared agent payload. ECS and MicroVM now deliver every payload through v2 authenticated bootstrap and a single-object signed URL. Registry content may exceed 4,096 bytes without expanding the MicroVM hook reference; the 8 MiB task-document cap still applies. Local integration coverage resolves a large MCP asset, runs the real payload assembly, verifies S3 bytes and the small hook reference, then separately exercises the Python hook mapper and local `.mcp.json` loader. That proves delivery/configuration, not a live connection to the remote tool. See [COMPUTE.md](/sample-autonomous-cloud-coding-agents/architecture/compute) for runtime behavior. + ## 3. Substrate mapping (the core design) Agent Registry answers *"what servers/skills exist, find me one"* (discovery metadata + semantic search). ABCA needs *"give me the exact runtime config to load this pinned asset."* These are two different objects, and the service validates the discovery object against the official schemas (an MCP record's body must be a valid MCP `server.json`, not our `.mcp.json`). So **every record carries BOTH a discovery descriptor AND ABCA's runtime payload.** diff --git a/docs/src/content/docs/architecture/Repo-onboarding.md b/docs/src/content/docs/architecture/Repo-onboarding.md index 03d04de47..2da230aec 100644 --- a/docs/src/content/docs/architecture/Repo-onboarding.md +++ b/docs/src/content/docs/architecture/Repo-onboarding.md @@ -7,7 +7,7 @@ title: Repo onboarding Before users can submit tasks for a repository, that repository must be onboarded to the platform. Onboarding registers the repo and produces a per-repo configuration that the orchestrator uses at task time: compute strategy, model, credentials, networking, and pipeline customizations. If a user submits a task for a non-onboarded repo, the API returns `422 REPO_NOT_ONBOARDED`. - **Use this doc for:** the Blueprint construct interface, RepoConfig schema, override precedence, compute strategy interface, and pipeline customization model. -- **For practical usage:** see [Quick Start](../guides/QUICK_START.mdx) for onboarding your first repo and [User Guide](/sample-autonomous-cloud-coding-agents/using/overview) for per-repo overrides. +- **For practical usage:** see [Quick Start](/sample-autonomous-cloud-coding-agents/getting-started/quick-start) for onboarding your first repo and [User Guide](/sample-autonomous-cloud-coding-agents/using/overview) for per-repo overrides. - **Related docs:** [ORCHESTRATOR.md](/sample-autonomous-cloud-coding-agents/architecture/orchestrator) for how the orchestrator consumes blueprint config, [COMPUTE.md](/sample-autonomous-cloud-coding-agents/architecture/compute) for compute backends, [SECURITY.md](/sample-autonomous-cloud-coding-agents/architecture/security) for custom step trust boundaries. ## Why onboarding? @@ -36,11 +36,11 @@ For operators with IAM access to the deployed stack, `bgagent repo onboard` and | **Soft-delete** | `status=removed` + 30-day TTL on stack removal | Same semantics on `offboard` | | **Custom runtime / token IAM** | `additionalRuntimeArns` / `additionalSecretArns` in CDK | Stored in the row, but orchestrator IAM still requires a CDK deploy | | **Cedar, egress, pipeline steps** | Supported via construct props | Not exposed — use CDK | -| **Audit trail** | CloudFormation change set + deploy logs | CLI stdout only today (see [ADR-017](/sample-autonomous-cloud-coding-agents/architecture/adr-017-operator-cli-repo-onboarding)) | +| **Audit trail** | CloudFormation change set + deploy logs | CLI stdout only today (see [ADR-017](/sample-autonomous-cloud-coding-agents/decisions/adr-017-operator-cli-repo-onboarding)) | Use the CLI path for quick day-2 registration with platform defaults (runtime ARN, GitHub token secret). Use CDK when the repo needs durable infrastructure, custom IAM, Cedar policies, egress rules, or pipeline customization. The onboard command prints notes explaining which platform defaults apply and when a redeploy is still required. -See also: [Using the CLI — operator commands](/sample-autonomous-cloud-coding-agents/using/overview#operator-commands-stack-admin) and [ADR-017](/sample-autonomous-cloud-coding-agents/architecture/adr-017-operator-cli-repo-onboarding). +See also: [Using the CLI — operator commands](/sample-autonomous-cloud-coding-agents/using/using-the-cli#operator-commands-stack-admin) and [ADR-017](/sample-autonomous-cloud-coding-agents/decisions/adr-017-operator-cli-repo-onboarding). ### Blueprint construct diff --git a/docs/src/content/docs/architecture/Security.md b/docs/src/content/docs/architecture/Security.md index d7a6d04e1..e78c0c972 100644 --- a/docs/src/content/docs/architecture/Security.md +++ b/docs/src/content/docs/architecture/Security.md @@ -41,16 +41,22 @@ Three authentication mechanisms protect the platform, matching its input channel **Authorization** is user-scoped: any authenticated user can submit tasks, but users can only view and cancel their own tasks (`user_id` enforcement). Both webhook and API key management enforce ownership with 404 (not 403) to avoid leaking resource existence. A platform API key inherits its creator's `user_id`, so a webhook created via a key is attributed to that owner exactly as an interactive session would be. Key scopes gate which routes a key may call (Phase 1: `webhooks:manage`); reserved scopes (`tasks:read`, `tasks:cancel`) are validated but not yet wired to any route. -**Agent credentials** - GitHub access currently uses a PAT stored in Secrets Manager. The orchestrator reads the secret at hydration time and passes it to the agent runtime. The model never receives the token in its context. Planned: replace the shared PAT with a GitHub App via AgentCore Identity Token Vault, providing per-task, repo-scoped, short-lived tokens (see [GitHub issues](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues), the [ADR-016](/sample-autonomous-cloud-coding-agents/architecture/adr-016-pluggable-identity-and-auth) two-seam design, and the [IDENTITY_AND_AUTH.md](/sample-autonomous-cloud-coding-agents/architecture/identity-and-auth) worked examples). +**Agent credentials** - GitHub access currently uses a PAT stored in Secrets Manager. The orchestrator reads the secret at hydration time and passes it to the agent runtime. The model never receives the token in its context. Planned: replace the shared PAT with a GitHub App via AgentCore Identity Token Vault, providing per-task, repo-scoped, short-lived tokens (see [GitHub issues](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues), the [ADR-016](/sample-autonomous-cloud-coding-agents/decisions/adr-016-pluggable-identity-and-auth) two-seam design, and the [IDENTITY_AND_AUTH.md](/sample-autonomous-cloud-coding-agents/architecture/identity-and-auth) worked examples). **Per-session IAM scoping** - The agent does not use its long-lived compute role (the AgentCore Runtime `ExecutionRole`, ECS Fargate task role, or Lambda MicroVMs execution role) for tenant data. Instead, at task startup it assumes a per-task **SessionRole** via `sts:AssumeRole` with session tags `{user_id, repo, task_id}`, and uses the resulting short-lived credentials for all DynamoDB and S3 tenant-data access. The SessionRole's policies self-constrain on those tags: -- **DynamoDB**: item access on the four `task_id`-partitioned tables (task, events, approvals, nudges) is gated by a `dynamodb:LeadingKeys` condition equal to `${aws:PrincipalTag/task_id}`, so a session can read or write only its own task's rows. `Scan` is not granted (it ignores leading-keys). `task_id` is the isolation boundary because it is the base-table partition key — `LeadingKeys` cannot bind to a GSI partition key such as `user_id`. +- **DynamoDB**: item access on the four `task_id`-partitioned tables (task, events, approvals, nudges) is gated by a `dynamodb:LeadingKeys` condition equal to `${aws:PrincipalTag/task_id}` and requires that context key to be present. `Scan` is not granted. The main task table permits reads and only attribute-scoped `UpdateItem` writes: the agent can report status, heartbeat, results and approval details, but cannot replace/delete the row or edit coordinator-owned start receipts, capacity reservations, owner identity or compute handles. The reviewed write list is `cdk/src/constructs/agent-task-write-attributes.json`; Python tests exercise the current writers against it. Events and nudges retain task-scoped item writes. Approval records permit reads and transaction condition checks only. `LeadingKeys` binds the base-table partition key, not a GSI key such as `user_id`. +- **Approval requests**: workers call an IAM-signed API path restricted to their task tag. Its trusted handler can create `PENDING` or conditionally record `TIMED_OUT`; human decisions, notification markers and retention fields are unavailable to the worker. IAM binds the caller to the tagged task path. The transaction guards against concurrent task ownership/state changes and checks MicroVM worker leases. See the [approval trust boundary](/sample-autonomous-cloud-coding-agents/architecture/cedar-hitl-gates#121-trust-boundaries) and [upgrade procedure](/sample-autonomous-cloud-coding-agents/getting-started/deployment-guide#upgrading-approval-permissions). - **S3**: trace writes and attachment reads are scoped to the `/${aws:PrincipalTag/user_id}/` object prefix. -The compute role retains only non-tenant access (Bedrock model invocation — already ARN-scoped; CloudWatch Logs; the GitHub PAT secret, read once before the SessionRole is assumed; AgentCore Memory) plus `sts:AssumeRole`/`sts:TagSession` on the SessionRole. Because the agent runs under credentials that are themselves an assumed role, its `AssumeRole` is *role chaining* — capped at one hour regardless of the role's max session duration — so the agent uses a **refreshable** credential provider that re-assumes before expiry (tasks can run up to the 8-hour `maxLifetime`). The design is backend-agnostic: the same SessionRole and agent code serve all three backends (AgentCore Runtime, ECS Fargate, Lambda MicroVMs). A compromised agent session is therefore confined to its own task's data, enforced at the IAM layer rather than by application-code conventions. The policy structure (the `dynamodb:LeadingKeys` condition on `${aws:PrincipalTag/task_id}`, per-user S3 prefixes, and `Scan` exclusion) is asserted by CDK template tests, and the refreshable-credential and session-tag flow by agent unit tests; the matching-tag → allow / mismatched-or-absent-tag → deny behaviour was additionally confirmed once via the IAM policy simulator during development. +The compute role retains only non-tenant access (Bedrock model invocation — already ARN-scoped; CloudWatch Logs; the GitHub PAT secret, read once before the SessionRole is assumed; AgentCore Memory) plus `sts:AssumeRole`/`sts:TagSession` on the SessionRole. Because the agent runs under credentials that are themselves an assumed role, its `AssumeRole` is *role chaining* — capped at one hour regardless of the role's max session duration — so the agent uses a **refreshable** credential provider that re-assumes before expiry (tasks can run up to the 8-hour `maxLifetime`). The design is backend-agnostic: the same SessionRole and agent code serve all three backends (AgentCore Runtime, ECS Fargate, Lambda MicroVMs). Existing scoped credentials are constrained to their tags, but the compute role chooses those tags when assuming the role. The trust policy does not independently bind a worker to one task; this is not complete isolation against a compromised worker with ambient credentials. No worker needs direct access to the shared capacity counter. ECS's configuration without a session role retains unscoped task reads/reporting updates, with the same attribute restriction, but cannot enable approval gates; the construct rejects approval wiring without the session role and service URL. The policy structure and Python writer compatibility are tested locally. A historical IAM simulator check covered the original tag conditions; it does not verify the new attribute restriction or deployed transaction authorization. The [acceptance checklist](/sample-autonomous-cloud-coding-agents/verification/readme#live-acceptance-for-an-installation) includes effective-role authorization checks. + +**Lambda MicroVMs compute-role delta** — The compute role additionally reads the per-workspace channel-OAuth secrets (`bgagent-linear-oauth-*`, `bgagent-jira-oauth-*`) and its payload bucket's `bootstrap/*` manifests. Ambient task-object reads and payload-bucket listing are explicitly denied. It also has read-only `ec2:DescribeAvailabilityZones` for repository CDK synthesis; that API has no resource-level scope. + +Recorded live calls rejected source-conditioned MicroVM role trust and service-conditioned PassRole grants. The integration therefore omits those conditions on the affected paths, while restricting which roles can be passed and which resources they may access. Reintroduce conditions only after verifying service support. [ADR-021 §4](/sample-autonomous-cloud-coding-agents/decisions/adr-021-lambda-microvms-compute-backend#4-infra-and-iam-conditional-resources-behind-bootstrap-computetypes) records the role responsibilities and limits. + +**ECS/MicroVM task bootstrap (v2)** — The coordinator publishes non-secret deployment manifests and supplies a short-lived signed URL for exactly one task's payload. The worker reads the manifest with its ambient role, whose explicit deny outside its own `bootstrap/*` prevents a foreign public bucket from authorizing fake configuration. It verifies downloaded task identity and exact manifest/config agreement before installing MicroVM settings. The coordinator privately persists the URL outside TaskTable for retry and deletes both task objects at finalization. Worker roles cannot read/list task objects using their ambient credentials. A signed URL is a bearer capability: redact it in errors and keep it out of task rows and repository subprocess environments. Existing session-tag trust and other platform grants remain separate limitations. The [payload guide](/sample-autonomous-cloud-coding-agents/verification/645-payload-bootstrap) describes coordinated image/coordinator/policy upgrades; the [acceptance summary](/sample-autonomous-cloud-coding-agents/verification/readme) distinguishes recorded AWS checks from remaining validation. -**Lambda MicroVMs compute-role delta** - On this backend the compute role additionally holds prefix-scoped `secretsmanager:GetSecretValue` on the per-workspace channel-OAuth secrets (`bgagent-linear-oauth-*`, `bgagent-jira-oauth-*`), read access to the `/run` payload bucket (the task payload arrives as an S3 object, not as environment), and `ec2:DescribeAvailabilityZones` (`Resource: *`, read-only — EC2 describe actions have no resource-level scoping) for a CDK repo's own synth gate. It is also **the only compute role in the platform whose trust policy carries no confused-deputy condition**: the Lambda MicroVMs service populates no `aws:SourceAccount` or `aws:SourceArn` when it assumes the role, so a trust policy carrying one is unassumable — verified live, not assumed (connector creation failed deterministically, and `RunMicrovm` surfaced the same root cause as a misleading caller-side `iam:PassRole` denial). The compensating controls are that each of the three MicroVM roles can be passed to `lambda.amazonaws.com` only by a named principal — the orchestrator, via an `iam:PassRole` scoped to the execution role's exact ARN, and the CloudFormation deployment role for the build and connector-operator roles — that every other resource they reach is account-scoped by ARN apart from two justified `Resource: *` read/create-time statements, and that none of them holds `iam:*`, cross-account trust, or any `sts:AssumeRole` beyond the execution role's scoped hop to the per-task SessionRole. Full evidence, the two-arm PassRole experiment, and the alternatives considered are in [ADR-021 §4](/sample-autonomous-cloud-coding-agents/architecture/adr-021-lambda-microvms-compute-backend#4-infra-and-iam-conditional-resources-behind-bootstrap-computetypes). > Out of scope for this control and tracked separately as GitHub issues: replacing the shared GitHub PAT (GitHub App / Token Vault), binding credentials to the MicroVM via attestation, and scoping AgentCore Memory (namespace isolation by `actorId`/`sessionId` remains its boundary). diff --git a/docs/src/content/docs/architecture/Vision.md b/docs/src/content/docs/architecture/Vision.md index f003616d4..f6d8a5f5c 100644 --- a/docs/src/content/docs/architecture/Vision.md +++ b/docs/src/content/docs/architecture/Vision.md @@ -7,7 +7,7 @@ title: Vision This document states the long-term direction of **ABCA (Autonomous Background Coding Agents on AWS)** and the **tenets** that should guide design, implementation, and review. Use it when evaluating pull requests, RFCs, and ADRs: if a change clearly advances the vision and respects the tenets, it belongs; if it trades tenets away without an explicit, documented rationale, it needs more discussion. - **Use this doc for:** alignment checks in review — “does this fit where we are going?” -- **Not a substitute for:** [ARCHITECTURE.md](/sample-autonomous-cloud-coding-agents/architecture/architecture) (system shape), [GitHub issues](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues) (planned work and priorities), or [docs/decisions/](../decisions/) (specific accepted choices). +- **Not a substitute for:** [ARCHITECTURE.md](/sample-autonomous-cloud-coding-agents/architecture/architecture) (system shape), [GitHub issues](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues) (planned work and priorities), or [docs/decisions/](/sample-autonomous-cloud-coding-agents/decisions/readme) (specific accepted choices). ## Vision @@ -17,7 +17,7 @@ We are building toward **lights-sparse**, **graduated** autonomy (defined below) ### What "lights-sparse" means -**Lights-sparse** is project vocabulary (not general industry jargon): it names the autonomy posture ABCA targets today, drawn from the **software dark factory** analogy in the [introduction](/sample-autonomous-cloud-coding-agents/architecture/index). +**Lights-sparse** is project vocabulary (not general industry jargon): it names the autonomy posture ABCA targets today, drawn from the **software dark factory** analogy in the [introduction](/sample-autonomous-cloud-coding-agents/). - **Lights-out** (the analogy’s end state): humans set goals, policy, and constraints; production runs without people on the floor. - **Lights-sparse** (where teams are now): the **implementation loop** — edit code, run tests, open pull requests — is increasingly **unattended**, while **governance, merge authority, and production release** stay **supervised**. Humans are not at the keyboard for every step; they are still accountable for what ships. @@ -148,7 +148,7 @@ These are out of scope for the project vision. Proposals that primarily serve th | [SECURITY.md](/sample-autonomous-cloud-coding-agents/architecture/security) | Threat model and controls (tenets 4–5 in depth) | | [CEDAR_HITL_GATES.md](/sample-autonomous-cloud-coding-agents/architecture/cedar-hitl-gates) | HITL approval gates, pre-approve scopes, graduated in-run autonomy | | [INTERACTIVE_AGENTS.md](/sample-autonomous-cloud-coding-agents/architecture/interactive-agents) | Async UX, watch/nudge, notification plane, approval state machine | -| [docs/decisions/](../decisions/) | Recorded choices when tenets conflict or ambiguity is resolved | -| [docs/src/content/docs/index.md](/sample-autonomous-cloud-coding-agents/architecture/index) (synced intro) | Public-facing narrative including dark-factory attribute table | +| [docs/decisions/](/sample-autonomous-cloud-coding-agents/decisions/readme) | Recorded choices when tenets conflict or ambiguity is resolved | +| [docs/src/content/docs/index.md](/sample-autonomous-cloud-coding-agents/) (synced intro) | Public-facing narrative including dark-factory attribute table | When tenets and architecture principles overlap, **tenets win for review judgment**; **architecture and ADRs win for implementation detail** once a direction is chosen. diff --git a/docs/src/content/docs/architecture/Workflows.md b/docs/src/content/docs/architecture/Workflows.md index 255044f6b..4aa265e8c 100644 --- a/docs/src/content/docs/architecture/Workflows.md +++ b/docs/src/content/docs/architecture/Workflows.md @@ -10,7 +10,7 @@ The three former task types — `new_task`, `pr_iteration`, `pr_review` — are - **Use this doc for:** the workflow file schema, step-kind catalog, the agent-side step runner model, and how a `workflow_ref` flows from API to agent. - **Related docs:** [ARCHITECTURE.md](/sample-autonomous-cloud-coding-agents/architecture/architecture) for the deterministic-steps-wrapping-one-agentic-step model, [ORCHESTRATOR.md](/sample-autonomous-cloud-coding-agents/architecture/orchestrator) for the durable lifecycle the workflow runs inside, [REPO_ONBOARDING.md](/sample-autonomous-cloud-coding-agents/architecture/repo-onboarding) for the per-repo **Blueprint** (a distinct concept — see [Naming](#naming-workflow-vs-blueprint)), [CEDAR_HITL_GATES.md](/sample-autonomous-cloud-coding-agents/architecture/cedar-hitl-gates) for the policy engine a workflow's `agent_config` feeds, [SECURITY.md](/sample-autonomous-cloud-coding-agents/architecture/security) for tool tiers, and [API_CONTRACT.md](/sample-autonomous-cloud-coding-agents/architecture/api-contract) for the `workflow_ref` wire field. -- **Decision record:** [ADR-014](/sample-autonomous-cloud-coding-agents/architecture/adr-014-workflow-driven-tasks). +- **Decision record:** [ADR-014](/sample-autonomous-cloud-coding-agents/decisions/adr-014-workflow-driven-tasks). - **Tracking issue:** [#248](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/248). Pairs with the agent asset registry ([#246](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/246)) and attribution ([#245](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/245)). Scoped-down, current-architecture track of the broader AKW vision ([#99](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/99)). ## Background: what workflows replaced @@ -229,7 +229,7 @@ This second example is the **target shape** for repo-less execution — the acce ## The agent-side step runner -Per [ADR-014](/sample-autonomous-cloud-coding-agents/architecture/adr-014-workflow-driven-tasks), the runner is **agent-side**: it lives in the container and interprets `workflow.steps`. The orchestrator's durable shape (`admission-control → pre-flight → hydrate-context → start-session → await-agent-completion → finalize`) is unchanged — the workflow drives *what happens inside* the `RUNNING` state, not the platform lifecycle. This keeps the blast radius off durable orchestration and matches the issue's "executes steps in order *inside the container*." +Per [ADR-014](/sample-autonomous-cloud-coding-agents/decisions/adr-014-workflow-driven-tasks), the runner is **agent-side**: it lives in the container and interprets `workflow.steps`. The orchestrator's durable shape (`admission-control → pre-flight → hydrate-context → start-session → await-agent-completion → finalize`) is unchanged — the workflow drives *what happens inside* the `RUNNING` state, not the platform lifecycle. This keeps the blast radius off durable orchestration and matches the issue's "executes steps in order *inside the container*." ```python # agent/src/workflow/runner.py (shape, not final code) @@ -254,7 +254,7 @@ The step runner runs inside the compute substrate, which is **not** a throwaway - **Step completion is checkpointed; resume skips completed steps.** The runner records each step's outcome to a small `workflow_state.json` on the persistent mount (`/mnt/workspace`) as it goes. On resume (orchestrator re-invokes the same session, or — when shipped — a replacement worker rehydrates from the [S3-backed SDK session store](#relationship-to-portable-resume)), the runner reads that checkpoint, **skips already-completed deterministic steps** (`clone_repo` need not re-clone a populated `/workspace`; a completed `verify_build` is not re-run), and **resumes the agent loop** via the persisted SDK session UUID rather than restarting it from turn 0. This is the same property the orchestrator already relies on for session start being idempotent (pre-generated, reused session id). - **Side-effecting steps remain idempotent.** Independent of resume, `clone_repo`, `ensure_pr`, `post_review`, and `deliver_artifact` must tolerate a partial prior run (a resume can re-enter the step that was in flight when the worker died). Each documents its idempotency key — PR branch, review id, artifact S3 key = `task_id` — so re-entry reconciles rather than duplicates (today's `ensure_pr` already does this: it checks `gh pr view` before creating). - **`on_failure: continue` is forbidden after side effects** (validation rule 10). A failed `ensure_pr` (commits pushed, PR-create failed) must not reach a *succeeded* terminal — committed work with no PR and no compensation. `continue` is permitted only for non-side-effecting, advisory steps (e.g. an informational `verify_lint`). `skip_remaining` ends the workflow cleanly and runs terminal-outcome resolution against whatever completed; `fail` (default) is terminal `FAILED`. -- **Granularity boundary.** Resume is *workflow-step granular on the agent side*, not a new orchestrator-side durable checkpoint per step — the orchestrator still treats the whole session as one `await-agent-completion` step, so platform invariants stay agent-external ([ADR-014](/sample-autonomous-cloud-coding-agents/architecture/adr-014-workflow-driven-tasks)). What changes versus today is that the agent-side runner makes its *own* progress recoverable across a stop/resume, which today's monolithic `run_task` does not. +- **Granularity boundary.** Resume is *workflow-step granular on the agent side*, not a new orchestrator-side durable checkpoint per step — the orchestrator still treats the whole session as one `await-agent-completion` step, so platform invariants stay agent-external ([ADR-014](/sample-autonomous-cloud-coding-agents/decisions/adr-014-workflow-driven-tasks)). What changes versus today is that the agent-side runner makes its *own* progress recoverable across a stop/resume, which today's monolithic `run_task` does not. #### Relationship to portable resume @@ -320,7 +320,7 @@ Scope discipline for #248: **`github` is the only implemented provider** — add Read-only is enforced by Cedar hard-deny rules. **As of #248 Phase 2a** these key off the `context.read_only` attribute (`read_only_forbid_write`, `read_only_forbid_edit`), not a principal literal — and `read_only: true` *also* makes the runner drop `Write`/`Edit` from the SDK `allowed_tools` list. Two layers: - **Defense in depth.** `read_only: true` makes the runner drop `Write`/`Edit` from `allowed_tools` *and* sends `context.read_only == true` on every Cedar request — closing the earlier gap where read-only was enforced only by a Cedar string-match on the principal, not by the tool list. -- **Property-keyed enforcement (security-relevant — was precise, not hand-waved).** Read-only enforcement attaches to the *property*, not a per-task-type literal: the principal keeps the legacy `Agent::TaskAgent::""` identity scheme (audit/attribution only), while the two hard-deny rules forbid `Write`/`Edit` **whenever `context.read_only == true`**. So the deny applies uniformly to *every* read-only workflow — not just `coding/pr-review` — and there is no literal a new read-only workflow could fail to match. This was a deliberate, recorded behavior change (see [ADR-014](/sample-autonomous-cloud-coding-agents/architecture/adr-014-workflow-driven-tasks) addendum 2026-06-08), gated by the `contracts/cedar-parity/` fixtures (`read-only-forbid-write`, `read-only-forbid-edit`, `read-only-false-permits-write`) run against *both* the `cedarpy` and `cedar-wasm` engines. +- **Property-keyed enforcement (security-relevant — was precise, not hand-waved).** Read-only enforcement attaches to the *property*, not a per-task-type literal: the principal keeps the legacy `Agent::TaskAgent::""` identity scheme (audit/attribution only), while the two hard-deny rules forbid `Write`/`Edit` **whenever `context.read_only == true`**. So the deny applies uniformly to *every* read-only workflow — not just `coding/pr-review` — and there is no literal a new read-only workflow could fail to match. This was a deliberate, recorded behavior change (see [ADR-014](/sample-autonomous-cloud-coding-agents/decisions/adr-014-workflow-driven-tasks) addendum 2026-06-08), gated by the `contracts/cedar-parity/` fixtures (`read-only-forbid-write`, `read-only-forbid-edit`, `read-only-false-permits-write`) run against *both* the `cedarpy` and `cedar-wasm` engines. This is the migration step where an error *silently weakens* enforcement (the rule stops matching) rather than failing loudly. The original plan was to ship it as an isolated PR ahead of the Phase 2b workflow migrations; because 2b shipped first behind a `read_only ⇒ "pr_review"` principal bridge (so read-only was never unprotected), Phase 2a instead removes that bridge and lands the property-keyed rules + parity fixtures together on the #248 branch. See the ADR-014 addendum and [Phasing](#phasing). @@ -334,7 +334,7 @@ Registry-sourced `cedar_policy_modules` / `mcp_servers` are trusted content load ### Authorship & governance -A workflow file selects the agent's tool surface and policy posture, so **who may publish a `production` workflow is a trust decision, not a convenience**. Per [ADR-003](/sample-autonomous-cloud-coding-agents/architecture/adr-003-contribution-governance), publishing or promoting a first-party workflow follows the same issue → approval → review → merge path as any code change — a workflow YAML in `agent/workflows/**` is reviewed like code, and the synth-time validator (the [validation rules](#validation-rules)) is a required CI gate. When the registry (#246) makes workflows publishable out-of-band, publish/promote ACLs are Cedar-governed per #246 Phase 3; until then, the only way a `production` workflow exists is through a reviewed merge. The `description`/`guidance` discovery fields are author-controlled free text; when they feed an agent's workflow-*selection* context (Phase 4), they are treated as untrusted-external input and screened like other hydrated content. +A workflow file selects the agent's tool surface and policy posture, so **who may publish a `production` workflow is a trust decision, not a convenience**. Per [ADR-003](/sample-autonomous-cloud-coding-agents/decisions/adr-003-contribution-governance), publishing or promoting a first-party workflow follows the same issue → approval → review → merge path as any code change — a workflow YAML in `agent/workflows/**` is reviewed like code, and the synth-time validator (the [validation rules](#validation-rules)) is a required CI gate. When the registry (#246) makes workflows publishable out-of-band, publish/promote ACLs are Cedar-governed per #246 Phase 3; until then, the only way a `production` workflow exists is through a reviewed merge. The `description`/`guidance` discovery fields are author-controlled free text; when they feed an agent's workflow-*selection* context (Phase 4), they are treated as untrusted-external input and screened like other hydrated content. ## Wire contract: `workflow_ref` from API to agent @@ -460,7 +460,7 @@ So: JSON Schema = canonical shape, consumed not copied; cross-field rules = one ## Promotion is earned, not set -`status: production` is not a label an author flips — it is a state a version *earns* by passing its declared `promotion_gate`. This makes the promotion lifecycle (`draft → validated → production → deprecated`) a machine-checked quality gate rather than a human's say-so, and it slots directly onto the existing [tiered validation pyramid](/sample-autonomous-cloud-coding-agents/architecture/adr-013-tiered-validation-pyramid): +`status: production` is not a label an author flips — it is a state a version *earns* by passing its declared `promotion_gate`. This makes the promotion lifecycle (`draft → validated → production → deprecated`) a machine-checked quality gate rather than a human's say-so, and it slots directly onto the existing [tiered validation pyramid](/sample-autonomous-cloud-coding-agents/decisions/adr-013-tiered-validation-pyramid): | Workflow status | Gate that must pass | Validation tier (ADR-013) | |---|---|---| @@ -501,10 +501,10 @@ Adapted from the issue's phases (the issue framed Phase 1 as a `task_type` *alia | Phase | Deliverable | Primary files | |---|---|---| -| 0 | This design doc + [ADR-014](/sample-autonomous-cloud-coding-agents/architecture/adr-014-workflow-driven-tasks) + JSON Schema + step-runner skeleton | `docs/design/WORKFLOWS.md`, `docs/decisions/`, `agent/workflows/schema/` | +| 0 | This design doc + [ADR-014](/sample-autonomous-cloud-coding-agents/decisions/adr-014-workflow-driven-tasks) + JSON Schema + step-runner skeleton | `docs/design/WORKFLOWS.md`, `docs/decisions/`, `agent/workflows/schema/` | | 1 | Step runner + `default/agent-v1` + migrate `new_task` to a workflow file; introduce `workflow_ref` and **remove the `task_type` enum** end-to-end (API/CLI/agent); the single workflow validator + `contracts/workflow-validation/` golden corpus | `agent/src/workflow/`, `agent/workflows/coding/new-task-v1.yaml`, `cdk/src/handlers/`, `cli/src/`, `contracts/workflow-validation/` | | 2b | Migrate `pr_iteration`, `pr_review` onto workflows behind a `read_only ⇒ "pr_review"` principal bridge (read-only stays enforced by the existing literal rules throughout) | `agent/workflows/coding/*`, `agent/tests/` | -| 2a | **Cedar property-keyed read-only migration** — literal `"pr_review"` hard-deny → `context.read_only == true` rules (`read_only_forbid_write/edit`), threaded via `context.read_only`; removes the 2b bridge; adds `read-only-*` `contracts/cedar-parity/` fixtures verified on *both* engines. (Originally planned as an isolated PR ahead of 2b; reordered after 2b shipped first behind the bridge — see [ADR-014](/sample-autonomous-cloud-coding-agents/architecture/adr-014-workflow-driven-tasks) addendum.) | `agent/policies/`, `cdk/src/handlers/shared/builtin-policies.ts`, `contracts/cedar-parity/`, `agent/src/policy.py`, `agent/src/workflow/loader.py` | +| 2a | **Cedar property-keyed read-only migration** — literal `"pr_review"` hard-deny → `context.read_only == true` rules (`read_only_forbid_write/edit`), threaded via `context.read_only`; removes the 2b bridge; adds `read-only-*` `contracts/cedar-parity/` fixtures verified on *both* engines. (Originally planned as an isolated PR ahead of 2b; reordered after 2b shipped first behind the bridge — see [ADR-014](/sample-autonomous-cloud-coding-agents/decisions/adr-014-workflow-driven-tasks) addendum.) | `agent/policies/`, `cdk/src/handlers/shared/builtin-policies.ts`, `contracts/cedar-parity/`, `agent/src/policy.py`, `agent/src/workflow/loader.py` | | 3 | Repo-optional `web_research` workflow (the repo-optional refactor — see [the requires_repo note](#domain--requires_repo)) | `cdk/src/handlers/`, `agent/workflows/knowledge/` | | 4 | Registry-native workflows (#246); Blueprint workflow allow-list + `default_workflow`; inline/repo-local for dev | depends on #246 | @@ -514,7 +514,7 @@ Per #248, the following remain out of scope (deferred to #99 / separate issues): ## Open questions -These are genuine forks; the repo-optional items (1–2) were **prerequisites for Phase 3** and have been **resolved as recorded decisions in the [ADR-014](/sample-autonomous-cloud-coding-agents/architecture/adr-014-workflow-driven-tasks) addendum (2026-06-08)**, with the one implied schema reshape applied — so the Phase-0 schema is now **frozen**. They are kept here (struck-through) for traceability. +These are genuine forks; the repo-optional items (1–2) were **prerequisites for Phase 3** and have been **resolved as recorded decisions in the [ADR-014](/sample-autonomous-cloud-coding-agents/decisions/adr-014-workflow-driven-tasks) addendum (2026-06-08)**, with the one implied schema reshape applied — so the Phase-0 schema is now **frozen**. They are kept here (struck-through) for traceability. 1. ~~**Memory actorId for repo-less tasks.**~~ **RESOLVED (ADR-014 addendum):** per-user `actorId = user:{cognito_sub}` (caller-scoped, no cross-tenant bleed; mirrors the per-user trace prefix). Cross-workflow knowledge pooling is explicitly not adopted. **No schema field added** (fixed platform fallback, not author-configurable) — a Phase-3 `memory.py` change keys on `user:{user_id}` when `repo` is absent. Coordinate with [MEMORY.md](/sample-autonomous-cloud-coding-agents/architecture/memory). 2. ~~**Artifact delivery contract.**~~ **RESOLVED (ADR-014 addendum):** `deliver_artifact.target` is an **open string naming a registered Python deliverer** (`workflow/deliverers.py` → `DELIVERERS`), not a closed enum — new delivery methods are registered deliverers, not schema changes. Shared plumbing is **pinned**: task-scoped key `artifacts/{task_id}/`, a prefix-scoped SessionRole IAM grant, a per-artifact size limit, and `TaskDetail` URL surfacing; the SessionRole `repo` tenant tag gains a `workflow:{id}` repo-less form. Each deliverer declares the outcomes it `produces`; validator rule 11 consults that registry. Implementations land in Phase 3; only the contract is frozen here. @@ -529,4 +529,4 @@ This design is a scoped-down reconciliation of the unmerged AKW port on `origin/ Two refinements layered on top of that port are worth calling out, because they shape the schema: 1. **Discovery is separate from execution.** A workflow carries optional `description` / `guidance` fields — a human- and agent-readable selection surface for registry search and workflow-selection (#246) — kept distinct from the machine-facing `prompt`. -2. **Promotion is earned, not set** — see [Promotion is earned, not set](#promotion-is-earned-not-set). `production` is gated by a declared `promotion_gate`, reusing the [ADR-013](/sample-autonomous-cloud-coding-agents/architecture/adr-013-tiered-validation-pyramid) validation pyramid rather than being a label an author flips. +2. **Promotion is earned, not set** — see [Promotion is earned, not set](#promotion-is-earned-not-set). `production` is gated by a declared `promotion_gate`, reusing the [ADR-013](/sample-autonomous-cloud-coding-agents/decisions/adr-013-tiered-validation-pyramid) validation pyramid rather than being a label an author flips. diff --git a/docs/src/content/docs/customizing/Cedar-policies.md b/docs/src/content/docs/customizing/Cedar-policies.md index 6fe17e11a..a04540221 100644 --- a/docs/src/content/docs/customizing/Cedar-policies.md +++ b/docs/src/content/docs/customizing/Cedar-policies.md @@ -6,7 +6,7 @@ title: Cedar policy guide This guide is for **blueprint authors** — repo owners writing the Cedar policies that govern what tool calls the agent can make unattended versus which ones pause for human approval. -> **If you are a task submitter** looking for how approvals work at the CLI, see [User guide — Approval gates](/sample-autonomous-cloud-coding-agents/using/overview#approval-gates-cedar-hitl). This guide is about *writing* the rules that cause approvals. +> **If you are a task submitter** looking for how approvals work at the CLI, see [User guide — Approval gates](/sample-autonomous-cloud-coding-agents/using/approval-gates-cedar-hitl). This guide is about *writing* the rules that cause approvals. > > **For the full design** (fail-closed posture, engine internals, concurrency), see [Cedar HITL gates design doc](/sample-autonomous-cloud-coding-agents/architecture/cedar-hitl-gates). @@ -47,7 +47,7 @@ security: approvalGateCap: 50 # optional per-task gate budget (1–500, default 50) ``` -The built-in rule set is documented in [`agent/policies/hard_deny.cedar`](../../agent/policies/hard_deny.cedar) and [`agent/policies/soft_deny.cedar`](../../agent/policies/soft_deny.cedar). Run `bgagent policies list --repo owner/repo` against a deployed stack to see the effective rules for a repo. +The built-in rule set is documented in [`agent/policies/hard_deny.cedar`](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/agent/policies/hard_deny.cedar) and [`agent/policies/soft_deny.cedar`](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/agent/policies/soft_deny.cedar). Run `bgagent policies list --repo owner/repo` against a deployed stack to see the effective rules for a repo. ## Vocabulary @@ -67,11 +67,14 @@ Every tool call the agent makes is evaluated as a Cedar `(principal, action, res |---|---|---|---| | `@rule_id("...")` | **Yes on soft-deny** (recommended on hard-deny) | Unique kebab/snake-case identifier | Stable ID for `--pre-approve rule:X`, for audit events, and for `bgagent policies show --rule X`. Engine rejects duplicates at task start. | | `@tier("hard"\|"soft")` | **Yes** | Exactly one of `"hard"` or `"soft"` | Must match the file section. Mismatches fail task start. | -| `@approval_timeout_s("N")` | No | Integer seconds ≥ 30 | Per-rule timeout. Defaults to 300 s (overridable per-task via `--approval-timeout`). When multiple soft rules match, the engine picks the minimum. Values < 120 s emit a load-time warning; values < 30 s are rejected. Ignored on hard-deny. | +| `@approval_timeout_s("N")` | No | Integer seconds ≥ 30 | Optional per-rule decision deadline. If absent, uses the task setting, whose default `0` means no deadline. The shortest positive task/rule deadline wins. Values < 120 s emit a load-time warning; values < 30 s are rejected. Ignored on hard-deny. | | `@severity("low"\|"medium"\|"high")` | No | One of three | Displayed in the approval prompt. Default: `medium`. | | `@category("...")` | No | `destructive`, `network`, `filesystem`, `auth`, or free-form | Optional UX grouping. Not enforced. | -**Rule of thumb:** every soft-deny rule must have `@rule_id` and should set `@severity` + `@approval_timeout_s` explicitly. Users scanning `bgagent pending` lean on these fields to triage quickly. +**Rule of thumb:** every soft-deny rule must have `@rule_id` and should set +`@severity`. Add `@approval_timeout_s` only when the workflow needs a decision +deadline. Omit the annotation to let the task choose; unlike the task's `0` +setting, a zero-valued rule annotation is invalid. ## Common patterns @@ -159,12 +162,12 @@ Fix the blueprint, redeploy (or update the `blueprint.yaml` if you're using a pu ## Testing policies before shipping -Every repo blueprint is covered by **cross-engine parity fixtures** in [`contracts/cedar-parity/`](../../contracts/cedar-parity/). Before shipping a non-trivial rule change, drop a golden-file fixture that pins the expected `(decision, matching_rule_ids)` for a representative `(policies, input)` pair. Both the Python `cedarpy` engine and the TypeScript `@cedar-policy/cedar-wasm` engine run it — divergence fails CI. See the directory's README for the fixture schema. +Every repo blueprint is covered by **cross-engine parity fixtures** in [`contracts/cedar-parity/`](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/tree/main/contracts/cedar-parity). Before shipping a non-trivial rule change, drop a golden-file fixture that pins the expected `(decision, matching_rule_ids)` for a representative `(policies, input)` pair. Both the Python `cedarpy` engine and the TypeScript `@cedar-policy/cedar-wasm` engine run it — divergence fails CI. See the directory's README for the fixture schema. -For unit coverage of your own rules without the cross-engine guarantee, add a case to [`agent/tests/test_policy.py`](../../agent/tests/test_policy.py) using `PolicyEngine.evaluate_tool_use(...)`. +For unit coverage of your own rules without the cross-engine guarantee, add a case to [`agent/tests/test_policy.py`](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/agent/tests/test_policy.py) using `PolicyEngine.evaluate_tool_use(...)`. ## Where to look next - [`docs/design/CEDAR_HITL_GATES.md`](/sample-autonomous-cloud-coding-agents/architecture/cedar-hitl-gates) — full design: engine internals, fail-closed posture, late-approval races, concurrency. -- [`agent/policies/hard_deny.cedar`](../../agent/policies/hard_deny.cedar) + [`agent/policies/soft_deny.cedar`](../../agent/policies/soft_deny.cedar) — the built-in rule set, good starting point for copy-paste. -- [User guide — Approval gates](/sample-autonomous-cloud-coding-agents/using/overview#approval-gates-cedar-hitl) — the CLI side (`bgagent pending` / `approve` / `deny` / `policies`). +- [`agent/policies/hard_deny.cedar`](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/agent/policies/hard_deny.cedar) + [`agent/policies/soft_deny.cedar`](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/agent/policies/soft_deny.cedar) — the built-in rule set, good starting point for copy-paste. +- [User guide — Approval gates](/sample-autonomous-cloud-coding-agents/using/approval-gates-cedar-hitl) — the CLI side (`bgagent pending` / `approve` / `deny` / `policies`). diff --git a/docs/src/content/docs/decisions/Adr-002-least-privilege-bootstrap-policies.md b/docs/src/content/docs/decisions/Adr-002-least-privilege-bootstrap-policies.md index befa6db4b..2e39f10d8 100644 --- a/docs/src/content/docs/decisions/Adr-002-least-privilege-bootstrap-policies.md +++ b/docs/src/content/docs/decisions/Adr-002-least-privilege-bootstrap-policies.md @@ -68,7 +68,7 @@ The implementation is decomposed into 8 sub-issues, each independently reviewabl ## References -- [ADR-001](/sample-autonomous-cloud-coding-agents/architecture/adr-001-stacked-pull-requests) — delivery methodology (stacked PRs) +- [ADR-001](/sample-autonomous-cloud-coding-agents/decisions/adr-001-stacked-pull-requests) — delivery methodology (stacked PRs) - RFC #120 — parent issue with full design and sub-issue breakdown - `docs/design/DEPLOYMENT_ROLES.md` — current documentation (will become generated) - PR #46 — original policy derivation and validation methodology diff --git a/docs/src/content/docs/decisions/Adr-004-tabula-rasa-documentation.md b/docs/src/content/docs/decisions/Adr-004-tabula-rasa-documentation.md index 697d4ef34..6f818da23 100644 --- a/docs/src/content/docs/decisions/Adr-004-tabula-rasa-documentation.md +++ b/docs/src/content/docs/decisions/Adr-004-tabula-rasa-documentation.md @@ -48,7 +48,7 @@ Never force a novice to read expert material to proceed. Never force an expert t ### Self-contained references When referencing another document: -- State what the reader gets from it: "See [Deployment Guide](link) for AWS account setup (required before this step)" +- State what the reader gets from it: "See [Deployment Guide](/sample-autonomous-cloud-coding-agents/getting-started/deployment-guide) for AWS account setup (required before this step)" - Never assume the reader has read it - Never use "as mentioned above" — each section must stand alone after context compaction diff --git a/docs/src/content/docs/decisions/Adr-012-operational-knowledge-stack.md b/docs/src/content/docs/decisions/Adr-012-operational-knowledge-stack.md index 5c3399b3b..d12a38a8c 100644 --- a/docs/src/content/docs/decisions/Adr-012-operational-knowledge-stack.md +++ b/docs/src/content/docs/decisions/Adr-012-operational-knowledge-stack.md @@ -155,7 +155,7 @@ Organized by persona: ```markdown # Contributor Workflow -> Operationalizes [ADR-003](/sample-autonomous-cloud-coding-agents/architecture/adr-003-contribution-governance) +> Operationalizes [ADR-003](/sample-autonomous-cloud-coding-agents/decisions/adr-003-contribution-governance) ## For Planners - Issue quality bar (what makes an issue "ready") diff --git a/docs/src/content/docs/decisions/Adr-014-workflow-driven-tasks.md b/docs/src/content/docs/decisions/Adr-014-workflow-driven-tasks.md index 099bbc0e6..bf9a3068f 100644 --- a/docs/src/content/docs/decisions/Adr-014-workflow-driven-tasks.md +++ b/docs/src/content/docs/decisions/Adr-014-workflow-driven-tasks.md @@ -106,5 +106,5 @@ With both resolved and the one schema reshape applied, the Phase-0 schema is **f - [docs/design/REPO_ONBOARDING.md](/sample-autonomous-cloud-coding-agents/architecture/repo-onboarding) — the Blueprint construct and `step_sequence` model - [docs/design/CEDAR_HITL_GATES.md](/sample-autonomous-cloud-coding-agents/architecture/cedar-hitl-gates) — policy engine the `agent_config` feeds - Prior art: `origin/merge/akw-integration` (commit `9d066a8`) — AKW YAML registry and models (reconciled, scoped down) -- [ADR-013](/sample-autonomous-cloud-coding-agents/architecture/adr-013-tiered-validation-pyramid) — the validation pyramid the `promotion_gate` layers onto -- [ADR-005](/sample-autonomous-cloud-coding-agents/architecture/adr-005-feedback-loop) — the feedback loop that workflow trajectory-evolution would extend (future, out of scope) +- [ADR-013](/sample-autonomous-cloud-coding-agents/decisions/adr-013-tiered-validation-pyramid) — the validation pyramid the `promotion_gate` layers onto +- [ADR-005](/sample-autonomous-cloud-coding-agents/decisions/adr-005-feedback-loop) — the feedback loop that workflow trajectory-evolution would extend (future, out of scope) diff --git a/docs/src/content/docs/decisions/Adr-016-pluggable-identity-and-auth.md b/docs/src/content/docs/decisions/Adr-016-pluggable-identity-and-auth.md index d30b13960..ee2d596e9 100644 --- a/docs/src/content/docs/decisions/Adr-016-pluggable-identity-and-auth.md +++ b/docs/src/content/docs/decisions/Adr-016-pluggable-identity-and-auth.md @@ -94,7 +94,7 @@ The abstraction is intentionally a contract, not a forklift of credential handli - **Backend-agnostic.** One `resolve__token()` contract serves both the AgentCore Runtime backend (token arrives via the `WorkloadAccessToken` header) and the parked ECS backend (in-process boto3). The backend decides how the token arrives; the seam does not care. - **Incremental.** The rollout is phased and flag-gated, and the shared PAT fallback stays until the vault path is green. No big-bang cutover. -- **Consistent with ADR-014.** This is the credential-plane analog of [ADR-014](/sample-autonomous-cloud-coding-agents/architecture/adr-014-workflow-driven-tasks)'s provider-neutral `VcsProvider` seam: that one named GitHub-specific control-plane operations as instances of generic concepts; this one names the per-integration credential resolvers as instances of one outbound contract. +- **Consistent with ADR-014.** This is the credential-plane analog of [ADR-014](/sample-autonomous-cloud-coding-agents/decisions/adr-014-workflow-driven-tasks)'s provider-neutral `VcsProvider` seam: that one named GitHub-specific control-plane operations as instances of generic concepts; this one names the per-integration credential resolvers as instances of one outbound contract. ## Credential types @@ -164,7 +164,7 @@ Per-user `McpCredential` selection requires the Gateway to know *which task-user | P6 | **Trusted task-user identity propagation** (prerequisite for per-user MCP on the general plane, P5): specify + validate a user-scoped inbound identity the Gateway authorizer trusts, replacing the M2M JWT for per-user credential selection. | **Blocks per-user `McpCredential`.** Until done, MCP credentials are workspace-scoped at best. | | P7 | Jira + Slack `ChannelCredential` (same shape as P1); GitHub `GithubOauth2` behind a flag, retire the shared PAT; OBO `act`-claim delegation feeding #237. | Flag-gated; per-surface. | -**Substrate exception — `lambda-microvm` (2026-09-02):** P1's vault cannot be enabled on the MicroVM substrate. The two together synthesize 505 resources against CloudFormation's hard 500-resource limit (MicroVM alone 496, the vault alone 488), so `AgentStack` refuses the combination at synth, naming both context flags. The MicroVM wiring itself is complete — `platform_config` carries the workload name and the guest execution role holds the mint grant — so this is a capacity limit, not a design gap, and it lifts as soon as a subsystem moves into a nested stack. +**Lambda MicroVMs:** The vault uses the guest's compute execution role, with its workload name delivered in authenticated `platform_config`. The former resource-count refusal was removed under #857 after stack reductions and nesting; backend/image-mode tests now include vault and gateway combinations, and check each template's resource and size budgets. See the [Linear setup guide](/sample-autonomous-cloud-coding-agents/using/linear-setup-guide#using-the-vault-with-lambda-microvms). **Substrate independence (verified 2026-07-21, both proven live):** the vault path works on any compute. AgentCore Runtime injects the Workload Access Token as the `WorkloadAccessToken` header; ECS/Fargate/Lambda bootstrap it via `GetWorkloadAccessTokenForJWT(workloadName, userToken=)` against a **standalone** (non-service-linked) workload identity, then call `GetResourceOauth2Token`. Runtime-managed (service-linked) workload identities cannot self-vend, so the ECS path needs a manually-created workload identity. The runtime execution role today has `GetWorkloadAccessToken*` but **not** `GetResourceOauth2Token` — P1 adds it, plus `GetSecretValue` scoped to that surface's providers (`bedrock-agentcore-identity!default/oauth2/*`; see the implementation notes below for why this is narrower than the wildcard first anticipated here). @@ -246,7 +246,7 @@ revoking a grant — it changes who stores and refreshes the token, not whether - Issue [#215](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/215) — Bedrock billing attribution - Issue [#237](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/237) — governance planes / `abca.audit.v1` correlation block - Issue [#288](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/288) / PR [#302](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/302) — Jira integration (second provider through the resolver seam) -- [ADR-014](/sample-autonomous-cloud-coding-agents/architecture/adr-014-workflow-driven-tasks) — workflow-driven tasks; introduced the provider-neutral `VcsProvider` seam this ADR is the credential-plane analog of +- [ADR-014](/sample-autonomous-cloud-coding-agents/decisions/adr-014-workflow-driven-tasks) — workflow-driven tasks; introduced the provider-neutral `VcsProvider` seam this ADR is the credential-plane analog of - [IDENTITY_AND_AUTH.md](/sample-autonomous-cloud-coding-agents/architecture/identity-and-auth) — the worked use-cases, seams table, decision tree, and Linear before/after - [SECURITY.md](/sample-autonomous-cloud-coding-agents/architecture/security) — current auth posture, the shared-PAT limitation this ADR resolves - GitHub issues — per-repo GitHub credentials, layered credential derivation, delegation chain propagation (priority labels `P0`, `P1`, etc.) diff --git a/docs/src/content/docs/decisions/Adr-019-agentcore-gateway-tool-federation.md b/docs/src/content/docs/decisions/Adr-019-agentcore-gateway-tool-federation.md index a9e061b15..34f70f0af 100644 --- a/docs/src/content/docs/decisions/Adr-019-agentcore-gateway-tool-federation.md +++ b/docs/src/content/docs/decisions/Adr-019-agentcore-gateway-tool-federation.md @@ -9,9 +9,9 @@ title: Adr 019 agentcore gateway tool federation ## Context -ABCA gives the agent tools by writing a per-thread `.mcp.json` into the cloned repo before each run. `agent/src/channel_mcp.py` maps an inbound channel to a hosted MCP server entry. **Today the agent holds zero functional platform-managed MCP servers:** the single `CHANNEL_MCP_BUILDERS` entry is `jira`, and it is a **non-functional placeholder** — the headless agent cannot complete Atlassian's interactive OAuth 2.1 flow, so the live outbound path is the REST shim in `jira_reactions.py` (see [ADR-015](/sample-autonomous-cloud-coding-agents/architecture/adr-015-jira-integration)). Linear is **deterministic by decision** ([ADR-016](/sample-autonomous-cloud-coding-agents/architecture/adr-016-pluggable-identity-and-auth)): there is no Linear MCP entry — it was removed after it proved non-functional, and `strip_linear_mcp_servers()` scrubs any Linear MCP a repo commits to `.mcp.json` before the SDK loads it. When a real MCP tool *is* wired, its entry carries a credential the container holds for the whole task (e.g. a `Bearer ${...}` header), resolved by a per-integration `resolve__token()` in `config.py`. This pattern — for any tool ABCA does adopt — has four structural costs: +ABCA gives the agent tools by writing a per-thread `.mcp.json` into the cloned repo before each run. `agent/src/channel_mcp.py` maps an inbound channel to a hosted MCP server entry. **Today the agent holds zero functional platform-managed MCP servers:** the single `CHANNEL_MCP_BUILDERS` entry is `jira`, and it is a **non-functional placeholder** — the headless agent cannot complete Atlassian's interactive OAuth 2.1 flow, so the live outbound path is the REST shim in `jira_reactions.py` (see [ADR-015](/sample-autonomous-cloud-coding-agents/decisions/adr-015-jira-integration)). Linear is **deterministic by decision** ([ADR-016](/sample-autonomous-cloud-coding-agents/decisions/adr-016-pluggable-identity-and-auth)): there is no Linear MCP entry — it was removed after it proved non-functional, and `strip_linear_mcp_servers()` scrubs any Linear MCP a repo commits to `.mcp.json` before the SDK loads it. When a real MCP tool *is* wired, its entry carries a credential the container holds for the whole task (e.g. a `Bearer ${...}` header), resolved by a per-integration `resolve__token()` in `config.py`. This pattern — for any tool ABCA does adopt — has four structural costs: -1. **The tool credential lives in the container.** Every MCP entry injects a bearer token into the agent's environment. The token is in the blast radius of any prompt-injection or dependency compromise for the whole task. [ADR-016](/sample-autonomous-cloud-coding-agents/architecture/adr-016-pluggable-identity-and-auth) is unifying the *resolution* of these tokens, but not the fact that the resolved token still lands in the container for MCP. +1. **The tool credential lives in the container.** Every MCP entry injects a bearer token into the agent's environment. The token is in the blast radius of any prompt-injection or dependency compromise for the whole task. [ADR-016](/sample-autonomous-cloud-coding-agents/decisions/adr-016-pluggable-identity-and-auth) is unifying the *resolution* of these tokens, but not the fact that the resolved token still lands in the container for MCP. 2. **Tool wiring is bespoke per server.** Adding a tool means editing `channel_mcp.py`, adding a `CHANNEL_MCP_BUILDERS` entry, threading a credential resolver, and redeploying. There is no declarative "add a tool" path for an operator. @@ -21,13 +21,13 @@ ABCA gives the agent tools by writing a per-thread `.mcp.json` into the cloned r **Amazon Bedrock AgentCore Gateway is a fully managed service that addresses all four.** Gateway converts APIs, Lambda functions, Smithy models, OpenAPI specs, and remote MCP servers into MCP-compatible tools; aggregates multiple such **targets** behind **one virtual MCP server** (a single consolidated `tools/list`); and manages **both** inbound authentication (agent → gateway) and outbound authentication (gateway → target) as a managed concern. It is serverless and observable, supports **MCP session reuse** and **semantic tool search**, and — critically for ABCA — its inbound authorizer can be **AWS IAM (SigV4)** or a **CUSTOM_JWT** token, both of which are portable across compute substrates. -**This is the tool-plane complement to [ADR-016](/sample-autonomous-cloud-coding-agents/architecture/adr-016-pluggable-identity-and-auth).** ADR-016 unifies *who is the principal* (inbound) and *how do we resolve an outbound credential* (the `resolve__token()` seam) with AgentCore Identity as one backend. This ADR decides *where the agent's tools live and how it reaches them*: behind a single managed Gateway endpoint whose outbound leg is built **on** AgentCore Identity's credential providers — so the same vault ADR-016 adopts holds the tool credential, and it is injected gateway → target and **never enters the container**. +**This is the tool-plane complement to [ADR-016](/sample-autonomous-cloud-coding-agents/decisions/adr-016-pluggable-identity-and-auth).** ADR-016 unifies *who is the principal* (inbound) and *how do we resolve an outbound credential* (the `resolve__token()` seam) with AgentCore Identity as one backend. This ADR decides *where the agent's tools live and how it reaches them*: behind a single managed Gateway endpoint whose outbound leg is built **on** AgentCore Identity's credential providers — so the same vault ADR-016 adopts holds the tool credential, and it is injected gateway → target and **never enters the container**. **A spike has already de-risked the mechanism.** The reference branch `feat/agentcore-gateway-mcp` (on the upstream remote — not merged to `main`; its `docs/design/AGENTCORE_GATEWAY_MCP_SPIKE.md` records findings F0–F16) took Gateway end-to-end against a hosted MCP server and reached a working federated endpoint — **verdict GO**, live-proven and torn down. That spike is Linear-coupled and, per [#641](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/641), **a reference for the mechanism, not a base** — this ADR does not continue off it. It reuses the proven *mechanism* — per-workspace provisioning shape, M2M inbound token minting, the CDK role/grant pattern, the registry → `.mcp.json` plumbing — while structuring it around a **provider-agnostic** target + config model rather than Linear-specific code, so tools onboard without new platform code. The originating issue is [#641](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/641). **Near-term scope: lead with the simplest, most common target type — not the hardest.** [#641](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/641) is explicit that OAuth is *one option among several* and that most targets ABCA would add never touch it; the prior spike's mistake was exercising only a single 3LO-OAuth MCP-server target. So P1 onboards a **Lambda tool target** — outbound auth is the gateway execution role (IAM), **no stored credential and no consent flow at all** — which proves the end-to-end path (provisioning, substrate-portable inbound, aggregation, agent routing) on the *easiest* outbound leg. Simpler-and-common target types (IAM-signed HTTP/OpenAPI, API-key) follow; the demanding 3LO-OAuth remote-MCP path is exercised **last**, once the general model is proven, not first. -Linear and Jira MCP are **explicitly not in near-term scope.** Linear is deterministic by decision ([ADR-016](/sample-autonomous-cloud-coding-agents/architecture/adr-016-pluggable-identity-and-auth)) and stays that way — this ADR does not re-introduce an agent-side Linear MCP. Jira's live path is the REST shim ([ADR-015](/sample-autonomous-cloud-coding-agents/architecture/adr-015-jira-integration)); whether a gateway that owns the OAuth flow could *unbreak* the non-functional Jira MCP placeholder is recorded below as a **speculative later experiment**, explicitly not part of the deterministic reaction/orchestration paths. +Linear and Jira MCP are **explicitly not in near-term scope.** Linear is deterministic by decision ([ADR-016](/sample-autonomous-cloud-coding-agents/decisions/adr-016-pluggable-identity-and-auth)) and stays that way — this ADR does not re-introduce an agent-side Linear MCP. Jira's live path is the REST shim ([ADR-015](/sample-autonomous-cloud-coding-agents/decisions/adr-015-jira-integration)); whether a gateway that owns the OAuth flow could *unbreak* the non-functional Jira MCP placeholder is recorded below as a **speculative later experiment**, explicitly not part of the deterministic reaction/orchestration paths. > **Not to be confused with the input gateway.** [`INPUT_GATEWAY.md`](/sample-autonomous-cloud-coding-agents/architecture/input-gateway) describes ABCA's *inbound channel normalization* — turning CLI / Slack / webhook payloads into one internal message format on the way **in**. AgentCore Gateway here is the opposite direction: aggregating *outbound tool calls* the agent makes **out**. They share the word "gateway" and nothing else. @@ -117,21 +117,21 @@ Implementation lands as **multiple PRs** after this ADR: | P1 | **Lambda tool behind the gateway.** Context-gated CDK gateway (L2 `Gateway`) + service role, AWS_IAM (SigV4) inbound, one read-only Lambda target (`abca_repo_config`); outbound = gateway execution role, **no credential**. **Agent routing** is an **in-process SigV4 MCP bridge** (`agent/src/gateway_tools.py`), *not* a `.mcp.json` entry: the SDK's remote MCP client attaches only a **static** `Authorization` header, so it cannot sign SigV4 — the bridge signs each request in-process with the compute role's credentials (`service = bedrock-agentcore`) and proxies the model's call. The whole feature (CDK + agent bridge) is gated on `ABCA_TOOL_GATEWAY_URL` (set only under `--context enableToolGateway=true`), so default synth is byte-unchanged and the bridge is inert when the flag is off. | Green on microVM; default synth byte-unchanged; feature inert when gate off. | | P2 | **Substrate parity + a credentialed target.** IAM-SigV4 / JWT inbound proven on **both** microVM and ECS/Fargate; add one credentialed simple target (IAM-signed HTTP/OpenAPI or API-key vaulted via AgentCore Identity). | Both substrates pass; credential never enters the container. | | P3 | **Generalize to N targets.** Provider-agnostic target + config model keyed off the registry ([#246](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/246)); `bgagent gateway add-target/list-targets/sync/remove-target`; remaining simpler target types (Smithy, no-auth, OAuth 2LO). | Onboarding branches green; idempotent re-run. | -| P4 | **Hardening, search, and the hard auth path.** Semantic tool search evaluation; lifecycle + cleanup + `*_PENDING_AUTH` recovery; aggregation of ≥2 target types behind one gateway; the 3LO-OAuth remote-MCP path (the spike's territory). *Speculative experiment, separately gated:* whether a gateway that owns the OAuth flow can make the non-functional Jira MCP placeholder reachable — explicitly **not** re-introducing agent-side Linear/Jira MCP into the deterministic reaction/orchestration paths ([ADR-015](/sample-autonomous-cloud-coding-agents/architecture/adr-015-jira-integration), [ADR-016](/sample-autonomous-cloud-coding-agents/architecture/adr-016-pluggable-identity-and-auth)). | As needed by tool count / ops; Jira-MCP experiment records reachable-or-not. | +| P4 | **Hardening, search, and the hard auth path.** Semantic tool search evaluation; lifecycle + cleanup + `*_PENDING_AUTH` recovery; aggregation of ≥2 target types behind one gateway; the 3LO-OAuth remote-MCP path (the spike's territory). *Speculative experiment, separately gated:* whether a gateway that owns the OAuth flow can make the non-functional Jira MCP placeholder reachable — explicitly **not** re-introducing agent-side Linear/Jira MCP into the deterministic reaction/orchestration paths ([ADR-015](/sample-autonomous-cloud-coding-agents/decisions/adr-015-jira-integration), [ADR-016](/sample-autonomous-cloud-coding-agents/decisions/adr-016-pluggable-identity-and-auth)). | As needed by tool count / ops; Jira-MCP experiment records reachable-or-not. | ## Out of scope (this ADR) - **The provisioning, CLI, and agent-routing code.** This ADR records the decision; the code is the P1–P4 follow-up PRs on [#641](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/641). - **The registry schema and storage.** Owned by [#246](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/246) / PR [#548](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/548); this ADR consumes its MCP-server asset kind, it does not define a second registry. -- **Inbound principal verification and outbound credential resolution.** Owned by [ADR-016](/sample-autonomous-cloud-coding-agents/architecture/adr-016-pluggable-identity-and-auth); this ADR is the tool-plane consumer of both seams. -- **The Linear/Jira deterministic reaction paths.** `linear_reactions.py` / `jira_reactions.py` stay as-is and authoritative. This ADR does **not** re-introduce an agent-side Linear or Jira MCP: Linear stays deterministic ([ADR-016](/sample-autonomous-cloud-coding-agents/architecture/adr-016-pluggable-identity-and-auth), enforced by `strip_linear_mcp_servers()`), and Jira's live path stays the REST shim ([ADR-015](/sample-autonomous-cloud-coding-agents/architecture/adr-015-jira-integration)). The speculative Jira-MCP-via-gateway experiment (P4) is gated separately and touches neither path. +- **Inbound principal verification and outbound credential resolution.** Owned by [ADR-016](/sample-autonomous-cloud-coding-agents/decisions/adr-016-pluggable-identity-and-auth); this ADR is the tool-plane consumer of both seams. +- **The Linear/Jira deterministic reaction paths.** `linear_reactions.py` / `jira_reactions.py` stay as-is and authoritative. This ADR does **not** re-introduce an agent-side Linear or Jira MCP: Linear stays deterministic ([ADR-016](/sample-autonomous-cloud-coding-agents/decisions/adr-016-pluggable-identity-and-auth), enforced by `strip_linear_mcp_servers()`), and Jira's live path stays the REST shim ([ADR-015](/sample-autonomous-cloud-coding-agents/decisions/adr-015-jira-integration)). The speculative Jira-MCP-via-gateway experiment (P4) is gated separately and touches neither path. ## References - Issue [#641](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/641) — the feature this ADR records (AgentCore Gateway tool federation) - Issue [#246](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/246) / PR [#548](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/548) — central agent asset registry (the MCP-server asset kind this ADR's targets are an instance of) -- [ADR-016](/sample-autonomous-cloud-coding-agents/architecture/adr-016-pluggable-identity-and-auth) — pluggable identity and auth (the credential-plane seams this ADR consumes) -- [ADR-015](/sample-autonomous-cloud-coding-agents/architecture/adr-015-jira-integration) — Jira integration (the REST-shim precedent for a non-functional MCP placeholder) +- [ADR-016](/sample-autonomous-cloud-coding-agents/decisions/adr-016-pluggable-identity-and-auth) — pluggable identity and auth (the credential-plane seams this ADR consumes) +- [ADR-015](/sample-autonomous-cloud-coding-agents/decisions/adr-015-jira-integration) — Jira integration (the REST-shim precedent for a non-functional MCP placeholder) - `docs/design/AGENTCORE_GATEWAY_MCP_SPIKE.md` — the GO spike, findings F0–F16, and the provisioning recipe + gotchas. Lives only on the reference branch `feat/agentcore-gateway-mcp` (upstream remote); not merged to `main`. - [`INPUT_GATEWAY.md`](/sample-autonomous-cloud-coding-agents/architecture/input-gateway) — the *inbound channel* gateway, distinct from AgentCore Gateway - AgentCore Gateway — [AWS Bedrock AgentCore developer guide](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/gateway.html) and the [gateway target / outbound-auth references](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/gateway-target-MCPservers.html) for target types, the inbound/outbound auth matrix, and session configuration diff --git a/docs/src/content/docs/decisions/Adr-021-lambda-microvms-compute-backend.md b/docs/src/content/docs/decisions/Adr-021-lambda-microvms-compute-backend.md index fec83f5a4..a928ad2cf 100644 --- a/docs/src/content/docs/decisions/Adr-021-lambda-microvms-compute-backend.md +++ b/docs/src/content/docs/decisions/Adr-021-lambda-microvms-compute-backend.md @@ -4,433 +4,221 @@ title: Adr 021 lambda microvms compute backend # ADR-021: AWS Lambda MicroVMs as a third ComputeStrategy backend -> Number: candidate ADR-021 (ADR-020 is the highest accepted on `main`; ADR-018 is claimed by open PR [#548](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/548), ADR-019 by open PR [#663](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/663)). Numbers are never reused. If a lower number frees before merge, renumber and coordinate with those PRs. +> **Implementation status (2026-09-18): P1 and P2 are merged; P3 is in draft review.** Approval sleep/wake, retained requests, conversation/workspace recovery and nested infrastructure have live acceptance evidence. Reusable migration and final integration checks remain open; see [verification status](/sample-autonomous-cloud-coding-agents/verification/readme). This ADR defines P1–P3, not an official P4. **Status:** proposed **Date:** 2026-07-29 ## Context -ABCA selects a per-repo compute backend through the Blueprint's `compute_type` field (`cdk/src/handlers/shared/repo-config.ts`). Two backends exist today, resolved by `resolveComputeStrategy` (`cdk/src/handlers/shared/compute-strategy.ts`) behind a uniform `ComputeStrategy` interface (`startSession` / `pollSession` / `stopSession`): +ABCA selects compute per repository through Blueprint `compute_type`. AgentCore Runtime is the default; ECS Fargate supports larger workloads. [#645](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/645) adds Lambda MicroVMs for explicit lifecycle control and reduced compute usage during human approval waits. -- **AgentCore Runtime** (`agentcore`, default) — managed Firecracker MicroVM per session. Invoked via `InvokeAgentRuntime`; liveness is inferred from agent heartbeats in DynamoDB plus the FastAPI `/ping` endpoint (`agent/src/server.py`) — the strategy's `pollSession` is a stub that always reports `running`. Constraints: 2 GB image limit, no substrate-level suspend API exposed to the orchestrator. -- **ECS on Fargate** (`ecs`) — always-on Fargate task (16 vCPU / 120 GB, ARM64) for repos that exceed AgentCore's limits ([#596](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/596)). Invoked via `RunTask` in batch mode (bypassing the HTTP server); liveness via `DescribeTasks`. No suspend — an idle task blocked on an approval wait burns full compute the whole time. - -**AWS Lambda MicroVMs** (launched 2026-06-22) is a serverless Firecracker sandbox primitive that AWS positions explicitly for AI coding agents: VM-level isolation, snapshot-based near-instant launch, **suspend/resume with full memory + disk state preserved** (compute charges stop while suspended), a dedicated JWE-authenticated HTTPS endpoint per instance, lifecycle hooks (`/run`, `/suspend`, `/resume`, `/terminate`), and up to 8 hours per session. It is **not** classic Lambda: the 15-minute function cap does not apply, and [COMPUTE.md](/sample-autonomous-cloud-coding-agents/architecture/compute)'s "Lambda: poor fit" verdict refers to functions, not MicroVMs. - -[#645](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/645) proposes adding Lambda MicroVMs as a third `ComputeStrategy`. The fit is strong but not free — pre-implementation review of the service's lifecycle model surfaced real design tensions this ADR resolves: +Lambda MicroVMs are managed Firecracker virtual machines, separate from ordinary Lambda functions. A MicroVM can preserve memory and disk while suspended and can live for eight hours, including suspended time. The function service’s fifteen-minute limit does not apply. ### Capability comparison (delta rows only — full matrix in COMPUTE.md) -| | AgentCore Runtime | ECS on Fargate | **Lambda MicroVMs** | +| Capability | AgentCore Runtime | ECS Fargate | Lambda MicroVMs | |---|---|---|---| -| Isolation | MicroVM (managed) | Task-level (Firecracker) | MicroVM (Firecracker) | -| Max duration | 8 h | No cap | 8 h (running + suspended) — **verified** (`L-B430C318 = 8 hours`) | -| Suspend/resume | No orchestrator-visible API | No | **Yes** — explicit API + idle policy, state preserved, no compute charge while suspended. **Verified**: suspend reaches `SUSPENDED` in ~1 s with no `idlePolicy`; resume restores `RUNNING` in ~1 s with `microvmId` **and** `endpoint` byte-identical | -| Resources | AgentCore-managed | 16 vCPU / 120 GB / 20–200 GB disk | **Baseline 8 GiB RAM / 4 vCPU, auto-scaling to a 32 GiB / 16 vCPU peak**; 32 GB disk. `minimumMemoryInMiB` configures the BASELINE (max 8,192 MiB); the service scales vertically on demand — capacity is baseline-priced with 4× burst headroom | -| Packaging | ECR image ≤ 2 GB | ECR image, no hard cap | **Zip + Dockerfile in S3 → service-built snapshot image** (versioned, storage billed) | -| Invocation | `InvokeAgentRuntime` (SigV4) | `RunTask` + container overrides | `RunMicrovm` (image **ARN** required — a bare name is rejected) → dedicated HTTPS endpoint + JWE token (`CreateMicrovmAuthToken`, ≤ 60 min TTL) | -| Liveness | Agent heartbeat + `/ping` | `DescribeTasks` | MicroVM state (RUNNING / SUSPENDED / TERMINATED) via control-plane API **and** agent heartbeat (see sub-decision 1) | -| Session storage | `/mnt/workspace` FUSE (no `flock()`) | Ephemeral disk | Native disk in snapshot — **survives suspend/resume, `flock()` works** | -| Architecture | ARM64 | ARM64 | ARM64 (Graviton) | -| Regions (launch) | Broad | Broad | 5 (us-east-1/2, us-west-2, eu-west-1, ap-northeast-1) | - -> Rows marked **verified** were discharged empirically on 2026-07-31 (us-east-1); see `docs/verification/645-p1-lambda-microvm-runbook.md`. Two originally documented constants were **refuted** by that run and are corrected throughout this ADR: the `runHookPayload` cap (16 KB → 4 KB) and the claim that `RunMicrovm` accepts a bare image name (it does not). -> -> **Source hierarchy for service facts.** Where sources disagree, the higher tier wins, and the tier is recorded next to the claim: -> -> 1. **Live boundary probe** — a request the service actually accepted or rejected in this account/Region. Strongest, and the only tier that can refute the others. (Both corrections above came from here.) -> 2. **Modeled constraints** — the CLI/SDK request schema and enumerated allowed values (`--generate-cli-skeleton`, `ARM_64`, `ENABLED|DISABLED`). Authoritative about request *shape*; silent about runtime behaviour. -> 3. **Service developer guide** — including its sizing/scaling tables. Authoritative about *semantics* a probe cannot see, which is exactly how the memory row was fixed: a probe can only show that 32,768 MiB is rejected as a `minimumMemoryInMiB` value; only the guide explains that the field is a BASELINE and that the service scales to a 32 GiB peak on its own. A boundary probe is the strongest evidence about a boundary and says nothing about what the boundary *means*. -> 4. **SDK docstrings** — generated, and demonstrably stale here: `runHookPayload` documents "Maximum: 16,384 bytes" against an enforced 4,096. -> 5. **Launch blogs / skills / toolkit material** — orientation only; never load-bearing on its own. -> -> **Omitted API fields mean "service default", never "none".** Two live findings drove this rule: leaving `ingressNetworkConnectors` unset attaches a PUBLIC `HTTP_INGRESS` connector, and leaving `/ready` out of `hooks` makes the image un-creatable once any lifecycle hook is enabled. So for any security-relevant field, the desired posture must be **requested explicitly** and the test must assert the **outcome** (the ARN present in the request, the hook enabled) rather than the omission (`expect(field).toBeUndefined()`) — an omission assertion passes just as happily when the service is silently choosing something wider. -> -> On the **memory** row specifically: `CreateMicrovmImage` enumerates the baseline sizes a base image supports — `[512, 1024, 2048, 4096, 8192]` MiB for `al2023-1` — and rejects anything else, which is why the construct validates against that list at synth. The 32 GiB / 16 vCPU peak is reached by the service's own vertical scaling, not by asking for it. The account memory quota (`L-CD1C0CC4`, 1024 GB, "burst up to 4×") is an aggregate across MicroVMs, not a per-VM limit; note that concurrency arithmetic should be done against the PEAK, not the baseline, since that is what a busy fleet can actually consume. +| Packaging | ECR image, 2 GB limit | ECR image | ZIP + Dockerfile in S3 → versioned snapshot | +| Duration | Eight hours | No task duration cap | Eight hours, running + suspended | +| Explicit suspend/resume in ABCA | Unsupported | Unsupported | Control-plane APIs; memory and disk retained | +| Storage | Ephemeral disk plus preview persistent FUSE mount | Configurable ephemeral disk | 32 GB native disk; supports `flock()` | +| Sizing | Service-managed | Configurable; larger sustained workloads | ABCA baseline 8,192 MiB; service guide lists up to 32 GiB / 16 vCPU | +| Invocation/liveness | Invoke API + agent heartbeat | RunTask/DescribeTasks | RunMicrovm/GetMicrovm + agent heartbeat | + +See [COMPUTE.md](/sample-autonomous-cloud-coding-agents/architecture/compute) for the full comparison and costs. Suspending stops compute charges, but snapshot storage and save/restore charges remain. A shorter sleep delay does not guarantee lower total cost. + +Recorded P1 probes established a 4,096-byte `runHookPayload` limit, an image-ARN requirement and accepted baseline values of 512, 1,024, 2,048, 4,096 and 8,192 MiB. Those observations override conflicting generated SDK descriptions for the tested account/Region. They do not measure guest-visible launch memory, vertical-scaling latency or sustained workload fit. The service guide’s capacity figures and live observations must remain distinguishable. ### Design tensions the strategy must resolve -1. **Idle detection is inbound-traffic-based; the ABCA agent is outbound-only.** MicroVM idle policies suspend when no traffic arrives at the *endpoint*. A busy agent running a 40-minute build receives no inbound traffic and would be suspended mid-work by a naive idle policy. Conversely, "no inbound traffic" is the agent's *normal* state. -2. **No self-suspend.** The agent cannot suspend its own MicroVM from inside; only an external `SuspendMicrovm` call can. Suspend decisions must be owned by the orchestrator — which aligns with the unified liveness model proposed in [#491](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/491). -3. **Snapshots bake state.** The image snapshot is captured once at build time; every MicroVM resumes from it. Secrets, tokens, and per-task identity must arrive at run time (`runHookPayload`, ≤ 4 KB, or fetched in the `/run` hook), never at image build. CSPRNG reseeding is a consequence of the same property and is scoped to **P3** — see the amended note under sub-decision 2 and the risk bullet, which record why the exposure is negligible today. -4. **Auth tokens are short-lived.** JWE tokens max out at 60 minutes; any orchestrator→agent HTTP interaction over the endpoint needs token refresh, unlike AgentCore's SigV4 invoke or ECS's no-endpoint model. -5. **Identity delta — narrower than it looks.** Most AgentCore services ABCA uses are standalone and substrate-portable: Memory is already consumed from ECS via an IAM grant plus `MEMORY_ID` (`EcsAgentCluster`), and Gateway (ADR-019) is portable by design (SigV4 inbound). The genuinely Runtime-coupled piece is the workload-access-token **delivery mechanism** (`runtimeUserId` → `WorkloadAccessToken` request header → `BedrockAgentCoreContext`, used by `resolve_linear_api_token()`), which has no MicroVM equivalent. The ECS backend already lives with this delta (env-var token delivery); MicroVMs inherit the same posture until the pluggable identity work ([#249](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/249), ADR-016) redesigns the seam. -6. **The service's defaults are not our posture.** Two of them, both discovered live: `RunMicrovm` attaches a **public** `HTTP_INGRESS` connector (and mints a public `*.lambda-microvm..on.aws` endpoint) when `ingressNetworkConnectors` is omitted, and `CreateMicrovmImage` **requires** the `/ready` hook whenever any lifecycle hook is enabled. Neither posture can be reached by leaving a field out — each needs an explicit control (see sub-decisions 3 and 4). +1. Traffic-based idle policies observe inbound endpoint traffic. A busy coding agent mostly sends outbound requests, so absence of inbound traffic cannot establish that it is idle. +2. Suspension needs an external controller. The agent can prepare a safe checkpoint, but the coordinator owns the service call. +3. Image snapshots share build-time process state. Credentials, task identity and deployment configuration must arrive after launch. +4. A worker’s eight-hour lifetime is shorter than an unanswered approval may remain useful. Durable task state must outlive the worker. +5. New packaging, IAM and lifecycle behavior need live checks; a successful CDK synth cannot validate service semantics. ## Decision -Adopt **AWS Lambda MicroVMs as a third, opt-in `ComputeStrategy` backend** named `lambda-microvm`, selected per repo via Blueprint `compute_type`. AgentCore remains the default. Five sub-decisions: +Add `lambda-microvm` as an opt-in `ComputeStrategy`. AgentCore remains the default. ### 1. Strategy shape: extend the interface with mandatory suspend/resume -`ComputeType` widens to `'agentcore' | 'ecs' | 'lambda-microvm'` (mirrored in `cli/src/types.ts` and the CLI's inline unions). `SessionHandle` gains a `{ strategyType: 'lambda-microvm', microvmId, endpoint }` variant — `microvmId` because every lifecycle API (`suspend-microvm`, `resume-microvm`, `terminate-microvm`, `get-microvm`, and `create-microvm-auth-token` — the latter not called in P1–P3, see sub-decision 3) takes only the MicroVM identifier, and `endpoint` because it is per-session (minted by `RunMicrovm`) and required for any orchestrator→agent HTTP interaction. Note the naming seam: the handle field is `microvmId` (matching `RunMicrovmResponse.microvmId`), while the request key on every lifecycle command is `microvmIdentifier` — the strategy is the only place that translates between the two. The image ARN is deliberately **not** in the handle: like the ECS task definition ARN, it is deployment-time configuration consumed by `startSession` (from construct-injected environment) and recorded in the session-start log entry for diagnostics, not per-session lifecycle state. - -**A second, sharper naming seam: `imageIdentifier` must be an ARN.** The name suggests a bare image name is acceptable — `create-microvm-image --name` takes one, and this ADR originally assumed `run-microvm --image-identifier` would too. It does not: a bare name is rejected with `ValidationException: Malformed ARN - doesn't start with 'arn:'`, and so is `list-microvm-image-builds --image-identifier ` (`Invalid ARN format`). The construct therefore resolves an operator-supplied name to its exact `arn:${Partition}:lambda:${Region}:${Account}:microvm-image:${Name}` ARN **once** — the same value it scopes the lifecycle IAM grant to — and injects THAT as `MICROVM_IMAGE_IDENTIFIER`. One derivation, two consumers, so a request field and an IAM resource can never disagree. The strategy validates the invariant and fails fast with the remedy, because the service's own error names neither the env var nor the fix. - -The `ComputeStrategy` interface gains **mandatory** `suspendSession(handle)` / `resumeSession(handle)` methods returning a typed result (`{ supported: false } | { supported: true }`-shaped, exact type at implementation time); the widening lands in P3 (see sub-decision 5), in one commit across all three strategies. Mandatory-with-explicit-stub is the codebase idiom, not optional methods: no behavioral interface in the codebase has an optional method, `AgentCoreComputeStrategy.pollSession` is already a mandatory explicit stub rather than an optional member, and the exhaustive-`never` switch culture means a fourth backend must make a compile-checked decision about its suspend semantics instead of silently falling through a `strategy.suspendSession?.()` feature-detection. The agentcore and ecs strategies return `unsupported` (not a silent success — a suspend that silently no-ops would let the orchestrator believe compute billing stopped when it did not); the orchestrator gates its suspend policy on the typed response, consistent with how `pollTaskStatus` already branches explicitly on `computeType`. - -**Poll semantics — the strategy reports, the orchestrator interprets.** `pollSession(handle)` receives only the session handle and cannot see task state, so the health rules must live where the DynamoDB status lives. `SessionStatus` gains a `'suspended'` variant; the strategy maps `GetMicrovm` state mechanically and the **orchestrator** cross-references against the task row — the same division of labor `finalPollState` already uses for ECS (substrate stopped + non-terminal DynamoDB status → failed) and `pollTaskStatus` uses for agentcore heartbeats: substrate `suspended` + task `AWAITING_APPROVAL` is healthy (orchestrator-intended suspend); `suspended` with any other task status is an anomaly to surface, not fail-fast; substrate terminal + non-terminal task status → classify failed. +All strategies implement `startSession`, `pollSession`, `stopSession`, `suspendSession` and `resumeSession`. The lifecycle methods return an explicit supported/unsupported result. AgentCore and ECS return unsupported without making suspension API calls; MicroVM bounds each control request. An accepted API request does not prove the transition completed. -**Liveness on this backend is substrate state AND agent heartbeat.** The substrate cross-check above answers one question — "is the VM still there?" — and P2 established that it is not sufficient on its own. `GetMicrovm` catches a MicroVM that *died*; it cannot catch a MicroVM that is alive and reporting `RUNNING` while the pipeline **inside the guest** is hung, deadlocked, or was OOM-killed. That state is not hypothetical or self-correcting on this substrate, and the P2 live run narrowed *why* without weakening the conclusion. The service does reap a VM whose run hook FAILS (a 4xx makes it terminate within ~12 s — see sub-decision 2), so the P1 evidence for this paragraph (a hook-less image sitting in `RUNNING` indefinitely with no `stateReason`) no longer describes an ABCA image. What the service reaps is a hook *result*; it has no view into the guest afterwards. So the surviving — and more realistic — hang case is a `/run` hook that returned **200** and a pipeline that then hung, deadlocked or was OOM-killed behind it: the substrate stays `RUNNING`, the service is satisfied, and nothing else notices. Left to the substrate check alone, such a task would burn the orchestrator's full ~8.5 h poll window — billing an 8-hour reservation — before the safety net fired. +The MicroVM handle records `microvmId`, `endpoint`, the actual launched `imageArn`/`imageVersion`, and verified `lifecycleProtocol` when available. Lifecycle request fields use `microvmIdentifier`. Save the known handle before optional image discovery so a failed capability check cannot lose cleanup ownership. Current deployment settings cannot establish an older worker’s capabilities. -The in-guest half is already being written: the agent updates `agent_heartbeat_at` on the task row unconditionally, with no backend awareness, so the timestamp exists on every substrate. Only the orchestrator's *reaction* to it was backend-scoped — `pollTaskStatus` evaluated staleness for `agentcore` alone — which is the gap P2 closes by extending it to `lambda-microvm`. The grace and stale thresholds are the SAME on both: the timestamp is written by the same pipeline code at the same cadence, so a backend-specific window would encode a difference that does not exist. The two signals stay complementary rather than redundant — the substrate check is the crash detector, the heartbeat is the hang detector — and the check remains scoped to task status `RUNNING`, which is what keeps a deliberately suspended VM during an approval wait (sub-decision 2, P3) from being read as a dead one. `ecs` is deliberately left out: `DescribeTasks` reports a real container exit *with an exit code* (OOM-kill included) and the ECS poll block already interprets it with its own patience counters, so adding the heartbeat there would give one backend two independently-tuned kill paths for the same failure. - -The service's `MicrovmState` enum has **six** members, not three, so the mapping is stated exhaustively (one line of rationale each, mirrored in the strategy's doc comment): - -| `MicrovmState` | `SessionStatus` | Why | +| Service state | Strategy result | Coordinator responsibility | |---|---|---| -| `PENDING` | `running` | Still booting; the same way ECS's `PENDING`/`PROVISIONING` map to `running`. | -| `RUNNING` | `running` | — | -| `SUSPENDING` | `suspended` | Already on its way to frozen; reporting `running` would tell the orchestrator compute is still progressing when it is not. Both suspend states land on a report the orchestrator treats as benign-or-anomalous depending on task status, never as failure. **Never observable in practice** — suspend reached `SUSPENDED` in under 1 s live — so it is mapped for completeness and nothing may *wait* for it. | -| `SUSPENDED` | `suspended` | — | -| `TERMINATING` | `completed` | Terminal-bound and carries no exit code, so "the substrate is gone" is all the strategy can honestly say. | -| `TERMINATED` | `completed` | Success vs failure is the orchestrator's call — it cross-references the DynamoDB status. **This is the load-bearing terminal signal**, not `NotFound` (see below). | -| *unrecognized* | `running` | A future service enum addition must never fail a healthy task; the strategy warns and keeps polling. | - -**`GetMicrovm` `ResourceNotFoundException` → `completed`, but as a LATE fallback.** This deliberately diverges from `ecs-strategy`, where `DescribeTasks` returning no task maps to `failed`. ECS keeps stopped tasks describable for roughly an hour, so a missing task there really is anomalous; a MicroVM is eventually reaped from the control plane **by design**, so `failed` would fail tasks that finished cleanly. - -What the live run corrected is the *timing*, not the mapping. `NotFound` is **not the near-term terminal signal**: a terminated MicroVM reported `TERMINATING` at +1 s, `TERMINATED` at +3 s, and was **still `TERMINATED` ~10 minutes later** and at every subsequent checkpoint — `ResourceNotFoundException` was never observed in that window. So the branch that actually fires in practice is `TERMINATED → completed` in the table above; the `NotFound` rule covers only a VM reaped after a long gap (a poll resumed after a crash, a stranded-task reconciler sweep). Both are required and neither substitutes for the other: without the `TERMINATED` row the orchestrator would poll a finished VM until its safety net fired, and without the `NotFound` rule a late sweep would classify a cleanly-finished task as a poll error. - -The divergence is safe because it does not weaken detection: the orchestrator still fails the task when a terminal report lands while the DynamoDB status is non-terminal, so a genuine mid-run disappearance is caught — it simply receives the substrate-failure classification instead of a misleading poll error. Because that cross-check acts on a status read earlier in the same poll cycle, the orchestrator **re-reads the task row before failing** (the normal shutdown order is "agent writes terminal status → agent exits → VM terminates", which a stale read would otherwise turn into a spurious failure); ECS buys the same protection with a five-consecutive-poll patience counter instead. - -Neither the mapping nor the NotFound rule is a health decision: both are mechanical restatements of substrate state, which is what keeps the "strategy reports, orchestrator interprets" split intact. - -Normative requirements (EARS, per [ADR-020](/sample-autonomous-cloud-coding-agents/architecture/adr-020-ears-requirements-syntax)): - -- When a task's Blueprint sets `compute_type: 'lambda-microvm'`, the orchestrator shall resolve the `LambdaMicrovmComputeStrategy` via `resolveComputeStrategy`. -- When `startSession` is invoked, the strategy shall call `RunMicrovm` with `maximumDurationInSeconds` set to 28 800 (the service maximum, matching AgentCore's 8-hour session cap and sitting inside the orchestrator's ~8.5 h safety-net poll window). -- When `startSession` is invoked, the strategy shall pass a fully-qualified MicroVM image ARN as `imageIdentifier`. -- If the configured image identifier is not an ARN, then the strategy shall fail the session start with an error naming the environment variable and the redeploy remedy, before performing any AWS call. -- When `startSession` returns, the orchestrator shall persist the MicroVM handle (`microvmId`, `endpoint`) in the task row's `compute_metadata` (the field `cancel-task.ts` already reads ECS handles from). -- The strategy shall omit `idlePolicy` on every `RunMicrovm` call, in every phase. -- The orchestrator shall be the sole initiator of suspension, via `suspendSession`. -- When `pollSession` observes MicroVM state `SUSPENDED` or `SUSPENDING`, the strategy shall report `suspended` without interpreting task state. -- When `pollSession` observes MicroVM state `TERMINATED` or `TERMINATING`, the strategy shall report `completed` (the observable terminal state persists for at least ~10 minutes, so `completed` shall not depend on the MicroVM being reaped). -- When `pollSession` observes a MicroVM state it does not recognize, the strategy shall report `running`. -- If `GetMicrovm` reports that the MicroVM does not exist, then the strategy shall report `completed`. -- If the strategy reports a terminal substrate state while the task's DynamoDB status is non-terminal, then the orchestrator shall re-read the task row and, if it is still non-terminal, classify the task as failed with a substrate-failure remedy. -- If the strategy reports `suspended` while the task's DynamoDB status is not `AWAITING_APPROVAL`, then the orchestrator shall surface an anomaly event and shall not fail-fast the task. -- While a `lambda-microvm` task's DynamoDB status is `RUNNING`, if the task's `agent_heartbeat_at` is stale (or absent past the grace window) by the same thresholds the orchestrator applies to `agentcore`, then the orchestrator shall treat the session as unhealthy and stop polling — the substrate `GetMicrovm` check shall remain the crash detector, and the heartbeat shall be the in-guest hang detector. -- The task-detail API response shall include `agent_heartbeat_at`, and the CLI shall surface it while the task is non-terminal (P2r2-F11: the field drove the orchestrator's hang detector but was never projected, so no operator could observe the signal — and its invisibility produced a wrong verification conclusion). -- The task-**summary** API response (`GET /v1/tasks`) shall also include `agent_heartbeat_at`, and `bgagent list` shall render it as an age column. Extending the field to the list response is a deliberate widening of the requirement above rather than an incidental one: the detail-only projection makes liveness a per-task question, and an operator checking a fleet of tasks one `bgagent status` at a time is exactly how the hung task P2r2-F11 describes went unnoticed. Same suppression rule as the detail view (terminal tasks and never-beaten tasks render a placeholder) so the two views cannot disagree. -- If `suspendSession` or `resumeSession` is invoked on a strategy that does not support suspension, then the strategy shall return an explicit unsupported result. -- When the agent process reaches a terminal state, the agent shall exit. -- When the orchestrator finalizes a `lambda-microvm` task, the orchestrator shall call `terminate-microvm` (termination shall not rely on any substrate timeout, and shall not rely on the MicroVM self-terminating — it does not). - -*On the omitted `idlePolicy`:* if the block is present all three fields are required, so omission is the unambiguous disabled state the invariant test asserts. This deliberately forgoes `suspendedDurationSeconds` — it lives *inside* `idlePolicy` and cannot be set without re-enabling the traffic-idle machinery — so the suspended-state bound is `maximumDurationInSeconds` plus orchestrator termination and the stranded-approval reconciler (see sub-decision 2). A tighter substrate-level suspended-TTL remains available later as an additive `idlePolicy` change if operators want it. *On the fixed `maximumDurationInSeconds`:* no wall-clock task budget exists in the platform (budgets are `max_turns` / `max_budget_usd`), so the value is parity with AgentCore's 8 h cap rather than derived policy; a Blueprint override can be added later if a real need appears. - -### 2. Lifecycle: suspend/resume reconciled with the agent-owned approval poll - -The headline economic win is suspend during **HITL approval waits** (Cedar approval gates, [CEDAR_HITL_GATES.md](/sample-autonomous-cloud-coding-agents/architecture/cedar-hitl-gates)): while a task waits on a human decision, the MicroVM is suspended (compute charges stop; memory/disk state — cloned repo, warm build caches — is preserved) and resumed when the decision lands. Under Cedar decision #6 the approval window is bounded (default 300 s, ceiling 1 h, timeout → deny), so the saving per gate is bounded at ~1 h of compute — real at 16 vCPU, and it makes any future extension of gate ceilings (the off-hours posture §14.8 deliberately defers) cheap on this backend. - -The handshake must respect the existing approval mechanics: the agent **discovers decisions itself** by polling DynamoDB (`_poll_for_decision`, monotonic timeout), the approve/deny Lambda writes only the decision rows, and `AWAITING_APPROVAL` holds the concurrency slot (Cedar decision #7). Nothing "delivers" an approval to the agent, and suspension freezes the agent's monotonic clock — so the design is: - -- **Suspend — orchestrator-owned.** The orchestrator's durable poll observes `AWAITING_APPROVAL` on a `lambda-microvm` task and calls `suspendSession` after a grace period, and only when the gate's remaining window exceeds grace + resume overhead (suspending a 30 s gate is pure loss). Suspend is a policy decision on a poll observation, not a user action. -- **Resume — inline in the approve/deny Lambdas, orchestrator poll as backstop.** After the transactional decision write commits, `ApproveTaskFn`/`DenyTaskFn` load the MicroVM handle from the task row's `compute_metadata` (persisted at session start — the same field `cancel-task.ts` reads ECS handles from) via a post-commit strongly-consistent `GetItem`, then call `resumeSession` **best-effort**: on failure they log a warning and write a resume-orphan task event; the decision response never fails on a compute error (the decision row is already durable). The orchestrator poll reconciles: decision row present + MicroVM still `SUSPENDED` → retry resume (idempotent). +| RUNNING or starting | `running` | Read task state and applicable heartbeat | +| SUSPENDING / SUSPENDED | `suspended` | Reconcile the approval and lifecycle intent | +| TERMINATING / TERMINATED / not found | `completed` | Re-read task state; distinguish completion, recoverable checkpoint and failure | +| Unknown future state | `running`, with warning | Continue bounded observation | - *Why inline rather than poll-only — codebase precedent:* resume-on-approve is structurally identical to task cancellation — a user-initiated, latency-sensitive action whose purpose is an immediate compute-lifecycle side effect. `cancel-task.ts` already resolves this exact tension: the API-plane handler invokes ECS `StopTask` / AgentCore `StopRuntimeSession` inline, best-effort (a failed stop logs a warning and the state transition stands; a `task_cancel_compute_orphan` event is written when no stoppable compute handle exists, `reason: missing_runtime_handle`) — with the conditional IAM wired in `task-api.ts`. The resume path goes one step further than the precedent by also writing the orphan event on *failed* resume calls, because a failed resume strands a suspended VM awaiting a decision — a stronger liveness consequence than a failed stop of an already-cancelled task. The alternative (orchestrator-poll-only resume) preserves single-owner lifecycle purity but pays up to a full poll interval (~30 s) of latency on every approval, and the purity argument was already litigated and declined for cancel. `approve-task.ts` is deliberately minimal today (security-critical ownership comparison, Cedar finding #6); the resume call is therefore added *after* the transaction commits, cannot alter the decision outcome, and carries one conditional `lambda:ResumeMicrovm` grant — the same blast-radius trade the cancel handler accepted in review. -- **Timeout under freeze — the agent re-bases on the wall clock it already owns.** The agent's monotonic gate timer freezes while suspended, so resuming near the deadline is not enough: the frozen timer would still hold its remaining budget and fire the deny minutes *after* the user-visible window — colliding with the approval row's TTL (`created_at + timeout_s + 120s`) and triggering the "row reaped → stranded" fallback on a healthy gate. Instead, the gate expires at **`min(monotonic budget, created_at + timeout_s)`**, evaluated on each poll iteration and on `/resume`. This is not a new principle: Cedar decision #6 is already "min wins" for timeouts, the wall-clock deadline is already durable in the approval row the agent itself writes (`created_at` is in the agent's own clock domain — no skew), and §13.12's late-approval race fix already establishes that the durable row is authoritative over the agent's local timer. Deny authority stays agent-side (the conditional `TIMED_OUT` write + ConsistentRead re-read race protection is untouched); the orchestrator's resume at `deadline − margin` is purely the wake-up mechanism, with no correctness role. -- **Backstops, not mechanisms.** `maximumDurationInSeconds` (mandatory on every `RunMicrovm`, pinned at 28 800 s — see sub-decision 1) is the substrate kill switch bounding running **and** suspended time; the orchestrator's finalization `terminate-microvm` is the active cleanup path; the stranded-approval reconciler retains its role for orphaned waits. No `idlePolicy`-based bound is used in any phase — see sub-decision 1's omit-`idlePolicy` invariant. +A strategy result is not a task outcome. A terminal substrate can leave a recoverable pending-approval checkpoint; otherwise a non-terminal task needs failure classification. Capacity release separately requires confirmed physical shutdown, not merely the strategy’s `completed` result. - **The active terminate is still mandatory on the SUCCESS path, and P2 sharpened why.** P1 concluded flatly that "nothing self-terminates": a hook-less MicroVM reached `RUNNING` in 12 s and stayed there with no `stateReason` through every checkpoint. P2 refuted that *for the failure path only* — with `run: ENABLED`, a run hook that answers 4xx makes the **service** terminate the VM within ~12 s, `stateReason: "Run lifecycle hook returned HTTP status 400. Please check your hook endpoint and application logs for more details."`, after which `suspend-microvm` correctly refuses it. That is a real improvement in cost posture and a direct benefit of declaring hooks (see also the failure-path row in the phasing table, sub-decision 3). - - It does **not** relieve the orchestrator of anything, because the two cases are disjoint. The service reaps a hook *result* it did not like; it has no view of the guest once the hook returned 200. So a task that starts normally — the overwhelming majority — has no service-side reaper at all, and a VM whose pipeline finished, crashed after `/run`, or hung is reaped by nobody but `TerminateMicrovm`. A leaked handle therefore remains a cost incident that bills until the 8 h cap; only the "the guest rejected its own payload" corner now cleans itself up. -- **Concurrency slot stays held** during suspend. Cedar decision #7's rationale ("container alive, consuming memory") weakens under suspend, and the harder replacement rationale — "AWS counts `SUSPENDED` MicroVMs toward the account memory quota, so releasing ABCA's slot would not free real capacity" — is **undischarged**: the suspended VM stayed in `list-microvms` at every checkpoint, but that only proves *listed*. `L-CD1C0CC4` (1024 GB, account-scoped) exposes no `UsageMetric`, `AWS/Usage` carries only `CallCount` per API, and no MicroVM memory metric exists in any namespace, so consumption is **not observable safely** — proving it would need a large concurrent fleet. The conclusion (hold the slot) stands as the conservative choice, not as a verified fact. Size the arithmetic against the 32 GiB **peak** rather than the 8 GiB baseline: a busy fleet scales up, so peak is what actually competes for the account quota. - -The agent's `/suspend` hook flushes progress events (durable writes before returning 200, within the 60 s hook budget); `/resume` reseeds CSPRNGs and refreshes cached credentials. - -**Amended (P2 review): CSPRNG reseeding is P3 scope, and the P2 exposure is negligible.** The original EARS requirement below implied a `/run`-time reseed had to land with P2; it did not, and shipping P2 with the requirement unmet-but-asserted was itself the defect. Measured exposure, which is what changed the scoping: - -- The only consumer of a non-cryptographic PRNG in `agent/src` is `progress_writer.py`'s `getrandbits(80)`, used for the random half of a **ULID**. `os.urandom` / `secrets` are not seeded from the snapshot at all, and nothing in the agent derives a key, token, or nonce from `random`. -- That ULID is a DynamoDB **sort key under a `task_id` partition**. A collision therefore needs two events in the *same task* at the *same millisecond* with the same 80 random bits — and identical PRNG state across two MicroVMs restored from one snapshot does not produce that, because the events are in different partitions. The worst case is a duplicate progress event within one task, not a security boundary. -- There is no credential exposure from the snapshot's PRNG state: build-role credentials are kept out of the snapshot structurally (the build hooks make zero AWS calls, so `boto3.DEFAULT_SESSION` is never populated — see `_aws_silent_log`), and per-task credentials arrive via `platform_config.agent_session_role_arn` at `/run`. - -So the reseed moves to P3 alongside `/suspend` + `/resume`, where a *resumed* VM — which really does continue with the exact PRNG state it was frozen with, repeatedly — makes it load-bearing rather than theoretical. P3 must reseed on both `/run` and `/resume`, and must not treat the ULID as the only consumer: any future use of `random` for anything security-relevant needs the reseed in place first. +AgentCore and MicroVM start a heartbeat writer every 45 seconds. Their RUNNING tasks use the same 120-second startup grace and 240-second stale threshold. ECS’s batch entrypoint does not start that writer, so applying this check to ECS would fail healthy tasks. Heartbeats establish writer liveness, not progress of every coding thread. Returning from approval to RUNNING refreshes the heartbeat atomically. Task detail/list APIs and CLI views expose the timestamp. Normative requirements (EARS): -- While a `lambda-microvm` task is in `AWAITING_APPROVAL` and the gate's remaining window exceeds the configured grace period plus resume overhead, the orchestrator shall call `suspendSession` after the grace period. -- When the approve or deny Lambda commits a decision for a `lambda-microvm` task, the Lambda shall load the MicroVM handle from `compute_metadata` and call `resumeSession` best-effort. -- If the inline resume fails, then the Lambda shall record a resume-orphan task event and shall still return the decision outcome. -- While a decision row exists and the MicroVM remains `SUSPENDED`, the orchestrator shall retry `resumeSession`. -- While a `lambda-microvm` task waits on an approval gate, the agent shall evaluate gate expiry as the earlier of its monotonic budget and the row's wall-clock deadline (`created_at + timeout_s`), on each poll iteration and on `/resume`. -- If gate expiry is reached without a decision, then the agent shall deny. -- If no decision arrives by the gate's wall-clock deadline minus the resume margin, then the orchestrator shall resume the MicroVM so the agent can evaluate expiry and fire the deny agent-side. - -### 3. Packaging: same agent image source, new build path - -The existing agent container (`agent/` Dockerfile, already ARM64) is repackaged as a zip + Dockerfile artifact in S3 and built into a versioned `MicrovmImage` via `CreateMicrovmImage`. The agent runs its existing FastAPI server (`agent/src/server.py`) — the MicroVM path uses the HTTP entrypoint like AgentCore, not ECS's batch bypass — plus the runtime lifecycle hooks (`/run`, `/suspend`, `/resume`, `/terminate`) and the `/ready` + `/validate` build hooks, all on the same port the server already listens on (8080, declared as the image's `hooks.port`). Runtime hooks are fast-notification only (1–60 s): `/run` validates the payload and starts the pipeline **asynchronously**, mirroring how the agent loop already runs in a background thread behind `/ping` on AgentCore. - -**Hook phasing — corrected: `/ready` + `/run` are both P1.** The original plan split *declaring* a hook from *serving* it, putting `/run`'s declaration in P1 and its implementation in P2. Live verification proved that split is **not a reachable service state**, on two independent counts: - -- `CreateMicrovmImage` rejects an image that enables any lifecycle hook without `/ready`: *"The ready (/ready) MicroVM image hook must be enabled when any MicroVM lifecycle hook (run, resume, suspend, or terminate) is enabled."* So a P1 image declaring only `/run` is **not creatable**. -- With `/ready` added but unserved, both chipset builds fail: *"Ready hook check failed: the application returned a client error (HTTP 4xx) response."* So a declared hook must be served in the same phase. -- And an image with **no** hooks at all — the only other creatable shape — cannot receive a payload: *"The run hook must be enabled in the MicroVM image to pass the run hook payload."* So deferring hooks entirely also defers the whole payload-delivery channel. - -The phasing is therefore: +- When a Blueprint selects `lambda-microvm`, the orchestrator shall use the MicroVM strategy and persist its handle in `compute_metadata`. +- Every launch shall use a full image ARN, `maximumDurationInSeconds=28800`, explicit `NO_INGRESS` and no `idlePolicy`. Invalid image configuration shall fail before launch. +- Before allowing suspension, the coordinator shall verify the actual launched image version and lifecycle protocol. +- When a suspended state conflicts with task state, the coordinator shall record and reconcile the anomaly within bounded recovery windows. +- When the task ends, the coordinator shall actively terminate its MicroVM. The service lifetime is a backstop, not the normal cleanup mechanism. +- Uncertain starts, lost replies and supervisor replay shall preserve the original worker identity, ownership and lifetime rather than launch duplicate workers. -| Hook | Declared by | Served by the agent | Notes | -|---|---|---|---| -| `/ready` | **P1** (construct enables `hooks.microvmImageHooks.ready`) | **P1** | MANDATORY, not a quality nicety — see above. A 200 proves uvicorn is bound and `server` imported cleanly (pulling in `pipeline` → `runner` → the policy engine), so a missing policy file fails the BUILD instead of the first task. **Since P2-F5 it also WARMS the snapshot** — the hook's 200 is what the service waits for before capturing the snapshot, making this the only place a warm page can be created, and the 225 MiB `claude` binary was cold in it (see the P2-F5 correction below). A required warm-up failure answers 503, so a snapshot that cannot exec the agent's own CLI fails the image build instead of every task. Still makes ZERO AWS calls, logging included (a `--version` exec is neither an AWS call nor a network call). | -| `/run` | **P1** (construct sets `hooks.microvmHooks.run`) | **P1** | The payload-delivery channel. Must be served in P1 because `/ready` forces hooks to exist at all, and a hook-less image cannot accept `runHookPayload`. Since P2 it is also the **platform-configuration** channel (see "Platform configuration delivery" below). | -| `/validate` | **P2** (construct sets `hooks.microvmImageHooks.validate`) | **P2** | An **image** (build-time) hook, and a **shallow self-check only**: server alive, hook routes registered, interpreter + contract sanity. It runs under the BUILD role, which deliberately holds no Bedrock / Secrets Manager / DynamoDB grants, so it must make **zero AWS API calls** and must not touch credential resolution — the "deeper warm-up assertions (Bedrock reachability, Memory access, tool availability)" this ADR originally assigned here are **not implementable**: every one of them would `AccessDenied` and fail every image build. They belong to the first task's own error handling. 200 when the checks pass, 503 while still initialising. | -| `/terminate` | **P2** (construct sets `hooks.microvmHooks.terminate`) | **P2** | Best-effort final flush: a final structured log line, then 200 — always, inside the hook budget, even with nothing running. It must **not** write terminal task status (the orchestrator finalizes the task and *then* calls `TerminateMicrovm`, so a terminate hook that wrote a status would race that finalization and could clobber the real outcome) and must not join the pipeline thread. There is nothing buffered to flush: `_ProgressWriter` does a synchronous `put_item` per event, so durability is per-write. "Always 200" also covers the BODY: the handler reads the raw request rather than a typed model, because a typed body is validated before the handler runs and would answer 422 to malformed JSON — a reported hook failure on a successful teardown. Safe to declare because `TerminateMicrovm` removes the VM with or without in-guest cooperation. **Correction (P2-F8):** the service sends `microvmId: ""` on this hook, unlike `/run` where it is populated, so an empty id is expected-normal and this hook cannot join the guest record to the control-plane one — `/run`'s accepted line carries that correlation instead. | -| `/suspend`, `/resume` | **P3** | **P3** | Declaring a runtime hook the agent does not serve fails the corresponding lifecycle transition, so each is declared only in the phase that implements it. P1 termination is the orchestrator's `TerminateMicrovm`, which needs no in-guest cooperation. | +### 2. Lifecycle: suspend/resume reconciled with the agent-owned approval poll -Consequence to state plainly, replacing the original "a P1-built MicroVM image is not runnable end to end": **a P1 image is creatable, launchable and payload-deliverable, but carries no smoke-parity guarantee.** P1 delivers the strategy, the construct, the roles/buckets/connectors, the image resource, the packaging script, and the `/ready` + `/run` endpoints — so a `lambda-microvm` task can start a MicroVM and hand it a payload. What P1 has **not** established is anything P2 owns: AgentCore Memory grants and `MEMORY_ID` delivery, the agent's non-secret env parity inside the snapshot, egress specifics from a running MicroVM, and heartbeat/progress behaviour end to end. No clone → change → PR run has happened on this substrate. P2 ("smoke parity") is the phase that closes that gap. The construct and the packaging script both surface exactly this at synth/run time (`abca:microvm-image-p1-smoke-unverified`) so an operator cannot mistake a launchable substrate for a verified one. +Unanswered approvals have no deadline by default (`approval_timeout_s=0`). Explicit task deadlines range from 30 to 3,600 seconds; a positive policy-rule deadline can also apply. The independent sleep preference defaults to 600 seconds per approval wait. `microvm_sleep_after_s=0` keeps the task awake. -**The `AWS::Lambda::MicrovmImage` L1 enforces the API's enums (P2-F2, live 2026-08-06).** This closes the one item P1 left explicitly open, and it closes it against the construct's own stated reasoning. CloudFormation's generated types make `cpuConfigurations[].architecture` and all four `hooks.*` fields plain strings and document no allowed values, from which P1 concluded that the CloudFormation surface takes a *hook path* while the API takes an `ENABLED`/`DISABLED` flag, and that both were correct for their own surface. CloudFormation refused the change set at **early validation** — the stack was never touched, so there was no rollback and no runtime symptom to trace back — on five values: +Automatic suspension also requires the deployment’s `microvm_approval_suspend_enabled` opt-in, which defaults false for new deployments. A live Parameter Store switch lets existing durable executions stop initiating new suspensions without changing their pinned Lambda environment. The verified normal deployment has this opt-in enabled. Turning it off does not abandon already-suspended workers. -``` -/aws/lambda-microvms/runtime/v1/run is not a valid enum value. Supported values: [DISABLED, ENABLED] - (at /Resources/…/Properties/Hooks/MicrovmHooks/Run) … and the same for Terminate, Ready, Validate -arm64 is not a valid enum value. Supported values: [ARM_64] - (at /Resources/…/Properties/CpuConfigurations/0/Architecture) -``` +The coordinator observes the exact pending gate and waits for the sleep delay. It skips sleep when a timed gate has too little time left before its wake margin. The guest holds a coding barrier, drains acknowledged progress and commits a checkpoint before accepting `/suspend`. Lifecycle HTTP responses explicitly close their connections before freeze to avoid reuse of a stale pooled connection; see the [transport evidence summary](/sample-autonomous-cloud-coding-agents/verification/readme#recorded-acceptance). -Three consequences. First, the CloudFormation surface is **identical** to the API surface, and the packaging script (`--cpu-configurations '[{"architecture":"ARM_64"}]'`, `--hooks '{"microvmHooks":{"run":"ENABLED",…}}'`) had it right all along. Second, the "CDK-managed (recommended)" bootstrap path was **non-functional** for the whole of P1 and P2 — the out-of-band `--create-image` script was the only working path — and no unit test, `cdk synth` or cdk-nag rule could see it, because the types accept any string. Third, **hook paths are not configurable on either surface**: the service calls fixed well-known routes (proved by the build and run logs, which POST to exactly the `/aws/lambda-microvms/runtime/v1/*` paths the agent serves), so the route constants in the construct are an agent-side cross-package contract ONLY and must never be sent as property values again. Also discharged in passing: the `microvmImageHooks` property name and nesting are correct — CloudFormation resolved `…/Hooks/MicrovmImageHooks/Ready` and objected only to its value. +Approval and denial handlers commit the decision first. They then read the current handle consistently, persist wake intent and request resume best-effort. A wake failure records diagnostics and does not undo the accepted decision. The durable supervisor retries and observes both service state and guest consumption of the decision; RUNNING alone does not prove the tool was released. -**A snapshot is only as warm as the pages touched before it was captured (P2-F5, live 2026-08-07).** This is the defect that stopped the P2 smoke run one step short of a pull request, and it is a property of the substrate rather than a bug in any one file. Every task failed at turn 0, reproducibly: +A sleeping worker retains its concurrency reservation. For a longer wait, ABCA verifies a complete, version-pinned conversation/workspace checkpoint, fences the attempt, confirms shutdown and releases the reservation. A later decision can admit one replacement through the original published coordinator. It restores Git state, required workspace files, the actual SDK conversation, exact pending tool inputs and cumulative usage with fresh scoped credentials. See the [continuation protocol](/sample-autonomous-cloud-coding-agents/architecture/orchestrator#retained-microvm-approvals). -``` -TimeoutExpired: Command '['claude', '--version']' timed out after 10 seconds -``` +Normative requirements (EARS): -The binary was fine — in the identical image, locally, `claude --version` answers `2.1.191 (Claude Code)` in under a second. It is a **225 MiB (236,305,136-byte) statically-linked ELF** that nothing had exec'd before the snapshot was taken, so on a guest restored ~50 s earlier the first `exec` had to fault all of those pages in from lazily-restored storage, and 10 s was not enough. `/ready` existed precisely so "the snapshot is taken with a warm server", and the snapshot was warm for uvicorn and stone cold for the binary that does all the work. +- While all sleep gates and guest safety checks hold, the coordinator shall suspend after the configured delay; it shall remain the sole initiator of suspension. +- When a decision is committed, the API shall preserve that outcome even if optional wake or replacement dispatch fails. +- A timed request shall retain its original wall-clock deadline and, within the original process, its monotonic cap. The earlier limit wins; wake/replacement shall not restart either decision window. +- Before an unanswered timed gate reaches its deadline, the coordinator shall wake the worker with the configured margin. For a retired attempt, it may resolve expiry atomically while admitting a replacement. +- Before releasing a worker reservation, the coordinator shall verify checkpoint integrity, fence ownership and confirm physical shutdown. An uncertain control request shall not release capacity. +- A replacement shall preserve the task/request/tool identity and usage totals, obtain fresh scoped credentials and use a new authenticated launch reference. +- Pending requests shall have no storage TTL. Cancellation or terminal cleanup shall close unanswered requests and preserve recorded decisions without extending an existing retention TTL. +- The guest shall reseed application PRNG state from OS entropy on `/run` and `/resume`; security-sensitive values shall continue to use cryptographic randomness. -Both halves of the fix are kept, because they answer different questions. `/ready` now **exec's the heavyweight binaries before returning 200** (`claude` required, `git`/`node` best-effort), which is the only mechanism that can make the shipped snapshot warm — and its own budget rises to 300 s, well inside the 3600 s build-hook window, because it now does work whose duration is a cold `exec`. Two structural rules keep that honest, because **per-command timeouts do not compose**: the required command runs FIRST with its own budget so no best-effort warm-up can starve the one that decides whether the snapshot is usable, and the best-effort ones then SHARE the remainder of a total warm-up ceiling that sits inside the hook budget with margin (240 s against 300 s). Without them, three commands at 120 s each would be 360 s — a fix for a runtime failure that produces a build failure instead — and a single hung optional command could hold up a 200 that the required warm-up had already earned. Separately, the version probe's timeout goes from 10 s to 60 s: a probe that exists to print a version string into a log line gains nothing from a tight bound and loses the whole task when it trips. The general rule this generalises to, and the reason it belongs in the ADR rather than only in a comment: **on this backend, a first-touch cost that other substrates pay during container start is deferred to the first task instead**, so anything large and lazily-loaded is a turn-0 hazard unless it is touched in `/ready`. +The platform verifies ownership and the exact approved action. It does not implement a separate semantic relevance checker; the agent decides whether the proposed work still makes sense. -**Payload delivery** reuses the ECS strategy's S3-pointer pattern, adapted to `runHookPayload` (**≤ 4 KB** — measured, see below): payloads that fit ride inline; the rest are uploaded by the strategy to a platform payload bucket (the ECS payload bucket pattern in `ecs-agent-cluster.ts`: orchestrator write access, compute-role read-only scoped to the bucket, lifecycle expiry on objects) with the S3 URI in `runHookPayload` in place of the payload itself — the MicroVM **execution role** holds the read grant, exactly as the ECS task role does today. Since P2 the hook body also carries `platform_config` in both branches, so `runHookPayload` is never *only* the URI (see the canonical shapes below). +### 3. Packaging: same agent image source, new build path -The cap is **4 096 bytes**, not the 16 384 the SDK documents. Measured exactly: 4 096 passes, 4 097 is rejected with *"Value at 'runHookPayload' failed to satisfy constraint: Member must have length less than or equal to 4096"*. Two consequences follow. First, the original threshold would have inlined every envelope between 4 097 and 16 384 bytes and had the service reject all of them. Second, and more structurally: **the S3-pointer path is now the dominant one, and inline is the exception.** A hydrated task payload (prompt + issue thread + repo context) essentially always exceeds 4 KB, so "small payloads ride inline" describes tiny repo-less prompts rather than the common case. The payload bucket is therefore not a rarely-exercised overflow valve but a required part of every normal task, which raises its lifecycle rule (`MICROVM_PAYLOAD_TTL_DAYS`) and the execution role's read grant from edge-case plumbing to load-bearing. +Package the existing ARM64 `agent/Dockerfile` and its local inputs into a deterministic ZIP. Managed builds use `microvm-images/agent-artifact-.zip`; deploy the digest with the managed base-image ARN/version. A changed object URI triggers CloudFormation to build a new image version. The first deployment may create only infrastructure so the artifact bucket exists before upload. An external image identifier supports out-of-band builds. -**Canonical wire shapes.** Three, and the producer (`lambda-microvm-strategy.ts`) emits exactly these: +All six hooks share the FastAPI listener on port 8080. AWS hook properties accept `ENABLED`/`DISABLED`, not route paths; the architecture enum is `ARM_64`. -| Where | Exact shape | +| Hook | Contract | |---|---| -| `runHookPayload`, inline branch | `{"agent_payload": {…}, "platform_config": {…}}` | -| `runHookPayload`, pointer branch | `{"agent_payload_s3_uri": "s3://…", "platform_config": {…}}` | -| the object at that S3 URI | `{…agent_payload fields…, "platform_config": {…}}` — the payload's own fields at the TOP level, with the config merged in beside them | - -Two asymmetries are deliberate and must not be "tidied" without changing both sides. First, **the S3 object is not the envelope**: the payload's fields sit at the top level (that is what P1 uploaded, before `platform_config` existed) rather than nested under an `agent_payload` key. Second, **`platform_config` is duplicated** on the pointer path — once beside the pointer, once inside the uploaded object. It costs a few hundred bytes and buys the property that the config is reachable whichever end of the fetch a reader looks at, which matters because it is the agent's only substitute for an env block. - -The agent's reader is deliberately more permissive than this contract: it also accepts an S3 object shaped like the envelope (`agent_payload` nested), and `platform_config` present in only one of the two places (the fetched object wins, the hook body is the fallback). Those are **defensive compatibility** for the independent deploy cadences of a snapshot image and the orchestrator Lambda — a tolerant reader, not an alternative contract. A producer must emit the three shapes above. - -**Platform configuration delivery (P2): payload-sourced, allowlisted, fail-closed.** The other two backends hand the agent its non-secret platform env at launch — AgentCore Runtime env vars, ECS container overrides — and there is no equivalent on this substrate: a MicroVM starts from a **snapshot**, so its process environment is whatever was frozen at *image build* time and is then replayed by every MicroVM launched from that image version. Baking the deployment's identifiers into the snapshot would make them **version-frozen**: a redeploy that renames a table, adds a bucket or rotates the session role would leave every existing image version describing a deployment that no longer exists, and the drift would surface as a task-time `ResourceNotFound` rather than a deploy-time error. So the values travel with the task instead: `platform_config` is a SIBLING of the payload — beside `agent_payload` in the inline branch, beside `agent_payload_s3_uri` in the pointer branch, and merged in beside the payload's own fields inside the S3 object (the canonical shapes above give each one exactly) — whose snake_case keys the agent installs into `os.environ` as their UPPER_SNAKE equivalents. A payload value therefore **wins** over any pre-existing/image value — the orchestrator is describing the live deployment, the snapshot is describing a past one. `platform_config` carries **non-secret identifiers only** (table and bucket names, secret ARNs, the session-role ARN); secrets are still fetched at `/run` time from Secrets Manager using those ARNs, so the snapshot-must-stay-secret-free requirement above is untouched. Per-task fields — `memory_id` and friends — stay inside `agent_payload`: `platform_config` configures the *process*, `agent_payload` describes the *task*. - -Two rules make it safe. First, **the allowlist fails closed**: the agent installs a fixed set of keys and *rejects the entire run* (HTTP 400, nothing spawned, not one key installed) if the block carries anything else. These values become environment variables of the process that spawns the agent's tool subprocesses, so an unrecognised key is an attempt to set an arbitrary variable in the agent (`AWS_ENDPOINT_URL`, `LD_PRELOAD`, `PATH`, …) — an injection attempt, not a forward-compatibility gap, which is why unknown keys are refused rather than filtered out. Second, **installation happens before any credential or pipeline initialisation** on the hook path: the very next step reads `GITHUB_TOKEN_SECRET_ARN` to resolve the GitHub token and `AGENT_SESSION_ROLE_ARN` to scope the task's credentials, so installing later would silently resolve the whole task against the snapshot's frozen env. The one call that must precede installation is the S3 payload fetch (the config is *inside* the fetched object), which therefore runs on the ambient compute role via the attributed platform client — and it is the ONLY one: the same rule covers **logging**, so every `/run` log line before the install is stdout-only. The CloudWatch writer would otherwise resolve credentials and pin a boto3 default session (region included) off whatever a snapshot happened to bake, which is the build-hook defect one phase later. Nothing is lost — in the intended deployment there is no baked `LOG_GROUP_NAME`, so those lines would have gone to stdout anyway, and the reason for every pre-install rejection also travels in the structured 4xx/5xx body the service surfaces. A **required subset** (task table, task-events table, GitHub token secret ARN, session-role ARN) is rejected as `…_INCOMPLETE` when missing or blank — a distinct wire code from the `…_INVALID` allowlist rejection, because the remedies differ (deployment wiring vs. producer bug). A `/run` envelope with *no* `platform_config` at all is still accepted, loudly warned: the image snapshot and the orchestrator Lambda deploy on independent cadences, and a new image must not require a same-instant orchestrator. The key set is a cross-package contract in `contracts/constants.json` (`microvm_platform_config`), consumed by the agent's `/run` hook and produced by the orchestrator, with shape and required-subset invariants enforced by `scripts/check-constants-sync.ts`. - -**No orchestrator→agent HTTP path exists in P1–P3**: payload arrives through the `/run` hook, all agent work is outbound, and therefore **no JWE auth tokens are minted at all** — token minting (and its ≤ 60 min TTL refresh problem) is deferred until a real consumer exists (e.g. operator shell access, [#391](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/391)). The `endpoint` stays in the `SessionHandle` because it is genuinely per-session state that becomes load-bearing the day such a consumer appears. But note the service does not agree by default: omitting `ingressNetworkConnectors` on `RunMicrovm` attaches a **public** `HTTP_INGRESS` connector, so the strategy passes the Lambda-managed `NO_INGRESS` connector explicitly on every launch (see sub-decision 4's security table). - -**Constraint accepted:** the configured **baseline is 8 GiB RAM / 4 vCPU** and the service scales vertically to a **32 GiB / 16 vCPU peak** on its own, with 32 GB of disk. So capacity is baseline-priced with 4× burst headroom — good for the bursty compile-and-test shape of an agent task — but the SUSTAINED ceiling is still 32 GiB, so repos that motivated the 120 GB ECS sizing stay on `ecs`. What the construct configures (and validates) is the baseline; the peak is not something a deployment asks for. - -Normative requirements (EARS). **Each requirement's own `(Pn)` tag is authoritative**; there is no blanket phase for the list. The tags are per-requirement because the original list *was* split P1/P2 on the assumption that a hook could be declared in one phase and served in a later one — which the service does not permit (see the phasing table above), so the hook-serving requirements collapsed into P1 while the P2 items below arrived with the P2 hooks and `platform_config`: - -- (P1) The image build shall not embed secrets, tokens, or per-task identity in the snapshot. -- (P1) If the task payload exceeds the 4 KB `runHookPayload` limit, then the strategy shall upload the payload to the platform payload bucket and pass its S3 URI in `runHookPayload` in place of the payload (from P2, alongside `platform_config`). -- (P1) The MicroVM execution role shall hold read-only access to the payload bucket, scoped to that bucket. -- (P1) Where no ingress is configured for a deployment, the strategy shall pass the Lambda-managed `NO_INGRESS` network connector on every `RunMicrovm` call (the field shall not be omitted). -- (P1) Where the image enables any MicroVM lifecycle hook, the image shall also enable the `/ready` hook and the agent shall serve it. -- (P1) When the `/run` hook receives the task payload, the agent shall validate it, start the pipeline asynchronously, and return HTTP 200 within the hook budget. -- (P1) If the `/run` hook payload cannot be resolved to a task payload, then the agent shall reject the hook with a client error and shall not start a pipeline. -- (P1) The agent shall not execute the clone→verify→PR pipeline on the hook path. -- (P1) The agent shall resolve credentials at `/run` time. -- (P1) Where a deployment configures a MicroVM image before smoke parity is verified, the platform shall warn that the backend has no smoke-parity guarantee. -- (P2) Where the image declares the `/validate` build hook, the agent shall serve it. -- (P2) The `/validate` hook shall make no AWS API calls. -- (P2) When the `/run` hook receives `platform_config`, the agent shall install only allowlisted keys into the environment before pipeline initialization. -- (P2) If `platform_config` carries a key that is not on the allowlist, then the agent shall reject the run with a 400 and shall install none of the block's keys. -- (P2) If a required `platform_config` key is missing, then the agent shall reject the run with a 400. -- (P2) Where a `platform_config` value and an image-baked environment value disagree, the agent shall use the `platform_config` value. -- (P2) Until `platform_config` is installed, the `/run` hook shall make no AWS API call other than the payload fetch, and shall log to stdout only. -- (P2) Where the image declares the `/terminate` hook, the agent shall return 200 within the hook budget for any request body — including a malformed, empty or absent one — and shall not write terminal task status. -- (P2) When the `/ready` hook runs, the agent shall exec the agent CLI binary before returning 200, so that its pages are resident when the snapshot is captured. -- (P2) If a required `/ready` warm-up does not complete successfully, then the agent shall report not-ready (HTTP 503) rather than allow the snapshot to be taken. -- (P2) The `/ready` hook shall make no AWS API call, warm-up included. - -### 4. Infra and IAM: conditional resources behind bootstrap `ComputeTypes` - -Mirroring the ECS pattern: a `compute-lambda-microvm` bootstrap policy (`cdk/src/bootstrap/policies/`) gated on the `ComputeTypes` CFN parameter; a CDK construct provisioning the build role, execution role (admitted to the per-session role via `AgentSessionRole.admitComputeRole`, which was designed for exactly this), the S3 artifact bucket wiring, and image build automation. Egress uses the platform VPC via egress network connectors so the DNS Firewall / security-group / flow-log stack in [COMPUTE.md](/sample-autonomous-cloud-coding-agents/architecture/compute) applies unchanged; ingress is suppressed with the Lambda-managed `NO_INGRESS` connector (no `SHELL_INGRESS` — it is noted as a candidate for [#391](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/391), operator session access, as a separate decision). - -Two networking facts the construct has to encode, both established live: - -- **A `VPC_EGRESS` connector requires an operator role.** CloudFormation's generated L1 types `operatorRole` as optional and this ADR originally assumed Lambda would manage the ENIs with its own service-linked role. It does not: the connector fails to create with *"NetworkConnectorOperatorRole is required for VPC_EGRESS connector type"*. The construct creates one role — trusting the bare `lambda.amazonaws.com` service principal (see the trust-policy fact below), carrying `AWSLambdaVPCAccessExecutionRole` plus the ENI / tag / private-IP actions that policy omits — and shares it across both connectors, since it manages interfaces rather than traffic. -- **The MicroVM-facing roles cannot carry a confused-deputy source condition.** All three (build, execution, connector operator) trust the bare `lambda.amazonaws.com` service principal with **no** `aws:SourceAccount` / `aws:SourceArn`, and that is a forced choice, not an oversight: the Lambda MicroVMs service presents no source key when it assumes them, so a trust policy carrying one is unassumable. Two symptoms of the one cause, both live 2026-08-06/07 and both blocking: - - - Both `AWS::Lambda::NetworkConnector` resources `CREATE_FAILED` **deterministically** (on a freshly deleted stack, so not propagation lag — which matters, because *"The service is unable to assume the provided NetworkConnectorOperatorRole. Please verify the trust policy on the role."* is also the classic propagation symptom and a re-run is the obvious wrong guess). Removing the condition → both created within a second. - - `RunMicrovm` failed with a **misleading `iam:PassRole` AccessDenied on the caller**, with the orchestrator's grant present, `simulate-principal-policy` returning `allowed`, no permissions boundary, and a temporary *unconditioned* `iam:PassRole` **also** denied. The real cause was the execution role's trust; removing its conditions made the next submission reach `RUNNING` in 6 s. So the service reports a role it cannot pass-and-assume as an identity-policy denial on the principal passing it. - - Recorded plainly because the fix looks like a regression to anyone applying the standard service-principal pattern — and because it *was* a regression in the other direction: P1's standalone-validated operator-role probe had no conditions and worked, and the P1 F2 fix then added them "to mirror the build/execution roles". `sts:TagSession` stays: the service needs both actions and it was never implicated. - -- **Neither can the `iam:PassRole` grants carry an `iam:PassedToService` condition — same root cause, identity side (P2r2-F9 + P2r2-F10, live 2026-08-07 run 2).** An earlier revision of this ADR recorded the opposite, that the identity-side condition "was exonerated" by run 1's elimination. **That was a false negative**, and its cause is worth recording because it is a general trap: run 1 tested the conditioned grant by *adding* a temporary unconditioned `iam:PassRole` and watching the task still fail — but the temporary grant remained attached through the later submissions that succeeded, so the conditioned grant was never once tested against a working trust policy. A contaminated control. - - Run 2 ran the clean experiment — same exact-ARN resource, same ~5-minute IAM settle, one variable. It removed the run-1 workaround **first** (submission 4: denied) and only then added the unconditioned grant back on the same resource (submission 5: `RUNNING`), which is the ordering run 1 got wrong: - - | Orchestrator `iam:PassRole` on the execution role | Result | - |---|---| - | exact ARN **+ `iam:PassedToService: lambda.amazonaws.com`** | **DENIED** (two independent submissions) | - | exact ARN, **no condition** | **`RUNNING` in 9 s** | - - The denial lands on the **caller**, which is what makes it so misleading — the statement names that exact ARN and `simulate-principal-policy` answers `allowed`: - - ``` - User: arn:aws:sts:::assumed-role/backgroundagent-dev-TaskOrchestratorOrchestratorFn-… - is not authorized to perform: iam:PassRole on resource: - arn:aws:iam:::role/backgroundagent-dev-LambdaMicrovmComputeExecutionRo-… - because no identity-based policy allows the iam:PassRole action - ``` - - And the same key blocks the *other* PassRole path, which run 1 never reached because the enum defect (P2-F2) stopped it earlier: CloudFormation could not pass the **build role** at `CreateMicrovmImage` under the bootstrap `infrastructure` policy's allowlisted `IAMPassRole`. Verbatim, so the diagnosis does not have to be taken on trust: +| `/ready` | Required when runtime hooks are enabled; execute required binary warm-up before snapshot capture | +| `/validate` | Check local readiness, routes and configuration contracts without AWS calls | +| `/run` | Authenticate/install launch configuration and start the pipeline asynchronously | +| `/terminate` | Close the local coding barrier, log and acknowledge any request body; do not join the pipeline or write terminal task status | +| `/suspend` | Drain acknowledged progress and commit the current safe checkpoint within the hook budget | +| `/resume` | Renew credentials and reconcile the original gate before releasing coding | - ``` - LambdaMicrovmComputeImage… CREATE_FAILED - User: arn:aws:sts:::assumed-role/cdk-hnb659fds-cfn-exec-role--us-east-1/AWSCloudFormation - is not authorized to perform: iam:PassRole on resource: - arn:aws:iam:::role/backgroundagent-dev-LambdaMicrovmComputeBuildRoleF0-… - because no identity-based policy allows the iam:PassRole action - (Service: LambdaMicrovms, Status Code: 403) - ``` +Warm-up budgets come from `contracts/constants.json`; the total guest budget must remain below the image hook timeout. Cold-binary startup exceeded the hook budget in an earlier image; warm-up moved this work into image preparation. Historical sizes and timings are not sizing guarantees for later builds. - Three pieces of evidence pin that to the *condition* rather than to a stale bootstrap or a wrong resource pattern: +**Authenticated v2 payload transport.** Every ECS and MicroVM task uses S3. The coordinator publishes a deployment manifest, conditionally creates the task payload and sends a short-lived signed URL for that one object. The serialized MicroVM reference must fit 4,096 bytes; payloads are bounded at 8 MiB and manifests at 16 KiB. - 1. the live `IaCRole-ABCA-Infrastructure` policy was byte-identical to this branch's `cdk/bootstrap/policies/infrastructure.json`, so `cdk bootstrap --force` would have changed nothing; - 2. `aws iam simulate-principal-policy --policy-source-arn --action-names iam:PassRole --resource-arns ` returned `allowed` **with** `--context-entries ContextKeyName=iam:PassedToService,ContextKeyValues=lambda.amazonaws.com,ContextKeyType=string` and `implicitDeny` with no context entry — so the resource pattern matches and the condition key is the only remaining variable; - 3. the **control**: the out-of-band `create-microvm-image` call passed the *same build role* to the *same service* successfully, using operator credentials that carry no such condition. The role's trust is therefore fine and the denial is genuinely caller-side. +| Location | Shape | +|---|---| +| Hook reference / ECS `AGENT_PAYLOAD_REF` | `{version:2, task_id, bootstrap_s3_uri, payload_url, expires_at}` | +| `bootstrap/.json` | `{version:2, backend, platform_config}` | +| `/payload.json` | `{version:2, task_id, agent_payload, platform_config}` | +| Private `/launch.json` | `{fingerprint, reference}` for coordinator replay | - So: **the Lambda MicroVMs service presents no usable value for `iam:PassedToService` on either PassRole path** (CloudFormation → build role at `CreateMicrovmImage`; orchestrator → execution role at `RunMicrovm`), exactly as it presents no `aws:SourceAccount` on the assume-role path. One root cause, two more symptoms. Both statements therefore drop the condition, and the fix is deliberately asymmetric so it stays contained: +The worker authenticates only its deployment’s `bootstrap/*` using ambient credentials. Other payload-bucket object reads and listing are explicitly denied; signed task downloads carry coordinator authorization. The task ID and configuration must exactly match the reference and authenticated manifest before installation. MicroVM `platform_config` contains allowlisted non-secret identifiers, including role/secret ARNs; it does not contain credentials. ECS already receives its deployment configuration through task settings, so its manifest config is empty. - - `task-orchestrator.ts` sid `MicrovmPassExecutionRole` — condition removed; the **exact execution-role ARN** is now the whole of the scoping, which is why that resource must never be relaxed to a prefix or `*`. - - a new sid `MicrovmPassRoles` in the **conditional** `compute-lambda-microvm` bootstrap policy — unconditioned `iam:PassRole` on the build- and connector-operator role **name prefixes only** (not the execution role, which CloudFormation never passes). The shared `infrastructure` `IAMPassRole` keeps its allowlist, so no other role in the stack loses that constraint, and an agentcore-only bootstrap never gains an unconditioned pass at all. **Operators must re-bootstrap** (bundle ≥ 1.6.0) for the CDK-managed image path to work. +The coordinator saves the exact reference for idempotent replay. Signed URLs last at most 900 seconds, bounded by known credential expiry, and initial creation requires at least 300 seconds. An expired saved reference fails without re-signing the same launch request. Finalization deletes task payload/launch objects; one-day asynchronous S3 expiry is the backstop. - If AWS documents the value the service does present, adding it to both statements restores the condition. `microvms.lambda.amazonaws.com`, `lambda-microvms.amazonaws.com` and `microvms.amazonaws.com` were all `implicitDeny` against the conditioned policy, so any one of them would serve as the allowlist entry if it turns out to be right. Note that CloudTrail carries **no `lambda-microvms` management events at all** today, so the value cannot be read out of a log — only confirmed by AWS or found by a bounded sweep. +Normative requirements (EARS): - Compensating controls, enumerated per role. The deployment-role's shared `IAMPassRole` grant is **name-prefix-scoped**, so it technically covers all three roles; in practice only the orchestrator actively invokes `iam:PassRole` on the execution role (CloudFormation never requests it). The other two roles are passed **to** the deployment role (not to themselves). +- Image builds shall not embed credentials, task identity or deployment-specific configuration. Build hooks shall avoid AWS clients and credential caches. +- The worker shall reject legacy envelopes, oversized/malformed data, mismatched provenance and unknown configuration keys before starting a pipeline or installing configuration. +- Before configuration installation, the worker shall make only the bootstrap manifest/payload reads and shall log diagnostics to stdout. +- Signed URLs shall not appear in ordinary logs, agent-readable task rows or repository subprocess environments. Downloads shall use the exact regional S3 HTTPS object without redirects/proxies and with bounded response sizes. +- Producers, images and IAM shall be upgraded together; incompatible workers must be drained before switching transport. - | Role | Who can pass it, and how that grant is scoped | - |---|---| - | **Execution role** | The **orchestrator Lambda only**, at `RunMicrovm` — `iam:PassRole` scoped to this role's **exact ARN**, no condition (`constructs/task-orchestrator.ts`, sid `MicrovmPassExecutionRole`). The referenced comment contains the authoritative two-arm experiment evidence that the condition is the true blocker (not a permissions gap or stale bootstrap). | - | **Build role** | The **CloudFormation deployment role**, at `CreateMicrovmImage` (the L1's `buildRoleArn`) — via the new `MicrovmPassRoles` statement, scoped to `role/backgroundagent-dev-LambdaMicrovmComputeBuild*`, no condition. Also whoever runs `package-microvm-artifact.sh --create-image` out of band, using their own credentials. | - | **Connector operator role** | The **CloudFormation deployment role**, at `AWS::Lambda::NetworkConnector` create/update (`operatorRole`) — the same statement, scoped to `role/backgroundagent-dev-LambdaMicrovmComputeConnector*`. | +The [payload contract](/sample-autonomous-cloud-coding-agents/verification/645-payload-bootstrap) and [live checks](/sample-autonomous-cloud-coding-agents/verification/readme) record the implementation and validation. This transport does not establish complete hostile-worker isolation: other platform grants remain, and a stolen signed URL is usable until expiry or revocation. - The rest of the posture: every resource these roles can reach is account-scoped by ARN **except two deliberate `Resource: '*'` statements** — `ec2:DescribeAvailabilityZones` on the execution role (EC2 describe actions have no resource-level scoping; read-only, no mutation, no data access, needed so a CDK target repo's `cdk synth` build gate can resolve AZ context on a fresh clone) and the connector operator role's ENI/tag/private-IP statement (`CreateNetworkInterface` is authorized before the ENI exists and the `Describe*` calls take no resource, which is why the AWS-managed VPC-access policy uses `*` too). Both are justified in the construct's cdk-nag `AwsSolutions-IAM5` suppressions, which is where a reviewer should check them rather than here. The Logs grants are prefix-scoped (`/aws/lambda-microvms/*` plus one named log group), i.e. wildcards inside a namespace, not `*`. Separately, the **orchestrator's** `lambda:PassNetworkConnector` is also `Resource: '*'` and unavoidably so — the AWS-managed connectors live in the `aws` account, outside any ARN we could enumerate (justified in `task-orchestrator.ts`, sid `MicrovmPassNetworkConnector`). Finally: none of the three roles holds `iam:*`, none has cross-account trust, and the only `sts:AssumeRole` any of them has is the execution role's, scoped to the per-task SessionRole. +No ABCA endpoint consumer exists in P1–P3. The platform grants no `CreateMicrovmAuthToken` permission and mints no JWE tokens. `NO_INGRESS` can still return an endpoint URL; an unauthenticated 403 verifies the authentication boundary, not valid-token reachability. - If AWS later populates a source key on this path, adding it to the shared principal fixes all three roles and both `sts` actions at once. +### 4. Infra and IAM: conditional resources behind bootstrap `ComputeTypes` -- **Build-time egress needs port 80; runtime does not.** `agent/Dockerfile` installs Debian packages and `apt-get` fetches over plain HTTP, so a 443-only egress path fails every snapshot build (`Could not connect to deb.debian.org:80 … exit code: 100`). Rather than widen the runtime posture, the construct provisions a **second, build-only** connector on the same private subnets with a 443 + 80 security group, referenced solely by the image resource and the packaging script. The agent at run time still has 443-only egress. +The backend adds build/runtime VPC connectors, build artifacts, launch payloads, logs, roles and a managed or external image. Its bootstrap policy is conditional on `ComputeTypes` including `lambda-microvm`. A VPC egress connector requires an operator role. Build egress permits ports 80/443 for package installation; runtime egress permits 443 through the platform VPC. -- Where the bootstrap `ComputeTypes` parameter includes `lambda-microvm`, the generated template shall attach the `IaCRole-ABCA-Compute-LambdaMicrovms` policy to the CloudFormation execution role. -- The orchestrator role shall receive only the MicroVM lifecycle actions it calls (`lambda:RunMicrovm`, `lambda:SuspendMicrovm`, `lambda:ResumeMicrovm`, `lambda:TerminateMicrovm`, `lambda:GetMicrovm` for `pollSession`, and `lambda:PassNetworkConnector`, which is required even for the default connectors), scoped to platform-created images. -- Where the `lambda-microvm` backend is enabled, the approve and deny Lambdas shall receive `lambda:ResumeMicrovm` and `lambda:GetMicrovm` — conditionally, mirroring the cancel handler's conditional `RUNTIME_ARN` wiring in `task-api.ts`. -- The trust policy of every MicroVM-facing role shall name `lambda.amazonaws.com` and shall carry no source-condition key (the service presents none; see the trust-policy fact above). -- The `iam:PassRole` grant the orchestrator uses for the MicroVM execution role shall carry no `iam:PassedToService` condition and shall be scoped to that role's exact ARN. -- Where the bootstrap `ComputeTypes` parameter includes `lambda-microvm`, the `IaCRole-ABCA-Compute-LambdaMicrovms` policy shall grant `iam:PassRole` without an `iam:PassedToService` condition, scoped to the MicroVM build- and connector-operator role name prefixes, and shall not extend that grant to the MicroVM execution role. -- The shared `IaCRole-ABCA-Infrastructure` `iam:PassRole` statement shall retain its `iam:PassedToService` allowlist. -- The MicroVM execution role shall hold `logs:CreateLogStream` and `logs:PutLogEvents` on the application log group whose name is delivered in `platform_config`, scoped to that log group. +**Trust and PassRole limitation.** Recorded live checks rejected `aws:SourceAccount`/`aws:SourceArn` conditions on the MicroVM-facing roles and `iam:PassedToService` on the MicroVM PassRole paths. The working roles trust `lambda.amazonaws.com` without those conditions; build/execution roles also allow `sts:TagSession`. IAM simulation with caller-supplied condition values did not prove that the service supplied those values. Reintroduce a condition only after verifying service support. -`lambda:CreateMicrovmAuthToken` is granted to no role in P1–P3 (no JWE consumer exists; see sub-decision 3). +| Role/action | Scope and responsibility | +|---|---| +| Coordinator lifecycle APIs | Configured image ARN and its version-qualified sibling; includes actual-version capability lookup | +| Coordinator `iam:PassRole` | Exact execution-role ARN, without `iam:PassedToService` | +| Deployment `iam:PassRole` | Backend-specific build/operator role name patterns; shared infrastructure allowlist remains intact | +| Approval/denial handlers | Observe/resume the configured image after committing a decision; dispatch parked continuations | +| Build role | Selected immutable artifact plus manual-build key; MicroVM log writes | +| Execution role | Bootstrap manifests, startup secrets, allowlisted models, Memory and logs; tenant data through the per-task SessionRole | +| Connector operator | Tested ENI/tag/private-IP permissions plus AWSLambdaVPCAccessExecutionRole | -**Cost attribution.** `cdk/src/main.ts` currently tags the whole stack with a single `compute_type` context value (default `agentcore`) — already imprecise with two backends, wrong with three. P1 must add backend-identifying cost-allocation tags on the MicroVM-specific resources (images, payload/artifact bucket wiring, log groups) and revisit the stack-level tag semantics (e.g. a `compute_types` list), keeping attribution consistent with [#645](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/645)'s cost/attribution acceptance criterion. +`lambda:PassNetworkConnector` has no resource-level authorization support and therefore uses `Resource: *`. The operator role also has wildcard ENI permissions; some mutations can be scoped in IAM, so their current tested wildcard is not evidence that narrower permissions are impossible. DescribeAvailabilityZones needs a wildcard for fresh CDK repository lookups. These exceptions and namespace wildcards are documented in the construct’s cdk-nag suppressions. -- Where a deployment enables the `lambda-microvm` backend, MicroVM-specific resources shall carry backend-identifying cost-allocation tags. +**Nested infrastructure.** `microvm_nested_stack` must be explicitly selected until flat-to-nested migration is verified. New and already-nested deployments should use `true`, putting MicroVM resources in a nested stack. Omission fails synthesis. The shared execution role stays in the parent to avoid a role-trust dependency cycle. Bootstrap 1.9.0 covers nested deployment roles. Before upgrading an existing flat installation, set and retain `microvm_nested_stack=false` until completing the reviewed overlap/drain migration, explicit image/coordinator pins and rollback checks described in the [migration prerequisites](/sample-autonomous-cloud-coding-agents/verification/645-p3-nested-stack). Reusable migration commands are still being completed. Moving construct paths alone is not a safe migration. Preserve `microvm_resource_name_prefix` after a migrated deployment. #### Security bar vs existing backends ([#645](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/645) acceptance criterion) -| Control | AgentCore | ECS | Lambda MicroVMs | Delta | -|---|---|---|---|---| -| Egress, runtime (DNS Firewall, TCP 443 SG, flow logs) | Platform VPC | Platform VPC | Platform VPC via egress network connector | None | -| Egress, image build | ECR build outside the platform VPC | ECR build outside the platform VPC | Platform VPC via a **separate build-only connector, TCP 443 + 80** (`apt-get` is plain HTTP) | New surface: build-time egress is wider than runtime egress by one port, on a connector no running MicroVM can use | -| Tenant-data scoping | Per-session role (`admitComputeRole`) | Per-session role | Per-session role, execution role admitted identically | None | -| Secrets delivery | Runtime env + Identity injection | Task env vars | Fetched at `/run`; never in snapshot | New surface: snapshot must stay secret-free (EARS req., sub-decision 3) | -| Non-secret platform config (table/bucket names, secret + role ARNs) | Runtime env vars | Task env vars | `platform_config` in the `/run` payload, installed into the process env | New surface: the values are attacker-relevant *as env vars* (`LD_PRELOAD`, `AWS_ENDPOINT_URL`), so the agent installs a fixed **allowlist** and rejects the whole run on any other key (EARS req., sub-decision 3) | -| Inbound exposure | None (SigV4 invoke only) | None (no endpoint) | **None — but only because the strategy passes `NO_INGRESS` explicitly.** The service default is a PUBLIC `HTTP_INGRESS` connector plus a public `*.lambda-microvm..on.aws` endpoint; no tokens are minted in P1–P3 either way | New surface **and** a new failure mode: "no inbound" is an active control, not an absence. Drop the `NO_INGRESS` argument and every agent MicroVM gets a public endpoint (EARS req., sub-decision 3) | -| IAM condition keys on the compute-role trust **and** on the `iam:PassRole` grants that hand it over | Trust pinned with `aws:SourceAccount`; `PassRole` under the allowlisted bootstrap statement | Trust pinned per-service; `PassRole` under the allowlisted bootstrap statement | **Neither is possible.** All three MicroVM-facing roles trust the bare `lambda.amazonaws.com` with no `aws:SourceAccount`/`aws:SourceArn`, **and** both `iam:PassRole` grants (orchestrator → execution role at `RunMicrovm`; CloudFormation → build role at `CreateMicrovmImage`) carry no `iam:PassedToService` — the service presents no usable value for any of those keys, and each condition is a hard blocker while present (live-verified, blocking, four times across two runs) | **Real, evidenced gap that does not close from our side, and it is wider than the trust policy alone.** `lambda.amazonaws.com` is shared with every other Lambda feature, so neither the account pin nor the passed-to-service pin is available on this path. Compensated per role (table in sub-decision 4): the **execution** role is passable by the **orchestrator only** (at `RunMicrovm`), restricted to its **exact ARN**; the **build** and **connector-operator** roles are passable by the CloudFormation deployment role under a new **conditional, per-backend, name-prefix-scoped** statement (`MicrovmPassRoles`, bootstrap ≥ 1.6.0) that deliberately excludes the execution role. The shared allowlisted `IAMPassRole` (`role/backgroundagent-dev-*`) is left intact to avoid widening the grant for ~30 other roles, so while it technically matches the execution role, only the orchestrator actively reaches for it. Resources are account-scoped by ARN apart from two justified `Resource: \'*\'` statements (`ec2:DescribeAvailabilityZones`; the operator role\'s ENI management — both carry cdk-nag IAM5 suppressions). No `iam:*`, no cross-account trust. Revisit if AWS ever documents the values the service presents; CloudTrail records no `lambda-microvms` events, so they cannot be read from logs | -| Per-task observability writes | Runtime writes to the vended APPLICATION_LOGS group | Task role writes to the task log group | Execution role writes to the SAME APPLICATION_LOGS group, granted against the group `platform_config` names (P2-F4) | None — but only after P2-F4: the name was delivered a phase before the grant, so the agent attempted the write and every per-task line (and `METRICS_REPORT`) was `AccessDenied`, degrading silently to guest stdout | -| Session isolation | MicroVM | Task-level | MicroVM (Firecracker) | None (≥ ECS) | -| State reuse | None | None | Snapshot shared across MicroVMs | New surface: CSPRNG reseed + credential refresh on `/run`/`/resume` — **P3 scope** (P2 exposure measured as negligible: sole `random` consumer is a ULID sort key under a `task_id` partition; `os.urandom`/`secrets` unaffected; no credential derives from `random`). Credential refresh IS in P2: per-task credentials arrive via `platform_config` at `/run` | -| Workload-token injection | Yes (Runtime-coupled) | No (env-var posture) | No (env-var posture) | Shared with ECS; deferred to [#249](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/249)/ADR-016 | -| Operator shell access | No | No | Not enabled (`SHELL_INGRESS` omitted; candidate for [#391](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/391)) | None by default | -| Auth-token minting | n/a | n/a | `CreateMicrovmAuthToken` granted to no role in any phase | Verified static-only: any principal holding the action *can* mint a working JWE, including against a `SUSPENDED` MicroVM, so the posture rests entirely on the grant being absent | +| Control | MicroVM posture | Difference to account for | +|---|---|---| +| Runtime egress | Platform VPC, DNS Firewall, HTTPS security group, flow logs | Separate build connector additionally allows HTTP | +| Tenant access | Task-scoped SessionRole | Shared compute-role permissions remain outside that boundary | +| Configuration/secrets | Authenticated runtime identifiers; credentials resolved after launch | Shared snapshot must not capture credentials or task identity | +| Ingress | Explicit NO_INGRESS; no platform token minting | Service defaults would select HTTP_INGRESS if the field were omitted | +| Trust/PassRole | Exact role/image scopes where supported | Source-condition limitations described above | +| Logs | Service image group plus platform APPLICATION_LOGS | Both namespaces need explicit grants | +| Retained work | Versioned, bounded, checksum-verified checkpoint | Protect stored conversation/files and clean terminal state | +| Workload identity | Runtime credentials and task-role refresh | Linear vault identifiers arrive through `platform_config`; the compute execution role mints tokens ([setup](/sample-autonomous-cloud-coding-agents/using/linear-setup-guide#using-the-vault-with-lambda-microvms)) | + +MicroVM-specific resources carry `abca:compute-backend=lambda-microvm` cost tags. A stack-wide compute tag cannot accurately attribute a mixed-backend deployment by itself. #### Regional availability enforcement -Lambda MicroVMs launched in 5 regions (us-east-1/2, us-west-2, eu-west-1, ap-northeast-1) and will expand. ABCA is a single-region deployment, so the constraint is binary per stack: either the stack region supports the backend or the backend does not exist there. Enforcement is layered — one static check where offline determinism is required, live probes everywhere else so the platform self-heals as AWS adds regions (`list-managed-microvm-images` is the documented read-only availability probe): +The launch-region list is us-east-1, us-east-2, us-west-2, eu-west-1 and ap-northeast-1. New Regions can be supported before the static list is updated: -| Stage | Mechanism | Check | -|---|---|---| -| CDK synth/deploy | Static region constant (single exported list, documented update path) | Synth fails when `ComputeTypes` includes `lambda-microvm` in an unlisted region; context-flag escape hatch for newly launched regions ahead of the constant update | -| Repo onboarding | Live probe from the CLI | `bgagent repo onboard --compute-type lambda-microvm` calls `list-managed-microvm-images` in the stack region and rejects with a remedy (supported-region list + suggest `agentcore`/`ecs`) | -| `bgagent platform doctor` | Live probe (precedent: `checkBedrockModel`) | Reports backend availability for the stack region whenever any active blueprint selects `lambda-microvm` | -| Orchestration (defense in depth) | Error classification | `startSession` failures from a missing regional endpoint classify to a typed remedy in `error-classifier.ts`, never a cryptic SDK error on the task | +- Synth rejects a concrete unlisted Region unless `microvm_region_override` is set. An unresolved Region defers to live checks. +- CLI onboarding and platform doctor probe `list-managed-microvm-images`. +- Runtime regional failures receive a configuration remedy instead of an opaque SDK error. -- If an operator onboards a repo with `compute_type: 'lambda-microvm'` and the availability probe fails for the stack region, then the CLI shall reject the onboarding with the supported-region list and alternative backends as the remedy. -- If `startSession` fails because the MicroVM service is unavailable in the stack region, then the orchestrator shall classify the failure with a configuration remedy and shall not retry. -- When the platform doctor runs in a deployment where any active blueprint selects `lambda-microvm`, the doctor shall probe MicroVM availability in the stack region and report the result. +### 5. Rollout: phased, AgentCore remains the default backend -### 5. Rollout: phased, default unchanged +| Phase | Delivered behavior | +|---|---| +| P1 | Strategy, infrastructure, bootstrap/types, minimal `/ready` + `/run` serving; merged in [#689](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/689) | +| P2 | Clone → change → PR, progress/logs/Memory, runtime configuration, `/validate` + `/terminate`; merged in [#733](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/733), with takeover follow-up validation | +| P3 | Approval-aware sleep/wake, credential renewal, original deadlines, retained requests, complete checkpoint/replacement recovery, nested deployment and live acceptance | + +Activate sleep only after verifying the deployed image and coordinator together. Keep a compatible published coordinator and explicit image pin for rollback. Normal acceptance includes the 600-second default, explicit expiry, new and existing task off-switch behavior, replacement, cleanup and preservation of unrelated infrastructure. -- **P1 — strategy + infra + minimal hook serving:** `LambdaMicrovmComputeStrategy` (start/poll/stop), CDK construct, bootstrap policy, types sync, unit + CDK assertion tests, and the agent's `/ready` + `/run` endpoints. No suspend yet. The image IS creatable and launchable and the payload DOES reach the agent — but there is **no smoke-parity guarantee** (sub-decision 3's phasing table). -- **P2 — smoke parity:** the agent serves `/terminate` + `/validate` and installs its platform env from the `/run` payload (see sub-decision 3's "Platform configuration delivery"); agent completes clone → change → PR on the backend with progress visible to `bgagent watch`; failure classification entries in `error-classifier.ts`; **AgentCore Memory parity** (IAM grant + `MEMORY_ID` delivery, following the `EcsAgentCluster` pattern — Memory is a standalone service already consumed cross-substrate, and omitting the grant silently no-ops cross-session learning); the agent's remaining non-secret env parity inside the snapshot. -- **P3 — suspend/resume:** the interface widening from sub-decision 1 (mandatory methods, all three strategies in one commit), HITL-wait suspend policy, inline resume in the approve/deny Lambda with orchestrator-poll reconciliation (sub-decision 2), timeout-under-freeze wall-clock handling; coordinate with [#491](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/491)'s unified liveness model and update Cedar decision #7's rationale note. -- **Out of scope:** replacing AgentCore as default; classic Lambda functions as a runtime; GPU; the Runtime-coupled workload-access-token injection path (delivery mechanism exists only on AgentCore Runtime; MicroVMs adopt the ECS env-var posture until [#249](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/249)/ADR-016 redesign the seam). Gateway integration is orthogonal: ADR-019/[#641](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/641) is substrate-portable by design and applies to this backend when it lands. +Approval responses are supported through the authenticated CLI and owner-authored `approve`/`deny` replies to Linear approval comments. Changing the default backend, GPU support, native Slack approval buttons and operator shell access remain outside this ADR. Future work needs its own scope; “P4” is not an approved phase here. ## Consequences -- (+) **Suspend/resume economics.** Tasks idling on approval waits stop billing compute while preserving full state — bounded at ~1 h per gate under the current Cedar ceiling (decision #6), and the enabler for cheap off-hours gate-ceiling extensions later (§14.8). Verified end to end at the substrate level: suspend and resume each complete in ~1 s, and `microvmId` **and** `endpoint` survive the cycle byte-identical, so a stored `SessionHandle` remains valid. -- (+) **VM-level isolation without cluster ops.** Firecracker isolation with no ECS cluster, task definition, or capacity management; one-session-per-MicroVM maps 1:1 onto ABCA's task model. -- (+) **Escapes AgentCore's 2 GB image limit and FUSE `flock()` workaround** — native disk in the snapshot supports `uv`/`mise` without the split-storage scheme. Note the size comparison must say WHICH measure it means: the same agent tree is 1.799 GB as an OCI image (629.7 MB compressed, i.e. under AgentCore's limit) but reports `codeInstallSizeInBytes` of 2.17 GiB as a MicroVM snapshot (i.e. over it). The two straddle the limit and are not interchangeable; memory/disk snapshot sizes are a third thing again and must not be summed into the comparison. -- (+) **Liveness becomes explicit.** Unlike AgentCore's stub `pollSession`, the strategy can report real substrate state, strengthening the [#491](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/491) unification. -- (−) **32 GiB sustained ceiling (8 GiB baseline + automatic 4× vertical scaling), 32 GB disk.** Not a successor to the ECS backend for heavy CI-parity builds; the platform now maintains three backends. -- (−) **Capacity is baseline-priced with burst headroom, which is a narrower promise than "32 GiB".** The deployment configures an 8 GiB / 4 vCPU baseline and the service scales to 32 GiB / 16 vCPU on demand — well matched to an agent task, which is idle-ish while waiting on the model and spiky during builds, and cheaper than reserving the peak. But it is *burst*, not a reservation: a workload that needs 32 GiB **sustained** is relying on scaling behaviour this ADR has not measured, and against ECS's 120 GB the gap for sustained-memory workloads is unchanged. So the value proposition remains the suspend economics, the observable control-plane state machine, and the absence of cluster ops — with capacity now a fair-to-good fit rather than a hard blocker. Repos with genuinely sustained heavy builds still belong on `ecs`. -- (−) **New packaging pipeline.** Zip + Dockerfile + service-side image builds with versioned snapshots (storage billed per version) alongside the existing ECR flow; image versions need lifecycle cleanup — including versions left behind by FAILED builds, and noting the last version of an image cannot be deleted individually (delete the image, which reaps it). -- (−) **The payload bucket is on the hot path, not the overflow path.** With a 4 KB `runHookPayload` cap, virtually every real task delivers its payload via S3, so the bucket, its TTL rule and the execution role's read grant are load-bearing for normal operation rather than an edge case (sub-decision 3). -- (−) **8-hour hard cap includes suspended time**, and with `idlePolicy` omitted there is no tighter substrate-level suspended-TTL — the suspended-state bound is `maximumDurationInSeconds` plus orchestrator termination and the stranded reconciler. A manually suspended VM was observed alive at 1 h with no TTL in sight (observation truncated there), so nothing contradicts this bound, but nothing narrows it either. Under today's 1 h gate ceiling it is comfortably sufficient; any future extension of gate ceilings must revisit the bound (an additive `idlePolicy` change) and give the orchestrator a checkpoint-and-restart path (push branch, new session) beyond the cap. -- (!) **Idle-policy foot-gun.** Traffic-based auto-suspend would freeze a busy outbound-only agent; the decision to disable auto-suspend must be enforced in code and covered by tests, not left to configuration discipline. -- (!) **Service defaults are not the desired posture.** Two live-caught cases (public `HTTP_INGRESS` by default; `/ready` mandatory) mean an omitted field on this backend does not mean "off" — it can mean "the service picks, and it picks wider than we want". Every new `RunMicrovm` / `CreateMicrovmImage` field should be assumed to have an opinionated default until checked. -- (!) **Nothing self-terminates on the paths that matter** — superseding P1's unqualified version of this bullet. With `run: ENABLED` the service DOES reap a VM whose run hook returns 4xx (~12 s, `stateReason: "Run lifecycle hook returned HTTP status 400."`, live-verified), so a guest that rejects its own payload cleans itself up. That is the only self-cleaning case: the service reaps a hook *result*, and once `/run` has answered 200 it has no view of the guest. A VM whose task finished, crashed after `/run`, or hung stays `RUNNING` and billing until the 8 h cap, so the orchestrator's `TerminateMicrovm` on finalize remains the only cleanup for normal operation and a leaked handle is still a cost incident. -- (!) **Snapshot uniqueness.** Shared memory snapshots mean every MicroVM restored from one image starts with identical PRNG state. **Re-scoped to P3 in the P2 review, with the exposure measured rather than assumed** (see the amendment under sub-decision 2): the sole `random` consumer in `agent/src` is a ULID sort key under a `task_id` partition, `os.urandom`/`secrets` are unaffected, and no credential or token derives from `random` — so the P2 exposure is a possible duplicate progress event, not a security boundary. It becomes load-bearing at P3, where a *resumed* VM continues from frozen state repeatedly; the reseed must land on `/run` and `/resume` together with those hooks. Asserting the requirement while leaving it unimplemented was the real defect, and this amendment is the fix. -- (−) **The agent stack template is at 98.6 % of CloudFormation's 1 MB limit** (985,886 bytes) and 486 of 500 resources with a MicroVM image configured — ~14 KB of headroom, i.e. roughly one more construct, and down from 98.4 % / ~16 KB one run earlier. Not caused by this backend (the MicroVM construct is ~6 KB of it) but reached by it, and it will block deploys for reasons that have nothing to do with MicroVMs. Tracked in [#735](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/735); the candidate remedies are `suppressTemplateIndentation` and a stack split. -- (!) **Snapshot WARMTH is a first-class property, not an optimisation.** A snapshot inherits only the pages something touched before it was captured, so a large lazily-loaded artifact — the 225 MiB `claude` binary, and anything similar added later — pays its first-touch cost on the *first task* instead of at container start. That cost failed every task at turn 0 in the P2 smoke run (P2-F5). Anything heavyweight added to the image must be exec'd in `/ready`, and any timeout guarding a first touch must be sized for a cold page fault rather than for the work itself. -- (!) **Regional availability (5 regions at launch, expanding)** — enforced in layers (synth-time static check, onboarding + doctor live probes, orchestration-time classification; see sub-decision 4). The static CDK constant is the one piece that rots as AWS expands; its update path and context-flag escape hatch are deliberate. -- (!) **Workload-token injection delta persists** (shared with the ECS backend) until [#249](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/249)/ADR-016 land; document it in the security bar comparison rather than blocking on it. Memory and Gateway are explicitly *not* deltas — both are standalone services consumed via IAM from any substrate. +- Approval waits can stop consuming compute without discarding the question or the saved work. Snapshot and checkpoint costs still apply. +- Native disk supports build-tool locking, and the service exposes explicit worker state without a cluster to operate. +- ABCA now maintains three backends and an additional artifact/snapshot lifecycle. Failed or unused image versions also need cleanup. +- Eight hours remains a per-worker limit. Retained approvals rely on verified retirement and replacement rather than extending a worker indefinitely. +- Suspended workers retain ABCA capacity until confirmed retirement. The service’s account-quota treatment of suspended memory was not established by the recorded probes. +- Memory baseline validation and published peak capacity are not workload benchmarks. Sustained heavy builds require measured sizing and may fit ECS better. +- A healthy heartbeat does not prove coding progress. Hook failures may cause service termination, but successful hook acceptance does not remove ABCA’s cleanup responsibility. +- Service error wording can hide transport errors. The pooled-hook mitigation has local and live evidence; historical internal dispatch traces remain unavailable. ## Testing -**P1 (start/poll/stop — no suspend):** - -- Unit tests for the strategy: start/poll/stop mapping (including `SessionStatus` `'suspended'` reported mechanically, without task-state interpretation), payload-size branching (inline vs S3 pointer, at the exact 4 096/4 097-byte boundary), the image-identifier-must-be-an-ARN guard, the explicit `NO_INGRESS` argument (including the blank-env-var fallback, which must never omit the field), error classification (`ServiceQuotaExceededException`, `ThrottlingException`, `ResourceNotFoundException`, regional-unavailability), the omit-`idlePolicy` invariant, and `maximumDurationInSeconds` fixed at 28 800. -- Agent tests: `/ready` returns 200 once the server is up and starts nothing; `/run` accepts both envelope shapes (inline and S3 pointer), starts the pipeline asynchronously through the same mapper `/invocations` uses, returns before the pipeline finishes, and rejects every unusable envelope with a named code before spawning; `/suspend` and `/resume` are NOT served (`/validate` and `/terminate` joined the served set in P2). -- Orchestrator tests: substrate-terminal + non-terminal task status → failed classification; `suspended` + non-`AWAITING_APPROVAL` status → anomaly event, no fail-fast; `compute_metadata` persisted with `microvmId`/`endpoint` after `startSession`. -- CDK assertions: MicroVM resources present only when `ComputeTypes` includes the backend; synth failure for unsupported regions (plus the context-flag escape hatch); memory size validated against the accepted list at synth; the connector operator role and its trust; two connectors with the build-only one carrying port 80 and the runtime one not; the `/ready` + `/run` hook declaration and the absence of the others; IAM actions scoped as specified (orchestrator lifecycle set; no `CreateMicrovmAuthToken` anywhere); payload-bucket grants (execution role read-only); backend cost-allocation tags; types-sync check covers the widened `ComputeType`. -- CLI tests: onboarding rejection with remedy when the availability probe fails; doctor check present when a blueprint selects the backend. -- P1 verification items (external service facts) — **executed 2026-07-31, us-east-1**; see `docs/verification/645-p1-lambda-microvm-runbook.md` for the full evidence. Discharged: `runHookPayload` limit (**4 096**, not 16 KB), the accepted baseline memory sizes (`[512…8192]` MiB — note the developer guide, not the probe, is what establishes that this is a BASELINE with a 32 GiB peak), image-identifier ARN requirement, IAM action names and the observed image-ARN shape, region probe behaviour, manual suspend/resume without `idlePolicy`, terminate timing and the `TERMINATED`-persists-≥10-min finding, the default public `HTTP_INGRESS`, and the `/ready` requirement. **Not** discharged: account-quota treatment of `SUSPENDED` MicroVMs (not observable safely), suspended TTL beyond 1 h (truncated), the vertical-scaling behaviour itself (no workload here approached the baseline, so the 4× peak is documented rather than observed), and the `AWS::Lambda::MicrovmImage` CloudFormation value shapes (never exercised — the run used the out-of-band script path; **discharged, and REFUTED, by the P2 run — see P2-F2 in sub-decision 3**). Record the closed answers in COMPUTE.md. - -**P2 (smoke parity):** - -- Agent tests: `/validate` returns 200 with its individual check results, 503 while initialising, reports a missing hook route / unsupported interpreter, starts nothing, and makes **zero AWS calls even with `LOG_GROUP_NAME` set** (asserted by poisoning the boto3 and CloudWatch-writer seams — the same assertion covers `/ready`); `/terminate` returns 200 with no body at all, with a malformed / non-object / wrong-content-type / whitespace-only body, when the body read itself fails, with a pipeline still running (without joining it), and when its own best-effort step raises — and never calls `task_state.write_terminal`; a structural assertion that the route carries no typed body param keeps the 422 from being reintroduced. -- `platform_config` tests: the allowlist and required subset are read from `contracts/constants.json` (the wire key set is additionally asserted literally, as the agent-side tripwire on a contract edit); an unknown key rejects the whole block with nothing installed; a non-object block and a non-string value are rejected; a blank/`null` optional value is skipped without clobbering an image value while a blank required value is rejected; a payload value beats a pre-existing env value; installation is observed to happen before the GitHub-token resolver runs and before any pipeline thread exists; the config is picked up from the inline envelope, from beside the S3 pointer, and from inside the fetched object (inner wins); an envelope with no `platform_config` is still accepted with a warning. -- Snapshot credential hygiene: a subprocess probe asserts that importing `server` and serving `/ready` + `/validate` imports neither `boto3` nor `botocore`, caches no `aws_session` session, and spawns no CloudWatch writer thread — the property that keeps a build-role credential chain and the build-time region out of the snapshot. -- `/run` pre-install silence: with a **baked `LOG_GROUP_NAME`** (the hostile case — without it the assertions pass vacuously) every AWS/credential seam (`boto3.client`/`Session`, the `aws_session` factories, `_debug_cw`/`_warn_cw`) is armed to raise until the install succeeds. Asserted on the accepted path, on all three rejection paths (bad envelope, `platform_config` invalid, `platform_config` incomplete) and on the failed-fetch 500 — where the seams stay armed for the whole request, because a rejected run installed nothing and so earns no AWS call. The permitted exception is asserted POSITIVELY: exactly one client is built pre-install, for `s3`, through the attributed factory. -- `/ready` warm-up tests (P2-F5): the hook exec's each configured binary exactly once with a generous timeout; `claude` is the only REQUIRED entry; a timeout, a missing binary, a non-zero exit and an unexpected `OSError` each produce **503 with the reason logged to stdout** rather than a 200 or a 500; a best-effort failure still reports ready; the warm-up makes zero AWS calls with `LOG_GROUP_NAME` baked. Plus the backstop half: the `claude --version` probe's bound is asserted to be ≥ 60 s and to be applied to the *exec* rather than to the PATH lookup, and a missing CLI warns instead of raising. -- CDK assertions (P2-F1/F2/F4): no source-condition key on any of the three MicroVM-facing role trusts, and no `aws:SourceAccount`/`aws:SourceArn` string anywhere in them; hook properties are `ENABLED` and the architecture is `ARM_64`, with a negative assertion that **no** hook route string appears anywhere in the rendered image resource; the agent hook routes are asserted against their own dedicated constant (the template no longer carries a path to compare); the execution role holds `logs:CreateLogStream`/`PutLogEvents` on the application log group and the two logs grants stay separate; the stack wires the SAME log group it delivers as `platform_config.log_group_name`. -- Smoke (gated like the ECS backend): clone → change → PR with `bgagent watch` progress; Memory write parity (no AccessDenied no-op). **Run 1 (2026-08-06) FAILED at `implement`, turn 0 — no PR. Run 2 (2026-08-07) PASSED: two tasks clone → change → commit → push → PR, `COMPLETED`, 12 turns / $0.279 / 153 s** (`docs/verification/645-p2-smoke-runbook.md`), which also discharged P2-F1, P2-F2, P2-F4, P2-F5 and the dual-signal-liveness item (45 s heartbeat cadence observed across a 181 s `RUNNING` window). **The row is not yet fully closed:** run 2 needed one live IAM workaround, and establishing why produced P2r2-F10 (the identity-side `iam:PassedToService`) and P2r2-F9 (its CloudFormation twin). Both are fixed in source above and neither has been re-exercised live, so what remains is a re-run on a re-bootstrapped account with no workarounds. - -**P3 (suspend/resume):** +The [verification summary](/sample-autonomous-cloud-coding-agents/verification/readme) distinguishes recorded live acceptance from open PR checks. Required coverage includes: -- Unit tests: suspend/resume mapping; agentcore/ecs `unsupported` stubs; approve/deny inline resume with handle loaded from `compute_metadata`. -- HITL lifecycle tests: inline resume failure leaves the decision outcome intact and records the orphan event; orchestrator backstop retries resume; gate expiry fires at `min(monotonic budget, created_at + timeout_s)` — including the suspend/resume case where the monotonic budget exceeds the wall-clock remainder — without disturbing the §13.12 late-approval race protection. -- Smoke: suspend/resume across a simulated approval wait preserving workspace state. +- Strategy state mapping, explicit unsupported results, ARN/Region validation, NO_INGRESS, omitted idlePolicy and bounded uncertain-start recovery. +- Hook readiness/warm-up, AWS-silent build hooks, authenticated payload installation, arbitrary terminate bodies and lifecycle connection closure. +- Decision/timeout/cancellation races, image capability, coding/progress barriers, original deadlines, durable replay and exact-attempt capacity ownership. +- Real S3 version/integrity/access checks, real DynamoDB transactions and actual SDK conversation/Git/workspace recovery after process and disk loss. +- Cloud approve/deny/expiry/cancellation, repeated sleep/wake, AWS credential renewal after actual expiry, retirement/replacement and final resource cleanup. +- Nested fresh deployment and overlapping migration, compatible image/coordinator rollback, normal CLI feedback and live off-switch acceptance. -**All phases:** docs sync for [COMPUTE.md](/sample-autonomous-cloud-coding-agents/architecture/compute) (new column distinguishing MicroVMs from classic Lambda) and [ORCHESTRATOR.md](/sample-autonomous-cloud-coding-agents/architecture/orchestrator) (liveness + suspend lifecycle). +Live evidence must distinguish service acknowledgment from completed guest recovery and synthetic handler checks from actual external-channel submissions. ## References -- Issue [#645](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/645) — originating RFC proposal -- Issue [#491](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/491) — unified liveness decision model (soft dependency, P3) -- Issue [#641](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/641) / ADR-019 (PR [#663](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/663)) — substrate-portable tool plane -- PR [#596](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/596) — ECS Fargate backend (pattern source for conditional wiring) -- [AWS Lambda MicroVMs](https://docs.aws.amazon.com/lambda/latest/dg/lambda-microvms-guide.html) — developer guide; [Running and using MicroVMs](https://docs.aws.amazon.com/lambda/latest/dg/microvms-launching.html) — lifecycle APIs and hooks -- [Agent Toolkit for AWS — aws-lambda-microvms skill](https://github.com/aws/agent-toolkit-for-aws/blob/main/skills/specialized-skills/serverless-skills/aws-lambda-microvms/SKILL.md) — operational constraints (no self-suspend, idle-policy semantics, snapshot uniqueness, size limits) -- [ADR-020](/sample-autonomous-cloud-coding-agents/architecture/adr-020-ears-requirements-syntax) — EARS syntax used for the normative requirements above -- [CEDAR_HITL_GATES.md](/sample-autonomous-cloud-coding-agents/architecture/cedar-hitl-gates) — approval-gate mechanics (decisions #6, #7) the suspend/resume handshake preserves; `cancel-task.ts` / `task-api.ts` — the inline best-effort + reconciler-backstop pattern the resume path mirrors -- [COMPUTE.md](/sample-autonomous-cloud-coding-agents/architecture/compute), [ORCHESTRATOR.md](/sample-autonomous-cloud-coding-agents/architecture/orchestrator) — design docs to be updated by the implementing PRs +- [Issue #645](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/645) — originating proposal +- [Issue #491](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/491) — unified liveness model +- [Issue #641](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/641) — substrate-portable tool plane +- [AWS Lambda MicroVMs guide](https://docs.aws.amazon.com/lambda/latest/dg/lambda-microvms-guide.html) and [lifecycle APIs/hooks](https://docs.aws.amazon.com/lambda/latest/dg/microvms-launching.html) +- [ADR-020](/sample-autonomous-cloud-coding-agents/decisions/adr-020-ears-requirements-syntax) — requirement syntax +- [Compute](/sample-autonomous-cloud-coding-agents/architecture/compute), [orchestrator](/sample-autonomous-cloud-coding-agents/architecture/orchestrator), [approval gates](/sample-autonomous-cloud-coding-agents/architecture/cedar-hitl-gates) diff --git a/docs/src/content/docs/decisions/Adr-022-agent-asset-registry.md b/docs/src/content/docs/decisions/Adr-022-agent-asset-registry.md index cc8ac39b3..a8c5a9808 100644 --- a/docs/src/content/docs/decisions/Adr-022-agent-asset-registry.md +++ b/docs/src/content/docs/decisions/Adr-022-agent-asset-registry.md @@ -16,10 +16,10 @@ This has three costs: 1. **Every new tool/skill/policy costs a deploy.** Rolling out a new MCP server to N repos is N Blueprint edits + a CDK deploy. Teams can't publish autonomously. 2. **No pin, no reproducibility.** Because assets aren't versioned, "the tool the agent used on 2026-05-01" can't be reconstructed from the task record. -3. **The vocabulary already anticipates a registry.** [ADR-014](/sample-autonomous-cloud-coding-agents/architecture/adr-014-workflow-driven-tasks) and [WORKFLOWS.md](/sample-autonomous-cloud-coding-agents/architecture/workflows) already: - - Coined `registry://kind/name` refs (grammar in [`agent/src/workflow/validator.py`](../../agent/src/workflow/validator.py) `_REGISTRY_REF`). +3. **The vocabulary already anticipates a registry.** [ADR-014](/sample-autonomous-cloud-coding-agents/decisions/adr-014-workflow-driven-tasks) and [WORKFLOWS.md](/sample-autonomous-cloud-coding-agents/architecture/workflows) already: + - Coined `registry://kind/name` refs (grammar in [`agent/src/workflow/validator.py`](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/agent/src/workflow/validator.py) `_REGISTRY_REF`). - Modeled `agent_config` asset kinds (`mcp_servers`, `skills`, `plugins`, `subagents`, `prompt_fragments`, `cedar_policy_modules`) 1:1 with the vocabulary this ADR needs. - - Designed a resolver interface as a **drop-in swap** — filesystem-backed today, registry-backed later ([`agent/src/workflow/loader.py:107`](../../agent/src/workflow/loader.py), WORKFLOWS.md §"Registry integration (#246)"). + - Designed a resolver interface as a **drop-in swap** — filesystem-backed today, registry-backed later ([`agent/src/workflow/loader.py:107`](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/agent/src/workflow/loader.py), WORKFLOWS.md §"Registry integration (#246)"). - Left validator rule 8 as a deferred check: every asset ref resolves — builtins today, registry refs when the registry lands. Issue [#246](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/246) proposes closing this gap with a **central versioned asset registry**: a catalog of typed, immutable-at-version artifact records that blueprints pin by `registry://kind/name@constraint`, that the orchestrator resolves at task start, and that the agent receives as a resolved bundle. Six acceptance criteria — asset kinds enumerated, publish+resolve with semver+immutability, blueprint reference of at least one kind, agent E2E for one kind, descriptor validation at publish, tests + docs. @@ -32,7 +32,7 @@ Two forces shape the decision: Adjacent decisions the ADR must respect but not re-open: - **[#381](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/381) — ADR↔Persona↔Skill graph.** #381 wires bidirectional frontmatter edges between ADRs, personas, and skills as `docs/` and plugin markdown, enforced by a parity linter. Both #246 and #381 mention "skills," but they mean different things: #381 is *documentation graph consistency*; #246 is a *runtime artifact catalog*. This ADR keeps them cleanly separated. -- **[ADR-014](/sample-autonomous-cloud-coding-agents/architecture/adr-014-workflow-driven-tasks).** Workflows are the registry's first `capability`-kind consumer; the `workflow_ref` field's resolution semantics are the model this ADR generalizes to all asset kinds. Workflows themselves stay filesystem-backed in the MVP — they already ship and are validated; migrating them to the registry is a separate follow-up, not part of #246. +- **[ADR-014](/sample-autonomous-cloud-coding-agents/decisions/adr-014-workflow-driven-tasks).** Workflows are the registry's first `capability`-kind consumer; the `workflow_ref` field's resolution semantics are the model this ADR generalizes to all asset kinds. Workflows themselves stay filesystem-backed in the MVP — they already ship and are validated; migrating them to the registry is a separate follow-up, not part of #246. ## Decision @@ -44,7 +44,7 @@ The ADR fixes the **contract**. The substrate ranking below was written to defer ### Sub-decisions -1. **URI grammar and kinds.** `registry:////@`. MVP kinds: `mcp_server`, `cedar_policy_module`, `skill`. Schema declares — but does not yet load — `plugin`, `subagent`, `prompt_fragment`, `capability` (`capability` = workflow, ADR-014 vocabulary). This grammar *extends* the shape pre-declared for ADR-014 at [`agent/src/workflow/validator.py`](../../agent/src/workflow/validator.py) `_REGISTRY_REF`, which admitted a 2-segment `registry://kind/name` form but not the `@` suffix or `_` (snake_case) in the kind segment. The implementation widens that lenient acceptance check and lands the **authoritative, strict grammar** in a dedicated parser mirrored byte-for-byte across both languages (`cdk/src/handlers/shared/registry/ref.ts` and `agent/src/registry/ref.py`); `validator.py` remains a lenient pre-flight admitting both forms. +1. **URI grammar and kinds.** `registry:////@`. MVP kinds: `mcp_server`, `cedar_policy_module`, `skill`. Schema declares — but does not yet load — `plugin`, `subagent`, `prompt_fragment`, `capability` (`capability` = workflow, ADR-014 vocabulary). This grammar *extends* the shape pre-declared for ADR-014 at [`agent/src/workflow/validator.py`](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/agent/src/workflow/validator.py) `_REGISTRY_REF`, which admitted a 2-segment `registry://kind/name` form but not the `@` suffix or `_` (snake_case) in the kind segment. The implementation widens that lenient acceptance check and lands the **authoritative, strict grammar** in a dedicated parser mirrored byte-for-byte across both languages (`cdk/src/handlers/shared/registry/ref.ts` and `agent/src/registry/ref.py`); `validator.py` remains a lenient pre-flight admitting both forms. **Short vs long forms (migration note).** [WORKFLOWS.md](/sample-autonomous-cloud-coding-agents/architecture/workflows) illustrates refs in a *short* 2-segment form with no constraint (`registry://prompt/web-research-workflow`, `registry://mcp/web-search-v1`, `registry://skill/research-synthesis-v1`). Those are **forward-declarations** — the WORKFLOWS spec explicitly marks `registry://` refs as "declared in the schema now but ignored by the runner until the registry (#246) can resolve them." This ADR's strict grammar is the *long* form: 3 segments (`//`), snake_case kinds (`mcp_server`, not `mcp`; `prompt_fragment`, not `prompt`; `cedar_policy_module`, not `cedar`), and a **mandatory** `@`. Only the long form resolves. The short form stays **lenient-only** — accepted syntactically by `validator.py`'s pre-flight so existing illustrative workflows don't fail validation, but *not resolvable* by the registry until rewritten to the long form. There is no automatic aliasing (`mcp` → `mcp_server`) at resolve time: a ref must be long-form to load. Migrating the WORKFLOWS examples to long form is doc-only cleanup tracked with the workflow/registry integration, not a blocker for #246. @@ -180,12 +180,12 @@ Regardless of substrate, the invariants above (semver, immutability, resolve-at- - Issue [#481](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/481) — capability descriptors. - Issue [#230](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/230) — event-rule packs (defer to registry Phase 3). - Issue [#99](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/99) — ToolBuilderAgent / meta-agent vision (out of scope; registry is a prerequisite). -- [ADR-003](/sample-autonomous-cloud-coding-agents/architecture/adr-003-contribution-governance) — contribution governance (publish/promote follows the same approval path). -- [ADR-014](/sample-autonomous-cloud-coding-agents/architecture/adr-014-workflow-driven-tasks) — workflow-driven tasks (defines the `registry://` grammar, resolver interface, and asset-kind vocabulary this ADR generalizes). +- [ADR-003](/sample-autonomous-cloud-coding-agents/decisions/adr-003-contribution-governance) — contribution governance (publish/promote follows the same approval path). +- [ADR-014](/sample-autonomous-cloud-coding-agents/decisions/adr-014-workflow-driven-tasks) — workflow-driven tasks (defines the `registry://` grammar, resolver interface, and asset-kind vocabulary this ADR generalizes). - [docs/design/WORKFLOWS.md](/sample-autonomous-cloud-coding-agents/architecture/workflows) §"Registry integration (#246)" — the workflow-side spec this ADR closes. -- [`agent/src/workflow/validator.py`](../../agent/src/workflow/validator.py) `_REGISTRY_REF` — the lenient pre-flight ref check (pre-dates this ADR; widened here to admit the strict form). +- [`agent/src/workflow/validator.py`](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/agent/src/workflow/validator.py) `_REGISTRY_REF` — the lenient pre-flight ref check (pre-dates this ADR; widened here to admit the strict form). - `cdk/src/handlers/shared/registry/ref.ts` / `agent/src/registry/ref.py` — the authoritative strict `registry://` grammar (two-language, parity-tested). -- [`agent/src/workflow/loader.py:107`](../../agent/src/workflow/loader.py) — Phase 4 deferral comment this ADR unblocks. +- [`agent/src/workflow/loader.py:107`](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/agent/src/workflow/loader.py) — Phase 4 deferral comment this ADR unblocks. - [AWS Agent Registry documentation](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/registry.html) — preferred substrate candidate. - [AWS Agent Registry key capabilities](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/registry-key-capabilities.html) — governance lifecycle, hybrid search, EventBridge integration. - [AWS Agent Registry: Migration from public preview](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/registry-faq.html) — the 2026-08-06 namespace migration referenced in risks. diff --git a/docs/src/content/docs/decisions/Adr-023-trusted-approval-writer.md b/docs/src/content/docs/decisions/Adr-023-trusted-approval-writer.md new file mode 100644 index 000000000..70fc974d7 --- /dev/null +++ b/docs/src/content/docs/decisions/Adr-023-trusted-approval-writer.md @@ -0,0 +1,72 @@ +--- +title: Adr 023 trusted approval writer +--- + +# ADR-023: Trusted approval writer and retained human decisions + +**Status:** proposed +**Date:** 2026-09-21 + +## Context + +Workers previously had direct write access to approval records. Restricting a +worker to its own task did not prevent it from replacing a pending request with +an approved record. MicroVM continuation makes the distinction between worker +authority and human consent especially important, but the same defect affects +ECS and AgentCore. + +A short approval deadline also couples human response time to compute lifetime. +Someone taking time to consider a request should not lose it merely because the +worker should stop consuming resources. + +## Decision + +Use a small IAM-authenticated API and Lambda as the worker-facing approval +writer. Its interface creates a pending request or closes one after a worker +timeout/failure. It cannot record human approval or denial, notification markers, +or retention TTLs. Human decisions remain in the owner-authenticated decision +handlers. Workers retain approval reads and transaction condition checks. + +The session role can invoke only the task path named by its `task_id` tag. +Creation atomically writes the request and transitions its task from running to +awaiting approval. The transaction checks the task owner/state and, for MicroVM, +the coordinator-owned worker lease. A stale worker cannot use its old lease to +create or close a request. + +Unanswered requests have no decision deadline by default. Explicit positive +timeouts remain available. Task cancellation, terminal failure and invalidated +worker execution can close requests independently. Compute lifetime is separate: +MicroVM can checkpoint and release a worker while retaining a request; ECS and +AgentCore currently cannot restore that waiting execution into a replacement. +Their task execution limits still apply. + +This implementation is included in the P3 review because retaining approvals +without protecting their decision records would preserve the security defect. +The ADR remains proposed for maintainer review. + +## Consequences + +- Existing deployments must pause submissions, drain old workers and deploy + matching infrastructure and images. Old images still attempt direct writes, + which the new IAM policy denies. See the + [upgrade procedure](/sample-autonomous-cloud-coding-agents/getting-started/deployment-guide#upgrading-approval-permissions). +- The service adds one signed request per creation/closure and becomes an + availability dependency. Failure denies permission to proceed; it never + becomes human consent. +- The service validates record shape, not the truth of worker-supplied policy + descriptions. A preview can be truncated and is not a proof of the full tool + input. `TIMED_OUT` currently also represents a worker polling failure; it is + not proof that a human deadline elapsed. Human `DENIED` remains distinct. +- The compute role still chooses session tags when assuming the session role. + This API protects the decision-writing boundary but does not provide full + isolation from a compromised worker retaining ambient compute credentials. +- Linear consent is read back from Linear and must come from the mapped owner, + excluding bots and the saved OAuth token's own identity. Other webhook paths + still need a separate review of the legacy shared OAuth/signing-secret bundle. + +## References + +- [MicroVM backend, issue #645](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/645) +- [Implementation and security review, PR #904](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/pull/904) +- [Approval trust boundaries](/sample-autonomous-cloud-coding-agents/architecture/cedar-hitl-gates#121-trust-boundaries) +- [ADR-021: Lambda MicroVM compute backend](/sample-autonomous-cloud-coding-agents/decisions/adr-021-lambda-microvms-compute-backend) diff --git a/docs/src/content/docs/developer-guide/Contributing.md b/docs/src/content/docs/developer-guide/Contributing.md index 86f5f0b3e..eef135d34 100644 --- a/docs/src/content/docs/developer-guide/Contributing.md +++ b/docs/src/content/docs/developer-guide/Contributing.md @@ -18,9 +18,9 @@ Describe what you intend to contribute. This avoids duplicate work and gives mai ### 2. Set up your environment -Follow the [Quick Start](./docs/guides/QUICK_START.mdx) to clone, install, and build the project. See the [Developer guide](/sample-autonomous-cloud-coding-agents/developer-guide/introduction) for local testing and the development workflow. +Follow the [Quick Start](/sample-autonomous-cloud-coding-agents/getting-started/quick-start) to clone, install, and build the project. See the [Developer guide](/sample-autonomous-cloud-coding-agents/developer-guide/introduction) for local testing and the development workflow. -Use **[AGENTS.md](/sample-autonomous-cloud-coding-agents/architecture/agents)** to understand where to make changes (CDK vs CLI vs agent vs docs), which tests to extend, and common pitfalls (generated docs, mirrored API types, `mise` tasks). Package-specific detail lives in **`AGENTS.md`** under `cdk/`, `cli/`, `agent/`, and `docs/`. +Use **[AGENTS.md](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/AGENTS.md)** to understand where to make changes (CDK vs CLI vs agent vs docs), which tests to extend, and common pitfalls (generated docs, mirrored API types, `mise` tasks). Package-specific detail lives in **`AGENTS.md`** under `cdk/`, `cli/`, `agent/`, and `docs/`. ### 3. Implement your change @@ -32,7 +32,7 @@ Guidelines: - If you change API types in `cdk/src/handlers/shared/types.ts`, update `cli/src/types.ts` to match. - If you change docs sources (`docs/guides/`, `docs/design/`), run `mise //docs:sync` so generated content stays in sync. - For significant features, add a design document to `docs/design/`. -- For cross-cutting or hard-to-reverse decisions, add an ADR to `docs/decisions/` (see [ADR README](/sample-autonomous-cloud-coding-agents/architecture/readme)). +- For cross-cutting or hard-to-reverse decisions, add an ADR to `docs/decisions/` (see [ADR README](/sample-autonomous-cloud-coding-agents/decisions/readme)). ### 4. Commit @@ -102,6 +102,19 @@ PRs labeled `auto-approve` are approved automatically by the `auto-approve` work If `prek install` fails with "refusing to install hooks with `core.hooksPath` set", another tool owns your hooks. Either unset it (`git config --unset-all core.hooksPath`) or integrate these checks into your hook manager. +The build's transaction tests use DynamoDB Local through +`ABCA_DDB_LOCAL_ENDPOINT` (a loopback HTTP endpoint). CI starts a Docker service +pinned by image digest and supplies its mapped port to both CDK and Python tests; +those tests fail if `CI=true` without an endpoint. To reproduce locally, start +DynamoDB Local with `-inMemory -sharedDb`, bind port 8000 to loopback, and run +`ABCA_DDB_LOCAL_ENDPOINT=http://127.0.0.1:8000 mise run build`. + +The pinned SDK continuation probe also runs in CI. Locally enable it with +`ABCA_TEST_SDK_CONTINUATION=1` when running +`agent/tests/test_continuation_sdk_probe.py`. It uses a local simulated Bedrock +endpoint and synthetic credentials; it does not require an AWS account or model +billing. + ## Versioning The project uses semantic versioning based on [Conventional Commits](https://www.conventionalcommits.org/en/v1.0.0/): diff --git a/docs/src/content/docs/developer-guide/Installation.md b/docs/src/content/docs/developer-guide/Installation.md index 5a01813c6..449345a84 100644 --- a/docs/src/content/docs/developer-guide/Installation.md +++ b/docs/src/content/docs/developer-guide/Installation.md @@ -2,7 +2,7 @@ title: Installation --- -Follow the [Quick Start](./QUICK_START.mdx) to clone, install, deploy, and submit your first task. It covers prerequisites, toolchain setup, deployment, PAT configuration, Cognito user creation, and a smoke test. +Follow the [Quick Start](/sample-autonomous-cloud-coding-agents/getting-started/quick-start) to clone, install, deploy, and submit your first task. It covers prerequisites, toolchain setup, deployment, PAT configuration, Cognito user creation, and a smoke test. This section covers what the Quick Start does not: troubleshooting, local testing, and the development workflow. @@ -152,7 +152,7 @@ For the full list, see `agent/README.md`. For how the model default is layered, ### Deployment -Follow the [Quick Start](./QUICK_START.mdx) steps 3-6 for first-time deployment. For subsequent deploys after code changes: +Follow the [Quick Start](/sample-autonomous-cloud-coding-agents/getting-started/quick-start) steps 3-6 for first-time deployment. For subsequent deploys after code changes: ```bash mise run build diff --git a/docs/src/content/docs/developer-guide/Introduction.md b/docs/src/content/docs/developer-guide/Introduction.md index 5b259c4ca..8a81f45b7 100644 --- a/docs/src/content/docs/developer-guide/Introduction.md +++ b/docs/src/content/docs/developer-guide/Introduction.md @@ -14,6 +14,6 @@ The repository is organized around four main pieces: - **Infrastructure as code** in AWS CDK under `cdk/src/` - stacks, constructs, and handlers that define and deploy the platform on AWS. - **Documentation site** under `docs/` - source guides/design docs plus the generated Astro/Starlight documentation site. - **CLI package** under `cli/` - the `bgagent` command-line client used to authenticate, submit tasks, and inspect task status/events. -- **Claude Code plugin** under `docs/abca-plugin/` - a [Claude Code plugin](https://docs.anthropic.com/en/docs/claude-code/plugins) with guided skills and agents for setup, deployment, task submission, and troubleshooting. See the [plugin README](/sample-autonomous-cloud-coding-agents/architecture/readme) for details. +- **Claude Code plugin** under `docs/abca-plugin/` - a [Claude Code plugin](https://docs.anthropic.com/en/docs/claude-code/plugins) with guided skills and agents for setup, deployment, task submission, and troubleshooting. See the [plugin README](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/docs/abca-plugin/README.md) for details. > **Tip:** If you use Claude Code, run `claude --plugin-dir docs/abca-plugin` from the repo root. The plugin's `/setup` skill walks you through the entire setup process interactively. \ No newline at end of file diff --git a/docs/src/content/docs/developer-guide/Repository-preparation.md b/docs/src/content/docs/developer-guide/Repository-preparation.md index 0877d39b1..e267b0656 100644 --- a/docs/src/content/docs/developer-guide/Repository-preparation.md +++ b/docs/src/content/docs/developer-guide/Repository-preparation.md @@ -2,7 +2,7 @@ title: Repository preparation --- -The [Quick Start](./QUICK_START.mdx) covers the basic setup: forking a sample repo, creating a PAT, registering a Blueprint, and storing the token in Secrets Manager. This section covers what you need beyond that. +The [Quick Start](/sample-autonomous-cloud-coding-agents/getting-started/quick-start) covers the basic setup: forking a sample repo, creating a PAT, registering a Blueprint, and storing the token in Secrets Manager. This section covers what you need beyond that. ### Pre-flight checks @@ -13,7 +13,7 @@ Permission requirements vary by task type: - `new_task` and `pr_iteration` require Contents (read/write) and Pull requests (read/write). - `pr_review` only needs Triage or higher since it does not push branches. -Classic PATs with `repo` + `read:org` scopes also work and are required when fine-grained tokens cannot reach the target repo (collaborator access, cross-org repos). See [agent/README.md](/sample-autonomous-cloud-coding-agents/architecture/readme#github-pat--minimal-permissions) for when to use which token type. +Classic PATs with `repo` + `read:org` scopes also work and are required when fine-grained tokens cannot reach the target repo (collaborator access, cross-org repos). See [agent/README.md](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/agent/README.md#github-pat--minimal-permissions) for when to use which token type. ### Quick setup (single repo) @@ -79,7 +79,7 @@ The default image (`agent/Dockerfile`) includes Python, Node 24 (LTS), `git`, `g A blueprint can declare its own `security.cedarPolicies` rules on top of the built-in hard/soft-deny starter set. Hard-deny rules absolutely block a tool call; soft-deny rules pause the agent and ask a human before proceeding. -See the [Cedar policy guide](/sample-autonomous-cloud-coding-agents/customizing/cedar-policies) for the full authoring reference — vocabulary (`execute_bash`, `write_file`, `context.command`, `context.file_path`), annotations (`@rule_id`, `@tier`, `@approval_timeout_s`, `@severity`, `@category`), worked examples, multi-match rules, and cross-engine parity testing with [`contracts/cedar-parity/`](../../contracts/cedar-parity/) fixtures. +See the [Cedar policy guide](/sample-autonomous-cloud-coding-agents/customizing/cedar-policies) for the full authoring reference — vocabulary (`execute_bash`, `write_file`, `context.command`, `context.file_path`), annotations (`@rule_id`, `@tier`, `@approval_timeout_s`, `@severity`, `@category`), worked examples, multi-match rules, and cross-engine parity testing with [`contracts/cedar-parity/`](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/tree/main/contracts/cedar-parity) fixtures. ### Other options diff --git a/docs/src/content/docs/developer-guide/Where-to-make-changes.md b/docs/src/content/docs/developer-guide/Where-to-make-changes.md index 395e15315..7ba8d0713 100644 --- a/docs/src/content/docs/developer-guide/Where-to-make-changes.md +++ b/docs/src/content/docs/developer-guide/Where-to-make-changes.md @@ -12,4 +12,4 @@ Before editing, decide which part of the monorepo owns the behavior. This keeps | Agent runtime | `agent/` | Bundled into the image CDK deploys; run `mise run quality` in `agent/` or root build. | | Docs (source) | `docs/guides/`, `docs/design/` | After edits, run **`mise //docs:sync`** or **`mise //docs:build`**. Do not edit `docs/src/content/docs/` directly. | -For a concise duplicate of this table, common pitfalls, and a CDK test file map, see **[AGENTS.md](/sample-autonomous-cloud-coding-agents/architecture/agents)** at the repo root (oriented toward automation-assisted contributors). Package-specific detail lives in **`AGENTS.md`** under `cdk/`, `cli/`, `agent/`, and `docs/`. \ No newline at end of file +For a concise duplicate of this table, common pitfalls, and a CDK test file map, see **[AGENTS.md](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/AGENTS.md)** at the repo root (oriented toward automation-assisted contributors). Package-specific detail lives in **`AGENTS.md`** under `cdk/`, `cli/`, `agent/`, and `docs/`. \ No newline at end of file diff --git a/docs/src/content/docs/getting-started/Deployment-guide.md b/docs/src/content/docs/getting-started/Deployment-guide.md index 068612dc5..6dc802e44 100644 --- a/docs/src/content/docs/getting-started/Deployment-guide.md +++ b/docs/src/content/docs/getting-started/Deployment-guide.md @@ -4,35 +4,38 @@ title: Deployment guide # Deployment guide -This guide covers deploying ABCA into an AWS account, including compute backend choices, scale-to-zero characteristics, and the complete AWS service inventory. For day-to-day development workflow, see the [Developer guide](/sample-autonomous-cloud-coding-agents/developer-guide/introduction). For a quick first deployment, see the [Quick start](./QUICK_START.mdx). For least-privilege IAM deployment roles, see [DEPLOYMENT_ROLES.md](/sample-autonomous-cloud-coding-agents/architecture/deployment-roles). +This guide covers deploying ABCA into an AWS account, including compute backend choices, scale-to-zero characteristics, and the complete AWS service inventory. For day-to-day development workflow, see the [Developer guide](/sample-autonomous-cloud-coding-agents/developer-guide/introduction). For a quick first deployment, see the [Quick start](/sample-autonomous-cloud-coding-agents/getting-started/quick-start). For least-privilege IAM deployment roles, see [DEPLOYMENT_ROLES.md](/sample-autonomous-cloud-coding-agents/architecture/deployment-roles). ## Architecture overview -ABCA deploys as a **single CDK stack** (`backgroundagent-dev`) containing all platform resources. The stack uses a `ComputeStrategy` interface to support three compute backends within the same stack: +ABCA deploys through a **root CDK stack** (`backgroundagent-dev`) and nested stacks for parts of the platform, including registry infrastructure. A `ComputeStrategy` interface supports three compute backends: | Aspect | AgentCore (default) | ECS Fargate (opt-in) | Lambda MicroVMs (experimental) | |--------|--------------------|--------------------|--------------------| | **Compute** | Bedrock AgentCore Runtime (Firecracker MicroVMs) | ECS Fargate containers | AWS Lambda MicroVMs | -| **Resources** | 2 vCPU, 8 GB RAM, 2 GB max image size | 2 vCPU, 4 GB RAM | 8 GB baseline / 32 GB peak memory | +| **Resources** | 2 vCPU, 8 GB RAM, 2 GB max image size | Build: 4 vCPU / 16 GiB; planning: 2 vCPU / 8 GiB (configurable) | 8 GiB baseline / 32 GiB peak memory | | **Orchestration** | Durable Lambda (checkpoint/replay) | Same durable Lambda via `ComputeStrategy` | Same durable Lambda via `ComputeStrategy` | | **Agent mode** | FastAPI server (HTTP invocation) | Batch (run-to-completion) | FastAPI server (lifecycle hooks) | | **Startup** | ~10s (warm MicroVM) | ~60-180s (Fargate cold start) | ~6s to `RUNNING` (live-measured) | | **Max duration** | 8 hours (AgentCore service limit) | 9 hours (orchestrator `executionTimeout`) | 8 hours (`maximumDurationInSeconds`) | -All backends are orchestrated by the same durable Lambda function. The `ComputeStrategy` interface abstracts `startSession()`, `pollSession()`, and `stopSession()` -- the ECS strategy calls `ecs:RunTask` / `ecs:DescribeTasks` / `ecs:StopTask` directly from the Lambda. No Step Functions are used. +All backends are orchestrated by the same durable Lambda function. The `ComputeStrategy` interface abstracts `startSession()`, `pollSession()`, and `stopSession()` -- the ECS strategy calls `ecs:RunTask` / `ecs:DescribeTasks` / `ecs:StopTask` directly from the Lambda. Task orchestration does not use Step Functions; other platform features may use them. -ECS Fargate is currently **opt-in** -- the `EcsAgentCluster` construct is present in the stack code but commented out. To enable it, uncomment the ECS blocks in `cdk/src/stacks/agent.ts`. +ECS Fargate is **opt-in**. Deploy with `--context compute_type=ecs`; the stack enables `EcsAgentCluster` from that context flag. ### Lambda MicroVMs backend (experimental) -> **Not for production.** `lambda-microvm` carries no smoke-parity guarantee for an unattended deployment. Keep production repositories on `agentcore` or `ecs`. Synth emits an unsuppressible warning to this effect whenever the backend is selected. Design detail: [COMPUTE.md](/sample-autonomous-cloud-coding-agents/architecture/compute) and [ADR-021](/sample-autonomous-cloud-coding-agents/architecture/adr-021-lambda-microvms-compute-backend). +> **Not for production.** `lambda-microvm` carries no smoke-parity guarantee for an unattended deployment. Keep production repositories on `agentcore` or `ecs`. Synth emits a verification warning whenever a MicroVM image is configured; selecting the backend without an image emits a separate setup warning. Design detail: [COMPUTE.md](/sample-autonomous-cloud-coding-agents/architecture/compute) and [ADR-021](/sample-autonomous-cloud-coding-agents/decisions/adr-021-lambda-microvms-compute-backend). -Selecting it is a synth-time context flag: +For a new installation, select the backend and nested layout: ```bash -mise //cdk:deploy -- --context compute_type=lambda-microvm +mise //cdk:deploy -- --context compute_type=lambda-microvm --context microvm_nested_stack=true ``` +Existing flat installations must retain `microvm_nested_stack=false` until +completing the [resource migration](/sample-autonomous-cloud-coding-agents/verification/645-p3-nested-stack). + **You must re-bootstrap first.** This is the single most common way this backend fails, and the failure does not look like a configuration problem: 1. Check the bootstrap policy bundle already deployed in the account: @@ -57,9 +60,11 @@ mise //cdk:deploy -- --context compute_type=lambda-microvm Operational notes specific to this backend: -- **Nothing self-terminates.** A MicroVM whose task finished, crashed, or hung stays `RUNNING` and billing until the 8-hour cap. The orchestrator calls `TerminateMicrovm` on finalize, and the heartbeat-staleness check catches a hung guest inside a healthy VM -- but a leaked handle is a cost incident. The one exception: the service reaps a VM whose `/run` hook returns 4xx (~12s). +- **Nothing self-terminates.** A MicroVM whose task finished, crashed, or hung stays `RUNNING` and billing until the 8-hour cap. The orchestrator calls `TerminateMicrovm` on finalize, and the heartbeat-staleness check detects loss of the in-guest heartbeat writer (a pipeline hang can leave that writer running) -- but a leaked handle is a cost incident. The one exception: the service reaps a VM whose `/run` hook returns 4xx (~12s). - **Logs** land in `/aws/lambda-microvms/`. Guest stdout goes there too, which is the fallback path when the agent cannot reach the application log group. -- **Deployment identifiers are not baked into the image.** The snapshot carries no configuration; table names, secret ARNs, and the per-task session-role ARN arrive in the `/run` payload as a `platform_config` block. A version-skewed orchestrator that does not send it is refused rather than run with tenant scoping disabled. +- **Deployment identifiers are not baked into the image.** Current table names, secret ARNs and session-role ARN arrive through the v2 IAM-authenticated manifest and signed task document. Old unsigned envelopes are refused. Deploy matching coordinator code, worker images and IAM with admissions paused and old tasks drained; the repository runbook `docs/verification/645-payload-bootstrap.md` records the procedure and pending live checks. +- **Registry tools share the runtime network restriction.** Remote HTTP/SSE MCP assets need reachable HTTPS/443 endpoints. AgentCore and ECS defaults also block remote non-443 ports. `stdio` programs run locally but their outbound calls remain restricted; resolution does not test connectivity. The image builder's 80/443 access does not widen runtime egress. See [REGISTRY.md](/sample-autonomous-cloud-coding-agents/architecture/registry). +- **Logging failures have a fallback record.** Debug/warn CloudWatch failures emit `cloudwatch_write_failed` to stdout with writer/task/error class, without another AWS call or sensitive log text. This is not a configured metric/alarm; verify collection in the guest log stream while the VM is running. ### Optional Agent Registry @@ -245,6 +250,66 @@ Triggers via `workflow_run` when `build.yml` completes successfully. The pipelin ## Known deployment issues +### Upgrading approval permissions + +Worker approval creation and timeout writes now use an IAM-authenticated +service. Deploy the matching agent image and CDK together: old workers write +directly to DynamoDB and cannot create new gates after those permissions are removed. + +1. Pause submissions from the CLI, integrations and schedules during the upgrade. + Let existing tasks finish, or have their owners cancel them. Include tasks + awaiting approval, suspended MicroVMs and retained continuations; an empty + running-container list does not prove the deployment has drained. +2. Build the agent from the same revision as the CDK. For an externally managed + MicroVM image, publish that build and select its new version before resuming + submissions. A suspended VM keeps its old code. +3. Deploy the stack. Check that the SessionRole has approval-table reads and + condition checks only, plus `execute-api:Invoke` restricted to its task tag. + CDK supplies `APPROVAL_REQUESTS_API_URL` to all three compute backends. + Custom ECS constructs must provide both the SessionRole and service URL; + approval wiring without them is rejected before deployment. A MicroVM + manifest missing the URL is rejected before the worker starts. + If an AgentCore environment was edited to remove the URL, the worker reports + `APPROVAL_REQUESTS_API_URL is required for cloud approval requests` at its + next gate rather than attempting a direct DynamoDB write. Redeploy the + matching stack and image. +4. Submit a test task that triggers a known approval rule on each enabled backend. + Verify that the request appears, an owner decision resumes it, and an explicit + deadline records `TIMED_OUT` without overwriting a human decision. Then resume + normal submissions. + +If an old worker survives the upgrade, its next approval write fails closed. +Existing rows remain readable; do not restore direct writes to work around a stale +image. Roll forward with the matching image. Rolling back IAM restores the original +approval-record vulnerability and requires a deliberate operator decision. + +### Scheduled maintenance stack + +Concurrency repair, admission-queue pickup, stranded-task repair and pending-upload +cleanup run in the `ConcurrencyMaintenance` nested stack. MicroVM continuation +recovery also runs there when an image is configured. An upgrade recreates the +stateless functions, roles and schedules; their task tables and storage stay in +the parent stack. This applies to every compute backend, including AgentCore, +and replaces roughly twenty resources, depending on enabled features. +CloudFormation creates the new schedules before deleting the old ones, so both +can fire during the update. Task mutations and the continuation scan cursor use +conditional writes to tolerate that overlap. The drain above is required for +the approval permission/image upgrade, not a scheduling gap. + +Review the replacements in the change set. Existing Lambda log groups remain +under their old generated names; use the new function's log group for post-upgrade +invocations and retain the old groups when investigating earlier runs. +Both flat and nested MicroVM layouts support the optional +tool gateway and Linear Identity vault without exceeding the template budget. + +New MicroVM installations must explicitly select `microvm_nested_stack=true` +(bootstrap bundle 1.9.0). +Before upgrading an existing flat MicroVM deployment, save +`"microvm_nested_stack": false` in its CDK context or pass +`--context microvm_nested_stack=false` on every deploy. Keep this escape hatch +until completing the [resource migration](/sample-autonomous-cloud-coding-agents/verification/645-p3-nested-stack). +Omitting the setting fails synthesis. Selecting `true` does not migrate existing flat resources. + ### AgentCore unsupported Availability Zones **Affects:** Fresh deploys in accounts whose default Availability Zones don't line up with the zones AgentCore supports for the region. @@ -302,7 +367,7 @@ aws ec2 describe-subnets --filters "Name=vpc-id,Values=" \ --query 'Subnets[].[SubnetId,AvailabilityZone,AvailabilityZoneId]' --output text ``` -Be aware that destroying a VPC whose subnets held AgentCore ENIs can take 20–40 minutes while AWS reclaims them (see the `DELETE_FAILED` note in the [quick start](./QUICK_START.mdx) troubleshooting table). +Be aware that destroying a VPC whose subnets held AgentCore ENIs can take 20–40 minutes while AWS reclaims them (see the `DELETE_FAILED` note in the [quick start](/sample-autonomous-cloud-coding-agents/getting-started/quick-start) troubleshooting table). ### DNS Query Log Config replacement cascade (upgrading from pre-v0.5) @@ -365,11 +430,11 @@ For users without AWS CLI access. ## Related docs -- [Quick start](./QUICK_START.mdx) -- Zero-to-first-PR in 6 steps. +- [Quick start](/sample-autonomous-cloud-coding-agents/getting-started/quick-start) -- Zero-to-first-PR in 6 steps. - [Developer guide](/sample-autonomous-cloud-coding-agents/developer-guide/introduction) -- Local development, testing, repository onboarding. - [User guide](/sample-autonomous-cloud-coding-agents/using/overview) -- API reference, CLI usage, task management. - [DEPLOYMENT_ROLES.md](/sample-autonomous-cloud-coding-agents/architecture/deployment-roles) -- Least-privilege IAM policies for CloudFormation execution. - [COST_MODEL.md](/sample-autonomous-cloud-coding-agents/architecture/cost-model) -- Per-task costs, cost guardrails, cost at scale. - [COST_ATTRIBUTION.md](/sample-autonomous-cloud-coding-agents/getting-started/cost-attribution) -- Operator FinOps setup for per-user/per-repo Bedrock chargeback (Cost Explorer / CUR 2.0, invocation-log forensics). - [COMPUTE.md](/sample-autonomous-cloud-coding-agents/architecture/compute) -- Compute backend architecture and trade-offs. -- [ADR-021](/sample-autonomous-cloud-coding-agents/architecture/adr-021-lambda-microvms-compute-backend) -- Lambda MicroVMs backend decision, phased rollout, and live-verification evidence. +- [ADR-021](/sample-autonomous-cloud-coding-agents/decisions/adr-021-lambda-microvms-compute-backend) -- Lambda MicroVMs backend decision, phased rollout, and live-verification evidence. diff --git a/docs/src/content/docs/getting-started/Quick-start.mdx b/docs/src/content/docs/getting-started/Quick-start.mdx index b46a6113a..010ce7ef4 100644 --- a/docs/src/content/docs/getting-started/Quick-start.mdx +++ b/docs/src/content/docs/getting-started/Quick-start.mdx @@ -170,7 +170,7 @@ If you have not yet added `blueprintRepo` to `cdk/cdk.json`, go back to [Registe :::note[Collaborator or cross-org repos?] -Fine-grained tokens only work for repos you own (or orgs that have opted in). If you're a collaborator on someone else's repo, create a **classic PAT** with `repo` + `read:org` scopes instead. See [agent/README.md](/sample-autonomous-cloud-coding-agents/architecture/readme#github-pat--minimal-permissions) for details. +Fine-grained tokens only work for repos you own (or orgs that have opted in). If you're a collaborator on someone else's repo, create a **classic PAT** with `repo` + `read:org` scopes instead. See [agent/README.md](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/agent/README.md#github-pat--minimal-permissions) for details. ::: @@ -480,7 +480,7 @@ node lib/bin/bgagent.js deny \ --reason "Don't force-push shared branches; open a revert PR instead" ``` -The task transitions back to `RUNNING` immediately on a decision. The denial reason is injected into the agent's context so it can adapt rather than retry the same tool call. If no decision arrives within the rule's timeout (300 s by default), the gate is treated as a denial with `timed_out` as the reason. +A decision lets the task continue; a sleeping or retired MicroVM first wakes or restores its saved work. The denial reason is passed to the agent so it can adapt. Requests have no decision deadline by default. If an explicit task or rule deadline expires, the gate is treated as a denial with `timed_out` as the reason. Worker lifetime limits remain separate. If you want a task to run without interactive gates (e.g. an unattended overnight job), pre-approve the scopes you trust up-front: @@ -492,7 +492,7 @@ node lib/bin/bgagent.js submit --repo owner/repo --issue 42 \ --pre-approve write_path:tests/** ``` -Hard-deny rules (no `@tier("soft")` annotation) are always enforced — `--pre-approve` only short-circuits soft-deny rules. For the full command reference see [User guide — Approval gates](/sample-autonomous-cloud-coding-agents/using/overview#approval-gates-cedar-hitl); for authoring your own rules see the [Cedar policy guide](/sample-autonomous-cloud-coding-agents/customizing/cedar-policies). +Hard-deny rules (no `@tier("soft")` annotation) are always enforced — `--pre-approve` only short-circuits soft-deny rules. For the full command reference see [User guide — Approval gates](/sample-autonomous-cloud-coding-agents/using/approval-gates-cedar-hitl); for authoring your own rules see the [Cedar policy guide](/sample-autonomous-cloud-coding-agents/customizing/cedar-policies). ## What happened behind the scenes diff --git a/docs/src/content/docs/using/Approval-gates-cedar-hitl.md b/docs/src/content/docs/using/Approval-gates-cedar-hitl.md index b28882b5d..834372cfd 100644 --- a/docs/src/content/docs/using/Approval-gates-cedar-hitl.md +++ b/docs/src/content/docs/using/Approval-gates-cedar-hitl.md @@ -15,10 +15,27 @@ When a rule marked `@tier("soft")` matches a tool call: 1. The agent stops before invoking the tool. 2. A row is atomically written to the approvals table and the task status flips to `AWAITING_APPROVAL`. 3. A progress event (`approval_requested`) is emitted so `bgagent watch` shows the gate in real time. -4. The task waits for your decision up to the rule's timeout (default 300 s, configurable per-rule and per-task). +4. The task waits for your decision without an automatic deadline by default. An explicit per-task or policy-rule timeout can limit that window. 5. On approval, the agent proceeds; on denial, the deny reason is best-effort injected back into the agent's context so it can adapt; on timeout, the gate is treated as a denial with `timed_out` as the reason. -A decision is recorded at most once per request. Replaying approve/deny on the same `(task_id, request_id)` is idempotent. +A decision is recorded at most once per request. A repeated decision cannot change an already closed request. + +### Responding in Linear + +For a Linear task, the bot posts an **Approval needed** comment with the action +and reason. Reply **approve** or **deny** in that comment's thread. You do not +need a task ID, request ID, or bot mention. Your Linear account must be linked to +the ABCA account that submitted the task. + +`approve` allows the displayed action once. The bot acknowledges the saved +decision. An old thread cannot approve a newer request; replies to closed requests +explain that no new decision was recorded. Use the CLI for broader approval scopes +or a denial reason. Editing an existing comment does not submit a decision. + +This works on every compute backend. A sleeping MicroVM wakes to receive the +decision; AgentCore and ECS receive it through their existing approval wait. +Sleep does not set the approval deadline. Requests have no automatic deadline by +default, and an explicitly configured deadline still applies. ### Listing pending approvals @@ -26,7 +43,7 @@ A decision is recorded at most once per request. Replaying approve/deny on the s node lib/bin/bgagent.js pending ``` -Lists every approval across your tasks that is currently awaiting your decision. The default text output gives you the `request_id`, tool, severity, the reason the rule matched, the tool-input preview, the expiry time, and ready-to-run `approve` / `deny` command lines. Pipe through `--output json` for scripting. +Lists every approval across your tasks that is currently awaiting your decision. The default text output gives you the `request_id`, tool, severity, the reason the rule matched, the tool-input preview, the deadline or “no automatic expiry,” and ready-to-run `approve` / `deny` command lines. The JSON `expires_at` is `null` when there is no deadline. Cancelled or completed tasks no longer have answerable requests. ```text 1 pending approval(s): @@ -38,7 +55,7 @@ Lists every approval across your tasks that is currently awaiting your decision. rules: force_push_any preview: git push --force origin feature/xyz created: 2026-05-13T12:04:12Z - expires: 2026-05-13T12:09:12Z (timeout_s=300) + expires: no automatic expiry approve: bgagent approve 01KN37PZ77P1W19D71DTZ15X6X 01R... deny: bgagent deny 01KN37PZ77P1W19D71DTZ15X6X 01R... --reason "..." ``` @@ -98,10 +115,30 @@ node lib/bin/bgagent.js submit --repo owner/repo --issue 42 \ --pre-approve tool_type:Bash \ --pre-approve write_path:tests/** -# Per-task timeout override (platform default is 300s) +# Optional ten-minute decision deadline (the default has no deadline) node lib/bin/bgagent.js submit --repo owner/repo --issue 42 --approval-timeout 600 ``` `--pre-approve` can be repeated up to the platform limit (see `bgagent submit --help` for the current cap). Valid scope forms are the same as the `approve --scope` table above. Hard-deny rules are still enforced — `--pre-approve` only short-circuits soft-deny rules. -`--approval-timeout` sets the task-wide default; a rule with its own `@approval_timeout_s` annotation still takes the minimum of the two. \ No newline at end of file +`--approval-timeout 0` keeps unanswered requests available. A positive setting +limits the decision window; the shortest positive deadline from the task and +matching policy rules wins. Zero does not disable a policy rule's explicit +deadline. Cancelling the task closes its requests. + +For Lambda MicroVM tasks, `--microvm-sleep-after 600` selects the default +10-minute delay; `--microvm-sleep-after 120` selects two minutes and +`--microvm-sleep-after off` keeps the worker awake. The delay starts when each +approval request is created. Waking for approval, denial, or an approaching +deadline remains automatic. Sleep never starts a new approval timer. + +An unanswered request can outlive its MicroVM. After a longer wait, ABCA saves +the workspace and conversation, stops the worker, and releases its capacity. +Your answer can then start a replacement when capacity is available. A request +remaining open does not mean its old computer must stay alive. Sleeping saves +compute charges but adds snapshot save/restore charges and wake-up time; short +pauses can cost more than staying awake. The API equivalent is +`microvm_sleep_after_s` (zero means off); task +details return the saved setting. Automatic suspension is disabled by default for +new deployments. An operator enables it after [verifying the deployed image and +coordinator](/sample-autonomous-cloud-coding-agents/verification/readme#live-acceptance-for-an-installation). \ No newline at end of file diff --git a/docs/src/content/docs/using/Jira-setup-guide.md b/docs/src/content/docs/using/Jira-setup-guide.md index ddc12a4ba..1c621ca6b 100644 --- a/docs/src/content/docs/using/Jira-setup-guide.md +++ b/docs/src/content/docs/using/Jira-setup-guide.md @@ -107,7 +107,7 @@ Comments are advisory and best-effort: network/auth failures are logged and swal > (`mcp.atlassian.com`) requires an interactive, browser-based OAuth 2.1 flow > and cannot connect from a headless agent. Forge provides the supported app > actor through `api.asApp().requestJira(...)`. See -> [ADR-015](/sample-autonomous-cloud-coding-agents/architecture/adr-015-jira-integration). +> [ADR-015](/sample-autonomous-cloud-coding-agents/decisions/adr-015-jira-integration). Inbound admission (webhook → task) is Jira-specific and has no DynamoDB Streams consumer of its own. Ordinary **terminal** status comments are delivered by the shared fan-out plane's DynamoDB Streams consumer (`dispatchToJira`). For comment-triggered iterations, fan-out matures standalone status comments while the orchestration reconciler matures child-iteration comments before restacking dependents. diff --git a/docs/src/content/docs/using/Linear-setup-guide.md b/docs/src/content/docs/using/Linear-setup-guide.md index 6460bf507..e78572f25 100644 --- a/docs/src/content/docs/using/Linear-setup-guide.md +++ b/docs/src/content/docs/using/Linear-setup-guide.md @@ -15,7 +15,7 @@ Set up the ABCA Linear integration so that applying a label to a Linear issue tr ## How it works -You create a Linear OAuth app and authorize it on your workspace. When someone adds the trigger label to an issue in a mapped project, Linear fires a webhook at ABCA; the receiver verifies the HMAC signature, looks up the workspace, resolves a Linear API token, and creates a task. The agent clones the repo, makes the change, opens a PR, and comments back on the issue. +You create a Linear OAuth app and authorize it on your workspace. When someone adds the trigger label to an issue in a mapped project, Linear sends a webhook to ABCA. The receiver verifies its signature and invokes a processor, which resolves the workspace token and creates a task. The agent clones the repo, makes the change, opens a PR, and reports back on the issue. The app is installed with `actor=app`, so everything ABCA writes is attributed to the app rather than to whoever clicked Authorize. @@ -25,10 +25,10 @@ One of two places, chosen automatically at setup time: | | When it's used | What's stored | |---|---|---| -| **AgentCore Identity vault** | The stack was deployed with `--context enableLinearIdentityVault=true` | Nothing long-lived. AgentCore holds the refresh token and mints short-lived access tokens on demand. | +| **AgentCore Identity vault** | The stack was deployed with `--context enableLinearIdentityVault=true` | AgentCore manages the OAuth grant and token refresh. ABCA retains the OAuth client credentials and workspace/webhook metadata in Secrets Manager. | | **Secrets Manager** | Otherwise — including regions where AgentCore Identity isn't available | An OAuth token bundle in `bgagent-linear-oauth-`, refreshed and rotated by ABCA. | -The vault is unavailable on the `lambda-microvm` substrate — see [Not available with `compute_type=lambda-microvm`](#not-available-with-compute_typelambda-microvm) below. +The vault can be enabled with AgentCore, ECS or Lambda MicroVM compute. `bgagent linear setup` picks whichever the deployment supports and tells you which one it used. There is no flag. If the vault isn't available it prints one line and continues on Secrets Manager: @@ -38,11 +38,16 @@ AgentCore Identity not available in us-east-1 — using Secrets Manager. A workspace that started on Secrets Manager and later moves to the vault **keeps** its Secrets Manager token as a fallback. A workspace onboarded straight onto the vault has no such token by design — it needs the vault to be reachable. -When a workspace's authorization dies, ABCA records it on the registry row and publishes to the stack's operational alert topic. That topic has **no subscribers unless you deployed with `alertEmail`**, so set it if you want to hear about a dead workspace rather than discover it from `bgagent platform doctor`. +When a workspace's authorization dies, ABCA records it on the registry row and publishes to the stack's operational alert topic. A new topic starts without subscribers; configure `alertEmail` or subscribe another destination to receive those alerts. A revoked legacy Secrets Manager fallback does not by itself mean the active vault grant is revoked. -#### Not available with `compute_type=lambda-microvm` +#### Using the vault with Lambda MicroVMs -The vault and the Lambda MicroVMs substrate cannot be enabled on the same stack. Together they synthesize 505 CloudFormation resources against a hard limit of 500 (MicroVM alone is 496, the vault alone 488), so `cdk deploy` refuses the combination by name at synth rather than failing partway through. Use the vault on the `agentcore` or `ecs` substrate; a MicroVM stack stays on Secrets Manager until the stack reclaims room. +Deploy with `compute_type=lambda-microvm`, `enableLinearIdentityVault=true` and an explicit `microvm_nested_stack` value (`true` for new/already-nested installations; retain `false` for existing flat installations until migration). The coordinator sends the workload identity name through authenticated `platform_config`; the guest uses its compute execution role to obtain a Linear token. Credentials are not baked into the MicroVM image. + +When upgrading an existing MicroVM deployment, rebuild the guest image too: +the coordinator and guest must both support the vault configuration fields. + +The runtime resolves the vault in its AWS Region. A grant in another Region or under another workload identity does not automatically carry over. The old resource-count guard ([#857](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/issues/857)) has been replaced with configuration, permission and deployment-budget checks. #### One workload identity per stack @@ -212,7 +217,7 @@ Pick the teammate from the list of human members. You get a one-time code (24h T They need an ABCA account first. If they don't have one: -1. **Admin** runs `bgagent admin invite-user teammate@example.com` to create their Cognito user (see [User guide → Joining an existing deployment](/sample-autonomous-cloud-coding-agents/using/overview#joining-an-existing-deployment) for the full Cognito-side flow). +1. **Admin** runs `bgagent admin invite-user teammate@example.com` to create their Cognito user (see [User guide → Joining an existing deployment](/sample-autonomous-cloud-coding-agents/using/authentication#joining-an-existing-deployment) for the full Cognito-side flow). 2. **Teammate** pastes the bundle + temp password from the admin into: ```bash diff --git a/docs/src/content/docs/using/Roles.md b/docs/src/content/docs/using/Roles.md index f4c45ab7e..7ff440b1b 100644 --- a/docs/src/content/docs/using/Roles.md +++ b/docs/src/content/docs/using/Roles.md @@ -13,6 +13,6 @@ There are four lifecycle roles. They are often the same person early on, but the | **Repo onboarder** | Runs `bgagent linear onboard-project` (or registers a Blueprint via CDK) to wire a repo into the platform | As needed; any authenticated user | | **Teammate** | Runs `bgagent configure` once + `bgagent submit` / Linear or Jira label / Slack mention from then on | Daily user | -If you're a teammate joining an existing deployment, jump to [Joining an existing deployment](#joining-an-existing-deployment) below. +If you're a teammate joining an existing deployment, jump to [Joining an existing deployment](/sample-autonomous-cloud-coding-agents/using/authentication#joining-an-existing-deployment) below. -If you're standing up a new deployment from scratch, see the [Developer guide](/sample-autonomous-cloud-coding-agents/developer-guide/introduction) first, then come back here for the [admin onboarding flow](#get-stack-outputs). \ No newline at end of file +If you're standing up a new deployment from scratch, see the [Developer guide](/sample-autonomous-cloud-coding-agents/developer-guide/introduction) first, then come back here for the [admin onboarding flow](/sample-autonomous-cloud-coding-agents/using/authentication#get-stack-outputs). \ No newline at end of file diff --git a/docs/src/content/docs/using/Task-lifecycle.md b/docs/src/content/docs/using/Task-lifecycle.md index e416e934d..38b444ce0 100644 --- a/docs/src/content/docs/using/Task-lifecycle.md +++ b/docs/src/content/docs/using/Task-lifecycle.md @@ -28,7 +28,7 @@ The orchestrator uses Lambda Durable Functions to manage the lifecycle durably - | `SUBMITTED` | Task accepted; orchestrator invoked asynchronously | | `HYDRATING` | Orchestrator passed admission control; assembling the agent payload | | `RUNNING` | Agent session started and actively working on the task | -| `AWAITING_APPROVAL` | Agent paused at a Cedar HITL gate; waiting for your `approve` or `deny` decision. See [Approval gates](#approval-gates-cedar-hitl). | +| `AWAITING_APPROVAL` | Agent paused at a Cedar HITL gate; waiting for your `approve` or `deny` decision. See [Approval gates](/sample-autonomous-cloud-coding-agents/using/approval-gates-cedar-hitl). | | `FINALIZING` | Agent session ended; task is wrapping up (post-session hooks / PR finalization) before reaching a terminal state | | `COMPLETED` | Agent finished and created a PR (or determined no changes were needed) | | `FAILED` | Something went wrong - pre-flight check failed, concurrency limit reached, guardrail blocked the content, or the agent encountered an error | diff --git a/docs/src/content/docs/using/Using-the-cli.md b/docs/src/content/docs/using/Using-the-cli.md index 7b6c29a9f..affbfaac9 100644 --- a/docs/src/content/docs/using/Using-the-cli.md +++ b/docs/src/content/docs/using/Using-the-cli.md @@ -114,7 +114,8 @@ Created: 2026-04-01T00:39:51.271Z | `--max-budget` | Maximum cost budget in USD (0.01–100). Overrides per-repo Blueprint default. No default limit. | | `--idempotency-key` | Idempotency key for deduplication. | | `--trace` | Enable detailed tracing: raises progress preview cap to 4 KB and uploads full NDJSON trajectory to S3 on completion. Download with `bgagent trace download`. | -| `--approval-timeout` | Cedar HITL per-task approval timeout in seconds (default 300). A matching rule with its own `@approval_timeout_s` annotation still takes the minimum. See [Approval gates](#approval-gates-cedar-hitl). | +| `--approval-timeout` | Cedar HITL decision window: `0` (default) keeps unanswered requests available; a positive value sets a 30–3,600 second deadline. A matching rule can require a shorter positive deadline. See [Approval gates](/sample-autonomous-cloud-coding-agents/using/approval-gates-cedar-hitl). | +| `--microvm-sleep-after` | Seconds to wait for approval before putting a Lambda MicroVM to sleep (default 600 = 10 minutes; 0–3600 accepted). Use `off` to keep it awake. Requires the deployment's automatic-sleep feature to be enabled; does not change approval deadlines or affect other compute backends. | | `--pre-approve` | Cedar HITL scope to approve up-front (repeatable). Same scope forms as `bgagent approve --scope`. Hard-deny rules are always enforced. | | `--wait` | Poll until the task reaches a terminal status. | | `--output` | Output format: `text` (default) or `json`. | diff --git a/docs/src/content/docs/verification/645-p3-lifecycle-diagnostics.md b/docs/src/content/docs/verification/645-p3-lifecycle-diagnostics.md new file mode 100644 index 000000000..bda5d7655 --- /dev/null +++ b/docs/src/content/docs/verification/645-p3-lifecycle-diagnostics.md @@ -0,0 +1,95 @@ +--- +title: 645 p3 lifecycle diagnostics +--- + +# MicroVM lifecycle diagnostics + +An approval being saved, AWS accepting Resume, and the guest resuming work are +three different events. Diagnose each independently; an accepted API response +is not a completed wake. + +## Correlate coordinator and guest logs + +Search coordinator/approval Lambda logs and the MicroVM image log group using +`task_id` and `microvm_id`. `request_id` identifies the approval gate; +`aws_request_id` identifies an AWS call; `hook_id` identifies one guest HTTP hook. + +| Coordinator record | Meaning | +|---|---| +| `MicroVM observed after approval decision` | State and saved lifecycle intent observed after the decision. | +| `MicroVM wake request started/requested after approval decision` | Wake dispatch and acknowledgment, including receipt and elapsed time when available. | +| `MicroVM lifecycle request started/acknowledged/failed` | Suspend/Resume operation and result. | +| `MicroVM supervisor observation changed` | Task/worker/gate state, recovery timing, failure counters and outcome. | +| `MicroVM reached a terminal state with a substrate reason` | Service reason and worker/image identity. | +| `Lambda MicroVM termination requested` | Cleanup request and receipt. | + +Unchanged supervisor observations are deduplicated across durable replay. +Recovery timeouts retain their original start; investigating or retrying must +not reset them. + +For guest logs, select the deployed image's `/aws/lambda-microvms/` +log group and run a CloudWatch Logs Insights query such as: + +```text +fields @timestamp, event, action, stage, callback_stage, code, http_status, + hook_id, request_id, pid, phase, elapsed_ms, late, + error_type, aws_error_code, aws_request_id +| filter microvm_id = "REPLACE_WITH_WORKER_ID" +| sort @timestamp asc +| limit 500 +``` + +| Guest event | Meaning | +|---|---| +| `microvm_hook_started` | Handler entry, before body reading. | +| `microvm_hook_stage` | The next potentially blocking operation. | +| `microvm_hook_stage_finished` / `microvm_hook_stage_failed` | Operation completion or safe error metadata. | +| `microvm_hook_finished` | Selected handler status/code; not proof of service receipt. | + +`stage` records the last entered operation, such as credential refresh or approval +identity reconciliation. `late: true` means a callback finished after the handler +had already returned; it cannot turn a timed-out wake into success. The coding +barrier remains responsible for preventing tools after an uncertain wake. + +## Checkpoint failure codes + +The checkpoint diagnostic `code` narrows the failing operation; it does not +prove that saved data is corrupt. + +| Code | Meaning | +|---|---| +| `checkpoint_failed` | No narrower classification; inspect the accompanying stage and message. | +| `checkpoint_invalid_json` | Checkpoint JSON could not be encoded or decoded. | +| `checkpoint_sdk_unverified` | SDK version or required accounting interface is not verified. | +| `checkpoint_sdk_timeout` | SDK transcript acknowledgment or accounting request timed out. | +| `checkpoint_storage_unverified` | A save could not be verified by reading back the exact data; keep the source worker. | +| `checkpoint_storage_unavailable` | Storage configuration is unavailable or the saved version could not be read. | + +Workspace capture and restore can also report specific codes such as +`disk_pressure` or `git_timeout`. Preserve the reported code and stage when +escalating; do not replace them with a generic checkpoint label. + +## Diagnose a failed wake + +1. Locate the saved decision/intent and actual Resume acknowledgment or failure. + Retain the original generation, deadlines and API request ID. +2. Find guest hook entry. If absent, account for log delivery and retention before + inferring anything about the listener or process. +3. If entered, inspect the final stage/code. Credential-refresh `AccessDenied` + points to renewal permissions; identity-read failures point to task/gate + reconciliation. A timeout identifies the outstanding operation. +4. Check subsequent guest progress, coordinator outcome, worker termination and + capacity release. A saved approval or successful cleanup does not establish + that the approved tool ran. +5. Retain UTC timestamps, task/worker identifiers, exact image/coordinator + versions, service state reason, AWS receipts and relevant sanitized logs. + +`MICROVM_RESUME_HOOK_FAILED` identifies recognized service resume-hook failures. +Preserve the raw service reason for diagnosis. The wording “connection was +refused” alone does not prove a closed listener: historical guest observations +support a stale pooled-connection race, while service-side dispatch traces remain +unavailable. Lifecycle responses explicitly close connections before freeze. + +Diagnostics omit hook bodies, tool arguments, approval contents, credentials, +raw exception messages and SDK response bodies. Do not attach signed payload +URLs or conversation/workspace checkpoints when escalating an incident. diff --git a/docs/src/content/docs/verification/645-p3-nested-stack.md b/docs/src/content/docs/verification/645-p3-nested-stack.md new file mode 100644 index 000000000..0f4c56866 --- /dev/null +++ b/docs/src/content/docs/verification/645-p3-nested-stack.md @@ -0,0 +1,78 @@ +--- +title: 645 p3 nested stack +--- + +# Nested MicroVM infrastructure + +This is a CloudFormation infrastructure split, not a virtual machine running +inside another virtual machine. Fresh nested deployment and a deployment-specific +migration were verified; reusable migration commands remain unfinished. + +## Resource ownership + +`AgentStack` creates the `Microvm` child (`LambdaMicrovmStack`). It owns the +managed image when configured, build/runtime network connectors and security +groups, artifact/payload buckets, logs, and build/operator roles. + +The execution role remains at `LambdaMicrovmCompute/ExecutionRole` in the parent. +This preserves its logical ID and avoids a dependency cycle through SessionRole +trust. Parent `Microvm*` outputs retain the names used by packaging and consumers. +Names derive from the concrete parent deployment name, not the child stack token. + +## Configuration + +| Setting | Behavior | +|---|---| +| `microvm_nested_stack` | Required when MicroVM is enabled: use `true` for new or already-nested deployments; retain `false` for existing flat deployments until completing the reviewed migration. Omission fails synthesis. | +| `microvm_resource_name_prefix` | Supplies distinct names for overlapping nested resources; preserve it after migration. It does not retain old resources or permissions by itself. | +| `microvm_managed_image_version` | Pins new tasks to an explicitly verified image version. Without a pin, selection follows the latest active version. | +| `microvm_approval_suspend_enabled` | Defaults to `false`; enable only after testing the deployed image/coordinator. Disabling new sleep preserves wake and cleanup. | + +Nested deployment requires bootstrap bundle **1.9.0** or later. See +[deployment roles](/sample-autonomous-cloud-coding-agents/architecture/deployment-roles) and the +[artifact packaging instructions](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/cdk/scripts/README.md). +Build a new image before switching its runtime pin; building alone must not +silently change the image used by a pinned coordinator. Retain a compatible +published coordinator and exact image version for rollback. + +## Existing flat deployments + +Before the first upgrade, save `"microvm_nested_stack": false` in the deployment's +CDK context or pass `--context microvm_nested_stack=false` on every deploy. +Omitting the setting fails synthesis. This temporary requirement protects upgrades +that previously omitted the setting; it does not inspect live resources or perform +a migration. New and already-nested installations should explicitly set `true`. +Do not reuse older synthesized assemblies: regenerate them with this version. + +Do not deploy the nested template directly over a flat deployment. CloudFormation +sees removed parent resources and newly created child resources; names can collide +and bucket auto-delete handlers can erase artifacts or pending payloads. + +The tested provider rejected native `AWS::Lambda::MicrovmImage` stack refactoring +with an unsupported tag-schema error even though the preview succeeded. Do not +assume resource import/refactoring is supported from a successful preview. + +The required overlap migration has these stages. This checklist is not yet a +runnable migration command: + +1. Preserve the exact deployed template/cloud assembly, image and published + coordinator versions. Capture current resource identities and configuration. + Keep `microvm_nested_stack=false` on ordinary updates before migration. +2. Create the child with distinct names while retaining the old resources and + their permissions. Build and verify the new image before switching consumers. + Reject intermediate templates that exceed CloudFormation's 500-resource limit. +3. Deploy compatible approval handlers/coordinator before producers of retained + requests. Preserve permissions for old workers, including the old payload + access that P3's new bootstrap deny would otherwise block. +4. Switch consumers with an explicit image pin. Verify normal tasks, retained + approvals, sleep/wake and rollback while both resource sets are available. +5. Retire old resources only after old workers, Durable executions and uncertain + starts are accounted for and old tasks/leases are drained. An idle inventory + snapshot alone is not an admission fence. Preserve checkpoint data and any + resources still needed by published coordinator versions. + +Setting a concurrency counter to zero is not a reliable pause for asynchronous +producers. A rollback must restore compatible code, image selection and IAM +without deleting pending requests or saved work. The reusable tool must enforce +these prerequisites; migration template and permission helpers alone do not +constitute that tool. Track remaining acceptance in [verification status](/sample-autonomous-cloud-coding-agents/verification/readme#open-pr-checks). diff --git a/docs/src/content/docs/verification/645-payload-bootstrap.md b/docs/src/content/docs/verification/645-payload-bootstrap.md new file mode 100644 index 000000000..f722d4e98 --- /dev/null +++ b/docs/src/content/docs/verification/645-payload-bootstrap.md @@ -0,0 +1,83 @@ +--- +title: 645 payload bootstrap +--- + +# Trusted task delivery for ECS and MicroVM + +The coordinator delivers each task through an IAM-authenticated deployment +manifest and a signed URL for one task object. This is application startup +configuration, distinct from the CDK infrastructure bootstrap. + +## Contract and permissions + +Version 2 is defined in `contracts/constants.json` under `payload_bootstrap`. +The [ADR wire contract](/sample-autonomous-cloud-coding-agents/decisions/adr-021-lambda-microvms-compute-backend#3-packaging-same-agent-image-source-new-build-path) +and [security design](/sample-autonomous-cloud-coding-agents/architecture/security) describe its trust boundary. + +| Object | Purpose | +|---|---| +| `bootstrap/.json` | Non-secret backend/platform configuration; coordinator writes, worker reads with its ambient role. | +| `/payload.json` | Task instructions/configuration; downloaded only through the signed capability. | +| `/launch.json` | Private coordinator replay record containing the exact saved reference. | + +The worker's explicit S3 deny outside its own `bootstrap/*` prevents a foreign +public bucket from supplying fake configuration. A content hash alone does not +establish authorship. Workers cannot list their payload bucket or use ambient +credentials to read task objects. The coordinator needs `ListBucket` to distinguish +missing launch records from access denial. + +Downloaded task identity and configuration must match the authenticated manifest. +Old unsigned envelopes are rejected. All payload sizes use S3; the serialized +MicroVM reference must fit 4,096 bytes. Manifest/payload limits are 16 KiB/8 MiB. +MicroVM receives the reference through `runHookPayload`; ECS uses +`AGENT_PAYLOAD_REF` and removes it before repository subprocesses start. + +A signed URL is a bearer credential. Never log it or store it in a worker-readable +task row. HTTPS downloads reject redirects, environment proxies, alternate hosts +and invalid task paths. The runtime needs regional S3 HTTPS and DNS connectivity. + +## Retry and cleanup + +- Conditional writes keep task instructions immutable. Read-back reconciles a + write that succeeded but lost its response. +- Retries reuse the exact saved URL and Run client token. Re-signing would change + the request under the same token. An expired record fails explicitly. +- Signed lifetime is at most 900 seconds and can end earlier when temporary + credentials expire. Initial creation refuses a known lifetime under 300 seconds. +- A timeout does not prove no worker started. Reconcile the existing handle or + uncertain start before considering replacement; the application replay window + is not a service idempotency-retention guarantee. +- Finalization deletes payload and launch objects on a best-effort basis; bucket + lifecycle deletion is an asynchronous backstop. The coordinator refreshes + identical manifest bytes on preparation so old deployment settings remain usable. +- Resume continues saved task state with refreshed credentials; it does not + download the task again using an expired launch URL. + +## Coordinated upgrades + +Coordinator code, worker images/task definitions and IAM must implement the same +contract. Keep the previous deployable artifacts and inspect the change set. +For an upgrade that cannot support overlapping versions, pause actual admission +sources and drain tasks, approvals and uncertain starts before switching these +components together. A concurrency counter alone does not pause webhook/queue +admission. Do not delete and recreate storage to perform an upgrade. + +For overlapping flat-to-nested migration, preserve old-worker permissions until +those workers drain; follow the [migration prerequisites](/sample-autonomous-cloud-coding-agents/verification/645-p3-nested-stack). +Rollback also requires compatible code, images and IAM. Never reuse a task ID +with conflicting or expired stored launch instructions. + +## Verification + +Local producer/consumer tests cover conflicts, lost committed replies, malformed +or oversized input, wrong task/configuration, URL redaction and cleanup. These +cannot prove effective AWS permissions. In a representative deployment, verify: + +- Own manifest and signed payload succeed; foreign manifests, ambient payload + reads, bucket listing and worker mutations are denied. +- Modified signatures/paths, expired credentials/URLs and revoked objects fail + without starting the pipeline or exposing a capability in logs. +- Competing preparation and lost S3/Run replies preserve one exact launch request. +- Finalization removes both task objects and releases only confirmed capacity. + +See [recorded acceptance and remaining checks](/sample-autonomous-cloud-coding-agents/verification/readme). diff --git a/docs/src/content/docs/verification/Readme.md b/docs/src/content/docs/verification/Readme.md new file mode 100644 index 000000000..7ba3c667f --- /dev/null +++ b/docs/src/content/docs/verification/Readme.md @@ -0,0 +1,105 @@ +--- +title: Readme +--- + +# Lambda MicroVM verification + +For maintainers reviewing/testing the MicroVM backend and operators deploying, +migrating or diagnosing it. [ADR-021](/sample-autonomous-cloud-coding-agents/decisions/adr-021-lambda-microvms-compute-backend) +explains the design; the [user guide](/sample-autonomous-cloud-coding-agents/using/approval-gates-cedar-hitl) +explains approval and sleep options for people submitting tasks. +Detailed deployment transcripts, temporary worker identifiers and investigation +diaries are archived outside the repository. These documents are not test runners. + +- [Task payload delivery](/sample-autonomous-cloud-coding-agents/verification/645-payload-bootstrap): authorization, retries and coordinated upgrades. +- [Nested infrastructure](/sample-autonomous-cloud-coding-agents/verification/645-p3-nested-stack): configuration and migration prerequisites. +- [Lifecycle diagnostics](/sample-autonomous-cloud-coding-agents/verification/645-p3-lifecycle-diagnostics): locating and interpreting failed wakes. +- [Continuation design](/sample-autonomous-cloud-coding-agents/architecture/orchestrator#retained-microvm-approvals): checkpoint ownership, retirement and replacement. + +## Recorded acceptance + +AWS checks on September 14–18, 2026 exercised the following behavior in the tested +deployments. They do not certify a different image, account, Region or upgrade. + +| Area | Observed result | +|---|---| +| Task delivery | Signed payload downloads, invalid input rejection, expiry/revocation, immutable preparation and lost-reply recovery passed. | +| Approval lifecycle | Approve, deny, explicit expiry and cancellation passed; unanswered requests stayed available without a default deadline. | +| Sleep/wake | Repeated wakes, the 600-second default and the live sleep-off switch passed with compatible image/coordinator versions. | +| Credential renewal | A wait exceeding one hour was followed by successful AWS access with renewed task credentials. | +| Continuation | Conversation and Git/workspace recovery, replacement admission, usage limits and capacity release passed. | +| External integrations | Repository work and remote MCP access passed across sleep; an actual Linear submission exercised MicroVM compute with AgentCore Identity vault. | +| Linear decisions | Native threaded `approve` and `deny` replies passed on MicroVM and AgentCore. Both MicroVMs were suspended before the replies; exact decisions, tool results, thread acknowledgements and eventual capacity release were verified. | +| Infrastructure | Fresh nested deployment and a deployment-specific overlapping migration passed, including compatible rollback and old-resource cleanup. | +| Other backends | ECS and AgentCore approval/cancellation and scoped-access checks passed in their tested deployments. | + +The wake correction sends `Connection: close` in lifecycle responses before +freeze. Local transport controls reproduced failure on an old connection; +long-sleep controls and the corrected live flows passed. Service-side traces +for the historical failures remain unavailable, so their exact transport error +is not established for every worker. + +The full local build after the resource-budget fix passed: 5,560 CDK tests, +2,170 agent tests and 1,005 CLI tests, plus compile, lint, contracts, docs and +synthesis. The widest parent stacks use 489 resources for ECS and 488 for +MicroVM, including synth metadata, within the unchanged 490-resource budget. +Concurrency maintenance now has its own nested stack. Upgrading recreates that +stateless repair function and schedule; task and concurrency tables stay in the +parent stack. + +## Open PR checks + +- Finish reusable flat-to-nested migration commands and independently test an + upgrade from current `main` on the same deployment. The earlier bespoke + migration is not a substitute for that acceptance. +- Deploy the latest review fixes and verify CLI/API approve and deny on a + `PARKED` task, including immediate replacement admission. Earlier Linear + acceptance does not exercise the decision API functions' configuration. + +## Reproduce local checks + +From the repository root, with dependencies installed: + +```bash +mise run build +MISE_EXPERIMENTAL=1 mise //cdk:testf -- 'microvm|migration|payload-bootstrap' +``` + +For focused worker checks: + +```bash +cd agent +uv run pytest tests/test_microvm_*.py tests/test_continuation_*.py \ + tests/test_approval_retention.py tests/test_payload_bootstrap.py --no-cov +ABCA_TEST_SDK_CONTINUATION=1 uv run pytest tests/test_continuation_sdk_probe.py --no-cov +``` + +The last command opts into the pinned real SDK/CLI probe with a deterministic +loopback model; it does not launch a cloud worker. Optional DynamoDB Local tests +require their documented local service. Mocks do not establish effective AWS IAM. +The standalone cloud acceptance harness and raw receipts remain outside this PR. +The repository's narrower [launch, payload and replay probes](https://github.com/aws-samples/sample-autonomous-cloud-coding-agents/blob/main/cdk/test/live/README.md) +have separate inspection and execution commands. + +## Live acceptance for an installation + +Record source commit, image ARN/version, coordinator version, configuration and +UTC test window. Use an owned test repository and identity. Enable suspension +only after verifying that image and coordinator together. + +| Exercise | Required observation | +|---|---| +| Normal task | Repository tools run; terminal status, payload cleanup and capacity release agree. | +| Approval after sleep | The saved decision reaches the exact pending tool; verify guest recovery and tool output, not only the Resume API receipt. | +| Linear replies | Submit real issues on each backend. Reply `approve` or `deny` to the exact approval comment as its linked owner; verify the saved decision source, same-thread acknowledgement and allowed/blocked tool result. For MicroVM, observe suspension before replying. | +| Deny, expiry, cancellation | No denied/cancelled tool runs; expiry uses the original deadline; compute and capacity are cleaned up. | +| Default and disabled sleep | Omitted override uses 600 seconds; task-level off and deployment off prevent new suspension while wake/cleanup remain available. | +| Credential expiry | Sleep past the original credential lifetime, then perform actual task-scoped AWS operations. | +| Retirement/replacement | Confirm old-worker shutdown before capacity release; one replacement restores files/conversation and preserves approval identity and usage. | +| CLI/API replacement | After the task reaches `PARKED`, approve or deny through the CLI/API. Verify replacement admission starts from that decision without waiting for the scheduled sweep, and the exact pending tool follows the decision. | +| Upgrade/rollback | Preserve unrelated resource identities, old in-flight work and recoverable checkpoints; test compatible code/image/policy rollback. | + +Include effective-role checks for own-task access and denial of cross-task data, +foreign bootstrap manifests and ambient payload reads. Keep credentials, signed +URLs, prompts and checkpoints out of diagnostic attachments. Clean up test +workers, executions, task data and owned infrastructure after the run. diff --git a/docs/verification/645-p1-lambda-microvm-runbook.md b/docs/verification/645-p1-lambda-microvm-runbook.md deleted file mode 100644 index 12d1d56dd..000000000 --- a/docs/verification/645-p1-lambda-microvm-runbook.md +++ /dev/null @@ -1,2198 +0,0 @@ -# ADR-021 P1 Lambda MicroVM verification runbook - -Working verification document for issue #645 / PR #689 on branch -`feat/645-lambda-microvm-p1`. This is for a **new, CDK-bootstrapped sandbox -account in which ABCA has never been deployed**. It is not a production deploy -guide. - -`docs/scripts/sync-starlight.mjs` mirrors only selected guides, `docs/design/`, -`docs/decisions/`, `CONTRIBUTING.md`, and assets. It does **not** mirror -`docs/verification/`, so this file intentionally stays here. - -## ⚠️ This runbook predates the Stage D fixes - -The pass recorded under “Live execution results” below ran against the code as of -`0505f914`, and its findings (F1–F14) were then fixed. The **instructions** above -each step have been updated only where a fix renamed something an instruction -asserts on. The following instructions are known to still describe the pre-fix -world and must be adjusted by whoever runs this again: - -| Step | Stale instruction | Post-fix reality | -|---|---|---| -| 1.2 | ~~unescaped `ParameterValue=agentcore,lambda-microvm`; bootstrap without `--force`~~ | **CORRECTED IN PLACE** — the comma is now escaped (`agentcore\,lambda-microvm`) and both bootstrap invocations carry `--force`. F14/F10 are the evidence; nothing is left to adjust here | -| 2.2 | "a 443-only security group" and one `AWS::Lambda::NetworkConnector` | **two** security groups (443-only runtime, 443 + 80 build) and **two** connectors, plus a connector operator role; a seventh output `MicrovmBuildEgressConnectorArns` | -| 2.3 | `list-stack-resources --query "…|[0]"` | needs `--no-paginate` (F14) | -| 4.3 | `export IMAGE_VERSION=1` | the service returns `1.0`; there are two builds per version (one per chipset) | -| 5.1 / 5.9 | `--image-identifier "$IMAGE_NAME"` | an **ARN** is required (F3); use the `imageArn` the script prints | -| 5.9 | 16 384 / 16 385-byte probes | the real boundary is **4 096 / 4 097** (F6) | -| Phase 5 | "the hook-less P1 image" framing | the image now declares AND serves `/ready` + `/run`, so hook behaviour is a different experiment | - -## Important P1 and tooling limits - -- P1 provisions and can start the substrate, and — **as of the Stage D fixes** — - its image declares AND the agent serves `/ready` + `/run`, so the image is - creatable, launchable and payload-deliverable. What P1 has no guarantee of is - smoke parity (Memory grants, snapshot env parity, egress specifics from a - running MicroVM, heartbeats). *Pre-fix, this bullet read "not runnable end to - end" because the plan was to declare `/run` in P1 and serve it in P2; the live - run proved that is not a reachable service state (F1).* -- P1's orchestrator role intentionally has only `RunMicrovm`, `GetMicrovm`, - `TerminateMicrovm`, `PassNetworkConnector`, and the required `iam:PassRole`. - It does **not** have `SuspendMicrovm`, `ResumeMicrovm`, or - `CreateMicrovmAuthToken`. Therefore use the orchestrator role for the P1 IAM - checks when it can be assumed, but use the sandbox administrator identity for - the manual suspend/resume experiment. This is a deliberate P1/brief mismatch. -- The repository's AWS SDK model is - `@aws-sdk/client-lambda-microvms@3.1098.0`. It verifies operation names, - request keys, state enums, and `delete-microvm-image-version`. The local - `aws-cli/2.35.8` does **not** recognize `aws lambda-microvms`; consequently all - `aws lambda-microvms ...` commands below are **best-effort CLI spellings - derived from that SDK model and the repository packaging script**, not locally - CLI-validated. No minimum AWS CLI release containing this service could be - established. The executor must install a CLI build for which - `aws lambda-microvms help` succeeds. Do not proceed with image/lifecycle work - merely because `aws --version` is newer than 2.35.8. -- The packaging-script model drift is resolved: its direct service request now - uses SDK 3.1098.0's `ARM_64` architecture and `ENABLED|DISABLED` hook-state - shape (with port and timeout), rather than CloudFormation's `arm64` and hook - path strings. The CDK L1 remains intentionally unchanged because its generated - CloudFormation types accept string values and document no architecture/hook - allowed-value constraint. Step 4.2 still captures CLI help/input skeleton as - a live check because the local CLI cannot validate this service offline. -- Commands are run from the repository root. `mise` is primary. Commands marked - **raw fallback** are only for a machine without `mise`. -- **Run the whole thing under `set -o pipefail`.** Several steps pipe a command - through `tee`; without `pipefail` the pipeline reports `tee`'s exit status and a - failed command looks like a success. The 2026-07-31 pass recorded `EXIT=0` for a - `package-microvm-artifact.sh` run that had actually failed service validation - (F14). The script now also prints an explicit - `!! package-microvm-artifact.sh FAILED (exit N) !!` marker on any failure, so - the teed log carries the truth either way — but set the option anyway: - - ```bash - set -o pipefail - ``` - -## Variables and evidence directory - -**Purpose:** make every subsequent command target one account, Region, and stack. - -```bash -export AWS_REGION=us-east-1 -export AWS_DEFAULT_REGION="$AWS_REGION" -export CDK_DEFAULT_REGION="$AWS_REGION" -export STACK_NAME=backgroundagent-dev -export EXPECTED_BRANCH=feat/645-lambda-microvm-p1 -export EVIDENCE_DIR="/tmp/abca-645-p1-$(date -u +%Y%m%dT%H%M%SZ)" -mkdir -p "$EVIDENCE_DIR" -``` - -If using a profile, also `export AWS_PROFILE=`. Supported -Regions are `us-east-1`, `us-east-2`, `us-west-2`, `eu-west-1`, and -`ap-northeast-1`; this runbook defaults to `us-east-1`. - -**Expected:** the directory exists and all variables print non-empty. - -**Record:** variable values and evidence-directory path. - -**ADR-021 item:** regional availability enforcement and reproducibility. - ---- - -## Phase 0 — Preflight - -### 0.1 Verify identity, branch, and virgin account - -**Purpose:** prevent deploying to the wrong account/branch and fail fast if this -is not the assumed first ABCA deployment. - -```bash -aws sts get-caller-identity | tee "$EVIDENCE_DIR/caller-identity.json" -export ACCOUNT_ID="$(aws sts get-caller-identity --query Account --output text)" -test "$(git branch --show-current)" = "$EXPECTED_BRANCH" -git status --short --branch | tee "$EVIDENCE_DIR/git-status.txt" - -if aws cloudformation describe-stacks --stack-name "$STACK_NAME" \ - >"$EVIDENCE_DIR/unexpected-existing-stack.json" 2>"$EVIDENCE_DIR/stack-absence.txt"; then - echo "STOP: $STACK_NAME already exists; this runbook requires a virgin account." >&2 - exit 1 -fi -``` - -The expected CloudFormation error is `ValidationError: Stack ... does not -exist`. Any other error (especially `AccessDenied`) is not proof of absence; -stop and fix credentials. - -**Expected:** account is the intended sandbox, branch test passes, only known -local files are shown, and `backgroundagent-dev` is absent. `CDKToolkit` may and -normally will exist. - -**Record:** account ID, caller ARN, branch SHA (`git rev-parse HEAD`), status, -and the exact stack-absence response. - -**ADR-021 item:** clean first-deploy/substrate bootstrap path. - -### 0.2 Record tools and install dependencies - -**Purpose:** capture the exact client/service models and prepare CDK and CLI. - -```bash -aws --version 2>&1 | tee "$EVIDENCE_DIR/aws-version.txt" -node --version | tee "$EVIDENCE_DIR/node-version.txt" -npx cdk --version | tee "$EVIDENCE_DIR/cdk-version.txt" -mise --version | tee "$EVIDENCE_DIR/mise-version.txt" -python3 --version | tee "$EVIDENCE_DIR/python-version.txt" -zip -v | tee "$EVIDENCE_DIR/zip-version.txt" -rsync --version | tee "$EVIDENCE_DIR/rsync-version.txt" - -MISE_EXPERIMENTAL=1 mise run install -``` - -**Raw fallback if `mise` is absent:** - -```bash -yarn install --check-files -``` - -Then perform the mandatory service-client gate: - -```bash -aws lambda-microvms help >"$EVIDENCE_DIR/lambda-microvms-help.txt" -node -p "require('./node_modules/@aws-sdk/client-lambda-microvms/package.json').version" \ - | tee "$EVIDENCE_DIR/lambda-microvms-sdk-version.txt" -``` - -**Expected:** install succeeds; SDK version is `3.1098.0`; the help command -lists at least `run-microvm`, `get-microvm`, `suspend-microvm`, -`resume-microvm`, `terminate-microvm`, `create-microvm-auth-token`, and -`delete-microvm-image-version`. If help fails, stop and update/install the AWS -CLI distribution that exposes the preview/new service. - -**Record:** every version and whether service help was available. - -**ADR-021 item:** empirical IAM/API action-name verification. - ---- - -## Phase 1 — Least-privilege bootstrap - -### 1.1 Inspect the existing bootstrap - -**Purpose:** determine whether the standard bootstrap lacks ABCA's generated -template and custom compute policies. - -```bash -aws cloudformation describe-stacks --stack-name CDKToolkit \ - | tee "$EVIDENCE_DIR/cdktoolkit-before.json" -aws cloudformation get-template --stack-name CDKToolkit \ - --query TemplateBody --output text >"$EVIDENCE_DIR/cdktoolkit-template-before.txt" -python3 - "$EVIDENCE_DIR/cdktoolkit-template-before.txt" <<'PY' -import pathlib, sys -s = pathlib.Path(sys.argv[1]).read_text() -for needle in ("ComputeTypes", "IaCRoleABCAComputeLambdaMicrovms"): - print(needle, "present" if needle in s else "ABSENT") -PY -``` - -**Expected:** a standard bootstrap may report both markers absent. That is the -reason for the next step, not a failure. - -**Record:** template markers and current CDKToolkit parameters. - -**ADR-021 item:** conditional bootstrap policy exists only with the custom -template. - -### 1.2 Re-bootstrap, then set the CloudFormation parameter - -**Purpose:** replace the standard administrator bootstrap with the repository's -generated least-privilege template, then enable both AgentCore and Lambda -MicroVM deployment permissions. `cdk bootstrap` has no `--parameters`; CDK -context is not a substitute for this CloudFormation parameter. - -```bash -# `--force` is REQUIRED on an already-bootstrapped account: without it the CDK CLI -# refuses to replace the default template and exits 0 ("Bootstrap stack already -# exists, containing 'AWS CDK: Default Resources'. Not overwriting it…"), leaving -# AdministratorAccess attached while looking like a success (F10). Note also that -# BootstrapVariant stays 'AWS CDK: Default Resources' afterwards, so every future -# non-forced bootstrap refuses again. -MISE_EXPERIMENTAL=1 mise //cdk:bootstrap -- --force - -# The comma MUST be backslash-escaped. The CLI's shorthand parser otherwise splits -# on it and rejects the call (F14): -# "Invalid type for parameter Parameters[0].ParameterValue, -# value: ['agentcore', 'lambda-microvm'], type: , -# valid types: " -# The comment above `[tasks.bootstrap]` in cdk/mise.toml shows the same escaping. -aws cloudformation update-stack \ - --stack-name CDKToolkit \ - --use-previous-template \ - --capabilities CAPABILITY_NAMED_IAM \ - --parameters 'ParameterKey=ComputeTypes,ParameterValue=agentcore\,lambda-microvm' -aws cloudformation wait stack-update-complete --stack-name CDKToolkit -aws cloudformation describe-stacks --stack-name CDKToolkit \ - --query 'Stacks[0].Parameters' \ - | tee "$EVIDENCE_DIR/cdktoolkit-parameters-after.json" -``` - -**Raw fallback if `mise` is absent** (run from `cdk/`): - -```bash -npx tsx scripts/generate-bootstrap-artifacts.ts -npx tsx scripts/generate-bootstrap-template.ts -# --force for the same reason as above (F10). -npx cdk bootstrap --template bootstrap/bootstrap-template.yaml --force -``` - -Then run the same `aws cloudformation update-stack` parameter dance above, -including the escaped comma. - -**Expected:** `ComputeTypes` is exactly `agentcore,lambda-microvm`. The custom -template replaces default `AdministratorAccess` with generated ABCA policies. - -**Record:** update stack ID/events and final parameters. - -**ADR-021 item:** “where `ComputeTypes` includes `lambda-microvm`, attach -`IaCRole-ABCA-Compute-LambdaMicrovms`.” - -### 1.3 Verify policy creation and attachment - -**Purpose:** prove the MicroVM CloudFormation permissions are attached to the -actual execution role. - -```bash -export CFN_EXEC_ROLE="$(aws cloudformation describe-stack-resource \ - --stack-name CDKToolkit \ - --logical-resource-id CloudFormationExecutionRole \ - --query StackResourceDetail.PhysicalResourceId --output text)" - -aws iam list-policies --scope Local \ - --query "Policies[?contains(PolicyName, 'IaCRole-ABCA-Compute-LambdaMicrovms')].[PolicyName,Arn]" \ - --output table | tee "$EVIDENCE_DIR/microvm-bootstrap-policy.txt" -aws iam list-attached-role-policies --role-name "$CFN_EXEC_ROLE" \ - | tee "$EVIDENCE_DIR/cfn-exec-attached-policies.json" -``` - -**Expected:** one generated policy whose name contains -`IaCRole-ABCA-Compute-LambdaMicrovms` exists and its ARN is attached to -`$CFN_EXEC_ROLE`. - -**Record:** execution role name, policy ARN, and attachments. - -**ADR-021 item:** conditional bootstrap policy and verified IAM action names. - -**Optional negative deliberately skipped:** a scratch-qualifier bootstrap with -only `agentcore` would create another bootstrap stack, buckets, ECR repository, -roles, and policies merely to prove a template condition already covered by CDK -tests. It is not cheap enough for the core pass and complicates teardown. Run it -only if specifically requested, and destroy every scratch bootstrap resource. - ---- - -## Phase 2 — Substrate-only deploy (no image context) - -### 2.1 Synthesize and deploy the bootstrap state - -**Purpose:** verify the intended first-deploy state: connector, buckets, roles, -and logs exist while no image or orchestrator image configuration exists. - -```bash -MISE_EXPERIMENTAL=1 mise //cdk:synth -- \ - "$STACK_NAME" --context compute_type=lambda-microvm \ - 2>&1 | tee "$EVIDENCE_DIR/substrate-synth.txt" - -MISE_EXPERIMENTAL=1 mise //cdk:deploy -- \ - "$STACK_NAME" --require-approval never \ - --context compute_type=lambda-microvm \ - 2>&1 | tee "$EVIDENCE_DIR/substrate-deploy.txt" -``` - -**Raw fallback if `mise` is absent** (run from `cdk/`): - -```bash -npx cdk synth "$STACK_NAME" --context compute_type=lambda-microvm -npx cdk deploy "$STACK_NAME" --require-approval never --context compute_type=lambda-microvm -``` - -**Expected:** synth includes warning ID -`abca:microvm-image-not-provisioned`; deploy completes. This is intentionally -not `abca:microvm-image-p1-smoke-unverified` yet because no image is configured. -(The 2026-07-31 pass observed the pre-fix id `abca:microvm-image-p1-not-runnable`; -the warning was renamed when F1 was fixed.) - -**Record:** warning, deployment duration, stack ID/status, and failures/retries. - -**ADR-021 item:** conditional substrate and explicit no-image first-deploy -warning. - -### 2.2 Resolve exact outputs and resources - -**Purpose:** prove the script-facing substrate contract and capture physical IDs. - -```bash -aws cloudformation describe-stacks --stack-name "$STACK_NAME" \ - --query 'Stacks[0].Outputs' | tee "$EVIDENCE_DIR/stack-outputs-substrate.json" -aws cloudformation list-stack-resources --stack-name "$STACK_NAME" \ - | tee "$EVIDENCE_DIR/stack-resources-substrate.json" - -for key in ComputeSubstrate MicrovmArtifactBucketName MicrovmArtifactObjectKey \ - MicrovmBuildRoleArn MicrovmExecutionRoleArn MicrovmEgressConnectorArns \ - MicrovmLogGroupName; do - aws cloudformation describe-stacks --stack-name "$STACK_NAME" \ - --query "Stacks[0].Outputs[?OutputKey=='$key'].OutputValue | [0]" --output text -done | tee "$EVIDENCE_DIR/microvm-output-values.txt" -``` - -**Expected:** `ComputeSubstrate=lambda-microvm`; all six `Microvm...` outputs -are populated; artifact key is `microvm-images/agent-artifact.zip`. Stack -resources include two S3 buckets, build/execution roles, a 443-only security -group, `/aws/lambda-microvms/...` log group, and -`AWS::Lambda::NetworkConnector`; no `AWS::Lambda::MicrovmImage` exists. - -**Record:** outputs and physical IDs. - -**ADR-021 item:** construct resources, egress connector, build/execution roles, -artifact/payload buckets. - -### 2.3 Verify no orchestrator `MICROVM_*` environment - -**Purpose:** prove partial image configuration is not injected. - -```bash -export ORCHESTRATOR_FN="$(aws cloudformation list-stack-resources \ - --stack-name "$STACK_NAME" \ - --query "StackResourceSummaries[?ResourceType=='AWS::Lambda::Function' && contains(LogicalResourceId, 'TaskOrchestrator')].PhysicalResourceId | [0]" \ - --output text)" -aws lambda get-function-configuration --function-name "$ORCHESTRATOR_FN" \ - --query 'Environment.Variables' | tee "$EVIDENCE_DIR/orchestrator-env-no-image.json" -aws lambda get-function-configuration --function-name "$ORCHESTRATOR_FN" \ - --query 'Environment.Variables' --output json \ - | python3 -c 'import json,sys; d=json.load(sys.stdin); print([k for k in d if k.startswith("MICROVM_")])' -``` - -**Expected:** the final line is `[]`. - -**Record:** function name and environment-key list (do not publish environment -values if later deployments add sensitive configuration). - -**ADR-021 item:** reject tasks when deployed without an image; all-or-nothing -strategy configuration. - -### 2.4 Verify backend cost tags - -**Purpose:** verify deployed, taggable construct resources carry -`abca:compute-backend=lambda-microvm`. - -```bash -aws resourcegroupstaggingapi get-resources \ - --tag-filters Key=abca:compute-backend,Values=lambda-microvm \ - --query 'ResourceTagMappingList[].ResourceARN' --output text \ - | tee "$EVIDENCE_DIR/microvm-tagged-resource-arns.txt" - -aws cloudformation get-template --stack-name "$STACK_NAME" \ - --query TemplateBody --output json >"$EVIDENCE_DIR/deployed-template.json" -python3 - "$EVIDENCE_DIR/deployed-template.json" <<'PY' -import json, pathlib, sys -t = json.loads(pathlib.Path(sys.argv[1]).read_text()) -for logical_id, r in t["Resources"].items(): - if "LambdaMicrovmCompute" not in logical_id: - continue - tags = r.get("Properties", {}).get("Tags") - print(logical_id, r["Type"], tags if tags is not None else "NOT-TAGGABLE/NO-TAGS") -PY -``` - -**Expected:** every taggable construct resource (buckets, roles, security group, -log group, and connector where supported) shows the backend tag. Generated -policies/bucket policies are not independently taggable resources. Save any -service that omits tags as a defect rather than silently accepting it. - -**Record:** tagged ARNs and the per-logical-resource template report. - -**ADR-021 item:** backend-identifying cost-allocation tags. - ---- - -## Phase 3 — Region gate (synth only) - -### 3.1 Reject an unsupported Region - -**Purpose:** prove static fail-fast enforcement without deploying there. - -```bash -env AWS_REGION=eu-central-1 AWS_DEFAULT_REGION=eu-central-1 \ - CDK_DEFAULT_REGION=eu-central-1 CDK_DEFAULT_ACCOUNT="$ACCOUNT_ID" \ - MISE_EXPERIMENTAL=1 mise //cdk:synth -- \ - "$STACK_NAME" --context compute_type=lambda-microvm \ - >"$EVIDENCE_DIR/unsupported-region-synth.txt" 2>&1 && { - echo "ERROR: unsupported-region synth unexpectedly succeeded" >&2; exit 1; - } -``` - -**Raw fallback if `mise` is absent** (run from `cdk/`): - -```bash -env AWS_REGION=eu-central-1 AWS_DEFAULT_REGION=eu-central-1 \ - CDK_DEFAULT_REGION=eu-central-1 CDK_DEFAULT_ACCOUNT="$ACCOUNT_ID" \ - npx cdk synth "$STACK_NAME" --context compute_type=lambda-microvm -``` - -**Expected:** failure names `eu-central-1`, all five supported Regions, and -`--context microvm_region_override=true`. - -**Record:** complete stderr. - -**ADR-021 item:** static unsupported-Region synth failure. - -### 3.2 Exercise the escape hatch - -**Purpose:** prove newly launched Regions can bypass only the static list. - -```bash -env AWS_REGION=eu-central-1 AWS_DEFAULT_REGION=eu-central-1 \ - CDK_DEFAULT_REGION=eu-central-1 CDK_DEFAULT_ACCOUNT="$ACCOUNT_ID" \ - MISE_EXPERIMENTAL=1 mise //cdk:synth -- \ - "$STACK_NAME" --context compute_type=lambda-microvm \ - --context microvm_region_override=true \ - 2>&1 | tee "$EVIDENCE_DIR/unsupported-region-override-synth.txt" -``` - -**Expected:** synth succeeds with warning `abca:microvm-region-override`. - -**Record:** warning and exit status. - -**ADR-021 item:** Region-list escape hatch. - -**Gotcha:** `src/main.ts` reads `CDK_DEFAULT_REGION`, not merely `AWS_REGION`. -An unresolved/region-agnostic CDK token skips the static check by design. The -commands set both account and Region to force the real test. Synth makes no -MicroVM control-plane calls, but CDK context/asset bundling may still require -valid AWS credentials and the bootstrap version parameter. - ---- - -## Phase 4 — Package and build an image - -### 4.1 Select a managed base image - -**Purpose:** pin a real regional base-image ARN/version rather than guessing. - -```bash -aws lambda-microvms list-managed-microvm-images \ - | tee "$EVIDENCE_DIR/managed-images.json" -export BASE_IMAGE_ARN="$(aws lambda-microvms list-managed-microvm-images \ - --query 'items[0].imageArn' --output text)" -aws lambda-microvms list-managed-microvm-image-versions \ - --image-identifier "$BASE_IMAGE_ARN" \ - | tee "$EVIDENCE_DIR/managed-image-versions.json" -# NEWEST FIRST (measured 2026-07-31): items[0] is the latest version, items[-1] -# is the OLDEST. The original `items[-1]` here selected version 0 instead of 1. -export BASE_IMAGE_VERSION="$(aws lambda-microvms list-managed-microvm-image-versions \ - --image-identifier "$BASE_IMAGE_ARN" \ - --query 'items[0].imageVersion' --output text)" -test -n "$BASE_IMAGE_ARN" && test "$BASE_IMAGE_ARN" != None -test -n "$BASE_IMAGE_VERSION" && test "$BASE_IMAGE_VERSION" != None -``` - -**Expected:** the regional probe succeeds and returns at least one ARN/version. -Ordering is newest-to-oldest, so `items[0]` is correct; inspect the timestamps -and explicitly export the desired version if the installed CLI ever differs. - -**Record:** complete catalogs and selected pair. - -**ADR-021 item:** live regional availability probe and managed base-image API. - -### 4.2 Package, upload, and start the out-of-band build - -**Purpose:** exercise the actual script interface and avoid slow CloudFormation -iteration while still using CDK-created bucket, role, connector, and logs. - -Before running it, capture the installed CLI's authoritative request shape: - -```bash -aws lambda-microvms create-microvm-image help \ - >"$EVIDENCE_DIR/create-microvm-image-help.txt" -aws lambda-microvms create-microvm-image --generate-cli-skeleton input \ - >"$EVIDENCE_DIR/create-microvm-image-skeleton.json" -``` - -Confirm that `hooks.microvmHooks.run` is `ENABLED` and -`cpuConfigurations[].architecture` is `ARM_64`, matching SDK 3.1098.0. The CDK -L1 request is a separate CloudFormation surface and legitimately retains its -generated path/string shape. If the installed CLI skeleton differs from the SDK -model or rejects the script request, stop this phase, save the parser/service -error as a model-drift defect, and mark later image/runtime steps blocked. - -```bash -export IMAGE_NAME="${STACK_NAME}-abca-agent" -export BUILD_STARTED_AT="$(date -u +%Y-%m-%dT%H:%M:%SZ)" -cdk/scripts/package-microvm-artifact.sh \ - --stack-name "$STACK_NAME" \ - --create-image \ - --base-image-arn "$BASE_IMAGE_ARN" \ - --base-image-version "$BASE_IMAGE_VERSION" \ - --image-name "$IMAGE_NAME" \ - 2>&1 | tee "$EVIDENCE_DIR/package-and-create-image.txt" -# Explicit status check — `| tee` reports tee's status, so a bare `$?` here lies -# unless `set -o pipefail` is on (see "Important P1 and tooling limits"). -test "${PIPESTATUS[0]}" -eq 0 -``` - -The script requires `aws`, `zip`, `python3`, and `rsync`; reads outputs -`MicrovmArtifactBucketName`, `MicrovmArtifactObjectKey`, -`MicrovmBuildRoleArn`, `MicrovmBuildEgressConnectorArns`, -`MicrovmEgressConnectorArns`, and `MicrovmLogGroupName`; stages root -`Dockerfile`, `agent/`, and `contracts/`; uploads the zip; and calls -`create-microvm-image` with ARM64, 8,192 MiB, `/ready` **and** `/run` enabled on -port 8080, the **build-time** egress connector (443 + 80), and the backend tag. - -**Expected:** upload succeeds, create returns/starts image version `1.0` in this -virgin image name, and the output contains the conspicuous “P1 image is runnable -but NOT smoke-verified” reminder — printed BOTH before and after the create call, -so a failing create cannot swallow it. - -**Record:** artifact size printed by the script, S3 object size from the command -below, create response, and exact banner. - -```bash -export ARTIFACT_BUCKET="$(aws cloudformation describe-stacks --stack-name "$STACK_NAME" \ - --query "Stacks[0].Outputs[?OutputKey=='MicrovmArtifactBucketName'].OutputValue | [0]" --output text)" -export ARTIFACT_KEY="$(aws cloudformation describe-stacks --stack-name "$STACK_NAME" \ - --query "Stacks[0].Outputs[?OutputKey=='MicrovmArtifactObjectKey'].OutputValue | [0]" --output text)" -aws s3api head-object --bucket "$ARTIFACT_BUCKET" --key "$ARTIFACT_KEY" \ - | tee "$EVIDENCE_DIR/artifact-head.json" -``` - -**ADR-021 item:** zip+Dockerfile packaging plane, no secret build inputs, build -role/connector/log wiring, and the `abca:microvm-image-p1-smoke-unverified` -warning (renamed from `…-p1-not-runnable` when F1 was fixed: the image IS -creatable, launchable and payload-deliverable now that `/ready` + `/run` are -declared AND served — what is unverified is smoke parity). - -### 4.3 Poll build and image status to ACTIVE - -**Purpose:** capture real snapshot build duration, state, and component sizes. - -```bash -export IMAGE_VERSION=1 -while :; do - aws lambda-microvms list-microvm-image-builds \ - --image-identifier "$IMAGE_NAME" --image-version "$IMAGE_VERSION" \ - | tee "$EVIDENCE_DIR/image-builds-latest.json" - STATE="$(aws lambda-microvms list-microvm-image-builds \ - --image-identifier "$IMAGE_NAME" --image-version "$IMAGE_VERSION" \ - --query 'items[0].buildState' --output text)" - date -u '+%Y-%m-%dT%H:%M:%SZ buildState='"$STATE" - case "$STATE" in SUCCESSFUL) break;; FAILED) exit 1;; esac - sleep 30 -done - -export BUILD_ID="$(aws lambda-microvms list-microvm-image-builds \ - --image-identifier "$IMAGE_NAME" --image-version "$IMAGE_VERSION" \ - --query 'items[0].buildId' --output text)" -aws lambda-microvms get-microvm-image-build \ - --image-identifier "$IMAGE_NAME" --image-version "$IMAGE_VERSION" \ - --build-id "$BUILD_ID" | tee "$EVIDENCE_DIR/image-build-final.json" - -while :; do - aws lambda-microvms get-microvm-image-version \ - --image-identifier "$IMAGE_NAME" --image-version "$IMAGE_VERSION" \ - | tee "$EVIDENCE_DIR/image-version-latest.json" - STATUS="$(aws lambda-microvms get-microvm-image-version \ - --image-identifier "$IMAGE_NAME" --image-version "$IMAGE_VERSION" \ - --query status --output text)" - test "$STATUS" = ACTIVE && break - sleep 30 -done -export BUILD_FINISHED_AT="$(date -u +%Y-%m-%dT%H:%M:%SZ)" -``` - -**Expected:** build states may include `PENDING` and `IN_PROGRESS`, then -`SUCCESSFUL`; image-version `state` becomes `SUCCESSFUL` and `status` becomes -`ACTIVE`. On failure, save `stateReason` and tail the exact output log group: - -```bash -export MICROVM_LOG_GROUP="$(aws cloudformation describe-stacks --stack-name "$STACK_NAME" \ - --query "Stacks[0].Outputs[?OutputKey=='MicrovmLogGroupName'].OutputValue | [0]" --output text)" -aws logs tail "$MICROVM_LOG_GROUP" --since 2h -``` - -**Record:** start/finish timestamps, duration, all states/reasons, build ID, and -`snapshotBuild.memorySnapshotSizeInBytes`, `codeInstallSizeInBytes`, and -`diskSnapshotSizeInBytes`. Compare **code-install size** to AgentCore's 2 GB -container-image limit, while clearly noting that memory/disk snapshots are not -equivalent to an OCI image and must not be summed into a misleading comparison. -Also record the reported image resources/disk facts; the SDK exposes minimum -memory but no explicit disk-capacity field, so verify the 32 GB disk claim in -the service quota/console and report “not exposed” if that remains true. - -**ADR-021 item:** buildability, final image size, 2 GB-limit narrative, and disk -quota external fact. - -### 4.4 Redeploy against the built image and inspect IAM/env - -**Purpose:** hand the out-of-band image to the orchestrator and prove exact-image -IAM scoping. - -```bash -MISE_EXPERIMENTAL=1 mise //cdk:deploy -- \ - "$STACK_NAME" --require-approval never \ - --context compute_type=lambda-microvm \ - --context microvm_image_identifier="$IMAGE_NAME" \ - --context microvm_image_version="$IMAGE_VERSION" \ - 2>&1 | tee "$EVIDENCE_DIR/image-configured-deploy.txt" -``` - -**Expected:** synth/deploy emits `abca:microvm-image-p1-smoke-unverified` (the -2026-07-31 pass saw the pre-fix id `abca:microvm-image-p1-not-runnable`; the -warning was renamed when F1 was fixed). The orchestrator now has -`MICROVM_IMAGE_IDENTIFIER` — **a full image ARN**, not a bare name (F3) — -`MICROVM_IMAGE_VERSION`, `MICROVM_EXECUTION_ROLE_ARN`, -`MICROVM_EGRESS_CONNECTOR_ARNS`, `MICROVM_PAYLOAD_BUCKET`, and -`MICROVM_INGRESS_CONNECTOR_ARNS` carrying the Lambda-managed `NO_INGRESS` -connector (F7 — the pre-fix build had no ingress variable at all). - -```bash -aws lambda get-function-configuration --function-name "$ORCHESTRATOR_FN" \ - --query 'Environment.Variables' --output json \ - | python3 -c 'import json,sys; d=json.load(sys.stdin); print({k:d[k] for k in d if k.startswith("MICROVM_")})' \ - | tee "$EVIDENCE_DIR/orchestrator-microvm-env.txt" - -export ORCHESTRATOR_ROLE_ARN="$(aws lambda get-function-configuration \ - --function-name "$ORCHESTRATOR_FN" --query Role --output text)" -export ORCHESTRATOR_ROLE="${ORCHESTRATOR_ROLE_ARN##*/}" -aws iam list-role-policies --role-name "$ORCHESTRATOR_ROLE" \ - | tee "$EVIDENCE_DIR/orchestrator-inline-policy-names.json" -export ORCH_POLICY_NAME="$(aws iam list-role-policies --role-name "$ORCHESTRATOR_ROLE" \ - --query 'PolicyNames[0]' --output text)" -aws iam get-role-policy --role-name "$ORCHESTRATOR_ROLE" \ - --policy-name "$ORCH_POLICY_NAME" \ - | tee "$EVIDENCE_DIR/orchestrator-inline-policy.json" -``` - -The inline policy's physical name is CDK-generated and therefore cannot be -hard-coded; the actual name to fetch is `$ORCH_POLICY_NAME` returned by -`list-role-policies` (normally the role's `DefaultPolicy`). If more than one is -listed, fetch each and select the document containing `Sid=MicrovmLifecycle`. - -**Expected:** `MicrovmLifecycle` grants exactly `lambda:RunMicrovm`, -`lambda:GetMicrovm`, and `lambda:TerminateMicrovm` against exactly -`arn:...:microvm-image:$IMAGE_NAME` and its `:` suffix sibling; -`MicrovmPassNetworkConnector` has `lambda:PassNetworkConnector` on `*`; no -`SuspendMicrovm`, `ResumeMicrovm`, or `CreateMicrovmAuthToken` exists. - -**Record:** warning, environment-key/value map, role/policy names, statements, -and exact image ARN format observed. - -**ADR-021 item:** the `abca:microvm-image-p1-smoke-unverified` warning (formerly -`…-p1-not-runnable`), exact-ARN lifecycle IAM, no-JWE grant, and all-or-nothing -environment wiring — which now includes `MICROVM_INGRESS_CONNECTOR_ARNS`, always -injected, carrying the `NO_INGRESS` control. - ---- - -## Phase 5 — Manual lifecycle and empirical checklist - -### 5.0 Resolve launch inputs and IAM identity mode - -**Purpose:** use deployed values and distinguish true role-policy evidence from -admin-only lifecycle evidence. - -```bash -export EXECUTION_ROLE_ARN="$(aws cloudformation describe-stacks --stack-name "$STACK_NAME" \ - --query "Stacks[0].Outputs[?OutputKey=='MicrovmExecutionRoleArn'].OutputValue | [0]" --output text)" -export EGRESS_CONNECTORS="$(aws cloudformation describe-stacks --stack-name "$STACK_NAME" \ - --query "Stacks[0].Outputs[?OutputKey=='MicrovmEgressConnectorArns'].OutputValue | [0]" --output text)" - -aws sts assume-role --role-arn "$ORCHESTRATOR_ROLE_ARN" \ - --role-session-name abca-645-verification \ - >"$EVIDENCE_DIR/orchestrator-assume-role.json" \ - 2>"$EVIDENCE_DIR/orchestrator-assume-role-error.txt" || true -``` - -Lambda execution-role trust normally allows only `lambda.amazonaws.com`, so -operator assumption is expected to fail unless the sandbox has an explicit -test trust path. **Do not modify production-like trust just for this run.** If -assumption succeeds, open a subshell with those temporary credentials for steps -5.1, 5.2, 5.7, and 5.8(a/b), and mark evidence `ORCHESTRATOR_ROLE`. Otherwise -run as sandbox admin and mark scoping observations `ADMIN — advisory`; the -static inline-policy inspection in 4.4 remains authoritative. - -Example temporary-credential subshell setup: - -```bash -read AWS_ACCESS_KEY_ID AWS_SECRET_ACCESS_KEY AWS_SESSION_TOKEN < **Stale premise.** This step was written against the pre-fix world, where the -> plan was to declare `/run` in P1 and serve it in P2. F1 proved that is not a -> reachable service state, and the agent now serves `/ready` + `/run`. Two -> consequences for a re-run: (i) the image is no longer hook-less, so -> `--run-hook-payload` is ACCEPTED rather than rejected — the interesting -> observation becomes whether the payload reaches the pipeline, not what a -> hook-less VM does; and (ii) `--image-identifier` needs the image **ARN**, not -> `$IMAGE_NAME` (F3). What still holds unchanged, and is worth re-confirming, is -> everything below about the state enum, the 28,800-second bound, the -> omit-`idlePolicy` invariant, and the default public ingress. - -```bash -export RUN_STARTED_EPOCH="$(date +%s)" -aws lambda-microvms run-microvm \ - --image-identifier "$IMAGE_NAME" \ - --image-version "$IMAGE_VERSION" \ - --execution-role-arn "$EXECUTION_ROLE_ARN" \ - --egress-network-connectors "$EGRESS_CONNECTORS" \ - --run-hook-payload '{"verification":"issue-645-p1"}' \ - --maximum-duration-in-seconds 28800 \ - | tee "$EVIDENCE_DIR/run-microvm.json" -export MICROVM_ID="$(python3 -c 'import json,os; print(json.load(open(os.environ["EVIDENCE_DIR"] + "/run-microvm.json"))["microvmId"])')" -export MICROVM_ENDPOINT="$(python3 -c 'import json,os; print(json.load(open(os.environ["EVIDENCE_DIR"] + "/run-microvm.json"))["endpoint"])')" -``` - -Do **not** pass `--idle-policy`. Immediately poll: - -```bash -for delay in 0 2 5 10 20 30 60; do - sleep "$delay" - date -u '+%Y-%m-%dT%H:%M:%SZ' - aws lambda-microvms get-microvm --microvm-identifier "$MICROVM_ID" || true -done | tee "$EVIDENCE_DIR/hookless-state-timeline.txt" -``` - -**Expected:** `run-microvm` should return a `microvmId`, endpoint, image ARN, -version, and initial state if `/run` is asynchronous at control-plane return. -The eventual state is deliberately **not prescribed**: record whether it reaches -`RUNNING`, remains `PENDING`, becomes `TERMINATING/TERMINATED`, or disappears, -plus `stateReason` and elapsed time. If `run-microvm` itself rejects because the -hook returns 404/times out, record that exact exception and timing. This result -feeds the P2 hook/startup design. - -**Record:** full response, endpoint (not an auth token), all states/reasons, -time-to-first-state/time-to-terminal, and relevant log lines. - -**ADR-021 item:** real state enum, 28,800-second bound, and omit-`idlePolicy` -invariant. (The original "P1-not-runnable premise" this step was written to probe -no longer exists — see the Phase 5 row of the stale-instruction table.) - -### 5.2 Explicit `get-microvm` state mapping - -**Purpose:** validate the six SDK states used by strategy mapping. - -```bash -aws lambda-microvms get-microvm --microvm-identifier "$MICROVM_ID" \ - | tee "$EVIDENCE_DIR/get-microvm.json" -``` - -**Expected:** observed values come from `PENDING`, `RUNNING`, `SUSPENDING`, -`SUSPENDED`, `TERMINATING`, `TERMINATED`; a reaped ID returns -`ResourceNotFoundException`. - -**Record:** every distinct state actually observed and any unknown state. - -**ADR-021 item:** mechanical state mapping and future-enum safety premise. - -### 5.3 Manual suspend without `idlePolicy` - -**Purpose:** determine whether explicit suspend works independently of traffic -idle policy and whether a hook-less image survives long enough to suspend. - -Use the sandbox admin identity because P1's orchestrator correctly lacks this -permission: - -```bash -date -u '+%Y-%m-%dT%H:%M:%SZ suspend-request' -aws lambda-microvms suspend-microvm --microvm-identifier "$MICROVM_ID" \ - | tee "$EVIDENCE_DIR/suspend-microvm.json" -for i in 1 2 3 4 5 6; do - aws lambda-microvms get-microvm --microvm-identifier "$MICROVM_ID" || true - sleep 10 -done | tee "$EVIDENCE_DIR/suspend-state-timeline.txt" -``` - -**Expected:** if the VM reached a suspendible state, observe -`SUSPENDING → SUSPENDED`. A conflict/not-found caused by the failed `/run` hook -is a valid P1 result but means 5.4–5.6 cannot discharge TTL/resume empirically; -mark those **BLOCKED-BY-P1-HOOKLESS-IMAGE**, do not invent an answer. - -**Record:** identity, response/exception, states, and suspend latency. - -**ADR-021 item:** manual suspend without idle policy. - -### 5.4 Suspended TTL experiment - -**Purpose:** determine the default lifetime of a manually suspended VM when -`idlePolicy` (and therefore `suspendedDurationSeconds`) is omitted. - -Only run after observing `SUSPENDED`: - -```bash -for seconds in 0 900 3600 14400; do - sleep "$seconds" - printf '\ncheckpoint_after_sleep_seconds=%s at %s\n' "$seconds" "$(date -u +%Y-%m-%dT%H:%M:%SZ)" - aws lambda-microvms get-microvm --microvm-identifier "$MICROVM_ID" || true -done | tee "$EVIDENCE_DIR/suspended-ttl-checkpoints.txt" -``` - -The sleeps are incremental (0, then 15 min, then +1 h, then +4 h). A time-boxed -executor may truncate after 15 min or 1 h, but must say so. Never leave the VM -past the 28,800-second maximum; Phase 8 terminates it. - -**Expected:** unknown by design. Record whether it stays `SUSPENDED`, terminates, -or becomes NotFound and at what wall-clock age. `maximumDurationInSeconds=28800` -bounds the worst case even if no separate suspended TTL exists. - -**Record:** complete checkpoints, truncation, start age, and terminal time. - -**ADR-021 item:** P1 external fact — manual-suspend default TTL without -`idlePolicy`. - -### 5.5 Quota treatment while suspended - -**Purpose:** test the rationale that a suspended VM still holds account memory -quota. - -```bash -aws service-quotas list-services \ - --query "Services[?contains(ServiceName, 'Lambda')].[ServiceCode,ServiceName]" \ - --output table | tee "$EVIDENCE_DIR/lambda-service-codes.txt" -aws service-quotas list-service-quotas --service-code lambda \ - --query "Quotas[?contains(QuotaName, 'MicroVM') || contains(QuotaName, 'microVM') || contains(QuotaName, 'memory')]" \ - | tee "$EVIDENCE_DIR/microvm-service-quotas.json" -aws lambda-microvms list-microvms --image-identifier "$IMAGE_NAME" \ - | tee "$EVIDENCE_DIR/microvms-while-suspended.json" -``` - -Also open the Lambda MicroVM **account quota / memory utilization view** in the -AWS console before launch, while RUNNING, while SUSPENDED, and after termination; -capture values/timestamps. The SDK 3.1098 model has no “get account memory -usage” operation, and `service-quotas` normally reports limits rather than live -consumption. If the console has no utilization view and quota is too high to -safely saturate with a second 32 GiB VM, record **NOT OBSERVABLE SAFELY** rather -than launching VMs until failure. - -**Expected:** the suspended VM remains in `list-microvms`; the load-bearing -claim is discharged only if the account quota view continues counting its -32,768 MiB after `SUSPENDED`. - -**Record:** quota names/codes/limits, list response, and four utilization -snapshots. Distinguish “listed” from “proven to consume quota.” - -**ADR-021 item:** account-memory quota treatment of suspended VMs / concurrency -slot-held rationale. - -### 5.6 Resume and state transition - -**Purpose:** verify manual resume and preserved lifecycle identity. - -```bash -aws lambda-microvms resume-microvm --microvm-identifier "$MICROVM_ID" \ - | tee "$EVIDENCE_DIR/resume-microvm.json" -for i in 1 2 3 4 5 6; do - aws lambda-microvms get-microvm --microvm-identifier "$MICROVM_ID" || true - sleep 10 -done | tee "$EVIDENCE_DIR/resume-state-timeline.txt" -``` - -**Expected:** unknown for P1 because no `/resume` hook is declared and `/run` -may already have failed. Record whether resume succeeds and transitions to -`RUNNING`, conflicts, terminates, or disappears, and whether ID/endpoint remain -stable. - -**Record:** response, transitions, latency, ID/endpoint stability. - -**ADR-021 item:** empirical resume lifecycle input for P3 design. - -### 5.7 Terminate and observe reaping - -**Purpose:** validate active cleanup and the `NotFound → completed` strategy -mapping premise. - -```bash -date -u '+%Y-%m-%dT%H:%M:%SZ terminate-request' -aws lambda-microvms terminate-microvm --microvm-identifier "$MICROVM_ID" \ - | tee "$EVIDENCE_DIR/terminate-microvm.json" -for delay in 0 2 5 10 20 30 60 120; do - sleep "$delay" - date -u '+%Y-%m-%dT%H:%M:%SZ' - aws lambda-microvms get-microvm --microvm-identifier "$MICROVM_ID" || true -done | tee "$EVIDENCE_DIR/terminate-reap-timeline.txt" -``` - -**Expected:** `TERMINATING`/`TERMINATED` may be visible, followed by -`ResourceNotFoundException`. Exact reaping time is empirical. - -**Record:** transition and seconds from terminate request to NotFound, including -exception name/message/status code. - -**ADR-021 item:** explicit terminate path and `ResourceNotFoundException → -completed` mapping. - -### 5.8 IAM negative tests (only under assumed orchestrator role) - -**Purpose:** prove the deployed role cannot escape its exact image or mint JWE -tokens. - -Run only if step 5.0 successfully assumed the orchestrator role. Otherwise mark -**SKIPPED — LAMBDA ROLE TRUST DOES NOT ALLOW OPERATOR ASSUMPTION** and rely on -4.4's policy document. - -```bash -aws lambda-microvms run-microvm \ - --image-identifier "${IMAGE_NAME}-different" \ - --execution-role-arn "$EXECUTION_ROLE_ARN" \ - --egress-network-connectors "$EGRESS_CONNECTORS" \ - --maximum-duration-in-seconds 60 \ - 2>&1 | tee "$EVIDENCE_DIR/iam-negative-different-image.txt" - -aws lambda-microvms create-microvm-auth-token \ - --microvm-identifier "$MICROVM_ID" \ - --expiration-in-minutes 5 \ - --allowed-ports '[{"port":8080}]' \ - 2>&1 | tee "$EVIDENCE_DIR/iam-negative-auth-token.txt" -``` - -**Expected:** both return `AccessDeniedException`. If the different-name test -returns `ResourceNotFoundException`, it does not prove exact-ARN denial; create -or use a real second sandbox image only if already available, otherwise mark -that sub-check inconclusive. The `allowed-ports` CLI union syntax is -best-effort; the verified SDK request is -`{microvmIdentifier, expirationInMinutes, allowedPorts:[{port:8080}]}`. A local -argument-parser error is not an IAM result. - -**Record:** assumed caller ARN and full errors. - -**ADR-021 item:** exact-image ARN scoping and no -`CreateMicrovmAuthToken`/no-JWE posture. - -### 5.9 Direct service boundary: 16,384 vs 16,385 bytes - -**Purpose:** independently verify the documented `runHookPayload` service cap. -This is direct-service validation; the ABCA strategy already routes envelopes -larger than 16,384 bytes to S3. - -```bash -python3 - <<'PY' -from pathlib import Path -Path('/tmp/abca-payload-16384.txt').write_bytes(b'x' * 16384) -Path('/tmp/abca-payload-16385.txt').write_bytes(b'x' * 16385) -PY -wc -c /tmp/abca-payload-16384.txt /tmp/abca-payload-16385.txt - -aws lambda-microvms run-microvm \ - --image-identifier "$IMAGE_NAME" --image-version "$IMAGE_VERSION" \ - --execution-role-arn "$EXECUTION_ROLE_ARN" \ - --egress-network-connectors "$EGRESS_CONNECTORS" \ - --run-hook-payload file:///tmp/abca-payload-16384.txt \ - --maximum-duration-in-seconds 60 \ - | tee "$EVIDENCE_DIR/run-payload-16384.json" - -aws lambda-microvms run-microvm \ - --image-identifier "$IMAGE_NAME" --image-version "$IMAGE_VERSION" \ - --execution-role-arn "$EXECUTION_ROLE_ARN" \ - --egress-network-connectors "$EGRESS_CONNECTORS" \ - --run-hook-payload file:///tmp/abca-payload-16385.txt \ - --maximum-duration-in-seconds 60 \ - 2>&1 | tee "$EVIDENCE_DIR/run-payload-16385.txt" -``` - -Immediately terminate the MicroVM returned by the accepted request (if any). - -**Expected:** 16,384 bytes is accepted; 16,385 bytes is rejected, likely with -`ValidationException`. Record the actual exception rather than treating the -predicted name as normative. Confirm the installed CLI expands `file://` to file -contents; if it passes the literal URI, repeat with command substitution and -record the client behavior. - -**Record:** byte counts, both complete responses, exception name/message/status, -and cleanup ID. - -**ADR-021 item:** exact 16 KB `runHookPayload` boundary. - ---- - -## Phase 6 — CLI behavior - -### 6.1 Build and point the CLI at the stack - -**Purpose:** use the repository CLI's real resolution rules: operator commands -take `--region`/`--stack-name`; configured Region is a fallback; API commands -also need Cognito configuration/login. - -```bash -MISE_EXPERIMENTAL=1 mise //cli:build -bgagent() { node cli/lib/bin/bgagent.js "$@"; } -bgagent configure --stack-name "$STACK_NAME" --region "$AWS_REGION" -bgagent platform outputs --stack-name "$STACK_NAME" --region "$AWS_REGION" -``` - -**Expected:** config is written under `${BGAGENT_CONFIG_DIR:-$HOME/.bgagent}`; -stack outputs resolve. No Cognito login is needed for operator AWS commands. - -**Record:** CLI build result and redacted outputs. - -**ADR-021 item:** deploy/CLI substrate discovery contract. - -### 6.2 Onboard, probe, inspect, and clean up a dummy row - -**Purpose:** exercise the live managed-image probe, doctor check, and runtime -grouping without submitting a task. - -```bash -export DUMMY_REPO=verification-only/issue-645 -bgagent repo onboard "$DUMMY_REPO" \ - --compute-type lambda-microvm \ - --stack-name "$STACK_NAME" --region "$AWS_REGION" \ - --output json | tee "$EVIDENCE_DIR/cli-onboard.json" - -bgagent platform doctor --stack-name "$STACK_NAME" --region "$AWS_REGION" \ - --output json | tee "$EVIDENCE_DIR/cli-doctor.json" || true -bgagent runtime status --stack-name "$STACK_NAME" --region "$AWS_REGION" \ - --output json | tee "$EVIDENCE_DIR/cli-runtime-status.json" - -bgagent repo offboard "$DUMMY_REPO" \ - --stack-name "$STACK_NAME" --region "$AWS_REGION" \ - --output json | tee "$EVIDENCE_DIR/cli-offboard.json" -``` - -**Expected:** onboarding's `ListManagedMicrovmImages` probe passes and the row -has `compute_type=lambda-microvm`; doctor contains -`lambda_microvm_availability`; runtime status groups it under -`lambda_microvm_substrates`; offboard marks it removed. Doctor may still exit -non-zero in a virgin sandbox because its GitHub secret is an unpopulated -placeholder or no real repo/token/model access exists—record those independent -failures. - -**SKIP-IF-UNCONFIGURED:** if the sandbox principal lacks DDB or MicroVM catalog -read permission, record the IAM gap and skip the write. Onboarding itself does -not require a valid GitHub token, but a meaningful doctor pass and any task do. - -**Record:** probe result, row, doctor checks, grouping, and cleanup row status. - -**ADR-021 item:** onboarding region probe and doctor availability check. - ---- - -## Phase 7 — OPTIONAL negative task path - -### 7.1 Submit only in a fully configured sandbox - -**Purpose:** observe orchestrator classification and terminate-on-finalize, not -to claim P2 smoke parity. - -This is **OPTIONAL / usually DEFERRED-TO-P2-ENV**. A virgin account is missing a -Cognito user/login, populated GitHub token secret, and a genuinely accessible -onboarded repository unless the executor configures all three. Do not create -those merely for this P1 substrate pass. - -If they already exist: - -```bash -bgagent login -bgagent repo onboard --compute-type lambda-microvm \ - --stack-name "$STACK_NAME" --region "$AWS_REGION" -bgagent submit --repo --task "P1 negative: observe hook-less MicroVM failure" -# Use the returned task ID: -bgagent watch -``` - -**Expected:** the backend does not complete agent work. Capture the task's -failure classification/remedy, persisted `compute_metadata` (`microvmId` and -endpoint), orchestrator logs, and whether finalization calls -`TerminateMicrovm`. If no complete setup exists, write -`DEFERRED-TO-P2-ENV — missing Cognito user/login, GitHub token, and/or real repo -onboarding`. - -**Record:** task ID/evidence or exact deferral reason. - -**ADR-021 item:** defense-in-depth failure classification, handle persistence, -and terminate fire; this is not P2 clone→change→PR smoke parity. - ---- - -## Phase 8 — Teardown - -### 8.1 Terminate every MicroVM - -**Purpose:** stop compute billing before image/stack deletion. - -```bash -aws lambda-microvms list-microvms --image-identifier "$IMAGE_NAME" \ - | tee "$EVIDENCE_DIR/microvms-before-teardown.json" -for id in $(aws lambda-microvms list-microvms --image-identifier "$IMAGE_NAME" \ - --query 'items[].microvmId' --output text); do - aws lambda-microvms terminate-microvm --microvm-identifier "$id" || true -done -sleep 30 -aws lambda-microvms list-microvms --image-identifier "$IMAGE_NAME" \ - | tee "$EVIDENCE_DIR/microvms-after-terminate.json" -``` - -**Expected:** no nonterminal VM remains; wait/retry if necessary. - -**Record:** IDs and final states. - -**ADR-021 item:** explicit cleanup rather than relying on the eight-hour bound. - -### 8.2 Delete out-of-band image versions and image - -**Purpose:** stop snapshot storage charges. The verified operation/CLI command -name is `delete-microvm-image-version` with `imageIdentifier` and -`imageVersion`. - -```bash -aws lambda-microvms list-microvm-image-versions --image-identifier "$IMAGE_NAME" \ - | tee "$EVIDENCE_DIR/image-versions-before-delete.json" -for version in $(aws lambda-microvms list-microvm-image-versions \ - --image-identifier "$IMAGE_NAME" --query 'items[].imageVersion' --output text); do - aws lambda-microvms delete-microvm-image-version \ - --image-identifier "$IMAGE_NAME" --image-version "$version" -done -aws lambda-microvms delete-microvm-image --image-identifier "$IMAGE_NAME" -``` - -**Expected:** all versions enter deletion and the image is deleted. If the -service requires deleting the parent first/last or waiting between operations, -follow the returned conflict remedy and record actual order. - -**Record:** versions, responses, final NotFound/list absence. - -**ADR-021 item:** versioned image lifecycle cleanup. - -### 8.3 Destroy ABCA; retain bootstrap - -**Purpose:** remove the platform and recurring network costs while preserving -the account's reusable CDK bootstrap. - -```bash -MISE_EXPERIMENTAL=1 mise //cdk:destroy -- \ - "$STACK_NAME" --force \ - --context compute_type=lambda-microvm \ - --context microvm_image_identifier="$IMAGE_NAME" \ - --context microvm_image_version="$IMAGE_VERSION" -aws cloudformation wait stack-delete-complete --stack-name "$STACK_NAME" -aws cloudformation describe-stacks --stack-name CDKToolkit \ - --query 'Stacks[0].[StackStatus,Parameters]' \ - | tee "$EVIDENCE_DIR/bootstrap-left-in-place.json" -``` - -**Raw fallback if `mise` is absent** (run from `cdk/`): - -```bash -npx cdk destroy "$STACK_NAME" --force \ - --context compute_type=lambda-microvm \ - --context microvm_image_identifier="$IMAGE_NAME" \ - --context microvm_image_version="$IMAGE_VERSION" -``` - -**Expected:** `backgroundagent-dev` is absent; `CDKToolkit` remains -`CREATE_COMPLETE`/`UPDATE_COMPLETE` with the custom policies. VPC teardown can -lag while service-managed ENIs are reclaimed; wait and retry rather than -force-deleting resources past CloudFormation. - -**Record:** destroy duration/events, leftovers, and final bootstrap status. - -**ADR-021 item:** clean substrate/resource lifecycle. - -**Cost note:** this pass can accrue MicroVM running minutes, snapshot/image -storage, S3 artifact storage, NAT gateway hourly/data charges, VPC endpoint -hourly charges, CloudWatch logs, and brief Lambda/DynamoDB/API usage. Suspended -VMs should stop compute charges but may retain billed snapshot storage and -account memory quota; the experiment determines the latter. NAT gateways and -endpoints continue charging until stack deletion. - ---- - -## Live execution results — 2026-07-31, account , us-east-1 - -Executed against real AWS. Evidence directory: -`/tmp/abca-645-p1-20260731T184822Z`. Wall clock 18:48Z → 23:07Z (4 h 19 min). - -**Execution deviations from the runbook as written** (each is itself a result): - -1. `mise` is **not installed** on the executor; every **raw fallback** was used. -2. The account is **not virgin overall** — `serverless-api-powertools`, - `BuildingServerlessAPIs`, `aws-sam-cli-managed-default`, and `CDKToolkit` - pre-existed. `backgroundagent-dev` was absent, so the ABCA-specific - first-deploy premise held. -3. `docker` is absent; `finch` 1.x (`CDK_DOCKER=finch`) built the AgentCore - container asset. -4. Phase 2 could not deploy from unmodified sources. The stack was deployed from - a **hand-patched cloud assembly** (`/tmp/cdkout-p1*`, a build artifact — no - repository source file was modified). Two patches, both forced by live-service - rejections recorded below: a MicroVM connector **operator role**, and moving - subnets off `us-east-1a`. -5. A temporary **port-80 egress rule** (`sgr-07ed1fa48ef38467a`) was added to the - construct's security group to get any image to build at all (see 4.3). -6. Suspend-TTL observation was **truncated at ~1 h** (runbook allows this). - -### Phase 0 - -**0.1** — Account ``, caller -`arn:aws:sts:::assumed-role/AdminConsoleAccess/aamorosi-Isengard` -(administrator). Branch `feat/645-lambda-microvm-p1`, SHA -`0505f914fd7093cccd067b6346a24e1c40e50643`. Untracked: `docs/verification/`, -`opencode.json`. Stack absence returned exactly the expected error: - -``` -An error occurred (ValidationError) when calling the DescribeStacks operation: Stack with id backgroundagent-dev does not exist -``` - -**0.2** — `aws-cli/2.36.13 Python/3.14.6 Darwin/25.5.0 source/arm64`; -`node v24.16.0`; `cdk 2.1129.0`; `mise NOT INSTALLED`; `Python 3.9.6`; -`Zip 3.0`; **`openrsync` (protocol 29, "rsync 2.6.9 compatible")** — the -packaging script's `rsync -a --exclude` usage worked unmodified on macOS. -`@aws-sdk/client-lambda-microvms` = **3.1098.0**. - -`aws lambda-microvms help` **succeeded (exit 0)** and lists **24** commands -including all seven the runbook requires. The CLI command list is an **exact -match** to the SDK 3.1098.0 command list (24 vs 24), including -`create-microvm-shell-auth-token`. **No CLI/SDK operation-name drift.** - -`create-microvm-image --generate-cli-skeleton input` **confirms the packaging -script's shape and refutes the CDK L1 shape**: - -- `cpuConfigurations[].architecture` — help documents exactly one allowed value: - `ARM_64`. The L1's `'arm64'` is not a documented value. -- `hooks` — `{"port": integer, "microvmHooks": {"run": "DISABLED"|"ENABLED", - "runTimeoutInSeconds": integer, ...}}`. There is **no hook-path field at all**; - the L1's `run: '/run'` path string has no counterpart in the service model. -- `run-microvm` skeleton confirms `idlePolicy` - `{maxIdleDurationSeconds, suspendedDurationSeconds, autoResumeEnabled}` and - `runHookPayload` as a plain string. - -The CFN-vs-API question is therefore **half-adjudicated**: the API side is -settled, but the `AWS::Lambda::MicrovmImage` CFN path was never exercised, -because the construct only synthesizes it when `microvm_base_image_arn` + -`microvm_base_image_version` context is supplied, and the runbook's Phase 4 uses -the out-of-band script path. **The CFN value shapes remain untested** — and 4.2 -below shows the request would be rejected on hook semantics regardless of shape. - -### Phase 1 - -**1.1** — `CDKToolkit` `CREATE_COMPLETE`, created 2025-11-24, `BootstrapVariant` -= `AWS CDK: Default Resources`. Markers: `ComputeTypes` **ABSENT**, -`IaCRoleABCAComputeLambdaMicrovms` **ABSENT**, `AdministratorAccess` -**present** — exactly the standard-bootstrap starting state the step predicts. - -**1.2 — DEFECT (runbook + `mise //cdk:bootstrap`): the re-bootstrap is a silent -no-op on an already-bootstrapped account.** Verbatim: - -``` -Bootstrap stack already exists, containing 'AWS CDK: Default Resources'. Not overwriting it with a template containing 'ABCA: Least-Privilege Bootstrap' (use --force if you intend to overwrite) -✅ Environment aws:///us-east-1 bootstrapped (no changes). -``` - -Exit status **0**. A pass that trusts this would proceed to 1.3 with -`AdministratorAccess` still attached. `--force` was required. A **durable** -consequence: after the forced bootstrap, `BootstrapVariant` **remains** -`AWS CDK: Default Resources` (the CDK CLI re-sends the previous value rather than -the template default `ABCA: Least-Privilege Bootstrap`), so **every future -non-forced `mise //cdk:bootstrap` will refuse again**. - -**1.2 — DEFECT (runbook): the `ComputeTypes` parameter command as written is -rejected.** Verbatim: - -``` -An error occurred (ParamValidation): Parameter validation failed: -Invalid type for parameter Parameters[0].ParameterValue, value: ['agentcore', 'lambda-microvm'], type: , valid types: -``` - -The CLI shorthand parser splits on the comma. The escaped form -`ParameterValue=agentcore\,lambda-microvm` works — which is exactly what the -comment above `[tasks.bootstrap]` in `cdk/mise.toml` already shows; the runbook -dropped the escapes. Final state: `UPDATE_COMPLETE`, `ComputeTypes` = -`agentcore,lambda-microvm`. - -**1.3 — PASS.** `CFN_EXEC_ROLE` = -`cdk-hnb659fds-cfn-exec-role--us-east-1`. Exactly one local policy -matched: `cdk-hnb659fds-IaCRole-ABCA-Compute-LambdaMicrovms--us-east-1`, -and it is attached. Attachments are the five ABCA policies (Application, -Infrastructure, Observability, Compute-Agentcore, Compute-LambdaMicrovms) and -**no `AdministratorAccess`** — the template's replacement works. The -`LambdaMicrovms` statement grants 19 actions, all image/version/build/ -managed-catalog/network-connector, including `lambda:PassNetworkConnector`. -The optional scratch-qualifier negative was deliberately skipped as the runbook -directs. - -### Phase 2 - -**2.1 — synth PASS, deploy BLOCKED THREE TIMES.** Synth emitted -`abca:microvm-image-not-provisioned` and **not** -`abca:microvm-image-p1-not-runnable`, exactly as specified. Incidental synth -warnings worth noting: `Template size is approaching limit: 893273/1000000` and -`Number of resources: 463 is approaching allowed maximum of 500`. - -*Blocker A (environmental, not an ABCA defect).* The stack carries one Docker -image asset (`agent/Dockerfile`) and no container builder was installed. With -`finch`, the `gh-builder` stage failed twice on upstream flakiness: - -``` -pkg/mod/github.com/cli/go-gh/v2@v2.13.0/internal/yamlmap/yaml_map.go:8:2: unrecognized import path "gopkg.in/yaml.v3": reading https://gopkg.in/yaml.v3?go-get=1: 502 Proxy Error -``` - -`gopkg.in` alternated 200/502 from the host too. A direct `finch build` then -succeeded. **Data point for 4.3's size narrative:** the AgentCore container -image is **1.799 GB uncompressed / 629.7 MB compressed**. - -*Blocker B — **the P1 substrate cannot deploy from unmodified sources**.* -`AWS::Lambda::NetworkConnector` `CREATE_FAILED`, verbatim: - -``` -Resource handler returned message: "NetworkConnectorOperatorRole is required for VPC_EGRESS connector type (Service: Lambda, Status Code: 400, Request ID: 04726267-6c61-4ff5-bb1d-302122e9f955) (SDK Attempt Count: 1)" (RequestToken: 1a1652c3-2166-e865-c49d-6cdb5927bbfe, HandlerErrorCode: InvalidRequest) -``` - -This **directly refutes an explicit design assumption** stated in -`cdk/src/constructs/lambda-microvm-compute.ts` (~line 467): - -> `operatorRole` is left unset so Lambda manages the ENIs with its own -> service-linked role rather than a role we would have to trust. - -The generated L1 also marks `operatorRole` optional -(`readonly operatorRole?: string`) with no note that `VPC_EGRESS` requires it. -An independent probe stack (`abca645-connector-probe`) confirmed the minimal -working recipe: a role trusting `lambda.amazonaws.com` with -`AWSLambdaVPCAccessExecutionRole` plus `ec2:CreateNetworkInterface` / -`DeleteNetworkInterface` / `DescribeNetworkInterfaces` / `DescribeSubnets` / -`DescribeVpcs` / `DescribeSecurityGroups` / `CreateTags` / -`AssignPrivateIpAddresses` / `UnassignPrivateIpAddresses` / -`Describe|ModifyNetworkInterfaceAttribute` → connector `CREATE_COMPLETE`. - -*Blocker C (AgentCore, account-AZ-specific, blocks any ABCA deploy here).* - -``` -Resource handler returned message: "Agent runtime creation failed with status: CREATE_FAILED for runtime: backgroundagentdevRuntimeCC6E3A5A-yiKm9OEVPo. Reason: The following subnets are in unsupported availability zones in region us-east-1: subnet-02b0221802f3fee10 in us-east-1a (ID: use1-az6). Supported availability zones are: use1-az4, use1-az1, use1-az2" -``` - -This account maps `us-east-1a` → `use1-az6`. `AgentVpc` does not constrain AZ -selection, so CDK's default two-AZ pick lands on an AZ AgentCore rejects. -Patched `us-east-1a` → `us-east-1c` (`use1-az2`). - -*Two further teardown/iteration gotchas.* (i) Rollback itself failed once: -`Validation failed during DeleteMemory: Memory is in transitional state -CREATING. Cannot delete memory.` — `AWS::BedrockAgentCore::Memory` cannot be -deleted while creating, leaving `ROLLBACK_FAILED`; a plain `delete-stack` -cleared it. (ii) Post-synth template edits are **silently ignored** if the -template's S3 asset object already exists: the object key is the pre-edit content -hash recorded in `*.assets.json`, so `cdk-assets` skips the upload and CFN -re-uses the stale template. The stale object must be deleted. - -Successful deploy: `CREATE_COMPLETE`, 19:41:53Z → 19:55:35Z = **13 min 42 s**, -464 resources. - -**2.2 — PASS.** `ComputeSubstrate=lambda-microvm`; all six `Microvm…` outputs -populated; artifact key exactly `microvm-images/agent-artifact.zip`. The 13 -`LambdaMicrovmCompute` resources are: artifact + payload buckets (each with a -bucket policy and an auto-delete custom resource), build role + policy, execution -role + policy, `AWS::EC2::SecurityGroup sg-0e662dc0d6f6e9ade`, -`AWS::Logs::LogGroup /aws/lambda-microvms/backgroundagent-dev-abca-agent`, and -`AWS::Lambda::NetworkConnector nc-132ede11-cb63-4dfa-b75b-6a4713023c1a`. -**No `AWS::Lambda::MicrovmImage`** — correct for this state. The security group -has exactly one rule: egress `tcp/443 → 0.0.0.0/0`, *"Allow HTTPS egress (GitHub -API, AWS services)"*. (That single rule is what breaks the image build — 4.3.) - -**2.3 — PASS.** `MICROVM_*` keys = `[]` (14 env keys total). -**DEFECT (runbook): the `ORCHESTRATOR_FN` command is broken by pagination.** -With 464 resources, `list-stack-resources --query "…|[0]"` applies the query -**per page** and printed five lines (`None None None None `), which then -failed `get-function-configuration` on the multi-line value. Needs -`--no-paginate` (or local parsing). Resolved value: -`backgroundagent-dev-TaskOrchestratorOrchestratorFn-gM2sydgNVf1V`. - -**2.4 — PASS.** All **six** taggable construct resources carry -`abca:compute-backend=lambda-microvm`: security group, network connector, log -group, both buckets, and — verified via `iam list-role-tags` — both roles. -`resourcegroupstaggingapi` returned only **5** ARNs; **IAM roles are simply not -returned by that API**, which is an API coverage gap, not a missing tag. -Bucket policies and auto-delete custom resources are not independently taggable, -as the step anticipates. - -### Phase 3 - -**3.1 — PASS** (exit 1). Verbatim: - -``` -Error: AWS Lambda MicroVMs are not available in eu-central-1. The lambda-microvm compute backend is enabled (--context compute_type=lambda-microvm) but the stack Region is not one of: us-east-1, us-east-2, us-west-2, eu-west-1, ap-northeast-1. Either deploy the stack into a supported Region, drop the backend (--context compute_type=agentcore or ecs), or — if AWS has since launched Lambda MicroVMs in eu-central-1 — bypass this static check with --context microvm_region_override=true and add eu-central-1 to LAMBDA_MICROVM_SUPPORTED_REGIONS in cdk/src/handlers/shared/microvm-regions.ts. -``` - -Names the Region, all five supported Regions, and the override flag. - -**3.2 — PASS** (exit 0) with `abca:microvm-region-override` (and, correctly, the -`abca:microvm-image-not-provisioned` warning still present). - -### Phase 4 - -**4.1 — PASS, with a runbook selector bug.** Exactly **one** managed base image -exists in us-east-1: `arn:aws:lambda:us-east-1:aws:microvm-image:al2023-1`, with -versions `1` (created 2026-07-21) and `0` (2026-06-17). **Ordering is -newest-first**, so the runbook's `items[-1].imageVersion` returns **`0`** (the -older version); `items[0]` returns `1`. The runbook's own warning about ordering -is therefore *load-bearing here*. Selected `1`; the service echoes it back as -`baseImageVersion: "1.0"`. - -**4.2 — Artifact plane PASS; image creation FAILED TWICE on service -validation.** The script read all five outputs, staged `Dockerfile` + `agent/` + -`contracts/`, printed **`584K artifact`**, and uploaded successfully (S3 -`ContentLength` **597305**, SSE `AES256`). - -Then, on the exact request the script builds, verbatim: - -``` -An error occurred (ValidationException) when calling the CreateMicrovmImage operation: The ready (/ready) MicroVM image hook must be enabled when any MicroVM lifecycle hook (run, resume, suspend, or terminate) is enabled. The ready hook signals when the application has finished initializing so the snapshot is taken in a ready state. -``` - -This **refutes the P1 hook-phasing plan directly**. The construct comments state -`/ready` and `/validate` are *"omitted in P1 because the agent does not implement -them yet: configuring a `/validate` endpoint that 404s would fail every image -build"* — but the service **will not accept `run: ENABLED` without -`ready: ENABLED`**. "Declare `/run` in P1, serve it in P2" is not a reachable -state. - -Consequences of that failure: **the conspicuous "P1 image is NOT runnable end to -end" banner was never printed**, because the script's banner heredoc comes after -the `create-microvm-image` call. Also, the runbook's `2>&1 | tee` pipeline -reported `EXIT=0` while the script had failed — the tee status masks it. - -Retrying with `/ready` enabled surfaced the **second** rejection: - -``` -An error occurred (ValidationException) when calling the CreateMicrovmImage operation: The requested memory size of 32768 MiB is not supported by base MicroVM image arn:aws:lambda:us-east-1:aws:microvm-image:al2023-1. Supported memory sizes in MiB are: [512, 1024, 2048, 4096, 8192]. -``` - -The script's and construct's `32768` MiB (`DEFAULT_MINIMUM_MEMORY_MIB`, -documented as *"the service ceiling"*) is **not accepted**. The ceiling for this -base image is **8192 MiB (8 GiB)**, a quarter of what ADR-021 claims. - -**4.3 — Image ACTIVE only after two more corrections; the P1 hook shape is -UNBUILDABLE.** - -*Attempt 1* (hooks omitted, 8192 MiB): create succeeded, returning -`imageVersion: "1.0"` — **not `1`**, contradicting the runbook's -`IMAGE_VERSION=1` and the script's `--image-version 1` guidance. Both builds then -**FAILED** with `stateReason: "The container image build failed."` The exact -root cause, from `/aws/lambda-microvms/backgroundagent-dev-abca-agent`: - -``` -Could not connect to deb.debian.org:80 (146.75.38.132), connection timed out -E: Unable to locate package curl -E: Unable to locate package git -E: Unable to locate package build-essential -E: Package 'gnupg' has no installation candidate -ERROR: process "/bin/sh -c apt-get update && ... apt-get install -y --no-install-recommends curl git build-essential ca-certificates gnupg ..." did not complete successfully: exit code: 100 -``` - -**The construct's 443-only security group makes the agent image unbuildable.** -`agent/Dockerfile` runs `apt-get`, which uses HTTP on **port 80**; DNS resolution -worked, so only the port is the problem. Adding a temporary port-80 egress rule -(`sgr-07ed1fa48ef38467a`) fixed it immediately. - -*Attempt 2* (`update-microvm-image`, port 80 open) produced version **`2.0`**, -both builds `SUCCESSFUL`, `state=SUCCESSFUL`, `status=ACTIVE` in -20:11:18Z → 20:17:09Z = **5 min 51 s**. - -Further API-shape observations: - -- **Two builds per image version**, one per `chipsetGeneration` (`3` and `4`, - `chipset: GRAVITON`). The runbook's `items[0].buildState` inspects only one. -- `list-microvm-image-builds --image-identifier ` → - `ValidationException: Invalid ARN format: backgroundagent-dev-abca-agent`. - **An ARN is required.** `--image-version` accepts either `1` or `1.0`. -- `snapshotBuild` is returned by **`get-microvm-image-build`**, not by - `get-microvm-image-version` (which returned `snapshotBuild: null`). - -Sizes (`buildId 2033b7d8-1aa2-44a4-b174-3fc4bffebcea`, GRAVITON gen 4): - -| Field | Bytes | Human | -|---|---|---| -| `codeInstallSizeInBytes` | 2,334,748,672 | **2.17 GiB** | -| `memorySnapshotSizeInBytes` | 1,216,577,536 | 1.13 GiB | -| `diskSnapshotSizeInBytes` | 37,089,280 | 35.4 MiB | - -**Code-install size (2.17 GiB) exceeds AgentCore's 2 GB container-image limit**, -while the equivalent OCI image built locally was 1.799 GB (629.7 MB compressed). -So the same agent tree is *over* the AgentCore ceiling when measured as MicroVM -code-install and *under* it as an OCI image — the two are not interchangeable -measures, and the ADR narrative should say which one it means. Memory and disk -snapshots are deliberately **not** summed into that comparison. - -**Disk capacity: NOT EXPOSED.** No disk quota appears in `service-quotas` -(full list under 5.5) and no image/version field reports disk capacity. The -32 GB disk claim remains unverified; the 32 GB *memory* claim is refuted (8 GiB). - -*The decisive experiment.* A second image (`…-abca-agent-hooks`) was created with -the **exact P1 hook shape plus the service-mandated `/ready`**, and rebuilt after -port 80 was open. Both builds **FAILED**: - -``` -Ready hook check failed: the application returned a client error (HTTP 4xx) response -``` - -The agent **does** answer on port 8080 (an HTTP 4xx, not a connection failure), -but does not implement `/ready`. Combined with 4.2: **a P1 image that declares -`/run` cannot be built at all**, and an image that omits hooks cannot receive a -`runHookPayload` (5.1). P1 as specified is not merely "not runnable end to end" — -its image is **not creatable**. - -**4.4 — PASS, exactly as designed.** `UPDATE_COMPLETE`. Synth emitted -`abca:microvm-image-p1-not-runnable` with the full expected text. Orchestrator -env is exactly the five variables and **no ingress variable**: - -``` -MICROVM_EGRESS_CONNECTOR_ARNS = arn:aws:lambda:us-east-1::network-connector:nc-132ede11-cb63-4dfa-b75b-6a4713023c1a -MICROVM_EXECUTION_ROLE_ARN = arn:aws:iam:::role/backgroundagent-dev-LambdaMicrovmComputeExecutionRo-ZJu8Y1ybJt1N -MICROVM_IMAGE_IDENTIFIER = backgroundagent-dev-abca-agent -MICROVM_IMAGE_VERSION = 1.0 -MICROVM_PAYLOAD_BUCKET = backgroundagent-dev-lambdamicrovmcomputepayloadbuc-en08fimmvu6h -``` - -Inline policy `TaskOrchestratorOrchestratorFnServiceRoleDefaultPolicyDECF0D43`: - -- `Sid: MicrovmLifecycle` — exactly `lambda:RunMicrovm`, `lambda:GetMicrovm`, - `lambda:TerminateMicrovm` on exactly - `arn:aws:lambda:us-east-1::microvm-image:backgroundagent-dev-abca-agent` - and `…:backgroundagent-dev-abca-agent:*`. **Observed image ARN format matches** - the Service Authorization Reference pattern the construct derives. -- `Sid: MicrovmPassNetworkConnector` — `lambda:PassNetworkConnector` on `*`. -- `Sid: MicrovmPassExecutionRole` — `iam:PassRole` scoped to the execution role - with `iam:PassedToService = lambda.amazonaws.com`. -- **Zero** `SuspendMicrovm` / `ResumeMicrovm` / `CreateMicrovmAuthToken` actions - anywhere in the role (only attached managed policy is - `AWSLambdaBasicDurableExecutionRolePolicy`). - -**Critical mismatch:** `MICROVM_IMAGE_IDENTIFIER` is a **bare name**, and -`RunMicrovm` rejects bare names (5.1). The construct's comment asserts the -opposite — *"`imageIdentifier` may legitimately be a bare image NAME (that is -what `create-microvm-image --name` returns and what `run-microvm ---image-identifier` accepts)"*. `lambda-microvm-strategy.ts:236` passes that env -var straight through, so the P1 orchestrator would fail at `RunMicrovm`. - -### Phase 5 - -**5.0 — ADMIN — advisory.** AssumeRole denied, verbatim: - -``` -An error occurred (AccessDenied) when calling the AssumeRole operation: User: arn:aws:sts:::assumed-role/AdminConsoleAccess/aamorosi-Isengard is not authorized to perform: sts:AssumeRole on resource: arn:aws:iam:::role/backgroundagent-dev-TaskOrchestratorOrchestratorFnS-Bd7rBa2V6Jwf -``` - -Trust policy allows only `Service: lambda.amazonaws.com`. Trust was **not** -modified. All Phase 5 lifecycle evidence is therefore **admin-identity**; -4.4's static policy inspection remains the authoritative scoping evidence. - -**5.1 — Hook-less behaviour MEASURED: the VM runs and stays running.** -Three requests, in order: - -1. Bare name → `ValidationException: Malformed ARN - doesn't start with 'arn:'` -2. ARN + `--run-hook-payload` on the hook-less image → - `ValidationException: The run hook must be enabled in the MicroVM image to pass the run hook payload` -3. ARN, no payload → **success**: - -```json -{ "microvmId": "microvm-b44b69d9-f23b-30d1-97d1-a4ac558cfb5c", - "state": "PENDING", - "endpoint": ".lambda-microvm.us-east-1.on.aws", - "imageArn": "arn:aws:lambda:us-east-1::microvm-image:backgroundagent-dev-abca-agent", - "imageVersion": "2.0", - "maximumDurationInSeconds": 28800, - "ingressNetworkConnectors": ["arn:aws:lambda:us-east-1:aws:network-connector:aws-network-connector:HTTP_INGRESS"], - "egressNetworkConnectors": ["arn:aws:lambda:us-east-1::network-connector:nc-132ede11-cb63-4dfa-b75b-6a4713023c1a"] } -``` - -`maximumDurationInSeconds=28800` accepted; `idlePolicy` omitted as required. -Timeline: `RUNNING` at **+12 s**, and still `RUNNING` at +15/+21/+32/+53/+84/ -**+145 s**, `stateReason` always `None`. It did **not** terminate, stall in -`PENDING`, or disappear. A hook-less MicroVM is a healthy idle VM that bills. - -**Security finding — unrequested public ingress.** The service **auto-attached** -`arn:aws:lambda:us-east-1:aws:network-connector:aws-network-connector:HTTP_INGRESS` -even though `--ingress-network-connectors` was never passed, and returned a -public `*.lambda-microvm.us-east-1.on.aws` endpoint. ADR-021's "no ingress" -posture is not what the service defaults to; it is a *default-on* public HTTP -ingress that P1 neither requests nor suppresses. - -**5.2 — five of six states observed.** `PENDING`, `RUNNING`, `SUSPENDED`, -`TERMINATING`, `TERMINATED`. **`SUSPENDING` was never observable** — suspend -reached `SUSPENDED` in under 1 s. No unknown/unmapped state appeared. -`ResourceNotFoundException` for a reaped ID was **not** reached (5.7). - -**5.3 — PASS.** Admin `suspend-microvm` on a hook-less image with no -`idlePolicy`: **empty response body**, `SUSPENDED` at **+1 s**, stable across -+12/+23/+34/+44/+55 s. Explicit suspend works independently of any idle policy, -and a hook-less image is perfectly suspendible. - -**5.4 — Suspended TTL: survived the full observation window; TRUNCATED at ~1 h.** -Suspended at 20:21:12Z. `maximumDurationInSeconds` was 28800 and `startedAt` -stayed 20:18:05Z throughout. - -| Checkpoint | Wall clock | Suspended age | `get-microvm` state | `list-microvms` | Account quota view | -|---|---|---|---|---|---| -| start | 20:21:13Z | 1 s | `SUSPENDED` | present | `L-CD1C0CC4` = 1024 GB (limit only) | -| +15 min | 20:36:27Z | 915 s | `SUSPENDED` | present | 1024 GB (unchanged) | -| +45 min | 21:06:12Z | 2700 s | `SUSPENDED` | present | 1024 GB (unchanged) | -| ~+1 h (cap) | 21:21:29Z | 3617 s | `SUSPENDED` | present | 1024 GB (unchanged) | - -**No suspended TTL was observed within 1 h.** The runbook's 4-hour checkpoint was -**not** run (time-boxed, as the runbook permits). So the answer is bounded, not -final: *a manually suspended MicroVM with no `idlePolicy` survives at least -1 h 0 min 17 s (3617 s)*; whether a TTL exists between 1 h and the 8 h -`maximumDurationInSeconds` bound is **still open**. The VM was terminated in -Phase 8 rather than left to expire. - -**5.5 — Suspended VM stays listed; quota consumption NOT PROVABLE.** -`service-quotas list-service-quotas --service-code lambda` returned 24 MicroVM -quotas. The load-bearing ones: - -| Code | Name | Value | -|---|---|---| -| `L-CD1C0CC4` | Max allocated memory | **1024 Gigabytes** | -| `L-B430C318` | Max Execution Duration of a MicroVM (in Hours) | **8** | -| `L-942E56BE` | Number of MicroVM images | 100 | -| `L-F8BECE9C` | Versions per MicroVM Image | 50 | -| `L-72E0D058` | Number of concurrent MicroVM image builds | 10 | -| `L-535CA9B6` / `L-91B95582` | Rate / burst of `RunMicrovm` | 5 / 5 | -| `L-90045317` / `L-139F9A48` | Rate / burst of `SuspendMicrovm` | 2 / 2 | -| `L-118C44B3` / `L-25EEC0A4` | Rate / burst of `ResumeMicrovm` | 5 / 5 | -| `L-74787B8A` / `L-2CCA0501` | Rate / burst of `TerminateMicrovm` | 10 / 10 | -| `L-7712260B` / `L-D65D9F16` | Rate / burst of `CreateMicrovmAuthToken` | 50 / 50 | -| `L-772D8D8F` … `L-19741F6D` | Concurrent connections per 1/2/4/8/16 vCPU MicroVM | 8 / 16 / 32 / 64 / 128 | - -`L-B430C318 = 8 hours` independently confirms the 28,800-second bound. There is -**no disk quota at all**, corroborating "disk capacity not exposed". -`L-CD1C0CC4` is `QuotaAppliedAtLevel: ACCOUNT`, described as *"The maximum amount -of memory that can be allocated across all MicroVMs per account per region. -Customers can burst up to 4x this limit."* - -**Utilization is not observable.** `get-service-quota` returns **no -`UsageMetric`** for `L-CD1C0CC4`; `AWS/Usage` exposes only `CallCount` per API -name (`RunMicrovm`, `GetMicrovm`, `CreateMicrovmImage`, …) and **no memory -metric**; there is no MicroVM metric in `AWS/Lambda` and no MicroVM CloudWatch -namespace. Saturating the quota would require ~128 × 8 GiB VMs. - -**Verdict: LISTED, NOT PROVEN.** The suspended VM remained in `list-microvms` at -every checkpoint, but the claim *"a suspended VM still holds account memory -quota"* is **NOT OBSERVABLE SAFELY** in this account and remains undischarged. -Note also that the rationale was written for 32,768 MiB per VM against this -1024 GB quota; at the real 8192 MiB ceiling the arithmetic changes by 4×. - -**5.6 — Resume PASSES with no `/resume` hook declared.** Run on a second VM -(`microvm-d4992cba-92a5-328b-b534-d40ef8715ab3`) so the TTL observation was not -disturbed. `resume-microvm` returned an empty body; state `RUNNING` at **+1 s** -and stable through +55 s. **`microvmId` and `endpoint` were byte-identical before -and after** suspend/resume — lifecycle identity is preserved, so a stored -`SessionHandle` survives a suspend/resume cycle. - -**5.7 — Terminate is fast; `NotFound` did NOT arrive.** Terminate requested -20:26:09Z: `TERMINATING` at **+1 s**, `TERMINATED` at **+3 s**, then still -`TERMINATED` at +9/+20/+41/+72/+133/**+254 s**, and still `TERMINATED` at the -+15-min checkpoint (~10 min after termination) and in `list-microvms` at every -later checkpoint. **`ResourceNotFoundException` was never observed.** The -strategy's `NotFound → completed` mapping is not wrong, but it is **not the -near-term signal**: for at least ~10 minutes the observable terminal state is -`TERMINATED`, so `TERMINATED` must map to completed on its own. - -**5.8 — SKIPPED as a role-scoped test** (Lambda role trust does not allow -operator assumption — 5.0). Advisory admin-identity results: - -- Different image ARN → `ResourceNotFoundException: No active version found for - MicroVM image arn:aws:lambda:us-east-1::microvm-image:backgroundagent-dev-abca-agent-different`. - **Inconclusive** for exact-ARN denial, exactly as the runbook predicts. Note - the message is about *no active version*, not a missing image. -- `create-microvm-auth-token --expiration-in-minutes 5 --allowed-ports - '[{"port":8080}]'` → **SUCCEEDED**. The CLI union syntax is valid, and the - request shape matches the SDK. It returned a genuine JWE under key - `X-aws-proxy-auth` (header `{"kid":"9e81880a-…","alg":"dir","enc":"A256GCM"}`), - **minted against a `SUSPENDED` MicroVM**. So tokens are mintable at will by any - principal holding the action, including for suspended VMs; the no-JWE posture - rests entirely on the orchestrator role omitting the action, which 4.4 verified - statically. - -**5.9 — The 16 KB boundary is WRONG: the real cap is 4096.** Both runbook -payloads were rejected identically: - -``` -An error occurred (ValidationException) when calling the RunMicrovm operation: 1 validation error detected: Value at 'runHookPayload' failed to satisfy constraint: Member must have length less than or equal to 4096 -``` - -Re-measured at the real boundary: **4096 bytes passes length validation** (and -then fails only the hook-enabled check), **4097 bytes is rejected** with the -message above. The CLI **does** expand `file://` (a literal URI would have been -33 bytes and passed). Length validation runs **before** the hook-enabled check, -which is why the boundary was measurable on a hook-less image at all. - -`lambda-microvm-strategy.ts:84` sets `RUN_HOOK_PAYLOAD_LIMIT_BYTES = 16_384` and -inlines anything `<= 16384`; ADR-021 states `≤ 16 KB` in four places. Any -envelope between 4,097 and 16,384 bytes would be inlined by the strategy and -**rejected by the service**. No MicroVM was created by these calls, so no -cleanup was needed. - -### Phase 6 - -**6.1 — PASS.** `npx tsc -p cli/tsconfig.json` built cleanly (mise fallback). -`bgagent configure` wrote config (`BGAGENT_CONFIG_DIR=/tmp/abca-645-bgagent`); -`platform outputs` resolved all seven MicroVM outputs. No Cognito login was -needed for operator AWS commands, as documented. - -**6.2 — PASS.** `repo onboard verification-only/issue-645 --compute-type -lambda-microvm` succeeded (the live `ListManagedMicrovmImages` probe passed) with -`status: active`, `compute_type: lambda-microvm`. -`platform doctor` returned `passed: true` with **all six** checks passing, -including `lambda_microvm_availability` — *"Managed MicroVM images are available -in us-east-1"*. (`github_token` also passed: the CFN-generated secret holds a -32-character generated placeholder, which the check cannot distinguish from a -real PAT — worth noting, since the runbook expected doctor to fail here.) -`runtime status` grouped the row under `lambda_microvm_substrates` with -`used_by_repos: ["verification-only/issue-645"]`. -`repo offboard` set `status: removed` with a TTL. Nothing was skipped. - -### Phase 7 - -**7.1 — DEFERRED-TO-P2-ENV — missing Cognito user/login and real repo -onboarding.** The user pool `us-east-1_sYU3Rftw6` has **zero users**, so -`bgagent login` is impossible, and no genuinely accessible repository is -onboarded; the platform GitHub secret is a 32-character generated placeholder. -None of these were created, per the step's instruction. Independently, the task -path could not have produced meaningful classification evidence: `RunMicrovm` -rejects the bare-name identifier the orchestrator injects (4.4 / 5.1), so every -task would fail with a `ValidationException` at launch rather than at -"session start" as the P1 narrative predicts. - -### Phase 8 — teardown (executed) - -**8.1** — `list-microvms` before teardown showed -`microvm-b44b69d9-…` `SUSPENDED` and `microvm-d4992cba-…` `TERMINATED`. -`terminate-microvm` on the suspended VM (a `SUSPENDED` VM terminates directly, -no resume required) → both `TERMINATED`; no non-terminal VM remained. - -**8.2** — Both out-of-band images deleted (`list-microvm-images` now returns -**empty**). **Correction to the runbook's loop:** the *last remaining* version -cannot be deleted individually — -`ValidationException: This is the last version. Please delete the entire image` — -so the correct order is *delete every version except the last, then delete the -image*, which reaps the final version. A version delete in flight also puts the -image in `UPDATING` and makes concurrent calls fail with -`ConflictException: MicroVM Image is already in state: UPDATING` and -`ValidationException: Cannot delete MicroVM image in its current state: `; -both cleared on retry after ~40 s. Versions reaped: `1.0` + `2.0` on -`…-abca-agent` and `1.0` + `2.0` on `…-abca-agent-hooks` — note that **failed -builds still create versions that must be reaped**. `get-microvm-image` on the -first image now returns -`ResourceNotFoundException: MicroVMImage not found for MicroVMImageID: `. - -**8.3 — Stack deletion is INCOMPLETE: `DELETE_FAILED`, blocked by leaked -AgentCore ENIs. All billable resources are confirmed gone.** - -`cdk destroy` ran 21:24:34Z and deleted 460 of 464 resources. It then failed: - -``` -The following resource(s) failed to delete: [AgentVpcRuntimeSG96507CD0, AgentVpcPrivateSubnet1Subnet8051BB57, AgentVpcPrivateSubnet2SubnetC66971D0]. -resource sg-05f50b9950d41572e has a dependent object (Service: Ec2, Status Code: 400 …) -Resource handler returned message: "The subnet 'subnet-0befbffccbb83b718' has dependencies and cannot be deleted. (Service: Ec2, Status Code: 400 …)" -``` - -Cause: two ENIs of `InterfaceType: agentic_ai` (AgentCore-managed, requester -`AROA…[redacted]:[redacted]`) — `eni-080353b2356328ed7` and -`eni-04500acbfed377f4a` — remained `in-use` in the private subnets. The runbook's -own gotcha ("VPC teardown can lag while service-managed ENIs are reclaimed; wait -and retry rather than force-deleting resources past CloudFormation") was -followed: **three** `delete-stack` retries spread over ~1 h 40 min (21:48Z, -22:37Z, 23:05Z) all returned `DELETE_FAILED`, and the ENIs were still `in-use` -each time. A direct `delete-network-interface` was attempted once for diagnosis -and correctly refused (`InvalidParameterValue: Network interface … is currently -in use.`); nothing was force-deleted past CloudFormation. - -Final residual state of `backgroundagent-dev` — **4 resources, all zero-cost**: -`AWS::EC2::VPC AgentVpcA6796801` (`vpc-0a12c1a64cc960c6a`), two private subnets, -and one security group, plus the two AgentCore ENIs holding them. - -**Billing is stopped.** Verified after teardown: - -- MicroVMs: both `TERMINATED`; `list-microvm-images` empty (no snapshot storage). -- **NAT gateways: none in the ABCA VPC** (`nat-0c6cdbee97699f4e8` deleted - 21:24:43Z). The two `available` NAT gateways in the account belong to - pre-existing `vpc-01c9984d163d2965e` and were **not** created or touched by - this run. -- **VPC endpoints in the ABCA VPC: none.** -- ABCA S3 buckets: none (auto-delete custom resources emptied them). -- Elastic IPs: no unattached (billable) addresses. -- `/aws/lambda-microvms/backgroundagent-dev-abca-agent`: deleted. - -Retry command for whoever picks this up (should succeed once AgentCore releases -the ENIs): - -```bash -aws cloudformation delete-stack --stack-name backgroundagent-dev -aws cloudformation wait stack-delete-complete --stack-name backgroundagent-dev -``` - -**Verification-only resources created and removed:** the -`abca645-connector-probe` stack (deleted — `Stack with id -abca645-connector-probe does not exist`) and the temporary port-80 egress rule -`sgr-07ed1fa48ef38467a` (removed with its security group when the stack deleted -it). The local `finch` VM was stopped. - -**Bootstrap retained as instructed:** `CDKToolkit` `UPDATE_COMPLETE`, -`ComputeTypes = agentcore,lambda-microvm`, with all five ABCA policies attached -to `cdk-hnb659fds-cfn-exec-role--us-east-1` and **no -`AdministratorAccess`**. - ---- - -## Results table (fill this in) - -Use one row per material observation; add rows as needed. - -| Step | Expected | Observed | ADR item discharged | Feeds back to design? | -|---|---|---|---|---| -| Setup | Vars + evidence dir | `us-east-1`, `backgroundagent-dev`, evidence `/tmp/abca-645-p1-20260731T184822Z`. `mise` absent → all **raw fallbacks** used | Reproducibility | No | -| 0.1 | Virgin account, correct branch | Account ``, admin `AdminConsoleAccess/aamorosi-Isengard`, branch OK, SHA `0505f914`. `backgroundagent-dev` absent (`ValidationError … does not exist`). **Account not virgin overall** — 3 unrelated stacks + `CDKToolkit` pre-existed; ABCA itself never deployed, so the premise holds | Clean first-deploy path | No | -| 0.2 | CLI exposes Lambda MicroVMs; SDK 3.1098.0 | `aws lambda-microvms help` **exit 0**, 24 commands, all 7 required present. CLI command list is an **exact match** to SDK 3.1098.0 (24 vs 24). CLI 2.36.13, cdk 2.1129.0, node 24.16.0, python 3.9.6, **openrsync** (worked). Skeleton confirms `ARM_64`-only and `ENABLED\|DISABLED` hooks with a single `hooks.port` and **no hook-path field** | API/action-name verification | **Yes — CFN L1 shapes (`arm64`, `run:'/run'`) have no counterpart in the service model** | -| 1.2 | Custom template replaces admin bootstrap | **DEFECT: `cdk bootstrap --template …` is a silent no-op** on an already-bootstrapped account (`Not overwriting it with a template containing 'ABCA: Least-Privilege Bootstrap' (use --force …)`, **exit 0**). `--force` required. `BootstrapVariant` then *stays* `AWS CDK: Default Resources`, so every future non-forced bootstrap refuses again. **DEFECT: the runbook's `--parameters` shorthand is rejected** (`Invalid type for parameter Parameters[0].ParameterValue, value: ['agentcore', 'lambda-microvm'] … valid types: `); the escaped `agentcore\,lambda-microvm` works | Conditional bootstrap IAM | **Yes — `mise //cdk:bootstrap` + runbook/docs** | -| 1.3 | Custom policy attached for `agentcore,lambda-microvm` | `ComputeTypes=agentcore,lambda-microvm`; exactly one `IaCRole-ABCA-Compute-LambdaMicrovms` policy, attached to `cdk-hnb659fds-cfn-exec-role-…`; 5 ABCA policies, **no `AdministratorAccess`**; 19 MicroVM/connector actions incl. `PassNetworkConnector` | Conditional bootstrap IAM | No | -| 2.1 (synth) | No-image warning | `abca:microvm-image-not-provisioned` present, `…p1-not-runnable` correctly absent. Incidental: template 893273/1000000, 463/500 resources | First-deploy bootstrap state | No | -| 2.1 (deploy) | Substrate deploys | **BLOCKED — cannot deploy from unmodified sources.** `AWS::Lambda::NetworkConnector` CREATE_FAILED: `"NetworkConnectorOperatorRole is required for VPC_EGRESS connector type (… Status Code: 400 …)" HandlerErrorCode: InvalidRequest`. Refutes the construct's stated *"`operatorRole` is left unset so Lambda manages the ENIs with its own service-linked role"*. Deployed only after patching in an operator role (+ moving off `us-east-1a`): `CREATE_COMPLETE` in **13 min 42 s** | Conditional substrate — **only with a code fix** | **YES — construct must create + pass an operator role; L1 marks it optional** | -| 2.1 (deploy, 2nd) | — | **AgentCore blocker:** `The following subnets are in unsupported availability zones in region us-east-1: subnet-… in us-east-1a (ID: use1-az6). Supported availability zones are: use1-az4, use1-az1, use1-az2`. This account maps `us-east-1a`→`use1-az6`; `AgentVpc` does not constrain AZs | — | **YES — `AgentVpc` should pin AgentCore-supported AZs** | -| 2.1 (rollback) | — | Rollback itself failed: `Validation failed during DeleteMemory: Memory is in transitional state CREATING. Cannot delete memory.` → `ROLLBACK_FAILED`; plain `delete-stack` cleared it | — | Minor — yes (`AgentMemory` delete retry) | -| 2.2 | Outputs/resources present | `ComputeSubstrate=lambda-microvm`; all six `Microvm…` outputs populated; key `microvm-images/agent-artifact.zip`; 2 buckets, build+execution roles, **443-only SG** (`sg-0e662dc0d6f6e9ade`, one rule tcp/443), `/aws/lambda-microvms/…` log group, `AWS::Lambda::NetworkConnector`; **no `AWS::Lambda::MicrovmImage`** | Conditional substrate/config | No | -| 2.3 | No `MICROVM_*` env | `[]` ✓ (14 env keys). **DEFECT: the runbook's `ORCHESTRATOR_FN` query is broken by pagination** — 464 resources → the query ran per page and returned 5 values (`None None None None `), breaking the next call. Needs `--no-paginate` | Reject-without-image config | Yes — runbook | -| 2.4 | Construct resources have backend tag | All **6** taggable resources tagged `abca:compute-backend=lambda-microvm` (SG, connector, log group, 2 buckets, 2 roles). `resourcegroupstaggingapi` returned only 5 — **IAM roles are not returned by that API** (coverage gap, not a missing tag); confirmed via `iam list-role-tags` | Cost attribution | No | -| 3.1–3.2 | Unsupported failure; override warning | 3.1 exit 1 naming `eu-central-1`, all five Regions, and `--context microvm_region_override=true`. 3.2 exit 0 with `abca:microvm-region-override` (plus the not-provisioned warning) | Region gate/escape hatch | No | -| 4.1 | Live regional probe | Exactly **one** base image: `…:aws:microvm-image:al2023-1`, versions `1` and `0`. **Ordering is newest-first, so the runbook's `items[-1]` selects the OLDER version `0`**; `items[0]`=`1` is correct. Service echoes `baseImageVersion: "1.0"` | Regional availability probe | Yes — runbook selector | -| 4.2 (artifact) | Script uploads | Staged + zipped + uploaded fine on macOS/openrsync: script printed `584K artifact`, S3 `ContentLength 597305`, SSE `AES256` | Packaging plane | No | -| 4.2 (create) | Image created; P1 banner shown | **FAILED:** `The ready (/ready) MicroVM image hook must be enabled when any MicroVM lifecycle hook (run, resume, suspend, or terminate) is enabled.` → **the P1 banner was never printed** (it comes after the failing call), and the runbook's `2>&1 \| tee` reported `EXIT=0`, masking the failure | **NOT discharged** — packaging + operator warning | **YES — "declare `/run` in P1, serve it in P2" is not a reachable state** | -| 4.2 (memory) | 32,768 MiB accepted | **FAILED:** `The requested memory size of 32768 MiB is not supported by base MicroVM image …al2023-1. Supported memory sizes in MiB are: [512, 1024, 2048, 4096, 8192].` Real ceiling **8192 MiB (8 GiB)**, ¼ of the documented figure | Refutes sizing premise | **YES — `DEFAULT_MINIMUM_MEMORY_MIB` and ADR-021's "32 GB RAM"** | -| 4.3 (build 1) | Build successful | **FAILED** (`The container image build failed.`). Root cause in the log group: `Could not connect to deb.debian.org:80 (146.75.38.132), connection timed out` → `E: Unable to locate package curl/git/build-essential` → `exit code: 100`. **The construct's 443-only SG makes the agent image unbuildable** (`apt-get` needs port 80; DNS was fine) | **NOT discharged** without a fix | **YES — SG must allow 80, or the Dockerfile must not use HTTP apt** | -| 4.3 (build 2) | Image ACTIVE | After adding a temporary port-80 egress rule: version **`2.0`** `state=SUCCESSFUL`, `status=ACTIVE` in **5 min 51 s**, both builds `SUCCESSFUL` | Buildability (with fixes) | No | -| 4.3 (shapes) | `IMAGE_VERSION=1` | **Version is `1.0`, not `1`.** **Two builds per version** (`chipsetGeneration` 3 and 4, GRAVITON) — the runbook's `items[0]` checks only one. `list-microvm-image-builds --image-identifier ` → `ValidationException: Invalid ARN format: …` (**ARN required**). `snapshotBuild` lives on `get-microvm-image-build`, **not** on the version (which returns `null`) | State/shape mapping | Yes — runbook + script | -| 4.3 (sizes) | Size vs 2 GB narrative | `codeInstallSizeInBytes` **2,334,748,672 (2.17 GiB) — exceeds AgentCore's 2 GB container-image limit**; `memorySnapshotSizeInBytes` 1,216,577,536 (1.13 GiB); `diskSnapshotSizeInBytes` 37,089,280 (35.4 MiB). Same tree as an OCI image = 1.799 GB (629.7 MB compressed). Snapshots deliberately not summed | Sizing narrative | **Yes — state which measure the 2 GB comparison uses** | -| 4.3 (disk) | Verify 32 GB disk | **NOT EXPOSED** — no disk quota in `service-quotas`, no disk field on image/version. 32 GB disk unverified; 32 GB *memory* refuted | Disk external fact | Yes | -| 4.3 (P1 shape) | — | **Decisive:** the exact P1 hook shape + service-mandated `/ready` **FAILED both builds**: `Ready hook check failed: the application returned a client error (HTTP 4xx) response`. The agent *does* answer on 8080 but not `/ready`. **A P1 image declaring `/run` is not creatable at all** | Refutes P1 premise | **YES — P1/P2 hook phasing** | -| 4.4 | Env present; exact image IAM; no JWE grant | **PASS exactly as designed.** `abca:microvm-image-p1-not-runnable` emitted. Exactly 5 `MICROVM_*` vars, no ingress var. `MicrovmLifecycle` = exactly `RunMicrovm`/`GetMicrovm`/`TerminateMicrovm` on `…:microvm-image:backgroundagent-dev-abca-agent` + `:*`; `MicrovmPassNetworkConnector` on `*`; `MicrovmPassExecutionRole` with `iam:PassedToService=lambda.amazonaws.com`; **zero** Suspend/Resume/AuthToken actions | Least privilege | No | -| 4.4 (identifier) | Bare name accepted by RunMicrovm | **REFUTED.** `MICROVM_IMAGE_IDENTIFIER` is the bare name `backgroundagent-dev-abca-agent`; `run-microvm` with a bare name → `ValidationException: Malformed ARN - doesn't start with 'arn:'`. `lambda-microvm-strategy.ts:236` passes it straight through, so P1 would fail at launch | Refutes construct comment | **YES — inject the image ARN, not the name** | -| 5.0 | Assume-role identity or trust denial | **ADMIN — advisory.** `AccessDenied … not authorized to perform: sts:AssumeRole on resource: …TaskOrchestratorOrchestratorFnS-Bd7rBa2V6Jwf`; trust = `lambda.amazonaws.com` only. Trust not modified | Verification confidence | No | -| 5.1 | Hook-less behavior measured, not assumed | **Measured: it runs and keeps running.** `RUNNING` at **+12 s**, still `RUNNING` at +145 s, `stateReason` always `None` — no terminate, no stall, no disappearance. `maximumDurationInSeconds=28800` accepted, `idlePolicy` omitted. A payload on a hook-less image is rejected: `The run hook must be enabled in the MicroVM image to pass the run hook payload` | P1/P2 phase boundary | **Yes — P2 startup/hooks** | -| 5.1 (ingress) | No ingress | **Service auto-attached `…:aws:network-connector:aws-network-connector:HTTP_INGRESS`** with a public `*.lambda-microvm.us-east-1.on.aws` endpoint, though none was requested. "No ingress" is not the service default | Security posture | **YES — P1 must suppress or accept default public ingress** | -| 5.2 | Actual state enum values recorded | Observed 5 of 6: `PENDING`, `RUNNING`, `SUSPENDED`, `TERMINATING`, `TERMINATED`. **`SUSPENDING` never observable** (<1 s). No unknown state. `ResourceNotFoundException` not reached | State mapping | Yes (see 5.7) | -| 5.3 | Manual suspend without idle policy | **PASS.** Admin suspend on a hook-less image, no `idlePolicy`: **empty response body**, `SUSPENDED` at **+1 s**, stable | Explicit suspend external fact | **Yes — P3 lifecycle** | -| 5.4 | Suspended TTL/checkpoint result | **No TTL within 1 h.** `SUSPENDED` at start / +15 min / +45 min / +1 h (3617 s); `startedAt` and `maximumDurationInSeconds=28800` unchanged. **TRUNCATED at ~1 h**; the 4 h checkpoint was NOT run, so a TTL between 1 h and the 8 h bound is **still open** | Partially — bounded below only | **Yes — timeout policy** | -| 5.5 | Suspended quota consumption proven/inconclusive | **LISTED, NOT PROVEN — NOT OBSERVABLE SAFELY.** Suspended VM present in `list-microvms` at every checkpoint. `L-CD1C0CC4 Max allocated memory = 1024 GB` (ACCOUNT, "burst up to 4x"), **no `UsageMetric`**; `AWS/Usage` has only `CallCount`; no MicroVM memory metric anywhere. Proving it needs ~128 × 8 GiB VMs. `L-B430C318 = 8 hours` independently confirms the 28,800 s bound; **no disk quota exists** | **NOT discharged** | **Yes — concurrency policy; rationale was sized on 32 GiB/VM, real is 8 GiB** | -| 5.6 | Resume transitions/result | **PASS with no `/resume` hook declared.** `RUNNING` at **+1 s**, empty response body; **`microvmId` and `endpoint` byte-identical** across suspend→resume, so a stored `SessionHandle` survives | Resume external fact | **Yes — P3 hooks/reconciliation** | -| 5.7 | Terminate→NotFound timing | `TERMINATING` **+1 s** → `TERMINATED` **+3 s**, then `TERMINATED` at +254 s and still `TERMINATED` ~10 min later and at every later checkpoint. **`ResourceNotFoundException` never observed** | Partially — terminate path yes, `NotFound` mapping no | **YES — `TERMINATED` must map to completed; `NotFound` is not the near-term signal** | -| 5.8 | Different image/JWE denied under role | **SKIPPED — LAMBDA ROLE TRUST DOES NOT ALLOW OPERATOR ASSUMPTION.** Advisory (admin): different image ARN → `ResourceNotFoundException: No active version found for MicroVM image …-different` (**inconclusive**, as predicted). `create-microvm-auth-token … --allowed-ports '[{"port":8080}]'` **SUCCEEDED** as admin, returning a real JWE under `X-aws-proxy-auth` (`{"alg":"dir","enc":"A256GCM"}`) **against a SUSPENDED VM**; CLI union syntax valid. No-JWE posture rests solely on 4.4's role omission | Static only (4.4) | Yes — tokens are mintable for suspended VMs | -| 5.9 | 16,384 accepted; 16,385 rejected | **REFUTED — the cap is 4096, not 16,384.** Both 16,384 and 16,385 → `Value at 'runHookPayload' failed to satisfy constraint: Member must have length less than or equal to 4096`. Re-measured: **4096 passes, 4097 rejected**. CLI does expand `file://`. Length validation precedes the hook check. `RUN_HOOK_PAYLOAD_LIMIT_BYTES = 16_384` would inline 4,097–16,384-byte envelopes that the service rejects | **Refutes the documented boundary** | **YES — strategy threshold + ADR-021 (4 places)** | -| 6.1–6.2 | Onboard probe, doctor, grouping, cleanup | **PASS, nothing skipped.** `npx tsc` build clean; `platform outputs` resolved all 7 MicroVM outputs; onboard probe passed with `compute_type=lambda-microvm`, `status=active`; doctor `passed: true` with all 6 checks including `lambda_microvm_availability` ("Managed MicroVM images are available in us-east-1"); `runtime status` grouped under `lambda_microvm_substrates`; offboard `status=removed` + TTL. Note `github_token` **passed** on a 32-char generated placeholder | CLI regional enforcement | Minor — yes (`github_token` can't detect a placeholder) | -| 7.1 | Negative task evidence or explicit deferral | **DEFERRED-TO-P2-ENV — missing Cognito user/login, real GitHub token, and real repo onboarding.** User pool `us-east-1_sYU3Rftw6` has **zero users**; secret is a 32-char placeholder. Not created, per the step. Independently moot: `RunMicrovm` rejects the bare-name identifier, so tasks would fail at launch, not at session start | Not discharged (by design) | **Yes — P2 env** | -| 8.1–8.2 | VMs/images gone | **PASS.** Both MicroVMs `TERMINATED` (a `SUSPENDED` VM terminates directly, no resume needed). `list-microvm-images` **empty**. **Correction:** the last remaining version cannot be deleted alone (`This is the last version. Please delete the entire image`) — delete all but the last, then the image. Concurrent calls during a version delete give `ConflictException: MicroVM Image is already in state: UPDATING`. **Failed builds still create versions that must be reaped** (`1.0`+`2.0` on both images) | Versioned image lifecycle | Yes — runbook loop order | -| 8.3 | Stack gone; bootstrap retained | **PARTIAL — `DELETE_FAILED`.** 460/464 resources deleted; VPC + 2 private subnets + 1 SG remain, blocked by two leaked AgentCore ENIs (`InterfaceType: agentic_ai`) still `in-use` after 3 retries over ~1 h 40 min. **All billable resources confirmed gone** (no ABCA NAT gateway, no VPC endpoints, no buckets, no images/VMs, no unattached EIPs); residual 4 resources are zero-cost. Nothing force-deleted past CFN. `CDKToolkit` retained `UPDATE_COMPLETE`, `ComputeTypes=agentcore,lambda-microvm`, 5 ABCA policies, no `AdministratorAccess`. Verification-only extras (`abca645-connector-probe`, temp rule `sgr-07ed1fa48ef38467a`) removed | Partially — cleanup blocked by AgentCore | **YES — AgentCore ENI reclaim blocks clean `cdk destroy`** | - -## Findings summary - -Live run, 2026-07-31, account ``, `us-east-1`, branch -`feat/645-lambda-microvm-p1` @ `0505f914`. Evidence: -`/tmp/abca-645-p1-20260731T184822Z`. - -**Headline:** the *P1 substrate* is broadly correct — conditional bootstrap IAM, -outputs, tags, region gate, warnings, exact-ARN least privilege, and the CLI -surface all behave as designed. But **P1 cannot deploy, cannot build its image, -and cannot launch a MicroVM from unmodified sources**: five independent -live-service rejections had to be worked around to get any empirical result, and -three documented ADR-021 constants (32 GB memory, 16 KB payload, bare-name image -identifier) are **wrong**. - -### Items discharged (behaved exactly as designed) - -1. **0.2** — CLI/SDK operation names: `aws lambda-microvms` exposes 24 commands, - an exact match to SDK 3.1098.0. No action-name drift; the packaging script's - `ARM_64` / `ENABLED` shapes are confirmed correct against the live model. -2. **1.3** — Conditional bootstrap IAM: `IaCRole-ABCA-Compute-LambdaMicrovms` is - created and attached only with `ComputeTypes` including `lambda-microvm`, and - `AdministratorAccess` really is replaced. -3. **2.2** — Substrate contract: `ComputeSubstrate=lambda-microvm`, all six - `Microvm…` outputs, both buckets, both roles, 443-only SG, `/aws/lambda-microvms/` - log group, network connector, and **no** `AWS::Lambda::MicrovmImage`. -4. **2.3** — All-or-nothing config: zero `MICROVM_*` env vars without an image. -5. **2.4** — Cost tags: all six taggable construct resources carry - `abca:compute-backend=lambda-microvm`. -6. **3.1 / 3.2** — Region gate and escape hatch, verbatim as specified. -7. **4.1** — Live regional availability probe works. -8. **4.2 (artifact half)** — Packaging plane: zip+Dockerfile staging, no secret - build inputs, correct bucket/key, works on macOS with `openrsync`. -9. **4.4** — Least privilege: exactly `RunMicrovm`/`GetMicrovm`/`TerminateMicrovm` - on exactly the image ARN + `:*`; `PassNetworkConnector`; scoped `iam:PassRole`; - **zero** `SuspendMicrovm`/`ResumeMicrovm`/`CreateMicrovmAuthToken`. Both - no-image and image-configured warnings fire correctly. -10. **5.3** — Manual suspend works without any `idlePolicy` (`SUSPENDED` in ~1 s). -11. **5.6** — Manual resume works with no `/resume` hook declared; `microvmId` - **and** `endpoint` are preserved, so `SessionHandle` survives a cycle. -12. **5.7 (terminate half)** — Explicit terminate is near-instant - (`TERMINATING` +1 s → `TERMINATED` +3 s). -13. **6.1 / 6.2** — CLI: outputs discovery, live `ListManagedMicrovmImages` - onboarding probe, `lambda_microvm_availability` doctor check, - `lambda_microvm_substrates` grouping, and offboard all pass. -14. **8.1 / 8.2** — MicroVM and image cleanup paths work (with the version-order - correction below). - -### Items contradicting design assumptions — `feeds-back-to-design: YES` - -Ordered by severity. - -**F1. The P1 image is not creatable at all** (blocks the entire P1 premise). -`create-microvm-image` with the script's/construct's hook shape: - -``` -ValidationException: The ready (/ready) MicroVM image hook must be enabled when any MicroVM lifecycle hook (run, resume, suspend, or terminate) is enabled. The ready hook signals when the application has finished initializing so the snapshot is taken in a ready state. -``` - -And with `/ready` added as demanded, both builds fail: - -``` -Ready hook check failed: the application returned a client error (HTTP 4xx) response -``` - -ADR-021's hook-phasing plan ("declare `/run` in P1, serve it in P2; omit `/ready` -and `/validate` because the agent does not implement them") is **not a reachable -service state**. The only creatable P1 image is one with **no hooks at all**, and -such an image **cannot accept a `runHookPayload`** -(`The run hook must be enabled in the MicroVM image to pass the run hook -payload`) — so P1's payload-delivery path cannot function either. The agent does -answer on port 8080 (HTTP 4xx, not a connection refusal), so serving `/ready` -is the unblocking change. - -**F2. The substrate cannot deploy: the network connector requires an operator -role.** - -``` -"NetworkConnectorOperatorRole is required for VPC_EGRESS connector type (Service: Lambda, Status Code: 400, Request ID: 04726267-6c61-4ff5-bb1d-302122e9f955)" HandlerErrorCode: InvalidRequest -``` - -This refutes the explicit comment in `lambda-microvm-compute.ts` (~L467): -*"`operatorRole` is left unset so Lambda manages the ENIs with its own -service-linked role rather than a role we would have to trust."* The generated L1 -also mis-signals it as optional (`readonly operatorRole?: string`). Proven fix -(validated standalone): a role trusting `lambda.amazonaws.com` with -`AWSLambdaVPCAccessExecutionRole` + `ec2:CreateNetworkInterface` / -`DeleteNetworkInterface` / `DescribeNetworkInterfaces` / `DescribeSubnets` / -`DescribeVpcs` / `DescribeSecurityGroups` / `CreateTags` / -`AssignPrivateIpAddresses` / `UnassignPrivateIpAddresses` / -`Describe|ModifyNetworkInterfaceAttribute`. - -**F3. `RunMicrovm` requires an image ARN; the orchestrator injects a bare name.** - -``` -ValidationException: Malformed ARN - doesn't start with 'arn:' -``` - -`MICROVM_IMAGE_IDENTIFIER` is set to `backgroundagent-dev-abca-agent` and -`lambda-microvm-strategy.ts:236` passes it straight to `RunMicrovm`. The -construct's comment claims *"`run-microvm --image-identifier` accepts"* bare -names — it does not. Every P1 task would fail at launch. The same applies to -`list-microvm-image-builds` (`ValidationException: Invalid ARN format: …`), which -the packaging script's operator instructions also get wrong. Note the construct -already derives the correct ARN for IAM, so the fix is to inject that ARN. - -**F4. The 443-only security group makes the agent image unbuildable.** - -``` -Could not connect to deb.debian.org:80 (146.75.38.132), connection timed out -E: Unable to locate package curl / git / build-essential -… did not complete successfully: exit code: 100 -``` - -`agent/Dockerfile` runs `apt-get`, which uses **HTTP/80**; the construct's SG -allows only 443 (DNS resolution succeeded, so the port is the sole cause). Either -the SG must allow 80 for build-time egress, or the Dockerfile must use an -HTTPS apt transport/mirror. Opening port 80 made the build succeed immediately. - -**F5. Memory: 32,768 MiB is rejected; the real ceiling is 8,192 MiB.** - -``` -ValidationException: The requested memory size of 32768 MiB is not supported by base MicroVM image arn:aws:lambda:us-east-1:aws:microvm-image:al2023-1. Supported memory sizes in MiB are: [512, 1024, 2048, 4096, 8192]. -``` - -`DEFAULT_MINIMUM_MEMORY_MIB = 32768` is documented in the construct as *"the -service ceiling"*, and ADR-021 states **32 GB RAM** in at least three places -(comparison table, constraint paragraph, consequences). The real ceiling for the -only available base image (`al2023-1`) is **8 GiB — one quarter**. This -materially changes the "MicroVMs target the default-sized workload" positioning -*and* the 5.5 concurrency arithmetic (which was sized on 32 GiB/VM). - -**F6. `runHookPayload` cap is 4096 bytes, not 16,384.** - -``` -ValidationException: 1 validation error detected: Value at 'runHookPayload' failed to satisfy constraint: Member must have length less than or equal to 4096 -``` - -Boundary measured exactly: **4096 passes, 4097 fails**. Both of the runbook's -16 KB probes failed. `RUN_HOOK_PAYLOAD_LIMIT_BYTES = 16_384` -(`lambda-microvm-strategy.ts:84`) would inline every envelope from 4,097 to -16,384 bytes and the service would reject all of them; ADR-021 repeats `≤ 16 KB` -in four places, including the P1 requirement statement. - -**F7. `RunMicrovm` attaches a default public HTTP ingress connector.** -Without `--ingress-network-connectors`, the response contained: - -``` -"ingressNetworkConnectors": ["arn:aws:lambda:us-east-1:aws:network-connector:aws-network-connector:HTTP_INGRESS"] -``` - -plus a public `*.lambda-microvm.us-east-1.on.aws` endpoint. ADR-021's "no -orchestrator→agent HTTP path / no ingress in P1–P3" posture is **not the service -default**; P1 either has to suppress it explicitly or document that a public -ingress endpoint exists on every agent MicroVM. - -**F8. `NotFound` is not the near-term terminal signal.** After terminate, the VM -was `TERMINATED` at +3 s and **still `TERMINATED` ~10 min later** and at every -subsequent checkpoint; `ResourceNotFoundException` was **never observed**. The -strategy's `NotFound → completed` mapping is fine as a fallback, but `TERMINATED` -must map to completed in its own right or the orchestrator will poll a terminal -VM indefinitely. Relatedly, **`SUSPENDING` is not observable** (suspend reaches -`SUSPENDED` in <1 s), so any state machine that waits for `SUSPENDING` will hang. - -**F9. Hook-less MicroVMs run indefinitely and bill.** A P1-style hook-less image -reaches `RUNNING` in **12 s** and stays `RUNNING` (no `stateReason`, no -self-termination) up to its 8 h `maximumDurationInSeconds`. The P1 narrative -("a launch may return an ID and endpoint and then fail its hook or terminate") -is wrong in the safe direction for correctness and wrong in the *expensive* -direction for cost: nothing fails, so nothing cleans up. Active -`TerminateMicrovm` is mandatory, not belt-and-braces. - -**F10. `mise //cdk:bootstrap` silently no-ops on an already-bootstrapped -account** — `Not overwriting it with a template containing 'ABCA: Least-Privilege -Bootstrap' (use --force if you intend to overwrite)` with **exit 0**. Worse, after -a forced bootstrap `BootstrapVariant` remains `AWS CDK: Default Resources`, so -the refusal recurs forever. Any operator following ADR-002's documented flow on an -existing account keeps `AdministratorAccess`. - -**F11. `AgentVpc` picks AZs AgentCore rejects.** -`The following subnets are in unsupported availability zones in region us-east-1: -subnet-… in us-east-1a (ID: use1-az6). Supported availability zones are: -use1-az4, use1-az1, use1-az2`. AZ *names* are account-scoped, so this is a -latent first-deploy failure for any account whose `us-east-1a` maps to `use1-az6`. - -**F12. AgentCore leaks ENIs and blocks `cdk destroy`.** Two -`InterfaceType: agentic_ai` ENIs stayed `in-use` for >1 h 40 min after runtime -deletion, leaving the stack `DELETE_FAILED` with VPC/subnets/SG undeletable. -Also `AWS::BedrockAgentCore::Memory` cannot be deleted while `CREATING` -(`Validation failed during DeleteMemory: Memory is in transitional state -CREATING`), which turned one rollback into `ROLLBACK_FAILED`. - -**F13. `codeInstallSizeInBytes` = 2.17 GiB exceeds AgentCore's 2 GB image -limit**, while the same tree as an OCI image is 1.799 GB (629.7 MB compressed). -The ADR's "compare to AgentCore's 2 GB limit" narrative must say *which* measure -it means, because the two straddle the limit. - -**F14. Runbook/tooling defects found in execution** (lower severity, but they -would silently corrupt a future pass): - -- The `--parameters 'ParameterKey=ComputeTypes,ParameterValue=agentcore,lambda-microvm'` - form is rejected (`Invalid type for parameter … valid types: `); - needs `agentcore\,lambda-microvm`, as `cdk/mise.toml` already shows. -- `list-stack-resources --query "…|[0]"` is **paginated** at 464 resources and - returned five values, breaking `ORCHESTRATOR_FN`. Needs `--no-paginate`. -- `list-managed-microvm-image-versions` is **newest-first**, so `items[-1]` - selects the *older* version. -- Image version is **`1.0`**, not `1`, so `IMAGE_VERSION=1` is wrong. -- There are **two builds per version** (GRAVITON gen 3 and 4); `items[0]` checks - only one. -- `snapshotBuild` comes from `get-microvm-image-build`, not - `get-microvm-image-version` (which returns `null`). -- The script's failure is masked by `2>&1 | tee` (reported `EXIT=0`), and its - "P1 image is NOT runnable" banner never prints because it sits *after* the - `create-microvm-image` call that fails. -- Teardown: the **last** image version cannot be deleted individually - (`This is the last version. Please delete the entire image`). -- Post-synth cloud-assembly template edits are ignored if the template's S3 - asset object already exists (key = pre-edit content hash). - -### Items skipped, blocked, or inconclusive - -| Item | Verdict | Reason | -|---|---|---| -| 1.3 optional negative (scratch-qualifier bootstrap) | **SKIPPED** | The runbook directs skipping it; not cheap enough and complicates teardown. | -| CFN `AWS::Lambda::MicrovmImage` value shapes (`arm64`, `run:'/run'`) | **NOT TESTED** | The construct only synthesizes the L1 with `microvm_base_image_arn`/`_version` context, which the runbook's Phase 4 does not use. The API side is settled (0.2) and the request would be rejected on hook semantics anyway (F1), so the CFN-vs-API shape question stays **open**. | -| 5.4 suspended TTL beyond 1 h | **TRUNCATED / OPEN** | Time-boxed at ~1 h as the runbook permits. `SUSPENDED` held at start/+15/+45/+60 min (3617 s). The 4 h checkpoint was not run; a TTL between 1 h and the 8 h bound remains unknown. Bounded result: **survives ≥ 1 h with no `idlePolicy`**. | -| 5.5 suspended VM consumes account memory quota | **NOT OBSERVABLE SAFELY / UNDISCHARGED** | The VM stays in `list-microvms`, but that only proves *listed*. `L-CD1C0CC4` (1024 GB, ACCOUNT) exposes **no `UsageMetric`**; `AWS/Usage` has only `CallCount`; no MicroVM memory metric exists in any namespace; no console utilization view is reachable from a CLI-only session. Proving it would need ~128 × 8 GiB VMs. | -| 5.8 IAM negatives under the orchestrator role | **SKIPPED — LAMBDA ROLE TRUST DOES NOT ALLOW OPERATOR ASSUMPTION** | `AccessDenied … sts:AssumeRole`; trust is `lambda.amazonaws.com` only and was deliberately not modified. 4.4's static policy inspection is authoritative. | -| 5.8 exact-ARN denial sub-check | **INCONCLUSIVE** | As the runbook predicts, a different image name returns `ResourceNotFoundException: No active version found for MicroVM image …-different`, not `AccessDenied`. Advisory only (admin identity). | -| 5.8 no-JWE posture | **STATIC ONLY** | As admin, `create-microvm-auth-token` **succeeded**, returning a real JWE (`{"alg":"dir","enc":"A256GCM"}`, key `X-aws-proxy-auth`) — **against a `SUSPENDED` VM**. The posture depends entirely on the role omitting the action. | -| 7.1 negative task path | **DEFERRED-TO-P2-ENV** | Cognito pool `us-east-1_sYU3Rftw6` has **zero users**, no real repo onboarded, GitHub secret is a 32-char generated placeholder. Not created, per the step. Independently moot given F3. | -| 8.3 stack deletion | **DELETE_FAILED (billing stopped)** | Leaked AgentCore ENIs (F12). 4 zero-cost resources remain; all billable resources verified gone. Retry command recorded in 8.3. | -| Disk capacity / 32 GB disk claim | **NOT EXPOSED** | No disk quota in `service-quotas`, no disk field on image or version. Unverified. | - -### Elapsed and approximate cost - -**Elapsed:** 18:48Z → 23:07Z = **4 h 19 min** wall clock. Of that, ~1 h was the -suspend-TTL observation (run concurrently with the CLI phase, IAM checks, the -16 KB probes, and the second image build, per instructions — no idle waiting); -~1 h 15 min was consumed by the three blocked deploys plus rollbacks/redeploys; -~1 h 40 min was teardown retries. - -**Approximate cost: well under US$10, dominated by NAT gateways and VPC -endpoints, not by MicroVMs.** - -| Item | Quantity | Est. | -|---|---|---| -| NAT gateways (2 × $0.045/h) | ~2.3 h summed across 4 stack lifetimes | ~$0.21 + trivial data | -| Interface VPC endpoints (7 × $0.01/h × 2 AZ) | ~1.6 h | ~$0.22 | -| MicroVM runtime | VM1 ~3 min `RUNNING` + ~2 h 45 min `SUSPENDED`; VM2 ~3 min | < $0.50 (suspended compute is not billed; snapshot storage was < 3 h) | -| MicroVM image builds | 6 builds (2 versions × 2 chipsets × 2 images), ~6 min each | low single-digit $ at most | -| Snapshot/image storage | ~3.5 GiB × 2 images × < 3 h | negligible | -| AgentCore runtime | created 4×, never invoked | negligible | -| S3 / DynamoDB / Lambda / API GW / Cognito / Secrets / logs | brief, mostly idle | < $1 | -| ECR container asset (retained in bootstrap) | 630 MB stored | ~$0.06/month ongoing | - -The 8 h `maximumDurationInSeconds` worst case was never approached; both VMs were -explicitly terminated. - -### Deliberately left in place - -1. **`CDKToolkit`** — retained as the runbook instructs, now carrying - `ComputeTypes=agentcore,lambda-microvm` and the five ABCA least-privilege - policies **instead of `AdministratorAccess`**. ⚠️ This is a change to a - *shared* account: other CDK apps in `` now deploy through the - ABCA-scoped execution role. The original standard-bootstrap template is - captured at `$EVIDENCE_DIR/cdktoolkit-template-before.txt` if it needs - restoring. -2. **`backgroundagent-dev` in `DELETE_FAILED`** — VPC `vpc-0a12c1a64cc960c6a`, - two private subnets, one security group, and two leaked AgentCore ENIs. Zero - cost; retry `delete-stack` once AgentCore releases the ENIs. -3. **Bootstrap S3/ECR assets** — including the 630 MB agent container image, - normal bootstrap content. -4. **Service-vended log groups** — `/aws/bedrock-agentcore/runtimes/…` and - `/aws/lambda/backgroundagent-dev-…` created outside CloudFormation. - -Untouched and **not** created by this run: the pre-existing stacks -(`serverless-api-powertools`, `BuildingServerlessAPIs`, -`aws-sam-cli-managed-default`) and the two `available` NAT gateways in -`vpc-01c9984d163d2965e`. - -### Recommended follow-up before P1 merges - -F1, F2, F3, F4, F5, and F6 are each independently sufficient to make the -`lambda-microvm` backend non-functional. F1 (the `/ready` requirement) is the one -that changes the *shape* of the phase plan rather than a constant, so it should be -adjudicated first: either P1 grows a minimal `/ready` (and `/run`) responder, or -P1 ships a hook-less image and explicitly defers all payload delivery to P2. - -## Report back - -Attach or summarize the evidence paths, then provide the filled results table to -the orchestrator. Draft it as a comment for issue #645; **do not post it until -the orchestrator reviews it**. Highlight every `BLOCKED`, `INCONCLUSIVE`, -`SKIPPED`, and `DEFERRED-TO-P2-ENV` result and explicitly separate observed AWS -behavior from expectations inferred from the SDK model. diff --git a/docs/verification/645-p2-smoke-runbook.md b/docs/verification/645-p2-smoke-runbook.md deleted file mode 100644 index d7d94e4e5..000000000 --- a/docs/verification/645-p2-smoke-runbook.md +++ /dev/null @@ -1,1905 +0,0 @@ -# ADR-021 P2 Stage D — live smoke runbook - -Working verification document for issue #645 on branch -`feat/645-lambda-microvm-p2` @ `3a4b61a97b22b7bcdd9832101f8d61a077fbf103`. -Companion to [`645-p1-lambda-microvm-runbook.md`](./645-p1-lambda-microvm-runbook.md), -whose command incantations this run reuses wholesale (escaped `ComputeTypes` -comma, `bootstrap --force`, newest-first version ordering, image-version `N.0` -spelling, two builds per version, delete-all-but-last then delete-image, -teardown-as-finally). - -`docs/scripts/sync-starlight.mjs` does not mirror `docs/verification/`, so this -file intentionally stays here and is not part of the Starlight site. - -**Live run:** 2026-08-06 22:38Z → 2026-08-07 01:26Z, account ``, -`us-east-1`. Evidence directory: `/tmp/abca-645-p2-20260806`. - -**Teardown: complete.** All MicroVMs terminated, image + version deleted, both -live IAM workarounds reverted, Cognito user deleted, secret died with the stack, -96/100 stack resources deleted (4 zero-cost residual, #702), **all billable -resources confirmed gone**, and every global/host configuration restored. See -Phase 8. - -## What Stage D was for - -P2 (`ab4808c`) wired everything short of the live run. Its commit message names -what it could not assert: - -> Remaining for P2 completion: the live smoke run (clone -> change -> PR with -> `bgagent watch`) and the deferred empirical items (suspend TTL >1h, -> SUSPENDED-vs-quota, `microvmImageHooks` API spelling, `NO_INGRESS` ARN). - -Five jobs, and their verdicts: - -| # | Job | Verdict | -|---|---|---| -| 1 | **THE SMOKE** — clone → change → PR through `bgagent watch` | **BLOCKED at `implement`, turn 0** — reproducible. Clone, branch, `/run`, `platform_config`, progress events and terminate all work; `claude --version` times out (P2-F5). **No PR was created.** | -| 2 | Adjudicate the CFN `AWS::Lambda::MicrovmImage` shape P1 left half-open | **DISCHARGED — the L1 is REFUTED on 5 values** (P2-F2) | -| 3 | `microvmImageHooks` spelling | **DISCHARGED both ways** — property name/nesting correct, hook *values* must be `ENABLED`/`DISABLED` (P2-F2); the API request built by the packaging script is correct and all four hooks were accepted and served | -| 4 | `NO_INGRESS` ARN name | **DISCHARGED** — the injected ARN is right and does suppress P1 F7's default public ingress | -| 5 | Extend suspend-TTL bound; re-probe SUSPENDED-vs-quota | **TTL extension SKIPPED** (time-boxed, see Skipped); quota **re-confirmed NOT OBSERVABLE** | - -**Headline:** the P2 substrate is much closer than P1 — the image builds with all -four hooks, the MicroVM launches, `/run` installs `platform_config`, the repo -clones, progress events stream to `bgagent watch`, and finalization terminates -the VM. But **four independent defects had to be worked around to get that far**, -and the run still ends one step short of a PR on a fifth. None of the four is -visible to `cdk synth`, `mise //cdk:test`, or any unit test: every one is a -live-service contract mismatch. - ---- - -## Variables - -```bash -set -o pipefail # P1's hard-won lesson: `| tee` masks failures -export AWS_PROFILE=aamorosi+workshops-AdminConsoleAccess -export AWS_REGION=us-east-1 -export AWS_DEFAULT_REGION="$AWS_REGION" -export CDK_DEFAULT_REGION="$AWS_REGION" -export CDK_DEFAULT_ACCOUNT="" -export STACK_NAME=backgroundagent-dev -export EXPECTED_BRANCH=feat/645-lambda-microvm-p2 -export SCRATCH_REPO=dreamorosi/batch-sync-triage -export SMOKE_USER="" -export CDK_DOCKER=finch # no docker on this box (P1 deviation 3) -export EVIDENCE_DIR=/tmp/abca-645-p2-20260806 -export BGAGENT_CONFIG_DIR=/tmp/abca-645-p2-bgagent -``` - -### ⚠️ zsh trap that bit this run - -The executor's shell is **zsh**, where `$VAR:latest` is parsed as the -`${VAR:l}` *lowercase modifier*, silently producing `…atest`. It cost one -mis-diagnosed container push. **Always brace: `${VAR}:latest`.** Likewise zsh has -no `PIPESTATUS` (it is `$pipestatus`, 1-indexed) and no `timeout(1)` — P1's -`test "${PIPESTATUS[0]}" -eq 0` silently evaluates to empty here. Use -`set -o pipefail` plus a plain `$?`. - ---- - -## Execution deviations (each is itself a result) - -1. **AZ pinning was required — P1 F11 is UNFIXED on this branch.** - `agent-vpc.ts` still does `maxAzs: 2` with no AZ constraint, and this account - still maps `us-east-1a → use1-az6`, which AgentCore rejects. Rather than - re-derive a finding P1 already recorded, the gitignored CDK context cache - (`cdk/cdk.context.json`, a build artifact — `git status` stayed clean) was - trimmed to lead with `us-east-1b` (`use1-az1`) + `us-east-1c` (`use1-az2`). - Original saved to `$EVIDENCE_DIR/cdk.context.json.ORIGINAL` and **restored at - teardown**. With the pin, AgentCore Memory and Runtime created cleanly, so - this is the whole of F11's remaining impact. -2. **The AgentCore container asset could not be pushed; an existing ECR manifest - was retagged instead.** `finch` (v1.17.2, no docker on this box) *built* the - image fine but every `finch push`/`finch pull` against ECR failed instantly - with `no basic auth credentials`. Root cause established: `finch push` shells - into the Lima VM as `limactl shell finch sudo -E nerdctl push`, and the VM's - `DOCKER_CONFIG` (`/home/aamorosi.guest/.finch-vm-config/config.json`) carries - `credsStore: finchhost`, a helper that cannot resolve this account's - Isengard `credential_process`. A host-side `finch login` succeeded and wrote - to `~/.finch/config.json`, but the VM never consulted it. Since the ECR image - is **only** consumed by the AgentCore runtime — which this run never invokes, - because the smoke runs on the MicroVM substrate built from the S3 zip by the - Lambda MicroVMs service — the required asset tag was added to an existing - manifest with `aws ecr put-image`. **This does not touch the MicroVM image - under test.** Consequence to be honest about: the deployed AgentCore runtime - carries a P1-era agent image. Nothing in this runbook depends on it. - *(All global config touched during that investigation — `~/.finch/config.json` - and the VM's docker config — was reverted; see Teardown.)* -3. **The connector operator role's trust policy had to be patched to deploy at - all** (P2-F1) — via a `/tmp` cloud-assembly patch, P1's technique. No - repository source file was modified. -4. **Two IAM workarounds were applied live to get past P2-F3** (execution-role - trust condition, plus a temporary unconditioned `iam:PassRole` that turned out - to be unnecessary). Both reverted; see Teardown. -5. **The CDK-managed `CfnMicrovmImage` path was attempted first, as briefed, and - failed** (P2-F2). The out-of-band `--create-image` script path was used - instead — the same path P1 used, and currently the only one that works. -6. **The suspend-TTL extension beyond P1's 1 h floor was skipped** (time-boxed). -7. **One self-inflicted error is recorded rather than hidden:** the GitHub PAT was - first written to Secrets Manager with a trailing newline (`gh auth token | - … --secret-string file:///dev/stdin`), which produced a real clone failure. - Corrected; see 3.2. - ---- - -## Phase 0 — Preflight - -### 0.1 Identity, region, branch - -`aws sts get-caller-identity` → -`arn:aws:sts:::assumed-role/AdminConsoleAccess/aamorosi-Isengard`, -account ``. Branch `feat/645-lambda-microvm-p2`, SHA -`3a4b61a97b22b7bcdd9832101f8d61a077fbf103`. Untracked: `docs/verification/`, -`opencode.json` — the same two P1 saw. - -> **Credential note.** The profile's default region is `eu-west-1`, so -> `AWS_REGION`/`AWS_DEFAULT_REGION` must be exported explicitly for every -> command. A bare `aws` call in this account goes to the wrong Region. -> Mid-run the credentials expired once; re-pinning `AWS_PROFILE` (which resolves -> through `credential_process` and auto-refreshes) fixed it. **No global AWS -> config was read or modified at any point.** - -`backgroundagent-dev` was **absent**, exactly as P1's Phase 8 hoped: - -``` -An error occurred (ValidationError) when calling the DescribeStacks operation: Stack with id backgroundagent-dev does not exist -``` - -**P1 F12's stack half is therefore RESOLVED**: P1 ended with `backgroundagent-dev` -in `DELETE_FAILED` (VPC + 2 subnets + 1 SG pinned by two leaked `agentic_ai` -ENIs) and a recorded retry command. AgentCore did eventually release them and the -stack is gone. The three unrelated pre-existing stacks -(`serverless-api-powertools`, `BuildingServerlessAPIs`, -`aws-sam-cli-managed-default`) plus `CDKToolkit` remain. - -### 0.2 Tooling - -| Tool | Version | Note | -|---|---|---| -| `aws-cli` | 2.36.13 | `aws lambda-microvms help` **exit 0**, **24 commands** — identical to P1 | -| `mise` | 2026.8.1 | **present this time** (P1 ran entirely on raw fallbacks) | -| node | v24.16.0 | | -| python3 | 3.9.6 | system python; the *guest* runs 3.13.13 | -| `zip` | 3.0 | | -| `rsync` | openrsync (protocol 29) | packaging script works unmodified | -| docker | **absent** | | -| `finch` | v1.17.2 | builds fine, **cannot push to ECR** (deviation 2) | - -### 0.3 Bootstrap — no re-bootstrap needed - -`CDKToolkit` `UPDATE_COMPLETE` (last updated 2026-07-31T18:54:48Z, i.e. P1's run), -and it already carries what P2 needs: - -- `ComputeTypes` = **`agentcore,lambda-microvm`** ✓ -- `BootstrapVariant` = **`ABCA: Least-Privilege Bootstrap`** -- bootstrap version SSM parameter = `32` -- `cdk-hnb659fds-cfn-exec-role--us-east-1` carries exactly the **five** - ABCA policies (Application, Infrastructure, Observability, Compute-Agentcore, - **Compute-LambdaMicrovms**) and **no `AdministratorAccess`**. - -So `bootstrap --force` was **not** re-run. - -> **P1 F10 is partly self-healing — correction to the P1 record.** P1 reported as -> a "durable consequence" that `BootstrapVariant` *stays* `AWS CDK: Default -> Resources` after a forced bootstrap, so every future non-forced bootstrap -> refuses forever. It now reads `ABCA: Least-Privilege Bootstrap`. The reason is -> the very next command in P1's own runbook: `update-stack -> --use-previous-template` with **only** `ParameterKey=ComputeTypes` supplied -> resets every unspecified parameter to its **template default** (that is -> CloudFormation's documented behaviour absent `UsePreviousValue=true`), and the -> ABCA template's default for `BootstrapVariant` is the ABCA string. F10's -> *first* half (the silent `exit 0` no-op without `--force`) stands; the -> "recurs forever" half does not. - -### 0.4 Managed base image - -Unchanged from P1: exactly **one** managed base image in `us-east-1`, -`arn:aws:lambda:us-east-1:aws:microvm-image:al2023-1`, versions `1` -(2026-07-21) and `0` (2026-06-17), **newest first**, so `items[0]` = `1`. -Selected version `1`; the service echoes `baseImageVersion: "1.0"`. - ---- - -## Phase 1 — Deploy - -### 1.1 Substrate synth — PASS - -`abca:microvm-image-not-provisioned` emitted with the full remedy text; -`abca:microvm-image-p1-smoke-unverified` correctly **absent**. Both incidental -warnings grew since P1 and one is close to the wall: - -| Metric | P1 | P2 | Limit | -|---|---|---|---| -| Template size | 893,273 | **977,650** (substrate) / **983,796** (with image) | 1,000,000 | -| Resource count | 463 | **485** / **486** | 500 | - -**Feeds back to design:** the template is at **98.4 %** of the 1 MB -CloudFormation limit with the image configured. That is ~16 KB of headroom — -roughly one more construct. `suppressTemplateIndentation` or a stack split is -now a near-term requirement, not a nicety. - -### 1.2 Substrate deploy — BLOCKED, then PASS after a trust-policy patch - -**Attempts 1–2 failed on the ECR push** (deviation 2), a purely local tooling -problem. **Attempts 3–4 failed identically on the network connectors** — this is -P2-F1: - -``` -Resource handler returned message: "The service is unable to assume the provided NetworkConnectorOperatorRole. Please verify the trust policy on the role. (Service: Lambda, Status Code: 400, Request ID: dbe1d2f4-dd0c-4319-8f5d-b4be4f076843) (SDK Attempt Count: 1)" (RequestToken: d6e21be9-9857-f452-bf12-3b93f89a4c75, HandlerErrorCode: InvalidRequest) -``` - -Both `LambdaMicrovmCompute/EgressConnector` and -`LambdaMicrovmCompute/BuildEgressConnector` `CREATE_FAILED`. Attempt 4 ran -against a **freshly deleted stack**, so this is **deterministic, not IAM -propagation lag** — an important distinction, because that error message is the -classic propagation symptom and a re-run is the obvious (wrong) first guess. - -**Attempt 5** deployed a `/tmp` cloud assembly with exactly one edit — the -`aws:SourceAccount` condition removed from -`LambdaMicrovmComputeConnectorOperatorRole`'s `AssumeRolePolicyDocument`: - -```json -{"Statement":[{"Action":"sts:AssumeRole","Effect":"Allow","Principal":{"Service":"lambda.amazonaws.com"}}],"Version":"2012-10-17"} -``` - -Both connectors went `CREATE_IN_PROGRESS → Resource creation Initiated` within a -second. **`CREATE_COMPLETE` 23:13:43Z → 23:27:03Z = 13 min 20 s**, 485 resources -(P1's comparable figure: 13 min 42 s / 464 resources). - -*Two P1 gotchas recurred verbatim and their remedies still work:* - -- `AWS::BedrockAgentCore::Memory` cannot be deleted while `CREATING` - (`Validation failed during DeleteMemory: Memory is in transitional state - CREATING. Cannot delete memory.`) → rollback ends `ROLLBACK_FAILED`; a plain - `aws cloudformation delete-stack` clears it (~3 min). **P1 F12's Memory half is - unfixed.** -- Post-synth template edits are silently ignored unless the template's S3 asset - object is deleted first (the object key is the *pre-edit* content hash). - -### 1.3 Substrate outputs — PASS - -`ComputeSubstrate = lambda-microvm`; all **seven** MicroVM outputs populated -(the six P1 had, plus `MicrovmBuildEgressConnectorArns`); artifact key exactly -`microvm-images/agent-artifact.zip`. Runtime and build egress connectors are -distinct ARNs, as designed. - -### 1.4 No-image orchestrator env — PASS, and a correction to P1 F14 - -`MICROVM_*` keys = **`[]`** (23 env keys total). All-or-nothing holds. - -> **P1 F14's `--no-paginate` remedy is WRONG and silently truncates.** P1 -> concluded that `list-stack-resources --query "…|[0]"` needs `--no-paginate`. -> With 485 resources, `--no-paginate` returns **only the first page** — it -> yielded 6 Lambda functions and *no* `TaskOrchestrator`, resolving -> `ORCHESTRATOR_FN` to the literal string `None` and producing a confident -> `Function not found: …:function:None`. That is worse than P1's original bug, -> because the original at least printed five values and broke loudly. **The -> correct approach is to let the CLI paginate (its default) and filter -> client-side**, which returns all 485: -> ```bash -> aws cloudformation list-stack-resources --stack-name "$STACK_NAME" --output json \ -> | python3 -c 'import json,sys; print([r["PhysicalResourceId"] for r in json.load(sys.stdin)["StackResourceSummaries"] if "TaskOrchestrator" in r["LogicalResourceId"] and r["ResourceType"]=="AWS::Lambda::Function"])' -> ``` - -**The new `agentPlatformConfig` env vars are present — 11 of 13.** All seven -`agentPlatformConfig` fields plus the four inherited ones landed on the -orchestrator: - -| `platform_config` key | Orchestrator env var | Present | -|---|---|---| -| `task_table_name` | `TASK_TABLE_NAME` | ✓ | -| `task_events_table_name` | `TASK_EVENTS_TABLE_NAME` | ✓ | -| `github_token_secret_arn` | `GITHUB_TOKEN_SECRET_ARN` | ✓ | -| `agent_session_role_arn` | `AGENT_SESSION_ROLE_ARN` | ✓ | -| `task_approvals_table_name` | `TASK_APPROVALS_TABLE_NAME` | ✓ | -| `nudges_table_name` | `NUDGES_TABLE_NAME` | ✓ | -| `log_group_name` | `LOG_GROUP_NAME` | ✓ | -| `artifacts_bucket_name` | `ARTIFACTS_BUCKET_NAME` | ✓ | -| `trace_artifacts_bucket_name` | `TRACE_ARTIFACTS_BUCKET_NAME` | ✓ | -| `aws_sdk_ua_app_id` | `AWS_SDK_UA_APP_ID` | ✓ (`uksb-wt64nei4u6#backgroundagent-dev`) | -| `anthropic_default_haiku_model` | `ANTHROPIC_DEFAULT_HAIKU_MODEL` | ✓ (`us.anthropic.claude-haiku-4-5-20251001-v1:0`) | -| `linear_oauth_secret_arn` | `LINEAR_OAUTH_SECRET_ARN` | — (per-workspace, CLI-created; correctly absent) | -| `jira_oauth_secret_arn` | `JIRA_OAUTH_SECRET_ARN` | — (same) | - -The two absentees are **not** in the contract's `required` list, so the -producer's optional-key handling is exercised and correct. - -*Incidental:* `ARTIFACTS_BUCKET_NAME` and `TRACE_ARTIFACTS_BUCKET_NAME` resolve -to the **same** bucket (`…-traceartifactsbucket8cbd5207-pwjbv0diorys`). Not a -MicroVM issue, but worth a glance — two distinct `platform_config` keys carrying -one bucket makes the trace/artifact separation notional. - -### 1.5 The CDK-managed `CfnMicrovmImage` path — **REFUTED** (P2-F2) - -This is the question P1 left open, and it is now closed. Synth produced the -warning correctly (`abca:microvm-image-p1-smoke-unverified`) and the exact L1 -under test: - -```json -"CpuConfigurations": [{ "Architecture": "arm64" }], -"Hooks": { - "Port": 8080, - "MicrovmHooks": { "Run": "/aws/lambda-microvms/runtime/v1/run", "RunTimeoutInSeconds": 60, - "Terminate": "/aws/lambda-microvms/runtime/v1/terminate", "TerminateTimeoutInSeconds": 15 }, - "MicrovmImageHooks": { "Ready": "/aws/lambda-microvms/runtime/v1/ready", "ReadyTimeoutInSeconds": 60, - "Validate": "/aws/lambda-microvms/runtime/v1/validate", "ValidateTimeoutInSeconds": 60 } -} -``` - -CloudFormation **rejected it at change-set early validation** — the stack was -never touched, so there was no rollback: - -``` -Early validation failed for change set cdk-deploy-change-set: -backgroundagent-dev/LambdaMicrovmCompute/Image (AWS::Lambda::MicrovmImage LambdaMicrovmComputeImage16B48539) - /aws/lambda-microvms/runtime/v1/run is not a valid enum value. Supported values: [DISABLED, ENABLED] (at - /Resources/LambdaMicrovmComputeImage16B48539/Properties/Hooks/MicrovmHooks/Run) - /aws/lambda-microvms/runtime/v1/terminate is not a valid enum value. Supported values: [DISABLED, ENABLED] (at - /Resources/LambdaMicrovmComputeImage16B48539/Properties/Hooks/MicrovmHooks/Terminate) - arm64 is not a valid enum value. Supported values: [ARM_64] (at - /Resources/LambdaMicrovmComputeImage16B48539/Properties/CpuConfigurations/0/Architecture) - /aws/lambda-microvms/runtime/v1/ready is not a valid enum value. Supported values: [DISABLED, ENABLED] (at - /Resources/LambdaMicrovmComputeImage16B48539/Properties/Hooks/MicrovmImageHooks/Ready) - /aws/lambda-microvms/runtime/v1/validate is not a valid enum value. Supported values: [DISABLED, ENABLED] (at - /Resources/LambdaMicrovmComputeImage16B48539/Properties/Hooks/MicrovmImageHooks/Validate) -``` - -**Five rejected values, and the CFN surface is identical to the API surface.** -This refutes the construct's explicit reasoning — *"The CDK L1 remains -intentionally unchanged because its generated CloudFormation types accept string -values and document no architecture/hook allowed-value constraint"*. The types -accept strings; the **service** enforces the enum, at change-set time. - -Two useful corollaries: - -- **The `microvmImageHooks` *spelling* is CORRECT.** The errors are scoped to - `…/Hooks/MicrovmImageHooks/Ready` and `…/Validate`, so CFN resolved the - property name and its children. Only the *values* are wrong. Combined with 2.1 - below (the API accepted the identical structure), the naming question is - discharged in both directions. -- **Hook paths are not configurable anywhere.** Neither CFN nor the API takes a - path; both take `ENABLED`/`DISABLED`. The service calls fixed well-known routes - — 2.2 proves they are exactly the `/aws/lambda-microvms/runtime/v1/*` strings - the agent serves. So `RUN_HOOK_PATH` and friends are correct *as route - constants for the agent* and simply must not be sent as property values. - -### 1.6 Wired deploy (out-of-band image) — PASS - -After the image existed (Phase 2), redeploying with -`--context microvm_image_identifier=` completed in **3 min 33 s** -(00:06:00Z → 00:09:33Z), `UPDATE_COMPLETE`, warning -`abca:microvm-image-p1-smoke-unverified` emitted. **All six `MICROVM_*` vars -present — all-or-nothing WITH the image confirmed:** - -``` -MICROVM_EGRESS_CONNECTOR_ARNS = arn:aws:lambda:us-east-1::network-connector:nc-d306f00f-1bd0-45ea-9457-0fcec0dab2a4 -MICROVM_EXECUTION_ROLE_ARN = arn:aws:iam:::role/backgroundagent-dev-LambdaMicrovmComputeExecutionRo-tF8Idpc9aT0R -MICROVM_IMAGE_IDENTIFIER = arn:aws:lambda:us-east-1::microvm-image:backgroundagent-dev-abca-agent -MICROVM_IMAGE_VERSION = 1.0 -MICROVM_INGRESS_CONNECTOR_ARNS = arn:aws:lambda:us-east-1:aws:network-connector:aws-network-connector:NO_INGRESS -MICROVM_PAYLOAD_BUCKET = backgroundagent-dev-lambdamicrovmcomputepayloadbuc-bctl5ej8aazr -``` - -**P1 F3 is FIXED:** `MICROVM_IMAGE_IDENTIFIER` is a full ARN, not a bare name. - ---- - -## Phase 2 — Image - -### 2.1 `create-microvm-image` — PASS with all four hooks - -`package-microvm-artifact.sh` (no `--create-image`) staged and uploaded first: -**904K artifact / 922,056 bytes in S3, SSE `AES256`** (P1: 584K / 597,305 — the -P2 `server.py` growth). Then with `--create-image`, exit **0**, and the service -echoed the request back: - -```json -"cpuConfigurations": [{ "architecture": "ARM_64" }], -"resources": [{ "minimumMemoryInMiB": 8192 }], -"hooks": { - "port": 8080, - "microvmHooks": { "run": "ENABLED", "runTimeoutInSeconds": 60, - "terminate": "ENABLED", "terminateTimeoutInSeconds": 15 }, - "microvmImageHooks": { "ready": "ENABLED", "readyTimeoutInSeconds": 60, - "validate": "ENABLED", "validateTimeoutInSeconds": 60 } -}, -"state": "CREATING", "imageVersion": "1.0", -"imageArn": "arn:aws:lambda:us-east-1::microvm-image:backgroundagent-dev-abca-agent" -``` - -**`microvmImageHooks` with `ready` + `validate`, and `microvmHooks` with `run` + -`terminate`, are all ACCEPTED.** P1 could not test this: it never got a -four-hook image created (F1), and its `/ready`-only attempt failed the build. - -The script's P2 reminder banner printed both before and after the call, and -`8192 MiB` was accepted — P1 F5's ceiling holds. - -### 2.2 Build — PASS in 4 min 35 s, `/ready` **and** `/validate` served - -Two builds per version again (`chipsetGeneration` 3 and 4), both `SUCCESSFUL`; -`state=SUCCESSFUL`, `status=ACTIVE`. **00:00:01Z → 00:04:36Z = 4 min 35 s** -(P1: 5 min 51 s). - -The decisive evidence, from `/aws/lambda-microvms/backgroundagent-dev-abca-agent` -— **this is the item P1 F1 blocked entirely**: - -``` -[server/build-hook] /ready hook: server is up, reporting ready for snapshot -INFO: 127.0.0.1:52856 - "POST /aws/lambda-microvms/runtime/v1/ready HTTP/1.1" 200 OK -[server/build-hook] /validate hook: ok (python=3.13.13, platform_config_keys=13, warnings=0) -INFO: 127.0.0.1:50860 - "POST /aws/lambda-microvms/runtime/v1/validate HTTP/1.1" 200 OK -``` - -(both lines twice, once per chipset). Note what this proves beyond "the hooks -work": the service POSTs to **exactly** the `/aws/lambda-microvms/runtime/v1/*` -paths, confirming the fixed-route model inferred in 1.5; `/validate` reports -`platform_config_keys=13`, so the cross-package contract loaded inside the -snapshot; and `warnings=0`, so the baked-secret scan found nothing. - -**P1 F4 is FIXED:** `apt-get` reached `deb.debian.org` over port 80 through the -dedicated build connector — the build log shows `Get:… http://deb.debian.org/…` -succeeding. No temporary security-group rule was needed this time. - -`agent/Dockerfile` also now installs Go tooling; the build log shows -`go: downloading …` completing, so build-time egress is sufficient. - -### 2.3 Sizes — P1 F13 reconfirmed and slightly worse - -| Field | Bytes | Human | vs P1 | -|---|---|---|---| -| `codeInstallSizeInBytes` | 2,342,203,392 | **2.18 GiB** | 2.17 GiB | -| `memorySnapshotSizeInBytes` | 1,223,421,952 | 1.14 GiB | 1.13 GiB | -| `diskSnapshotSizeInBytes` | 34,959,360 | 33.3 MiB | 35.4 MiB | - -`codeInstallSizeInBytes` still **exceeds AgentCore's 2 GB container-image -limit**, while the same tree as an OCI image is 1.803 GB / 631.1 MB compressed -(measured locally). P1 F13's "say which measure you mean" recommendation stands. -`snapshotBuild` is still only on `get-microvm-image-build`, not on the version. - ---- - -## Phase 3 — Platform user, secret, repo onboarding - -### 3.1 Cognito user — PASS, no first-login dance needed - -`bgagent admin invite-user --stack-name … --region …` -created the user with a **permanent** password (`UserStatus: CONFIRMED`, not -`FORCE_CHANGE_PASSWORD`), wrote credentials to -`$BGAGENT_CONFIG_DIR/invites/…txt` mode `0600`, and printed a `configure` -bundle. `bgagent configure --stack-name` + `bgagent login --username …` → -`Login successful.` **This is the answer to "find the exact bgagent commands": -`admin invite-user` is the whole first-login flow** — there is no -`RespondToAuthChallenge` step to drive, which is why P1's 7.1 blocker -("user pool has zero users") is a two-command fix, not an obstacle. - -### 3.2 GitHub PAT into the stack secret — PASS on the second attempt - -The token was piped from `gh auth token` straight into -`aws secretsmanager put-secret-value` and **never** written to disk, a log, or -this file. - -**Operator error worth recording, because it produced a convincing false -defect.** `gh auth token` emits a trailing newline and -`--secret-string file:///dev/stdin` stores it verbatim, so the secret was **41 -bytes**. My own verification used `.strip()` and reported "length 40", hiding it. -The task then failed inside the guest with: - -``` -RuntimeError: clone failed (non-transient): Post "https://api.github.com/graphql": net/http: invalid … -``` - -— i.e. an invalid HTTP header value, because the token carried `\n`. Rewriting -the secret stripped (40 bytes, no trailing whitespace) fixed the clone -immediately. - -*Verification lesson:* assert on the **raw** secret, never a stripped copy: - -```bash -aws secretsmanager get-secret-value --secret-id "$SECRET_ARN" --output json \ - | python3 -c 'import json,sys; s=json.load(sys.stdin)["SecretString"]; print(len(s), s!=s.strip())' -``` - -*Minor robustness observation (not a defect found by this run's design):* the -agent passes the secret value through to `gh`/`git` unstripped, so any -whitespace an operator introduces surfaces as a confusing `net/http` error rather -than "your token looks malformed". A `.strip()` at the token resolver would turn -a 20-minute misdiagnosis into a non-event. - -### 3.3 Repo onboarding — PASS, gate and probe both behaved - -`bgagent repo onboard dreamorosi/batch-sync-triage --compute-type lambda-microvm`: - -```json -{ "repo": "dreamorosi/batch-sync-triage", "status": "active", - "compute_type": "lambda-microvm", "onboarded_at": "2026-08-07T00:10:43.376Z" } -``` - -Both guards fired as designed and in the documented order: the **ComputeSubstrate -gate** passed because the stack output reads `lambda-microvm`, and the live -**`ListManagedMicrovmImages` availability probe** passed. The command also -printed the two ADR-021 advisory notes, including the smoke-unverified warning — -correct, and still accurate at the end of this run. - -`bgagent platform doctor` → `passed: true`, **all 7 checks**, including -`Managed MicroVM images are available in us-east-1` and — because 3.2 had already -run — `GitHubTokenSecretArn contains a token value`. (P1 noted this check cannot -distinguish a real PAT from the 32-char generated placeholder; that is still -true, it just happens to be a true positive here.) - ---- - -## Phase 4 — THE SMOKE - -Five submissions. Each failure moved the boundary forward, so all five are -recorded. - -| # | Task ID | Outcome | Finding | -|---|---|---|---| -| 1 | `01KZCRY70HRBP236GECR768JJX` | `FAILED` — `RunMicrovm … AccessDeniedException … iam:PassRole` | P2-F3 | -| 2 | `01KZCS6451HRPSXAG33Z4R5XRV` | identical, **with an unconditioned `iam:PassRole` attached** → so PassRole was never the real problem | P2-F3 | -| 3 | `01KZCSCD6PKDZNC8WRTFNRQG3H` | identical, 3 min after the IAM change → **not propagation lag** | P2-F3 | -| 4 | `01KZCSNM8MD4MB17ZTW1VDB6PY` | **`RUNNING`** after removing the execution-role trust condition, then `clone failed … net/http: invalid` | P2-F3 root cause proven; 3.2 token bug | -| 5 | `01KZCSVRHZXVHXQ4T29XYZKBAM` | **`RUNNING` → clone OK → branch OK → `implement` failed at turn 0** | **P2-F5** | -| 5r | `01KZCT8SWZ1DDZC7P0RYS982ES` | identical to #5 — **reproducible** | P2-F5 | - -### 4.1 P2-F3 — the `iam:PassRole` deny that was really a trust-policy deny - -``` -Session start failed: Error: MicroVM RunMicrovm failed: AccessDeniedException: User: arn:aws:sts:::assumed-role/backgroundagent-dev-TaskOrchestratorOrchestratorFnS-kyEY8iz3mrcm/backgroundagent-dev-TaskOrchestratorOrchestratorFn-huSs3tbuFbJs is not authorized to perform: iam:PassRole on resource: arn:aws:iam:::role/backgroundagent-dev-LambdaMicrovmComputeExecutionRo-tF8Idpc9aT0R because no identity-based policy allows the iam:PassRole action -``` - -The message is actively misleading, and the diagnosis is the useful part of this -section. The orchestrator's inline policy **does** carry the grant, exactly as -P1 4.4 recorded it: - -```json -{ "Sid": "MicrovmPassExecutionRole", "Effect": "Allow", "Action": "iam:PassRole", - "Resource": "arn:aws:iam:::role/backgroundagent-dev-LambdaMicrovmComputeExecutionRo-tF8Idpc9aT0R", - "Condition": { "StringEquals": { "iam:PassedToService": "lambda.amazonaws.com" } } } -``` - -Elimination sequence: - -1. Attached a **temporary unconditioned** `iam:PassRole` on the same resource → - **still denied** (submission 2). So the `iam:PassedToService` condition was - *not* the cause. -2. `aws iam get-role` → **no permissions boundary**, path `/`. -3. `aws iam simulate-principal-policy … --action-names iam:PassRole` → - **`allowed`**, matching my temporary statement. So IAM itself says yes. -4. Waited 3 minutes and resubmitted → **still denied** (submission 3). Not - propagation. -5. Read the **target** role's trust policy — and found the same - `aws:SourceAccount` condition that had just broken the network connectors: - -```json -{"Effect":"Allow","Principal":{"Service":"lambda.amazonaws.com"},"Action":"sts:AssumeRole", - "Condition":{"StringEquals":{"aws:SourceAccount":""}}} -{"Effect":"Allow","Principal":{"Service":"lambda.amazonaws.com"},"Action":"sts:TagSession", - "Condition":{"StringEquals":{"aws:SourceAccount":""}}} -``` - -6. Removed both conditions → **submission 4 reached `RUNNING` in 6 seconds.** - -**So `RunMicrovm` reports a role it cannot pass-and-assume as an `iam:PassRole` -identity-policy denial on the *caller*.** P2-F1 and P2-F3 are therefore **one -root cause with two symptoms**: the Lambda MicroVMs service does not present -`aws:SourceAccount` when assuming ABCA's MicroVM-facing roles, so every trust -policy carrying that confused-deputy condition is unassumable. - -### 4.2 THE SMOKE (submission 5) — `bgagent watch`, verbatim - -``` -Watching task 01KZCSVRHZXVHXQ4T29XYZKBAM... (Ctrl+C to stop) -[5:28:03 PM] ★ repo_setup_complete: branch=bgagent/01KZCSVRHZXVHXQ4T29XYZKBAM/add-a-codeowners-file-at-the-repository-root-conta build_before=False -[5:28:04 PM] ★ step:implement:start -[5:28:15 PM] ★ step:implement:failed -[5:28:15 PM] ★ agent_execution_complete: status=error turns=0 -Task 01KZCSVRHZXVHXQ4T29XYZKBAM failed. timeout: Build/tests didn't finish in time (timed out) — Workflow run_agent step failed: TimeoutExpired: Command '['claude', '--version']' timed out after 10 seconds -``` - -**Time-to-RUNNING: ~6 s.** Submitted 00:27:09Z, `started_at` -`2026-08-07T00:27:11`, VM `startedAt` 00:27:12Z. That is materially faster than -AgentCore or ECS cold start and is the backend's main selling point — worth -recording as the one positive performance result of the run. - -**`/run` accepted the envelope, and `platform_config` was installed** — the -required observation, from the guest log group: - -``` -[server/run-pre-config] /run hook received: microvm_id='microvm-8ad29e93-99c2-3ccc-b079-12f8bdab2936' bytes=2120 -[server/debug] /run hook installed platform_config env: ['AGENT_SESSION_ROLE_ARN', 'ANTHROPIC_DEFAULT_HAIKU_MODEL', 'ARTIFACTS_BUCKET_NAME', 'AWS_SDK_UA_APP_ID', 'GITHUB_TOKEN_SECRET_ARN', 'LOG_GROUP_NAME', 'NUDGES_TABLE_NAME', 'TASK_APPROVALS_TABLE_NAME', 'TASK_EVENTS_TABLE_NAME', 'TASK_TABLE_NAME', 'TRACE_ARTIFACTS_BUCKET_NAME'] -[server/debug] /run hook accepted task_id='01KZCSVRHZXVHXQ4T29XYZKBAM' microvm_id='microvm-8ad29e93-99c2-3ccc-b079-12f8bdab2936' -``` - -Everything in the P2 delivery design is confirmed here: **2,120 bytes**, so the -envelope went **inline** and stayed under the real 4,096-byte cap (P1 F6); the -installed set is **exactly the 11 available keys**, names only, no values; the -pre-install line is stdout-only (`run-pre-config`) and the post-install line is -the first to reach CloudWatch, exactly as the snapshot-credential-hygiene work -intended. - -**Progress events streamed** — `repo_setup_complete`, `step:implement:start`, -`step:implement:failed`, `agent_execution_complete` all arrived live in `watch`. - -**Clone → change → PR got exactly one step:** clone ✓, branch ✓, -`build_before=False` ✓ … then `implement` died at turn 0. **No commit, no push, -no PR.** - -**Heartbeats: NOT observed.** `agent_heartbeat_at` was `None` at every poll -across all six submissions. The task was `RUNNING` for only ~12 s and the agent -bumps the heartbeat every 45 s, so it never had a chance to fire. **The -dual-signal liveness path is therefore NOT discharged** — see Skipped. - -### 4.3 P2-F5 — `claude --version` times out in the guest - -``` -[00:28:04] AGENT claude-agent-sdk version: 0.2.110 -[00:28:15] ERROR step 'implement' handler raised: TimeoutExpired: Command '['claude', '--version']' timed out after 10 seconds -[00:28:15] WORKFLOW step 'implement' failed (on_failure=fail) — workflow FAILED -``` - -`METRICS_REPORT`: `"turns": 0, "duration_s": 12.7, "code_changed": null, -"pr_url": null, "memory_written": true`. - -The probe is `agent/src/runner.py:476`: - -```python -["claude", "--version"], capture_output=True, text=True, timeout=10 -``` - -Characterisation, to separate "broken binary" from "slow substrate": - -- **In the identical image, locally under finch: `claude --version` → `2.1.191 - (Claude Code)` in under 1 second.** So the binary, its symlink and `PATH` are - all fine. -- `/usr/bin/claude` → `../lib/node_modules/@anthropic-ai/claude-code/bin/claude.exe`, - a **236,305,136-byte (225 MiB) statically-linked ELF**. -- The MicroVM had been restored from a snapshot ~50 s earlier and had already - done a full `git clone` over the network. - -The consistent reading is **lazy snapshot hydration**: the first `exec` of a -225 MiB binary that was never touched before the snapshot was taken must fault -its pages in from lazily-restored storage, and that exceeds 10 s. Two -independent fixes suggest themselves, and the second is the more interesting -because P2 already built the mechanism and then deliberately declined to use it: - -1. Raise / make backend-aware the `timeout=10` (it is a *liveness probe for a - version string* — a tight bound buys nothing). -2. **Warm `claude` in the `/ready` build hook so it lands in the memory - snapshot.** P2's `/ready` deliberately does the minimum — `"server is up, - reporting ready for snapshot"` — and its own docstring explains that `/ready` - exists so "the snapshot is taken with a warm server". The snapshot is warm for - *uvicorn* and stone cold for the 225 MiB binary that does all the work. - -**`build_passed: true, lint_passed: false`** in the same report is incidental -noise from the scratch repo (`mise ERROR no tasks defined …`, `unknown command: -lint`) and unrelated to the substrate. - -### 4.4 P2-F4 — the execution role cannot write to the application log group - -Found while reading logs for F5, and independent of it. Every structured agent -log line to the log group that `platform_config` itself delivers is denied: - -``` -[server/debug/self] CloudWatch write failed: AccessDeniedException: … User: arn:aws:sts:::assumed-role/backgroundagent-dev-LambdaMicrovmComputeExecutionRo-tF8Idpc9aT0R/Lambda-microvmsExecutor-86cfecce-… is not authorized to perform: logs:CreateLogStream on resource: arn:aws:logs:us-east-1::log-group:/aws/vendedlogs/bedrock-agentcore/runtime/APPLICATION_LOGS/backgroundagent-dev:log-stream:server_debug/01KZCSVRHZXVHXQ4T29XYZKBAM because no identity-based policy allows the logs:CreateLogStream action -``` - -Same denial for the `metrics/` stream, so `METRICS_REPORT` never lands -either. P2 wired `log_group_name` into `platform_config` (making the agent -*attempt* the write) but the execution role's logs grant is scoped to -`/aws/lambda-microvms/*` only — the construct's own cdk-nag suppression says so: -*"the `MICROVM_LOG_GROUP_PREFIX/*` namespace"*. - -Not fatal — the agent degrades to stdout, which the MicroVM log group captures, -which is why this run could be debugged at all. But it is a genuine smoke-parity -hole: on `lambda-microvm`, the platform's canonical per-task observability -streams are empty, and anything reading them (rather than the guest's stdout) -sees nothing. Exactly the class of grant P2 added for Bedrock, Secrets Manager -and Memory — this one was missed. - ---- - -## Phase 5 — Post-smoke verification - -### 5.1 Finalization called `TerminateMicrovm` — PASS - -`get-microvm` on the smoke VM: - -``` -microvmId = microvm-8ad29e93-99c2-3ccc-b079-12f8bdab2936 -state = TERMINATED -stateReason= Success. -startedAt = 2026-08-06T17:27:12.144000-07:00 -``` - -and the in-guest breadcrumb fired, 200 OK: - -``` -[server/debug] /terminate hook: {"active_pipeline_threads": 0, "background_pipeline_failed": false, "event": "microvm_terminate", "microvm_id": "", "timestamp": "2026-08-07T00:28:42.926428+00:00"} -INFO: 127.0.0.1:37364 - "POST /aws/lambda-microvms/runtime/v1/terminate HTTP/1.1" 200 OK -``` - -So `/terminate` is served, returns 200, reports a clean pipeline state, and does -not write terminal task status. **Note `"microvm_id": ""`** — the hook body's -`microvmId` arrives empty, unlike `/run` where it is populated. Cosmetic, but it -defeats the hook's stated purpose of joining the guest record to the -control-plane one, and the `/run` line has to carry that correlation instead -(which it does). - -**P1 F8 reconfirmed:** the VM sat at `TERMINATED` for the whole observation -window; `ResourceNotFoundException` never arrived. `TERMINATED` must map to -completed in its own right. - -### 5.2 `NO_INGRESS` — DISCHARGED - -``` -ingressNetworkConnectors = ['arn:aws:lambda:us-east-1:aws:network-connector:aws-network-connector:NO_INGRESS'] -egressNetworkConnectors = ['arn:aws:lambda:us-east-1::network-connector:nc-d306f00f-1bd0-45ea-9457-0fcec0dab2a4'] -``` - -The ARN P2 injects is **correct and effective**: the service accepted it and -`HTTP_INGRESS` (P1 F7's unrequested default public ingress) is **gone**. P1's -security finding is mitigated. - -**One caveat worth carrying forward:** a `NO_INGRESS` VM **still returns a public -endpoint hostname** (`.lambda-microvm.us-east-1.on.aws`). -So the presence of an `endpoint` in a `RunMicrovm` response is *not* evidence of -reachability, and any doc or alarm that treats "endpoint exists" as "ingress is -open" will be wrong in both directions. - -### 5.3 Dual-signal liveness — NOT DISCHARGED - -`agent_heartbeat_at` stayed `None` throughout. The heartbeat interval is 45 s and -no task was `RUNNING` for more than ~13 s, because F5 kills every task at turn 0. -**This item is blocked behind P2-F5, not independently testable.** The -`heartbeatLivenessApplies` switch is `RUNNING`-scoped and never got a -long-enough `RUNNING` window to exercise. - ---- - -## Phase 6 — Lifecycle extras - -### 6.1 A run-hook 4xx **self-terminates** the VM — supersedes P1 F9 - -Launching a second MicroVM from the image with **no** `runHookPayload` (the image -has `run: ENABLED`) produced a result P1 could not see, because P1's only -creatable image was hook-less: - -``` -state = TERMINATED -stateReason = Run lifecycle hook returned HTTP status 400. Please check your hook endpoint and application logs for more details. -``` - -Terminal within ~12 s of launch, and `suspend-microvm` then correctly refused: -`The MicroVM … has been terminated and its state cannot be changed.` - -**P1 F9 said "hook-less MicroVMs run indefinitely and bill", and warned that -nothing self-cleans.** With `/run` ENABLED that is no longer true: **the service -reaps a VM whose run hook returns 4xx.** This materially improves the cost -posture and is a direct benefit of declaring hooks. Active `TerminateMicrovm` -remains correct for the *success* path, but the *failure* path now has a -service-side backstop. - -Incidental confirmation: `executionRoleArn` must be a **full ARN** — -a bare role name is rejected with -`Member must satisfy regular expression pattern: arn:aws[a-z\-]*:iam::[0-9]{12}:role/?…`. - -### 6.2 Suspend / resume on a hook-ENABLED image — PASS - -To get a VM that stays `RUNNING`, a **hand-built valid envelope** (713 bytes: -`agent_payload` with `task_id`/`repo_url`/`task_description`/`resolved_workflow` -plus the four required `platform_config` keys) was passed as `runHookPayload`. -`/run` returned 200 and the VM held `RUNNING`. This is itself a useful result: -**the documented envelope shape is reproducible by hand from the contract alone.** - -| Step | Latency | State | Notes | -|---|---|---|---| -| launch → `RUNNING` | ~11 s | `RUNNING`, `stateReason` `None` | | -| `suspend-microvm` | **~2 s** | `SUSPENDED` | `/suspend` hook **DISABLED** in the image | -| `resume-microvm` | **~3 s** | `RUNNING` | `/resume` hook **DISABLED** | -| `terminate-microvm` | **~3 s** | `TERMINATED`, `stateReason` `Success.` | `/terminate` hook fired, 200 OK | - -`microvmId` **and** `endpoint` -(`.lambda-microvm.us-east-1.on.aws`) were -**byte-identical** before and after the suspend/resume cycle, and `startedAt` / -`maximumDurationInSeconds` (28800) never moved. - -**This extends P1 5.3/5.6 to a hook-enabled image:** suspend and resume work -*without* `/suspend` and `/resume` being declared, so P3's interface widening is -not gated on the hooks — a stored `SessionHandle` survives a cycle here too. -`SUSPENDING` was again never observable. - -### 6.3 Quota — re-probed, still NOT OBSERVABLE - -`L-CD1C0CC4` "Max allocated MicroVM memory" = **1024 Gigabytes**, and critically -`UsageMetric` = **`null`**. `AWS/Usage` still exposes **only** `CallCount` per API -name (`GetMicrovm`, `CreateMicrovmImage`, `ListMicrovmImages`, …) and **no memory -metric in any namespace**. **P1 5.5's verdict is unchanged: the claim "a -suspended VM still holds account memory quota" remains NOT OBSERVABLE SAFELY and -undischarged.** Proving it would still need ~128 concurrent 8 GiB VMs. - ---- - -## Phase 7 — Scratch repo - -**Nothing to clean up, and no PR to leave open.** `dreamorosi/batch-sync-triage` -is byte-for-byte as it was found: - -- Branches: `main` + the five pre-existing `dependabot/*`. **No `bgagent/*` - branch was ever pushed** — the agent created the branch locally in the guest - and died at `implement` before any commit or push. -- PRs: the same five open dependabot PRs (#1–#5) from 2025-11-24. **No PR was - created by this run.** - -So the instruction to leave the CODEOWNERS PR open is moot: **P2-F5 prevented any -PR from existing.** That absence is the single most important line in this -document. - ---- - -## Phase 8 — Teardown (executed as a finally-block) - -### 8.1 MicroVMs — all TERMINATED - -Nine MicroVMs existed across the run (six from the six task submissions, plus -the two Phase 6 lifecycle VMs and one orchestrator retry). **All nine were -already `TERMINATED`** at teardown — every one either finalized by the -orchestrator's `TerminateMicrovm`, service-reaped after a run-hook 4xx (6.1), or -explicitly terminated in 6.2. No VM needed chasing, and none ever approached the -8 h bound. - -### 8.2 Image and versions — deleted - -Only one version existed (`1.0`, `ACTIVE`), so `delete-microvm-image` alone was -sufficient and reaped it. `list-microvm-images` → `{"items": []}`; -`get-microvm-image` → -`ResourceNotFoundException: MicroVMImage not found for MicroVMImageID: …`. -P1's ordering correction (delete all but the last version, then the image) was -therefore not exercised, but is not contradicted. - -*Incidental:* one `ListMicrovmImages` call returned `502 Bad Gateway (reached max -retries: 2)` while the image was `DELETING`; it succeeded 30 s later. Worth -retrying rather than treating as a failure. - -### 8.3 Live IAM workarounds — reverted - -- Temporary unconditioned `iam:PassRole` inline policy - (`abca645p2-verification-passrole`) **deleted** from the orchestrator role. -- The execution role's trust policy **restored** to its deployed form, i.e. with - both `aws:SourceAccount` conditions back (verified by re-reading it). The - defect is left exactly as the branch produces it. - -### 8.4 Cognito user — deleted - -`bgagent admin delete-user ` → -`✓ Deleted Cognito user`. `list-users` then returned **empty**. The local invite -file containing its password was `rm`'d. The GitHub-token secret was left to die -with the stack (it did). - -### 8.5 Stack — DELETE_FAILED, 96/100 deleted, identical to P1's residual - -Three delete attempts, and each failed differently — that progression is itself -the finding: - -| Attempt | Duration | Outcome | -|---|---|---| -| 1 | 00:55:35Z → 01:04:28Z (8 min 53 s) | `DELETE_FAILED` — `AWS::BedrockAgentCore::Runtime`: `"Request timed out while deleting AWS::BedrockAgentCore::Runtime"`, `HandlerErrorCode: NotStabilized` | -| 2 (the briefed retry) | 01:04:55Z → 01:06:00Z (1 min) | `DELETE_FAILED` — same resource, **different error**: `"Access denied for operation 'AWS::BedrockAgentCore::Runtime'."`, `HandlerErrorCode: AccessDenied` | -| 3 (informed) | 01:07:02Z → 01:24:47Z (17 min 45 s) | `DELETE_FAILED` — but the Runtime, Memory and both IAM roles **did** delete; only the VPC set remains | - -Attempt 3 was not a blind third retry: between attempts, -`bedrock-agentcore-control list-agent-runtimes` returned **empty**, proving the -runtime was already gone server-side and the failures were a handler -stabilization/authorization artifact rather than a real leftover. Acting on that -evidence took the residual from 7 resources to 4. - -**Final residual — 4 resources, all zero-cost, exactly P1's set:** - -``` -CREATE_COMPLETE AWS::EC2::VPC vpc-072fddf653ccdcfc4 -DELETE_FAILED AWS::EC2::Subnet subnet-0e6a6a0ed18100c8a -DELETE_FAILED AWS::EC2::Subnet subnet-02d91450e51cf72a0 -DELETE_FAILED AWS::EC2::SecurityGroup sg-00997a58c1f4c5775 -``` - -``` -resource sg-00997a58c1f4c5775 has a dependent object (Service: Ec2, Status Code: 400 …) -Resource handler returned message: "The subnet 'subnet-0e6a6a0ed18100c8a' has dependencies and cannot be deleted. (Service: Ec2, Status Code: 400 …)" HandlerErrorCode: InvalidRequest -``` - -Cause: the same two AgentCore-managed ENIs, still `in-use` — -`eni-04911e11e08d670f9` and `eni-04a1c5a27966ee08b`, both -`InterfaceType: agentic_ai`. **This is #702 / P1 F12 reproducing verbatim.** -Nothing was force-deleted past CloudFormation. P1's experience says these are -eventually released (its stack is gone now — see 0.1), so the retry command is: - -```bash -aws cloudformation delete-stack --stack-name backgroundagent-dev -aws cloudformation wait stack-delete-complete --stack-name backgroundagent-dev -``` - -### 8.6 Billing confirmed stopped - -| Check | Result | -|---|---| -| NAT gateways in the ABCA VPC | **none** (the 2 `available` ones belong to pre-existing `vpc-01c9984d163d2965e` and were not touched — same as P1) | -| VPC endpoints in the ABCA VPC | **none** | -| ABCA S3 buckets | **none** | -| Unattached (billable) EIPs | **none** | -| MicroVM images / non-terminated VMs | **none** | -| AgentCore runtimes / memories | **none** | -| `/aws/lambda-microvms/*` log groups | **none** | - -### 8.7 Environment restored — nothing global left modified - -- `cdk/cdk.context.json` **restored** to the original six-AZ list (gitignored - build artifact; `git status` shows only `docs/verification/` and the - pre-existing `opencode.json`). -- `~/.finch/config.json` **restored** to `credsStore: osxkeychain` with **no - stored credential**, and the Lima VM's `DOCKER_CONFIG` restored to - `{"credsStore":"finchhost"}`; the `/root/.docker/config.json` written during - the push investigation was removed. **The short-lived ECR token written during - that investigation is no longer on disk anywhere.** -- The ECR tag this run added (`34fbc1d4…`) was removed with - `batch-delete-image`, leaving the three pre-existing tags and the underlying - manifest exactly as found. -- The finch VM was stopped, as P1 did. -- **No AWS global configuration was read or modified at any point** - (`~/.aws/config`, `~/.aws/credentials` untouched); `~/.cdk.json` was never - created. All credential and Region selection was via environment variables in - the run's own shell. - -### 8.8 Deliberately retained - -1. **`CDKToolkit`** — `UPDATE_COMPLETE`, `ComputeTypes = agentcore,lambda-microvm`, - `BootstrapVariant = ABCA: Least-Privilege Bootstrap`, five ABCA policies, no - `AdministratorAccess`. Unchanged by this run. ⚠️ Still the shared-account - caveat P1 raised: other CDK apps in `` deploy through the - ABCA-scoped execution role. -2. **`backgroundagent-dev` in `DELETE_FAILED`** — the 4 zero-cost resources in 8.5. -3. **Bootstrap S3/ECR assets** — normal bootstrap content, including the three - pre-existing agent container images. -4. **Service-vended log groups** created outside CloudFormation - (`/aws/bedrock-agentcore/runtimes/…`, `/aws/lambda/backgroundagent-dev-…`). -5. **`dreamorosi/batch-sync-triage`** — untouched (Phase 7). - -### 8.9 Scratch repo — verified unchanged - -Branches: `main` + the five pre-existing `dependabot/*`. Open PRs: #1–#5, all -dependabot. **No `bgagent/*` branch, no smoke PR.** There was nothing to close, -nothing to delete, and — because of P2-F5 — nothing to leave open for the -operator to look at. - ---- - -## Findings summary - -Live run 2026-08-06/07, account ``, `us-east-1`, branch -`feat/645-lambda-microvm-p2` @ `3a4b61a9`. Evidence: -`/tmp/abca-645-p2-20260806`. - -**Verdict on Stage D's primary objective: the smoke did NOT pass. No pull request -was created.** The run got to `clone ✓ → branch ✓ → implement ✗ (turn 0)`, -reproducibly, and it took four separate live workarounds to get even that far. -Every one of the five defects below is invisible to synth, unit tests and -`cdk-nag`; all five are live-service contract mismatches, which is precisely what -Stage D exists to surface. - -### Items DISCHARGED (behaved as designed) - -1. **`microvmImageHooks` spelling — both directions.** The API accepted - `microvmImageHooks:{ready,validate}` + `microvmHooks:{run,terminate}` with - timeouts, and CloudFormation resolved the identical property path - (`…/Hooks/MicrovmImageHooks/Ready`). The open P2 item is closed. -2. **All four hooks are served and exercised.** `/ready` 200 and `/validate` 200 - (`python=3.13.13, platform_config_keys=13, warnings=0`) during the build, on - both chipsets; `/run` 200 with the envelope; `/terminate` 200 at teardown. - **P1 F1 — "a P1 image that declares `/run` is not creatable at all" — is - fully fixed.** -3. **`NO_INGRESS` ARN.** Correct, accepted, and it suppresses P1 F7's default - public `HTTP_INGRESS`. (Caveat: a public endpoint hostname is still returned.) -4. **`platform_config` delivery end to end.** 2,120-byte envelope inline under - the real 4,096-byte cap; exactly the 11 available keys installed; names-only - logging; pre-install logging correctly stdout-only. -5. **Image buildability.** 4 min 35 s, two builds per version, `ACTIVE`; 8192 MiB - accepted; `apt-get`/Go egress over port 80 through the dedicated build - connector. **P1 F4 (443-only SG) and F5 (32 GiB memory) are fixed.** -6. **`MICROVM_IMAGE_IDENTIFIER` is a full ARN — P1 F3 fixed**, and all six - `MICROVM_*` vars appear together (all-or-nothing, both directions). -7. **`agentPlatformConfig` wiring.** 11 of 13 keys on the orchestrator; the two - absent ones are optional per-workspace secrets, so the optional-key path is - also verified. -8. **CLI end to end.** `admin invite-user` (permanent password — no first-login - challenge to drive), `configure --stack-name`, `login`, `repo onboard - --compute-type lambda-microvm` (ComputeSubstrate gate **and** live - `ListManagedMicrovmImages` probe both pass), `platform doctor` 7/7, - `submit`, `watch` streaming progress events. -9. **Finalization terminates the VM** (`TERMINATED`, `stateReason: Success.`). -10. **Suspend/resume on a hook-enabled image** with `/suspend` + `/resume` - undeclared: ~2 s / ~3 s, `microvmId` **and** `endpoint` preserved. -11. **Time-to-RUNNING ≈ 6 s** — the backend's headline advantage, measured. -12. **P1 F12's stack half resolved**: the leaked AgentCore ENIs were eventually - released and P1's `DELETE_FAILED` stack is gone. - -### Items CONTRADICTING design assumptions — `feeds-back-to-design: YES` - -**P2-F1 + P2-F3 (ONE root cause, two symptoms, both blocking). The -`aws:SourceAccount` confused-deputy condition makes ABCA's MicroVM-facing roles -unassumable by the Lambda MicroVMs service.** - -*Symptom A — the substrate cannot deploy.* Both `AWS::Lambda::NetworkConnector` -resources `CREATE_FAILED`, deterministically, on a freshly deleted stack: - -``` -"The service is unable to assume the provided NetworkConnectorOperatorRole. Please verify the trust policy on the role. (Service: Lambda, Status Code: 400, Request ID: dbe1d2f4-dd0c-4319-8f5d-b4be4f076843)" HandlerErrorCode: InvalidRequest -``` - -Removing the condition from `ConnectorOperatorRole`'s trust → both connectors -create within a second. Note this is a **regression introduced by the P1 F2 -fix**: P1's validated probe role trusted `lambda.amazonaws.com` with **no -conditions**, and the construct's comment says trust "mirrors the build/execution -roles (`lambda.amazonaws.com` + `aws:SourceAccount`)" — which is exactly what -breaks it. - -*Symptom B — no task can start.* `RunMicrovm` fails with a misleading -`iam:PassRole` denial **on the caller**, even though the orchestrator's grant is -present and `simulate-principal-policy` returns `allowed` and there is no -permissions boundary. Proven by elimination (unconditioned `iam:PassRole` still -denied; 3-minute wait still denied); removing the **execution role's** trust -conditions made the very next submission reach `RUNNING` in 6 s. - -Both the build role and the execution role carry the same pattern, so the fix is -one decision applied consistently: **drop `aws:SourceAccount` from every -MicroVM-facing role's trust policy, or find the condition key the service does -present** (`aws:SourceArn`/`aws:SourceAccount` are simply not populated on this -path). Note the *identity-side* `iam:PassedToService: lambda.amazonaws.com` -condition on the orchestrator was never shown to be wrong — it was exonerated by -step 1 and can stay. - -**P2-F2. The `CfnMicrovmImage` L1 is rejected by CloudFormation on five values — -the CDK-managed image path does not work at all.** Verbatim early-validation -output is in 1.5. `arm64` must be `ARM_64`; all four hooks must be -`ENABLED`/`DISABLED`, not paths. This **refutes the construct's stated reasoning** -that the L1 could keep CloudFormation's "path/string shape" because the generated -types "document no architecture/hook allowed-value constraint" — the service -enforces the enum at change-set time. Consequences: (a) the documented -"CDK-managed (recommended)" bootstrap path in -`cdk/scripts/package-microvm-artifact.sh` is currently non-functional and the -"out-of-band alternative" is the *only* working path; (b) hook paths are not -configurable on either surface, so the `*_HOOK_PATH` constants are agent route -constants only and must never be sent as property values. - -**P2-F4. The MicroVM execution role cannot write to the application log group -that `platform_config` tells the agent to use.** `logs:CreateLogStream` denied on -`/aws/vendedlogs/bedrock-agentcore/runtime/APPLICATION_LOGS/backgroundagent-dev` -for both the `server_debug/` and `metrics/` streams, because -the role's logs grant is scoped to `/aws/lambda-microvms/*`. P2 delivered -`log_group_name` (so the agent *attempts* the write) without the matching grant — -the same omission class the P2 Bedrock/Secrets/Memory grants were added to fix. -Non-fatal (stdout fallback lands in the MicroVM log group) but the canonical -per-task observability streams are empty on this backend, including -`METRICS_REPORT`. - -**P2-F5. `claude --version` times out after 10 s inside the MicroVM, failing -every task at turn 0. This is the blocker that prevented the PR.** Reproduced on -two consecutive submissions: - -``` -TimeoutExpired: Command '['claude', '--version']' timed out after 10 seconds -``` - -`agent/src/runner.py:476` uses `timeout=10`. The binary is fine: in the identical -image locally it answers `2.1.191 (Claude Code)` in **under 1 s**. It is a -**225 MiB (236,305,136-byte) statically-linked ELF** whose pages were never -touched before the snapshot was taken, exec'd on a guest restored ~50 s earlier — -consistent with lazy snapshot hydration. Two fixes: raise/make backend-aware the -timeout (a version-string probe gains nothing from a tight bound), and — more -interestingly — **warm `claude` in the `/ready` build hook**, which exists -precisely so "the snapshot is taken with a warm server" but currently warms only -uvicorn while leaving the 225 MiB binary that does all the work cold. - -**P2-F6. P1 F9 is superseded, in the safe direction.** With `run: ENABLED`, a -run-hook 4xx makes the **service** terminate the VM -(`stateReason: "Run lifecycle hook returned HTTP status 400."`) within ~12 s. P1's -"hook-less MicroVMs run indefinitely and bill; nothing self-cleans" no longer -describes this image. Worth correcting in ADR-021, because it changes the -cost-risk argument for the failure path. - -**P2-F7. Template size is at 98.4 % of the CloudFormation 1 MB limit** -(983,796/1,000,000) and 486/500 resources with the image configured. ~16 KB of -headroom — roughly one more construct before deploys start failing for reasons -unrelated to MicroVMs. - -**P2-F8. Runbook/tooling corrections (lower severity, but each would corrupt a -future pass).** - -- **P1 F14's `--no-paginate` remedy is wrong and silently truncating.** At 485 - resources it returns only the first page, resolved `ORCHESTRATOR_FN` to the - literal `None`, and produced `Function not found: …:function:None`. Let the CLI - paginate and filter client-side. -- **P1 F10's "recurs forever" half is wrong**: `BootstrapVariant` now reads - `ABCA: Least-Privilege Bootstrap`, because P1's own - `update-stack --use-previous-template` with only `ComputeTypes` supplied resets - unspecified parameters to template defaults. The silent-`exit 0` half stands. -- **`--execution-role-arn` requires a full ARN** (regex in 6.1); a bare role name - is rejected. -- **`zsh`**: `"$VAR:latest"` is the `${VAR:l}` lowercase modifier (cost one - mis-diagnosed push); no `PIPESTATUS` (it is `$pipestatus`, 1-indexed), so P1's - `test "${PIPESTATUS[0]}" -eq 0` silently evaluates empty; no `timeout(1)`. -- **Verify secrets on the raw value, never a `.strip()`ed copy** — a trailing - newline from `gh auth token` produced a convincing false clone defect (3.2). -- **`finch` cannot push to ECR on this box**: `finch push` runs - `limactl shell finch sudo -E nerdctl push` and the VM's `DOCKER_CONFIG` uses - the `finchhost` creds helper, which cannot resolve an Isengard - `credential_process`; a host-side `finch login` does not help. Both `push` and - `pull` fail with `no basic auth credentials`. -- The `/terminate` hook body's **`microvmId` arrives empty** (`"microvm_id": ""`), - defeating the hook's stated guest↔control-plane correlation purpose. -- **`ARTIFACTS_BUCKET_NAME` and `TRACE_ARTIFACTS_BUCKET_NAME` resolve to the same - bucket**, making that separation notional. - -### Items SKIPPED, BLOCKED, or INCONCLUSIVE - -| Item | Verdict | Reason | -|---|---|---| -| **clone → change → PR (the point of Stage D)** | **FAILED — no PR** | P2-F5, reproduced twice. Stopped at two retries as briefed. | -| Dual-signal liveness / `agent_heartbeat_at` fresh during RUN | **BLOCKED behind P2-F5** | Heartbeat interval is 45 s; no task stayed `RUNNING` beyond ~13 s. `agent_heartbeat_at` was `None` on all six submissions. Not independently testable until F5 is fixed. | -| Suspend TTL beyond P1's ≥1 h floor | **SKIPPED (time-boxed)** | Would have added 2 h+ of wall clock for a marginal bound after ~2.5 h already spent on five blocking defects. P1's result stands unchanged: **survives ≥ 1 h 0 min 17 s with no `idlePolicy`**; a TTL between 1 h and the 8 h `maximumDurationInSeconds` bound remains **OPEN**. | -| SUSPENDED VM consumes account memory quota | **STILL NOT OBSERVABLE — undischarged** | Re-probed: `L-CD1C0CC4` = 1024 GB with `UsageMetric: null`; `AWS/Usage` has only `CallCount`; no memory metric in any namespace. Identical to P1 5.5. | -| CDK-managed `CfnMicrovmImage` deploy | **REFUTED (P2-F2)** | Fell back to the out-of-band script path, as the brief directed. | -| AgentCore container asset push | **WORKED AROUND** | finch/ECR auth (deviation 2). ECR manifest retagged; irrelevant to the MicroVM image under test. | -| P1 F11 (AgentCore-unsupported AZ) | **STILL UNFIXED** | `agent-vpc.ts` unchanged; worked around via the gitignored AZ context cache, restored at teardown. | -| P1 F12 Memory-delete half | **STILL UNFIXED** | `AWS::BedrockAgentCore::Memory` still undeletable while `CREATING`; still turns a rollback into `ROLLBACK_FAILED`; plain `delete-stack` still clears it. | -| Orchestrator-role IAM negatives | **NOT ATTEMPTED** | P1 5.0 established the Lambda trust does not allow operator assumption; trust was not modified for this purpose. | - -### Elapsed and approximate cost - -**Elapsed:** 22:38Z → 01:26Z ≈ **2 h 48 min**. Roughly: ~35 min on the finch/ECR -push dead end and the ECR-retag workaround; ~40 min on the five deploy attempts -plus two `ROLLBACK_FAILED`/`delete-stack` cycles; ~10 min on image create+build; -~35 min on the six submissions and the P2-F3 elimination sequence; ~10 min on -lifecycle extras; ~30 min on teardown (three delete attempts); the remainder on -evidence capture and this document. - -**Approximate cost: well under US$5**, dominated as in P1 by NAT/VPC-endpoint -hours rather than MicroVMs. - -| Item | Quantity | Est. | -|---|---|---| -| NAT gateway (1 × $0.045/h) | ~1.6 h across 3 stack lifetimes | ~$0.08 | -| Interface VPC endpoints (7 × $0.01/h × 2 AZ) | ~1.6 h | ~$0.22 | -| MicroVM runtime | 9 VMs, all short-lived (~12 s to ~3 min each); longest suspended window ~1 min | < $0.10 | -| MicroVM image builds | 2 builds (1 version × 2 chipsets), ~4.5 min each | low single-digit cents | -| Snapshot/image storage | ~3.6 GiB × ~1 h | negligible | -| AgentCore runtime + Memory | created 3×, **never invoked** | negligible | -| Bedrock | **zero model tokens** — every task died before turn 1 | $0 | -| S3 / DynamoDB / Lambda / API GW / Cognito / Secrets / logs | brief, mostly idle | < $1 | - -The 8 h `maximumDurationInSeconds` was never approached; every VM was either -explicitly terminated or service-reaped. - -### Recommended follow-up before P2 is called complete - -Ordered by what unblocks what: - -1. **P2-F1/F3** (`aws:SourceAccount` on MicroVM-facing role trust) — nothing - deploys or runs without this. One decision, three roles. -2. **P2-F5** (`claude --version` timeout) — nothing *completes* without this. It - is the only thing between this run and a PR, and the `/ready`-warming option - is worth considering on its merits rather than just raising the timeout. -3. **P2-F2** (L1 enum values) — the documented recommended path is dead until - fixed; five one-word changes plus the tests that assert the old strings. -4. **P2-F4** (application-log-group grant) — cheap, and it is what makes the next - failure debuggable through the platform rather than through guest stdout. -5. **P2-F7** (template size) — unrelated to MicroVMs but it will bite soon. -6. Re-run Stage D after 1–2. The dual-signal-liveness item and the suspend-TTL - extension both need a task that stays `RUNNING` for minutes, which only F5's - fix provides. - ---- - -## Stage D-redux (run 2) - -Narrow re-run on branch `feat/645-lambda-microvm-p2` @ -`b927d1d6a58e2b040c2ed4ce4e9f1dc9be9fc981` — the commit that fixed P2-F1..F5 -against run 1's evidence. **Purpose: convert "fixed-against-evidence" into -"re-exercised live", and get the pull request.** - -**Live run:** 2026-08-07 02:58Z → 04:55Z (≈1 h 57 min), account ``, -`us-east-1`. Evidence directory: `/tmp/abca-645-p2r2-20260806`. - -**Teardown: complete.** Live IAM workaround reverted, Cognito user deleted, image -+ version deleted, ECR retag removed, AZ context cache restored, 481/485 stack -resources deleted (4 zero-cost residual, #702), **all billable resources confirmed -gone**, and no global/host configuration touched. Both smoke PRs left open as -briefed. See 2.12. - -> **Provenance note — HEAD moved mid-run, from outside this run.** Preflight -> confirmed `HEAD = b927d1d` with a clean tree at 02:58Z. At **03:04Z** an -> unrelated user commit landed on the branch — `045722c fix(deps): refresh -> js-yaml lock entry to 4.3.1`, **`yarn.lock` only, 3 insertions / 3 deletions**. -> No `git` write command was issued by this run, and `b927d1d` remains an -> ancestor of `HEAD`. **It cannot have affected any finding:** `mise run install` -> / `yarn install` was never re-run, so `cdk/node_modules` still reflected -> `b927d1d`'s lock for every synth and deploy; the MicroVM artifact is built from -> `agent/` + `contracts/` + `agent/Dockerfile` (Python/uv), which the commit does -> not touch. Every verdict below is therefore against `b927d1d`'s tree. - -### 2.0 Headline - -> **THE SMOKE PASSED. `https://github.com/dreamorosi/batch-sync-triage/pull/6`** -> — clone → change → commit → push → PR, `COMPLETED`, 12 turns, $0.279, 153 s. -> Run 1's single most important line ("no PR was created") is retired. - -But the PR required **one live IAM workaround**, and establishing *why* produced -the run's most consequential result: **P2-F3 is NOT fixed, and run 1's -exoneration of its identity-side condition was a false negative.** Two of the -five P2 fixes are fully discharged, two are discharged, one is refuted, and one -brand-new blocking defect was found on the path run 1 never reached. - -| Fix | Run-1 verdict | Run-2 live verdict | -|---|---|---| -| **P2-F1** (connector trust) | blocking | ✅ **DISCHARGED** — substrate deployed **first try**, zero workarounds, 0 `CREATE_FAILED` in 485 resources | -| **P2-F2** (`ARM_64`/`ENABLED` enums) | REFUTED at change-set validation | ✅ **DISCHARGED** — early validation **passed**, resource reached `CREATE_IN_PROGRESS` | -| **P2-F3** (`RunMicrovm` PassRole) | blocking; "trust was the sole cause" | ❌ **NOT FIXED (P2r2-F10)** — isolated to the *identity-side* `iam:PassedToService` condition run 1 explicitly exonerated | -| **P2-F4** (application-log grant) | blocking observability | ✅ **DISCHARGED** — `server_debug/`, `metrics/` **and** `trajectory/` streams exist, with content | -| **P2-F5** (`claude` warm-up) | **the blocker** — no PR | ✅ **DISCHARGED** — cold `claude` measured at **17–38 s**, warm **0.1 s**, `/ready` 200, no 503 | -| *(new)* **P2r2-F9** | not reachable in run 1 | ❌ CDK-managed image path blocked: bootstrap `IAMPassRole` denies the build role to CloudFormation | -| *(new)* **P2r2-F11** | mis-attributed in run 1 | ⚠️ `agent_heartbeat_at` is never projected into the API response | -| Dual-signal liveness | BLOCKED behind F5 | ✅ **DISCHARGED** — 45 s cadence observed live over a 181 s `RUNNING` window | - -### 2.1 Deltas from run 1's setup - -```bash -export SMOKE_USER="" -export EVIDENCE_DIR=/tmp/abca-645-p2r2-20260806 -export BGAGENT_CONFIG_DIR=/tmp/abca-645-p2r2-bgagent -``` - -Everything else — the zsh traps, `set -o pipefail`, explicit `AWS_REGION`, -newest-first version ordering — carried over unchanged and all of it still -applies. Two run-1 notes paid for themselves immediately: the **raw-secret -assertion** (2.6) and **client-side pagination** for `TaskOrchestrator` lookup. - -`bgagent` was invoked as `node cli/lib/bin/bgagent.js` after -`mise //cli:compile` (there is no linked binary in this tree). - -### 2.2 Preflight — run 1's residual had to be cleared first, and it did not clear itself - -`backgroundagent-dev` was still `DELETE_FAILED` with run 1's exact 4-resource -residual, and **the two `agentic_ai` ENIs were still `in-use` 1 h 32 min later** -(`eni-04911e11e08d670f9`, `eni-04a1c5a27966ee08b`, both requester `amazon-aws`, -`InstanceOwnerId: amazon-aws`), while -`bedrock-agentcore-control list-agent-runtimes` and `list-memories` both returned -**empty**. So the ENIs outlive the resources that created them by a wide margin. - -Run 1's documented retry command was executed verbatim and **failed again after -17 min 17 s** (02:58:41Z → 03:15:58Z): - -``` -The following resource(s) failed to delete: [AgentVpcRuntimeSG96507CD0, AgentVpcPrivateSubnet1Subnet8051BB57, AgentVpcPrivateSubnet2SubnetC66971D0]. -``` - -**Correction to run 1's §8.5 advice.** Run 1 concluded from P1's experience that -"these are eventually released, so the retry command is `delete-stack`". That is -true on a multi-day horizon and **useless on a same-session horizon** — a -17-minute retry that fails identically is not a remedy. The remedy that works is -`--retain-resources`, which cleared the stack record in **33 seconds**: - -```bash -aws cloudformation delete-stack --stack-name backgroundagent-dev \ - --retain-resources AgentVpcA6796801 AgentVpcPrivateSubnet1Subnet8051BB57 \ - AgentVpcPrivateSubnet2SubnetC66971D0 AgentVpcRuntimeSG96507CD0 -``` - -Note the VPC (`CREATE_COMPLETE`, never attempted) must be listed alongside the -three `DELETE_FAILED` children or the delete fails on it. Cost: an orphaned -zero-cost VPC + 2 subnets + 1 SG, now outside CloudFormation's knowledge (2.12). - -*Unchanged from run 1:* `CDKToolkit` `UPDATE_COMPLETE`, `ComputeTypes = -agentcore,lambda-microvm`, `BootstrapVariant = ABCA: Least-Privilege -Bootstrap`, bootstrap SSM version `32`, five ABCA policies, no -`AdministratorAccess` — so **no re-bootstrap was run**. That decision turns out -to matter; see P2r2-F9. - -**Managed base image** — still exactly one in `us-east-1`, -`arn:aws:lambda:us-east-1:aws:microvm-image:al2023-1`, versions `1` (newest -first) and `0`. Selected `1`; the service echoed `baseImageVersion: "1.0"`. - -**P1 F11 is still UNFIXED** (`agent-vpc.ts` is still `maxAzs: 2` with no AZ -constraint, `us-east-1a` is still `use1-az6`), so the gitignored -`cdk/cdk.context.json` AZ cache was trimmed to lead with `us-east-1b`/`us-east-1c` -again, saved to `$EVIDENCE_DIR/cdk.context.json.ORIGINAL`, and **restored at -teardown**. `git status` stayed clean throughout. - -### 2.3 The AgentCore container asset — retag workaround reused, and the tag had moved - -`finch` still cannot push to ECR (run 1 deviation 2 — the Lima VM's -`credsStore: finchhost` cannot resolve an Isengard `credential_process`). The -required asset tag was **`7a71005f…`, not run 1's `34fbc1d4…`**, because -`b927d1d` grew `agent/src/server.py`; run 1's tag had been correctly removed at -its teardown. The same `aws ecr put-image` retag was applied to the existing -manifest `sha256:cdf5436a…`: - -``` -{"imageDigest": "sha256:cdf5436ab8d9d17bcd5b555ad22c52b4e6d6622f1fc0373cdc97a73c3eb8e6a4", - "imageTag": "7a71005fe6f520a2741cdd1fb2a47ffa919b13e63476226ec39bc66eb4c150c5"} -``` - -Same caveat, restated because it is easy to lose: **the deployed AgentCore -runtime therefore carries a stale agent image, and nothing in this run depends on -it** — the smoke runs on the MicroVM substrate built from the S3 zip. The finch VM -was never started this run, and `~/.finch/config.json` was never touched -(verified still `credsStore: osxkeychain` with no stored credential). - -### 2.4 Substrate deploy — P2-F1 DISCHARGED, first try, zero workarounds - -Synth: `abca:microvm-image-not-provisioned` emitted, `…-p1-smoke-unverified` -correctly absent. Deploy `03:19:25Z → 03:33:59Z = 14 min 34 s`, **`EXIT=0` on the -first attempt**, 485 resources. - -**This is the whole of P2-F1's verification and it is unambiguous.** Run 1 needed -five attempts and a `/tmp` cloud-assembly trust-policy patch to get here. Run 2 -needed none: - -``` -2026-08-07T03:21:47Z LambdaMicrovmComputeEgressConnector9C36AAC2 CREATE_IN_PROGRESS -2026-08-07T03:21:48Z LambdaMicrovmComputeBuildEgressConnector3B762F80 CREATE_IN_PROGRESS -2026-08-07T03:26:05Z LambdaMicrovmComputeEgressConnector9C36AAC2 CREATE_COMPLETE -2026-08-07T03:26:06Z LambdaMicrovmComputeBuildEgressConnector3B762F80 CREATE_COMPLETE -``` - -and a scan of every stack event for `*FAILED` returned **`NONE`**. The live trust -policies on both the execution and build roles read exactly as source intends — -`lambda.amazonaws.com`, `sts:AssumeRole` + `sts:TagSession`, **no conditions**. - -All seven MicroVM outputs populated; `ComputeSubstrate = lambda-microvm`; -artifact key exactly `microvm-images/agent-artifact.zip`. - -**P2-F7 reconfirmed and marginally worse:** 979,867 B / 485 resources (substrate), -**985,886 B / 486** with the image — **98.6 %** of the 1 MB limit, ~14 KB of -headroom. Down from run 1's ~16 KB. - -### 2.5 The CDK-managed image path — P2-F2 DISCHARGED, then blocked by a NEW defect (P2r2-F9) - -The synthesized L1 now carries exactly the five values CloudFormation rejected in -run 1: - -```json -"CpuConfigurations": [{ "Architecture": "ARM_64" }], -"Hooks": { - "Port": 8080, - "MicrovmHooks": { "Run": "ENABLED", "RunTimeoutInSeconds": 60, - "Terminate": "ENABLED", "TerminateTimeoutInSeconds": 15 }, - "MicrovmImageHooks": { "Ready": "ENABLED", "ReadyTimeoutInSeconds": 300, - "Validate": "ENABLED", "ValidateTimeoutInSeconds": 60 } -} -``` - -**P2-F2 is DISCHARGED.** Change-set early validation **passed** — zero -`not a valid enum value` errors, no `Early validation failed` — and the resource -progressed to `CREATE_IN_PROGRESS`. Run 1's §1.5 refutation is fully answered and -the enum fix is correct on the CloudFormation surface. - -It then failed on something run 1 could never have seen, because run 1 never got -past early validation: - -``` -LambdaMicrovmComputeImage16B48539 CREATE_FAILED -Resource handler returned message: "User: arn:aws:sts:::assumed-role/cdk-hnb659fds-cfn-exec-role--us-east-1/AWSCloudFormation is not authorized to perform: iam:PassRole on resource: arn:aws:iam:::role/backgroundagent-dev-LambdaMicrovmComputeBuildRoleF0-9FxjQbiJC3px because no identity-based policy allows the iam:PassRole action (Service: LambdaMicrovms, Status Code: 403, Request ID: f8c45ab5-49c4-42fd-ac61-a6a3ac00dc26) (SDK Attempt Count: 1)" HandlerErrorCode: AccessDenied -``` - -`UPDATE_ROLLBACK_COMPLETE`; the substrate survived intact (all seven outputs still -populated), so this cost one deploy cycle and nothing else. - -**Diagnosis — and it is NOT a stale bootstrap.** The live -`…IaCRole-ABCA-Infrastructure…` policy is **byte-identical to -`cdk/bootstrap/policies/infrastructure.json` on this branch**, so -`bootstrap --force` would change nothing. Its `IAMPassRole` statement is: - -```json -{ "Sid": "IAMPassRole", "Effect": "Allow", "Action": "iam:PassRole", - "Resource": "arn:aws:iam::*:role/backgroundagent-dev-*", - "Condition": { "StringEquals": { "iam:PassedToService": [ - "lambda.amazonaws.com", "ecs-tasks.amazonaws.com", "ecs.amazonaws.com", - "apigateway.amazonaws.com", "logs.amazonaws.com", "bedrock.amazonaws.com", - "bedrock-agentcore.amazonaws.com", "events.amazonaws.com", - "vpc-flow-logs.amazonaws.com" ] } } } -``` - -The resource pattern matches the build role. `simulate-principal-policy` on the -deploy role returns **`allowed`** with -`iam:PassedToService=lambda.amazonaws.com` and **`implicitDeny`** with no context -— so the resource is right and the *condition* is the only remaining variable. - -**The control that makes this airtight:** the out-of-band `--create-image` path -(2.6) passed **the same build role** to the same service successfully, using -operator credentials that carry no such condition. So the build role's trust is -fine and the service can assume it; the denial is genuinely caller-side. - -This directly contradicts the construct's own accounting, which asserts -`buildRole` is passed "via the bootstrap `infrastructure` policy's `IAMPassRole`" -— live, it is not. - -*Incidental:* **CloudTrail has no `lambda-microvms` events at all** for this run -(`lookup-events` across the window returns nothing for `CreateMicrovmImage` / -`RunMicrovm`). The service appears not to emit management events yet, which is why -the actual `iam:PassedToService` value cannot simply be read out of a log — every -determination in this section had to be made by elimination. - -### 2.6 Image + build — P2-F5 DISCHARGED, with numbers - -Artifact: **935,874 bytes**, SSE `AES256` (run 1: 922,056). `--create-image` exit -**0**; the service echoed `ARM_64`, `minimumMemoryInMiB: 8192`, all four hooks -`ENABLED`, and `readyTimeoutInSeconds: 300`. - -Build `03:42:53Z → 03:47:23Z = 4 min 30 s`, both chipsets -(`GRAVITON` generation `3` and `4`) `SUCCESSFUL`, `state=SUCCESSFUL`, -`status=ACTIVE`. **The warm-up cost the build ~22 s** versus run 1's 4 min 35 s -hook-less-warm-up baseline — an outstanding trade for what it buys. - -**The decisive evidence, verbatim** (per build stream): - -``` -[server/build-hook] /ready hook: server is up, warming the snapshot before it is taken -[server/build-hook] /ready hook: warmed 'claude' in 17.2s (version='2.1.191 (Claude Code)') -[server/build-hook] /ready hook: warmed 'claude' in 37.8s (version='2.1.191 (Claude Code)') -[server/build-hook] /ready hook: warmed 'git' in 0.0s (version='git version 2.47.3') -[server/build-hook] /ready hook: warmed 'claude' in 0.1s (version='2.1.191 (Claude Code)') -[server/build-hook] /ready hook: warmed 'node' in 3.0s (version='v24.19.0') -[server/build-hook] /ready hook: reporting ready for snapshot -INFO: 127.0.0.1:33946 - "POST /aws/lambda-microvms/runtime/v1/ready HTTP/1.1" 200 OK -[server/build-hook] /validate hook: ok (python=3.13.13, platform_config_keys=13, warnings=0) -INFO: 127.0.0.1:53154 - "POST /aws/lambda-microvms/runtime/v1/validate HTTP/1.1" 200 OK -``` - -**Four independent things are proven here, and three of them are new:** - -1. **The lazy-hydration diagnosis was correct, and now it is measured.** A cold - `exec` of the 225 MiB `claude` binary takes **17.1–37.8 s**. Run 1 inferred - this from a 10 s timeout; run 2 has the number. **`timeout=10` could never - have passed** — it was 2–4× short. -2. **The warm-up works.** The same binary in the same guest answers in **0.1 s** - once its pages are faulted in. That is the mechanism doing exactly what the - `/ready` docstring claims. -3. **The service issues `/ready` three times per build, and the first two run - concurrently.** Both concurrent calls warm `claude` simultaneously and contend - (17.2 s and 37.8 s in the same stream). Worst-case single call ≈ 37.8 + 0.0 + - 6.4 ≈ **44 s**. So the 120 s required budget carries ~3.2× margin and the move - from `readyTimeoutInSeconds` 60 → 300 was not merely defensive — a 60 s hook - budget would have had ~16 s of slack against a contended cold start. -4. **No 503, ever.** The "required warm-up failed" path never fired, and - `/validate` still reports `platform_config_keys=13, warnings=0`. - -*Not re-captured:* the 2.3 size table (`codeInstallSizeInBytes` etc.) — the image -was deleted at teardown before those fields were read. Run 1's figures stand; -this run adds nothing to or against P1 F13. - -**Wired deploy** with `--context microvm_image_identifier=`: -`03:48:34Z → 03:53:37Z = 5 min 3 s`, `UPDATE_COMPLETE`. - -**`MICROVM_*` is FIVE vars this run, not run 1's six**, and that is **correct by -design**: `MICROVM_IMAGE_VERSION` is emitted only when the optional -`microvm_image_version` context is supplied (`agent.ts:245` → -`task-orchestrator.ts:467`), and the construct documents its absent state as -"let the service pick". Run 1 recorded six because it supplied the version. -Not a regression — but worth stating, because "all-or-nothing" is asserted of the -`MICROVM_*` block and the version is the one member that legitimately opts out. - -`platform_config` wiring reconfirmed at **11 of 13** keys, with -`LINEAR_OAUTH_SECRET_ARN` / `JIRA_OAUTH_SECRET_ARN` correctly absent. - -**P2-F4's grant is present on the live execution role** — the second statement is -new versus run 1: - -```json -{ "Action": ["logs:CreateLogStream","logs:PutLogEvents"], - "Resource": "arn:aws:logs:us-east-1::log-group:/aws/vendedlogs/bedrock-agentcore/runtime/APPLICATION_LOGS/backgroundagent-dev:*", - "Effect": "Allow" } -``` - -### 2.7 Platform user, secret, onboarding — all PASS - -`admin invite-user` → `CONFIRMED` with a permanent password; `configure ---stack-name`; `login` → `Login successful.` Exactly run 1's two-command flow. - -**The PAT was written correctly this time** — `gh auth token | tr -d '\n\r'` into -`--secret-string file:///dev/stdin`, then asserted on the **raw** value per run -1's §3.2 lesson: - -``` -len 40 differs_from_stripped False prefix gho_ -``` - -Run 1's self-inflicted 41-byte defect did not recur. `repo onboard ---compute-type lambda-microvm` → `status: active`, both the ComputeSubstrate gate -and the live `ListManagedMicrovmImages` probe passing; `platform doctor` → **7/7**. - -### 2.8 THE SMOKE — five submissions, and the last two are an experiment - -| # | Task ID | Orchestrator `iam:PassRole` grant in force | Settle | Outcome | -|---|---|---|---|---| -| 1 | `01KZD5SZ55N707GMWY14P09JGR` | **source only** (exact ARN + `iam:PassedToService`) | ~2.5 min | `FAILED` — PassRole denial | -| 2 | `01KZD61FHYH3QT774PFETDNVKY` | + unconditioned, exact ARN | 20 s | `FAILED` — identical | -| 3 | `01KZD6D95P5VXJTJ06BFJXF8J6` | + unconditioned, `backgroundagent-dev-*` | 3.5 min | ✅ **`RUNNING` in 5 s → `COMPLETED` → PR #6** | -| 4 | `01KZD71DJSTGZK2A91MZJ8GFHJ` | **source only** (workaround removed) | **5 min** | `FAILED` — identical → **control** | -| 5 | `01KZD7D7HD3CT2APBT3DH12XFE` | + unconditioned, **exact ARN** (source's resource) | **5 min** | ✅ **`RUNNING` in 9 s → `COMPLETED` → PR #7** | - -**Submissions 4 and 5 are the isolation experiment, and they are the most -important result in this document.** Same exact-ARN resource as source, same -5-minute settle, one variable — the `iam:PassedToService: lambda.amazonaws.com` -condition. With it: denied. Without it: `RUNNING` in 9 s. See P2r2-F10. - -#### 2.8.1 THE SMOKE (submission 3), `bgagent watch`, verbatim highlights - -``` -[9:07:17 PM] ★ repo_setup_complete: branch=bgagent/01KZD6D95P5VXJTJ06BFJXF8J6/add-a-codeowners-file-at-the-repository-root-conta build_before=False -[9:07:18 PM] ★ step:implement:start -[9:08:14 PM] ★ pre_approvals_loaded {"scopes":[],"count":0} -[9:08:31 PM] Turn #1 (claude-opus-4-8, 0 tool calls) - Text: This is a simple, well-defined task. Let me create the CODEOWNERS file. -[9:08:37 PM] ▶ Write: {'file_path': '/workspace/01KZD6D95P5VXJTJ06BFJXF8J6/CODEOWNERS', 'content': '* @dreamorosi\n'} -[9:09:27 PM] ▶ Bash: git add CODEOWNERS && git commit -m "chore(github): add CODEOWNERS file" && git push -u origin bgagent/… -[9:09:43 PM] ▶ Bash: gh pr create --repo dreamorosi/batch-sync-triage --head bgagent/… --base main --title "chore(github): add CODEO… -[9:09:45 PM] ◀ Bash: https://github.com/dreamorosi/batch-sync-triage/pull/6 -[9:09:50 PM] Cost: $0.2792 (1947 in / 3554 out tokens) -[9:09:50 PM] ★ step:implement:succeeded -[9:09:50 PM] ★ agent_execution_complete: status=success turns=22 -[9:09:50 PM] ★ pr_created: https://github.com/dreamorosi/batch-sync-triage/pull/6 -Task 01KZD6D95P5VXJTJ06BFJXF8J6 completed. -``` - -Final record: `status COMPLETED`, `duration_s 153.4`, `cost_usd 0.279`, -`turns_completed 12`, `build_passed true`, `lint_passed false`, -`session_id microvm-082ebde7-…`. **Time-to-`RUNNING` 5 s**, confirming run 1's -headline performance result on a task that actually finishes. - -*Minor inconsistency worth a glance:* `watch` prints -`agent_execution_complete: turns=22` while the persisted record and -`METRICS_REPORT` both say `turns: 12`. Two different notions of "turn" (SDK -messages vs. counted iterations) surfaced under one word in the same stream. - -*Incidental, both tasks:* `lint_passed: false` is scratch-repo noise -(`tsc: not found`, `biome` schema mismatch, `mise ERROR no tasks defined`), and -the agent correctly reasoned about it as pre-existing before proceeding. - -#### 2.8.2 P2-F4 — DISCHARGED with content, not just stream existence - -The brief's test was whether `server_debug/` now *exists*. It does, and -so does more: - -``` -metrics/01KZD6D95P5VXJTJ06BFJXF8J6 -server_debug/01KZD6D95P5VXJTJ06BFJXF8J6 -server_debug/server -trajectory/01KZD6D95P5VXJTJ06BFJXF8J6 -``` - -All three per-task streams were `AccessDenied` in run 1. `server_debug` carries -the `/run` breadcrumbs — including the **11-key** `platform_config` install line, -names only — and `metrics/` carries the `METRICS_REPORT` that run 1 lost -entirely: - -```json -{"event": "METRICS_REPORT", "status": "success", "agent_status": "success", - "pr_url": "https://github.com/dreamorosi/batch-sync-triage/pull/6", - "build_passed": true, "lint_passed": false, "cost_usd": 0.27916625, - "turns": 12, "duration_s": 153.4, "task_id": "01KZD6D95P5VXJTJ06BFJXF8J6", - "disk_before": "167.5 KB", "disk_after": "261.0 MB", …} -``` - -**P2-F4 is fully discharged**, and the platform's canonical per-task -observability is no longer empty on this backend. - -#### 2.8.3 Dual-signal liveness — DISCHARGED, and run 1's verdict was partly an artifact - -Polled **against DynamoDB** during submission 5's `RUNNING` window: - -| Wall clock | Status | Running for | `agent_heartbeat_at` | Freshness | -|---|---|---|---|---| -| 04:25:11Z | `RUNNING` | 71 s | `2026-08-07T04:24:47Z` | 24 s | -| 04:25:38Z | `RUNNING` | 99 s | `2026-08-07T04:25:32Z` | **6 s** | -| 04:26:05Z | `RUNNING` | 126 s | `2026-08-07T04:25:32Z` | 34 s | -| 04:26:33Z | `RUNNING` | 153 s | `2026-08-07T04:26:17Z` | 16 s | -| 04:27:00Z | `COMPLETED` | 181 s | `2026-08-07T04:26:17Z` | 44 s | - -Successive values are `04:24:47 → 04:25:32 → 04:26:17`: **exactly the 45 s -`_HEARTBEAT_INTERVAL_SECONDS` cadence**, freshness never worse than 34 s while -`RUNNING`. **The dual-signal liveness path is DISCHARGED on `lambda-microvm`** — -`heartbeatLivenessApplies` has a real, fresh signal to read. Progress events -streamed live in `watch` concurrently (2.8.1). - -**But this only became visible by reading DynamoDB directly.** `bgagent status` -reported `agent_heartbeat_at = None` on every poll of *both* completed tasks — -including submission 3, whose stored value was `2026-08-07T04:09:39Z`, 12 s before -`completed_at`. Cause: `toTaskDetail` (`cdk/src/handlers/shared/types.ts:786`) -never maps the field, though `TaskRecord` declares it at line 95 and -`orchestrator.ts` consumes it for liveness. See P2r2-F11 — and note this makes -run 1's "heartbeats NOT observed" partly a measurement artifact rather than a -pure consequence of P2-F5. - -### 2.9 Lifecycle — PASS - -Both smoke MicroVMs finalized cleanly: `state TERMINATED`, `stateReason -Success.`, `maximumDurationInSeconds 28800` never approached. `NO_INGRESS` -reconfirmed on both, and **run 1's caveat holds** — a `NO_INGRESS` VM still -returns a public endpoint hostname -(`76cfed33-…lambda-microvm.us-east-1.on.aws`), so "an endpoint exists" remains no -evidence of reachability. - -`list-microvms` showed **11** VMs, all `TERMINATED` — run 1's 9 plus this run's 2, -so the list is cumulative across runs and none of run 1's ever resurfaced. - -### 2.10 Not attempted this run - -Deliberately out of the narrow scope: suspend/resume and suspend-TTL (run 1 §6.2 -covered a hook-enabled image), the SUSPENDED-vs-quota probe (still -`UsageMetric: null` territory), and run 1's §6.1 run-hook-4xx self-termination -(P2-F6 — already recorded, and `b927d1d` corrected the ADR for it). - -### 2.11 Scratch repo — TWO PRs left open - -``` -#7 docs(contributors): add CONTRIBUTORS.md | bgagent/01KZD7D7HD3CT2APBT3DH12XFE/… -#6 chore(github): add CODEOWNERS file | bgagent/01KZD6D95P5VXJTJ06BFJXF8J6/… -#1–#5 pre-existing dependabot PRs -``` - -**#6 is the briefed smoke and is left open as instructed.** **#7 is a by-product -of the 2.8 isolation experiment** (submission 5 needed a task that would actually -run, and a second distinct file avoided colliding with #6) and is left open -alongside it rather than closed, so the evidence for the experiment survives. -Two `bgagent/*` branches were pushed; `main` is untouched. - -### 2.12 Teardown (executed as a finally-block) - -| Step | Result | -|---|---| -| **Live IAM workaround** | `abca645p2r2-verification-passrole` **deleted** from the orchestrator role; `list-role-policies` shows only the CDK-managed default policy. **No trust policy was modified at any point this run** (verified: execution role still `sts:AssumeRole` + `sts:TagSession`, conditionless, exactly as source produces). | -| **MicroVMs** | All 11 `TERMINATED` before teardown began; none needed chasing. | -| **Image + version** | Single version `1.0`; `delete-microvm-image` → `DELETING`, reaped it. | -| **Cognito user** | `admin delete-user` → `✓ Deleted`; `list-users` → `[]`; invite file `rm`'d. | -| **Secret** | Left to die with the stack. | -| **`cdk/cdk.context.json`** | **Restored** to the original six-AZ list; `git status` shows only `docs/verification/` + pre-existing `opencode.json`. | -| **ECR retag** | `7a71005f…` removed with `batch-delete-image`; pre-existing tags and the underlying manifest untouched. | -| **finch** | Never started this run; `~/.finch/config.json` never touched (still `credsStore: osxkeychain`, no stored credential). | -| **Global/host config** | **`~/.aws/*` never read or modified.** All credential and Region selection via environment variables in the run's own shell. `~/.cdk.json` never created. | - -**Stack — two delete attempts, ending at run 1's exact residual:** - -| Attempt | Window | Outcome | -|---|---|---| -| 1 | 04:29:02Z → 04:37:24Z (8 min 22 s) | `DELETE_FAILED` — `AWS::BedrockAgentCore::Runtime`: `"Request timed out while deleting AWS::BedrockAgentCore::Runtime"`, `HandlerErrorCode: NotStabilized`. **19** resources left. | -| 2 (informed) | 04:37:50Z → 04:55:17Z (17 min 27 s) | `DELETE_FAILED` — but the Runtime, Memory, all IAM roles, all DynamoDB tables, the S3 bucket and the secret **did** delete. **4** resources left. | - -Attempt 2 was not a blind retry: `bedrock-agentcore-control list-agent-runtimes` -returned **empty** first, proving the runtime was already gone server-side and the -failure was a handler stabilization artifact. **Run 1's attempt-3 technique -reproduced exactly and is confirmed as the right procedure** — it took the -residual from 19 to 4. - -**Final residual — 4 resources, all zero-cost, identical in shape to run 1 and P1:** - -``` -CREATE_COMPLETE AWS::EC2::VPC vpc-07dad8897791f477b -DELETE_FAILED AWS::EC2::Subnet subnet-09a1ff1f7568d05ef -DELETE_FAILED AWS::EC2::Subnet subnet-068397b48a25bf13f -DELETE_FAILED AWS::EC2::SecurityGroup sg-05e3f48665c47b358 -``` - -Pinned by two fresh `agentic_ai` ENIs (`eni-0942f1b0b3b6553ea`, -`eni-07876963c2818da63`, both `in-use`). **#702 / P1 F12 reproducing for the third -consecutive run.** Nothing was force-deleted past CloudFormation. - -**Billing confirmed stopped:** - -| Check | Result | -|---|---| -| NAT gateways in either orphaned ABCA VPC | **none** | -| VPC endpoints | **none** | -| ABCA S3 buckets | **none** | -| Unattached (billable) EIPs | **none** | -| MicroVM images / non-`TERMINATED` VMs | **none** | -| AgentCore runtimes / memories | **none** | -| `/aws/lambda-microvms/*` log groups | **none** | -| ABCA DynamoDB tables / GitHub-token secret | **none** | - -**Deliberately retained:** `CDKToolkit` (unchanged — still the shared-account -caveat P1 raised); bootstrap S3/ECR assets; service-vended log groups created -outside CloudFormation; the two scratch-repo PRs (2.11); and **two** orphaned -zero-cost VPC sets — run 2's four resources above (still inside the -`DELETE_FAILED` stack) plus **run 1's**, which `--retain-resources` moved outside -CloudFormation entirely (`vpc-072fddf653ccdcfc4`, `subnet-0e6a6a0ed18100c8a`, -`subnet-02d91450e51cf72a0`, `sg-00997a58c1f4c5775`). - -A best-effort hand cleanup of run 1's set was attempted and **refused** — -`DependencyViolation` on both subnets, the SG and the VPC, because -`eni-04911e11e08d670f9` and `eni-04a1c5a27966ee08b` were **still `in-use` -3 h 30 min after run 1's teardown**. Nothing was forced. Retry for both sets: - -```bash -aws cloudformation delete-stack --stack-name backgroundagent-dev # run 2's set -# run 1's set is no longer CFN-managed: -aws ec2 delete-subnet --subnet-id subnet-0e6a6a0ed18100c8a -aws ec2 delete-subnet --subnet-id subnet-02d91450e51cf72a0 -aws ec2 delete-security-group --group-id sg-00997a58c1f4c5775 -aws ec2 delete-vpc --vpc-id vpc-072fddf653ccdcfc4 -``` - -### 2.13 Findings summary - -Live run 2026-08-07 02:58Z → 04:55Z, account ``, `us-east-1`, branch -`feat/645-lambda-microvm-p2` @ `b927d1d6`. Evidence: -`/tmp/abca-645-p2r2-20260806`. - -**Verdict on the primary objective: THE SMOKE PASSED. -`https://github.com/dreamorosi/batch-sync-triage/pull/6`.** Clone → change → -commit → push → PR, `COMPLETED` in 153 s for $0.28, with progress events and a -live heartbeat. That retires run 1's headline. **It required one live IAM -workaround, and pinning down why is the run's most valuable output.** - -#### Fixes CONVERTED to "re-exercised live" - -1. **P2-F1 — DISCHARGED.** Substrate deployed **first try**, `EXIT=0`, - 14 min 34 s, **zero workarounds**, both `AWS::Lambda::NetworkConnector` - resources `CREATE_COMPLETE`, and **no `*FAILED` event anywhere** in 485 - resources. Run 1 needed five attempts and a cloud-assembly patch. -2. **P2-F2 — DISCHARGED.** Change-set **early validation passed** with `ARM_64` - and four `ENABLED` hooks; the resource reached `CREATE_IN_PROGRESS`. Run 1's - five-value refutation is answered on the CloudFormation surface. (The path is - still blocked downstream — P2r2-F9 — but *not* on the enums.) -3. **P2-F4 — DISCHARGED with content.** `server_debug/`, - `metrics/` **and** `trajectory/` all exist; the - `METRICS_REPORT` run 1 lost entirely now lands. -4. **P2-F5 — DISCHARGED, and now quantified.** Cold `claude` exec measured at - **17.1–37.8 s** (so `timeout=10` was 2–4× short and could never have passed); - **0.1 s once warm**; `/ready` 200, no 503; `/validate` still - `platform_config_keys=13, warnings=0`; whole build 4 min 30 s, i.e. the - warm-up cost ~22 s. **New empirical detail:** the service calls `/ready` - **three times per build, the first two concurrently**, so two cold `claude` - execs contend — worst single call ≈ 44 s. The 120 s required budget and the - `readyTimeoutInSeconds` 60 → 300 move are both correctly sized; 60 s would - have left ~16 s of slack. -5. **Dual-signal liveness — DISCHARGED** (run 1: BLOCKED). `agent_heartbeat_at` - advanced `04:24:47 → 04:25:32 → 04:26:17` — exact 45 s cadence — across a - 181 s `RUNNING` window, freshness ≤ 34 s. - -#### Items CONTRADICTING design assumptions — `feeds-back-to-design: YES` - -**P2r2-F10 (BLOCKING). P2-F3 is NOT fixed. The orchestrator's *identity-side* -`iam:PassedToService: lambda.amazonaws.com` condition — the one run 1 explicitly -exonerated and `b927d1d` deliberately kept — is a second, independent blocker.** - -Isolated by a clean two-arm experiment, same exact-ARN resource as source, same -5-minute settle, one variable: - -| Orchestrator grant | Result | -|---|---| -| exact ARN **+ `iam:PassedToService: lambda.amazonaws.com`** (source as written) | **DENIED** — submissions 1 *and* 4 | -| exact ARN, **no condition** | **`RUNNING` in 9 s** — submission 5 | - -``` -Session start failed: Error: MicroVM RunMicrovm failed: AccessDeniedException: User: arn:aws:sts:::assumed-role/backgroundagent-dev-TaskOrchestratorOrchestratorFnS-a7sP6rFzoIkU/backgroundagent-dev-TaskOrchestratorOrchestratorFn-p41lJqwFmxNG is not authorized to perform: iam:PassRole on resource: arn:aws:iam:::role/backgroundagent-dev-LambdaMicrovmComputeExecutionRo-pZQWXvKsITBa because no identity-based policy allows the iam:PassRole action -``` - -The trust half of run 1's fix was **necessary but not sufficient**. Run 1's -exoneration was a **false negative with an identifiable cause**: its temporary -unconditioned `iam:PassRole` was attached at §4.1 step 1 and, per its own §8.3, -**was still attached through submissions 4 and 5** — the ones that reached -`RUNNING`. So run 1 never tested the conditioned grant against a *working* trust, -and attributed the whole effect to the trust change. - -Two source statements must therefore change, and the comment in -`task-orchestrator.ts` asserting the condition "was EXONERATED live … so it -stays" must be reversed: - -- `cdk/src/constructs/task-orchestrator.ts`, sid `MicrovmPassExecutionRole` -- `cdk/bootstrap/policies/infrastructure.json`, sid `IAMPassRole` (P2r2-F9) - -**P2r2-F9 (BLOCKING, same root cause). The CDK-managed image path is dead one -step later than run 1 thought: CloudFormation cannot pass the build role.** -Verbatim `CREATE_FAILED` in 2.5. **This is not a stale bootstrap** — the live -policy is byte-identical to this branch's -`cdk/bootstrap/policies/infrastructure.json`. `simulate-principal-policy` returns -`allowed` for `iam:PassedToService=lambda.amazonaws.com` and `implicitDeny` -without it, so the resource pattern is right and the condition is the variable; -and the out-of-band `--create-image` call **succeeded passing the same build -role**, proving the trust is fine and the denial is caller-side. - -**P2r2-F9 and P2r2-F10 are ONE root cause with TWO symptoms** — precisely the -shape of run 1's P2-F1/F3, one layer in: **the Lambda MicroVMs service does not -present `iam:PassedToService: lambda.amazonaws.com` on either `PassRole` path** -(CloudFormation → build role at `CreateMicrovmImage`; orchestrator → execution -role at `RunMicrovm`). Consequence for the docs: the "CDK-managed (recommended)" -bootstrap path in `package-microvm-artifact.sh` is **still** non-functional and -`--create-image` is **still** the only working path — for a new reason. - -**P2r2-F11. `agent_heartbeat_at` is written and consumed correctly but is -invisible through the API.** `toTaskDetail` -(`cdk/src/handlers/shared/types.ts:786`) does not map it, though `TaskRecord` -declares it (line 95) and `orchestrator.ts` reads it for liveness; `cli/src/types.ts` -has no such field at all. So `bgagent status` / `watch` report `None` even when -DynamoDB holds a 6-second-old value. Cheap to fix, and worth fixing because **it -already caused a wrong conclusion**: run 1 recorded "heartbeats NOT observed" and -attributed it wholly to P2-F5. - -**P2r2-F12. Run 1's §8.5 stack-delete retry advice does not work on a -same-session horizon, and the ENI leak is worse than #702 records.** Run 1's -documented `delete-stack` retry failed **identically after 17 min 17 s**, with the -`agentic_ai` ENIs still `in-use` 1 h 32 min after run 1's teardown and *no* -AgentCore runtimes or memories in existence. At the end of *this* run those same -two ENIs were **still `in-use` 3 h 30 min on**, and a hand `delete-subnet` / -`delete-security-group` / `delete-vpc` sweep was refused with -`DependencyViolation` on all four. So the leak is not "slow" — it is **unbounded -relative to a developer's session**, and it now compounds: each run strands -another VPC set. The working escape hatch is `--retain-resources` (33 s), which -must list the `CREATE_COMPLETE` VPC alongside the three `DELETE_FAILED` children: - -```bash -aws cloudformation delete-stack --stack-name backgroundagent-dev \ - --retain-resources AgentVpcA6796801 AgentVpcPrivateSubnet1Subnet8051BB57 \ - AgentVpcPrivateSubnet2SubnetC66971D0 AgentVpcRuntimeSG96507CD0 -``` - -#702 should carry both the unbounded-hold evidence and this escape hatch, rather -than have each run rediscover them. Run 1's attempt-3 technique (verify -`list-agent-runtimes` is empty, *then* retry) is separately **confirmed correct** -— it took this run's residual from 19 resources to 4. - -**P2r2-F13. `MICROVM_IMAGE_VERSION` is legitimately optional** — five -`MICROVM_*` vars this run vs run 1's six, because the optional -`microvm_image_version` context was not supplied. Correct by design -(`task-orchestrator.ts:467`), but it qualifies the "all-or-nothing `MICROVM_*`" -claim, which should name the version as the one member that may be absent. - -**P2r2-F14. `turns` is reported inconsistently in one stream.** `watch` prints -`agent_execution_complete: turns=22`; the persisted record and `METRICS_REPORT` -both say `12`. - -**P2-F7 reconfirmed, worse.** 985,886 B / 486 resources = **98.6 %** of the 1 MB -limit, ~14 KB of headroom (run 1: ~16 KB). - -**CloudTrail blind spot.** No `lambda-microvms` management events at all, so -`PassRole` context values cannot be read from logs — every determination in -2.5/2.8 had to be by elimination. Worth knowing before the next person tries. - -**Still unfixed from earlier runs:** P1 F11 (AZ constraint — worked around again -via the gitignored AZ cache, restored at teardown); the finch→ECR push failure; -P1 F8 (`TERMINATED` is terminal, `ResourceNotFoundException` never arrives). - -#### Recommended follow-up - -1. **P2r2-F10 + P2r2-F9 together** — drop or correct `iam:PassedToService` on - both the orchestrator statement and the bootstrap `IAMPassRole`. Nothing runs - without the first; the recommended image path stays dead without the second. - Determining the value the service *does* present needs either AWS - confirmation or a bounded candidate sweep — `microvms.lambda.amazonaws.com`, - `lambda-microvms.amazonaws.com` and `microvms.amazonaws.com` are all - `implicitDeny` against the current policy, so any of them would work as the - allow-list entry if it is the right one. Reverse the "EXONERATED … so it - stays" comment while you are there. -2. **P2r2-F11** — one line in `toTaskDetail` plus the CLI type; it is what makes - the liveness signal observable to the people who need it. -3. **P2r2-F12** — put the `--retain-resources` escape hatch in #702. -4. **P2-F7** — `suppressTemplateIndentation` or a stack split; ~14 KB left. -5. **A third Stage D is NOT needed for P2-F1/F2/F4/F5 or dual-signal liveness** — - all five are now live-verified. The next run's scope is P2r2-F9/F10 plus the - still-deferred empirical items (suspend TTL > 1 h, SUSPENDED-vs-quota). - -### 2.14 Elapsed and cost - -**Elapsed:** 02:58Z → 04:55Z ≈ **1 h 57 min**, of which ~18 min was the failed -delete retry, ~15 min the substrate deploy, ~6 min the refuted CDK-managed image -attempt, ~10 min image create+build, ~5 min the wired deploy, ~30 min the five -submissions and the isolation experiment (mostly IAM settle waits), ~26 min the -two teardown delete attempts, and the remainder evidence capture. - -**Approximate cost: well under US$5**, dominated as always by NAT/VPC-endpoint -hours. Unlike run 1 this run actually spent Bedrock tokens: **$0.478 total across -two completed tasks** ($0.279 + $0.199). Two MicroVMs ran ~2.5 min each; two -image builds ~4.5 min each; one NAT gateway and 7×2 interface endpoints for -~1.6 h across a single stack lifetime. - -The 8 h `maximumDurationInSeconds` was never approached; every VM was explicitly -terminated by the orchestrator's finalization. diff --git a/docs/verification/645-p3-lifecycle-diagnostics.md b/docs/verification/645-p3-lifecycle-diagnostics.md new file mode 100644 index 000000000..9288934c7 --- /dev/null +++ b/docs/verification/645-p3-lifecycle-diagnostics.md @@ -0,0 +1,91 @@ +# MicroVM lifecycle diagnostics + +An approval being saved, AWS accepting Resume, and the guest resuming work are +three different events. Diagnose each independently; an accepted API response +is not a completed wake. + +## Correlate coordinator and guest logs + +Search coordinator/approval Lambda logs and the MicroVM image log group using +`task_id` and `microvm_id`. `request_id` identifies the approval gate; +`aws_request_id` identifies an AWS call; `hook_id` identifies one guest HTTP hook. + +| Coordinator record | Meaning | +|---|---| +| `MicroVM observed after approval decision` | State and saved lifecycle intent observed after the decision. | +| `MicroVM wake request started/requested after approval decision` | Wake dispatch and acknowledgment, including receipt and elapsed time when available. | +| `MicroVM lifecycle request started/acknowledged/failed` | Suspend/Resume operation and result. | +| `MicroVM supervisor observation changed` | Task/worker/gate state, recovery timing, failure counters and outcome. | +| `MicroVM reached a terminal state with a substrate reason` | Service reason and worker/image identity. | +| `Lambda MicroVM termination requested` | Cleanup request and receipt. | + +Unchanged supervisor observations are deduplicated across durable replay. +Recovery timeouts retain their original start; investigating or retrying must +not reset them. + +For guest logs, select the deployed image's `/aws/lambda-microvms/` +log group and run a CloudWatch Logs Insights query such as: + +```text +fields @timestamp, event, action, stage, callback_stage, code, http_status, + hook_id, request_id, pid, phase, elapsed_ms, late, + error_type, aws_error_code, aws_request_id +| filter microvm_id = "REPLACE_WITH_WORKER_ID" +| sort @timestamp asc +| limit 500 +``` + +| Guest event | Meaning | +|---|---| +| `microvm_hook_started` | Handler entry, before body reading. | +| `microvm_hook_stage` | The next potentially blocking operation. | +| `microvm_hook_stage_finished` / `microvm_hook_stage_failed` | Operation completion or safe error metadata. | +| `microvm_hook_finished` | Selected handler status/code; not proof of service receipt. | + +`stage` records the last entered operation, such as credential refresh or approval +identity reconciliation. `late: true` means a callback finished after the handler +had already returned; it cannot turn a timed-out wake into success. The coding +barrier remains responsible for preventing tools after an uncertain wake. + +## Checkpoint failure codes + +The checkpoint diagnostic `code` narrows the failing operation; it does not +prove that saved data is corrupt. + +| Code | Meaning | +|---|---| +| `checkpoint_failed` | No narrower classification; inspect the accompanying stage and message. | +| `checkpoint_invalid_json` | Checkpoint JSON could not be encoded or decoded. | +| `checkpoint_sdk_unverified` | SDK version or required accounting interface is not verified. | +| `checkpoint_sdk_timeout` | SDK transcript acknowledgment or accounting request timed out. | +| `checkpoint_storage_unverified` | A save could not be verified by reading back the exact data; keep the source worker. | +| `checkpoint_storage_unavailable` | Storage configuration is unavailable or the saved version could not be read. | + +Workspace capture and restore can also report specific codes such as +`disk_pressure` or `git_timeout`. Preserve the reported code and stage when +escalating; do not replace them with a generic checkpoint label. + +## Diagnose a failed wake + +1. Locate the saved decision/intent and actual Resume acknowledgment or failure. + Retain the original generation, deadlines and API request ID. +2. Find guest hook entry. If absent, account for log delivery and retention before + inferring anything about the listener or process. +3. If entered, inspect the final stage/code. Credential-refresh `AccessDenied` + points to renewal permissions; identity-read failures point to task/gate + reconciliation. A timeout identifies the outstanding operation. +4. Check subsequent guest progress, coordinator outcome, worker termination and + capacity release. A saved approval or successful cleanup does not establish + that the approved tool ran. +5. Retain UTC timestamps, task/worker identifiers, exact image/coordinator + versions, service state reason, AWS receipts and relevant sanitized logs. + +`MICROVM_RESUME_HOOK_FAILED` identifies recognized service resume-hook failures. +Preserve the raw service reason for diagnosis. The wording “connection was +refused” alone does not prove a closed listener: historical guest observations +support a stale pooled-connection race, while service-side dispatch traces remain +unavailable. Lifecycle responses explicitly close connections before freeze. + +Diagnostics omit hook bodies, tool arguments, approval contents, credentials, +raw exception messages and SDK response bodies. Do not attach signed payload +URLs or conversation/workspace checkpoints when escalating an incident. diff --git a/docs/verification/645-p3-nested-stack.md b/docs/verification/645-p3-nested-stack.md new file mode 100644 index 000000000..8470fa3bf --- /dev/null +++ b/docs/verification/645-p3-nested-stack.md @@ -0,0 +1,74 @@ +# Nested MicroVM infrastructure + +This is a CloudFormation infrastructure split, not a virtual machine running +inside another virtual machine. Fresh nested deployment and a deployment-specific +migration were verified; reusable migration commands remain unfinished. + +## Resource ownership + +`AgentStack` creates the `Microvm` child (`LambdaMicrovmStack`). It owns the +managed image when configured, build/runtime network connectors and security +groups, artifact/payload buckets, logs, and build/operator roles. + +The execution role remains at `LambdaMicrovmCompute/ExecutionRole` in the parent. +This preserves its logical ID and avoids a dependency cycle through SessionRole +trust. Parent `Microvm*` outputs retain the names used by packaging and consumers. +Names derive from the concrete parent deployment name, not the child stack token. + +## Configuration + +| Setting | Behavior | +|---|---| +| `microvm_nested_stack` | Required when MicroVM is enabled: use `true` for new or already-nested deployments; retain `false` for existing flat deployments until completing the reviewed migration. Omission fails synthesis. | +| `microvm_resource_name_prefix` | Supplies distinct names for overlapping nested resources; preserve it after migration. It does not retain old resources or permissions by itself. | +| `microvm_managed_image_version` | Pins new tasks to an explicitly verified image version. Without a pin, selection follows the latest active version. | +| `microvm_approval_suspend_enabled` | Defaults to `false`; enable only after testing the deployed image/coordinator. Disabling new sleep preserves wake and cleanup. | + +Nested deployment requires bootstrap bundle **1.9.0** or later. See +[deployment roles](../design/DEPLOYMENT_ROLES.md) and the +[artifact packaging instructions](../../cdk/scripts/README.md). +Build a new image before switching its runtime pin; building alone must not +silently change the image used by a pinned coordinator. Retain a compatible +published coordinator and exact image version for rollback. + +## Existing flat deployments + +Before the first upgrade, save `"microvm_nested_stack": false` in the deployment's +CDK context or pass `--context microvm_nested_stack=false` on every deploy. +Omitting the setting fails synthesis. This temporary requirement protects upgrades +that previously omitted the setting; it does not inspect live resources or perform +a migration. New and already-nested installations should explicitly set `true`. +Do not reuse older synthesized assemblies: regenerate them with this version. + +Do not deploy the nested template directly over a flat deployment. CloudFormation +sees removed parent resources and newly created child resources; names can collide +and bucket auto-delete handlers can erase artifacts or pending payloads. + +The tested provider rejected native `AWS::Lambda::MicrovmImage` stack refactoring +with an unsupported tag-schema error even though the preview succeeded. Do not +assume resource import/refactoring is supported from a successful preview. + +The required overlap migration has these stages. This checklist is not yet a +runnable migration command: + +1. Preserve the exact deployed template/cloud assembly, image and published + coordinator versions. Capture current resource identities and configuration. + Keep `microvm_nested_stack=false` on ordinary updates before migration. +2. Create the child with distinct names while retaining the old resources and + their permissions. Build and verify the new image before switching consumers. + Reject intermediate templates that exceed CloudFormation's 500-resource limit. +3. Deploy compatible approval handlers/coordinator before producers of retained + requests. Preserve permissions for old workers, including the old payload + access that P3's new bootstrap deny would otherwise block. +4. Switch consumers with an explicit image pin. Verify normal tasks, retained + approvals, sleep/wake and rollback while both resource sets are available. +5. Retire old resources only after old workers, Durable executions and uncertain + starts are accounted for and old tasks/leases are drained. An idle inventory + snapshot alone is not an admission fence. Preserve checkpoint data and any + resources still needed by published coordinator versions. + +Setting a concurrency counter to zero is not a reliable pause for asynchronous +producers. A rollback must restore compatible code, image selection and IAM +without deleting pending requests or saved work. The reusable tool must enforce +these prerequisites; migration template and permission helpers alone do not +constitute that tool. Track remaining acceptance in [verification status](./README.md#open-pr-checks). diff --git a/docs/verification/645-payload-bootstrap.md b/docs/verification/645-payload-bootstrap.md new file mode 100644 index 000000000..57add082a --- /dev/null +++ b/docs/verification/645-payload-bootstrap.md @@ -0,0 +1,79 @@ +# Trusted task delivery for ECS and MicroVM + +The coordinator delivers each task through an IAM-authenticated deployment +manifest and a signed URL for one task object. This is application startup +configuration, distinct from the CDK infrastructure bootstrap. + +## Contract and permissions + +Version 2 is defined in `contracts/constants.json` under `payload_bootstrap`. +The [ADR wire contract](../decisions/ADR-021-lambda-microvms-compute-backend.md#3-packaging-same-agent-image-source-new-build-path) +and [security design](../design/SECURITY.md) describe its trust boundary. + +| Object | Purpose | +|---|---| +| `bootstrap/.json` | Non-secret backend/platform configuration; coordinator writes, worker reads with its ambient role. | +| `/payload.json` | Task instructions/configuration; downloaded only through the signed capability. | +| `/launch.json` | Private coordinator replay record containing the exact saved reference. | + +The worker's explicit S3 deny outside its own `bootstrap/*` prevents a foreign +public bucket from supplying fake configuration. A content hash alone does not +establish authorship. Workers cannot list their payload bucket or use ambient +credentials to read task objects. The coordinator needs `ListBucket` to distinguish +missing launch records from access denial. + +Downloaded task identity and configuration must match the authenticated manifest. +Old unsigned envelopes are rejected. All payload sizes use S3; the serialized +MicroVM reference must fit 4,096 bytes. Manifest/payload limits are 16 KiB/8 MiB. +MicroVM receives the reference through `runHookPayload`; ECS uses +`AGENT_PAYLOAD_REF` and removes it before repository subprocesses start. + +A signed URL is a bearer credential. Never log it or store it in a worker-readable +task row. HTTPS downloads reject redirects, environment proxies, alternate hosts +and invalid task paths. The runtime needs regional S3 HTTPS and DNS connectivity. + +## Retry and cleanup + +- Conditional writes keep task instructions immutable. Read-back reconciles a + write that succeeded but lost its response. +- Retries reuse the exact saved URL and Run client token. Re-signing would change + the request under the same token. An expired record fails explicitly. +- Signed lifetime is at most 900 seconds and can end earlier when temporary + credentials expire. Initial creation refuses a known lifetime under 300 seconds. +- A timeout does not prove no worker started. Reconcile the existing handle or + uncertain start before considering replacement; the application replay window + is not a service idempotency-retention guarantee. +- Finalization deletes payload and launch objects on a best-effort basis; bucket + lifecycle deletion is an asynchronous backstop. The coordinator refreshes + identical manifest bytes on preparation so old deployment settings remain usable. +- Resume continues saved task state with refreshed credentials; it does not + download the task again using an expired launch URL. + +## Coordinated upgrades + +Coordinator code, worker images/task definitions and IAM must implement the same +contract. Keep the previous deployable artifacts and inspect the change set. +For an upgrade that cannot support overlapping versions, pause actual admission +sources and drain tasks, approvals and uncertain starts before switching these +components together. A concurrency counter alone does not pause webhook/queue +admission. Do not delete and recreate storage to perform an upgrade. + +For overlapping flat-to-nested migration, preserve old-worker permissions until +those workers drain; follow the [migration prerequisites](./645-p3-nested-stack.md). +Rollback also requires compatible code, images and IAM. Never reuse a task ID +with conflicting or expired stored launch instructions. + +## Verification + +Local producer/consumer tests cover conflicts, lost committed replies, malformed +or oversized input, wrong task/configuration, URL redaction and cleanup. These +cannot prove effective AWS permissions. In a representative deployment, verify: + +- Own manifest and signed payload succeed; foreign manifests, ambient payload + reads, bucket listing and worker mutations are denied. +- Modified signatures/paths, expired credentials/URLs and revoked objects fail + without starting the pipeline or exposing a capability in logs. +- Competing preparation and lost S3/Run replies preserve one exact launch request. +- Finalization removes both task objects and releases only confirmed capacity. + +See [recorded acceptance and remaining checks](./README.md). diff --git a/docs/verification/README.md b/docs/verification/README.md new file mode 100644 index 000000000..0492a2411 --- /dev/null +++ b/docs/verification/README.md @@ -0,0 +1,101 @@ +# Lambda MicroVM verification + +For maintainers reviewing/testing the MicroVM backend and operators deploying, +migrating or diagnosing it. [ADR-021](../decisions/ADR-021-lambda-microvms-compute-backend.md) +explains the design; the [user guide](../guides/USER_GUIDE.md#approval-gates-cedar-hitl) +explains approval and sleep options for people submitting tasks. +Detailed deployment transcripts, temporary worker identifiers and investigation +diaries are archived outside the repository. These documents are not test runners. + +- [Task payload delivery](./645-payload-bootstrap.md): authorization, retries and coordinated upgrades. +- [Nested infrastructure](./645-p3-nested-stack.md): configuration and migration prerequisites. +- [Lifecycle diagnostics](./645-p3-lifecycle-diagnostics.md): locating and interpreting failed wakes. +- [Continuation design](../design/ORCHESTRATOR.md#retained-microvm-approvals): checkpoint ownership, retirement and replacement. + +## Recorded acceptance + +AWS checks on September 14–18, 2026 exercised the following behavior in the tested +deployments. They do not certify a different image, account, Region or upgrade. + +| Area | Observed result | +|---|---| +| Task delivery | Signed payload downloads, invalid input rejection, expiry/revocation, immutable preparation and lost-reply recovery passed. | +| Approval lifecycle | Approve, deny, explicit expiry and cancellation passed; unanswered requests stayed available without a default deadline. | +| Sleep/wake | Repeated wakes, the 600-second default and the live sleep-off switch passed with compatible image/coordinator versions. | +| Credential renewal | A wait exceeding one hour was followed by successful AWS access with renewed task credentials. | +| Continuation | Conversation and Git/workspace recovery, replacement admission, usage limits and capacity release passed. | +| External integrations | Repository work and remote MCP access passed across sleep; an actual Linear submission exercised MicroVM compute with AgentCore Identity vault. | +| Linear decisions | Native threaded `approve` and `deny` replies passed on MicroVM and AgentCore. Both MicroVMs were suspended before the replies; exact decisions, tool results, thread acknowledgements and eventual capacity release were verified. | +| Infrastructure | Fresh nested deployment and a deployment-specific overlapping migration passed, including compatible rollback and old-resource cleanup. | +| Other backends | ECS and AgentCore approval/cancellation and scoped-access checks passed in their tested deployments. | + +The wake correction sends `Connection: close` in lifecycle responses before +freeze. Local transport controls reproduced failure on an old connection; +long-sleep controls and the corrected live flows passed. Service-side traces +for the historical failures remain unavailable, so their exact transport error +is not established for every worker. + +The full local build after the resource-budget fix passed: 5,560 CDK tests, +2,170 agent tests and 1,005 CLI tests, plus compile, lint, contracts, docs and +synthesis. The widest parent stacks use 489 resources for ECS and 488 for +MicroVM, including synth metadata, within the unchanged 490-resource budget. +Concurrency maintenance now has its own nested stack. Upgrading recreates that +stateless repair function and schedule; task and concurrency tables stay in the +parent stack. + +## Open PR checks + +- Finish reusable flat-to-nested migration commands and independently test an + upgrade from current `main` on the same deployment. The earlier bespoke + migration is not a substitute for that acceptance. +- Deploy the latest review fixes and verify CLI/API approve and deny on a + `PARKED` task, including immediate replacement admission. Earlier Linear + acceptance does not exercise the decision API functions' configuration. + +## Reproduce local checks + +From the repository root, with dependencies installed: + +```bash +mise run build +MISE_EXPERIMENTAL=1 mise //cdk:testf -- 'microvm|migration|payload-bootstrap' +``` + +For focused worker checks: + +```bash +cd agent +uv run pytest tests/test_microvm_*.py tests/test_continuation_*.py \ + tests/test_approval_retention.py tests/test_payload_bootstrap.py --no-cov +ABCA_TEST_SDK_CONTINUATION=1 uv run pytest tests/test_continuation_sdk_probe.py --no-cov +``` + +The last command opts into the pinned real SDK/CLI probe with a deterministic +loopback model; it does not launch a cloud worker. Optional DynamoDB Local tests +require their documented local service. Mocks do not establish effective AWS IAM. +The standalone cloud acceptance harness and raw receipts remain outside this PR. +The repository's narrower [launch, payload and replay probes](../../cdk/test/live/README.md) +have separate inspection and execution commands. + +## Live acceptance for an installation + +Record source commit, image ARN/version, coordinator version, configuration and +UTC test window. Use an owned test repository and identity. Enable suspension +only after verifying that image and coordinator together. + +| Exercise | Required observation | +|---|---| +| Normal task | Repository tools run; terminal status, payload cleanup and capacity release agree. | +| Approval after sleep | The saved decision reaches the exact pending tool; verify guest recovery and tool output, not only the Resume API receipt. | +| Linear replies | Submit real issues on each backend. Reply `approve` or `deny` to the exact approval comment as its linked owner; verify the saved decision source, same-thread acknowledgement and allowed/blocked tool result. For MicroVM, observe suspension before replying. | +| Deny, expiry, cancellation | No denied/cancelled tool runs; expiry uses the original deadline; compute and capacity are cleaned up. | +| Default and disabled sleep | Omitted override uses 600 seconds; task-level off and deployment off prevent new suspension while wake/cleanup remain available. | +| Credential expiry | Sleep past the original credential lifetime, then perform actual task-scoped AWS operations. | +| Retirement/replacement | Confirm old-worker shutdown before capacity release; one replacement restores files/conversation and preserves approval identity and usage. | +| CLI/API replacement | After the task reaches `PARKED`, approve or deny through the CLI/API. Verify replacement admission starts from that decision without waiting for the scheduled sweep, and the exact pending tool follows the decision. | +| Upgrade/rollback | Preserve unrelated resource identities, old in-flight work and recoverable checkpoints; test compatible code/image/policy rollback. | + +Include effective-role checks for own-task access and denial of cross-task data, +foreign bootstrap manifests and ambient payload reads. Keep credentials, signed +URLs, prompts and checkpoints out of diagnostic attachments. Clean up test +workers, executions, task data and owned infrastructure after the run. diff --git a/scripts/check-constants-sync.ts b/scripts/check-constants-sync.ts index 536427486..5a6a5bdf0 100644 --- a/scripts/check-constants-sync.ts +++ b/scripts/check-constants-sync.ts @@ -49,9 +49,16 @@ const POLICY_PY = path.join(REPO_ROOT, 'agent/src/policy.py'); const JIRA_REACTIONS_PY = path.join(REPO_ROOT, 'agent/src/jira_reactions.py'); const SERVER_PY = path.join(REPO_ROOT, 'agent/src/server.py'); const CONFIG_PY = path.join(REPO_ROOT, 'agent/src/config.py'); -const PYTHON_CONSUMERS = [POLICY_PY, JIRA_REACTIONS_PY, SERVER_PY, CONFIG_PY]; +const PAYLOAD_BOOTSTRAP_PY = path.join(REPO_ROOT, 'agent/src/payload_bootstrap.py'); +const MICROVM_HTTP_PY = path.join(REPO_ROOT, 'agent/src/microvm_http.py'); +const PAYLOAD_BOOTSTRAP_TS = path.join(REPO_ROOT, 'cdk/src/handlers/shared/payload-bootstrap.ts'); +const PYTHON_CONSUMERS = [ + POLICY_PY, JIRA_REACTIONS_PY, SERVER_PY, CONFIG_PY, PAYLOAD_BOOTSTRAP_PY, MICROVM_HTTP_PY, +]; const MICROVM_COMPUTE_TS = path.join(REPO_ROOT, 'cdk/src/constructs/lambda-microvm-compute.ts'); -const TS_CONSUMERS = [MICROVM_COMPUTE_TS]; +const MICROVM_IMAGE_CAPABILITY_TS = path.join(REPO_ROOT, 'cdk/src/handlers/shared/microvm-image-capability.ts'); +const MICROVM_STRATEGY_TS = path.join(REPO_ROOT, 'cdk/src/handlers/shared/strategies/lambda-microvm-strategy.ts'); +const TS_CONSUMERS = [MICROVM_COMPUTE_TS, MICROVM_IMAGE_CAPABILITY_TS, MICROVM_STRATEGY_TS]; /** Env var names must be UPPER_SNAKE — they are installed into a process env. */ const ENV_NAME_PATTERN = /^[A-Z][A-Z0-9_]*$/; @@ -118,6 +125,14 @@ const OWNED_PYTHON_PATTERNS: ReadonlyArray<{ name: string; regex: RegExp }> = [ name: '_READY_WARMUP_REQUIRED_TIMEOUT_SECONDS', regex: /^\s*_READY_WARMUP_REQUIRED_TIMEOUT_SECONDS\s*(?::\s*(?:int|float))?\s*=\s*-?\d+\b/m, }, + { + name: 'LIFECYCLE_HANDLER_BUDGET_S', + regex: /^\s*LIFECYCLE_HANDLER_BUDGET_S\s*(?::\s*(?:int|float))?\s*=\s*-?\d+\b/m, + }, + { + name: 'LIFECYCLE_HOOK_TIMEOUT_S', + regex: /^\s*LIFECYCLE_HOOK_TIMEOUT_S\s*(?::\s*(?:int|float))?\s*=\s*-?\d+\b/m, + }, ]; /** @@ -131,6 +146,26 @@ const OWNED_PYTHON_PATTERNS: ReadonlyArray<{ name: string; regex: RegExp }> = [ * the same way the Python ones are. */ const OWNED_TS_PATTERNS: ReadonlyArray<{ name: string; regex: RegExp }> = [ + { + name: 'MICROVM_MAX_DURATION_SECONDS', + regex: /^\s*(?:export\s+)?const\s+MICROVM_MAX_DURATION_SECONDS\s*(?::\s*number)?\s*=\s*-?\d[\d_]*\b/m, + }, + { + name: 'LIFECYCLE_HOOK_TIMEOUT_SECONDS', + regex: /^\s*(?:export\s+)?const\s+LIFECYCLE_HOOK_TIMEOUT_SECONDS\s*(?::\s*number)?\s*=\s*-?\d+\b/m, + }, + { + name: 'AGENT_HOOK_PORT', + regex: /^\s*(?:export\s+)?const\s+AGENT_HOOK_PORT\s*(?::\s*number)?\s*=\s*-?\d+\b/m, + }, + { + name: 'MICROVM_LIFECYCLE_PROTOCOL', + regex: /^\s*(?:export\s+)?const\s+MICROVM_LIFECYCLE_PROTOCOL\s*(?::\s*string)?\s*=\s*(?:String\(\s*)?["'\d]/m, + }, + { + name: 'MICROVM_IMAGE_PROTOCOL_ENV', + regex: /^\s*(?:export\s+)?const\s+MICROVM_IMAGE_PROTOCOL_ENV\s*(?::\s*string)?\s*=\s*["']/m, + }, { name: 'READY_HOOK_TIMEOUT_SECONDS', regex: /^\s*(?:export\s+)?const\s+READY_HOOK_TIMEOUT_SECONDS\s*(?::\s*number)?\s*=\s*-?\d+\b/m, @@ -184,6 +219,34 @@ function main(): number { ready_hook_timeout_seconds: number; warmup_total_budget_seconds: number; warmup_required_timeout_seconds: number; + lifecycle_hook_timeout_seconds: number; + lifecycle_handler_budget_seconds: number; + }; + microvm_lifecycle?: { + protocol_version: number; + image_protocol_env: string; + hook_port: number; + maximum_duration_seconds: number; + }; + microvm_continuation?: { + version: number; + lease_key_prefix: string; + object_key_prefix: string; + max_manifest_bytes: number; + max_workspace_bytes: number; + max_conversation_bytes: number; + park_after_seconds: number; + retirement_margin_seconds: number; + verified_sdk_version: string; + }; + payload_bootstrap?: { + version: number; + manifest_prefix: string; + launch_filename: string; + max_manifest_bytes: number; + max_payload_bytes: number; + url_ttl_seconds: number; + minimum_url_lifetime_seconds: number; }; }; try { @@ -223,7 +286,7 @@ function main(): number { if (agc.default < agc.min) invariantErrors.push('approval_gate_cap.default must be >= min'); if (agc.max < agc.default) invariantErrors.push('approval_gate_cap.max must be >= default'); if (ats.min <= 0) invariantErrors.push('approval_timeout_s.min must be > 0'); - if (ats.default < ats.min) invariantErrors.push('approval_timeout_s.default must be >= min'); + if (ats.default !== 0 && ats.default < ats.min) invariantErrors.push('approval_timeout_s.default must be 0 or >= min'); if (ats.max < ats.default) invariantErrors.push('approval_timeout_s.max must be >= default'); if (jiraAppActor.min_secret_length < 32) { invariantErrors.push('jira_app_actor.min_secret_length must be >= 32'); @@ -287,12 +350,11 @@ function main(): number { invariantErrors.push('microvm_platform_config.required contains a duplicate'); } - // ARN pinning (ADR-021 P2, review B5). `arn_keys` names the values the agent - // pins to its own partition/account before installing them into the env that - // resolves credentials and fetches secrets; `account_anchor_key` names the ARN - // that supplies the expected partition/account. Both are validated here as well - // as at agent import time, because a malformed entry would silently WIDEN what - // the agent accepts from a network payload. + // ARN consistency (ADR-021 P2, review B5). `arn_keys` names the values checked + // against the partition/account supplied by `account_anchor_key`. That anchor + // comes from the same payload; agreement does not establish deployment identity. + // Validate both here and at agent import time so a new ARN field cannot + // silently skip the existing consistency check. if (!Array.isArray(mpc.arn_keys) || mpc.arn_keys.length === 0) { invariantErrors.push('microvm_platform_config.arn_keys must be a non-empty array'); } else { @@ -306,6 +368,14 @@ function main(): number { if (new Set(mpc.arn_keys).size !== mpc.arn_keys.length) { invariantErrors.push('microvm_platform_config.arn_keys contains a duplicate'); } + for (const [key, envName] of Object.entries(envByKey)) { + if ((key.endsWith('_arn') || (typeof envName === 'string' && envName.endsWith('_ARN'))) + && !mpc.arn_keys.includes(key)) { + invariantErrors.push( + `microvm_platform_config: ARN-shaped key "${key}" is missing from arn_keys`, + ); + } + } if (!mpc.arn_keys.includes(mpc.account_anchor_key)) { invariantErrors.push( 'microvm_platform_config.account_anchor_key must be one of arn_keys', @@ -327,10 +397,37 @@ function main(): number { // for a runtime failure becomes a build failure. Both halves live here precisely // so the relationship is checkable; this is the check. const mhb = json.microvm_hook_budgets; + const lifecycle = json.microvm_lifecycle; + if (!lifecycle || !Number.isSafeInteger(lifecycle.protocol_version) || lifecycle.protocol_version <= 0 + || !Number.isInteger(lifecycle.hook_port) || lifecycle.hook_port < 1 || lifecycle.hook_port > 65535 + || !Number.isInteger(lifecycle.maximum_duration_seconds) + || lifecycle.maximum_duration_seconds <= 0 || lifecycle.maximum_duration_seconds > 28_800 + || typeof lifecycle.image_protocol_env !== 'string' || !/^ABCA_MICROVM_[A-Z0-9_]+$/.test(lifecycle.image_protocol_env)) { + invariantErrors.push('microvm_lifecycle requires a positive protocol version, valid hook port, duration within 1–28800 seconds and ABCA_MICROVM_ marker name'); + } + const continuation = json.microvm_continuation; + const sdkPins = [...fs.readFileSync(path.join(REPO_ROOT, 'agent/pyproject.toml'), 'utf8') + .matchAll(/^\s*"claude-agent-sdk==([^"]+)"/gm)]; + if (sdkPins.length !== 1 || sdkPins[0][1] !== continuation?.verified_sdk_version) { + invariantErrors.push('claude-agent-sdk pin must match microvm_continuation.verified_sdk_version; verify checkpoint compatibility before upgrading both'); + } + if (!continuation || !Number.isSafeInteger(continuation.version) || continuation.version <= 0 + || continuation.lease_key_prefix !== 'worker-lease#' || continuation.object_key_prefix !== 'continuations/' + || !Number.isSafeInteger(continuation.max_manifest_bytes) || continuation.max_manifest_bytes <= 0 + || !Number.isSafeInteger(continuation.max_workspace_bytes) || continuation.max_workspace_bytes <= 0 + || !Number.isSafeInteger(continuation.max_conversation_bytes) || continuation.max_conversation_bytes <= 0 + || !Number.isInteger(continuation.park_after_seconds) || continuation.park_after_seconds <= 0 + || !Number.isInteger(continuation.retirement_margin_seconds) || continuation.retirement_margin_seconds <= 0 + || !lifecycle || continuation.park_after_seconds + continuation.retirement_margin_seconds >= lifecycle.maximum_duration_seconds + || !/^\d+\.\d+\.\d+$/.test(continuation.verified_sdk_version)) { + invariantErrors.push('microvm_continuation requires stable prefixes, a positive version/size, verified SDK version and retirement within the worker lifetime'); + } const BUDGET_FIELDS = [ 'ready_hook_timeout_seconds', 'warmup_total_budget_seconds', 'warmup_required_timeout_seconds', + 'lifecycle_hook_timeout_seconds', + 'lifecycle_handler_budget_seconds', ] as const; if (!mhb || BUDGET_FIELDS.some(field => !Number.isInteger(mhb[field]))) { console.error( @@ -355,6 +452,33 @@ function main(): number { 'best-effort ones something to share)', ); } + if (mhb.lifecycle_handler_budget_seconds >= mhb.lifecycle_hook_timeout_seconds) { + invariantErrors.push( + 'microvm_hook_budgets.lifecycle_handler_budget_seconds must be < ' + + 'lifecycle_hook_timeout_seconds (pause/wake must leave time to answer)', + ); + } + + const bootstrap = json.payload_bootstrap; + const bootstrapNumbers = [ + 'version', 'max_manifest_bytes', 'max_payload_bytes', + 'url_ttl_seconds', 'minimum_url_lifetime_seconds', + ] as const; + if (!bootstrap || bootstrapNumbers.some(key => !Number.isInteger(bootstrap[key]) || bootstrap[key] <= 0) + || typeof bootstrap.manifest_prefix !== 'string' || !/^[a-z][a-z0-9_-]*\/$/.test(bootstrap.manifest_prefix) + || typeof bootstrap.launch_filename !== 'string' || !/^[a-z][a-z0-9_-]*\.json$/.test(bootstrap.launch_filename) + || bootstrap.launch_filename === 'payload.json') { + invariantErrors.push('payload_bootstrap must define positive integer bounds and distinct safe object paths'); + } else if (bootstrap.minimum_url_lifetime_seconds > bootstrap.url_ttl_seconds + || bootstrap.max_manifest_bytes > bootstrap.max_payload_bytes) { + invariantErrors.push('payload_bootstrap minimum lifetime/manifest size exceeds its corresponding maximum'); + } + // Check that both implementations still read the shared block. Literal copies + // would allow a security cap or protocol version to drift across languages. + if (!/CONTRACT\s*=\s*SHARED_CONSTANTS\["payload_bootstrap"\]/.test(fs.readFileSync(PAYLOAD_BOOTSTRAP_PY, 'utf8')) + || !/PAYLOAD_BOOTSTRAP\s*=\s*constants\.payload_bootstrap/.test(fs.readFileSync(PAYLOAD_BOOTSTRAP_TS, 'utf8'))) { + invariantErrors.push('payload_bootstrap consumers must read the shared contract'); + } if (invariantErrors.length > 0) { console.error(`Semantic invariant violations in ${CONSTANTS_JSON}:\n`); diff --git a/scripts/check-types-sync.ts b/scripts/check-types-sync.ts index 162590276..fc10a13dd 100644 --- a/scripts/check-types-sync.ts +++ b/scripts/check-types-sync.ts @@ -91,6 +91,7 @@ const CDK_ONLY_ALLOWLIST = new Set([ 'PendingApprovalRecord', 'ApprovedApprovalRecord', 'DeniedApprovalRecord', + 'CancelledApprovalRecord', // Persisted cancellation row; CLI consumes summaries, not DDB records. 'TimedOutApprovalRecord', 'StrandedApprovalRecord', 'NudgeRecord', diff --git a/yarn.lock b/yarn.lock index c32f94a3f..c14d03d4b 100644 --- a/yarn.lock +++ b/yarn.lock @@ -663,6 +663,20 @@ "@smithy/types" "^4.17.2" tslib "^2.6.2" +"@aws-sdk/client-ssm@^3.1078.0": + version "3.1132.0" + resolved "https://registry.yarnpkg.com/@aws-sdk/client-ssm/-/client-ssm-3.1132.0.tgz#b2c731fa3708366e52b647e1ae16057239c9061a" + integrity sha512-PLFPGor3lEvrhpgt9QOfFY1KBqIaoR+aSUgZ0EtzfVJ8KamNX5Lg63iXc+PBMJJpDTI7pQ2Ko1VIbcOgGr3K0A== + dependencies: + "@aws-sdk/core" "^3.978.0" + "@aws-sdk/credential-provider-node" "^3.972.83" + "@aws-sdk/types" "^3.974.5" + "@smithy/core" "^3.33.3" + "@smithy/fetch-http-handler" "^5.7.2" + "@smithy/node-http-handler" "^4.11.3" + "@smithy/types" "^4.17.2" + tslib "^2.6.2" + "@aws-sdk/client-sts@^3.1078.0": version "3.1119.0" resolved "https://registry.yarnpkg.com/@aws-sdk/client-sts/-/client-sts-3.1119.0.tgz#2b080c257fc5f4549f2cb743546cbb5ae9309034" @@ -720,6 +734,20 @@ bowser "^2.11.0" tslib "^2.6.2" +"@aws-sdk/core@^3.978.0": + version "3.978.0" + resolved "https://registry.yarnpkg.com/@aws-sdk/core/-/core-3.978.0.tgz#115e5d2e03edd380fe270d4617db0f4cc75869d6" + integrity sha512-2yX9LUmxPklVjSGTb8dfnWRJSiFQ3TeH2nn7G1mdKHTfnabzF0+gfrS8rYfLWmZrQ8A3mEcxMJjRc51dL5KWaA== + dependencies: + "@aws-sdk/types" "^3.974.5" + "@aws-sdk/xml-builder" "^3.972.40" + "@aws/lambda-invoke-store" "^0.3.0" + "@smithy/core" "^3.33.3" + "@smithy/signature-v4" "^5.6.12" + "@smithy/types" "^4.17.2" + bowser "^2.11.0" + tslib "^2.6.2" + "@aws-sdk/credential-provider-env@^3.972.55": version "3.972.55" resolved "https://registry.yarnpkg.com/@aws-sdk/credential-provider-env/-/credential-provider-env-3.972.55.tgz#1129acf8860db362a30a02a531387fd399448268" @@ -753,6 +781,17 @@ "@smithy/types" "^4.17.2" tslib "^2.6.2" +"@aws-sdk/credential-provider-env@^3.972.71": + version "3.972.71" + resolved "https://registry.yarnpkg.com/@aws-sdk/credential-provider-env/-/credential-provider-env-3.972.71.tgz#98261bde85bd0f2e2d4a7ba97dce679072e16aff" + integrity sha512-JN+JHruYZw3GUZB8YGAlDk4wTDPOEAEEdEzj5nS0xodWR4smzHsN7PnK2j6IeOsDIj2aqua5DSbhXl9Gtf90FQ== + dependencies: + "@aws-sdk/core" "^3.978.0" + "@aws-sdk/types" "^3.974.5" + "@smithy/core" "^3.33.3" + "@smithy/types" "^4.17.2" + tslib "^2.6.2" + "@aws-sdk/credential-provider-http@^3.972.57": version "3.972.57" resolved "https://registry.yarnpkg.com/@aws-sdk/credential-provider-http/-/credential-provider-http-3.972.57.tgz#4f73bb2f0e03525ab11e001ac18c39cb120dd14c" @@ -792,6 +831,19 @@ "@smithy/types" "^4.17.2" tslib "^2.6.2" +"@aws-sdk/credential-provider-http@^3.972.73": + version "3.972.73" + resolved "https://registry.yarnpkg.com/@aws-sdk/credential-provider-http/-/credential-provider-http-3.972.73.tgz#879466b23dcaec4bee635abf1103e1f5f35d122d" + integrity sha512-uyYYnJOnlis8uQzaYGPd7N1JoioCoNpXgnkXYixsWJXHXgXyYi8WXJSDfofxJeWfQIGWLe2Nwyq60Uc7MZdVOg== + dependencies: + "@aws-sdk/core" "^3.978.0" + "@aws-sdk/types" "^3.974.5" + "@smithy/core" "^3.33.3" + "@smithy/fetch-http-handler" "^5.7.2" + "@smithy/node-http-handler" "^4.11.3" + "@smithy/types" "^4.17.2" + tslib "^2.6.2" + "@aws-sdk/credential-provider-ini@^3.972.62": version "3.972.62" resolved "https://registry.yarnpkg.com/@aws-sdk/credential-provider-ini/-/credential-provider-ini-3.972.62.tgz#a2ca7fc7a1899c0a4a38e501baa666cf1449f2e9" @@ -849,6 +901,25 @@ "@smithy/types" "^4.17.2" tslib "^2.6.2" +"@aws-sdk/credential-provider-ini@^3.973.16": + version "3.973.16" + resolved "https://registry.yarnpkg.com/@aws-sdk/credential-provider-ini/-/credential-provider-ini-3.973.16.tgz#b51db2069b915d563e05969cac845f624aed1dec" + integrity sha512-i++ly+0Uxa+u3ebSSyr0S/3CFhFJDxCXT3+Zj+mW2bXenEx5bKGCdTIKFu39SgXBNhWDjex/8cXUx9MUTMCrTw== + dependencies: + "@aws-sdk/core" "^3.978.0" + "@aws-sdk/credential-provider-env" "^3.972.71" + "@aws-sdk/credential-provider-http" "^3.972.73" + "@aws-sdk/credential-provider-login" "^3.972.78" + "@aws-sdk/credential-provider-process" "^3.972.71" + "@aws-sdk/credential-provider-sso" "^3.973.15" + "@aws-sdk/credential-provider-web-identity" "^3.972.77" + "@aws-sdk/nested-clients" "^3.997.45" + "@aws-sdk/types" "^3.974.5" + "@smithy/core" "^3.33.3" + "@smithy/credential-provider-imds" "^4.4.16" + "@smithy/types" "^4.17.2" + tslib "^2.6.2" + "@aws-sdk/credential-provider-login@^3.972.61": version "3.972.61" resolved "https://registry.yarnpkg.com/@aws-sdk/credential-provider-login/-/credential-provider-login-3.972.61.tgz#f5ba696bae9af3505ae5f3e2a70d7e216d0d37f6" @@ -885,6 +956,18 @@ "@smithy/types" "^4.17.2" tslib "^2.6.2" +"@aws-sdk/credential-provider-login@^3.972.78": + version "3.972.78" + resolved "https://registry.yarnpkg.com/@aws-sdk/credential-provider-login/-/credential-provider-login-3.972.78.tgz#c1c7ca497201a7f2bd6ba7683fa3a896c4e218f8" + integrity sha512-eUtswnXu0+Ii9ieRK+0L7aPFV3Z/dnW2VntJzjBP9xs8s+8p5nBNuymIXtXwZ+5r5+XJP3e32nMkuZ/r0HozEA== + dependencies: + "@aws-sdk/core" "^3.978.0" + "@aws-sdk/nested-clients" "^3.997.45" + "@aws-sdk/types" "^3.974.5" + "@smithy/core" "^3.33.3" + "@smithy/types" "^4.17.2" + tslib "^2.6.2" + "@aws-sdk/credential-provider-node@^3.972.61", "@aws-sdk/credential-provider-node@^3.972.64": version "3.972.64" resolved "https://registry.yarnpkg.com/@aws-sdk/credential-provider-node/-/credential-provider-node-3.972.64.tgz#282456c24bad616faefe2a6c68701913f8289c10" @@ -953,6 +1036,23 @@ "@smithy/types" "^4.17.2" tslib "^2.6.2" +"@aws-sdk/credential-provider-node@^3.972.83": + version "3.972.83" + resolved "https://registry.yarnpkg.com/@aws-sdk/credential-provider-node/-/credential-provider-node-3.972.83.tgz#f4e16b3a1d5d0237460c1b532de37eb6505985dc" + integrity sha512-jdso7ejzfRnatxMUZK4S/U6KbaDPCvfIV4XL+IQAPFDBt5rj5Fq595euqlK8Le4lNCMFR9oUpt+1l0aMgaayOQ== + dependencies: + "@aws-sdk/credential-provider-env" "^3.972.71" + "@aws-sdk/credential-provider-http" "^3.972.73" + "@aws-sdk/credential-provider-ini" "^3.973.16" + "@aws-sdk/credential-provider-process" "^3.972.71" + "@aws-sdk/credential-provider-sso" "^3.973.15" + "@aws-sdk/credential-provider-web-identity" "^3.972.77" + "@aws-sdk/types" "^3.974.5" + "@smithy/core" "^3.33.3" + "@smithy/credential-provider-imds" "^4.4.16" + "@smithy/types" "^4.17.2" + tslib "^2.6.2" + "@aws-sdk/credential-provider-process@^3.972.55": version "3.972.55" resolved "https://registry.yarnpkg.com/@aws-sdk/credential-provider-process/-/credential-provider-process-3.972.55.tgz#bebd30d0065ca34f64e14a870f54a09f23cf11e2" @@ -986,6 +1086,17 @@ "@smithy/types" "^4.17.2" tslib "^2.6.2" +"@aws-sdk/credential-provider-process@^3.972.71": + version "3.972.71" + resolved "https://registry.yarnpkg.com/@aws-sdk/credential-provider-process/-/credential-provider-process-3.972.71.tgz#5a2ff72bcc2d7a658ca6472193ea33437c8b5a4e" + integrity sha512-lYmXJa4gvq4xN1lrT5NiP5vIYYKcGWAdj8y+8o6dlcateB5eF3Dn8DtmjjHKfMBrTPAMr2pebIiX/UOj8c1/UA== + dependencies: + "@aws-sdk/core" "^3.978.0" + "@aws-sdk/types" "^3.974.5" + "@smithy/core" "^3.33.3" + "@smithy/types" "^4.17.2" + tslib "^2.6.2" + "@aws-sdk/credential-provider-sso@^3.972.61": version "3.972.61" resolved "https://registry.yarnpkg.com/@aws-sdk/credential-provider-sso/-/credential-provider-sso-3.972.61.tgz#8c15846c65d2f03bfca3347626e1fcd1a0f40726" @@ -1025,6 +1136,19 @@ "@smithy/types" "^4.17.2" tslib "^2.6.2" +"@aws-sdk/credential-provider-sso@^3.973.15": + version "3.973.15" + resolved "https://registry.yarnpkg.com/@aws-sdk/credential-provider-sso/-/credential-provider-sso-3.973.15.tgz#6a1391ca239f6d505dcb55ea6655b8c14fad5af7" + integrity sha512-6Jhcf4v0pSFdjk1EW2kvzuEBKD+UZ2uNcHUIglKKLndD20YhvkL2kdmDOV5/j4mYuWWwe/a1FQ1aomU86/Cg5Q== + dependencies: + "@aws-sdk/core" "^3.978.0" + "@aws-sdk/nested-clients" "^3.997.45" + "@aws-sdk/token-providers" "3.1129.0" + "@aws-sdk/types" "^3.974.5" + "@smithy/core" "^3.33.3" + "@smithy/types" "^4.17.2" + tslib "^2.6.2" + "@aws-sdk/credential-provider-web-identity@^3.972.61": version "3.972.61" resolved "https://registry.yarnpkg.com/@aws-sdk/credential-provider-web-identity/-/credential-provider-web-identity-3.972.61.tgz#1938cc425e6156acd673484f10b437c5c4c85158" @@ -1061,6 +1185,18 @@ "@smithy/types" "^4.17.2" tslib "^2.6.2" +"@aws-sdk/credential-provider-web-identity@^3.972.77": + version "3.972.77" + resolved "https://registry.yarnpkg.com/@aws-sdk/credential-provider-web-identity/-/credential-provider-web-identity-3.972.77.tgz#5e288111937fcb6f9f02236fa153095d6e24d22f" + integrity sha512-uylIQSUWpfLuH2LovxEEfwzJGM/SabLOfLMg6YXu/E8jJEKUdpdILCVCQCdFvHyu/7dLJOHPMfrSwduxO56NkQ== + dependencies: + "@aws-sdk/core" "^3.978.0" + "@aws-sdk/nested-clients" "^3.997.45" + "@aws-sdk/types" "^3.974.5" + "@smithy/core" "^3.33.3" + "@smithy/types" "^4.17.2" + tslib "^2.6.2" + "@aws-sdk/dynamodb-codec@^3.973.26", "@aws-sdk/dynamodb-codec@^3.973.29": version "3.973.29" resolved "https://registry.yarnpkg.com/@aws-sdk/dynamodb-codec/-/dynamodb-codec-3.973.29.tgz#ea72996b2e1f2c687254c69c357fa4278e42f1da" @@ -1211,6 +1347,20 @@ "@smithy/types" "^4.17.2" tslib "^2.6.2" +"@aws-sdk/nested-clients@^3.997.45": + version "3.997.45" + resolved "https://registry.yarnpkg.com/@aws-sdk/nested-clients/-/nested-clients-3.997.45.tgz#3ed0e5fcd25c889025b2379a9434594fe7056a4a" + integrity sha512-mooq9Q+jLa18VoM7HouczmslZU60iiB0aKc/Ztnq/luIL1ud0z4DnYprLR/ZO1gp331S9tJctM1HZr7u6YKBXQ== + dependencies: + "@aws-sdk/core" "^3.978.0" + "@aws-sdk/signature-v4-multi-region" "^3.996.46" + "@aws-sdk/types" "^3.974.5" + "@smithy/core" "^3.33.3" + "@smithy/fetch-http-handler" "^5.7.2" + "@smithy/node-http-handler" "^4.11.3" + "@smithy/types" "^4.17.2" + tslib "^2.6.2" + "@aws-sdk/s3-presigned-post@^3.1078.0": version "3.1081.0" resolved "https://registry.yarnpkg.com/@aws-sdk/s3-presigned-post/-/s3-presigned-post-3.1081.0.tgz#ed27961af8e772e23745e7ff34f34083bc69c291" @@ -1314,6 +1464,18 @@ "@smithy/types" "^4.17.2" tslib "^2.6.2" +"@aws-sdk/token-providers@3.1129.0": + version "3.1129.0" + resolved "https://registry.yarnpkg.com/@aws-sdk/token-providers/-/token-providers-3.1129.0.tgz#a4a11a789d7259874ccf3733e538fc70f8a289e4" + integrity sha512-Sbl3rpzQdsG4ZK2zh0JWUYyZPKKorJlVOddA2T0DVbKJFrsW8J6wgnslxxUH04+WaBMr4A1HzJZvZX0xUvkniA== + dependencies: + "@aws-sdk/core" "^3.978.0" + "@aws-sdk/nested-clients" "^3.997.45" + "@aws-sdk/types" "^3.974.5" + "@smithy/core" "^3.33.3" + "@smithy/types" "^4.17.2" + tslib "^2.6.2" + "@aws-sdk/types@^3.222.0", "@aws-sdk/types@^3.973.15": version "3.973.15" resolved "https://registry.yarnpkg.com/@aws-sdk/types/-/types-3.973.15.tgz#98a4860bed33c32c7088924d0ab52f9eabbdf7c3"