Repository navigation
Run a US build stage on Modal - #987
Merged
Merged
Conversation
tools/modal_us_stage.py is a Modal app that runs one stage of the ACS local release tool (materialize, calibrate, qa, finalize, package, or all) off this machine. The image is a shallow clone of the plan's pushed commit synced from that tree's own uv.lock (--frozen), asserted clean so the tool's git identity names the real code vintage. Inputs come from a content-addressed Modal volume or the Hugging Face Hub at an explicit revision and are verified against the plan's sha256 before the tool starts. The tool's state is mirrored to a runs volume and every file in it is listed with its sha256 in a receipt. The default invocation is a cheap check (image, clone, environment, the pinned tool's own argument parser on the built argv, input digests, the run's prior state); --run executes. tools/modal_us_stage_plan.py is the pure, stdlib-only half (plan validation, argv, resource classes sized from the #974 measured peaks, hashing, mirroring, receipts, receipt verification) with unit tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
How a plan pins a run (commit, inputs by digest, allowlisted options), what the image, input staging, state mirroring and receipts do, the resource classes and their list-price cost from the #974 measured peaks, the exact commands (digest, upload, validate, check, run, fetch, verify), the #974 materialize replay as the first acceptance test, data placement, and what is not covered yet. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lans runner-smoke is an inline tool (no file in the pinned tree, so it runs on any pushed commit) that imports the synced environment, reads every staged input and writes one state file, which proves the whole run path: inputs staged and verified from the volume and the Hub, state mirrored to the runs volume, receipt written. It runs in the check-sized class. max_wall_seconds bounds a stage's cost below the resource class's hard timeout: the runner stops the tool when it passes and the receipt records stopped_at_budget and FAILED. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What ran on Modal on 22 September (volumes, the #974 replay check with its one correct refusal, the smoke run and its locally verified receipt), the max_wall_seconds budget, the smoke plan, and the exact commands for the heavy replay that has not run yet. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The inline smoke records the policyengine-us version, which the engine-free fast and wheels lanes do not install; the argv-shape test still runs there. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lize Materialize on the state SOI surface at build commit 767312d (branch overnight-acs-local-20260923), with that night's staging H5 and summary, the chronicle_us_b571381 feed and the PUMA ladder pinned by digest, the local run's peak-limit environment, and a 26,400 s wall budget (about $8.90 at list price for the heavy class, under $10 with staging and mirroring). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The receipt's cost estimate covered only the tool's wall, which leaves out staging the inputs (a 10.7 GB staging H5 is copied and hashed), hashing the state tree and mirroring it to the runs volume. The run now records the container's wall and its list-price cost next to the tool's. The runner identity also records the platform, the CPUs the container sees and any *_NUM_THREADS settings, so a Modal run can be compared with a local one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Modal restarts a preempted function on the same input, from scratch, whatever retries says. The 23 September acceptance run was preempted after 54 minutes and restarted with a fresh max_wall_seconds, so the plan's cost cap did not hold across attempts. Each attempt now writes runs/<run_id>/attempts/<stage>-<utc>.json when it starts and rewrites it every 120 seconds. A restart charges the time of every earlier attempt of the same plan that never wrote a receipt to max_wall_seconds and refuses to start with less than a minute left. The receipt lists those attempts and prices all of them; the check reports them and the time left. The runbook's preemption note said a preempted stage had to be run again by hand, which was wrong; it now describes the restart and the ledger. Verified on Modal with the runner smoke (runner-smoke-cadaf418): the attempt record was written, marked finished with its receipt, and the receipt carried the new fields. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every validation pattern now uses fullmatch. With re.match and a "$" anchor, a commit, run_id, branch, repo_url, env key, digest or file name followed by a newline was accepted; each such plan failed later, but the plan check should refuse it. The commit is also shell-quoted where the image build interpolates it. A non-string tool or stage raised TypeError (validate printed a traceback and exited 1); it is now a PlanError (REFUSED, exit 2). An env key under an allowlisted prefix is refused when its name looks like a credential (KEY, TOKEN, SECRET, PASSW, SIGNING, CREDENTIAL): a plan's values are copied into every receipt. POPULACE_LEDGER_API_KEY and MICROCOSM_UK_TERMINAL_GATE_SIGNING_KEY were accepted before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The image checks the plan's branch name out with `checkout -B <branch> FETCH_HEAD`, which puts any name on any commit, and the release tool records that name in its code identity. The runbook said the tool therefore recorded "the real sha and branch"; the branch was only plan-declared. Each container now proves it before running anything: a commits-only fetch (--filter=tree:0) of the branch from the plan's repo into a scratch bare repository, then `merge-base --is-ancestor <commit> <tip>`. The pinned clone is not touched, and the check runs at container start rather than in the cached image layer, so it is true as of the run. The check reports an unverified branch as a problem and the run refuses it; the git block of the receipt records branch_verified, branch_tip and how it was decided. Fetching microcosm's branch history this way took 0.4 s and 4.8 MB locally. Tests import the app against a stub modal module and run the check against a local remote: tip, ancestor, a commit on another branch, and a missing branch. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The tool ran with the container's whole environment, so an attached Hub secret put HF_TOKEN in it; HF_HUB_OFFLINE=1 kept huggingface_hub from using it, but nothing else did. The runner needs the token only to download inputs. tool_environment() now builds the tool's environment: the container's without any variable whose name looks like a credential (KEY, TOKEN, SECRET, PASSW, SIGNING, CREDENTIAL), HF_HUB_OFFLINE=1, and the plan's allowlisted overrides last (re-validated). The check's environment probe and argument parse use the same environment. The receipt lists the names removed (runner.tool_env_removed), never values. The runbook said the tool ran offline only when no secret was attached; it always does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The state-surface materialize at 767312d ran on Modal next to the local run from the same inputs. The check passed, and the Modal run did not finish. Modal preempted it at 54 minutes and again at 2 hours 4 minutes, restarting from zero each time, and it was stopped by hand at about $3.62 (workspace billing report). Per chunk it ran two to five times slower than local (149 to 355 s against 66 to 80 s), CPU-bound on one core under gVisor, and held 15 to 24 GB more RSS. The runbook records these numbers and the local run identity and checkpoint digests for the next comparison. It also replaces the unmeasured 94 GB estimate with the 77.9 GB local peak. It documents the opt-in "nonpreemptible" plan field, priced at three times list. That code (NONPREEMPTIBLE_PRICE_MULTIPLIER, Plan.nonpreemptible, run_stage_heavy_nonpreemptible and run_stage_light_nonpreemptible, and their tests) landed in e70b231. A concurrent commit in the same worktree picked it up from this session's uncommitted edit, so that commit's message does not describe it. The runbook also drops an unobserved claim about out-of-memory restarts. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… attempt The heartbeat stopped before the state was hashed and mirrored, "so that no commit lands mid-mirror"; Modal commits volumes in the background every few seconds anyway, so that bought nothing, and a preemption during hashing or mirroring (many minutes for a large state) went uncharged. The comment also bounded the undercount at one heartbeat, which left out the cold start before _run_stage. The heartbeat now runs until the attempt's final record, and the comment lists what is actually uncounted: cold start and image load, up to one interval after the last record (more after a failed write), and the preemption grace period. Each attempt now ends once with an outcome: "receipt" (finished), "refused" by the lock or the budget (finished, not charged), or "error" for any exception after its first record (unfinished, so charged, like a preemption). A lock serializes the writes, commits and reloads the heartbeat and the main thread make, and a late heartbeat can no longer overwrite the final record. Records carry Modal's input and call ids. The ledger is also the run's lock. An attempt refuses to start while an earlier attempt of the same run, any stage or plan, is still writing its record. A recent record is ambiguous (Modal restarts a preempted input within moments), so the attempt waits two heartbeats and reads again: a record that moved is running, one that did not is dead. A later attempt is left to refuse itself. Volumes have no atomic lock, so this is best effort. No override: a dead attempt stops blocking after one wait, and a running one must not be raced. The check reports recent records without waiting. Receipts are written under a temporary name and renamed, and record the attempt id. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…receipt mirror_tree copied each file straight over its target with copy2, so a preemption mid-copy could leave a truncated file on the runs volume, and a preemption between files a mix of new and old files. Nothing checked the state a restarted or later stage pulled; the runbook said a cut-short attempt's work "is lost", which was not true of one cut short while mirroring. Each file is now copied to a hidden ".<name>.mirror-partial" in its own directory and renamed over the target. Listings skip partials and the next mirror removes them, so no half-written file is ever part of a state tree. After pulling, a stage verifies the state against the run's latest receipt (by finished_at, since receipt names begin with the stage) with strict verify_receipt, and refuses on any difference, before staging inputs; with no receipt the state must be empty. The check hashes the volume's state in place and reports the same problems. The receipt names the receipt the pulled state matched (prior_state_verified_against). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review of the opt-in "nonpreemptible" plan field, whose code landed in e70b231 and whose first test landed in d0bab32 (both swept in from a concurrent session's uncommitted edit; neither message says so). The 3x multiplier matches modal.com/docs/guide/preemption and modal.com/pricing (both read 23 September), the installed Modal client (1.3.0) accepts nonpreemptible= on app.function, and the arithmetic in that test holds. Nothing needed fixing; three gaps are now tested: - the check class refuses nonpreemptible (it was untested); - every class and placement a valid plan can produce has a runner in RUNNERS whose Modal options carry that class's cpu, memory and timeout, retries=0 and the placement (the stub modal now keeps decorated functions and records their options); - every attempt of a non-preemptible plan, including earlier ones, is priced at 3x, and switching the flag is a new plan digest. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
16 tasks
The local entrypoint exited nonzero only on a nonzero return code, so a stage stopped at its max_wall_seconds budget by a tool that exits 0 on the stop signal reported success although its receipt says FAILED. It now fails on any status other than COMPLETED, says why, and prints stopped_at_budget and the all-attempts cost with the rest of the brief. Also: the pricing comment had the floor backwards (an estimate from the request is exact inside the request and a floor above it); the plan module says a Hub revision may be a branch name, the bytes being pinned by sha256; and the app's docstring lists what the check now covers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Runbook, against the evidence in _recovered/scratch-backup/893/overnight-20260923/ (status file, local materialize run.log, Modal client log): - Per-chunk figures: Modal 144.0 to 354.5 s against local 36.8 to 80.1 s (mean 66.3 over 80 chunks), 3.3 to 7.6 times per chunk and 3.93 times over chunks 1 to 29 (6,832 s against 1,737 s); loading 280.6 and 254.6 s against 79.3 s. It said 149 to 355 s against 66 to 80 s and "two to five times". The local wall is 5,604.5 s (5,569 s in the tool). The attempt's app id, launch time, runner commit and stop time are recorded, and that it left no receipt and no state. - The projection for an uninterrupted Modal materialize is now derived from the observed pace: 5.3 to 6.9 hours, $6.40 to $8.40 preemptible or $19 to $25 non-preemptible (it said 3.5 to 7 hours, $4 to $9 or $13 to $26). - A new "Preemption and restarts" section describes the ledger's outcomes and what is charged, the budget (a budget stop writes a FAILED receipt, so a relaunch of the same plan gets its whole budget), the lock and its wait, atomic mirroring and the pulled-state check, and recommends "nonpreemptible": true for heavy stages expected to run longer than about an hour, with the arithmetic: at the 0.67 preemptions an hour the run saw, a preemptible attempt is more likely than not to be cut short past 1.03 hours, and non-preemptible pays for itself in expected cost from 2.8 hours. The sample is two events; the table is labelled an order of magnitude. - How a run works: credential env refusal, input hashing (volume inputs while copied, Hub inputs after download), Hub branch revisions, the pulled-state check and the new receipt fields. The acceptance plan set four peak-limit variables that do not reach materialize at 767312d: the two MICROCOSM_ names are read nowhere, and the two POPULACE_ names only set defaults for with_optional_acs_spine and _preflight_staging_export, which only the staging builder calls. They are removed, which changes the plan sha256 from 5ec595c7 to d1341ec0; the runbook records that the 23 September attempt ran the earlier version (b2e34fe). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The runbook and the refusal message said any `*_NUM_THREADS` override was allowed, but `_ENV_KEY` admitted only OMP, MKL, OPENBLAS and NUMEXPR, so a plan overriding BLIS_NUM_THREADS (which Modal itself sets in the container) was refused with a message saying it was allowed. The allowlist is now the explicit `THREAD_ENV_KEYS` tuple, BLIS included; the refusal message and the runbook list the five names, and tests pin every listed name as allowed and an unlisted *_NUM_THREADS name as refused. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis
marked this pull request as ready for review
September 23, 2026 12:45
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Epic #956, acceleration item E. Main had no Modal code:
git grep -i modal origin/mainfinds only prose, and the one earlier app,tools/modal_rebuild_us_sparse.pyonmodal-sparse-rebuild(48277ea, populace era), was never merged. This PR adds the smallest working path for running a heavy US stage on Modal instead of queueing it on the one 128 GiB machine.What it adds
tools/modal_us_stage.py: a Modal app that runs one stage oftools/build_us_acs_local_release.py(materialize,calibrate,qa,finalize,packageorall). It also runs an inlinerunner-smoketool.uv sync --all-packages --extra us --frozenfrom that tree's ownuv.lock, on Python 3.14. The build assertsHEADequals the commit and that the tree is clean. Each container also proves that the plan's branch contains the commit, becausecheckout -Bwould put any name on any commit. It does a commits-only fetch of the branch into a scratch repository, thenmerge-base --is-ancestor. The check and the run refuse an unverified branch, so the sha and branch the tool records are both real.cas/sha256/<digest>/<name>) or from the Hugging Face Hub at an explicit revision. Each is verified against the plan's sha256 before the tool starts.HF_HUB_OFFLINE=1and without any credential-looking environment variable.HF_TOKENfrom an attached Hub secret is used only by the runner, to download inputs. The receipt lists the names removed, never the values.uv.lockdigests, argv, wall time, peak RSS and list-price cost estimates. Before a stage uses a run's state, it verifies the state against the run's latest receipt with strictverify-receiptand refuses any difference.retriessays. Each attempt writesruns/<run_id>/attempts/<stage>-<id>.jsonwhen it starts and rewrites it every 120 s until it ends, including hashing and mirroring. It ends with an outcome:receipt,refused(by the lock or budget) orerror. A restart charges every unfinished attempt of the same plan tomax_wall_seconds. A budget stop writes a FAILED receipt, so relaunching the same plan gets the whole budget again. The ledger is also the run's lock. An attempt refuses to start while an earlier attempt of the run is still writing its record. A recent record gets one 270 s wait to tell a running attempt from one preemption just killed."nonpreemptible": true: runs the heavy or light class on Modal's non-preemptible placement at 3x list price (modal.com/docs/guide/preemption).validateand the receipts price it that way; the check class refuses it._parse_argson the built argv, and the input digests. It also covers the run's state against its latest receipt, the attempts charged to the budget, and any attempt that may still be running.--runexecutes the stage and fails unless the receipt says COMPLETED.tools/modal_us_stage_plan.py: the pure, stdlib-only half. It validates plans and refuses:fullmatch);It also builds the argv, sets resource classes from the measured Record the full-scale local ACS hours rebuild of 2026-09-22 (#765 build evidence) #974 peaks, handles hashing, mirroring, receipts and the ledger's charging and lock logic, and enforces a
max_wall_secondsbudget. Its CLI hasvalidate,digest(sha256 plus the exactmodal volume putline) andverify-receipt.packages/microcosm-build/tests/test_us_modal_stage_plan_tool.py: 129 tests. They cover argv, receipts, refusals, budget, ledger charging and lock, atomic mirroring, pulled-state verification, env stripping and non-preemptible pricing. The app's helpers are imported against a stubmodalmodule, so CI runs them without the Modal client. The branch check runs against a local git remote. CI runs the file in therestandus-amlanes (tools/ci_test_groups.py --verifyok).docs/us-modal-stage-runbook.md, plus three plans:docs/us-modal-stage-example-plan.json, a replay of Record the full-scale local ACS hours rebuild of 2026-09-22 (#765 build evidence) #974 materialize atcadaf418;docs/us-modal-stage-smoke-plan.json;docs/us-modal-stage-acceptance-20260923-plan.json, the 23 September state-surface materialize at767312d6.The acceptance plan launched with four peak-limit env variables that do not reach materialize at
767312d6. They were removed afterwards, and the runbook records that the attempt ran the earlier version (plan sha2565ec595c7…, nowd1341ec0…).Nothing here uploads to the Hub, touches
latest.jsonor notifies anyone. Publication staystools/publish_release.sh.History note: three commits were made in a worktree shared with a second session. e70b231 also carries that session's
nonpreemptiblecode and d0bab32 its first test, though neither message says so; 9fa2b15 and 5d5c6e9 describe them. They were pushed before this was caught, so the history was not rewritten. e70b231 alone fails one test that d0bab32 updates.Running a stage
Sizing and cost
The classes are sized from the #974 measured peaks (totals SOI surface):
On the state SOI surface, the 23 September build measured materialize locally at a 77.9 GB peak and 5,604.5 s of wall. On Modal the same stage held 15 to 24 GB more RSS at the same chunk. It ran 3.3 to 7.6 times slower per chunk, and 3.93 times slower over chunks 1 to 29 (see below). Modal wall time and cost are therefore well above these local figures.
Modal list prices, read from modal.com/pricing on 22 September: $0.0000131 per core-second and $0.00000222 per GiB-second, billed on the higher of request and use. At those prices the heavy class costs about $1.21 an hour, or $3.63 non-preemptible. Materialize comes to about $1.71 at the local wall time and at most about $9.70 at the 8-hour timeout;
max_wall_secondscan cap it lower.The acceptance run was preempted twice in about 2.97 hours of running, about 0.67 preemptions an hour. At that rate a preemptible attempt longer than about an hour is more likely than not to be cut short, and each cut restarts it from zero. Non-preemptible placement pays for itself in expected cost from about 2.8 hours. The runbook therefore recommends
"nonpreemptible": truefor heavy stages expected to run longer than about an hour, with the arithmetic. Two events is a small sample, so the runbook presents the numbers as an order of magnitude.What ran on Modal (policyengine workspace, 22 September)
microcosm-us-stage-inputsandmicrocosm-us-stage-runs, and uploaded three small Record the full-scale local ACS hours rebuild of 2026-09-22 (#765 build evidence) #974 inputs: the ladder (447 KB), the staging summary (498 KB) and thechronicle_us_b571381feed (165 MB).cadaf418. The synced environment has policyengine-us 2.2.1 and policyengine-core 3.32.5, the same versions Record the full-scale local ACS hours rebuild of 2026-09-22 (#765 build evidence) #974 records, and the tree was clean. The pinned tool parsed the argv, and the three inputs were verified on the volume. The only problem reported was the refusal ofstaging_h5, which was not uploaded (10.7 GB).runner-smokecompleted in 7 s. It verified two volume inputs and one public Hub file (policyengine/populace-us@85a1ccb0…/latest.json), wrote its state file and receipt, andverify-receipt --strictpassed on the fetched state.Total compute was one image build and a few minutes on 2-core, 8 GiB containers. The changes since 9fa2b15 (branch verification, the lock, atomic mirroring, pulled-state verification, env stripping) have been tested locally but have not yet run on Modal.
Acceptance attempt on Modal (23 September)
docs/us-modal-stage-acceptance-20260923-plan.jsonran the overnight state-surface materialize (commit767312d6) on Modal next to the local run, from the same inputs. It ran as appap-mzp14wEyVLIFlYX5qQQxFZon runner commit1d80287af, before the attempt ledger existed. The 10.7 GB staging H5 went up by digest in 255 s. The check passed: clean clone, policyengine-us 2.2.1 and core 3.32.5 as local, argv parsed, and all four inputs verified on the volume.modal app stopat 05:02:22 EDT, because a third attempt would have taken the total past $10. The run left no receipt and no state. It cost about $3.62 at list price, per Modal's workspace billing report.targets_sha256d843209746bdcf96a831fa90d6fa2fe7aa2969fa32290db35afa9889f7c5d377, 5,604.5 s wall, 77.94 GB peak. Its checkpoint digests are in the runbook for the next comparison.So the heavy acceptance replay is still open. At the observed pace, one uninterrupted Modal materialize on this surface takes about 5.3 to 6.9 hours, against 1.6 hours locally. That is $6.40 to $8.40 preemptible if nothing preempts it, or $19 to $25 non-preemptible. The path works end to end up to the engine pass, but it does not yet make this stage faster or cheaper. Fanning the 80 chunks out across containers might, if the chunks are independent, which has not been checked.
Not yet covered
run_identity,targets.jsonor checkpoint digests have been compared with a local run. The Record the full-scale local ACS hours rebuild of 2026-09-22 (#765 build evidence) #974 replay should reproducetargets_sha2560f447ce9…and 1,247 targets. At the Modal pace above it needs about 5.5 hours: about $6.70 preemptible, or about $20 non-preemptible, which the runbook recommends for a run that long.ToolSpecand a class sized from a full-scale measured peak.tools/build_us_puf_support_base.py). It is not registered, because it takes the processed and restricted PUF files, and whether those may go on Modal needs a decision first.Test plan
pytest packages/microcosm-build/tests/test_us_modal_stage_plan_tool.py: 129 passed at f8e8538 (run on an isolatedgit archiveextract of that commit)uv run ruff check .andruff format --checkon the touched filestools/ci_test_groups.py --verifytargets_sha256🤖 Generated with Claude Code