Let compute offers name their provider instead of a connection - #245
Conversation
lc runs inside one Slurm environment, so a connection's `context` (`--clusters`) is gone and each provider reaches exactly one native authority: this machine, or the cluster the Slurm client reaches by default. That left connections a pure indirection, and the namespace UUID worse than redundant: Slurm discovery never filtered on it, so two catalogs with different UUIDs gave one job two IDs and looked for its TLS material in the wrong directory. - Offers carry `provider`; worker launch settings move into each offer's `config`; `connection_root`, the one setting read outside planning, is the catalog's top-level key. - Cluster IDs encode the provider instead of the namespace, so a full ID routes without a catalog lookup. Private files live under `<connection_root>/local/` and `<connection_root>/slurm/`. - Discovery queries `Catalog.providers`: the providers offers name, plus local always, so disabled local clusters stay inspectable and stoppable. - `launch --dry-run --json` reports `provider`; `status --json` errors are keyed by provider. No migration: catalogs with `connections:` fail validation, and clusters launched before this change should be stopped first. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
✅ Eval
lc statusConfusion & pain points (Claude analysis)Confusion & pain points
Full trace: |
Review follow-ups to the provider-named offers change: - Resolve connection_root once when the catalog loads, so a relative, '..' or unexpandable root is a catalog error naming the field instead of a traceback or a discovery failure. Providers take the resolved root, and launch plans report it the same way for local and Slurm. - Reject an offer whose provider is not in PROVIDERS at load; a typo used to load and then leave every discovery incomplete. - Keep each Slurm submission's log and attempts under <root>/slurm/<token>, created before submitting, so connecting from a catalog with another root says so instead of "not started yet". A timed-out status --wait names its last connection error. - Give the built-in local offer the GPUs CUDA_VISIBLE_DEVICES exposes, read as CUDA reads it, with no hardware probe. The local shortcut takes an offer whole unless --gpus 0, and an explicit local GPU offer plans only when the mask lists enough devices. - Build one provider per name per command, share offer-setting string validation, and quote sbatch's output when it is not a job ID. - Drop version markers from compute formats (catalog version, cluster ID, schema_version in compute JSON, the v1 in Slurm labels): these resources are ephemeral, so a format change only strands allocations that end on their own. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The container smoke tests have failed on every branch since 2026-10-02 (#245, `fix/sandbox-elf-interpreter`, `smoke-findings`) with: ``` error: No download found for request: cpython-3.13.16-linux-x86_64-gnu Error: building at STEP "RUN UV_PYTHON_INSTALL_DIR=/opt/python uv python install 3.13.16 …" ``` ## Why `lc init` pins `.python-version` to the exact interpreter running lc. The image build then installs that pin with the uv copied from `UV_IMAGE`. setup-uv gives CI the newest patch of each Python line, which since 2026-10-02 means 3.11.17 and 3.13.16. The pinned uv 0.12.5 predates both. The 3.12 job only passed because 3.12.14 is still known; 3.12.15 is out and would have broken it next. ## Change `UV_IMAGE` → `ghcr.io/astral-sh/uv:0.12.23@sha256:61d393e44e249f2e4b526b6c7ddcecce245946826e608e11c93ad4f5bba55b21`. - **Digest:** this is the manifest-list digest, covering `linux/amd64` and `linux/arm64`. I read it from ghcr's registry API and confirmed it by hashing the index body. The same method reproduces the old 0.12.5 pin (`e85be844…`) exactly. - **Version:** 0.12.22 is the first release that knows the new patches; 0.12.23 is the current one. ## Effect on projects This is an engine constant whose bump is meant to be visible (see the layer-6 invariants in CLAUDE.md): - **Containerized projects** get a new image tag, so `lc build` rebuilds, and a new `env_version`. Existing outputs read `behind`; nothing is remade. - **Direct-mode projects** are unaffected. ## Verification - **Pinned images, run by digest:** 0.12.5 lists none of cpython 3.11.17, 3.12.15 or 3.13.16; 0.12.23 lists all three. - **`tests/test_container_smoke.py` under Python 3.13.16 (podman):** the old pin fails with CI's exact error and image tag (`lc-env-712be96723c9091c`). The new pin passes all 6, including materialize in the image and `datalad rerun` on a clone that starts with no file contents. - **Under Python 3.11.17:** 6 passed. - **`test_image.py`, `test_container.py`, `test_identity.py`:** 96 passed. ## Note This will happen again whenever a Python patch release is newer than the pinned uv: CI moves to the newest patch, and the image's uv does not. This PR only restores CI; a structural fix is a separate decision. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
A catalog's own local offers already replaced the built-in one, so local.resources and local.time were a second way to say what an offer says, with a rule forbidding the two together. Customizing local compute is now an explicit local offer, and the catalog's only local key is allow_local, which blocks local launch and execution while keeping inspection and termination. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- One ProviderName type checks membership in PROVIDERS for offers and cluster IDs alike, so Compute.provider indexes the registry directly and Identity.decode stops re-checking what the fields already check. - The Slurm job-name pattern is built from its prefix and the shared cluster-name pattern, and one attempt_directory serves both the bootstrap that writes connection material and connect that reads it. - Drop the launch-plan guards no caller could trip: Compute routes each plan to the one provider that made it. - allow_local keeps what the catalog says; local_disabled_reason already folds in the NERSC guard. - Request.parse treats a missing gpus as CPU only, and each provider reads its config strings once. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This simplifies
compute.yaml: noconnectionssection, no namespace UUID, nocontext, noversion, and nolocalsettings block. Each offer names itsproviderdirectly, and the only local key left isallow_local.Before:
After:
What changed
Catalog shape
context(which became--clusters) is gone.localis this machine;slurmis the cluster the Slurm client reaches by default. That left connections as pure indirection.python,scratch_root,cwd,interface,task_slots_per_node) live in each offer'sconfig.connection_rootis the one top-level key, because discovery and connecting read it outside planning.allow_localreplaceslocal: {enabled, resources, time}. An explicit local offer already replaced the built-in one, solocal.resources/local.timewere a second way to say the same thing, plus a rule forbidding both together. To change the local budget or time limits, write a local offer.allow_local: false(and the NERSC login-node guard, which is unchanged) blocks local launch and execution; local clusters stay inspectable and stoppable.Checked when the catalog loads
connection_rootis expanded and resolved once. A relative,..or unexpandable root (such as~nosuchuser/…) is now a catalog error naming the field. Before, it crashed with a traceback or showed up as a discovery failure on every command. Providers take the resolved root.providermust be one lc supports (local,slurm). Before, a typo such asslurmmloaded fine and then made every discovery incomplete, which blocked every launch, local ones included.Local GPUs by default
CUDA_VISIBLE_DEVICES, counted the way CUDA reads the mask (it stops at the first invalid entry, so-1means none). Nothing probes hardware; without a mask, the offer is CPU-only.lc compute launchwithout--gpustakes a local offer whole, GPUs included.--gpus 0takes it without them. A remote request (--cpus/--memory) is still CPU-only unless--gpusis given.Slurm and status
<connection_root>/slurm/<token>/, created before submitting. Connecting from a catalog with a differentconnection_rootnow says the job was launched under another root, instead of "scheduler has not started yet". Native inspection anddownwork from any catalog; connecting needs the launching catalog's root, as for local.lc compute status --waitincludes the last connection error.id -uonce. Local and Slurm share one check for offer setting strings, and a non-numericsbatchreply is quoted in the error.Unversioned compute formats
version, no version element in cluster IDs, noschema_versionin compute JSON output, and Slurm labels areJobName=lc-<name>andComment=lightcone:kind=dask:token=<hex>. A format change only strands allocations that end on their own. Manifests keep theirschema_version, since they are committed.lc-…without lc's comment makes discovery incomplete, naming the job, until it is renamed or ends. Full cluster IDs keep working.Also
launch --dry-run --jsonreportsprovider, andstatus --jsonkeys itserrorsby provider.Breaking (pre-release, no migration)
version,connectionsorlocalkeys fail validation, and the error names the key.Docs
Not up to date with the follow-up commits.
docs/user/cluster.md,docs/cli/compute.md,docs/api/compute.mdanddocs/user/troubleshooting.mdstill showversion: 1,local.enabled/local.resources, explicit-only local GPUs,schema_version, thelc-v1-labels andsubmissions/<token>/paths. Their catalog examples no longer validate. The eval prompt and CLAUDE.md (layer-7 invariants and Recorded decisions) are updated.Tests
uv run pytest: 1158 passed before the merge frommain;test_image.pyandtest_compute.pywere re-run after it (220 passed).ruffandmypy --strictare clean. Nothing was submitted at NERSC.-1, UUIDs and invalid entries), the local shortcut with and without--gpus 0, connecting from another root, the last connection error on astatus --waittimeout, andallow_localvalidation.--clustersscoping, the catalogversion, and thelocalsettings block.🤖 Generated with Claude Code