Skip to content

Let compute offers name their provider instead of a connection - #245

Merged
EiffL merged 5 commits into
mainfrom
remove-compute-context
Oct 3, 2026
Merged

EiffL merged 5 commits into
mainfrom
remove-compute-context

Conversation

@EiffL

@EiffL EiffL commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

This simplifies compute.yaml: no connections section, no namespace UUID, no context, no version, and no local settings block. Each offer names its provider directly, and the only local key left is allow_local.

Before:

version: 1
local: {resources: {cpus: 4, memory: 8}}
connections:
  perlmutter:
    namespace: 9d0c0fc5-9be8-407a-a3ec-f17c4110b162
    provider: slurm
    context: perlmutter
offers:
  - name: batch
    connection: perlmutter
    resources: {cpus: 256, memory: 480}
    max_nodes: 16
    time: {default: 1h, max: 12h}
    config: {account: myproject, qos: regular, constraint: cpu}

After:

offers:
  - name: batch
    provider: slurm
    resources: {cpus: 256, memory: 480}
    max_nodes: 16
    time: {default: 1h, max: 12h}
    config: {account: myproject, qos: regular, constraint: cpu}
  - name: local            # optional: replaces the built-in local offer
    provider: local
    resources: {cpus: 4, memory: 8}
    max_nodes: 1
    time: {idle: 30m}

What changed

Catalog shape

  • One native authority per provider. lc runs inside one Slurm environment, so context (which became --clusters) is gone. local is this machine; slurm is the cluster the Slurm client reaches by default. That left connections as pure indirection.
  • Worker launch settings (python, scratch_root, cwd, interface, task_slots_per_node) live in each offer's config. connection_root is the one top-level key, because discovery and connecting read it outside planning.
  • allow_local replaces local: {enabled, resources, time}. An explicit local offer already replaced the built-in one, so local.resources/local.time were a second way to say the same thing, plus a rule forbidding both together. To change the local budget or time limits, write a local offer. allow_local: false (and the NERSC login-node guard, which is unchanged) blocks local launch and execution; local clusters stay inspectable and stoppable.

Checked when the catalog loads

  • connection_root is expanded and resolved once. A relative, .. or unexpandable root (such as ~nosuchuser/…) is now a catalog error naming the field. Before, it crashed with a traceback or showed up as a discovery failure on every command. Providers take the resolved root.
  • An offer's provider must be one lc supports (local, slurm). Before, a typo such as slurmm loaded fine and then made every discovery incomplete, which blocked every launch, local ones included.

Local GPUs by default

  • On Linux, the built-in local offer takes its GPUs from CUDA_VISIBLE_DEVICES, counted the way CUDA reads the mask (it stops at the first invalid entry, so -1 means none). Nothing probes hardware; without a mask, the offer is CPU-only.
  • lc compute launch without --gpus takes a local offer whole, GPUs included. --gpus 0 takes it without them. A remote request (--cpus/--memory) is still CPU-only unless --gpus is given.
  • An explicit local GPU offer is only used when the mask lists at least as many devices as it declares. This replaces the old check that the mask was merely nonempty.

Slurm and status

  • Each submission's log and per-attempt TLS files live under <connection_root>/slurm/<token>/, created before submitting. Connecting from a catalog with a different connection_root now says the job was launched under another root, instead of "scheduler has not started yet". Native inspection and down work from any catalog; connecting needs the launching catalog's root, as for local.
  • A timed-out lc compute status --wait includes the last connection error.
  • Each command builds one provider per name, so Slurm runs id -u once. Local and Slurm share one check for offer setting strings, and a non-numeric sbatch reply is quoted in the error.

Unversioned compute formats

  • Allocations are ephemeral, so compute formats carry no version: no catalog version, no version element in cluster IDs, no schema_version in compute JSON output, and Slurm labels are JobName=lc-<name> and Comment=lightcone:kind=dask:token=<hex>. A format change only strands allocations that end on their own. Manifests keep their schema_version, since they are committed.
  • One cost: a Slurm job of your own named lc-… without lc's comment makes discovery incomplete, naming the job, until it is renamed or ends. Full cluster IDs keep working.

Also

  • Cluster IDs carry the provider instead of the namespace, so a full ID routes without a catalog lookup.
  • launch --dry-run --json reports provider, and status --json keys its errors by provider.
  • Fixes a latent bug: Slurm discovery never filtered on the namespace, so two catalogs with different UUIDs gave one job two IDs.

Breaking (pre-release, no migration)

  • Catalogs with version, connections or local keys fail validation, and the error names the key.
  • Old cluster IDs no longer decode, Slurm job names and comments changed, and private files moved. Stop running clusters before upgrading.

Docs

Not up to date with the follow-up commits. docs/user/cluster.md, docs/cli/compute.md, docs/api/compute.md and docs/user/troubleshooting.md still show version: 1, local.enabled/local.resources, explicit-only local GPUs, schema_version, the lc-v1- labels and submissions/<token>/ paths. Their catalog examples no longer validate. The eval prompt and CLAUDE.md (layer-7 invariants and Recorded decisions) are updated.

Tests

uv run pytest: 1158 passed before the merge from main; test_image.py and test_compute.py were re-run after it (220 passed). ruff and mypy --strict are clean. Nothing was submitted at NERSC.

  • New tests cover catalog errors for bad roots and a misspelled provider, a resolved root handed to one provider per command, mask-derived local GPUs (including -1, UUIDs and invalid entries), the local shortcut with and without --gpus 0, connecting from another root, the last connection error on a status --wait timeout, and allow_local validation.
  • Tests for things that no longer exist are removed: the reserved connection name, namespace uniqueness, --clusters scoping, the catalog version, and the local settings block.

🤖 Generated with Claude Code

lc runs inside one Slurm environment, so a connection's `context`
(`--clusters`) is gone and each provider reaches exactly one native
authority: this machine, or the cluster the Slurm client reaches by
default. That left connections a pure indirection, and the namespace UUID
worse than redundant: Slurm discovery never filtered on it, so two
catalogs with different UUIDs gave one job two IDs and looked for its TLS
material in the wrong directory.

- Offers carry `provider`; worker launch settings move into each offer's
  `config`; `connection_root`, the one setting read outside planning, is
  the catalog's top-level key.
- Cluster IDs encode the provider instead of the namespace, so a full ID
  routes without a catalog lookup. Private files live under
  `<connection_root>/local/` and `<connection_root>/slurm/`.
- Discovery queries `Catalog.providers`: the providers offers name, plus
  local always, so disabled local clusters stay inspectable and stoppable.
- `launch --dry-run --json` reports `provider`; `status --json` errors are
  keyed by provider.

No migration: catalogs with `connections:` fail validation, and clusters
launched before this change should be stopped first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

✅ Eval

Metric Value
Outputs check success
Agent run success
Turns 8
Tool calls 6
Cost $0.12
Agent wall time 0m42s
Model claude-sonnet-5-5
lc status
  mode:    direct
  sandbox: landlock (fs: declared, network: allowed)
  crate:   up to date with the outputs

  · current  baseline/best_fit        8961049
  · current  baseline/hubble_diagram  8961049
  · current  baseline/residuals       8961049

3 current
Confusion & pain points (Claude analysis)

Confusion & pain points

  • Wasted launch probe. The agent ran lc compute launch --wait 2>&1 | tail -3 while inspecting the data. The digest does not show the output, so I can't say whether it succeeded. Later it ran lc materialize local, so it relied on an allocation named local that came from that launch. Piping through tail -3 hid the output that said the name was local. A cluster name the agent has to guess is a gap: lc materialize could print the available cluster name in its refusal. A bare lc compute launch could also say, in its last line, "use lc materialize local".

  • No scaffold for the spec edits. The agent patched astra.yaml with a Python script of string replace calls. It added format, inputs, decisions and the recipe commands, and it matched the existing text exactly. This worked only because the spec stubs had a predictable shape. A mismatch would have left the spec unchanged with no error. Validation passed, but the script has no way to confirm each replacement actually applied. Better tooling would be an astra command that edits or fills output fields, or a skill example that shows the required format: and inputs: fields.

  • Several steps in one command, with the output truncated. The agent chained uv add, validate, git commit, materialize and --check into a single Bash call and truncated the output with tail. A failure in any middle step would have been hard to attribute. It worked this time. The --check result was visible only in a truncated snippet ("up to date b…").

  • Pinned the validator version by hand. The agent ran uvx astra-tools@0.2.18 validate instead of calling an astra CLI directly. That suggests astra was not on PATH or the agent didn't know it was. The environment should provide astra, or the skill should say which invocation to use. The pinned 0.2.18 also looks guessed, with nothing in the trace showing where it came from.

  • License and crate setup done by sed. The agent edited pyproject.toml with sed to add license = "CC-BY-4.0". Its first expression, s/^license.*//, blanks any existing license line instead of replacing it. lc init could scaffold a license field, or the docs could say that setting [project].license turns crate maintenance on. The agent evidently knew this, since it mentions ro-crate-metadata.json.

Full trace: agent-trace artifact on this run.

Review follow-ups to the provider-named offers change:

- Resolve connection_root once when the catalog loads, so a relative,
  '..' or unexpandable root is a catalog error naming the field instead
  of a traceback or a discovery failure. Providers take the resolved
  root, and launch plans report it the same way for local and Slurm.
- Reject an offer whose provider is not in PROVIDERS at load; a typo
  used to load and then leave every discovery incomplete.
- Keep each Slurm submission's log and attempts under
  <root>/slurm/<token>, created before submitting, so connecting from a
  catalog with another root says so instead of "not started yet". A
  timed-out status --wait names its last connection error.
- Give the built-in local offer the GPUs CUDA_VISIBLE_DEVICES exposes,
  read as CUDA reads it, with no hardware probe. The local shortcut
  takes an offer whole unless --gpus 0, and an explicit local GPU offer
  plans only when the mask lists enough devices.
- Build one provider per name per command, share offer-setting string
  validation, and quote sbatch's output when it is not a job ID.
- Drop version markers from compute formats (catalog version, cluster
  ID, schema_version in compute JSON, the v1 in Slurm labels): these
  resources are ephemeral, so a format change only strands allocations
  that end on their own.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
EiffL added a commit that referenced this pull request Oct 3, 2026
The container smoke tests have failed on every branch since 2026-10-02
(#245, `fix/sandbox-elf-interpreter`, `smoke-findings`) with:

```
error: No download found for request: cpython-3.13.16-linux-x86_64-gnu
Error: building at STEP "RUN UV_PYTHON_INSTALL_DIR=/opt/python uv python install 3.13.16 …"
```

## Why

`lc init` pins `.python-version` to the exact interpreter running lc.
The image build then installs that pin with the uv copied from
`UV_IMAGE`. setup-uv gives CI the newest patch of each Python line,
which since 2026-10-02 means 3.11.17 and 3.13.16. The pinned uv 0.12.5
predates both. The 3.12 job only passed because 3.12.14 is still known;
3.12.15 is out and would have broken it next.

## Change

`UV_IMAGE` →
`ghcr.io/astral-sh/uv:0.12.23@sha256:61d393e44e249f2e4b526b6c7ddcecce245946826e608e11c93ad4f5bba55b21`.

- **Digest:** this is the manifest-list digest, covering `linux/amd64`
and `linux/arm64`. I read it from ghcr's registry API and confirmed it
by hashing the index body. The same method reproduces the old 0.12.5 pin
(`e85be844…`) exactly.
- **Version:** 0.12.22 is the first release that knows the new patches;
0.12.23 is the current one.

## Effect on projects

This is an engine constant whose bump is meant to be visible (see the
layer-6 invariants in CLAUDE.md):

- **Containerized projects** get a new image tag, so `lc build`
rebuilds, and a new `env_version`. Existing outputs read `behind`;
nothing is remade.
- **Direct-mode projects** are unaffected.

## Verification

- **Pinned images, run by digest:** 0.12.5 lists none of cpython
3.11.17, 3.12.15 or 3.13.16; 0.12.23 lists all three.
- **`tests/test_container_smoke.py` under Python 3.13.16 (podman):** the
old pin fails with CI's exact error and image tag
(`lc-env-712be96723c9091c`). The new pin passes all 6, including
materialize in the image and `datalad rerun` on a clone that starts with
no file contents.
- **Under Python 3.11.17:** 6 passed.
- **`test_image.py`, `test_container.py`, `test_identity.py`:** 96
passed.

## Note

This will happen again whenever a Python patch release is newer than the
pinned uv: CI moves to the newest patch, and the image's uv does not.
This PR only restores CI; a structural fix is a separate decision.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
EiffL and others added 2 commits October 3, 2026 12:16
A catalog's own local offers already replaced the built-in one, so
local.resources and local.time were a second way to say what an offer
says, with a rule forbidding the two together. Customizing local compute
is now an explicit local offer, and the catalog's only local key is
allow_local, which blocks local launch and execution while keeping
inspection and termination.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- One ProviderName type checks membership in PROVIDERS for offers and
  cluster IDs alike, so Compute.provider indexes the registry directly
  and Identity.decode stops re-checking what the fields already check.
- The Slurm job-name pattern is built from its prefix and the shared
  cluster-name pattern, and one attempt_directory serves both the
  bootstrap that writes connection material and connect that reads it.
- Drop the launch-plan guards no caller could trip: Compute routes each
  plan to the one provider that made it.
- allow_local keeps what the catalog says; local_disabled_reason already
  folds in the NERSC guard.
- Request.parse treats a missing gpus as CPU only, and each provider
  reads its config strings once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@EiffL
EiffL merged commit 5166fdc into main Oct 3, 2026
8 of 9 checks passed
@EiffL
EiffL deleted the remove-compute-context branch October 3, 2026 20:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant