Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
66 changes: 61 additions & 5 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -952,10 +952,10 @@ none; formats may, `tar.gz`), never `Path.stem`. The old
`_HASH_EXCLUDE` is gone with the directory that made it necessary: the
manifest cannot be inside the thing it describes any more.

**Dask owns the ordering.** Every task is submitted with its upstream
futures as arguments, so the dependency order, the parallelism, and the
scheduling all fall out of the argument graph. There is no ready-set loop
and no hand-rolled topological sort in the execution path.
**Dask owns the ordering.** Tasks that may execute receive upstream futures
or already-current `TaskResult` values as arguments, so dependency order,
parallelism, and scheduling fall out of the argument graph. There is no ready-set
loop and no hand-rolled topological sort in the execution path.
`Graph.order()` exists for the read-only walk, which has to classify a
task after everything upstream of it — and for submitting in an order
where a task's upstream handles already exist.
Expand Down Expand Up @@ -1165,7 +1165,7 @@ its sync touches only ignored paths — so both modes run one order.)

**A run fetches its declared inputs; the read-only verbs never do.**
`materialize` batch-runs `git annex get` over the graph's in-tree
declared inputs before anything hashes (driver-side — the storage
declared inputs before workers hash or execute (driver-side — the storage
invariant that nobody is ever asked to run an annex command by hand),
so a bytes-free clone materializes straight to up-to-date. A failed
fetch is a *warning*, never a refusal: independent tasks still run and
Expand Down Expand Up @@ -1727,6 +1727,10 @@ and use it for every native ownership check and filter.
**Local compute needs no setup.** An absent implicit `~/.lightcone/compute.yaml`
selects a built-in local catalog: one CPU, 1 GiB, one node, fast startup, 30-minute
default and two-hour maximum lifetime. It writes no catalog and starts no cluster.
GPU offers require an explicit catalog and, for local launches, a nonempty
`CUDA_VISIBLE_DEVICES` mask on Linux. No GPU auto-discovery. Local GPU capacity and
model labels are configured, not hardware-verified; allocations do not reserve
devices exclusively against other host programs or allocations.
Configured catalogs replace it completely; missing explicit paths and invalid
files are errors. Execution still requires an explicitly launched cluster's name or ID.

Expand All @@ -1753,6 +1757,17 @@ Connection names live only in the catalog's mapping keys. Use validated `replace
for updates and the explicit `as_dict` allowlists for public output. Preserve
duplicate-key rejection in YAML; providers validate their own `launch` and `config`.

**Allocation syntax follows SkyPilot without depending on SkyPilot.** Compute
CPU/memory requests support exact quantities or `+` minimums. Compute memory
uses binary units: bare `32`, `32GB`, and `32GiB` agree. Catalog resources use
one `accelerators: NAME[:COUNT]` or a one-entry mapping; CLI `--gpus A100:4`,
`A100`, or generic `GPU:4` selects an exact positive whole count, while `0`
means CPU only. Type matching is case-insensitive; no GPU `+`, fractions, or
global model alias registry. Local accelerator labels are trusted configuration.
Named Slurm offers must map their public label to the site's GRES type through
`config.gpu_type`; generic `GPU` offers may omit that setting. Preserve native
evidence in observations.

**Configured compute roots may be filesystem aliases.** Resolve connection and
scratch roots before appending managed namespace, submission, or attempt paths.
Keep symlink rejection within those managed paths and enforce private directory
Expand All @@ -1779,6 +1794,47 @@ Any driver failure while tasks are outstanding (a failed commit included, not
only a cluster error) carries `compute.UNSTOPPED`, the one wording for "the
allocation was not stopped and unreported tasks may still be running".

**Recipe resources use standard Dask admission.** Preserve ASTRA `recipe.resources`
in `plan.Task` as raw mappings so `status` and `--check` remain independent of
executor support. Parse `TaskResources` at execution admission: whole CPUs, memory
bytes, and whole GPU counts. Reuse the read-only classification walk before
admission: known current/behind outputs become values without Dask submission.
Validate tasks that may execute, including dependents of potentially rebuilt
outputs; workers recheck actual upstream digests. Normalize worker budgets once
with `worker_capacities`, then pass reservations explicitly to submission. Recipe
`time_limit` is unsupported and must fail explicitly; allocation walltime remains supported.
Workers advertise CPU/MEMORY/GPU; tasks reserve their declarations. Omitted RAM
adds no memory reservation; CPU requests and task slots govern concurrency.
Probes reserve all whole-worker budgets.
GPU recipes reserve the worker's full GPU budget, one GPU recipe at a time, and
inherit its whole allocation mask. Recipe `gpus` is a minimum capacity requirement,
not a visibility limit; it defaults to zero and does not select a model. Recipe
memory retains ASTRA units (`8Gi` binary, `8GB` decimal, no bare quantities),
independently of compute's SkyPilot units.
Thread slots remain a separate concurrency cap. Reservations are cooperative, not
per-command OS CPU/RAM limits or BLAS thread counts. Local Nanny defaults keep
OMP/MKL/OPENBLAS threads at one unless the launch environment overrides them; Slurm
uses direct workers and the job environment. Unsupported disk/model requests and
fractional CPU/GPU counts fail explicitly. Exact bytes are shared in `units.py`;
allocation durations are parsed in `compute.model`, with error conversion only
at the CLI request boundary.

**GPU visibility comes from the allocation, not device discovery.** No CUDA probe,
UUID inventory, MIG detection, model verification, or custom Dask worker. Local
launch freezes the externally supplied nonempty CUDA mask and optional device
order. Slurm validates native GPU counts, preserves its mask, and sets
`CUDA_DEVICE_ORDER=PCI_BUS_ID`. `exec_policy(use_gpus=True)` inherits that whole
mask; CPU commands get an empty one. Never mutate the reusable worker's environment.
Direct GPU policies grant native NVIDIA character nodes; OS permissions and cgroups
remain authoritative. The host must initialize NVIDIA character devices including
UVM before launch; lc neither loads drivers nor creates nodes. Standalone GPU
reruns need an explicit CUDA mask in their own environment. Container GPU execution
supports podman-hpc `--gpu` only. Explicit GPU recipes on ordinary Docker/Podman
fail before image preparation; probes use CPU policy with a diagnostic note while
retaining their whole-worker reservation. CPU containers remain supported on all
runtimes and set `NVIDIA_VISIBLE_DEVICES=void`. Physical GPU execution remains
unvalidated; tests check real subprocess masks and native argv without GPU hardware.

**One catalog selector, `LC_COMPUTE_CONFIG` (2026-09).** `lc compute --config`
was removed: `run` and `materialize` resolve clusters through the catalog too,
and a per-invocation override on one command group launched allocations those
Expand Down
77 changes: 72 additions & 5 deletions docs/api/compute.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ It owns no service, registry, or saved current-cluster selection.

| Symbol | Contract |
|---|---|
| `Request.parse(...)` | Common exact/minimum CPU and memory requests, node count, walltime, startup class. |
| `Request.parse(...)` | Exact/minimum CPU and memory requests, exact accelerator type/count, node count, walltime, startup class. |
| `Catalog.load(path)` | Ordered fixed shapes and stable connection namespaces; use the built-in local catalog only when the implicit default file is absent. |
| `Compute.plan(request, *, name=None)` | Select an eligible offer and freeze its native launch settings and optional name without allocation. |
| `Compute.launch(plan)` | Check names across native authorities, generate one if omitted, submit once, and return a self-contained `Identity`. |
Expand All @@ -17,14 +17,15 @@ It owns no service, registry, or saved current-cluster selection.
| `Provider` | `plan`, `launch`, `discover`, `inspect`, `connect`, `terminate`. |

`Catalog.load()` defaults to `~/.lightcone/compute.yaml`. When that implicit file
is absent, the built-in catalog exposes one `local` offer: one CPU, 1 GiB, one node,
fast startup, 30-minute default and two-hour maximum lifetime. It creates no
configuration file or allocation. Configured catalogs replace it completely.
is absent, the built-in catalog exposes a `local` offer: one CPU, 1 GiB, one node,
fast startup, 30-minute default and two-hour maximum lifetime. GPU offers require
an explicit catalog. It creates no configuration file or allocation. Configured
catalogs replace it completely.
Missing paths selected through an argument or `LC_COMPUTE_CONFIG`, unreadable
files, and invalid catalogs remain errors. Stable connection namespaces let
separate invocations discover and attach to the same local allocations.

`model.py` defines the shared Pydantic models: `Connection`, `Offer`, `Resources`,
`model.py` defines the shared Pydantic models: `Connection`, `Offer`, `Resources`, `Accelerator`,
`TimeLimits`, `Startup`, `Request`, `Identity`, `LaunchPlan`, and `Snapshot`.
`Catalog` validates YAML directly into these objects, which providers also use.
Unknown common fields are rejected; schema errors identify paths such as
Expand All @@ -39,6 +40,14 @@ keeps the configured `default` and `max` duration strings and exposes
`default_seconds` and `max_seconds`. `Startup.class_` corresponds to YAML `class`.
Connection names exist only as catalog mapping keys, referenced by `Offer.connection`.

Compute memory accepts bare GiB quantities and SkyPilot-style binary units:
`32`, `32GB`, and `32GiB` agree. CPU and memory requests accept a trailing `+`.
`Accelerator` accepts one `NAME[:COUNT]` or one-entry mapping, such as `A100:4`
or `{A100: 4}`, and serializes to that mapping. Counts are exact positive integers;
type matching is case-insensitive, and the generic name `GPU` accepts any model.
No accelerator registry or model alias expansion is maintained. Slurm's
`config.gpu_type` maps a named catalog accelerator to its native GRES type.

Model constructors take keyword arguments. `replace(...)` validates updates;
`model_dump()` and `model_validate()` support internal roundtrips without changing
units. Keep the explicit `as_dict()` methods for public CLI output so internal
Expand Down Expand Up @@ -113,6 +122,64 @@ Dask chooses the workers and handles dependencies; invocation-specific keys prev
unintended reuse across commands. There is no worker-selection layer, per-worker
preflight orchestration, source fingerprinting, or login-node guard. Driver-side
preparation and the existing task runtime/sandbox checks remain in their owners.

Workers advertise standard Dask `CPU`, `MEMORY`, and `GPU` resources; memory is measured
in bytes. `engine.execution_resources.TaskResources` validates ASTRA's
`recipe.resources` into whole CPUs, bytes, and a whole GPU count at
execution admission. `plan.Task` preserves the ASTRA mapping so read-only
classification does not impose executor restrictions. `worker_capacities(workers)`
normalizes advertised budgets once; `requirements(capacities)` checks that one
worker can satisfy a task and returns its `Client.submit` resource dictionary.
Omitted memory adds no `MEMORY` reservation. `whole_worker=True` reserves CPU,
memory, and GPUs for a probe. Recipe GPU counts default to zero; a GPU recipe
reserves the full GPU budget of a fitting worker
and inherits its whole allocation mask. The requested count is a minimum, not a
visibility limit. This serializes GPU recipes per worker without device assignment.
Unsupported disk/type requests and fractional CPU/GPU counts fail before execution.

The driver reuses the read-only classification walk before admission. Known
current or unrefreshed behind outputs become `TaskResult` values, without Dask
submission or resource reservations. Tasks that may execute, including dependents
of potentially rebuilt outputs, have their resource requests validated before
preparation. Workers recheck actual upstream digests and may still skip a reserved
task if its inputs prove unchanged. Allocation and task requests share byte
conversion utilities; their models remain distinct because allocation selection supports minimum quantities
and node counts. Standard Dask scheduling accounts for
concurrent CPU, memory, and GPU reservations; Dask execution-thread counts remain a
separate concurrency cap. Reservations do not impose hard limits on recipe
subprocesses. Recipe `time_limit` is unsupported and explicitly refused before
preparation or execution; allocation walltime remains supported.

Recipe memory remains ASTRA-style: `8Gi` is binary, `8GB` is decimal, and units
are required. Allocation memory follows the compute convention above; keep the
two parsers' contracts explicit even though they share exact byte arithmetic in
`units.py`. Allocation duration parsing stays in `compute.model.duration`, raising
`ValueError` for Pydantic; `Request.parse` converts it to `ComputeError`.

## GPU allocation and visibility

Local GPU offers require Linux and an explicit nonempty `CUDA_VISIBLE_DEVICES`.
Planning freezes that mask and optional `CUDA_DEVICE_ORDER`; launch passes them
to the worker unchanged. Count and model are catalog declarations, not hardware
observations. The built-in local offer remains CPU-only.

Slurm requests native GPU GRES and validates `SLURM_GPUS_ON_NODE` before
advertising the worker's GPU budget. Bootstrap preserves Slurm's CUDA mask and
sets `CUDA_DEVICE_ORDER=PCI_BUS_ID`. There is no CUDA probe, device inventory, or
custom Dask worker.

The sandbox's `use_gpus` policy option inherits the worker's mask for GPU commands
and supplies an empty mask for CPU commands, without modifying the reusable
worker's environment. Direct GPU policies grant native NVIDIA character devices.
Container GPU execution uses podman-hpc's `--gpu`. Explicit GPU recipes on ordinary
Docker or Podman are refused before image preparation; probes use a CPU policy and
report that GPU access is unavailable while retaining their whole-worker reservation.
Native permissions and cgroups remain authoritative. NVIDIA devices, including UVM,
must already exist; policy construction does not load drivers or create devices.
See [GPU deployment requirements](../user/cluster.md#gpu-allocations).

## Execution output and teardown

`output.py` transports byte chunks through standard Dask events so detached
workers' output reaches the invoking CLI. It uses the borrowed client's event
topic, which the schedulers lc launches drop as soon as the client disconnects
Expand Down
4 changes: 2 additions & 2 deletions docs/api/container.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,10 +18,10 @@ Sources: `src/lightcone/engine/image.py`,
| `image.tag(root)` | `lc-env-<16 hex>` over the rendered Containerfile *and* the identity document. |
| `image.archive_path(root, tag)` | `.datalad/environments/<tag>/image` — the `datalad containers-add` layout. |
| `container.build(root)` | Build + save + commit, idempotent; returns `(Runtime, "built" \| "present")`. |
| `container.runtime_for_run(root, *, build)` | One function, two strictnesses: `lc build`/materialize-preflight may build and commit; the probe and worker only ever find, fetch, and load. |
| `container.runtime_for_run(root, *, build, use_gpus=False)` | Resolve the runtime, refusing unsupported explicit GPU requests before preparing the image. Materialize may build and commit; probes and reruns only find, fetch, and load. |
| `container.backend(...)` | The single construction point for the exec backend — the only mode branch. |
| `container.sync(...)` | The in-container environment converge: network on, project `:rw`, host uv cache mounted, into `.lightcone/venv`. |
| `Runtime` | Facts only — root/mode/name/tag/id/arch — never mechanism. |
| `Runtime` | Resolved execution facts; `supports_gpus` is true for direct mode and podman-hpc. |

## What must stay true

Expand Down
1 change: 1 addition & 0 deletions docs/api/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,7 @@ is responsibility and contract, not every signature.
| [`worker`](worker.md) | Making one output; the rerun entry point | impure |
| [`materialize`](materialize.md) | The driver: gates, scheduling, the save/restore loop, status | impure |
| [`compute`](compute.md) | Resource requests, native allocation lifecycle, borrowed Dask clients | impure |
| [`execution_resources`](compute.md) | Task resource admission | pure |
| [`sandbox`](sandbox.md) | The exec boundary: policy, backends, attestation, denials | mixed |
| [`image` & `container`](container.md) | The container hatch: declaration → image → archive → runtime | pure / impure |
| [`crate`](crate.md) | The publication view: the repo as an RO-Crate | pure |
Expand Down
10 changes: 7 additions & 3 deletions docs/api/materialize.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,16 +18,20 @@ driver's stderr, independently of success or failure, leaving stdout for the rep
| `check(root, targets, *, refresh)` | The same classification without executing, committing, or fetching. Exempt from the dirty refusal. |
| `status(root)` | The report: every output's state and provenance commit, plus the mode/image/sandbox header facts. |
| `MaterializeReport` / `StatusReport` | The JSON surfaces; `ok` and `up_to_date` first. |
| `cluster_for_run(cluster_id)` | Borrow the cluster; the submit/completed scheduler seam (`submit`, `completed`). |
| `cluster_for_run(cluster_id)` | Borrow the cluster; expose resource validation, submission, and completion. |
| `run_record(...)` / `datalad_run_subject(...)` | The commit message `datalad rerun` replays, and the one spelling of its subject line — shared with the foreign-write comparator, because two strings here would drift. |
| `_engine_requirement()` | How a record pins its engine: by version for a release, by source commit (hatch-vcs) for a dev build. |

## The run's order, and why

1. **Read-only project checks before connecting** — tool, committer, dirty-tree,
spec and lock errors do not require a reachable cluster to report.
spec and lock errors do not require a reachable cluster to report. The shared
classification walk identifies outputs already current or left behind.
2. **Explicit cluster before preparing the environment** — validate native
allocation identity and connect before fetching inputs or building an image.
allocation identity, connect, and validate CPU/memory/GPU requests for tasks
that may execute. Known skips become values without resource reservations;
dependents of potentially rebuilt outputs still need admission. Explicit GPU
recipes must also have a supported runtime before any image build.
The dirty refusal has already run: in
containerized mode the converge can commit an image archive, and
`dataset.save` commits the whole index; on a dirty tree the user's
Expand Down
Loading
Loading