Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
61 changes: 51 additions & 10 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -182,7 +182,7 @@ src/lightcone/ # namespace — NO __init__.py
├── compute/ # explicit allocations and borrowed Dask clients
│ ├── __init__.py # Compute: catalog, resolve, launch, status, down; connect()
│ ├── model.py # the shared Pydantic models and the Provider protocol
│ ├── catalog.py # compute.yaml, or the built-in local offer
│ ├── catalog.py # compute.yaml with local defaults and policy
│ ├── runtime.py # private files, TLS material, the scheduler config
│ ├── local.py # local provider: validated OS process identities
│ ├── local_runtime.py # the detached LocalCluster owner
Expand Down Expand Up @@ -1710,7 +1710,8 @@ identities are the allocation authority; standard Dask supplies execution state.
No Lightcone server, lifecycle database, custom Dask worker, or implicit allocation.

**Names are native labels, not a registry.** `launch --name analysis` chooses a name;
otherwise launch generates `lc-` plus 12 random hexadecimal characters. Plain stdout
otherwise the local shortcut uses `local`, and explicit resource requests generate
`lc-` plus 12 random hexadecimal characters. Plain stdout
contains only the name for shell capture; JSON retains the full immutable ID too.
Resolve names through fresh discovery and refuse missing, ambiguous, or incomplete
observations. Check existing names before submission, but do not claim atomic global
Expand All @@ -1724,15 +1725,25 @@ comment retention requires Slurm's `AccountingStoreFlags` to include `job_commen
Resolve the Slurm command user's UID through `id -u` on the same command runner,
and use it for every native ownership check and filter.

**Local compute needs no setup.** An absent implicit `~/.lightcone/compute.yaml`
selects a built-in local catalog: one CPU, 1 GiB, one node, fast startup, 30-minute
default and two-hour maximum lifetime. It writes no catalog and starts no cluster.
**Local compute needs no setup.** The built-in local offer provides detected usable
logical CPUs and RAM, one node, fast startup, a 30-minute default and two-hour
maximum lifetime. Loading the catalog writes no catalog and starts no cluster.
`lc compute launch` without CPU/memory flags selects only local offers and defaults
the name to `local`. `--wait` returns when the accepted allocation is ready; timeout
or startup failure retains its ID without resubmitting or terminating it.
Configured remote offers precede the built-in local offer in selection order.
Explicit local connections supply their own offers instead. `local.resources`
overrides the built-in CPU/RAM budget and cannot accompany explicit local connections.
`local.enabled: false` blocks local launch and execution while preserving inspection
and termination. Recognized NERSC login nodes disable local compute automatically;
other sites can disable it in their catalogs. Native permissions remain the
enforcement boundary.
GPU offers require an explicit catalog and, for local launches, a nonempty
`CUDA_VISIBLE_DEVICES` mask on Linux. No GPU auto-discovery. Local GPU capacity and
model labels are configured, not hardware-verified; allocations do not reserve
devices exclusively against other host programs or allocations.
Configured catalogs replace it completely; missing explicit paths and invalid
files are errors. Execution still requires an explicitly launched cluster's name or ID.
Missing explicit paths and invalid files are errors. Execution still requires an
explicitly launched cluster's name or ID.

**A Slurm connection needs no launch settings (2026-09).** Every `launch` key
defaults, and the defaults assume a home directory shared by login and compute
Expand Down Expand Up @@ -1869,9 +1880,12 @@ ended. Every scheduler lc launches runs under `runtime.SCHEDULER_CONFIG`, whose
zero `events-cleanup-delay` drops a departed client's forwarded recipe output
instead of holding it for Dask's default hour.

**No login-node guard or venue module (PR #226 review).** Explicit catalog
selection and native backend permissions determine allocation. Do not infer
permission from hostnames, NERSC_HOST, or inherited Slurm job variables. Both
**Block local compute on recognized NERSC login nodes (2026-09).** A nonempty
`NERSC_HOST` plus a short hostname matching `login[0-9]+` disables local launch
and execution, including explicit local offers and `local.enabled: true`.
Interactive compute nodes remain eligible; inherited `SLURM_JOB_ID` never exempts
a login node. Apply this policy at runtime without writing a configuration file.
Keep inspection, termination, and Slurm execution available. Both execution
commands submit ordinary tasks to the Dask scheduler without worker restrictions;
existing task runtime gates and sandbox checks remain. Do not add a per-worker
validation framework around ordinary Dask task submission. Shared project storage
Expand Down Expand Up @@ -2112,6 +2126,33 @@ unlinks before writing; a new tampering test should too.

### Recorded decisions

- **Guard NERSC login nodes by default (2026-09).** This reverses the earlier
no-login-node-guard decision: an unconfigured first launch must not allocate a
whole shared login node. Detect the documented NERSC environment marker and
login hostname together, without DNS queries or scheduler probes. A runtime
restriction also covers copied or incomplete catalogs; no generated file or
override flag is needed. Local compute in interactive compute-node sessions
remains available, subject to the configured local policy.

- **Local compute accompanies remote catalogs (2026-09).** The built-in offer
uses the host's usable CPU/RAM capacity and follows configured offers, replacing
the previous one-CPU/1-GiB fallback that disappeared when a catalog existed.
`local.resources` sets a smaller budget; explicit local connections use their
own offers. Login-node catalogs disable local launch and execution through
`local.enabled: false`. Inspection and termination stay available. One local
allocation per user per machine is enforced across catalogs and connection
roots by scanning the process table for a live owner before launch
(`local._running_owners`: a session leader of this user running
`_OWNER_ARGS`, the command launch and `_process` share). This replaced an
`flock` held for the owner's lifetime, because NERSC home filesystems do
not support `flock`. Accepted residue: overlapping launches can both start,
and the scan covers one PID namespace, so a container sharing the home
does not see the host's owner. A missing identity record is transient
("retry shortly"); an unreadable one is not, so that refusal names the PID.
The suite scopes the scan to its own temporary tree
(`local_allocation_scope`), so a developer's running cluster does not refuse
test launches.

- **The engine is the host's uv tool, never a project dependency**
(2026-08, reversing spec §2's engine-in-lock rule and deleting layer 3).
`lc init` scaffolds no `lightcone-cli` dependency, and there is no
Expand Down
3 changes: 1 addition & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,8 +32,7 @@ cd my-analysis
uv add numpy
# When you are done with your edits, commit:
git add -A && git commit -m "First analysis"
CLUSTER=$(lc compute launch --cpus 1 --memory 1)
lc compute status "$CLUSTER" --wait
CLUSTER=$(lc compute launch --wait)
# Generate outputs with full provenance tracking
lc materialize "$CLUSTER"
lc compute down "$CLUSTER"
Expand Down
48 changes: 41 additions & 7 deletions docs/api/compute.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@ It owns no service, registry, or saved current-cluster selection.
| Symbol | Contract |
|---|---|
| `Request.parse(...)` | Exact/minimum CPU and memory requests, exact accelerator type/count, node count, walltime, startup class. |
| `Catalog.load(path)` | Ordered fixed shapes and stable connection namespaces; use the built-in local catalog only when the implicit default file is absent. |
| `Catalog.load(path)` | Ordered fixed shapes and stable connection namespaces; apply local defaults or disable policy alongside configured offers. |
| `Compute.plan(request, *, name=None)` | Select an eligible offer and freeze its native launch settings and optional name without allocation. |
| `Compute.launch(plan)` | Check names across native authorities, generate one if omitted, submit once, and return a self-contained `Identity`. |
| `Compute.discover()` | Snapshots and per-connection errors, querying each authority once. |
Expand All @@ -16,15 +16,27 @@ It owns no service, registry, or saved current-cluster selection.
| `connect(cluster_id, timeout=10)` | Resolve a name or full ID; borrow a standard Dask client, closing the client but never the allocation. |
| `Provider` | `plan`, `launch`, `discover`, `inspect`, `connect`, `terminate`. |

`Catalog.load()` defaults to `~/.lightcone/compute.yaml`. When that implicit file
is absent, the built-in catalog exposes a `local` offer: one CPU, 1 GiB, one node,
fast startup, 30-minute default and two-hour maximum lifetime. GPU offers require
an explicit catalog. It creates no configuration file or allocation. Configured
catalogs replace it completely.
`Catalog.load()` defaults to `~/.lightcone/compute.yaml`. The built-in `local`
offer uses detected usable CPUs and RAM, one node, fast startup, and a 30-minute
default/two-hour maximum lifetime. `local.resources` overrides its CPU/RAM budget;
`local.enabled: false` blocks local launch and execution while retaining connections
for inspection and termination. Remote catalogs retain the implicit local offer
unless disabled. Explicit local connections supply their own offers instead and
cannot be combined with `local.resources`. GPU offers require explicit configuration.
Loading creates no configuration file or allocation.
The effective local policy also disables local offers on recognized NERSC login
nodes: nonempty `NERSC_HOST` and a short hostname matching `login[0-9]+`.
Explicit enablement and Slurm job environment variables do not override this
guard; interactive compute nodes remain eligible. Local planning, launch, and
execution check the same policy, while status and termination remain available.
Missing paths selected through an argument or `LC_COMPUTE_CONFIG`, unreadable
files, and invalid catalogs remain errors. Stable connection namespaces let
separate invocations discover and attach to the same local allocations.

`Compute.plan_local()` selects only local offers and defaults the name to `local`.
CLI `launch --wait` waits through `Compute.status` using the accepted immutable ID;
errors retain that ID without resubmission or termination.

`model.py` defines the shared Pydantic models: `Connection`, `Offer`, `Resources`, `Accelerator`,
`TimeLimits`, `Startup`, `Request`, `Identity`, `LaunchPlan`, and `Snapshot`.
`Catalog` validates YAML directly into these objects, which providers also use.
Expand Down Expand Up @@ -136,7 +148,8 @@ See [Slurm's accounting field documentation](https://slurm.schedmd.com/sacct.htm
Execution submits ordinary tasks through the borrowed client's `submit` method.
Dask chooses the workers and handles dependencies; invocation-specific keys prevent
unintended reuse across commands. There is no worker-selection layer, per-worker
preflight orchestration, source fingerprinting, or login-node guard. Driver-side
preflight orchestration, or source fingerprinting. The local login-node guard
does not restrict remote Slurm execution from a login node. Driver-side
preparation and the existing task runtime/sandbox checks remain in their owners.

Workers advertise standard Dask `CPU`, `MEMORY`, and `GPU` resources; memory is measured
Expand Down Expand Up @@ -208,6 +221,27 @@ command. A driver that exits before every task reports says so with
`UNSTOPPED`: closing a client cannot prove that a remote subprocess stopped. Probes preserve both streams;
materialization sends recipe output to stderr to leave stdout for its report.

Local allocations are limited to one per user on each machine, independent of
connection roots and namespaces. Before spawning, the launcher scans the process
table for a live owner of the same user: a session leader running `-P -m
lightcone.engine.compute.local_runtime <directory>`, which excludes workers forked
from it. The process table spans every catalog and connection root and needs no
file lock, which some shared home filesystems, NERSC's included, do not support.
The owner's directory argument locates its identity record, so a refused launch
names the running cluster, its connection root, and the catalog it was launched
with. A record not yet written means the owner is still starting; one that is
missing or unreadable after that never recovers, so the refusal names the
owner's PID instead. Two limits are accepted rather than closed with a lock:
launches that overlap can both pass the scan, and the scan covers one PID
namespace, so a container sharing the home directory does not see the host's
owner.

A startup pipe lets the owner proceed only after the launcher publishes its
identity and launch records. If the launcher dies before completing publication,
the pipe closes and the owner exits. Failures before identity
publication remove the launcher's private files; a published identity remains
inspectable after a startup failure.

Local teardown drains the allocation's validated process group rather than
assuming the owner's exit proves every child stopped. Boot UUID, UID, process
session and the exact command containing a random allocation token establish
Expand Down
7 changes: 5 additions & 2 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -169,8 +169,11 @@ config-blob id, never a tag.
`engine.compute` owns allocation lifecycle through a small provider protocol.
A YAML catalog supplies ordered resource offers and stable native service
namespaces. When the implicit default file is absent, a built-in local catalog
provides one CPU and 1 GiB without setup. GPU offers need an explicit catalog,
which replaces the built-in defaults. Missing explicit paths and invalid files
provides all detected usable CPUs and RAM without setup. The `local` policy can
override that budget or disable local compute. Remote catalogs retain the default
local offer; explicit local connections supply their own offers. A runtime guard
disables local compute on recognized NERSC login nodes while permitting interactive
compute nodes. GPU offers need an explicit catalog. Missing explicit paths and invalid files
remain errors. No catalog is written and
no allocation starts until `compute launch` resolves resources and submits once. Slurm queries
and validated local OS identities are authoritative for allocations; Dask is the
Expand Down
Loading
Loading