Skip to content

Streamline local compute launch with readiness waiting and policy controls - #232

Merged
EiffL merged 4 commits into
mainfrom
feat/local-compute-defaults
Sep 29, 2026
Merged

EiffL merged 4 commits into
mainfrom
feat/local-compute-defaults

Conversation

@EiffL

@EiffL EiffL commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Starting local compute currently requires CPU and memory flags, a separate readiness command, and a custom catalog for anything larger than 1 CPU / 1 GiB. lc compute launch --wait now creates a cluster named local using the machine's detected usable CPUs and RAM and returns when it is ready for execution.

  • Add launch --wait and --timeout, retaining the accepted cluster ID on timeout or startup failure without resubmitting or terminating it.
  • Add local.resources to override the default CPU/RAM budget and local.enabled: false to block local launches and execution. Inspection and termination remain available.
  • Keep the built-in local offer alongside remote offers, after configured offers in selection order. Existing explicit local connections use their own offers instead; the no-resource shortcut only selects local compute.
  • Enforce one local cluster per user per machine with an OS lock inherited by the detached owner, across concurrent launchers, names, catalogs, and connection roots.
  • Update CLI, API, architecture, getting-started documentation, and the eval prompt, including a Slurm configuration that disables local compute on login nodes. The eval now uses lc compute launch --wait and explains cluster reuse, timeout handling, resource overrides, and disabled-local policy.

The default local lifetime remains 30 minutes (two-hour maximum), and GPU offers remain explicit. CPU/RAM budgets remain cooperative scheduling limits.

Validation:

  • Compute, local lifecycle, Slurm adapter, and output test suites passed, including real competing launchers and a no-resource launch --wait reopened from another process.
  • Final common compute suite: 177 passed, including additional configuration compatibility and capacity checks.
  • ruff check src/ tests/ and mypy src/ passed.
  • git diff --check passed. Slurm validation uses simulated native commands; no live Slurm allocation was submitted.

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

✅ Eval

Metric Value
Outputs check success
Agent run success
Turns 10
Tool calls 8
Cost $0.14
Agent wall time 0m47s
Model claude-sonnet-5-5
lc status
  mode:    direct
  sandbox: landlock (fs: declared, network: allowed)
  crate:   up to date with the outputs

  · current  baseline/best_fit        aadc309
  · current  baseline/hubble_diagram  aadc309
  · current  baseline/residuals       aadc309

3 current
Confusion & pain points (Claude analysis)

Confusion & pain points

  • The run was essentially clean: it made no failed commands, no source reverse-engineering and no retries. The friction below was small and mostly self-inflicted.
  • Guessing at data semantics. The agent spent three exploratory head/awk calls working out the Union2.1 file layout. It found that column 5 was not a systematic error and that no covariance matrix was shipped. It then invented a flat SYS_FLOOR = 0.1 mag for stat_and_sys. That is a scientifically arbitrary placeholder standing in for what the spec's decision implies. The eval task or data README should either ship the covariance matrix or document the columns, so agents don't fabricate systematics. The agent did disclose the workaround in its final summary.
  • Blind spec editing through string replacement. The agent patched astra.yaml with a Python str.replace script instead of reading the scaffold's output stanzas. It included a no-op replace, which suggests it was unsure of the exact text. It got away with it because the validator passed. There is no visible astra command that adds format:, inputs: or decisions: to an output. Such a command, or a clearer example in the skill, would remove the blind text munging.
  • Validation went through uvx astra-tools@0.2.18 rather than an installed astra. The agent pinned a version it apparently knew or guessed. Whether astra was already on PATH is unclear from the trace. If it wasn't, the eval environment should provide it so agents don't have to guess a version.
  • The license side effect was unprompted. The agent added license = "CC-BY-4.0" to pyproject.toml only to get ro-crate-metadata.json. It picked the license itself, which asserts terms over someone else's data. That follows from the design, because the crate is turned on only by a license. Even so, the eval prompt should say whether a crate is expected and which license to use, so the agent doesn't have to choose one.

Full trace: agent-trace artifact on this run.

@EiffL EiffL left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Automated review of 85b8d8d + 016e522 against main (15 findings; tests/test_compute.py passes). Four were reproduced by running the code, as marked; the rest come from reading it.

Most findings come from two changes:

1. The built-in local offer is now merged into configured catalogs. A Slurm-only catalog can silently fall back to a whole-host local cluster, a local discovery error can block Slurm launches, and a remote offer named local makes the catalog invalid. This also contradicts the layer-7 invariant in CLAUDE.md, the lc compute --help text and the getting-started guide. Deciding whether configured catalogs should get the local offer at all settles most of this group.

2. The per-user lock at /tmp/lightcone-local-<uid>. An error raised during launch is reported as a lock failure. An interrupted launch can leave an undiscoverable owner holding the lock. The refusal doesn't name the owning cluster. Another user can pre-create the directory, and /tmp cleaners can unlink a held lock. Tests can't redirect it, and a stdio-closed launcher can lose the lock to DEVNULL. Moving the lock under the connection root and writing the owner's cluster ID into it would address several of these at once.

Smaller items: leftover code for pre-lock allocations, the local-disabled check duplicated in four places, and a cleanup assertion dropped from the failed-spawn test.

🤖 Generated with Claude Code

resources = self.local.resources or Resources.from_bytes(
cpus=CPU_COUNT, memory_bytes=MEMORY_LIMIT,
)
offers.append(Offer(

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remote-only catalogs silently gain a whole-host local offer (reproduced)

A request that no remote offer satisfies now falls through to a local allocation instead of failing.

Scenario: an existing Slurm catalog with no local: block, on a login node. lc compute launch --cpus 2+ --memory 1+ --startup fast skips the Slurm offer (startup class unknown) and plans offer=local on connection=local with every detected CPU and all RAM (confirmed with --dry-run). The same happens when a Slurm offer raises UnavailableOfferError. The researcher gets a LocalCluster on the login node sized to the whole machine rather than a "no configured offer matches" error. Before this change, configured catalogs replaced the built-in offer completely.

"lock_fd": lock_fd,
},
)
return identity

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

An interrupted launch orphans an owner that holds the lock

The except Exception just below does not catch KeyboardInterrupt, and nothing catches SIGTERM/SIGKILL. If the launcher is interrupted after Popen but before identity.json/launch.json are published (for example, Ctrl-C or an agent timeout during the fsync'd writes), there is no killpg. The owner's loop (local_runtime.py:52) polls for an identity until its itimer fires, which is 30 min to 2 h, or longer for explicit offers.

Discovery skips the directory because it has no identity.json, so lc compute down has nothing to address, yet every lc compute launch fails with "already running or starting… stop it with lc compute down". Before the lock existed this orphan did no harm.

Comment thread src/lightcone/engine/compute/local.py Outdated
try:
fcntl.flock(descriptor, fcntl.LOCK_EX | fcntl.LOCK_NB)
except BlockingIOError as exc:
raise ComputeError(

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The refusal names no cluster, so the suggested remedy often can't be followed

The lock covers every catalog and connection root, but discovery only sees the current catalog.

Scenario: a cluster is launched with LC_COMPUTE_CONFIG pointing at an explicit local connection (custom connection_root or namespace), or with the built-in namespace before the user switches to their own local namespace. In a shell using the other catalog, lc compute launch says "a local cluster is already running… stop it with lc compute down", but lc compute status prints "No allocations found." and down cannot resolve it ("namespace is absent from the compute catalog"). The user is blocked until walltime. evals/prompt.md tells agents to "inspect lc compute status and reuse it", which cannot work here.

Writing the owner's cluster ID into the lock file would let the refusal name it.

Comment thread src/lightcone/engine/compute/local.py Outdated
"stop it with lc compute down before launching another"
) from exc
yield descriptor
except OSError as exc:

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

except OSError wraps the yield, so launch-body errors are reported as lock failures (reproduced)

When writing launch.json fails with EDQUOT (home quotas are common on HPC), the error reads cannot acquire the local allocation lock: [Errno 122] Disk quota exceeded, and the token directory under the connection root is left behind. A PermissionError from discover()'s iterdir is mislabeled the same way.

Only the os.open/fstat/flock calls should be inside the OSError handler.

"connection name 'local' is reserved for the built-in local backend"
)
# Retain this authority even when disabled so existing allocations can be stopped.
connections["local"] = Connection(namespace=_LOCAL_NAMESPACE, provider="local")

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The implicit local connection is added even with local.enabled: false, so a local discovery error now blocks Slurm

Scenario: a Slurm-only catalog on a machine where ~/.lightcone/compute/22c84e48-… has a permission problem (for example, restored from backup, so private_directory refuses), or where a local process identity cannot be verified. Compute.discover() records an error under local. Compute.launch then refuses every Slurm launch ("cannot check cluster names while discovery is incomplete: local: …"), and lc materialize my-slurm-cluster cannot resolve the name. Remote-only catalogs could not hit this before.

Comment thread tests/test_compute_local.py Outdated
patch.setattr("lightcone.engine.compute.local.subprocess.Popen", fail)
with pytest.raises(ComputeError, match="cannot execute"):
_launch(provider)
# A failed spawn must release the singleton lock as well as its private files.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The failed-spawn test no longer asserts that the launcher's private files are removed

The previous assert set(provider.root.iterdir()) == before was dropped when the test was reordered, although the comment still claims the behavior. A regression that leaves the token directory, launch.json or scratch behind on a failed spawn (the case the except OSError comment above already hits) now passes. Only the lock release is checked, indirectly, by the next successful launch.

Comment thread src/lightcone/engine/compute/local.py Outdated
if plan.name is not None:
validate_name(plan.name)
with _allocation_lock() as lock_fd:
# Also recognize allocations launched before the lifetime lock existed.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

CLAUDE.md: backward-compatibility and dead code for pre-lock allocations

CLAUDE.md: "No backward-compatibility code. Nothing exists to honor the behavior of an older CLI…" and "No dead code."

This comment and the if self.discover(): below run an extra discovery on every launch whose only purpose is allocations launched before the lock existed. Likewise local_runtime.py:28 if "lock_fd" in launch: is always true, since both launch.json writes now include lock_fd.

"local.resources cannot be combined with explicit local connections; "
"set their offer resources instead"
)
if not explicit:

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

CLAUDE.md: a recorded layer-7 invariant is reversed without being updated

CLAUDE.md still says the built-in catalog is "one CPU, 1 GiB" and "Configured catalogs replace it completely", and the repository map says catalog.py # compute.yaml, or the built-in local offer. _with_local now sizes the offer to the whole host and merges it into configured catalogs. CLAUDE.md says reversals "land as Recorded decisions here", so as it stands the next contributor reads the opposite of the current behavior.

Comment thread docs/user/getting-started.md Outdated
Launch the built-in local offer; no compute configuration is needed. It provides
one CPU and 1 GiB for 30 minutes. Keep the returned ID in `CLUSTER` for this
all usable CPUs and RAM for 30 minutes. Keep the returned ID in `CLUSTER` for this
walkthrough. If you already have a compute catalog, its offers replace that default;

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This line and the lc compute --help text still say a configured catalog replaces the local offer

This sentence, and the compute group docstring in src/lightcone/cli/compute.py:62 ("The catalog is LC_COMPUTE_CONFIG, else ~/.lightcone/compute.yaml, else a built-in local offer"), are now false: the local offer is appended to configured catalogs unless explicit local connections exist or local.enabled: false. The docs rule is to document only what exists.

) -> LaunchPlan | None:
"""Match one shape and validate its provider without allocating anything."""
connection = self.catalog.connections[offer.connection]
if connection.provider == "local" and not self.catalog.local.enabled:

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The local-disabled policy is special-cased in four places, one of them dead

provider == "local" and not catalog.local.enabled appears in _plan_offer, plan_local, launch and connect, with the message string copied three times. The _plan_offer branch can never fire because Catalog._with_local already drops disabled local offers. A new execution entry point that skips compute.connect (for example, Compute.status, which calls provider.connect directly) silently escapes the policy. Enforcing it once, in Compute.provider or as a connection-level flag set by _with_local, would remove the copies.

EiffL and others added 2 commits September 29, 2026 09:44
Make detached local startup transactional with a parent-child pipe, clean unpublished failures, and preserve published identities. Protect the per-user lock, retain owner diagnostics, and handle closed stdio and the brief kernel teardown window after expiry.

Disable local launch and execution on recognized NERSC login nodes while permitting interactive compute nodes and retaining Slurm access, inspection, and termination. Clarify catalog errors, isolate lifecycle tests, and update the documentation.

Validation: 249 compute, local lifecycle, and output tests passed; Ruff, mypy, and git diff --check passed. Earlier Slurm and CLI validation passed with one bootstrap test passing on retry. No live NERSC allocation was tested; the documented home-filesystem flock limitation remains.
The one-local-cluster-per-user rule was an flock held in the account home
for the owner's lifetime, and NERSC home filesystems do not support flock.
A launch now scans the process table for a session leader of this user
running the owner command, which spans every catalog and connection root
with no file lock. The owner command is one shared constant for launch,
the scan and the identity check.

A refusal still names the running cluster and the catalog to stop it
with, now recorded in its identity file. A record not yet written means
the owner is starting ("retry shortly"); a missing or unreadable one
never recovers, so that refusal names the owner's PID. An unreadable
process table is a clean ComputeError.

Accepted: overlapping launches can both start, and the scan covers one
PID namespace. The suite scopes the scan to its own temporary tree, so a
developer's running cluster does not refuse test launches.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@EiffL
EiffL merged commit 8e35228 into main Sep 29, 2026
8 of 9 checks passed
@EiffL
EiffL deleted the feat/local-compute-defaults branch September 29, 2026 17:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant