Repository navigation
End local clusters when idle instead of at a fixed age - #234
Conversation
Local clusters expired 30 minutes after launch (two-hour maximum) even
while a recipe was running. The built-in local offer now has no walltime
and a 30-minute idle timeout, handed to the scheduler as Dask's own
idle_timeout: running or queued tasks keep it alive and new work restarts
the countdown, while connected clients and status polling do not.
- An offer's time is {default?, max?, idle?} and needs a default or an
idle; local.time configures the built-in offer like local.resources.
- --time stays an explicit hard walltime that ends the allocation even
during active work; with both limits, the first to fire wins.
- The owner hears the scheduler close through a SchedulerPlugin and ends
its session like the walltime path, so the machine's one local
allocation and its name are free at once (LocalCluster's own close
waits ~34 s on the departed scheduler).
- Slurm refuses time.idle and still needs time.default: it ends at the
native walltime.
- lc compute resources gains an IDLE column; plans report idle_seconds.
Closes #233
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
✅ Eval
lc statusConfusion & pain points (Claude analysis)Confusion & pain points
Full trace: |
- compute.connect submits one no-op task, so lc run / lc materialize preparation (annex fetch, image build, sync) starts with a full idle countdown; status polling still does not restart it. - The scheduler close hook records the reason and kills the session itself unless the owner asked for the close, so a scheduler that idles out before LocalCluster() returns can no longer strand the owner. - SCHEDULER_CONFIG pins idle-timeout to None: ambient Dask config reaches neither local nor Slurm schedulers. - The walltime timer is armed inside the try that records errors. - LaunchPlan.idle_seconds derives from the offer; clearer --time refusal, help text, and docstrings; wider test margins for slow CI. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Pushed fixes for the review findings. Each is a few lines using mechanisms that already existed; there's no new component.
Each fix was mutation-checked: removing it fails a test. Still not handled, and recorded in CLAUDE.md: preparation longer than the idle timeout, and an owner whose scheduler loop hangs when no |
Closes #233.
Before this change, local clusters expired 30 minutes after launch, with a two-hour maximum, even while a recipe was running. Now the built-in local offer has no walltime and a 30-minute idle timeout. A two-hour recipe finishes normally, and the cluster stops 30 minutes after the last task.
What changed
timeis now{default?, max?, idle?}. It must set adefaultor anidle, so every allocation can still end.{idle: 30m}.local.timereplaces it. It works likelocal.resources: it can't be combined with explicit local connections, whose offers carry their owntime.idleis passed to the scheduler as itsidle_timeout, so Dask'scheck_idledecides what counts as activity; there is no tracker of ours.scheduler_infocalls andclient.rundo not. Measured before relying on it.--timeis still an explicit hard walltime (SIGALRM kills the whole session) and ends the allocation even during active work. When both limits are set, whichever fires first ends it. The built-in offer has nomax, because capping--timemakes no sense when leaving it out means no limit at all.SchedulerPlugin.closehook, then SIGKILLs its session the same way the walltime path does. It deliberately skipsLocalCluster's own close, which I measured waiting about 34 s on the departed scheduler. During that wait the machine's one-local-cluster slot and the name would both stay taken.time.idleand still needstime.default, because it ends at the native walltime. An ignored idle timeout would misreport how the job ends.lc compute resourcesgains anIDLEcolumn. Cell padding drops from 2 to 1, because the table already truncated headers at 80 columns (DEFAU…,START…) before this change. Plans reportidle_secondsbesidetime_seconds, and the--timehelp says "hard walltime".lc compute resources, and the quoted refusal is from a real run.Tests
The new integration tests run a real detached owner and scheduler, with intervals of a few seconds:
Compute().statuspolls it. Discovery then comes back empty, and the same name relaunches.Each was mutation-checked and fails when its mechanism is removed:
idle_timeoutpassed to the scheduler;Unit tests cover the catalog (
local.time, a time with no way to end, refusal beside explicit local connections) and the Slurm refusal.As an end-to-end check, I ran
lc materializeon a real local cluster with a 5 s idle timeout. A 12 s recipe committed normally, the cluster ended about 5 s aftermaterializereturned, andlc compute launch --namereused the name straight away.uv run pytest: 1148 passed.ruffandmypy --strictare clean.Not from this change:
test_standard_bootstrap_starts_scheduler_and_worker_on_rank_zero_and_worker_on_rank_one[False]fails intermittently onmaintoo, withFutureCancelledError: scheduler-connection-lostin 1 of 5 runs here. It was deselected for the full run.Accepted limitations (recorded in CLAUDE.md)
🤖 Generated with Claude Code