Skip to content

End local clusters when idle instead of at a fixed age - #234

Merged
EiffL merged 2 commits into
mainfrom
local-idle-timeout
Sep 29, 2026
Merged

EiffL merged 2 commits into
mainfrom
local-idle-timeout

Conversation

@EiffL

@EiffL EiffL commented Sep 29, 2026

Copy link
Copy Markdown
Member

Closes #233.

Before this change, local clusters expired 30 minutes after launch, with a two-hour maximum, even while a recipe was running. Now the built-in local offer has no walltime and a 30-minute idle timeout. A two-hour recipe finishes normally, and the cluster stops 30 minutes after the last task.

What changed

  • Time limits: an offer's time is now {default?, max?, idle?}. It must set a default or an idle, so every allocation can still end.
    • The built-in local offer is {idle: 30m}.
    • local.time replaces it. It works like local.resources: it can't be combined with explicit local connections, whose offers carry their own time.
  • Uses Dask's own idle timeout. idle is passed to the scheduler as its idle_timeout, so Dask's check_idle decides what counts as activity; there is no tracker of ours.
    • Running, queued or unrunnable tasks keep the cluster alive, and any task transition restarts the countdown.
    • Connected clients, scheduler_info calls and client.run do not. Measured before relying on it.
  • Hard limit: --time is still an explicit hard walltime (SIGALRM kills the whole session) and ends the allocation even during active work. When both limits are set, whichever fires first ends it. The built-in offer has no max, because capping --time makes no sense when leaving it out means no limit at all.
  • Owner exit: the detached owner learns that the scheduler closed through a SchedulerPlugin.close hook, then SIGKILLs its session the same way the walltime path does. It deliberately skips LocalCluster's own close, which I measured waiting about 34 s on the departed scheduler. During that wait the machine's one-local-cluster slot and the name would both stay taken.
  • Slurm is unchanged: it refuses time.idle and still needs time.default, because it ends at the native walltime. An ignored idle timeout would misreport how the job ends.
  • CLI: lc compute resources gains an IDLE column. Cell padding drops from 2 to 1, because the table already truncated headers at 80 columns (DEFAU…, START…) before this change. Plans report idle_seconds beside time_seconds, and the --time help says "hard walltime".
  • Docs: CLI, user guide, API page, getting-started, the eval prompt, and a Recorded decision in CLAUDE.md are updated. Every changed catalog example was loaded through lc compute resources, and the quoted refusal is from a real run.

Tests

The new integration tests run a real detached owner and scheduler, with intervals of a few seconds:

  • An unused cluster idles out even while Compute().status polls it. Discovery then comes back empty, and the same name relaunches.
  • Two queued 3 s tasks on a cluster with a 4 s idle timeout both finish. New work submitted before expiry restarts the countdown.
  • An explicit walltime stops a running task even though the idle timeout is a minute.

Each was mutation-checked and fails when its mechanism is removed:

  • no idle_timeout passed to the scheduler;
  • the owner ignoring the scheduler's close;
  • the idle timeout treated as a hard deadline;
  • SIGALRM never set.

Unit tests cover the catalog (local.time, a time with no way to end, refusal beside explicit local connections) and the Slurm refusal.

As an end-to-end check, I ran lc materialize on a real local cluster with a 5 s idle timeout. A 12 s recipe committed normally, the cluster ended about 5 s after materialize returned, and lc compute launch --name reused the name straight away.

uv run pytest: 1148 passed. ruff and mypy --strict are clean.

Not from this change: test_standard_bootstrap_starts_scheduler_and_worker_on_rank_zero_and_worker_on_rank_one[False] fails intermittently on main too, with FutureCancelledError: scheduler-connection-lost in 1 of 5 runs here. It was deselected for the full run.

Accepted limitations (recorded in CLAUDE.md)

  • The idle clock starts with the scheduler. A worker that takes longer to start than the idle timeout never receives work. This doesn't matter at 30 minutes.
  • If the driver pauses between tasks, for example during a long annex commit, that time counts as idle.

🤖 Generated with Claude Code

Local clusters expired 30 minutes after launch (two-hour maximum) even
while a recipe was running. The built-in local offer now has no walltime
and a 30-minute idle timeout, handed to the scheduler as Dask's own
idle_timeout: running or queued tasks keep it alive and new work restarts
the countdown, while connected clients and status polling do not.

- An offer's time is {default?, max?, idle?} and needs a default or an
  idle; local.time configures the built-in offer like local.resources.
- --time stays an explicit hard walltime that ends the allocation even
  during active work; with both limits, the first to fire wins.
- The owner hears the scheduler close through a SchedulerPlugin and ends
  its session like the walltime path, so the machine's one local
  allocation and its name are free at once (LocalCluster's own close
  waits ~34 s on the departed scheduler).
- Slurm refuses time.idle and still needs time.default: it ends at the
  native walltime.
- lc compute resources gains an IDLE column; plans report idle_seconds.

Closes #233

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

✅ Eval

Metric Value
Outputs check success
Agent run success
Turns 9
Tool calls 7
Cost $0.14
Agent wall time 0m42s
Model claude-sonnet-5-5
lc status
  mode:    direct
  sandbox: landlock (fs: declared, network: allowed)
  crate:   up to date with the outputs

  · current  baseline/best_fit        a962714
  · current  baseline/hubble_diagram  a962714
  · current  baseline/residuals       a962714

3 current
Confusion & pain points (Claude analysis)

Confusion & pain points

  • The run was essentially clean. There were no errored tool calls. The agent went from spec to materialized outputs in 7 calls and about 42 s. The agent's remaining friction was small.

  • The agent had to guess at the astra.yaml shape. It filled in format, inputs, decisions and recipe.command for each output with a blind Python str.replace script, keyed on exact whitespace of the scaffold text. It could not check the edits until astra validate ran at the very end. That validation passed, so the guess was right. The rewrite was fragile, though. If the scaffold's indentation or wording had differed, the replacement would have silently done nothing. A per-field CLI edit command, or a scaffold with the format: / inputs: / decisions: stubs already present, would remove the need for this.

  • [project].license is set by a sed on pyproject.toml, with no guidance. The agent added license = "CC-BY-4.0" so the RO-Crate would be produced. Nothing in the task or the visible tooling says that a license is what turns crate maintenance on. The agent apparently learned this from the skill or from prior knowledge. The sed also relies on version = "0.0.1" being the exact line to anchor on. The eval prompt, or lc init's output, should say that the license key is what enables publication.

  • astra is called through uvx astra-tools@0.2.18 with a pinned version. The agent did not use an astra binary. That suggests astra is not on PATH in this environment, or the skill tells the agent to use uvx. Either way it is a network-dependent, version-pinned detour. If the harness expects astra validate, it should provide the binary.

  • The agent chose an unusual error model without a spec definition. It combined the statistical and systematic columns in quadrature, and its summary admits this probably overstates the errors (reduced χ² of 0.37). The spec and data README did not clearly define what stat_and_sys means for Union2.1's separate covariance matrix. This is a spec-clarity gap rather than a tool bug. The task or astra.yaml should say how systematics are to be applied.

  • The cluster lifecycle needed extra steps. The agent had to run lc compute launch --wait, then lc materialize local, then lc compute down local. Forgetting down leaves an allocation running until the idle timeout. The eval prompt could mention this sequence up front.

Full trace: agent-trace artifact on this run.

- compute.connect submits one no-op task, so lc run / lc materialize
  preparation (annex fetch, image build, sync) starts with a full idle
  countdown; status polling still does not restart it.
- The scheduler close hook records the reason and kills the session
  itself unless the owner asked for the close, so a scheduler that idles
  out before LocalCluster() returns can no longer strand the owner.
- SCHEDULER_CONFIG pins idle-timeout to None: ambient Dask config reaches
  neither local nor Slurm schedulers.
- The walltime timer is armed inside the try that records errors.
- LaunchPlan.idle_seconds derives from the offer; clearer --time refusal,
  help text, and docstrings; wider test margins for slow CI.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@EiffL

EiffL commented Sep 29, 2026

Copy link
Copy Markdown
Member Author

Pushed fixes for the review findings. Each is a few lines using mechanisms that already existed; there's no new component.

  • Driver setup no longer eats the idle window. compute.connect, which only lc run and lc materialize use, submits one no-op task. That restarts Dask's idle countdown, so fetching inputs, building the image and syncing start with the full timeout. status connects through the provider, so polling it still doesn't keep a cluster alive. A test pins both.
  • The owner can't be left running forever. When the scheduler closes and the owner didn't ask for it, the SchedulerPlugin.close hook now writes the reason to error.json and kills the session itself. That covers a scheduler that idles out before LocalCluster(...) returns, and lc compute status now shows why the cluster ended.
  • Ambient Dask config no longer leaks in. SCHEDULER_CONFIG sets distributed.scheduler.idle-timeout to None, so a Dask setting from the environment or ~/.config/dask reaches neither local nor Slurm schedulers.
  • Smaller fixes:
    • The walltime timer is now set inside the try that records errors, so an overflowing --time is recorded instead of failing silently.
    • LaunchPlan.idle_seconds is derived from the offer instead of stored.
    • The --time-over-max refusal names the time limits.
    • The help text says "by default".
    • The close hook has a docstring.
    • The corrected docs note that an interrupted driver can leave a recipe running, which the idle timeout then stops.
    • Test idle timeouts are 8 s and the test walltime is 10 s, so slow CI workers start in time.

Each fix was mutation-checked: removing it fails a test. uv run pytest: 1148 passed, with the pre-existing flaky Slurm bootstrap test deselected.

Still not handled, and recorded in CLAUDE.md: preparation longer than the idle timeout, and an owner whose scheduler loop hangs when no --time is set. Fixing either would need a new timer or component.

@EiffL EiffL left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@EiffL
EiffL merged commit cd7c539 into main Sep 29, 2026
10 of 11 checks passed
@EiffL
EiffL deleted the local-idle-timeout branch September 29, 2026 20:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Use idle timeout instead of default hard lifetime for local compute

1 participant