You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Local clusters currently expire after a fixed lifetime measured from launch (30 minutes by default), even when a recipe is still running. For interactive local use, the default should instead be 30 minutes without compute activity.
Proposed behavior
Default local clusters to a 30-minute idle timeout, configurable through the compute configuration.
Running or queued Dask tasks prevent idle expiry. New task activity resets the idle timer; merely connecting a client or checking status does not keep the cluster alive.
Keep lc compute launch --time DURATION as an optional explicit hard lifetime for local clusters. Without an explicit hard limit, active work can continue beyond the current 30-minute default and two-hour built-in maximum. If both limits are configured, either can end the allocation; an explicit hard limit can interrupt active work.
Add a local idle-timeout configuration setting and pass it to the Dask scheduler.
Separate idle timeout from hard walltime throughout planning, launch metadata, and the local runtime. The current mandatory deadline and SIGALRM would otherwise still stop active work after 30 minutes.
When Dask shuts down, make the detached allocation owner exit, clean up its workers, and release the singleton lock. Currently the owner waits independently of scheduler shutdown.
Ensure status/discovery reports the allocation as ended and a new local cluster can launch after idle expiry.
Update CLI help, user/API documentation, and the eval prompt to distinguish idle timeout from hard lifetime.
Acceptance criteria
An unused local cluster expires after the configured idle interval, including when no task has ever been submitted.
A task running longer than the idle interval completes without being interrupted by idle expiry; queued work also prevents idle expiry.
New work before expiry resets the timer; client connections and status polling alone do not.
Idle shutdown stops the managed allocation, releases the singleton lock, and permits a replacement cluster with the same name.
An explicit --time still enforces a hard deadline, including during active work, and compute down still stops the cluster.
Slurm continues to use native allocation walltime.
Use short configurable intervals in integration tests to cover real scheduler shutdown and detached-owner cleanup.
Local clusters currently expire after a fixed lifetime measured from launch (30 minutes by default), even when a recipe is still running. For interactive local use, the default should instead be 30 minutes without compute activity.
Proposed behavior
lc compute launch --time DURATIONas an optional explicit hard lifetime for local clusters. Without an explicit hard limit, active work can continue beyond the current 30-minute default and two-hour built-in maximum. If both limits are configured, either can end the allocation; an explicit hard limit can interrupt active work.lc compute downfor explicit shutdown and the one-local-cluster-per-user-per-machine limit introduced in Streamline local compute launch with readiness waiting and policy controls #232.For example, a recipe running for two hours should finish normally, and the cluster should shut down 30 minutes afterward if no further work arrives.
Implementation
Dask already supports
distributed.scheduler.idle-timeout. Reuse its task-activity detection rather than introducing a separate activity tracker. See the scheduler implementation.Acceptance criteria
--timestill enforces a hard deadline, including during active work, andcompute downstill stops the cluster.Use short configurable intervals in integration tests to cover real scheduler shutdown and detached-owner cleanup.