Skip to content

Supervise Slurm workers with Dask nannies - #231

Merged
EiffL merged 1 commit into
mainfrom
fix/slurm-worker-supervision
Sep 29, 2026
Merged

EiffL merged 1 commit into
mainfrom
fix/slurm-worker-supervision

Conversation

@EiffL

@EiffL EiffL commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Slurm workers currently run inside their rank processes, so a worker crash can take down rank zero's scheduler and Slurm can terminate the remaining ranks. Run each worker under a stock Dask Nanny to restart it independently, and use --kill-on-bad-exit=0 --wait=0 to prevent rank exits from terminating healthy ranks.

Nannies also supply Dask's default OpenMP, MKL, and OpenBLAS thread counts of one. Forward the effective values into container recipes and probes, preserving explicit launch overrides. Keep the existing CPU, memory, and GPU resource reservations, including after worker restarts.

This does not add per-recipe memory enforcement or subprocess fencing: site OOM policy can still end the allocation, and a surviving recipe subprocess can overlap Dask's retry. Serial per-output Git/annex commits remain unchanged. A Perlmutter deployment run is still needed.

Validation:

  • uv run pytest -q — 1,099 passed after rebasing onto current main.
  • uv run ruff check src/ tests/ — passed.
  • uv run mypy src/ — passed.
  • Real Dask/TLS subprocess tests kill workers on both ranks and verify scheduler/peer survival, replacement capacity, resource budgets, thread defaults, and shutdown. Container tests cover thread forwarding for recipes and probes.

@github-actions

Copy link
Copy Markdown
Contributor

✅ Eval

Metric Value
Outputs check success
Agent run success
Turns 8
Tool calls 6
Cost $0.13
Agent wall time 0m41s
Model claude-sonnet-5-5
lc status
  mode:    direct
  sandbox: landlock (fs: declared, network: allowed)
  crate:   up to date with the outputs

  · current  baseline/best_fit        419dcb4
  · current  baseline/hubble_diagram  419dcb4
  · current  baseline/residuals       419dcb4

3 current
Confusion & pain points (Claude analysis)

Confusion & pain points

  • Untested execution paths. The agent ran only the baseline universe (Nelder-Mead, stat+sys, z>0.01) and never exercised L-BFGS-B or the other decision options. Its own summary admits this, and lc materialize --check cannot catch it. bounds=[(50,100),(0,1)] and the stat_only/none branches shipped untested. Fix: have the eval prompt or harness require at least one non-baseline universe, so the multiverse path is actually verified.

  • Scaffold command failed on a missing directory. cat > scripts/common.py 2>/dev/null || mkdir scripts errored with "No such file or directory". The agent guessed the layout instead of checking it. The spec-created scripts/ directory was absent, and the recipe stubs in astra.yaml gave no hint where scripts belong. The agent recovered by re-running the command, so the cost was small. Fix: scaffold or docs could state the expected code layout. CLAUDE.md deliberately declines to fix one, so this is a documentation gap rather than a product bug.

  • Recipe wiring done by blind string replacement. The agent used s.replace() on astra.yaml to inject format, inputs, decisions and commands, keyed to exact whitespace of the stub. It passed validation, but it was fragile, and the agent never viewed the edited file. The eval's real friction is that format: is required by lc but only recommended by ASTRA. plan.build refuses a spec without it. Fix: astra validate or the skill should flag missing format up front, since lc's hard requirement is otherwise discovered only at materialize time.

  • Undocumented settings edited by hand. The agent inserted license = "CC-BY-4.0" into pyproject.toml with a sed one-liner to enable crate maintenance. It also pinned uvx astra-tools@0.2.18 to validate, so it either knew or guessed the astra CLI version and license opt-in. Nothing in the trace shows it reading docs for either. Fix: the skill or the lc init output should mention that [project].license turns on the RO-Crate.

  • Compute lifecycle is manual ceremony. The agent had to launch, status --wait and later down a named local cluster, with --cpus 1 --memory 1 --name analysis chosen without visible discussion of resources. Nothing failed, but the required explicit allocation for a trivial local run is boilerplate. Fix: a documented single-line local recipe in the skill, so agents don't have to work out the sequence.

Full trace: agent-trace artifact on this run.

@EiffL EiffL left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@EiffL
EiffL merged commit eb79ef4 into main Sep 29, 2026
10 of 11 checks passed
@EiffL
EiffL deleted the fix/slurm-worker-supervision branch September 29, 2026 13:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant