You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Slurm workers currently run inside their rank processes, so a worker crash can take down rank zero's scheduler and Slurm can terminate the remaining ranks. Run each worker under a stock Dask Nanny to restart it independently, and use --kill-on-bad-exit=0 --wait=0 to prevent rank exits from terminating healthy ranks.
Nannies also supply Dask's default OpenMP, MKL, and OpenBLAS thread counts of one. Forward the effective values into container recipes and probes, preserving explicit launch overrides. Keep the existing CPU, memory, and GPU resource reservations, including after worker restarts.
This does not add per-recipe memory enforcement or subprocess fencing: site OOM policy can still end the allocation, and a surviving recipe subprocess can overlap Dask's retry. Serial per-output Git/annex commits remain unchanged. A Perlmutter deployment run is still needed.
Validation:
uv run pytest -q — 1,099 passed after rebasing onto current main.
uv run ruff check src/ tests/ — passed.
uv run mypy src/ — passed.
Real Dask/TLS subprocess tests kill workers on both ranks and verify scheduler/peer survival, replacement capacity, resource budgets, thread defaults, and shutdown. Container tests cover thread forwarding for recipes and probes.
mode: direct
sandbox: landlock (fs: declared, network: allowed)
crate: up to date with the outputs
· current baseline/best_fit 419dcb4
· current baseline/hubble_diagram 419dcb4
· current baseline/residuals 419dcb4
3 current
Confusion & pain points (Claude analysis)
Confusion & pain points
Untested execution paths. The agent ran only the baseline universe (Nelder-Mead, stat+sys, z>0.01) and never exercised L-BFGS-B or the other decision options. Its own summary admits this, and lc materialize --check cannot catch it. bounds=[(50,100),(0,1)] and the stat_only/none branches shipped untested. Fix: have the eval prompt or harness require at least one non-baseline universe, so the multiverse path is actually verified.
Scaffold command failed on a missing directory.cat > scripts/common.py 2>/dev/null || mkdir scripts errored with "No such file or directory". The agent guessed the layout instead of checking it. The spec-created scripts/ directory was absent, and the recipe stubs in astra.yaml gave no hint where scripts belong. The agent recovered by re-running the command, so the cost was small. Fix: scaffold or docs could state the expected code layout. CLAUDE.md deliberately declines to fix one, so this is a documentation gap rather than a product bug.
Recipe wiring done by blind string replacement. The agent used s.replace() on astra.yaml to inject format, inputs, decisions and commands, keyed to exact whitespace of the stub. It passed validation, but it was fragile, and the agent never viewed the edited file. The eval's real friction is that format: is required by lc but only recommended by ASTRA. plan.build refuses a spec without it. Fix: astra validate or the skill should flag missing format up front, since lc's hard requirement is otherwise discovered only at materialize time.
Undocumented settings edited by hand. The agent inserted license = "CC-BY-4.0" into pyproject.toml with a sed one-liner to enable crate maintenance. It also pinned uvx astra-tools@0.2.18 to validate, so it either knew or guessed the astra CLI version and license opt-in. Nothing in the trace shows it reading docs for either. Fix: the skill or the lc init output should mention that [project].license turns on the RO-Crate.
Compute lifecycle is manual ceremony. The agent had to launch, status --wait and later down a named local cluster, with --cpus 1 --memory 1 --name analysis chosen without visible discussion of resources. Nothing failed, but the required explicit allocation for a trivial local run is boilerplate. Fix: a documented single-line local recipe in the skill, so agents don't have to work out the sequence.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Slurm workers currently run inside their rank processes, so a worker crash can take down rank zero's scheduler and Slurm can terminate the remaining ranks. Run each worker under a stock Dask Nanny to restart it independently, and use
--kill-on-bad-exit=0 --wait=0to prevent rank exits from terminating healthy ranks.Nannies also supply Dask's default OpenMP, MKL, and OpenBLAS thread counts of one. Forward the effective values into container recipes and probes, preserving explicit launch overrides. Keep the existing CPU, memory, and GPU resource reservations, including after worker restarts.
This does not add per-recipe memory enforcement or subprocess fencing: site OOM policy can still end the allocation, and a surviving recipe subprocess can overlap Dask's retry. Serial per-output Git/annex commits remain unchanged. A Perlmutter deployment run is still needed.
Validation:
uv run pytest -q— 1,099 passed after rebasing onto currentmain.uv run ruff check src/ tests/— passed.uv run mypy src/— passed.