From 869ae905779c4c366a45759be8676406a1aae85c Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Micka=C3=ABl=20R=C3=A9mond?= Date: Fri, 11 Sep 2026 08:55:54 +0200 Subject: [PATCH 01/31] fix(bin): rebalance portable parallel test lanes using CI timings (#4151) * ci: rebalance the portable parallel lanes on measured runner durations Both portable parallel lanes are capped at 10 minutes. Lane 1 was cancelled at that cap on every request raised on 2026-09-10 while lane 2 finished in about 3.5 minutes, so no request could go green. CONTRACT CLASS: RESTORE. The workflow already promises two duration-balanced lanes and the shard documentation already claims a measured wall; this re-establishes both against what the lanes now cost, and changes no lane count, no cap, and no scope of what runs. The counter-argument, so nobody has to take that on trust: two pieces here are genuinely new rather than restored, and either could be argued to make this a NEW-behavior change. `--list-scheduled` now ranks a parallel lane on measured durations where it previously handed every parallel script the serial default weight and returned an alphabetical order; and `--check-coverage` gains three reported fields. I classify the change RESTORE because both exist only to make the already-promised property checkable, but they are named here rather than folded into the restoration. === PART 1: THE TOTAL, AND HOW IT WAS OBTAINED === This section stands on its own. It establishes what the parallel set costs. It derives no packing; Part 2 does that, from this number. THE TOTAL: 828568 ms, about 13 min 49 s of serial work across the 24 scripts. Lane 1 held 624299 ms of it and lane 2 held 204269 ms, a 3.06:1 split. HOW IT WAS OBTAINED. The difficulty was that lane 1 had never finished, so its duration did not exist as a recorded figure anywhere and no timing artifact was expected for it. It turned out to be recoverable from the real lane without estimating, by two routes, across six CI runs on 2026-09-10 (34459949083, 34460760299, 34462530836, 34462758357, 34466966385, 34470382458): - Run 34462758357's lane-1 job finished its suite 18 s BEFORE the wall and uploaded a complete fm-test-timing-portable-parallel-1 artifact carrying all 11 scripts, FM_TEST_SUMMARY total=11 failed=0 duration_ms=598225. The upload step is if: always(), so the cancellation did not suppress it. This is one full, untruncated lane-1 measurement. - The five other lane-1 jobs were cancelled mid-suite, but each logs every script that had already finished as an FM_TEST_END duration_ms= marker. Those per-script records are complete measurements of completed scripts; only the script in flight at cancellation is lost, and it differs by run. Lane 2 completed in all six runs, so its scripts come from the six uploaded fm-test-timing-portable-parallel-2 artifacts. Every one of the 24 scripts therefore carries at least one untruncated measurement: 20 of them measured in all six runs, two in three or four runs, and two (fm-brief, fm-transition-lib, the tail of lane 1) in the single complete run. Each hint is the SLOWEST value that script reached, so the total is an upper envelope rather than an average. NO FIGURE IN IT IS DERIVED FROM A TRUNCATED LANE, and no lower bound was ever extrapolated into a total. THE ENVIRONMENT, AND WHETHER IT TRANSFERS. Every hint is a serial run of the real portable parallel lane on a GitHub ubuntu-latest runner, produced by the lane's own CI job. It transfers because it is not a proxy for the lane; it is the lane. Nothing in the total came from this machine or from any harness of mine. That mattered, and here is what it would have cost. A same-day macOS cross-check of the same scripts ran 1.7x to 5.0x slower with the ratio varying per script (fm-test-run 157420 ms against 92944 ms, fm-x-mode 67217 ms against 31870 ms, fm-composer-ghost 10521 ms against 2120 ms). Local timings therefore do not scale the lane, they REORDER it, so a packing derived from them would have balanced the wrong thing while looking clean. WHAT IT REPLACES, which is the root cause. The lanes were packed from the 2026-08-20 concurrent isolation proof: 24 candidates across four LOCAL workers. That record answers whether the candidates are isolation-safe, not how long a SERIAL CI lane runs, so it was structurally incapable of representing lane wall clock even when it was fresh. It was also never refreshed while the set grew about 3.2x. Both the wrong instrument and the staleness are fixed here: the hints now come from the lane itself and carry their run ids and date. === PART 2: THE SPLIT DERIVED FROM THAT TOTAL === Longest-processing-time assignment over those hints gives 414269 ms and 414299 ms, 30 ms apart, against 624299/204269 before. tests/fm-pi-primary-types.test.sh stays in lane 1 because that is the job which installs the Pi package, so ci.yml needs no step changes. === PART 3: DOES THE MARGIN SURVIVE MACHINE VARIANCE === Stated explicitly, because 6.90 min against a 10 min cap is 69% of cap before any variance is applied, and the cap covers the whole job rather than the suite. worst lane, script time 414299 ms 6.90 min job overhead, measured on the real lane ~18 s (see below) expected healthy job ~432300 ms 7.21 min x1.29 on the script time, plus overhead ~552400 ms 9.21 min cap 600000 ms 10.00 min room left after the multiplication ~47.6 s 7.9% of cap The 1.29x is the runner variance measured today on the SIBLING SERIAL lane, as supplied; it is not this lane's own figure. This lane family does have its own, and it is tighter: the six full lane-2 sums today span 192939 ms to 203451 ms, a spread of 1.054x. At that figure the worst lane lands near 7.58 min with about 2.4 min of room. I have used the LARGER, borrowed 1.29x for the verdict rather than the tighter one this lane actually shows, and note that the hints are already per-script maxima, so 1.29x on top is conservative twice over. THE MARGIN SURVIVES THE MULTIPLICATION, so this proceeds rather than stopping. The 18 s overhead is measured, not assumed: in run 34462758357 the lane-1 job ran 10 min 16 s against a 598.2 s suite, and lane 2 ran 3 min 21 s against a 192.9 s suite, a ~10 s difference that matches lane 1's extra Pi package install. The cap is unchanged, the lane count is unchanged, and nothing in the serial lane, its shard count, its guard or its hint table is touched. === PART 4: THE RECORDED FACT === The workflow comment no longer restates the shard wall as a literal, which is how "~1 min of serial sum" survived a 10x change without announcing it. It now points at bin/fm-test-run.sh --check-coverage, which prints parallel_max_ms, parallel_imbalance_ms and parallel_unhinted derived from the hint table, so the current number is computed on demand. The shard documentation carries the dated run ids, which route it was taken by, and the local cross-check that shows why local numbers are not admissible as hints. Two regressions pin what rotted: lane membership must be stored longest-measured-first, and the lanes must be fully hinted and packed within 5% of each other. Both were run against the old composition and both fail on it (420030 ms imbalance against a 624299 ms worst lane). The ordering assertion they replace named a specific script by hand and had itself gone stale. === PART 5: NAMED AND LEFT, OUTSIDE THIS REBALANCE === tests/fm-captain-hold-lifecycle.test.sh alone is 296481 ms, 36% of the whole set, so it is the floor of any two-lane split: no repacking can put a lane below it. After this rebalance the cap is about 1.45x the healthy lane where the sibling serial lane keeps roughly 2x. Nothing refuses a stale parallel hint the way PORTABLE_SERIAL_MAX_UNHINTED_PERCENT bounds the serial lane. parallel_unhinted is reported, not enforced, which is what let this drift for three weeks unnoticed. * fix(review): Restrict parallel scheduling hints to portable parallel lanes * fix(document): Clarify parallel lane scheduling and timing evidence --- .github/workflows/ci.yml | 10 ++- bin/fm-test-run.sh | 129 ++++++++++++++++++++++++++------ docs/fm-test-portable-shards.md | 67 ++++++++--------- tests/fm-test-run.test.sh | 95 +++++++++++++++++++++-- 4 files changed, 236 insertions(+), 65 deletions(-) diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 94e3d1121f3..dc681e1e0d0 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -47,8 +47,13 @@ jobs: tests-portable-parallel-1: name: Behavior portable parallel 1 runs-on: ubuntu-latest - # Measured shard wall is ~1 min of serial sum on proven scripts; this cap is - # a hang tripwire with margin, not the expected healthy end of the lane. + # This cap is intended as a hang tripwire, but the previous lane 1 reached + # it; the former "~1 min of serial sum" estimate no longer applies. + # Compare it with the derived hints from fm-test-run.sh --check-coverage + # and completed job timings, allowing for setup and runner-speed spread. + # A packed hint sum is not a measured job wall time or proof of headroom. + # Evidence and refresh procedure: docs/fm-test-portable-shards.md. + # Changes to this cap or the lane count require a separate scope decision. timeout-minutes: 10 steps: - uses: actions/checkout@v6 @@ -92,6 +97,7 @@ jobs: tests-portable-parallel-2: name: Behavior portable parallel 2 runs-on: ubuntu-latest + # Same timeout rationale as portable parallel shard 1 above. timeout-minutes: 10 steps: - uses: actions/checkout@v6 diff --git a/bin/fm-test-run.sh b/bin/fm-test-run.sh index fd0854a0df0..fcfa4b4f43e 100755 --- a/bin/fm-test-run.sh +++ b/bin/fm-test-run.sh @@ -18,6 +18,7 @@ # fm-test-run.sh --list --family # fm-test-run.sh --list --lane portable-parallel-1 # fm-test-run.sh --list-scheduled --family +# fm-test-run.sh --list-scheduled --lane portable-parallel-1 # fm-test-run.sh --list-families # fm-test-run.sh --list-concurrent-safe-families # fm-test-run.sh --concurrent-safe-family-jobs-max @@ -35,7 +36,11 @@ # tool this host could not exercise. # --list print selected script paths (one per line) and exit 0 # --list-scheduled -# print selected paths longest-hint-first and exit 0 +# print selected paths longest-hint-first and exit 0. +# Only --lane portable-parallel-1 or portable-parallel-2 uses +# parallel hints, falling back to serial weights if missing. +# Every other selection uses serial weights alone. +# Equal weights are ordered by path under LC_ALL=C. # --base with --changed, compare against this ref (default: origin/main) # --exclude-family # drop scripts whose primary family matches after selection @@ -59,7 +64,7 @@ # family proofs may impose a lower cap. Individually proven # scripts share one phase; scripts admitted only by a family # proof run in a separate phase for each family. Concurrent -# phases are ordered longest-hint-first. Unproven stateful +# phases use serial weights, longest-hint-first. Unproven stateful # scripts run serially after all concurrent phases. Default is # 1 (serial) except for plain --changed and a plain list of # script paths, which use the bounded automatic scheduler. @@ -114,7 +119,13 @@ # Family labels, the changed-file map, and production portable-shard composition # live in this script only (one owner). The proven-isolated candidate set remains # owned by bin/fm-test-isolation-proof.sh; portable parallel shards are a -# duration-balanced partition of that exact set (see docs/fm-test-portable-shards.md). +# duration-balanced partition of that exact set, packed from the measured hints +# in portable_parallel_weight_hints (see docs/fm-test-portable-shards.md). +# --check-coverage reports parallel_max_ms (the larger lane hint sum), +# parallel_imbalance_ms (the absolute difference between the sums), and +# parallel_unhinted (the number of members missing a parallel hint). +# These sums exclude unhinted members and are estimates, not measured job wall +# times. Missing parallel hints are reported without failing this guard. # # portable-serial stays strictly serial. Its CI shards (portable-serial-of) # split it across separate runners, so two of its stateful scripts still never @@ -464,41 +475,87 @@ tests/fm-x-mode.test.sh EOF } -# Portable parallel shard 1: LPT balance of the proven-isolated set using the -# current concurrent-proof durations in docs/fm-test-isolation-proof.json. -# Execution order is longest first so wall-clock stays near the balanced sum. +# Per-script serial CI duration hints, one " " per line, used to +# pack only the two portable parallel lanes. Measurement provenance and the +# refresh procedure are owned by docs/fm-test-portable-shards.md. +portable_parallel_weight_hints() { + cat <<'EOF' +tests/fm-arm-pretool-check.test.sh 30898 +tests/fm-backend-herdr.test.sh 22144 +tests/fm-brief.test.sh 1625 +tests/fm-captain-hold-lifecycle.test.sh 296481 +tests/fm-cd-pretool-check.test.sh 16964 +tests/fm-composer-ghost.test.sh 2120 +tests/fm-composer-lib.test.sh 4798 +tests/fm-crew-state.test.sh 11557 +tests/fm-ensure-agents-md.test.sh 901 +tests/fm-grok-harness.test.sh 6563 +tests/fm-herdr-lab.test.sh 6936 +tests/fm-lint.test.sh 164262 +tests/fm-pi-primary-types.test.sh 8624 +tests/fm-pr-merge.test.sh 111145 +tests/fm-review-diff.test.sh 2747 +tests/fm-send-popup-settle.test.sh 4939 +tests/fm-send-settle.test.sh 2051 +tests/fm-send-strict.test.sh 3861 +tests/fm-spawn-batch.test.sh 2265 +tests/fm-supervision-instructions.test.sh 297 +tests/fm-test-run.test.sh 92944 +tests/fm-tmux-submit-busy.test.sh 2477 +tests/fm-transition-lib.test.sh 99 +tests/fm-x-mode.test.sh 31870 +EOF +} + +# Sum the hints above for the scripts read on stdin, and report how many of +# them had no hint at all, as " ". +portable_parallel_lane_weight() { + awk ' + NR == FNR { if (NF) { hint[$1] = $2 } ; next } + NF { + if ($1 in hint) { total += hint[$1] } else { unhinted++ } + } + END { printf "%d %d\n", total + 0, unhinted + 0 } + ' <(portable_parallel_weight_hints) - +} + +# Portable parallel shard 1: LPT balance of the proven-isolated set over the +# hints above. Stored order agrees with this lane's --list-scheduled output. +# tests/fm-pi-primary-types.test.sh belongs to this lane because +# this is the parallel job that installs the Pi package; moving it needs that +# workflow step moved with it. list_portable_parallel_1() { cat <<'EOF' -tests/fm-x-mode.test.sh -tests/fm-cd-pretool-check.test.sh -tests/fm-captain-hold-lifecycle.test.sh -tests/fm-test-run.test.sh -tests/fm-composer-ghost.test.sh -tests/fm-grok-harness.test.sh tests/fm-lint.test.sh +tests/fm-pr-merge.test.sh +tests/fm-test-run.test.sh +tests/fm-cd-pretool-check.test.sh tests/fm-pi-primary-types.test.sh +tests/fm-grok-harness.test.sh +tests/fm-composer-lib.test.sh tests/fm-review-diff.test.sh +tests/fm-tmux-submit-busy.test.sh +tests/fm-composer-ghost.test.sh tests/fm-brief.test.sh -tests/fm-transition-lib.test.sh EOF } # Portable parallel shard 2: the complementary LPT half of the proven set. list_portable_parallel_2() { cat <<'EOF' -tests/fm-backend-herdr.test.sh +tests/fm-captain-hold-lifecycle.test.sh +tests/fm-x-mode.test.sh tests/fm-arm-pretool-check.test.sh +tests/fm-backend-herdr.test.sh tests/fm-crew-state.test.sh tests/fm-herdr-lab.test.sh -tests/fm-pr-merge.test.sh tests/fm-send-popup-settle.test.sh -tests/fm-tmux-submit-busy.test.sh -tests/fm-send-settle.test.sh tests/fm-send-strict.test.sh tests/fm-spawn-batch.test.sh -tests/fm-supervision-instructions.test.sh +tests/fm-send-settle.test.sh tests/fm-ensure-agents-md.test.sh -tests/fm-composer-lib.test.sh +tests/fm-supervision-instructions.test.sh +tests/fm-transition-lib.test.sh EOF } @@ -749,6 +806,16 @@ portable_serial_unhinted() { rm -rf "$tmp" } +portable_parallel_weight_for() { + local want=$1 ms + ms=$(portable_parallel_weight_hints | awk -v want="$want" '$1 == want { print $2; exit }') + if [ -n "$ms" ]; then + printf '%s\n' "$ms" + return 0 + fi + portable_serial_weight_for "$want" +} + portable_serial_weight_for() { local want=$1 path ms while read -r path ms; do @@ -877,6 +944,7 @@ select_lane() { run_coverage_guard() { local tmp missing extra a b shard unhinted serial_total + local p1_ms p1_unhinted p2_ms p2_unhinted parallel_max_ms parallel_imbalance_ms local -a saved_scripts=() tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-coverage.XXXXXX") @@ -1006,9 +1074,21 @@ run_coverage_guard() { fi fi - printf 'FM_TEST_COVERAGE ok total=%s parallel=%s serial=%s serial_shards=%s serial_unhinted=%s herdr=%s\n' \ + # Keep these estimates derived from the membership and hint owners; see the + # header for the distinction between packed weights and measured job time. + read -r p1_ms p1_unhinted <<<"$(list_portable_parallel_1 | portable_parallel_lane_weight)" + read -r p2_ms p2_unhinted <<<"$(list_portable_parallel_2 | portable_parallel_lane_weight)" + parallel_max_ms=$p1_ms + [ "$p2_ms" -le "$parallel_max_ms" ] || parallel_max_ms=$p2_ms + parallel_imbalance_ms=$((p1_ms - p2_ms)) + [ "$parallel_imbalance_ms" -ge 0 ] || parallel_imbalance_ms=$((-parallel_imbalance_ms)) + + printf 'FM_TEST_COVERAGE ok total=%s parallel=%s parallel_max_ms=%s parallel_imbalance_ms=%s parallel_unhinted=%s serial=%s serial_shards=%s serial_unhinted=%s herdr=%s\n' \ "$(wc -l <"$tmp/all" | tr -d ' ')" \ "$(wc -l <"$tmp/shards_union" | tr -d ' ')" \ + "$parallel_max_ms" \ + "$parallel_imbalance_ms" \ + "$((p1_unhinted + p2_unhinted))" \ "$(wc -l <"$tmp/serial" | tr -d ' ')" \ "$PORTABLE_SERIAL_SHARDS" \ "$unhinted" \ @@ -1936,7 +2016,14 @@ fi if [ "$LIST_ONLY" -eq 1 ] || [ "$LIST_SCHEDULED" -eq 1 ]; then if [ "$LIST_SCHEDULED" -eq 1 ]; then for s in "${SCRIPTS[@]+"${SCRIPTS[@]}"}"; do - printf '%s\t%s\n' "$(portable_serial_weight_for "$s")" "$s" + case "$MODE:$LANE" in + lane:portable-parallel-1|lane:portable-parallel-2) + printf '%s\t%s\n' "$(portable_parallel_weight_for "$s")" "$s" + ;; + *) + printf '%s\t%s\n' "$(portable_serial_weight_for "$s")" "$s" + ;; + esac done | LC_ALL=C sort -t"$(printf '\t')" -k1,1nr -k2,2 | cut -f2- else for s in "${SCRIPTS[@]+"${SCRIPTS[@]}"}"; do diff --git a/docs/fm-test-portable-shards.md b/docs/fm-test-portable-shards.md index cc069175698..45c3a5fd763 100644 --- a/docs/fm-test-portable-shards.md +++ b/docs/fm-test-portable-shards.md @@ -5,47 +5,40 @@ ## Verification inputs -The current candidate timings came from the 2026-08-20 concurrent proof recorded in [fm-test-isolation-proof.md](fm-test-isolation-proof.md). -The proof ran 24 candidates with four workers and no failures. +Balance hints come from serial runs of the real lanes on `ubuntu-latest`. +The concurrent isolation proof in [fm-test-isolation-proof.md](fm-test-isolation-proof.md) establishes concurrency safety, not serial CI duration. +Local timings are not interchangeable with CI timings: platform and machine load can affect each script differently and change their relative weights. -| duration_ms | script | +The retained hints are the slowest completed value each script reached across six CI runs on 2026-09-10: [34459949083](https://github.com/kunchenguid/firstmate/actions/runs/34459949083), [34460760299](https://github.com/kunchenguid/firstmate/actions/runs/34460760299), [34462530836](https://github.com/kunchenguid/firstmate/actions/runs/34462530836), [34462758357](https://github.com/kunchenguid/firstmate/actions/runs/34462758357), [34466966385](https://github.com/kunchenguid/firstmate/actions/runs/34466966385), and [34470382458](https://github.com/kunchenguid/firstmate/actions/runs/34470382458). +Shard 2 completed in all six, so its scripts come from the uploaded `fm-test-timing-portable-parallel-2` artifacts. +Shard 1 was cancelled at its job cap in five of the six, so its scripts come from the `FM_TEST_END duration_ms=` markers in each cancelled job's log, which record every script that finished before the cancellation, plus the one complete `fm-test-timing-portable-parallel-1` artifact from run 34462758357. +Observed maxima provide conservative packing weights, not an upper bound on future durations. + +The measurements cover all 24 candidates, with six samples per script except: + +| Samples | Scripts | |---:|---| -| 45356 | `tests/fm-backend-herdr.test.sh` | -| 35415 | `tests/fm-x-mode.test.sh` | -| 35095 | `tests/fm-captain-hold-lifecycle.test.sh` | -| 27529 | `tests/fm-arm-pretool-check.test.sh` | -| 20922 | `tests/fm-test-run.test.sh` | -| 17558 | `tests/fm-crew-state.test.sh` | -| 16582 | `tests/fm-cd-pretool-check.test.sh` | -| 9766 | `tests/fm-lint.test.sh` | -| 9562 | `tests/fm-herdr-lab.test.sh` | -| 6768 | `tests/fm-grok-harness.test.sh` | -| 6290 | `tests/fm-pr-merge.test.sh` | -| 5569 | `tests/fm-composer-ghost.test.sh` | -| 4563 | `tests/fm-send-popup-settle.test.sh` | -| 4021 | `tests/fm-tmux-submit-busy.test.sh` | -| 3544 | `tests/fm-composer-lib.test.sh` | -| 3025 | `tests/fm-send-strict.test.sh` | -| 2753 | `tests/fm-send-settle.test.sh` | -| 2166 | `tests/fm-review-diff.test.sh` | -| 1315 | `tests/fm-brief.test.sh` | -| 975 | `tests/fm-spawn-batch.test.sh` | -| 598 | `tests/fm-pi-primary-types.test.sh` | -| 513 | `tests/fm-ensure-agents-md.test.sh` | -| 331 | `tests/fm-supervision-instructions.test.sh` | -| 99 | `tests/fm-transition-lib.test.sh` | +| 4 | `tests/fm-lint.test.sh` | +| 3 | `tests/fm-pi-primary-types.test.sh`, `tests/fm-review-diff.test.sh` | +| 1 | `tests/fm-brief.test.sh`, `tests/fm-transition-lib.test.sh` | -## Parallel lanes +The two scripts with one sample are the tail of shard 1 that only the complete run reached. +Collect completed per-script measurements for every member before calculating a split. +A cancelled lane's elapsed duration is only a lower bound; its unfinished scripts have no completed duration for that invocation. +The complete historical run supplies tail-script hints, not a completion time for any later cancelled invocation or for the rebalanced jobs. -The two parallel lanes use longest-processing-time assignment from those measured durations. +## Parallel lanes -| Lane | Script count | Estimated duration | -|---|---:|---:| -| `portable-parallel-1` | 11 | 134295 ms (~134.3 s) | -| `portable-parallel-2` | 13 | 126020 ms (~126.0 s) | -| imbalance | | 8275 ms | +The two parallel lanes use longest-processing-time assignment over those hints. +[`bin/fm-test-run.sh`](../bin/fm-test-run.sh) holds the duration values in `portable_parallel_weight_hints` and the ordered memberships and lane-specific prerequisite constraints beside `list_portable_parallel_1` and `list_portable_parallel_2`. +Read the derived packing estimates with that runner's `--check-coverage`; its header and `--help` own the output fields and the selection-specific `--list-scheduled` weight rules. +The largest individual hint sets a lower bound on the estimated duration of any split, regardless of how evenly the remaining work is assigned. +The CI cap and its rationale are owned by [`.github/workflows/ci.yml`](../.github/workflows/ci.yml). -`bin/fm-test-run.sh` contains the exact ordered memberships in `list_portable_parallel_1` and `list_portable_parallel_2`. +[`tests/fm-test-run.test.sh`](../tests/fm-test-run.test.sh), in `test_portable_parallel_lanes_stay_duration_balanced`, requires every parallel member to have a hint and the lane sums to differ by no more than five percent of the larger sum. +Its scheduling regressions also check stored parallel lane order and preserve serial-weight scheduling for other selections. +These checks do not detect a script outgrowing an existing hint or establish measured job headroom. +Refresh `portable_parallel_weight_hints` with the slowest completed `duration_ms` per script from several green CI runs' `fm-test-timing-portable-parallel-*` artifacts whenever the parallel set gains scripts or a member grows materially. ## Portable serial remainder @@ -125,9 +118,9 @@ Portable shards, each portable serial shard, and the Herdr lane upload runner-ge | Lane | Bound | Rationale | |---|---|---| -| portable parallel 1/2 | job `timeout-minutes: 10` | The measured shard sums are about three minutes and the timeout is a hang tripwire. | +| portable parallel 1/2 | See [CI workflow](../.github/workflows/ci.yml) | The workflow owns the parallel cap rationale and its evidence limits. | | portable serial 1-5 | job `timeout-minutes: 30` | Current runners can take about 20 minutes; the 30-minute cap remains a hang tripwire while leaving margin for job setup and runner-speed spread. | | Herdr | family-run step `timeout-minutes: 20`; job `timeout-minutes: 75` backstop | Healthy runs finished around 7 minutes before this lane gained `fm-backend-herdr-focus-flash-e2e`, which measures about 2 minutes against a real lab locally, so the step bound is still the hang tripwire (cleanup and timing artifacts still upload) while the job cap stays a last-resort backstop. Refresh this figure from the lane's uploaded timing artifact. | -Timeouts are hang tripwires rather than expected healthy durations. +Timeouts are intended as hang tripwires; a passing coverage guard does not establish a healthy job duration. `.github/workflows/ci.yml` owns the exact numbers. diff --git a/tests/fm-test-run.test.sh b/tests/fm-test-run.test.sh index b90dfa2129d..35f3a8d4980 100755 --- a/tests/fm-test-run.test.sh +++ b/tests/fm-test-run.test.sh @@ -962,8 +962,65 @@ test_exclude_family() { pass "exclude-family drops the named primary family after selection" } +test_list_scheduled_proven_isolated_uses_serial_weights() { + local tmp + tmp=$(fm_test_tmproot fm-test-run-proven-schedule) + "$RUNNER" --list --proven-isolated | LC_ALL=C sort >"$tmp/expected" + "$RUNNER" --list-scheduled --proven-isolated >"$tmp/actual" \ + || fail "--list-scheduled --proven-isolated failed" + cmp -s "$tmp/expected" "$tmp/actual" \ + || fail "proven-isolated scheduling must break serial-default ties by path" + pass "proven-isolated scheduling ignores parallel hints" +} + +test_list_scheduled_non_lane_selections_use_serial_weights() { + local tmp repo script selection + local -a scripts=( + tests/fm-operational-input.test.sh + tests/fm-lint.test.sh + tests/fm-muse-harness.test.sh + tests/fm-captain-hold-lifecycle.test.sh + tests/fm-kimi-harness.test.sh + tests/fm-brief.test.sh + ) + tmp=$(fm_test_tmproot fm-test-run-non-lane-schedule) + repo="$tmp/repo" + mkdir -p "$repo/bin" "$repo/tests" + cp "$RUNNER" "$repo/bin/fm-test-run.sh" + for script in "${scripts[@]}"; do + printf '#!/usr/bin/env bash\nexit 0\n' >"$repo/$script" + chmod +x "$repo/$script" + done + git -C "$repo" init -q + git -C "$repo" add . + git -C "$repo" -c user.name=test -c user.email=test@example.invalid commit -qm baseline + for script in "${scripts[@]}"; do + printf '\n' >>"$repo/$script" + done + printf '%s\n' \ + tests/fm-muse-harness.test.sh \ + tests/fm-brief.test.sh \ + tests/fm-captain-hold-lifecycle.test.sh \ + tests/fm-lint.test.sh \ + tests/fm-kimi-harness.test.sh \ + tests/fm-operational-input.test.sh >"$tmp/expected" + for selection in family all changed scripts; do + case "$selection" in + family) set -- --family pure-contract-unit ;; + all) set -- --all ;; + changed) set -- --changed --base HEAD ;; + scripts) set -- "${scripts[@]}" ;; + esac + "$repo/bin/fm-test-run.sh" --list-scheduled "$@" >"$tmp/actual" \ + || fail "--list-scheduled $selection failed" + cmp -s "$tmp/expected" "$tmp/actual" \ + || fail "$selection scheduling must use serial hints and path-ordered default ties" + done + pass "family, all, changed, and script selections ignore parallel hints" +} + test_portable_shard_union_and_coverage_guard() { - local s1 s2 proven serial herdr all_count union_count overlap out first + local s1 s2 proven serial herdr all_count union_count overlap out lane s1=$("$RUNNER" --list --lane portable-parallel-1) s2=$("$RUNNER" --list --lane portable-parallel-2) proven=$("$RUNNER" --list --proven-isolated) @@ -991,13 +1048,38 @@ test_portable_shard_union_and_coverage_guard() { # No duplicates across the four partitions. [ "$(printf '%s\n' "$s1" "$s2" "$serial" "$herdr" | LC_ALL=C sort | uniq -d | wc -l | tr -d ' ')" = "0" ] \ || fail "lanes must not duplicate scripts" - # LPT order: first script of shard 1 is the longest proven script. - first=$(printf '%s\n' "$s1" | head -n 1) - [ "$first" = "tests/fm-x-mode.test.sh" ] \ - || fail "shard 1 must start with the longest proven script, got $first" + # LPT execution order, asserted against the runner's own measured schedule + # rather than against a script name: naming the current longest script here is + # what let the recorded lane duration go stale unnoticed in the first place. + for lane in portable-parallel-1 portable-parallel-2; do + [ "$("$RUNNER" --list --lane "$lane")" = "$("$RUNNER" --list-scheduled --lane "$lane")" ] \ + || fail "$lane membership must be stored longest-measured-first" + done pass "portable shard union, disjointness, and coverage guard hold" } +# The two parallel lanes are only "duration-balanced" while every member has a +# measured hint and the packing over those hints stays even. Both halves went +# unchecked until one lane grew past its CI job cap and was cancelled on every +# run, so assert them through the guard's own reported numbers. +test_portable_parallel_lanes_stay_duration_balanced() { + local out max imbalance unhinted + out=$("$RUNNER" --check-coverage) + unhinted=$(printf '%s\n' "$out" | sed -n 's/.*parallel_unhinted=\([0-9]*\).*/\1/p') + max=$(printf '%s\n' "$out" | sed -n 's/.*parallel_max_ms=\([0-9]*\).*/\1/p') + imbalance=$(printf '%s\n' "$out" | sed -n 's/.*parallel_imbalance_ms=\([0-9]*\).*/\1/p') + [ -n "$unhinted" ] && [ -n "$max" ] && [ -n "$imbalance" ] \ + || fail "coverage guard must report parallel_unhinted, parallel_max_ms, parallel_imbalance_ms: $out" + [ "$unhinted" = "0" ] \ + || fail "$unhinted proven-isolated scripts have no measured parallel hint, so the lanes are packed on a guess" + [ "$max" -gt 0 ] || fail "parallel_max_ms must be a positive packed duration, got $max" + # 5% of the worst lane: wide enough that one script's growth does not trip it, + # narrow enough that a lopsided partition cannot call itself balanced. + [ "$((imbalance * 20))" -le "$max" ] \ + || fail "parallel lanes differ by ${imbalance}ms against a ${max}ms worst lane, more than 5%" + pass "portable parallel lanes are fully hinted and packed within 5% of each other" +} + test_portable_serial_shards_partition_the_serial_lane() { local lanes count serial shard listed union dups shard_lane total cap lanes=$("$RUNNER" --list-lanes) @@ -1602,7 +1684,10 @@ test_a_run_that_ran_records_no_skip_reason test_live_guards_expect_a_capability_skip_class test_fail_on_gate_skip_token test_exclude_family +test_list_scheduled_proven_isolated_uses_serial_weights +test_list_scheduled_non_lane_selections_use_serial_weights test_portable_shard_union_and_coverage_guard +test_portable_parallel_lanes_stay_duration_balanced test_portable_serial_shards_partition_the_serial_lane test_portable_serial_hint_coverage_is_reported_and_bounded test_portable_serial_shard_lane_refusals From 0a4e4a27cb7e911f953ece097bf4e6fe0f67dec3 Mon Sep 17 00:00:00 2001 From: Christopher McKay <101884182+karotkriss@users.noreply.github.com> Date: Fri, 11 Sep 2026 10:44:39 -0400 Subject: [PATCH 02/31] fix(bin): stop claiming prose-mentioned PR URLs as a task's delivered PR (#4148) pr_for_task fell back to scraping the whole status log with tail -1, so any PR URL a worker ever mentioned in prose - including a scout citing someone else's PR - became the task's delivered PR in the parent-channel terminal report. Recorded meta pr= is now the only authoritative source, the fallback scrape accepts only a preferred terminal line in a mode's ready-signal shape (done: PR or done: PR checks green), and a scout never carries pr= at all. --- bin/fm-inactive-reconcile.sh | 20 +++++++++------ tests/fm-inactive-reconcile.test.sh | 39 ++++++++++++++++++++++++++--- 2 files changed, 48 insertions(+), 11 deletions(-) diff --git a/bin/fm-inactive-reconcile.sh b/bin/fm-inactive-reconcile.sh index 9c30a9074be..5cf22755626 100755 --- a/bin/fm-inactive-reconcile.sh +++ b/bin/fm-inactive-reconcile.sh @@ -311,15 +311,19 @@ meta_incarnation() { # printf 'legacy-%s\n' "$(sha256_text "$identity")" } -pr_for_task() { # [preferred-line] - local meta=$1 status=$2 preferred=${3:-} value +# The task's delivered PR. Recorded meta pr= is the only authoritative source; +# the fallback scrape accepts only a preferred terminal line in a mode's +# ready-signal shape (`done: PR ` or `done: PR checks green`), so a +# PR a worker merely mentioned in prose is never claimed as the delivery. +# A scout never delivers a PR, so it never carries one. +pr_for_task() { # [preferred-line] + local meta=$1 preferred=${2:-} value + [ "$(meta_field "$meta" kind)" != scout ] || return 0 value=$(meta_field "$meta" pr) if [ -z "$value" ] && [ -n "$preferred" ]; then value=$(printf '%s\n' "$preferred" \ - | grep -Eo 'https?://[^[:space:])"]+/pull/[0-9]+' | head -1 || true) - fi - if [ -z "$value" ] && [ -f "$status" ]; then - value=$(grep -Eo 'https?://[^[:space:])"]+/pull/[0-9]+' "$status" 2>/dev/null | tail -1 || true) + | sed -nE 's|^done: PR (https?://[^[:space:])"]+/pull/[0-9]+)( checks green)?$|\1|p' \ + | head -1 || true) fi clean_field "$value" } @@ -398,7 +402,7 @@ report_child_ledger_locked() { # status="$STATE/$id.status" last=$(child_terminal_ledger_line "$status") || return 0 state=$(status_line_verb "$last") - pr=$(pr_for_task "$meta" "$status" "$last") + pr=$(pr_for_task "$meta" "$last") incarnation=$(meta_incarnation "$meta") fingerprint=$(sha256_text "$incarnation|$id|$state|ledger|$last") previous=$(grep -v '^[[:space:]]*$' "$status" 2>/dev/null \ @@ -496,7 +500,7 @@ reconcile_direct_child_locked() { # "$MATE/data/scout/report.md" write_child "$MATE" boom 'failed: build broke' - write_child "$MATE" replaced-pr $'working: old PR https://example.test/owner/repo/pull/11\ndone: replacement PR https://example.test/owner/repo/pull/22' + write_child "$MATE" replaced-pr $'working: old PR https://example.test/owner/repo/pull/11\ndone: PR https://example.test/owner/repo/pull/22' awk '$0 !~ /^pr=/' "$MATE/state/replaced-pr.meta" > "$MATE/state/replaced-pr.meta.tmp" mv "$MATE/state/replaced-pr.meta.tmp" "$MATE/state/replaced-pr.meta" FM_FAKE_CREW_STATE='unknown' run_reconcile "$MATE" @@ -244,8 +244,8 @@ test_secondmate_ledger_delivery_carries_report_and_failure() { "$MAIN/state/mate.status" || fail "scout delivery lost its report pointer: $(cat "$MAIN/state/mate.status")" grep -Fxq "failed [key=$boom_key]: child boom failed: build broke pr=https://example.test/owner/repo/pull/1 mode=no-mistakes yolo=off" \ "$MAIN/state/mate.status" || fail "failed line was not delivered under the failed verb: $(cat "$MAIN/state/mate.status")" - grep -Fxq "done [key=$replaced_key]: child replaced-pr done: replacement PR https://example.test/owner/repo/pull/22 pr=https://example.test/owner/repo/pull/22 mode=no-mistakes yolo=off" \ - "$MAIN/state/mate.status" || fail "ledger fallback did not prefer the terminal line PR: $(cat "$MAIN/state/mate.status")" + grep -Fxq "done [key=$replaced_key]: child replaced-pr done: PR https://example.test/owner/repo/pull/22 pr=https://example.test/owner/repo/pull/22 mode=no-mistakes yolo=off" \ + "$MAIN/state/mate.status" || fail "ledger fallback did not prefer the terminal ready line PR: $(cat "$MAIN/state/mate.status")" printf 'working: retrying\ndone: fixed on retry\n' >> "$MATE/state/boom.status" FM_FAKE_CREW_STATE='unknown' run_reconcile "$MATE" boom_key=$(reported_outcome_key "$MATE" boom 'done') || fail "recovered receipt key missing" @@ -256,6 +256,38 @@ test_secondmate_ledger_delivery_carries_report_and_failure() { pass "ledger delivery carries the report pointer, the failed verb, and each new terminal line" } +# A PR URL a worker only ever mentioned in prose is never claimed as the +# task's delivered PR: without a recorded PR, only a terminal line in the +# ready-signal shape carries one, and a scout never carries one at all. +test_pr_field_requires_recorded_pr_or_ready_signal_line() { + local id prose_key ready_key scout_key + make_world pr-provenance; bind_secondmate local + write_child "$MATE" prose $'working: context in https://example.test/other/repo/pull/33\ndone: cleanup finished' + write_child "$MATE" ready 'done: PR https://example.test/owner/repo/pull/44 checks green' + write_child "$MATE" lookout 'done: PR https://example.test/owner/repo/pull/55' + for id in prose ready; do + awk '$0 !~ /^pr=/' "$MATE/state/$id.meta" > "$MATE/state/$id.meta.tmp" + mv "$MATE/state/$id.meta.tmp" "$MATE/state/$id.meta" + done + awk '{ sub(/^kind=ship$/, "kind=scout"); print }' "$MATE/state/lookout.meta" \ + > "$MATE/state/lookout.meta.tmp" + mv "$MATE/state/lookout.meta.tmp" "$MATE/state/lookout.meta" + FM_FAKE_CREW_STATE='unknown' run_reconcile "$MATE" + prose_key=$(reported_outcome_key "$MATE" prose 'done') || fail "prose receipt key missing" + ready_key=$(reported_outcome_key "$MATE" ready 'done') || fail "ready receipt key missing" + scout_key=$(reported_outcome_key "$MATE" lookout 'done') || fail "scout receipt key missing" + grep -Fxq "done [key=$prose_key]: child prose done: cleanup finished mode=no-mistakes yolo=off" \ + "$MAIN/state/mate.status" \ + || fail "a PR mentioned only in prose was claimed as the delivery: $(cat "$MAIN/state/mate.status")" + grep -Fxq "done [key=$ready_key]: child ready done: PR https://example.test/owner/repo/pull/44 checks green pr=https://example.test/owner/repo/pull/44 mode=no-mistakes yolo=off" \ + "$MAIN/state/mate.status" \ + || fail "a ready-signal terminal line did not carry its PR: $(cat "$MAIN/state/mate.status")" + grep -Fxq "done [key=$scout_key]: child lookout done: PR https://example.test/owner/repo/pull/55 mode=no-mistakes yolo=off" \ + "$MAIN/state/mate.status" \ + || fail "a scout's ready-looking line carried a PR claim: $(cat "$MAIN/state/mate.status")" + pass "pr= requires the recorded PR or a ready-signal terminal line, and never a scout" +} + # If a terminal ledger line lands while the authoritative state read is in # flight, the ledger path remains the single owner on the next poll. test_terminal_line_during_state_read_yields_to_ledger_delivery() { @@ -790,6 +822,7 @@ test_main_direct_terminal_presentation_receipt test_local_secondmate_delivers_terminal_ledger_line test_busy_child_does_not_starve_later_ledger_outcomes test_secondmate_ledger_delivery_carries_report_and_failure +test_pr_field_requires_recorded_pr_or_ready_signal_line test_terminal_line_during_state_read_yields_to_ledger_delivery test_terminal_line_after_inactive_delivery_is_not_reported_twice test_progress_after_inactive_delivery_starts_a_new_event From 8a61b50545881ef10df1faefb6c43801b1d005b1 Mon Sep 17 00:00:00 2001 From: Christoph Meise Date: Fri, 11 Sep 2026 17:11:49 +0200 Subject: [PATCH 03/31] fix(procevent): confirm reconcile launches and reclaim provably dead claims instead of counting a dead drop as started (#4212) * fix(procevent): stop a dead runner owning a source and reconcile reporting it The captain answered ten calls on a bearings board, the board accepted them, and nothing collected them. He had to answer all ten again in chat. A surface that presents as armed while being a dead drop is worse than one that visibly fails, because the answers looked recorded. Two independent defects, reproduced together in an isolated home where reconcile reports started=1 on every run while ownership never moves and no runner ever attaches. 1. reconcile counted a launch it never verified. detach_runner is fire-and-forget and discards the child's stderr, so a runner that died before it could claim was counted exactly like one that is listening. Launches are now confirmed - the source observed owned, or its runner record moved - before being reported as started; the rest are reported as failed= with a non-zero exit. The runner-record clause is what keeps a fast-completing source from being reported as a failure when it finished between two polls. One bounded window covers a whole cycle's launches, so a home full of broken sources costs the same wait as one. 2. A claim whose whole generation is provably gone could be refused forever. Reclaiming it ran cleanups over that dead generation's own leftovers, and any failure vetoed the claim - permanently, because none of those conditions clears on its own. Every one of those leftovers is keyed by the dead generation's claim token and a replacement always claims a fresh one, so none can collide with what replaces it. fm_procevent_claim_capture_reservation_reclaim_locked already said this for the reservation record; the staging file and the shape check on the registry directory recorded to hold it now take the same rule. Removing the claim record itself stays a hard precondition: two owners is the one outcome worse than none. Two smaller repairs to the same "registered is not listening" confusion: - `list` reported OWNER=none for a source nothing can claim. A reused PID whose process group survives reaches that state through the stale branch rather than the leaderless one, so it read as an idle source waiting to be started - the reassuring answer this surface gave while a board collected nothing. It now reports the orphaned state it shares. - reconcile relaunched into that same unclaimable state on every cycle, spawning a runner that could only die on the claim. docs/configuration.md already promised it preserves such a claim without starting a replacement; the code now does that and reports it as uncertain. This is NOT a third instance of today's two lock-identity defects (4e1bf9aa and its replayed predecessor). Those were wrong liveness predicates: a reused PID read as a live holder, then an exec'd holder read as dead. Here the predicate is right - the code correctly proves the owner dead and refuses the claim anyway, on a condition unrelated to liveness. Regression coverage, each failing on the parent commit for its own reason: - tests/fm-procevent.test.sh: a source that cannot start is reported as failed rather than started; a dead generation whose leftovers cannot be tidied no longer keeps owning its source (the parent reports a start while nothing ever runs); the existing reused-PID fixture now also asserts the orphaned listing and that no doomed relaunch is reported. - tests/fm-captain-hold-lifecycle.test.sh: a board answer reaches the keyed-answer intake through the runner end to end - durable capture, the wake, and the closed task carrying the captain's selection. This one passes on the parent, because that chain was never what broke. fm-procevent 100, fm-bearings-board 18, fm-captain-hold-lifecycle 50, fm-procevent-when 13 and fm-procevent-quota 18 pass; bin/fm-lint.sh and bin/fm-doc-audience-check.sh clean. tests/fm-extension-binding.test.sh has two failures identical on the parent commit (EACCES on package install in this sandbox) and unrelated to this change. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_016gxgshn5jkWJ3GEYWy7vTG * no-mistakes(review): confirm reconcile launches on durable launch stamps * no-mistakes(review): announce stranded sources and refuse bad confirm windows * no-mistakes(review): announce leaderless strands, bound confirm window, fix recovery docs * no-mistakes(review): announce unconfirmed launches once per episode, qualify start reclaim * no-mistakes(review): nonce launch-failed keys, refuse bad window at arm * no-mistakes(review): state only observed launch outcome, shorten episode nonce * no-mistakes(test): assert launch-failed headline not re-delivered, allow recovery wake * no-mistakes(document): docs: cover strand and launch-failure wakes in skill trigger and verification record * no-mistakes(lint): restructure SC2015 chain into explicit if-block * test(watch-triage): fix two timing-exposed defects the pipeline found Both surfaced in the no-mistakes test step on this branch, each failing one full run of tests/fm-watch-triage.test.sh; neither was accepted as a flake to retry past. 1. The new launch-failed delivery test assumed an already-surfaced key never wakes the watcher again. That is false: a fresh watcher legitimately re-surfaces any unacknowledged queue row through its downtime-recovery path ("check: rearm-resurface"), so the assertion failed whenever a re-arm landed between its two checks. The pipeline's own fix tolerated any wake lacking the repeated key's headline; this tightens it to exactly one tolerated reason, by its exact line, with a failure message that names the expectation so a reworded path reads as "the tolerated recovery path changed" rather than as a mystery - and so nobody restores the strict silence check. The positive assertion (a fresh-suffix key is delivered under its own headline) is unchanged. 2. seed_captured_procevent_result retired its source in the gap between the runner publishing its wake and releasing its claim, so retire read the exiting runner's ownership as uncertain and refused ("cannot confirm runner identity"). The fixture and retire path pre-date this branch; the confirm window returns reconcile closer to the moment of capture, which made the gap easier to hit. The fixture now waits, bounded, for the claim release the publish promises, with the reason at the wait. Verified on this head with tasks-axi on PATH: fm-watch-triage 113/113 with no skips, fm-procevent 106/106, fm-captain-hold-lifecycle 50/50, fm-watch-arm 15/15, fm-bearings-board 18/18, fm-procevent-when 13/13, fm-procevent-quota 18/18; bin/fm-lint.sh and bin/fm-doc-audience-check.sh exit 0. First attempt, no retries. * no-mistakes(document): docs: route stranded and launch-failed wakes in skill handling --------- Co-authored-by: Claude Opus 5 --- .agents/skills/process-event-sources/SKILL.md | 20 +- AGENTS.md | 2 +- bin/fm-procevent-lib.sh | 92 +++- bin/fm-procevent.sh | 282 ++++++++++- bin/fm-watch.sh | 60 ++- docs/configuration.md | 39 +- docs/verification/process-event-sources.md | 12 +- tests/fm-captain-hold-lifecycle.test.sh | 71 +++ tests/fm-procevent.test.sh | 472 ++++++++++++++++++ tests/fm-watch-arm.test.sh | 43 ++ tests/fm-watch-triage.test.sh | 117 +++++ 11 files changed, 1177 insertions(+), 33 deletions(-) diff --git a/.agents/skills/process-event-sources/SKILL.md b/.agents/skills/process-event-sources/SKILL.md index 4b2629113a6..a18e7b0eb3f 100644 --- a/.agents/skills/process-event-sources/SKILL.md +++ b/.agents/skills/process-event-sources/SKILL.md @@ -3,8 +3,10 @@ name: process-event-sources description: >- Agent-only procedure for registered process-to-event sources and their wakes. Use before arming a long-polling source firstmate owns, before registering a - deterministic condition->action watch, and on any - `procevent ` check wake. + deterministic condition->action watch, on any + `procevent ` check wake, and on any + `process-event source stranded` or `process-event source failed to start` + check wake. Owns the arming commands, the condition->action eligibility boundary, the durable result read, which wakes must be routed to their adapter instead of acknowledged generically, the handled acknowledgement contract, the one-owner @@ -17,7 +19,7 @@ metadata: # process-event-sources -Load this before arming a long-polling source, before registering a deterministic condition->action watch, and whenever a `check:` wake carries `procevent `. +Load this before arming a long-polling source, before registering a deterministic condition->action watch, whenever a `check:` wake carries `procevent `, and whenever the watcher headlines a `process-event source stranded` or `process-event source failed to start` wake. The runner exists so a blocking external process never holds firstmate's conversational turn. Firstmate registers a source, keeps working, and is woken when that process completes. @@ -31,6 +33,14 @@ For a Lavish review artifact firstmate owns (a live investigating scout should h bin/fm-procevent-lavish.sh arm ``` +Registering a source is not the same fact as listening to it: arming records the source, and a separate runner still has to pick it up. +After arming by hand, confirm `bin/fm-procevent.sh list` reports that source as `live`, and run `bin/fm-procevent.sh reconcile` when it does not. +Reconcile reports every launch that did not prove it took its claim within the confirm window as `failed=` and exits non-zero, so a source that cannot be started says so instead of looking armed, and it wakes you once per failure episode about it because the watcher discards that count; `start` does not fix that - if the source stays unowned, run `start` attached to read the runner's refusal, then check the source command and adapter binary the registration names, and if a later reconcile finds the source owned the episode closes on its own. +A source `list` reports as `orphaned` is one reconcile will not relaunch, because something may still be polling it; reconcile wakes you once about it, and that wake's payload says which of two recoveries applies. +If the claim's recorded pid is alive under a different identity, `bin/fm-procevent.sh start ` takes the source back once you have checked nothing is still polling it - provided the dead generation's reservation records can still be tidied; otherwise it refuses with `cannot claim source`. +If the runner itself died and its process group survives, `start` reports `already owned` and takes nothing back: verify whether the dead runner's polling child is still attached to the source, and once that group is empty the next reconcile reclaims the source on its own. +Nothing signals that group automatically. + When a source carries captain answers to captain-held tasks, bind it BEFORE arming it, so it can never produce an answer that has nowhere to go: ```sh @@ -107,6 +117,10 @@ Two rules the commands cannot enforce for you: : Never append a raw result to a task's status history; that log is a bounded event record, not a payload channel. : A source whose adapter returns a terminal verdict for the captured result has already retired itself, so an ended review needs no cleanup from you and produces no further wake. Retire any other finished source with the adapter's `retire`, which stays safe and idempotent even for one that already retired. Retirement stops future completions; it is independent of acknowledging a result already captured, which only `handled` does. +`process-event source stranded` or `process-event source failed to start` (queue keys `procevent::stranded:` and `procevent::launch-failed:-`) +: Nothing was captured: the source named in the payload is registered but nothing is confirmed to be collecting from it. There is no result file to read and no `handled` call to make; the ordinary drain acknowledgement consumes the row. +: The payload says which shape it is and what clears it. Follow it exactly as the arming section above describes - a `start` is named only for the reused-pid strand, a leaderless group is a human check and reclaims itself once its group is empty, and a launch that never proved its claim closes its own episode if a later cycle finds the source owned. + ## What the runner guarantees, exactly Supported by tests: diff --git a/AGENTS.md b/AGENTS.md index 6e4901b601c..906195aef1e 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -566,7 +566,7 @@ These skills are not captain-invocable; load them only at their precise triggers - `stuck-crewmate-recovery` - load when the session-start digest reports an ordinary direct report's endpoint dead or its metadata has no window, after a stale wake, looping pane, repeated confusion, an answered-by-brief question, an unresponsive crewmate, or a failed steer, and whenever a live worker reports its no-mistakes pipeline dead, unreachable, or timed out. - `secondmate-provisioning` - load before creating, seeding, validating, launching, handing backlog to, recovering, pushing inherited local material into, or retiring a secondmate home, and before editing `data/secondmates.md`. - `captain-hold-lifecycle` - load before treating an investigation or visual review as complete, before ending a visual review that exposed a captain decision, when recording or routing the captain's answer, and on any `RECORD DIVERGENCE` line from the wake drain. -- `process-event-sources` - load before arming a long-polling source, before registering a deterministic condition->action watch (do X as soon as Y is true), and on any `procevent ` check wake. +- `process-event-sources` - load before arming a long-polling source, before registering a deterministic condition->action watch (do X as soon as Y is true), on any `procevent ` check wake, and on any `process-event source stranded` or `process-event source failed to start` check wake. Never run a registered source's blocking command yourself in a conversational turn. - `fmx-respond` - load on an `x-mention ` `check:` wake to handle the mention, on an `x-mode-error ...` `check:` wake to report the Relay configuration blocker, on a `public-followup ...` `check:` wake or a startup-surfaced public commitment, and on any milestone or terminal wake for a Relay-linked task before posting its completion follow-up; relevant only when Relay is on. - `firstmate-codexapp` - load before coordinating a visible Codex Desktop thread, evaluating a Codex App backend request, or reconciling Codex Desktop host-tool smoke evidence for Firstmate work. diff --git a/bin/fm-procevent-lib.sh b/bin/fm-procevent-lib.sh index 266365dd526..f5fce33dee1 100644 --- a/bin/fm-procevent-lib.sh +++ b/bin/fm-procevent-lib.sh @@ -211,22 +211,50 @@ fm_procevent_launch_floor_seconds() { printf '%s\n' "$value" } -fm_procevent_launch_floor_reset_locked() { # +# How long reconcile waits for a runner it just detached to prove it took the +# source's claim. Confirmation reads durable evidence, so a healthy launch +# settles on the first poll and only a launch not yet proved spends the +# window. The default stays well below FM_POLL because bin/fm-watch.sh runs +# reconcile once per supervision cycle, and every launch of a cycle shares ONE +# window rather than taking a window each. +FM_PROCEVENT_LAUNCH_CONFIRM_DEFAULT_SECONDS=3 +FM_PROCEVENT_LAUNCH_CONFIRM_MIN_SECONDS=1 +FM_PROCEVENT_LAUNCH_CONFIRM_MAX_SECONDS=600 + +fm_procevent_launch_confirm_seconds() { + local value=${FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS-} + if [ -z "$value" ]; then + printf '%s\n' "$FM_PROCEVENT_LAUNCH_CONFIRM_DEFAULT_SECONDS" + return 0 + fi + case "$value" in ''|*[!0-9]*) return 1 ;; esac + [ "$value" -ge "$FM_PROCEVENT_LAUNCH_CONFIRM_MIN_SECONDS" ] || return 1 + [ "$value" -le "$FM_PROCEVENT_LAUNCH_CONFIRM_MAX_SECONDS" ] || return 1 + printf '%s\n' "$value" +} + +# The one place the launch-pacing stamp's name is constructed. Every writer, +# pruner and reader goes through here so the naming rule is stated once. +fm_procevent_launch_floor_stamp_path() { # local reg identity case "$3" in *:*) ;; *) return 1 ;; esac case "$3" in ''|*[!0-9:]*) return 1 ;; esac + fm_procevent_source_id_valid "$2" || return 1 reg=$(fm_procevent_registry_dir "$1") || return 1 identity=${3//:/-} - rm -f -- "$reg/$2.$identity.last-launch" + printf '%s\n' "$reg/$2.$identity.last-launch" +} + +fm_procevent_launch_floor_reset_locked() { # + local stamp + stamp=$(fm_procevent_launch_floor_stamp_path "$1" "$2" "$3") || return 1 + rm -f -- "$stamp" } fm_procevent_launch_floor_prune_locked() { # - local reg identity keep stamp - case "$3" in *:*) ;; *) return 1 ;; esac - case "$3" in ''|*[!0-9:]*) return 1 ;; esac + local reg keep stamp + keep=$(fm_procevent_launch_floor_stamp_path "$1" "$2" "$3") || return 1 reg=$(fm_procevent_registry_dir "$1") || return 1 - identity=${3//:/-} - keep="$reg/$2.$identity.last-launch" for stamp in "$reg/$2".*.last-launch "$reg/$2.last-launch"; do [ "$stamp" = "$keep" ] && continue [ -e "$stamp" ] || [ -L "$stamp" ] || continue @@ -235,12 +263,9 @@ fm_procevent_launch_floor_prune_locked() { # - local state=$1 id=$2 expected=$3 floor=$4 reg stamp identity registration current_identity status=0 - case "$expected" in *:*) ;; *) return 1 ;; esac - case "$expected" in ''|*[!0-9:]*) return 1 ;; esac + local state=$1 id=$2 expected=$3 floor=$4 reg stamp registration current_identity status=0 + stamp=$(fm_procevent_launch_floor_stamp_path "$state" "$id" "$expected") || return 1 reg=$(fm_procevent_registry_dir "$state") || return 1 - identity=${expected//:/-} - stamp="$reg/$id.$identity.last-launch" [ ! -L "$stamp" ] || return 1 [ ! -e "$stamp" ] || [ -f "$stamp" ] || return 1 perl -MTime::HiRes=clock_gettime,sleep,CLOCK_MONOTONIC -e ' @@ -607,6 +632,29 @@ fm_procevent_claim_generation_gone_locked() { && ! fm_procevent_group_alive "${FM_PROCEVENT_CLAIM_PID:-}" } +# fm_procevent_claim_undisplaceable_locked +# The single owner of "this stale claim is one no unattended caller may +# displace". True when a claim record is still present for the source and its +# generation is NOT provably gone. Call it only where +# fm_procevent_claim_state_locked has just returned 1, so the FM_PROCEVENT_CLAIM_* +# globals below describe this source: that same return also covers a source with +# no claim record at all, which leaves those globals holding whatever the +# previous load put there, so the record check has to travel with the generation +# check rather than being left to each caller. +# +# What the surviving process group means is why this refuses rather than +# relaunches. fm_procevent_group_alive probes the runner's OWN process group, +# and the runner leads that group with its polling source child inside it, so +# "the group still has members" can mean that child is still attached to the +# session the source collects from. Starting a replacement there puts a second +# destructive poller on one session, which drains and loses what the source was +# collecting. A source that needs a human beats a source that silently eats what +# it was supposed to deliver. +fm_procevent_claim_undisplaceable_locked() { # + [ -e "$(fm_procevent_claim_path "$1")" ] || return 1 + ! fm_procevent_claim_generation_gone_locked +} + # Capture-reservation cleanup for a claim being reclaimed. # # Reservation records are keyed by CLAIM TOKEN, and every replacement claims a @@ -707,6 +755,26 @@ fm_procevent_claim_acquire_locked() { if [ "$status" -eq 0 ]; then fm_procevent_claim_capture_reservation_reclaim_locked || status=1 fi + # Every cleanup above tidies leftovers that belong to the DEAD + # generation - its staging file and its capture reservation, both keyed + # by ITS claim token - and a replacement always claims a fresh token, + # so nothing a failed tidy-up leaves behind can collide with the + # generation that replaces it. + # fm_procevent_claim_capture_reservation_reclaim_locked already states + # that rule for the reservation record; the staging file takes the same + # rule here, and so does the shape check on the registry directory + # recorded to hold it, which only decides whether that removal is safe + # to attempt. Once the stale owner and the + # independently absent process group prove the whole generation gone, + # the documented ownership promise is already granted, so a failed + # tidy-up may leave litter and nothing more. Vetoing the claim instead + # is what leaves a provably dead runner owning the source permanently, + # where no reconcile, no retire and no fresh arm can displace it. + if [ "$status" -ne 0 ] && fm_procevent_claim_generation_gone_locked; then + status=0 + fi + # Two owners is the one outcome worse than none: never proceed on a + # claim record that is still there. [ "$status" -ne 0 ] || rm -f -- "$claim" || status=1 else status=1 diff --git a/bin/fm-procevent.sh b/bin/fm-procevent.sh index 93361b64552..ee31dd8b3be 100755 --- a/bin/fm-procevent.sh +++ b/bin/fm-procevent.sh @@ -47,6 +47,29 @@ # start a runner for any registered source that has no live owner. # This is liveness repair only - it never discovers results by # polling the source, because the child blocks on the source itself. +# A start is REPORTED only once it is confirmed: starting a runner is +# detached and its errors reach no caller, so a source that cannot +# start would otherwise be counted exactly like one that is +# listening, and a wedged source would go on presenting as armed. +# Every launch is counted as `started` only after the source is +# observed owned or its launch-pacing stamp has moved, `failed` +# otherwise, and any failure also makes this command exit non-zero. +# One bounded window covers a whole cycle's launches +# (FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS; docs/configuration.md). +# A launch that fails to confirm is also announced as a durable +# `check` wake, once per failure episode - keyed by the registration +# identity it ran under and ended by a later launch of that source +# confirming - because the supervision cycle discards the `failed=` +# count. The launch itself is retried every cycle exactly as before. +# A source whose claim nothing may automatically displace is not +# relaunched at all; it is counted `uncertain` and announced once per +# stranded claim generation as a durable `check` wake, because the +# supervision cycle discards this command's own output and exit +# status. The wake names what clears that strand: the `start` +# command for a reused pid whose group survives, or the check a +# human makes for a group that lost its leader, which `start` +# reports as owned and which the next cycle reclaims on its own +# once that group is empty. # handled Durably and idempotently record that a captured result has been # fully handled: . Prints "handled: id seq" # the first time for that exact source-and-sequence generation and @@ -346,6 +369,8 @@ adapter_self_announcing() { # source_file() { printf '%s/%s.source\n' "$REG" "$1"; } runner_file() { printf '%s/%s.runner\n' "$REG" "$1"; } staging_file() { printf '%s/.%s.%s.output\n' "$REG" "$1" "$2"; } +stranded_file() { printf '%s/.%s.stranded\n' "$REG" "$1"; } +launch_failed_file() { printf '%s/.%s.launch-failed\n' "$REG" "$1"; } # Let the source's own adapter apply and acknowledge one captured result. See # the header for why this exists and what each exit means. An already @@ -1195,8 +1220,107 @@ detach_runner() { # isolate_runner detach "$1" } +# Announce a source whose claim no unattended caller may displace, once per +# stranded claim generation. +# +# The supervision cycle runs this command with its output and its exit status +# both discarded, so a strand that only shows up in `list` as `orphaned` and in +# this command's `uncertain=` count reaches nobody. A durable `check` wake does +# reach firstmate through the ordinary queue, and it carries what clears the +# strand so acting on it needs no hunt. The caller supplies that part, because +# the two strand shapes clear differently and naming the wrong recovery would +# send someone to a command that reports `already owned` and changes nothing. +# +# The marker records the claim generation that was reported, so the same strand +# never wakes twice while a genuinely new claim still does - an alarm that +# repeats every supervision cycle is as unusable as one nobody gets. It is +# written before the wake and removed again if the wake does not land, so a +# failed announcement retries instead of being silently marked as delivered. +report_stranded_source() { # + local id=$1 token=$2 detail=$3 + case "$token" in ''|*[!A-Za-z0-9._-]*) return 1 ;; esac + [ -n "$detail" ] || return 1 + announce_source_once "$(stranded_file "$id")" "$token" \ + "procevent:$id:stranded:$token" \ + "check: process-event source $id is registered but nothing can arm it: $detail" +} + +# Announce a launch that reconcile could not confirm, once per failure episode. +# +# A launch that never proves it took the claim - a runner that died before +# claiming on unreadable argv, a missing adapter binary or a guard that refused +# to start, or one merely too slow under load - is relaunched every supervision +# cycle and reported `failed=` to a stdout that cycle discards: armed in +# appearance, a dead drop in fact, which is the incident with a different cause. +# Confirmation observes only that no claim and no launch stamp appeared inside +# the window, so this says exactly that and no more about why. An episode is +# keyed by the registration identity the launch ran under and ends when a later +# cycle finds the source owned or a launch confirms, so a second failure inside +# one episode announces nothing, a slow runner that arms later closes its own +# episode without a retraction, and a source that recovers and then fails again +# announces a new one. Nothing here changes what reconcile does about the launch +# itself: it keeps relaunching exactly as before, and this only says so once. +# +# The queue key carries a nonce beyond the episode: the watcher remembers every +# key it has surfaced for good, so a key made of the registration identity alone +# would be surfaced for the first episode only and every later episode of the +# same registration would sit in the queue unannounced. The marker records the +# episode and that nonce together, and the episode alone decides whether to +# announce. +report_launch_failure() { # + local id=$1 identity=$2 episode nonce + case "$identity" in ''|*[!0-9:]*) episode=unreadable ;; *) episode=${identity//:/-} ;; esac + nonce="$RANDOM$RANDOM" + announce_source_once "$(launch_failed_file "$id")" "$episode" \ + "procevent:$id:launch-failed:$episode-$nonce" \ + "check: process-event source $id is registered but its launch did not prove it took the source's claim within FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS, so nothing is confirmed to be collecting from it; reconcile reports that as failed= and keeps launching it every supervision cycle. If it stays that way, check the source command and the adapter binary the registration names, and run an attached bin/fm-procevent.sh start $id to reproduce a refusal on its stderr - the detached launch discards it, and a hand-run reconcile only counts it as failed=. A later cycle that finds the source owned ends this episode on its own, so a runner that was merely slow to claim needs nothing from you." \ + "$episode $nonce" +} + +# Shared marker discipline for the announcements above: holds the +# generation last reported as its first field, written before the wake and +# removed again if the wake does not land, so a failed announcement retries +# instead of being marked delivered, and the same generation never announces +# twice. A caller may store more after that field (the launch-failure nonce); +# only the first field decides. +announce_source_once() { # [marker-record] + local marker=$1 generation=$2 key=$3 payload=$4 record=${5:-$2} previous + previous=$(cat -- "$marker" 2>/dev/null || true) + [ "${previous%%[[:space:]]*}" != "$generation" ] || return 1 + (umask 077; printf '%s\n' "$record" > "$marker") || return 1 + if ! fm_wake_append check "$key" "$payload"; then + rm -f -- "$marker" + return 1 + fi + return 0 +} + +# The reused-pid strand: the recorded pid is alive under a different identity +# while the runner's process group still has members. The claim path does not +# consult the process group, so a deliberate `start` reclaims this - provided +# the dead generation's reservation records can still be tidied, because that +# tidy-up is only waived for a generation proven gone, and this one is not. +stranded_reused_pid_detail() { # + printf '%s' "its claim names a dead runner whose process group still has members, so reconcile preserves that claim and starts no replacement. Check that nothing is still polling the source, then reclaim it with: bin/fm-procevent.sh start $1 - that reclaims it provided the dead generation's reservation records can still be tidied, and otherwise refuses with: cannot claim source" +} + +# The leaderless strand: the runner leader is gone and its group still has +# members. `start` reports this as owned and reclaims nothing, and nothing +# automatic signals that group, so the only honest recovery to name is the +# check a human makes; an empty group reads as gone on the next cycle. +stranded_leaderless_detail() { # + printf '%s' "its runner died and its polling child may still be attached to the source's session, so reconcile preserves that claim and starts no replacement, and nothing automatic will touch that group. Verify whether anything is still polling $1; once that process group is empty, the next reconcile reclaims the source on its own." +} + cmd_reconcile() { - local rec id published started=0 stopped=0 uncertain=0 claim owner pid token identity claim_state stop_state + local rec id published started=0 stopped=0 uncertain=0 failed=0 claim owner pid token identity claim_state stop_state + local launch_identity launch_stamp launch_mark unconfirmed entry + local -a launched=() + # Rejected before anything is launched, and by name. A window this command + # cannot use makes every launch unconfirmable, so validating it later would + # report a fleet of perfectly healthy runners as `failed=` and blame nothing. + fm_procevent_launch_confirm_seconds >/dev/null \ + || die "FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS must be whole seconds from $FM_PROCEVENT_LAUNCH_CONFIRM_MIN_SECONDS to $FM_PROCEVENT_LAUNCH_CONFIRM_MAX_SECONDS" owner_lease_refresh published=$(publish_pending) @@ -1251,15 +1375,39 @@ cmd_reconcile() { if [ -f "$(source_file "$id")" ] && [ ! -L "$(source_file "$id")" ]; then fm_procevent_claim_state_locked "$id" claim_state=$? - if [ "$claim_state" -eq 1 ]; then + if [ "$claim_state" -eq 1 ] && fm_procevent_claim_undisplaceable_locked "$id"; then + # A stale claim whose process group still has members, which can mean + # the dead runner's polling child is still on the source's session + # (fm_procevent_claim_undisplaceable_locked owns that reasoning). + # Preserve the claim, start nothing, and say the cycle could not + # settle it, which is what this command already promises for the + # leaderless variant below. Only a deliberate `start` reclaims here, + # so report the strand durably rather than leaving it to whoever + # happens to run this command. + uncertain=$((uncertain + 1)) + report_stranded_source "$id" "$FM_PROCEVENT_CLAIM_TOKEN" \ + "$(stranded_reused_pid_detail "$id")" || true + elif [ "$claim_state" -eq 1 ]; then if ! cleanup_extension_registration_invocations_locked "$id"; then uncertain=$((uncertain + 1)) fm_procevent_source_lock_release "$id" continue fi + # Snapshot the launch-pacing stamp for the registration generation + # this launch will run under, while the source lock still keeps that + # registration from being replaced underneath it. The runner writes + # this stamp after it claims and before it runs the source command, + # and nothing removes it on the way out, so an advanced or newly + # appeared value is durable evidence the launch got going. + launch_identity=$(fm_pr_file_identity "$(source_file "$id")" 2>/dev/null) || launch_identity= + launch_mark= + if [ -n "$launch_identity" ] \ + && launch_stamp=$(fm_procevent_launch_floor_stamp_path "$STATE" "$id" "$launch_identity"); then + launch_mark=$(cat -- "$launch_stamp" 2>/dev/null || true) + fi fm_procevent_source_lock_release "$id" detach_runner "$id" - started=$((started + 1)) + launched+=("$id"$'\t'"$launch_identity"$'\t'"$launch_mark") continue elif [ "$claim_state" -eq 4 ]; then owner=$FM_PROCEVENT_CLAIM_HOME @@ -1277,15 +1425,119 @@ cmd_reconcile() { elif [ "$claim_state" -eq 3 ]; then # A leaderless group's generation is ambiguous under PID/PGID reuse, # so preserve its claim without signalling or starting a replacement. + # This is the ordinary crash shape, and `start` cannot clear it + # either, so it is announced the same way as the reused-pid strand + # above but naming what a human should check rather than a command. uncertain=$((uncertain + 1)) + report_stranded_source "$id" "$FM_PROCEVENT_CLAIM_TOKEN" \ + "$(stranded_leaderless_detail "$id")" || true elif [ "$claim_state" -eq 2 ]; then uncertain=$((uncertain + 1)) + elif [ "$claim_state" -eq 0 ]; then + # A live owner is the same evidence confirmation reads, however the + # runner was started, so it ends any launch-failure episode here. + rm -f -- "$(launch_failed_file "$id")" fi fi fm_procevent_source_lock_release "$id" done fi - printf 'reconciled: published=%s started=%s stopped=%s uncertain=%s\n' "$published" "$started" "$stopped" "$uncertain" + if [ "${#launched[@]}" -gt 0 ]; then + unconfirmed=$(confirm_launched_runners "${launched[@]}") \ + || unconfirmed=$(printf '%s\n' "${launched[@]}") + for entry in "${launched[@]}"; do + id=${entry%%$'\t'*} + launch_identity=${entry#*$'\t'} + launch_identity=${launch_identity%%$'\t'*} + if launch_entry_listed "$entry" "$unconfirmed"; then + failed=$((failed + 1)) + report_launch_failure "$id" "$launch_identity" || true + else + started=$((started + 1)) + rm -f -- "$(launch_failed_file "$id")" + fi + done + fi + printf 'reconciled: published=%s started=%s stopped=%s uncertain=%s failed=%s\n' \ + "$published" "$started" "$stopped" "$uncertain" "$failed" + [ "$failed" -eq 0 ] +} + +launch_entry_listed() { # + local entry=$1 line + while IFS= read -r line; do + [ "$line" = "$entry" ] && return 0 + done <<< "$2" + return 1 +} + +# Bounded confirmation that every runner just detached actually took its +# source's claim, printing every launch entry that did not, one per line. +# +# detach_runner is fire-and-forget and discards the child's stderr, so before +# this every failure inside _start - a refused claim above all - was still +# counted and reported as a start. That made a source that CANNOT start +# indistinguishable from one that had, which is exactly how a wedged review +# board goes on presenting as armed while collecting nothing. +# +# Two signals confirm a launch, and each covers what the other cannot see: +# ownership covers the runner still blocked on its source, which is the only +# evidence such a runner ever shows; the launch-pacing stamp covers the runner +# that claimed, ran and exited between two polls, because the runner writes that +# stamp after claiming and before running the source command and nothing removes +# it on the way out - only registration replacement does, which also changes the +# snapshotted identity this reads under. A runner that dies BEFORE claiming +# reaches neither, and that is the case this confirmation exists to catch; a +# runner merely slow to claim looks the same inside the window, which is why +# the failure this reports is "not proved within the window" and nothing more. +# +# Every launch shares ONE window rather than taking a window each, so a whole +# fleet of failing sources costs a watcher cycle the same bounded wait as one. +confirm_launched_runners() { # ... + local deadline window entry id rest identity before state stamp mark + local -a pending=("$@") remaining=() + window=$(fm_procevent_launch_confirm_seconds) || return 1 + # A zero-padded window is a valid value to its validator, which reads base 10; + # reading it as octal here would silently shorten the window or abort this + # subshell under `set -u` and report every launch as failed. + # SECONDS is an integer clock that can tick at any moment after this + # assignment, so a deadline of exactly SECONDS + window waits anywhere in + # [window - 1, window] and a healthy launch could be reported failed for + # losing a second it was promised. The extra second bounds the wait to + # [window, window + 1] instead: never less than configured. + deadline=$((SECONDS + 10#$window + 1)) + while :; do + remaining=() + for entry in "${pending[@]+"${pending[@]}"}"; do + id=${entry%%$'\t'*} + rest=${entry#*$'\t'} + identity=${rest%%$'\t'*} + before=${rest#*$'\t'} + state=1 + if fm_procevent_source_lock_try_acquire "$id"; then + fm_procevent_claim_state_locked "$id" + state=$? + fm_procevent_source_lock_release "$id" + fi + if [ "$state" -eq 0 ]; then + continue + fi + mark= + if [ -n "$identity" ] \ + && stamp=$(fm_procevent_launch_floor_stamp_path "$STATE" "$id" "$identity"); then + mark=$(cat -- "$stamp" 2>/dev/null || true) + fi + if [ -n "$mark" ] && [ "$mark" != "$before" ]; then + continue + fi + remaining+=("$entry") + done + pending=("${remaining[@]+"${remaining[@]}"}") + [ "${#pending[@]}" -gt 0 ] || break + [ "$SECONDS" -lt "$deadline" ] || break + sleep 0.05 + done + [ "${#pending[@]}" -eq 0 ] || printf '%s\n' "${pending[@]}" } # Stop a runner and the child it is blocked on. A runner started by reconcile is @@ -1489,6 +1741,8 @@ cmd_retire() { fi rm -f -- "$(source_file "$id")" rm -f -- "$(runner_file "$id")" + rm -f -- "$(stranded_file "$id")" + rm -f -- "$(launch_failed_file "$id")" fm_procevent_source_lock_release "$id" # A retired source produces no further answer, so drop any decision binding it # carried. Generic and idempotent: the binding owner is asked to forget this @@ -1653,7 +1907,7 @@ cmd_sweep_home() { } cmd_list() { - local rec id adapter owner pending + local rec id adapter owner pending claim_state owner_lease_refresh if ! fm_procevent_any_registered "$STATE"; then printf 'no sources registered\n' @@ -1666,7 +1920,23 @@ cmd_list() { adapter=$(read_adapter "$id" 2>/dev/null || echo '?') fm_procevent_source_lock_acquire "$id" || continue fm_procevent_claim_state_locked "$id" - case "$?" in 0) owner=live ;; 1) owner=none ;; 3) owner=orphaned ;; *) owner=uncertain ;; esac + claim_state=$? + # A stale claim whose process group still has members is exactly as + # undisplaceable as the leaderless group state 3 already reports, and a + # reused PID reaches it through state 1 rather than state 3. Reporting that + # as `none` reads like an idle source waiting to be started, which is the + # reassuring answer this whole surface gave while a board collected nothing. + case "$claim_state" in + 0) owner=live ;; + 1) + owner=none + if fm_procevent_claim_undisplaceable_locked "$id"; then + owner=orphaned + fi + ;; + 3) owner=orphaned ;; + *) owner=uncertain ;; + esac fm_procevent_source_lock_release "$id" pending=$(fm_procevent_pending "$STATE" | grep -c "/$id\." || true) printf '%-28s %-12s %-10s %s\n' "$id" "$adapter" "$owner" "$pending" diff --git a/bin/fm-watch.sh b/bin/fm-watch.sh index 218405d2142..2a8e02b735a 100755 --- a/bin/fm-watch.sh +++ b/bin/fm-watch.sh @@ -76,6 +76,22 @@ # and has not been surfaced yet; reported once per # captured generation, never again while that record # stays queued and never once it is acknowledged +# check: process-event source stranded: +# a registered process-to-event source has a claim +# reconcile will not displace and nothing collecting +# for it (bin/fm-procevent.sh reconcile queues it +# once per stranded claim generation); the queued +# payload names what clears it +# check: process-event source failed to start: +# a registered process-to-event source was launched by +# reconcile and did not prove it took the claim within +# the confirm window, so nothing is confirmed to be +# collecting for it and every cycle will relaunch it +# (bin/fm-procevent.sh reconcile queues it once per +# failure episode, and a later cycle that finds the +# source owned closes that episode); the queued +# payload names what to check. These three kinds are +# joined with `;` when more than one surfaces in a cycle # check: rejected unauthenticated state checks: # unsafe state checks were refused without execution # check: rejected unauthenticated PR poll retirement receipts: @@ -116,6 +132,10 @@ mkdir -p "$STATE" . "$SCRIPT_DIR/fm-push-transition-lib.sh" # shellcheck source=bin/fm-pr-lib.sh . "$SCRIPT_DIR/fm-pr-lib.sh" +# Only for the arm-time check on FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS below; +# the per-cycle reconcile itself runs as a separate process. +# shellcheck source=bin/fm-procevent-lib.sh +. "$SCRIPT_DIR/fm-procevent-lib.sh" # Single owner of durable merge-outcome publication, shared with # bin/fm-pr-merge.sh so self and poll origins use the same role-routed outcome. # The watcher still owns immediate delivery of its actionable poll result and @@ -1401,7 +1421,7 @@ procevent_surface_after_output() { } procevent_surface_queued() { - local key reason + local key reason captured="" stranded="" unstarted="" PROCEVENT_SURFACED= [ -s "$FM_WAKE_QUEUE" ] || return 0 fm_lock_acquire_wait "$FM_WAKE_QUEUE_LOCK" @@ -1409,12 +1429,30 @@ procevent_surface_queued() { case "$key" in procevent:*) ;; *) continue ;; esac [ -e "$(procevent_surfaced_marker "$key")" ] && continue PROCEVENT_SURFACED="$PROCEVENT_SURFACED $key" + # A stranded source or one whose launch never proved itself is the opposite + # of a captured result: nothing is collecting for it. Headlining either as + # a capture would present it as healthy, which is the shape of defect + # these wakes exist to surface. + case "$key" in + procevent:*:stranded:*) stranded="$stranded $key" ;; + procevent:*:launch-failed:*) unstarted="$unstarted $key" ;; + *) captured="$captured $key" ;; + esac done < <(fm_wake_queued_keys_locked check) if [ -z "$PROCEVENT_SURFACED" ]; then fm_lock_release "$FM_WAKE_QUEUE_LOCK" return 0 fi - reason="check: process-event result captured:$PROCEVENT_SURFACED" + reason="check:" + [ -z "$captured" ] || reason="$reason process-event result captured:$captured" + if [ -n "$stranded" ]; then + [ "$reason" = "check:" ] || reason="$reason;" + reason="$reason process-event source stranded:$stranded" + fi + if [ -n "$unstarted" ]; then + [ "$reason" = "check:" ] || reason="$reason;" + reason="$reason process-event source failed to start:$unstarted" + fi # shellcheck disable=SC2034 # Consumed by wake() in the separately linted transition owner. FM_WAKE_POST_OUTPUT_ACTION=procevent_surface_after_output wake "$reason" @@ -1690,6 +1728,24 @@ if [ "${BASH_SOURCE[0]}" != "$0" ]; then return 0 fi +# FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS is validated here, at arm time, and an +# unusable value refuses to arm. This is deliberately NOT symmetry with the +# tunables above, which this watcher only defaults and never validates. The +# reason is specific: every supervision cycle runs `fm-procevent.sh reconcile` +# with its output and exit status discarded, and reconcile refuses an unusable +# window by name before it launches anything. Under this watcher that refusal +# is invisible - every cycle would exit early, no source would ever start, and +# the whole home would sit disarmed while presenting as supervised. A watcher +# that refuses to arm is loud through an existing, independent, proven path: +# the liveness guard's WATCHER DOWN banner in firstmate's own session. The +# message shape is reconcile's own, so the operator reads one refusal in both +# places. The refusal goes to stdout because bin/fm-watch-arm.sh relays the +# child's stdout and recognises `watcher: FAILED` as the typed failure line. +if ! fm_procevent_launch_confirm_seconds >/dev/null; then + echo "watcher: FAILED - FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS must be whole seconds from $FM_PROCEVENT_LAUNCH_CONFIRM_MIN_SECONDS to $FM_PROCEVENT_LAUNCH_CONFIRM_MAX_SECONDS" + exit 1 +fi + if ! fm_lock_try_acquire "$WATCH_LOCK"; then BEAT="$STATE/.last-watcher-beat" if [ -n "${FM_LOCK_HELD_PID:-}" ]; then diff --git a/docs/configuration.md b/docs/configuration.md index eb79f0a64b6..8bcc89448b1 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -833,11 +833,19 @@ A live identity-matched owner is never displaced, and release removes only the e Every stop proves ownership before its first signal: the live runner's recorded process identity must match and it must still lead its process group. Once that stop has proved ownership and sent TERM, its own escalation to KILL checks only whether the proved group still has members; it does not re-read the leader's identity or group membership, which can change or become unreadable as TERM ends the leader. This proof belongs only to that stop's own escalation and cannot authorize another caller that encounters an unproved group. -A claim counts as reclaimable only when its owner is stale and an independent process-group check finds no members; a crashed leader or reused pid whose process group still has members cannot relax ownership cleanup, so reconcile preserves the claim without signalling the ambiguous group or starting a replacement. -If the leader dies to anything other than the stop's own signal, `retire`, `reconcile`, `sweep-home`, and the guard all refuse its surviving group permanently, and the source silently stops listening. -Whether that group may ever be signalled remains an open decision; the repaired guard does not close this gap. -Reclaiming a generation that IS gone is not gated on tidying its capture-reservation records. -Those records are keyed by claim token and every replacement claims a fresh one, so a leftover that can no longer be located - a state-root identity a claim recorded before its home was re-created, for example - is stale bytes rather than an ownership hazard. +A stale claim whose process group still has members is one `reconcile` never displaces, and the two shapes it comes in recover differently. +`reconcile` preserves such a claim without signalling the ambiguous group or starting a replacement: the group check probes the runner's own process group, which contains its polling source child, so surviving members can mean that child is still attached to the session the source collects from, and a replacement would put a second destructive poller on it. +`list` reports both shapes as `orphaned`. +When the recorded pid is alive under a different identity while the group still has members, the claim boundary itself does not consult the process group, so `bin/fm-procevent.sh start ` reclaims that claim provided the dead generation's reservation records can still be tidied, and otherwise refuses with `cannot claim source`; that tidy-up is waived only for a generation proven gone, which this one is not. +That hand-run command is the recovery path, taken by someone who has checked that nothing is still polling the source. +That asymmetry between the automatic path and the deliberate one is the design rather than an inconsistency, and it is not a claim-level invariant: nothing below `reconcile` enforces it. +When the leader itself is gone and its group still has members - the leader died to anything other than the stop's own signal - `start` does not reclaim the claim either: it reports `already owned` and changes nothing, and `retire`, `reconcile`, `sweep-home`, and the guard all refuse the surviving group permanently, so the source stops listening. +Recovery there is a human verifying whether the dead runner's polling child is still attached to the source; once that process group is empty the generation reads as gone and the next `reconcile` reclaims the source on its own. +Nothing automatic signals that group, and whether it may ever be signalled remains an open decision; the repaired guard does not close this gap. +Neither shape stops listening quietly: the first `reconcile` that strands a claim generation publishes a durable `check` wake naming the source and what clears it - the `start` command for the reused pid, the check to make for the leaderless group - and later cycles stay silent for that same generation while a genuinely new stranded claim announces again. +Reclaiming a generation that IS gone is not gated on tidying anything that generation left behind: its capture-reservation records, its staging file, or the registry directory a claim recorded for them. +Every one of those is keyed by claim token and every replacement claims a fresh one, so a leftover that can no longer be located or removed - a state-root identity a claim recorded before its home was re-created, or a recorded registry directory that no longer resolves to a directory - is stale bytes rather than an ownership hazard. +Making any of them a precondition is what leaves a provably dead runner owning its source permanently, because none of those conditions clears on its own. Ordinary release and reclamation still attempt reservation cleanup and require it unless both owner staleness and whole-group absence prove the generation gone. The narrow live-owner terminal-self-retirement path also attempts cleanup but tolerates its own still-in-flight reservation, which the runner removes on the normal end-of-capture path; exact home, PID, and claim-token ownership remains mandatory before the claim is released. If identity cannot be established before the first signal, or a surviving owned group cannot be proved stopped, the operation preserves the registration and claim for safe retry rather than adding a second owner. @@ -876,6 +884,26 @@ Scope is the owning state root and one runner generation, never a script or proc `FM_PROCEVENT_LAUNCH_FLOOR_SECONDS` (default 1, range 1..3600) is the minimum time between consecutive launches of one registration generation's stored command, bounding the launch rate of an immediately returning source during that lease window. The generation's first launch is immediate, later launches share its monotonic pacing timestamp, a timestamp from before a reboot is treated as expired, and replacing the registration starts a fresh pacing generation. +`FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS` (default 3, range 1..600) bounds how long `reconcile` waits for the runners it just started to prove they are running: never less than the configured value, and at most one second more, because the wait is measured on a whole-second clock. +Starting a runner is detached and its errors are not visible to the caller, so `reconcile` reports a start only after the source is observed owned or its launch-pacing stamp has advanced or appeared, and reports every unconfirmed launch as `failed=` and a non-zero exit instead. +Both signals are durable evidence a runner claimed: ownership is the only evidence a runner still blocked on its source ever shows, and the stamp - written after the claim and before the source command runs, and removed only by registration replacement - covers a runner that claimed, ran and exited between two polls. +A healthy launch therefore confirms on the first poll and the window only bounds a launch that has not yet proved itself - one that died before claiming, or one merely too slow to claim inside the window; confirmation cannot tell those apart, and a launch that proves itself on a later cycle closes its failure episode without a retraction wake. +All of a cycle's launches share one window, so a home full of sources that cannot start costs the same bounded wait as one. + +Keep this window well below `FM_POLL`. +`bin/fm-watch.sh` runs `reconcile` once per supervision cycle, so a source that cannot start makes every cycle wait up to the confirm window before the rest of that cycle runs. +Raising the confirm window lengthens every supervision cycle and delays wake delivery by up to that much. + +A source that can never start is reported as `failed=` with a non-zero exit on every `reconcile`, rather than counted as `started` and retried silently as though it were healthy, so a wedged source stays visible instead of presenting as armed. +That count reaches only whoever runs the command, because `bin/fm-watch.sh` discards `reconcile`'s output and exit status, so an unconfirmed launch is also announced through the wake queue: `reconcile` publishes a durable `check` wake (`procevent::launch-failed:-`) once per failure episode, and later cycles stay silent for that episode until a launch of that source confirms, after which a fresh failure announces again under a fresh key, because the watcher never re-surfaces a key it has already surfaced. +The announcement changes nothing about the launch: `reconcile` keeps relaunching the source every cycle exactly as before, and nothing is retried differently, throttled, or recovered from that signal. +The wake says only what was observed for that shape - the launch did not prove it took the claim within the window - and, if it stays that way, names the source command and adapter binary the registration names as what to check and the attached `bin/fm-procevent.sh start ` as what reproduces a refusal on stderr, where the detached launch discards it; a later cycle that finds the source owned ends the episode on its own, so a runner that was merely slow to claim needs nothing from the operator. +A source stranded on a claim nothing may automatically displace is announced the same way, once per stranded claim generation, as described above. +`bin/fm-watch.sh` surfaces both under their own headlines - `process-event source stranded` and `process-event source failed to start` - rather than as a captured result. + +A value this command cannot use is refused by name before anything is launched, the same way `FM_PROCEVENT_LAUNCH_FLOOR_SECONDS` and `FM_PROCEVENT_MAX_OUTPUT_BYTES` are refused, so a mistyped window can never present as a fleet of sources that cannot start. +`bin/fm-watch.sh` validates the same value when it arms and refuses to arm on an unusable one, naming the variable and the range: under a running watcher that refusal would otherwise repeat on every cycle into a discarded stdout and leave the whole home disarmed while presenting as supervised, whereas a watcher that will not arm is loud through the liveness guard. + `FM_PROCEVENT_MAX_OUTPUT_BYTES` (default 1048576) bounds a single captured result while the source runs; oversized output is drained but truncated with a stderr notice rather than staged or published whole or dropped. The runner proves exactly one durability boundary: output that reached the runner is stored at mode `0600` before any event referencing it is published, and a captured result with no durable handled acknowledgement remains eligible for bounded re-announcement across any number of drains and restarts, not only the crash window right after capture. @@ -969,6 +997,7 @@ FM_PROCEVENT_CLAIM_ROOT= # machine-wide source claim root; defaul FM_PROCEVENT_OWNER_LEASE_SECONDS=600 # how long a source runner keeps going with no activity in its owning home; 1..86400 FM_PROCEVENT_OWNER_CHECK_SECONDS=15 # a runner guard's detection interval, read twice per interval; 1..3600 FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=1 # minimum interval between launches of one registration generation's source command; 1..3600 +FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS=3 # how long reconcile waits for the runners it started to prove they are running; 1..600, keep well below FM_POLL FM_WHEN_OUTPUT_TAIL_BYTES=8192 # bound on the command-output tail inside one condition->action outcome document FM_CODEX_WATCH_CHECKPOINT=180 # seconds per foreground watcher checkpoint in Codex primary supervision FM_CREW_STATE_NM_TIMEOUT=10 # seconds allowed per no-mistakes query inside fm-crew-state.sh diff --git a/docs/verification/process-event-sources.md b/docs/verification/process-event-sources.md index 524676dab71..7e7c4312a28 100644 --- a/docs/verification/process-event-sources.md +++ b/docs/verification/process-event-sources.md @@ -86,7 +86,7 @@ Never at-least-once, no-loss, or lossless. ## What the runner does prove -Exercised by `tests/fm-procevent.test.sh` against a fake blocking source whose completion is a process event, not a timer; for the two supervision-delivery rows below, by `tests/fm-watch-triage.test.sh` driving a real `bin/fm-watch.sh` over a real capture; and for adapter-owned application, by `tests/fm-remote-reply.test.sh` driving the real remote-reply relay end to end in an isolated home: +Exercised by `tests/fm-procevent.test.sh` against a fake blocking source whose completion is a process event, not a timer; for the supervision-delivery and headline rows below, by `tests/fm-watch-triage.test.sh` driving a real `bin/fm-watch.sh` over a real capture and over queued strand and launch-failure keys, with `tests/fm-watch-arm.test.sh` covering the arm-time refusal; and for adapter-owned application, by `tests/fm-remote-reply.test.sh` driving the real remote-reply relay end to end in an isolated home: | Guarantee | How it is proven | | --- | --- | @@ -117,9 +117,13 @@ Exercised by `tests/fm-procevent.test.sh` against a fake blocking source whose c | attached owner continuity | a foreground `start` with a one-second lease remains alive beyond that lease while its caller stays attached, then captures normally when the blocking source completes | | owner-home lifetime and scope | a detached runner and its spawning descendant are observed reparented before an expired owner lease stops their whole process group and process churn; replacing the state directory at the same path cannot keep the old runner alive with a new lease because its recorded device/inode no longer matches, while an identical runner in an unchanged home whose reconcile cycle keeps its lease fresh remains alive | | launch pacing during owner-loss grace | an immediately returning source that attempts detached self-relaunches is held to the configured minimum interval between command launches and remains bounded until its expired owner lease stops the generation; replacement starts a fresh pacing generation, prunes prior pacing state, and prevents a superseded sleeping runner from recreating it | -| stale reclaim without displacement | concurrent contenders replacing one stale claim start exactly one runner, cross-home replacement removes the old generation's staging file from its recorded state directory, and a generation whose stale owner and independently empty process group prove it gone remains reclaimable when its recorded state-root identity can no longer be revalidated | -| crashed leader with a live group | `SIGKILL` on only the runner leader leaves its blocking child group alive; reconcile treats that leaderless group as ambiguous, preserves its claim without starting a replacement, and still reclaims a generation with no leader and no surviving group | -| PID-reuse safety | retirement refuses a live PID whose identity differs from the claim before signalling, and a surviving process group prevents stale-generation cleanup on both ordinary and failed reservation-removal paths | +| stale reclaim without displacement | concurrent contenders replacing one stale claim start exactly one runner, cross-home replacement removes the old generation's staging file from its recorded state directory, and a generation whose stale owner and independently empty process group prove it gone remains reclaimable when its recorded state-root identity can no longer be revalidated or its recorded registry directory no longer resolves to a directory, so `reconcile` reclaims it once, the replacement runs the source, and later cycles report nothing to do | +| confirmed launches only | `reconcile` counts a launch as `started` only after the source is observed owned or its launch-pacing stamp has moved: a registration that cannot start is reported `failed=` with a non-zero exit and its source still listed `none`, a source that claimed, ran and exited before confirmation looked is still `started`, a zero-padded confirm window reads as base 10, and an unusable `FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS` is refused by name before any runner is launched | +| launch failure announced once per episode | an unconfirmed launch queues one `check` wake keyed by source, registration identity and an episode nonce; a second failure in the same episode queues nothing, a confirmed launch queues no failure and closes the episode, a later failure opens a new episode under a fresh key, and a 64-character source id keeps that key within the watcher's marker bound | +| crashed leader with a live group | `SIGKILL` on only the runner leader leaves its blocking child group alive; reconcile treats that leaderless group as ambiguous, preserves its claim without starting or signalling anything, `start` runs nothing beside it, the strand is queued as one `check` wake keyed by source and claim token that a second cycle does not repeat, and reconcile still reclaims a generation with no leader and no surviving group | +| reused pid with a live group | a stale claim whose recorded pid is alive under a different identity while its process group still has members is listed `orphaned`, is never relaunched by `reconcile` across cycles, is announced once naming the `start` command that clears it, and `start` reclaims it while the dead generation's leftovers can be tidied and refuses with `cannot claim source`, replacing nothing, when they cannot | +| strand and failure headlines | a real `bin/fm-watch.sh` surfaces queued `stranded` and `launch-failed` keys under `process-event source stranded` and `process-event source failed to start` rather than as a captured result, joins a mixed cycle's headlines, never re-delivers a key it has already surfaced, and delivers each new failure episode; `bin/fm-watch-arm.sh` refuses to arm on an unusable confirm window, naming the variable and range, with no beacon and no running watcher | +| PID-reuse safety | retirement refuses a live PID whose identity differs from the claim before signalling, and a surviving process group keeps `reconcile` and `retire` from cleaning up the stale generation on both ordinary and failed reservation-removal paths; the reused-pid row above owns what a deliberate `start` does there | | coherent ownership reads | a claim replacement held inside the source boundary blocks `list` until one complete generation is visible | | retire-start exclusion | a queued start revalidates registration after the serialized retirement boundary and executes no child | | uncertain identity before the first signal | a live owner whose identity probe transiently fails is not signaled or released, and its registration remains for retry | diff --git a/tests/fm-captain-hold-lifecycle.test.sh b/tests/fm-captain-hold-lifecycle.test.sh index cac69503b6a..fcc3b923e1d 100755 --- a/tests/fm-captain-hold-lifecycle.test.sh +++ b/tests/fm-captain-hold-lifecycle.test.sh @@ -47,6 +47,19 @@ run_lavish() { # "$ROOT/bin/fm-procevent-lavish.sh" "$@" } +# The generic process-event runner, run against this suite's isolated home and +# its own claim root, so a review armed here can never contend with a real one. +run_procevent() { # + local home=$1 + shift + PATH="$home/fakebin:$PATH" REAL_TASKS_AXI="$TASKS_AXI_BIN" \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$home" \ + FM_STATE_OVERRIDE="$home/state" FM_DATA_OVERRIDE="$home/data" \ + FM_CONFIG_OVERRIDE="$home/config" \ + FM_PROCEVENT_CLAIM_ROOT="$home/procevent-claims" \ + "$ROOT/bin/fm-procevent.sh" "$@" +} + run_bearings() { # [extra args] local home=$1 shift @@ -2158,6 +2171,63 @@ test_legacy_identities_keep_working() { pass "legacy identities, metadata, bindings, and the shim keep working" } +# A board answer must reach the keyed-answer intake through the RUNNER, not just +# through a hand-fed `answers` call. The captain answered ten calls on a bearings +# board, the board accepted them, and nothing collected them: the source that +# collects a board is a supervised process, and while it was not running the +# board went on presenting as armed. Everything between the captured result and +# the closed task is asserted here end to end - capture, the wake that tells +# firstmate to look, and the recorded answer - because each of those was intact +# on its own while the chain as a whole delivered nothing. +test_board_answer_reaches_the_keyed_answer_intake() { + local home sid stub out queue show + home=$(make_home board-channel) + sid=lavish-b0a4d0000000f1e2 + fm_test_track_procevent_home "$home" "$home/procevent-claims" + + run_captain "$home" hold sample-board-call --title "Choose the sample board route" \ + --reason "captain board route choice pending" --repo sample >/dev/null \ + || fail "could not register the board call" + + # One published Lavish poll response carrying the captain's structured answer, + # in the shape the adapter's own reader parses: a declared field order, an + # indented CSV row, and the versioned answer context inside its prompt. + stub="$home/board-source.sh" + cat > "$stub" <<'SH' +#!/usr/bin/env bash +cat <<'OUT' +session: + status: feedback + session_ended: false +prompts[1]{tag,text,prompt}: + "choice","Take the north route","Context data: {\"schema\":\"fm-bearings-answer.v1\",\"question\":\"sample-board-call\",\"selection\":\"north\",\"note\":\"\"}" +OUT +SH + chmod +x "$stub" + + run_procevent "$home" register lavish "$sid" -- "$stub" >/dev/null \ + || fail "could not register the board source" + run_captain "$home" bind "$sid" >/dev/null \ + || fail "could not bind the board source to the keyed-answer intake" + + out=$(run_procevent "$home" start "$sid" 2>&1) \ + || fail "the board source runner did not complete: $out" + assert_contains "$out" "$sid.1.result" "the board answer was never durably captured: $out" + assert_contains "$out" "answers-fed: $sid" \ + "the captured board answer never reached the keyed-answer intake: $out" + + queue=$(cat "$home/state/.wake-queue" 2>/dev/null || true) + assert_contains "$queue" "check: procevent lavish $sid 1" \ + "the captured board answer produced no wake: $queue" + + show=$(tasks_in "$home" show sample-board-call --full) + assert_contains "$show" "state: done" "the board answer did not close the captain call" + assert_contains "$show" "north" "the board answer lost the captain's selection" + assert_contains "$show" "the captured result $sid sequence 1" \ + "the recorded answer did not name the board result that carried it" + pass "a board answer reaches the keyed-answer intake and wakes firstmate" +} + # The intake is channel-agnostic, so chat must reach it the same way a captured # review does - for a task-id key, and for a legacy composed identity. test_chat_channel_feeds_the_same_keyed_answer_intake() { @@ -3795,6 +3865,7 @@ test_reconcile_closes_with_evidence_or_keeps_the_call_open test_reconcile_outcomes_retry_partial_failures_once test_unbound_source_closes_no_hold test_legacy_identities_keep_working +test_board_answer_reaches_the_keyed_answer_intake test_chat_channel_feeds_the_same_keyed_answer_intake test_origin_slug_validation_precedes_path_construction test_status_resolution_over_an_open_hold_is_signalled diff --git a/tests/fm-procevent.test.sh b/tests/fm-procevent.test.sh index 48d373ab07d..4ac1d5f7557 100755 --- a/tests/fm-procevent.test.sh +++ b/tests/fm-procevent.test.sh @@ -53,6 +53,43 @@ pe_register() { # -- ... new_home() { mkdir -p "$1/state"; } wake_payloads() { awk -F '\t' '{print $5}' "$1/state/.wake-queue" 2>/dev/null; } +# The wake queue is a durable tab-separated record firstmate consumes: +# . These read the rows reconcile +# publishes for a source it stranded, keyed by that source and its claim +# generation. +stranded_wake_keys() { # + [ -e "$1/state/.wake-queue" ] || return 0 + awk -F '\t' -v id="$2" \ + '$3 == "check" && index($4, "procevent:" id ":stranded:") == 1 { print $4 }' \ + "$1/state/.wake-queue" +} +stranded_wake_count() { # + stranded_wake_keys "$1" "$2" | grep -c . || true +} +stranded_wake_payloads() { # + [ -e "$1/state/.wake-queue" ] || return 0 + awk -F '\t' -v id="$2" \ + '$3 == "check" && index($4, "procevent:" id ":stranded:") == 1 { print $5 }' \ + "$1/state/.wake-queue" +} +# The same rows for a launch reconcile could not confirm, keyed by that source +# and the registration identity the launch ran under. +launch_failed_wake_keys() { # + [ -e "$1/state/.wake-queue" ] || return 0 + awk -F '\t' -v id="$2" \ + '$3 == "check" && index($4, "procevent:" id ":launch-failed:") == 1 { print $4 }' \ + "$1/state/.wake-queue" +} +launch_failed_wake_count() { # + launch_failed_wake_keys "$1" "$2" | grep -c . || true +} +launch_failed_wake_payloads() { # + [ -e "$1/state/.wake-queue" ] || return 0 + awk -F '\t' -v id="$2" \ + '$3 == "check" && index($4, "procevent:" id ":launch-failed:") == 1 { print $5 }' \ + "$1/state/.wake-queue" +} + first_result() { # : print the first captured result, if any local g for g in "$1/state/procevent-inbox/$2".*.result; do @@ -1138,6 +1175,42 @@ assert_contains "$orphan_out" "started=0" \ [ "$(wc -l < "$ORPHAN_LOG" | tr -d ' ')" = 1 ] \ || fail "reconcile started a source beside an ambiguous leaderless group" assert_absent "$ORPHAN_OVERLAP" "no replacement source starts while the leaderless group remains" +# This is the ordinary crash shape, and it is refused permanently: `orphaned` +# in a listing and `uncertain=1` in output the supervision cycle discards +# reach nobody, so the strand has to announce itself durably, exactly once, +# under a key the watcher can tell apart from a captured result. +orphan_token=$(sed -n '3p' "$FM_PROCEVENT_CLAIM_ROOT/orphan-src.claim") +[ -n "$orphan_token" ] || fail "could not read the leaderless claim's token" +[ "$(stranded_wake_count "$HG" orphan-src)" = 1 ] \ + || fail "reconcile stranded a leaderless source without announcing it: $orphan_out" +[ "$(stranded_wake_keys "$HG" orphan-src)" = "procevent:orphan-src:stranded:$orphan_token" ] \ + || fail "the stranded wake is not keyed by source and claim generation: $(stranded_wake_keys "$HG" orphan-src)" +orphan_wake=$(stranded_wake_payloads "$HG" orphan-src) +assert_contains "$orphan_wake" "orphan-src" \ + "the leaderless stranded wake does not name the source it is about: $orphan_wake" +assert_contains "$orphan_wake" "polling" \ + "the leaderless stranded wake does not say what a human should check: $orphan_wake" +# `start` reports this claim as owned and reclaims nothing, so a wake that +# named it as the clearing command would send someone to a no-op. +case "$orphan_wake" in + *"start orphan-src"*) fail "the leaderless stranded wake names start as clearing it: $orphan_wake" ;; +esac +orphan_start=$(pe "$HG" start orphan-src 2>&1) +assert_contains "$orphan_start" "already owned" \ + "start displaced a leaderless group's claim: $orphan_start" +[ "$(wc -l < "$ORPHAN_LOG" | tr -d ' ')" = 1 ] \ + || fail "start ran the source beside an ambiguous leaderless group" +orphan_again=$(pe "$HG" reconcile) +assert_contains "$orphan_again" "started=0" \ + "the second cycle replaced an ambiguous leaderless generation: $orphan_again" +assert_contains "$orphan_again" "uncertain=1" \ + "the second cycle stopped reporting the claim it could not settle: $orphan_again" +[ "$(stranded_wake_count "$HG" orphan-src)" = 1 ] \ + || fail "reconcile re-announced the same leaderless strand: $orphan_again" +[ "$(wc -l < "$ORPHAN_LOG" | tr -d ' ')" = 1 ] \ + || fail "the second cycle started a source beside an ambiguous leaderless group" +kill -0 -"$orphan_leader" 2>/dev/null \ + || fail "announcing the strand signalled the leaderless process group" kill -KILL -"$orphan_leader" 2>/dev/null || true for _ in $(seq 1 50); do kill -0 -"$orphan_leader" 2>/dev/null || break; sleep 0.1; done kill -0 -"$orphan_leader" 2>/dev/null && fail "could not clean up the leaderless fixture group" @@ -1277,8 +1350,66 @@ sr4_out=$(pe "$HSR4" reconcile) sleep 0.5 [ "$(wc -l < "$SR4_LOG" | tr -d ' ')" = 1 ] \ || fail "reconcile started a replacement beside a reused pid's live group: $sr4_out" +# Ownership cannot move here by design, so a replacement could only die on the +# claim it cannot take - once per reconcile cycle, forever. +assert_contains "$sr4_out" "started=0" \ + "reconcile reported a start into a claim nothing can take: $sr4_out" +assert_contains "$sr4_out" "uncertain=1" \ + "reconcile did not report the claim it could not settle: $sr4_out" [ "$(sed -n '2p' "$sr4_claim")" = "$sr4_leader" ] \ || fail "reconcile replaced the reused-pid generation's claim" +# Nothing can take this source, so reporting it as unowned reads like an idle +# source waiting to be started - the reassuring answer this surface gave while a +# review board collected nothing. +sr4_owner=$(pe "$HSR4" list | awk '$1 == "reused-group-src" { print $3 }') +[ "$sr4_owner" = orphaned ] \ + || fail "a source no caller can claim is listed as '$sr4_owner'" +# `orphaned` in a listing and `uncertain=1` in output the supervision cycle +# discards reach nobody. The strand has to announce itself durably, exactly +# once, and say which command clears it. +[ "$(stranded_wake_count "$HSR4" reused-group-src)" = 1 ] \ + || fail "reconcile stranded a source without announcing it: $sr4_out" +# The key carries the source and its claim generation in a shape the watcher +# can tell apart from a captured result, so the strand is never headlined as one. +[ "$(stranded_wake_keys "$HSR4" reused-group-src)" = "procevent:reused-group-src:stranded:$(sed -n '3p' "$sr4_claim")" ] \ + || fail "the stranded wake is not keyed by source and claim generation: $(stranded_wake_keys "$HSR4" reused-group-src)" +sr4_wake=$(stranded_wake_payloads "$HSR4" reused-group-src) +assert_contains "$sr4_wake" "reused-group-src" \ + "the stranded wake does not name the source it is about: $sr4_wake" +assert_contains "$sr4_wake" "bin/fm-procevent.sh start reused-group-src" \ + "the stranded wake does not name the command that clears it: $sr4_wake" +# A wake nobody can silence is as unusable as one nobody gets: the same stranded +# generation must not re-announce on every supervision cycle. +sr4_again=$(pe "$HSR4" reconcile) +assert_contains "$sr4_again" "uncertain=1" \ + "the second cycle stopped reporting the claim it could not settle: $sr4_again" +[ "$(stranded_wake_count "$HSR4" reused-group-src)" = 1 ] \ + || fail "reconcile re-announced the same stranded generation: $sr4_again" +[ "$(wc -l < "$SR4_LOG" | tr -d ' ')" = 1 ] \ + || fail "the second cycle started a replacement beside a reused pid's live group: $sr4_again" +# The wake names `start` as the recovery, so run it against the state it will +# actually meet. The earlier end-to-end demonstration of that command used an +# UNDRIFTED fixture and therefore proved only the easy case; on this one the +# state root has drifted, so the dead generation's reservation records cannot +# be tidied, and the claim path waives that tidy-up only for a generation +# proven gone - which a surviving group is not. `start` must refuse here, keep +# the claim, and start no second source beside the live group, and the wake +# must have said so rather than promising a reclaim. +assert_contains "$sr4_wake" "cannot claim source" \ + "the stranded wake promises an unconditional reclaim: $sr4_wake" +set +e +sr4_start=$(pe "$HSR4" start reused-group-src 2>&1) +sr4_start_rc=$? +set -e +[ "$sr4_start_rc" -ne 0 ] \ + || fail "start reported success against a claim it could not tidy: $sr4_start" +assert_contains "$sr4_start" "cannot claim source" \ + "start did not refuse by name on the drifted reused-pid fixture: $sr4_start" +[ "$(sed -n '2p' "$sr4_claim")" = "$sr4_leader" ] \ + || fail "a refused start replaced the reused-pid generation's claim" +sleep 0.3 +[ "$(wc -l < "$SR4_LOG" | tr -d ' ')" = 1 ] \ + || fail "a refused start ran a second source beside a reused pid's live group: $(cat "$SR4_LOG")" set +e sr4_retire=$(pe "$HSR4" retire reused-group-src 2>&1) sr4_rc=$? @@ -1299,6 +1430,347 @@ kill -0 -"$sr4_leader" 2>/dev/null \ && fail "retirement left the restored reused-group fixture running" pass "a reused pid never makes its surviving process group reclaimable" +# --- the easy case the wake promises: an undrifted reused-pid claim ---------- +# Same strand, no state-root drift: the dead generation's reservation records +# can be tidied, so the attached `start` the wake names takes the claim and +# runs the source. The claim path does not consult the process group; that is +# the documented asymmetry between reconcile and a deliberate start. +HSR5="$TMP_ROOT/hsr5"; new_home "$HSR5" +SR5_TRIGGER="$TMP_ROOT/reused-plain-trigger" +SR5_LOG="$TMP_ROOT/reused-plain-executions" +pe_register "$HSR5" lavish reused-plain-src -- "$RACE_BLOCKER" "$SR5_LOG" "$SR5_TRIGGER" >/dev/null +pe "$HSR5" reconcile >/dev/null +wait_for "$FM_PROCEVENT_CLAIM_ROOT/reused-plain-src.claim" \ + || fail "undrifted reused-pid fixture never claimed its source" +wait_for "$SR5_LOG" || fail "undrifted reused-pid fixture source never started" +sr5_claim="$FM_PROCEVENT_CLAIM_ROOT/reused-plain-src.claim" +sr5_leader=$(sed -n '2p' "$sr5_claim") +awk 'NR == 4 { print "different-live-process-identity"; next } { print }' \ + "$sr5_claim" > "$sr5_claim.tmp" && mv "$sr5_claim.tmp" "$sr5_claim" +chmod 0600 "$sr5_claim" +sr5_out=$(pe "$HSR5" reconcile) +assert_contains "$sr5_out" "uncertain=1" \ + "reconcile did not strand the undrifted reused-pid claim: $sr5_out" +[ "$(stranded_wake_count "$HSR5" reused-plain-src)" = 1 ] \ + || fail "the undrifted strand was not announced: $sr5_out" +pe "$HSR5" start reused-plain-src > "$TMP_ROOT/reused-plain-start.out" 2>&1 & +sr5_start_pid=$! +wait_for_lines "$SR5_LOG" 2 \ + || fail "start did not reclaim the undrifted reused-pid claim: $(cat "$TMP_ROOT/reused-plain-start.out")" +[ "$(sed -n '2p' "$sr5_claim")" != "$sr5_leader" ] \ + || fail "start ran the source without taking the claim from the dead generation" +: > "$SR5_TRIGGER" +wait "$sr5_start_pid" \ + || fail "start failed after reclaiming the undrifted claim: $(cat "$TMP_ROOT/reused-plain-start.out")" +assert_contains "$(cat "$TMP_ROOT/reused-plain-start.out")" "captured:" \ + "the reclaiming start did not capture the source's result" +for _ in $(seq 1 50); do kill -0 -"$sr5_leader" 2>/dev/null || break; sleep 0.1; done +pe "$HSR5" retire reused-plain-src >/dev/null 2>&1 || true +pass "start reclaims a reused-pid claim whose leftovers can still be tidied" + +# --- a launch that cannot confirm is announced once per failure episode ------ +# `bin/fm-watch.sh` discards reconcile's `failed=` count and exit status, so a +# runner that dies before claiming - for any cause, not only the claim wedge - +# would be relaunched and reported failed every cycle with nobody told: armed +# in appearance, a dead drop in fact. The episode is keyed by the registration +# identity the launch ran under and ends when a launch of that source confirms, +# so the registration below is damaged and repaired IN PLACE to keep that +# identity fixed across the whole sequence. The wake changes nothing about the +# launch: every failing cycle below still relaunches and still reports failed. +HEP="$TMP_ROOT/hep"; new_home "$HEP" +EP_SOURCE_CMD="$TMP_ROOT/episode-source.sh" +cat > "$EP_SOURCE_CMD" <<'SH' +#!/usr/bin/env bash +printf 'episode result\n' +SH +chmod +x "$EP_SOURCE_CMD" +pe_register "$HEP" lavish episode-src -- "$EP_SOURCE_CMD" >/dev/null +EP_SOURCE="$HEP/state/procevent/episode-src.source" +cp "$EP_SOURCE" "$TMP_ROOT/episode-good.source" +awk '/^argv:$/ { print; exit } { print }' "$EP_SOURCE" > "$TMP_ROOT/episode-bad.source" \ + || fail "could not prepare the damaged episode registration" +ep_damage() { cat "$TMP_ROOT/episode-bad.source" > "$EP_SOURCE"; } +ep_repair() { cat "$TMP_ROOT/episode-good.source" > "$EP_SOURCE"; } +ep_reconcile() { # ; sets ep_out + local rc=0 + ep_out=$(FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS=2 pe "$HEP" reconcile) || rc=$? + assert_contains "$ep_out" "$1" "$3: $ep_out" + if [ "$2" -eq 1 ]; then + [ "$rc" -ne 0 ] || fail "$3 (reconcile exited 0): $ep_out" + else + [ "$rc" -eq 0 ] || fail "$3 (reconcile exited $rc): $ep_out" + fi +} +ep_damage +ep_reconcile "failed=1" 1 "a launch that never proved its claim was not reported failed" +[ "$(launch_failed_wake_count "$HEP" episode-src)" = 1 ] \ + || fail "a launch that could not confirm was not announced: $ep_out" +ep_key=$(launch_failed_wake_keys "$HEP" episode-src) +# -: the watcher remembers every key +# it has surfaced for good, so the identity alone would announce only the first +# episode of a registration (tests/fm-watch-triage.test.sh proves delivery). +[[ "$ep_key" =~ ^(procevent:episode-src:launch-failed:[0-9]+-[0-9]+)-[0-9]+$ ]] \ + || fail "the launch-failed wake is not keyed by source, registration identity and episode: $ep_key" +ep_episode_prefix=${BASH_REMATCH[1]} +ep_wake=$(launch_failed_wake_payloads "$HEP" episode-src) +assert_contains "$ep_wake" "episode-src" \ + "the launch-failed wake does not name the source it is about: $ep_wake" +# The payload may state only what confirmation observed: no claim proved +# inside the window. It cannot know whether the runner died or was slow, so it +# must not assert a cause, must not present `start` as the fix, and must say +# that a later cycle finding the source owned closes the episode by itself. +assert_contains "$ep_wake" "did not prove it took the source's claim within FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS" \ + "the launch-failed wake does not state what confirmation observed: $ep_wake" +assert_contains "$ep_wake" "attached bin/fm-procevent.sh start episode-src to reproduce a refusal" \ + "the launch-failed wake does not say start reproduces rather than fixes: $ep_wake" +assert_contains "$ep_wake" "adapter binary" \ + "the launch-failed wake does not name what to check: $ep_wake" +assert_contains "$ep_wake" "finds the source owned ends this episode on its own" \ + "the launch-failed wake does not say a slow runner closes its own episode: $ep_wake" +case "$ep_wake" in + *"never claimed"*|*"exited without"*|*"runner died"*) + fail "the launch-failed wake asserts a cause confirmation cannot observe: $ep_wake" ;; +esac +ep_reconcile "failed=1" 1 "the second cycle stopped relaunching a source that cannot start" +[ "$(launch_failed_wake_count "$HEP" episode-src)" = 1 ] \ + || fail "the same failure episode was announced twice: $ep_out" +ep_repair +ep_reconcile "started=1" 0 "a repaired source did not confirm" +assert_contains "$ep_out" "failed=0" "a repaired source was still reported failed: $ep_out" +[ "$(launch_failed_wake_count "$HEP" episode-src)" = 1 ] \ + || fail "a confirmed launch produced a launch-failed wake: $ep_out" +for _ in $(seq 1 100); do + [ -e "$FM_PROCEVENT_CLAIM_ROOT/episode-src.claim" ] || break + sleep 0.1 +done +[ ! -e "$FM_PROCEVENT_CLAIM_ROOT/episode-src.claim" ] \ + || fail "the confirmed episode runner never released its claim" +ep_damage +ep_reconcile "failed=1" 1 "a source that failed again after recovering was not reported failed" +[ "$(launch_failed_wake_count "$HEP" episode-src)" = 2 ] \ + || fail "a new failure episode after a confirmed launch was not announced: $ep_out" +# The earlier version of this assertion locked in ONE key for both episodes, +# which is exactly the collision that left every episode after the first +# unsurfaced: both keys must carry the same registration identity and still +# differ, or the watcher's seen marker for episode one suppresses episode two. +ep_key_again=$(launch_failed_wake_keys "$HEP" episode-src | sed -n '2p') +[ "$ep_key_again" != "$ep_key" ] \ + || fail "a new failure episode reused the first episode's queue key: $ep_key_again" +case "$ep_key_again" in + "$ep_episode_prefix"-*) ;; + *) fail "the second episode ran under a different registration identity: $ep_key_again (first: $ep_key)" ;; +esac +ep_repair +pe "$HEP" retire episode-src >/dev/null 2>&1 || true +pass "a launch that cannot confirm is announced once per failure episode" + +# --- the launch-failed key fits the watcher's seen marker at the id limit ---- +# bin/fm-watch.sh names the marker for a surfaced key `.seen-procevent-`, +# 16 + 2 * keylen bytes against NAME_MAX 255, so a key longer than 119 chars +# cannot be marked and its wake would re-surface every cycle. The longest id +# the validator accepts is 64 chars; the executed key for such an id must fit. +HLK="$TMP_ROOT/hlk"; new_home "$HLK" +LK_ID=$(printf 'k%.0s' $(seq 1 64)) +[ "${#LK_ID}" -eq 64 ] || fail "fixture invalid: long source id is ${#LK_ID} chars" +pe_register "$HLK" lavish "$LK_ID" -- "$EP_SOURCE_CMD" >/dev/null +LK_SOURCE="$HLK/state/procevent/$LK_ID.source" +if ! { awk '/^argv:$/ { print; exit } { print }' "$LK_SOURCE" > "$LK_SOURCE.tmp" \ + && cat "$LK_SOURCE.tmp" > "$LK_SOURCE" && rm -f -- "$LK_SOURCE.tmp"; }; then + fail "could not damage the long-id registration" +fi +lk_out=$(FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS=2 pe "$HLK" reconcile) || true +assert_contains "$lk_out" "failed=1" "the long-id launch was not reported failed: $lk_out" +lk_key=$(launch_failed_wake_keys "$HLK" "$LK_ID") +[ -n "$lk_key" ] || fail "the long-id launch failure was not announced: $lk_out" +[ "${#lk_key}" -le 119 ] \ + || fail "a 64-char source id yields a ${#lk_key}-char launch-failed key, which the watcher cannot mark: $lk_key" +pe "$HLK" retire "$LK_ID" >/dev/null 2>&1 || true +pass "a 64-char source id keeps the launch-failed key within the watcher's marker bound" + +# --- reconcile reports only launches it actually confirmed ------------------- +# The reported incident. A review board the captain had answered sat collecting +# nothing while `reconcile` reported a start on every run: `detach_runner` is +# fire-and-forget with the child's stderr discarded, so a runner that died +# before it could claim was counted exactly like one that is listening. A +# surface that presents as armed while being a dead drop is worse than one that +# visibly fails, because the answers look recorded. +# +# The damaged registration below makes the runner die BEFORE it claims, which +# is what keeps this deterministic: a runner that claims and then dies would +# race the confirmation either way, and the next reconcile cycle is what covers +# that case. +HUF="$TMP_ROOT/huf"; new_home "$HUF" +UF_TRIGGER="$TMP_ROOT/unstartable-trigger" +pe_register "$HUF" lavish unstartable-src -- "$BLOCKER" "$UF_TRIGGER" "unstartable" >/dev/null +UF_SOURCE="$HUF/state/procevent/unstartable-src.source" +if ! awk '/^argv:$/ { print; exit } { print }' "$UF_SOURCE" > "$UF_SOURCE.tmp"; then + fail "could not damage the unstartable registration" +fi +mv "$UF_SOURCE.tmp" "$UF_SOURCE" || fail "could not damage the unstartable registration" +chmod 0600 "$UF_SOURCE" +uf_rc=0 +uf_out=$(FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS=2 pe "$HUF" reconcile) || uf_rc=$? +assert_contains "$uf_out" "started=0" \ + "reconcile counted a runner that never started as a start: $uf_out" +assert_contains "$uf_out" "failed=1" \ + "reconcile did not report the launch it could not confirm: $uf_out" +[ "$uf_rc" -ne 0 ] || fail "reconcile reported success while a source could not start: $uf_out" +uf_owner=$(pe "$HUF" list | awk '$1 == "unstartable-src" { print $3 }') +[ "$uf_owner" = none ] || fail "the unstartable source reports an owner: $uf_owner" +pe "$HUF" retire unstartable-src >/dev/null 2>&1 || true +pass "reconcile reports a launch it could not confirm instead of counting it as a start" + +# --- a launch that finished before the first poll is still confirmed --------- +# Confirmation has to read evidence a finished runner leaves behind. A runner +# removes its own runner record on the way out, so a source that claims, runs +# and exits before confirmation looks at it once returns every transient signal +# to exactly what it was before the launch - and a good run gets reported as a +# failure, on every cycle, for a source that is working perfectly. +# +# The second registration is what makes that deterministic rather than a race: +# reconcile launches the fast source first, then blocks acquiring the held +# lock of the second source, and the holder is released only once the fast +# runner has captured its result and let go of both its claim and its runner +# record. Confirmation therefore starts strictly after the fast runner is gone. +HFC="$TMP_ROOT/hfc"; new_home "$HFC" +FC_FAST="$TMP_ROOT/fast-source.sh" +cat > "$FC_FAST" <<'SH' +#!/usr/bin/env bash +printf 'fast payload\n' +SH +chmod +x "$FC_FAST" +FC_TRIGGER="$TMP_ROOT/fast-hold-trigger" +pe_register "$HFC" lavish aa-fast-src -- "$FC_FAST" >/dev/null +pe_register "$HFC" lavish zz-hold-src -- "$BLOCKER" "$FC_TRIGGER" "held" >/dev/null +FC_READY="$TMP_ROOT/fast-hold-ready"; FC_RELEASE="$TMP_ROOT/fast-hold-release" +hold_source_lock zz-hold-src "$FC_READY" "$FC_RELEASE" +wait_for "$FC_READY" || fail "the fast-source fixture could not hold a source lock" +( + for _ in $(seq 1 600); do + if first_result "$HFC" aa-fast-src >/dev/null 2>&1 \ + && [ ! -e "$HFC/state/procevent/aa-fast-src.runner" ] \ + && [ ! -e "$FM_PROCEVENT_CLAIM_ROOT/aa-fast-src.claim" ]; then + break + fi + sleep 0.05 + done + : > "$FC_RELEASE" +) & +FC_RELEASER=$! +fc_rc=0 +fc_out=$(FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS=2 pe "$HFC" reconcile) || fc_rc=$? +wait "$FC_RELEASER" 2>/dev/null || true +wait "$HOLDER_PID" 2>/dev/null || true +first_result "$HFC" aa-fast-src >/dev/null \ + || fail "fixture invalid: the fast source never produced a result: $fc_out" +assert_contains "$fc_out" "started=2" \ + "reconcile did not report both launches as started: $fc_out" +assert_contains "$fc_out" "failed=0" \ + "reconcile reported a launch that ran to completion as a failure: $fc_out" +[ "$fc_rc" -eq 0 ] || fail "reconcile exited non-zero with every launch confirmed: $fc_out" +: > "$FC_TRIGGER" +pe "$HFC" retire aa-fast-src >/dev/null 2>&1 || true +pe "$HFC" retire zz-hold-src >/dev/null 2>&1 || true +pass "a launch that finished before confirmation looked is still reported as started" + +# --- a zero-padded confirm window is read as base 10 ------------------------- +# The window's validator reads base 10, so `08` is a value it accepts. Read as +# octal in arithmetic it is not a number at all, which under `set -u` takes the +# confirmation down with it and turns every launch of the cycle - including a +# perfectly healthy one - into a reported failure and a non-zero exit. +HZP="$TMP_ROOT/hzp"; new_home "$HZP" +ZP_TRIGGER="$TMP_ROOT/zeropad-trigger" +pe_register "$HZP" lavish zeropad-src -- "$BLOCKER" "$ZP_TRIGGER" "zeropad" >/dev/null +zp_rc=0 +zp_out=$(FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS=08 pe "$HZP" reconcile 2>/dev/null) || zp_rc=$? +assert_contains "$zp_out" "started=1" \ + "a zero-padded confirm window lost the launch reconcile started: $zp_out" +assert_contains "$zp_out" "failed=0" \ + "a zero-padded confirm window reported a healthy launch as failed: $zp_out" +[ "$zp_rc" -eq 0 ] || fail "a zero-padded confirm window made reconcile exit non-zero: $zp_out" +: > "$ZP_TRIGGER" +pe "$HZP" retire zeropad-src >/dev/null 2>&1 || true +pass "a zero-padded launch confirm window is honored as base 10" + +# --- an unusable confirm window is refused by name -------------------------- +# A window this command cannot use makes every launch unconfirmable. Reported +# from inside the confirmation it comes out as a fleet of healthy runners that +# all "could not start", blaming the sources instead of the typo. Every other +# tunable on this path - the launch floor, the output bound - refuses a bad +# value by name before anything runs, and so does this one. +HIW="$TMP_ROOT/hiw"; new_home "$HIW" +IW_TRIGGER="$TMP_ROOT/invalid-window-trigger" +pe_register "$HIW" lavish invalid-window-src -- "$BLOCKER" "$IW_TRIGGER" "window" >/dev/null +for iw_value in 5s 0 700; do + iw_rc=0 + iw_out=$(FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS="$iw_value" pe "$HIW" reconcile 2>&1) || iw_rc=$? + [ "$iw_rc" -ne 0 ] \ + || fail "reconcile accepted the unusable confirm window '$iw_value': $iw_out" + assert_contains "$iw_out" "FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS" \ + "the unusable confirm window '$iw_value' was not named by what refused it: $iw_out" + case "$iw_out" in + *failed=*) fail "the unusable confirm window '$iw_value' was blamed on the sources: $iw_out" ;; + esac + [ ! -e "$FM_PROCEVENT_CLAIM_ROOT/invalid-window-src.claim" ] \ + || fail "reconcile launched a runner before refusing the confirm window '$iw_value'" +done +# The refusal costs the source nothing: it still arms on the next run with a +# usable value. +iw_ok=$(pe "$HIW" reconcile) +assert_contains "$iw_ok" "started=1" \ + "the source did not arm once its confirm window was usable: $iw_ok" +assert_contains "$iw_ok" "failed=0" \ + "the source was reported as failed once its confirm window was usable: $iw_ok" +: > "$IW_TRIGGER" +pe "$HIW" retire invalid-window-src >/dev/null 2>&1 || true +pass "an unusable launch confirm window is refused by name instead of blamed on the sources" + +# --- a dead generation's untidyable leftovers never wedge ownership ---------- +# The same wedge as the state-root case above, reached through the sibling +# cleanups in the stale-claim branch rather than the capture reservation. Every +# one of them tidies leftovers keyed by the DEAD generation's claim token, so +# none can collide with the replacement, yet a failure in any of them used to +# refuse the claim outright - permanently, because the condition never clears on +# its own. Here the recorded registry directory no longer resolves to a +# directory at all, which is what a claim recorded before its home was replaced +# looks like. +HUW="$TMP_ROOT/huw"; new_home "$HUW" +UW_TRIGGER="$TMP_ROOT/untidyable-trigger" +UW_LOG="$TMP_ROOT/untidyable-executions" +pe_register "$HUW" lavish untidyable-src -- "$RACE_BLOCKER" "$UW_LOG" "$UW_TRIGGER" >/dev/null +UW_REG_FILE="$TMP_ROOT/untidyable-recorded-registry" +: > "$UW_REG_FILE" +uw_identity=$(bash -c '. "$1/bin/fm-pr-lib.sh"; fm_pr_file_identity "$2"' _ \ + "$ROOT" "$HUW/state/procevent/untidyable-src.source") \ + || fail "could not read the untidyable fixture registration identity" +UW_CLAIM="$FM_PROCEVENT_CLAIM_ROOT/untidyable-src.claim" +{ + printf '%s\n%s\nuntidyable-token\nuntidyable-identity\n' "$HUW" 999999 + printf '%s\n%s\nactive\n' "$UW_REG_FILE" "$uw_identity" + printf '%s\n%s\n%s\n%s\n%s\n' "$HUW/state" \ + "$(bash -c '. "$1/bin/fm-pr-lib.sh"; fm_pr_file_device "$2"' _ "$ROOT" "$HUW/state")" \ + "$(bash -c '. "$1/bin/fm-pr-lib.sh"; fm_pr_file_inode "$2"' _ "$ROOT" "$HUW/state")" \ + "$(id -u)" 755 +} > "$UW_CLAIM" +chmod 0600 "$UW_CLAIM" +kill -0 999999 2>/dev/null && fail "fixture invalid: the untidyable claim names a live pid" +kill -0 -999999 2>/dev/null && fail "fixture invalid: the untidyable claim's process group is alive" +uw_rc=0 +uw_out=$(pe "$HUW" reconcile) || uw_rc=$? +[ "$uw_rc" -eq 0 ] || fail "reconcile could not repair a provably dead generation: $uw_out" +# Reporting a start is not the same fact as listening, so prove the listening +# half first: before this fix reconcile reported exactly this start on every run +# while the dead generation kept the claim and nothing ever attached. +wait_for "$UW_LOG" || fail "reconcile reported a start but no replacement source ever ran: $uw_out" +uw_new=$(sed -n '2p' "$UW_CLAIM") +[ "$uw_new" != 999999 ] || fail "the dead generation kept owning the source: $uw_out" +kill -0 "$uw_new" 2>/dev/null || fail "the replacement runner did not take ownership: $uw_out" +assert_contains "$uw_out" "started=1" "reconcile did not report the replacement it started: $uw_out" +assert_contains "$uw_out" "failed=0" "reconcile could not confirm the replacement: $uw_out" +: > "$UW_TRIGGER" +pe "$HUW" retire untidyable-src >/dev/null +pass "a dead generation whose leftovers cannot be tidied never keeps owning its source" + HJ="$TMP_ROOT/hj"; new_home "$HJ" TORN_TRIGGER="$TMP_ROOT/torn-trigger" pe_register "$HJ" lavish torn-src -- "$BLOCKER" "$TORN_TRIGGER" "torn" >/dev/null diff --git a/tests/fm-watch-arm.test.sh b/tests/fm-watch-arm.test.sh index 49d350c538f..33cd245700a 100755 --- a/tests/fm-watch-arm.test.sh +++ b/tests/fm-watch-arm.test.sh @@ -799,8 +799,51 @@ test_downtime_marker_does_not_follow_symlink() { pass "watch-arm: downtime marker publication does not follow symlinks" } +# The watcher validates FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS when it arms and +# refuses to arm on an unusable value. Under a running watcher that value would +# make every per-cycle reconcile refuse by name into a discarded stdout, so no +# source would ever start and the home would sit disarmed while presenting as +# supervised; refusing to arm is loud through the liveness guard instead. This +# drives the real arm entry and asserts the arm STOPPED - non-zero exit, no +# started line, no lock holder, no beacon - and that its refusal names the +# variable, so a validator that merely returned false somewhere would not pass. +test_arm_refuses_an_unusable_launch_confirm_window() { + local dir home state fakebin armout status lock_pid + dir=$(make_case confirm-window-refusal) + home="$dir/home" + state="$dir/state" + fakebin="$dir/fakebin" + armout="$dir/arm.out" + mkdir -p "$home/data" + + PATH="$fakebin:$PATH" FM_HOME="$home" FM_STATE_OVERRIDE="$state" \ + FM_POLL=1 FM_SIGNAL_GRACE=0 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + FM_ARM_CONFIRM_TIMEOUT=5 FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS=5s \ + "$WATCH_ARM" > "$armout" 2>&1 & + ARM_PID=$! + wait_for_exit "$ARM_PID" 200 + status=$? + [ "$status" -ne 124 ] || fail "arm with an unusable confirm window never stopped: $(cat "$armout")" + [ "$status" -ne 0 ] || fail "arm reported success with an unusable confirm window: $(cat "$armout")" + grep -q '^watcher: FAILED' "$armout" \ + || fail "arm did not report the typed failure line: $(cat "$armout")" + grep -qF 'FM_PROCEVENT_LAUNCH_CONFIRM_SECONDS' "$armout" \ + || fail "the refusal did not name the variable: $(cat "$armout")" + grep -qF "must be whole seconds from 1 to 600" "$armout" \ + || fail "the refusal did not name the accepted range: $(cat "$armout")" + ! grep -q '^watcher: started' "$armout" \ + || fail "arm reported a started watcher despite the refusal: $(cat "$armout")" + [ ! -e "$state/.last-watcher-beat" ] \ + || fail "a refused watcher still published a liveness beacon" + lock_pid=$(cat "$state/.watch.lock/pid" 2>/dev/null || true) + [ -z "$lock_pid" ] || ! kill -0 "$lock_pid" 2>/dev/null \ + || fail "a refused watcher is still running as pid $lock_pid" + pass "watch-arm: an unusable launch confirm window refuses to arm by name" +} + test_attached_arm_reports_the_delivered_wake test_attached_arm_reports_the_delivered_wake_after_drain +test_arm_refuses_an_unusable_launch_confirm_window test_attached_arm_still_fails_on_a_wake_it_did_not_deliver test_rearm_resurfaces_durable_queue_and_remote_open_decision test_marker_publish_failure_retains_recovery_evidence diff --git a/tests/fm-watch-triage.test.sh b/tests/fm-watch-triage.test.sh index 8c1c8f0de49..c4f3bd428d0 100755 --- a/tests/fm-watch-triage.test.sh +++ b/tests/fm-watch-triage.test.sh @@ -4022,6 +4022,19 @@ seed_captured_procevent_result() { # sleep 0.1 i=$((i + 1)) done + # The runner publishes that wake BEFORE it releases its claim and exits, so a + # retire that lands in that gap reads the exiting runner's ownership as + # uncertain and refuses with "cannot confirm runner identity" - the pipeline + # saw exactly that under load. Wait, bounded, for the release the publish + # promises, so retire meets a source nothing owns instead of racing the + # runner's last milliseconds. The bound keeps a runner that never releases a + # real failure at retire rather than a hang here. + i=0 + while [ "$i" -lt 100 ]; do + [ -e "$dir/claims/delivery-src.claim" ] || break + sleep 0.1 + i=$((i + 1)) + done pe_case "$dir" retire delivery-src >/dev/null || return 1 [ -s "$dir/state/.wake-queue" ] } @@ -4125,6 +4138,108 @@ test_procevent_marker_keys_are_injective() { pass "complete process-event queue keys map to distinct seen markers" } +# The reason line is the headline firstmate reads before the payload. Every +# procevent:* key used to surface as "process-event result captured", which +# presents a source that is collecting NOTHING as a healthy capture - the exact +# shape of the incident these wakes exist to expose. These assertions read the +# reason the watcher actually printed, so a typo in either classifying glob +# fails here instead of silently falling back to the healthy-looking headline. +surface_once() { # [limit-ticks]: run one watcher to its wake, return its status + local dir=$1 out=$2 limit=${3:-100} pid + procevent_watch_bg "$dir" "$out" + pid=$! + wait_for_exit "$pid" "$limit" +} + +test_procevent_headlines_classify_queue_keys() { + local dir state out + dir=$(make_case procevent-headline-captured); state="$dir/state"; out="$dir/watch.out" + append_wake "$state" check "procevent:cap-src:1" "check: procevent lavish cap-src 1" + surface_once "$dir" "$out" || fail "a captured-result key was not surfaced: $(cat "$out")" + grep -F "check: process-event result captured: procevent:cap-src:1" "$out" >/dev/null \ + || fail "a captured result did not surface under its own headline: $(cat "$out")" + ! grep -F "source stranded" "$out" >/dev/null \ + || fail "a captured result was headlined as a strand: $(cat "$out")" + ! grep -F "failed to start" "$out" >/dev/null \ + || fail "a captured result was headlined as a failed start: $(cat "$out")" + FM_STATE_OVERRIDE="$state" "$DRAIN" >/dev/null 2>&1 || fail "captured headline fixture drain failed" + + dir=$(make_case procevent-headline-stranded); state="$dir/state"; out="$dir/watch.out" + append_wake "$state" check "procevent:str-src:stranded:tok-1" "check: process-event source str-src is registered but nothing can arm it" + surface_once "$dir" "$out" || fail "a stranded key was not surfaced: $(cat "$out")" + grep -F "check: process-event source stranded: procevent:str-src:stranded:tok-1" "$out" >/dev/null \ + || fail "a stranded source did not surface under its own headline: $(cat "$out")" + ! grep -F "result captured" "$out" >/dev/null \ + || fail "a stranded source was headlined as a captured result: $(cat "$out")" + FM_STATE_OVERRIDE="$state" "$DRAIN" >/dev/null 2>&1 || fail "stranded headline fixture drain failed" + + dir=$(make_case procevent-headline-joined); state="$dir/state"; out="$dir/watch.out" + append_wake "$state" check "procevent:cap2-src:1" "check: procevent lavish cap2-src 1" + append_wake "$state" check "procevent:str2-src:stranded:tok-2" "check: process-event source str2-src is registered but nothing can arm it" + surface_once "$dir" "$out" || fail "a mixed cycle was not surfaced: $(cat "$out")" + grep -F "check: process-event result captured: procevent:cap2-src:1; process-event source stranded: procevent:str2-src:stranded:tok-2" "$out" >/dev/null \ + || fail "a cycle with a capture and a strand did not carry both headlines joined: $(cat "$out")" + FM_STATE_OVERRIDE="$state" "$DRAIN" >/dev/null 2>&1 || fail "joined headline fixture drain failed" + pass "process-event queue keys surface under their own headlines" +} + +# Delivery, not queue rows, is what proves a launch-failure episode reaches +# firstmate. The watcher remembers every procevent key it has surfaced for +# good, so reconcile keys each episode with a fresh suffix beyond the +# registration identity: this test would fail if a second episode reused the +# first one's key, because the watcher would keep polling and never wake. +test_procevent_launch_failed_episodes_are_each_delivered() { + local dir state out status + dir=$(make_case procevent-launch-failed-episodes); state="$dir/state"; out="$dir/watch.out" + append_wake "$state" check "procevent:lf-src:launch-failed:1-2-100-7" \ + "check: process-event source lf-src is registered but its launch did not prove it took the claim" + surface_once "$dir" "$out" || fail "a launch-failed key was not surfaced: $(cat "$out")" + grep -F "check: process-event source failed to start: procevent:lf-src:launch-failed:1-2-100-7" "$out" >/dev/null \ + || fail "a failed launch did not surface under its own headline: $(cat "$out")" + ! grep -F "result captured" "$out" >/dev/null \ + || fail "a failed launch was headlined as a captured result: $(cat "$out")" + ack_stopped_cycle "$state" >/dev/null || fail "launch-failed fixture could not be handled and acknowledged" + + # The same key again is what a registration-identity-only key would produce + # for the next episode: already surfaced, so the process-event surface never + # delivers it under its headline again. A fresh watcher still recovers the + # unacknowledged queue row through the generic `check: rearm-resurface` + # path (the contract test_procevent_unacknowledged_result_redrains_until_handled + # proves), so what this asserts is the headline, not silence. + append_wake "$state" check "procevent:lf-src:launch-failed:1-2-100-7" \ + "check: process-event source lf-src is registered but its launch did not prove it took the claim" + : > "$out" + status=0 + surface_once "$dir" "$out" 30 || status=$? + case "$status" in + 124) ;; + 0) + # The one wake this tolerates is the recovery path named above, by its + # exact reason line. A wake for any other reason would mean either that + # the ordinary surface delivered the repeated key after all, or that + # something unrelated fired inside the window - and both are failures of + # exactly what this test guards, so neither may pass as "recovery". + grep -F 'check: rearm-resurface' "$out" >/dev/null \ + || fail "an already-surfaced launch-failed key woke the watcher, and the reason was not the one tolerated recovery path (expected the exact line 'check: rearm-resurface'; if that path was reworded, update this expectation, do not restore the strict silence check): $(cat "$out")" + ;; + *) fail "the watcher failed on an already-surfaced launch-failed key (status $status): $(cat "$out")" ;; + esac + ! grep -F "failed to start: procevent:lf-src:launch-failed:1-2-100-7" "$out" >/dev/null \ + || fail "an already-surfaced launch-failed key was delivered again under its headline: $(cat "$out")" + ack_stopped_cycle "$state" >/dev/null || fail "repeated-key fixture could not be handled and acknowledged" + + # A later episode of the same registration carries the same identity under a + # fresh suffix, and that one must be delivered. + append_wake "$state" check "procevent:lf-src:launch-failed:1-2-160-9" \ + "check: process-event source lf-src is registered but its launch did not prove it took the claim" + : > "$out" + surface_once "$dir" "$out" || fail "a second launch-failure episode was not surfaced: $(cat "$out")" + grep -F "check: process-event source failed to start: procevent:lf-src:launch-failed:1-2-160-9" "$out" >/dev/null \ + || fail "a second launch-failure episode did not surface under its own headline: $(cat "$out")" + ack_stopped_cycle "$state" >/dev/null || fail "second episode fixture could not be handled and acknowledged" + pass "every launch-failure episode is delivered under the failed-to-start headline" +} + install_marker_mv_fault() { # local dir=$1 REAL_MV=$(command -v mv) @@ -4768,6 +4883,8 @@ test_triage_log_size_cap_accepts_spaced_wc_counts test_procevent_captured_result_surfaces_proactively test_procevent_unacknowledged_result_redrains_until_handled test_procevent_marker_keys_are_injective +test_procevent_headlines_classify_queue_keys +test_procevent_launch_failed_episodes_are_each_delivered test_procevent_surface_serializes_with_drain test_procevent_surface_crash_boundaries test_procevent_marker_failure_exits_and_replays From e0d269e07318a5a80069ab6193b4a0af4c077a61 Mon Sep 17 00:00:00 2001 From: Pablo Ontiveros Date: Fri, 11 Sep 2026 09:14:44 -0600 Subject: [PATCH 04/31] fix(herdr): verify agent liveness at process level before trusting registration (#4191) * fix(herdr): verify agent registrations at process level before trusting them Herdr keeps a Pi registration (`agent get` -> agent=pi, agent_status=idle) after the Pi process has exited to a plain shell whenever a nested interactive shell sits under the pane's top shell, which is the crew shape `treehouse get` leaves behind. The pane classifier trusted that registration alone, so `fm-control.sh relaunch`, `fm-spawn.sh --relaunch`, and the crew-state recovery read all treated a shell-only pane as a live agent and refused recovery for as long as the record lived. The Herdr adapter now reads `pane process-info` plus the real process table through a shared harness-process classifier (bin/fm-agent-process-lib.sh, moved verbatim out of the tmux adapter so both backends mean the same thing by agent, shell, and other) before a registered agent counts as live. A registration over a shell-only pane is the new explicit `stale-agent` pane state, which the recovery-grade read maps to `dead`; husk detection, reclaim, presentation recovery, and session cleanup keep refusing it, so recovery reuses the pane and nothing gains close authority. A working record is verified the same way before the native busy verdict reports busy, so the recovery classifier never reports a shell-only pane as working. An unreadable process view reads unknown, trusting neither the registration nor its absence. Reproduced and measured on Herdr 0.9.0 with Pi 0.85.1 in an isolated lab; the new default-on live guard tests/fm-herdr-pi-stale-registration-live-e2e.test.sh exercises the real stale record, tests/fm-control-herdr-smoke.test.sh proves exit and relaunch through the control plane, and the portable suites pin the classifier over real processes. Fixes #4115. Duplicates: #3639, #3487, #2908, #3545. * no-mistakes(review): settle transient prompt helpers before trusting herdr process state * no-mistakes(review): drop stray codegraph file; read spaced comm whole in descendant walk * no-mistakes(review): untrack stray .codegraph/.gitignore * no-mistakes(review): untrack codegraph file; make spaced-path walk test discriminating * no-mistakes(review): untrack stray .codegraph/.gitignore * no-mistakes(review): untrack stray .codegraph/.gitignore re-added by fix round * no-mistakes(review): untrack stray .codegraph/.gitignore * no-mistakes(review): untrack codegraph file, drop dead control case, record process-info floor * no-mistakes(review): refuse stale-agent on fresh herdr spawn preflight Documented non-goal: fresh-spawn, reclaim, and presentation-recovery auto-recovery for a stale-agent pane is a separate design change, out of scope here, to be proposed upstream as its own issue if wanted. * no-mistakes(test): Fix herdr flake: don't misread transient empty foreground as unreadable * no-mistakes(document): Add fm-agent-process-lib.sh to scripts inventory * no-mistakes(fix): update remote herdr fixture to the real pane process-info shape The shared remote-secondmate herdr fixture still returned the old flat process-info body ({"result":{"process":{"name":...}}}). The process-level liveness classifier added for #4115 requires the real {"result":{"type":"pane_process_info","process_info":{...foreground_processes}}} shape and treated the old body as unreadable, so an already-launched remote endpoint's agent-state read failed and any relaunch attempt against it died with "remote endpoint state is unreadable; refusing duplicate launch" instead of reaching the state it was actually exercising (tests/fm-remote-secondmate-parent-binding.test.sh, tests/fm-remote-secondmate-lifecycle-e2e.test.sh). * no-mistakes(review): test: add empty-foreground regression test for herdr flake fix * no-mistakes(document): docs: register new stale-registration live-e2e test in herdr entry points --- bin/backends/herdr.sh | 267 +++++++++++++--- bin/backends/tmux.sh | 62 +--- bin/fm-agent-process-lib.sh | 106 +++++++ bin/fm-backend.sh | 18 +- bin/fm-crew-state.sh | 6 +- bin/fm-spawn.sh | 7 +- bin/fm-test-run.sh | 10 +- docs/herdr-backend.md | 18 +- docs/scripts.md | 1 + docs/tmux-backend.md | 1 + docs/verification/rovo.md | 2 +- docs/verification/runtime-backends.md | 79 ++++- tests/fm-backend-herdr.test.sh | 298 ++++++++++++++++++ tests/fm-control-herdr-smoke.test.sh | 111 ++++++- tests/fm-crew-state.test.sh | 65 +++- tests/fm-cursor-harness.test.sh | 12 +- ...fm-harness-liveness-drift-live-e2e.test.sh | 2 +- ...rdr-pi-stale-registration-live-e2e.test.sh | 159 ++++++++++ tests/fm-omp-harness.test.sh | 8 +- tests/fm-tmux-agent-liveness.test.sh | 6 +- tests/herdr-client-pair-fixture.sh | 4 + tests/remote-herdr-fixture.sh | 7 +- 22 files changed, 1110 insertions(+), 139 deletions(-) create mode 100644 bin/fm-agent-process-lib.sh create mode 100755 tests/fm-herdr-pi-stale-registration-live-e2e.test.sh diff --git a/bin/backends/herdr.sh b/bin/backends/herdr.sh index 2e1f2b98c4b..41254e26d0f 100644 --- a/bin/backends/herdr.sh +++ b/bin/backends/herdr.sh @@ -86,6 +86,13 @@ FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" # shellcheck source=bin/fm-transition-lib.sh . "$FM_BACKEND_HERDR_ROOT/bin/fm-transition-lib.sh" +# Shared, backend-neutral harness-process identity (bin/fm-agent-process-lib.sh): +# the same agent|shell|other vocabulary the tmux adapter proves liveness with, +# so a Herdr registration is verified against the pane's real processes by the +# same rule (fm_backend_herdr_pane_process_state). +# shellcheck source=bin/fm-agent-process-lib.sh +. "$FM_BACKEND_HERDR_ROOT/bin/fm-agent-process-lib.sh" + FM_BACKEND_HERDR_MIN_PROTOCOL=14 # events.subscribe (the native pane.agent_status_changed push stream) and its # subscription_event schema first shipped at protocol 16 (verified: herdr @@ -2058,37 +2065,190 @@ fm_backend_herdr_explicit_close_pane_confirmed() { # [ "$presence" = dead ] } +# fm_backend_herdr_pane_process_state: what the operating system says is +# running in , as one of agent|shell|other|unreadable, from `pane +# process-info` plus the real process table. This is the process-level proof +# fm_backend_herdr_pane_agent_state demands before it lets a registration count +# as a live agent (issue #4115), built on the same shape the tmux adapter uses: +# the foreground process group is authoritative, read through the shared +# classifier in bin/fm-agent-process-lib.sh. +# +# agent - a foreground process is a verified harness (any identity +# surface: kernel name, argv[0], or a node-bundle argument), or +# a verified harness is still a descendant of the pane shell +# outside the foreground group (suspended or backgrounded). A +# registered agent whose process still exists is never demoted. +# shell - every foreground process is a recognized shell AND no +# descendant of the pane shell is a verified harness: positive +# proof the pane is shell-only. The descendant walk is what makes +# this safe for the crew shape, where a nested `treehouse get` +# shell sits under the pane's top shell. +# other - the foreground group holds something that is neither: a tool +# the agent is running in its own process group, a pager, a +# stranger's process. Not a shell-only pane. An idle shell +# transiently hosts prompt helpers such as starship in its +# foreground group (the same shape the idle-shell proof settles +# on), so this verdict alone is resampled for the same bounded +# settle window and the first agent or shell reading wins; only +# an exhausted window keeps `other`. +# unreadable - process-info failed, described a different pane, named no +# shell pid, or the process table could not be read or does not +# contain the shell pid. An empty foreground-process list is NOT +# unreadable: it is the real, momentary shape of the exec-to- +# shell handoff (the harness process has exited but Herdr has +# not yet repopulated the foreground group), so it is treated +# like a shells-only foreground and settled by the same +# descendant-process check below. +# +# Verified on Herdr 0.9.0 (docs/verification/runtime-backends.md "Stale agent +# registration"): process-info's `.name` is the kernel process name (`node` for +# Pi, `zsh` for a shell), `.argv0` the argv[0] basename (`pi`), and `.argv` / +# `.cmdline` the full command line, so Pi is identified by argv[0] exactly as +# the tmux probe identifies it from `ps`. +fm_backend_herdr_pane_process_state() { # + local attempt=0 max_attempts=${FM_BACKEND_HERDR_IDLE_SHELL_PROOF_POLLS:-10} verdict + while :; do + verdict=$(fm_backend_herdr_pane_process_state_sample "$1" "$2") + [ "$verdict" = other ] || break + attempt=$((attempt + 1)) + [ "$attempt" -lt "$max_attempts" ] || break + sleep 0.1 + done + printf '%s' "$verdict" +} + +# fm_backend_herdr_pane_process_state_sample: one instantaneous observation +# for fm_backend_herdr_pane_process_state, which owns the verdict contract and +# the settle retry. +fm_backend_herdr_pane_process_state_sample() { # + local session=$1 pane_id=$2 info shell_pid count i pid name argv0 args verdict + local others=0 ps_bin rows + info=$(fm_backend_herdr_cli "$session" pane process-info --pane "$pane_id" 2>/dev/null) \ + || { printf 'unreadable'; return 0; } + printf '%s' "$info" | jq -e --arg pane "$pane_id" ' + .result.type == "pane_process_info" + and .result.process_info.pane_id == $pane + ' >/dev/null 2>&1 || { printf 'unreadable'; return 0; } + shell_pid=$(printf '%s' "$info" | jq -er \ + '.result.process_info.shell_pid | select(type == "number" and . > 1) | floor' 2>/dev/null) \ + || { printf 'unreadable'; return 0; } + count=$(printf '%s' "$info" | jq -er \ + '.result.process_info.foreground_processes | select(type == "array") | length' 2>/dev/null) \ + || { printf 'unreadable'; return 0; } + i=0 + while [ "$i" -lt "$count" ]; do + pid=$(printf '%s' "$info" | jq -r --argjson i "$i" \ + '.result.process_info.foreground_processes[$i].pid | select(type == "number") | floor' 2>/dev/null) + name=$(printf '%s' "$info" | jq -r --argjson i "$i" \ + '.result.process_info.foreground_processes[$i].name // empty' 2>/dev/null) + argv0=$(printf '%s' "$info" | jq -r --argjson i "$i" ' + .result.process_info.foreground_processes[$i] as $p + | (($p.argv // [])[0]) // $p.argv0 // empty' 2>/dev/null) + args=$(printf '%s' "$info" | jq -r --argjson i "$i" ' + .result.process_info.foreground_processes[$i] as $p + | $p.cmdline // (($p.argv // []) | join(" ")) // empty' 2>/dev/null) + verdict=$(fm_agent_process_classify "$name" "$argv0" "$args" "$pid") + case "$verdict" in + agent) printf 'agent'; return 0 ;; + shell) ;; + *) others=$((others + 1)) ;; + esac + i=$((i + 1)) + done + + # Nothing in the foreground is a harness. A foreground that is not purely + # shells is already `other`, whatever else the pane holds. Before calling a + # shells-only foreground a shell-only PANE, look for a harness that is still a + # descendant of the pane shell outside the foreground group; only its + # absence, read from the real process table, is proof of an agent-free pane. + [ "$others" -eq 0 ] || { printf 'other'; return 0; } + ps_bin=${FM_HERDR_PS_BIN:-ps} + command -v "$ps_bin" >/dev/null 2>&1 || { printf 'unreadable'; return 0; } + rows=$(LC_ALL=C "$ps_bin" -axo pid=,ppid=,comm= 2>/dev/null) || { printf 'unreadable'; return 0; } + printf '%s\n' "$rows" | awk -v shell="$shell_pid" '$1 == shell { found = 1 } END { exit(found ? 0 : 1) }' \ + || { printf 'unreadable'; return 0; } + while IFS=$'\t' read -r pid name; do + [ -n "$pid" ] || continue + args=$(LC_ALL=C "$ps_bin" -p "$pid" -o args= 2>/dev/null) || continue + args=${args#"${args%%[![:space:]]*}"} + argv0=${args%%[[:space:]]*} + if [ "$(fm_agent_process_classify "$name" "$argv0" "$args" "$pid")" = agent ]; then + printf 'agent' + return 0 + fi + done < in as one of -# dead|no-agent|live|unknown, purely from the JSON body of two read-only -# calls - never from process exit status, since a business-logic "not found" -# response is a normal, expected outcome here, not a call failure (real herdr -# 0.7.1 exits 1 for it; the canned-response test fakes exit 0; parsing only -# the JSON keeps this function correct against either). +# dead|no-agent|stale-agent|live|unknown, from the JSON body of two read-only +# calls plus, for a registered agent, the pane's process-level view - never +# from process exit status, since a business-logic "not found" response is a +# normal, expected outcome here, not a call failure (real herdr 0.7.1 exits 1 +# for it; the canned-response test fakes exit 0; parsing only the JSON keeps +# this function correct against either). # -# dead - `pane get` responds with error code pane_not_found: the pane -# itself is gone (closed, or its process died and herdr already -# reaped it - verified empirically: killing a pane's shell pid -# on a live server makes herdr immediately drop both the pane -# and its tab from `pane get`/`tab list`). -# no-agent - `pane get` succeeds (the pane structurally exists) but `agent -# get` responds with error code agent_not_found: nothing is -# registered in it - exactly what a herdr session-layout restore -# produces (verified empirically: `session stop` + fresh `herdr -# server` restart leaves the pane alive, agent_status "unknown", -# agent get -> agent_not_found - docs/herdr-backend.md "ID -# stability across a server restart"), and what a future -# `resume_agents_on_restore = false` restore would produce too -# (a plain shell, never an agent). -# live - `agent get` succeeds and reports a real agent_status (working, -# idle, done, or blocked - any registered value). An idle or -# blocked agent is still a genuine, still-registered agent, not -# a restored husk, so it is never a close-and-replace candidate. -# unknown - anything else: an unparseable/unexpected response from either -# call, or a `pane get` success whose own echoed pane_id does not -# round-trip (guards against misreading a herdr response shape -# change as "the pane exists"). The caller must fail safe toward -# refusal here, never toward closing - this is the conservative -# backstop the husk check depends on. +# dead - `pane get` responds with error code pane_not_found: the pane +# itself is gone (closed, or its process died and herdr already +# reaped it - verified empirically: killing a pane's shell pid +# on a live server makes herdr immediately drop both the pane +# and its tab from `pane get`/`tab list`). +# no-agent - `pane get` succeeds (the pane structurally exists) but `agent +# get` responds with error code agent_not_found: nothing is +# registered in it - exactly what a herdr session-layout restore +# produces (verified empirically: `session stop` + fresh `herdr +# server` restart leaves the pane alive, agent_status "unknown", +# agent get -> agent_not_found - docs/herdr-backend.md "ID +# stability across a server restart"), and what a future +# `resume_agents_on_restore = false` restore would produce too +# (a plain shell, never an agent). +# stale-agent - `agent get` reports a registered agent_status (working, idle, +# done, or blocked) but fm_backend_herdr_pane_process_state +# proves the pane is shell-only: the registered agent's process +# has exited and Herdr kept its registration (issue #4115; +# Herdr does not release a Pi registration on TUI shutdown when +# a nested shell sits under the pane's top shell, the crew +# shape). This is the explicit agent-free reason: the pane is +# recoverable, and the record it carries is not evidence of a +# running agent. No registered status outranks the process +# view, because a killed mid-turn agent leaves `working` +# behind just as a quit one leaves `idle`. +# live - `agent get` succeeds with a registered agent_status and the +# process-level view is `agent` or `other`: a harness process +# is running, or something that is not a bare shell is, so the +# registration keeps its authority. An idle or blocked agent +# is still a genuine, still-registered agent, not a restored +# husk, so it is never a close-and-replace candidate. +# unknown - anything else: an unparseable/unexpected response from +# either call, a `pane get` success whose own echoed pane_id +# does not round-trip (guards against misreading a herdr +# response shape change as "the pane exists"), or a registered +# agent whose process-level view is unreadable - the +# registration alone is no longer trusted, and its absence is +# not claimed either. The caller must fail safe toward refusal +# here, never toward closing - this is the conservative +# backstop the husk check depends on. fm_backend_herdr_pane_agent_state() { # local session=$1 pane_id=$2 out code presence status presence=$(fm_backend_herdr_pane_presence_state "$session" "$pane_id") @@ -2107,16 +2267,23 @@ fm_backend_herdr_pane_agent_state() { # fi status=$(printf '%s' "$out" | jq -r '.result.agent.agent_status // empty' 2>/dev/null) case "$status" in - working|idle|done|blocked) printf 'live' ;; + working|idle|done|blocked) ;; + *) printf 'unknown'; return 0 ;; + esac + case "$(fm_backend_herdr_pane_process_state "$session" "$pane_id")" in + agent|other) printf 'live' ;; + shell) printf 'stale-agent' ;; *) printf 'unknown' ;; esac } # fm_backend_herdr_tab_is_husk: true (0) only for the two conservative husk # states (dead, no-agent) fm_backend_herdr_pane_agent_state can positively -# confirm; live and unknown both refuse (1), so an inconclusive read never -# licenses closing anything. Restored-layout recovery depends on this -# fail-safe-toward-refusal behavior. +# confirm; live, stale-agent, and unknown all refuse (1), so an inconclusive +# read never licenses closing anything, and a stale registration - agent-free +# for RECOVERY, which reuses the pane - still never licenses closing it, because +# the shell it holds may be a nested worktree shell. Restored-layout recovery +# depends on this fail-safe-toward-refusal behavior. fm_backend_herdr_tab_is_husk() { # case "$(fm_backend_herdr_pane_agent_state "$1" "$2")" in dead|no-agent) return 0 ;; @@ -2152,8 +2319,10 @@ fm_backend_herdr_server_running_state() { # # fm_backend_herdr_agent_state: recovery-grade state for the same session-start # sweep as the tmux classifier. It reuses the husk classifier rather than # creating a second Herdr state machine: a structurally gone pane is `missing`, -# a confirmed agent-less pane is `dead`, a registered agent is `alive`, and an -# unexpected or failed API read is `unreadable`. +# a confirmed agent-less pane is `dead` - whether nothing is registered or a +# registration lingers over a shell-only pane (stale-agent, issue #4115) - a +# registered agent with a live process is `alive`, and an unexpected or failed +# API read is `unreadable`. # # One exception to that last case, and it is deliberately made HERE rather than # in the husk classifier: a read can fail because the recorded session's server @@ -2173,7 +2342,7 @@ fm_backend_herdr_agent_state() { # fm_backend_herdr_parse_target "$target" || { printf 'unreadable'; return 0; } case "$(fm_backend_herdr_pane_agent_state "$FM_BACKEND_HERDR_SESSION" "$FM_BACKEND_HERDR_PANE")" in dead) printf 'missing' ;; - no-agent) printf 'dead' ;; + no-agent|stale-agent) printf 'dead' ;; live) printf 'alive' ;; *) case "$(fm_backend_herdr_server_running_state "$FM_BACKEND_HERDR_SESSION")" in @@ -2511,7 +2680,7 @@ fm_backend_herdr_projection_reclaim_rollback() { # case "$state" in dead) return 0 ;; no-agent) ;; - live|unknown) return 1 ;; + live|stale-agent|unknown) return 1 ;; esac fm_backend_herdr_projection_close_pane_focus_preserving "$session" "$new_pane" no-agent || return 1 [ "$(fm_backend_herdr_pane_agent_state "$session" "$new_pane")" = dead ] @@ -2563,7 +2732,7 @@ fm_backend_herdr_projection_reclaim_task() { # &2 return 2 ;; - live|unknown) + live|stale-agent|unknown) echo "error: exact herdr presentation pane for $id is $state; refusing duplicate launch" >&2 return 1 ;; @@ -2612,7 +2781,7 @@ fm_backend_herdr_projection_reclaim_task() { # &2 return 1 @@ -2635,7 +2804,7 @@ fm_backend_herdr_projection_reclaim_task() { # &2 return 1 ;; @@ -2717,7 +2886,7 @@ fm_backend_herdr_projection_recovery_allows_flat() { # &2 return 1 ;; @@ -3244,10 +3413,22 @@ fm_backend_herdr_agent_status_raw() { # # gets real semantics" per the design report. See # fm_backend_herdr_classify_agent_status for the status->busy/idle/unknown # mapping. +# +# A `busy` verdict is proven at process level before it is reported: a +# lingering `working` registration over a shell-only pane (an agent killed +# mid-turn, issue #4115) reads `unknown`, never busy, so the recovery classifier +# cannot report a shell-only pane as working. Only the busy case pays the extra +# process read; idle and unknown are never trusted as busy by any consumer. fm_backend_herdr_busy_state() { # + local verdict fm_backend_herdr_target_ready "$1" || { printf 'unknown'; return 0; } - fm_backend_herdr_classify_agent_status \ - "$(fm_backend_herdr_agent_status_raw "$FM_BACKEND_HERDR_SESSION" "$FM_BACKEND_HERDR_PANE")" + verdict=$(fm_backend_herdr_classify_agent_status \ + "$(fm_backend_herdr_agent_status_raw "$FM_BACKEND_HERDR_SESSION" "$FM_BACKEND_HERDR_PANE")") + if [ "$verdict" = busy ] \ + && [ "$(fm_backend_herdr_pane_process_state "$FM_BACKEND_HERDR_SESSION" "$FM_BACKEND_HERDR_PANE")" = shell ]; then + verdict=unknown + fi + printf '%s' "$verdict" } # fm_backend_herdr_wait_for_working: poll :'s NATIVE diff --git a/bin/backends/tmux.sh b/bin/backends/tmux.sh index 42a87fcc49d..4477eb97423 100644 --- a/bin/backends/tmux.sh +++ b/bin/backends/tmux.sh @@ -22,10 +22,8 @@ . "$FM_BACKEND_LIB_DIR/fm-tmux-lib.sh" # shellcheck source=bin/fm-session-lock-lib.sh . "$FM_BACKEND_LIB_DIR/fm-session-lock-lib.sh" -# shellcheck source=bin/fm-cursor-lib.sh -. "$FM_BACKEND_LIB_DIR/fm-cursor-lib.sh" -# shellcheck source=bin/fm-gemini-lib.sh -. "$FM_BACKEND_LIB_DIR/fm-gemini-lib.sh" +# shellcheck source=bin/fm-agent-process-lib.sh +. "$FM_BACKEND_LIB_DIR/fm-agent-process-lib.sh" # fm_backend_tmux_resolve_bare_selector: the live-window-listing fallback for a # selector that is neither an explicit target nor a task selector routed @@ -154,50 +152,10 @@ fm_backend_tmux_current_command() { # tmux display-message -p -t "$1" '#{pane_current_command}' 2>/dev/null } -# fm_backend_tmux_classify_process_name: the single owner of the process-name -# vocabulary shared by every liveness signal below - `agent` for a verified -# harness, `shell` for an idle login/interactive shell, `other` for anything -# else. Keeping one classifier means the two independent name sources can never -# drift into disagreeing about what a given name means. -fm_backend_tmux_classify_process_name() { # [argv0] -> agent|shell|other - local path=$1 argv0=${2:-} base - base=${path##*/} - base=${base#-} - case "$base" in - # muse is anchored rather than globbed like its neighbours: its installed - # binary is muse-bin- (the launcher execs it, so the version is the - # live process name and changes on every auto-update), and unlike `claude` or - # `codex` the substring `muse` is a common English fragment - a *muse* glob - # would classify musescore or amuse as a live agent pane. The install path - # cannot carry it either: ~/.local/bin/muse-bin- has no `muse` path - # COMPONENT, so the fm_harness_path_name fallback below never fires for it. - muse|muse-bin-*) printf 'agent' ;; - # omp (Oh My Pi) is anchored for the same reason as muse: its live process - # name is the bare word `omp` (verified, omp 18.1.11) and a glob would claim - # unrelated commands such as ompd or comp. - *claude*|*codex*|*opencode*|*grok*|*kimi*|*rovo*|pi|pi-signed|pi-launcher|Pi|omp) printf 'agent' ;; - zsh|bash|sh|dash|ash|ksh|mksh|tcsh|csh|fish) printf 'shell' ;; - *) - if fm_harness_path_name "$path" >/dev/null || fm_harness_path_name "$argv0" >/dev/null; then - printf 'agent' - # cursor-agent runs as a bundled node script, so tmux reports the pane - # command as a bare `node` that no name pattern above can own, and its - # other installed name is the far-too-generic `agent` (verified live on - # cursor-agent 2026.08.11-e8db854: #{pane_current_command} is `node` while - # `ps -o comm=` carries the cursor-agent install path). Identity therefore - # comes from the narrowed structural rule in bin/fm-cursor-lib.sh, which - # demands Cursor's own name or install tree in the path or argv[0]. An - # unrelated `node` or `agent` matches nothing here and stays `other`, - # which the callers above fold into `ambiguous` rather than `dead`, so a - # stranger's node pane is never reported as an agent-free pane. - elif fm_cursor_process_matches "${path:-$argv0}" '' "$argv0"; then - printf 'agent' - else - printf 'other' - fi - ;; - esac -} +# The process-name classifier every liveness signal below feeds +# (fm_agent_process_classify_name) is owned by bin/fm-agent-process-lib.sh, +# shared with the Herdr adapter so both backends mean the same thing by +# `agent`, `shell`, and `other`. # fm_backend_tmux_foreground_comms: the kernel-side names of every process in # 's pane tty foreground process group, one full value per line. @@ -330,7 +288,7 @@ fm_backend_tmux_agent_state() { # while IFS= read -r name; do [ -n "$name" ] || continue fg_seen=1 - case "$(fm_backend_tmux_classify_process_name "$name")" in + case "$(fm_agent_process_classify_name "$name")" in agent) printf 'alive'; return 0 ;; shell) fg_shell=1 ;; *) fg_other=1 ;; @@ -342,7 +300,7 @@ EOF argv0s=$(fm_backend_tmux_foreground_argv0s "$target") while IFS= read -r name; do [ -n "$name" ] || continue - if [ "$(fm_backend_tmux_classify_process_name '' "$name")" = agent ]; then + if [ "$(fm_agent_process_classify_name '' "$name")" = agent ]; then printf 'alive' return 0 fi @@ -379,7 +337,7 @@ EOF printf 'unreadable' return 0 } - if [ "$(fm_backend_tmux_classify_process_name "$comm")" = agent ]; then + if [ "$(fm_agent_process_classify_name "$comm")" = agent ]; then printf 'alive' return 0 fi @@ -398,7 +356,7 @@ EOF case "$comm" in '') printf 'unreadable'; return 0 ;; esac - case "$(fm_backend_tmux_classify_process_name "$comm")" in + case "$(fm_agent_process_classify_name "$comm")" in shell) printf 'dead' ;; *) printf 'ambiguous' ;; esac diff --git a/bin/fm-agent-process-lib.sh b/bin/fm-agent-process-lib.sh new file mode 100644 index 00000000000..5943f1b2749 --- /dev/null +++ b/bin/fm-agent-process-lib.sh @@ -0,0 +1,106 @@ +#!/usr/bin/env bash +# Backend-neutral harness-process identity. +# Sourced by bin/backends/tmux.sh and bin/backends/herdr.sh. This file is +# sourced by scripts and has no side effects on source. +# +# Why one owner: every runtime backend that proves an agent is alive does it by +# attributing operating-system processes - the pane's foreground process group +# on tmux, Herdr's `pane process-info` view plus the pane shell's descendants +# on Herdr - and the two must agree on what a given process name means, or a +# harness one backend recognizes silently reads as a dead pane on the other. +# The classifier moved here verbatim from the tmux adapter, where it was born; +# docs/tmux-backend.md "Agent liveness probe" owns the empirical basis for the +# names below, and tests/fm-tmux-agent-liveness.test.sh plus +# tests/fm-harness-liveness-drift-live-e2e.test.sh keep them honest. + +# shellcheck source=bin/fm-session-lock-lib.sh +. "$(dirname -- "${BASH_SOURCE[0]}")/fm-session-lock-lib.sh" +# shellcheck source=bin/fm-gemini-lib.sh +. "$(dirname -- "${BASH_SOURCE[0]}")/fm-gemini-lib.sh" + +# fm_agent_process_classify_name: the single owner of the process-name +# vocabulary shared by every liveness signal - `agent` for a verified harness, +# `shell` for an idle login/interactive shell, `other` for anything else. +# Keeping one classifier means independent name sources (a kernel process +# name, an argv[0], a rendered pane title) can never drift into disagreeing +# about what a given name means. +fm_agent_process_classify_name() { # [argv0] -> agent|shell|other + local path=$1 argv0=${2:-} base + base=${path##*/} + base=${base#-} + case "$base" in + # muse is anchored rather than globbed like its neighbours: its installed + # binary is muse-bin- (the launcher execs it, so the version is the + # live process name and changes on every auto-update), and unlike `claude` or + # `codex` the substring `muse` is a common English fragment - a *muse* glob + # would classify musescore or amuse as a live agent pane. The install path + # cannot carry it either: ~/.local/bin/muse-bin- has no `muse` path + # COMPONENT, so the fm_harness_path_name fallback below never fires for it. + muse|muse-bin-*) printf 'agent' ;; + # omp (Oh My Pi) is anchored for the same reason as muse: its live process + # name is the bare word `omp` (verified, omp 18.1.11) and a glob would claim + # unrelated commands such as ompd or comp. + *claude*|*codex*|*opencode*|*grok*|*kimi*|*rovo*|pi|pi-signed|pi-launcher|Pi|omp) printf 'agent' ;; + zsh|bash|sh|dash|ash|ksh|mksh|tcsh|csh|fish) printf 'shell' ;; + *) + if fm_harness_path_name "$path" >/dev/null || fm_harness_path_name "$argv0" >/dev/null; then + printf 'agent' + # cursor-agent runs as a bundled node script, so tmux reports the pane + # command as a bare `node` that no name pattern above can own, and its + # other installed name is the far-too-generic `agent` (verified live on + # cursor-agent 2026.08.11-e8db854: #{pane_current_command} is `node` while + # `ps -o comm=` carries the cursor-agent install path). Identity therefore + # comes from the narrowed structural rule in bin/fm-cursor-lib.sh, which + # demands Cursor's own name or install tree in the path or argv[0]. An + # unrelated `node` or `agent` matches nothing here and stays `other`, + # which the callers fold into `ambiguous` rather than `dead`, so a + # stranger's node pane is never reported as an agent-free pane. + elif fm_cursor_process_matches "${path:-$argv0}" '' "$argv0"; then + printf 'agent' + else + printf 'other' + fi + ;; + esac +} + +# fm_agent_process_classify: one process, from every identity surface a +# backend can hand over, as agent|shell|other. Any single surface naming a +# verified harness carries `agent`, because a false negative is the one outcome +# that launches a duplicate agent onto a live worktree; `shell` needs every +# readable surface to agree the process is a shell; anything else is `other`. +# +# the kernel process name (ps comm, or Herdr's process-info .name): +# on Linux the exec name, on macOS argv[0] truncated to 16 bytes. +# argv[0] as the process reports it - a bare name or an install +# path, whichever the launcher used (empty when unknown). +# the flattened command line, read only for the node-bundle +# harnesses whose identity sits in argv[1] (bin/fm-gemini-lib.sh). +# [pid] when given, lets the Gemini rule read argv boundaries from the +# live process instead of the flattened line. +fm_agent_process_classify() { # [pid] -> agent|shell|other + local name=${1:-} argv0=${2:-} args=${3:-} pid=${4:-} by_name by_argv0 + by_name=$(fm_agent_process_classify_name "$name" "$argv0") + [ "$by_name" != agent ] || { printf 'agent'; return 0; } + if [ -n "$argv0" ]; then + # argv[0] is classified as a path in its own right, so a bare `pi` or a + # `-zsh` login name reads by basename and an install path by component. + by_argv0=$(fm_agent_process_classify_name "$argv0" "$argv0") + [ "$by_argv0" != agent ] || { printf 'agent'; return 0; } + else + by_argv0=$by_name + fi + if [ -n "$pid" ] && fm_gemini_pid_is_gemini "$pid"; then + printf 'agent' + return 0 + fi + if [ -n "$args" ] && fm_gemini_args_are_gemini "$args"; then + printf 'agent' + return 0 + fi + if [ "$by_name" = shell ] && [ "$by_argv0" = shell ]; then + printf 'shell' + else + printf 'other' + fi +} diff --git a/bin/fm-backend.sh b/bin/fm-backend.sh index c744e3557a1..bd41f1fe9d9 100644 --- a/bin/fm-backend.sh +++ b/bin/fm-backend.sh @@ -884,13 +884,17 @@ fm_backend_target_exists() { # [expected-label] # ambiguous - the endpoint exists but its process cannot be attributed. # unreadable - a target or inventory read failed or contradicted itself. # unverified - this backend has no recovery classifier. -# Only `dead` and `missing` license recovery. The tmux adapter requires a -# successful session inventory and returns `missing` only when it omits the -# exact window; the Herdr adapter reuses its strict husk classifier, then maps -# a positively stopped session server to `missing` only in this recovery-grade -# view. Zellij remains unverified because its secondmate ghost-tab and -# agent-process recovery path has not been empirically validated. Orca and cmux -# do not support secondmate spawns. +# Only `dead` and `missing` license recovery. Every `alive` is proven at +# process level through the shared classifier in bin/fm-agent-process-lib.sh, +# never from a registration or a rendered title alone. The tmux adapter +# requires a successful session inventory and returns `missing` only when it +# omits the exact window; the Herdr adapter reuses its strict husk classifier - +# which verifies a registered agent against `pane process-info` and the real +# process table, so a registration Herdr kept over a shell-only pane reads +# `dead` here (issue #4115) - then maps a positively stopped session server to +# `missing` only in this recovery-grade view. Zellij remains unverified because +# its secondmate ghost-tab and agent-process recovery path has not been +# empirically validated. Orca and cmux do not support secondmate spawns. fm_backend_agent_state() { # local backend=$1 target=$2 fm_backend_source "$backend" || { printf 'unverified'; return 0; } diff --git a/bin/fm-crew-state.sh b/bin/fm-crew-state.sh index aaaf0f8c6b8..1512cf83c0a 100755 --- a/bin/fm-crew-state.sh +++ b/bin/fm-crew-state.sh @@ -794,8 +794,10 @@ if ! pane_readable "$BACKEND_TARGET"; then # genuine server death - a socket-connection failure is NOT # covered by the unknown-never-death rule above). # dead - the endpoint exists but confidently has no agent (herdr's agent - # get answered agent_not_found; tmux's readable foreground process - # group is nothing but shells), still positive death evidence. + # get answered agent_not_found, or its registration lingers over a + # pane whose processes are nothing but shells - issue #4115; + # tmux's readable foreground process group is nothing but + # shells), still positive death evidence. # alive - the endpoint and its agent answered and only the heavy # scrollback read failed, so the live state is classified by the # normal flow below instead of being discarded. diff --git a/bin/fm-spawn.sh b/bin/fm-spawn.sh index cc73ec64fb2..35a701c7041 100755 --- a/bin/fm-spawn.sh +++ b/bin/fm-spawn.sh @@ -2539,8 +2539,13 @@ herdr_projection_existing_meta_allows_flat() { # } old_state=$(fm_backend_herdr_pane_agent_state "$old_session" "$old_pane") case "$old_state" in + # A stale registration over a shell-only pane is agent-free for RECOVERY + # (--relaunch reuses the pane, issue #4115), but the duplicate-launch + # corridor keeps refusing it like every other non-husk state, so a fresh + # spawn is refused here consistently with the reclaim and presentation + # gates downstream. dead|no-agent) return 0 ;; - live|unknown) + live|stale-agent|unknown) echo "error: existing herdr endpoint for $ID is $old_state; refusing duplicate launch" >&2 return 1 ;; diff --git a/bin/fm-test-run.sh b/bin/fm-test-run.sh index fcfa4b4f43e..ab76526d8a2 100755 --- a/bin/fm-test-run.sh +++ b/bin/fm-test-run.sh @@ -339,6 +339,7 @@ family_for_basename() { fm-harness-liveness-drift-live-e2e.test.sh|\ fm-muse-signals-live-e2e.test.sh|fm-rovo-signals-live-e2e.test.sh|\ fm-herdr-version-floor-live-e2e.test.sh|\ + fm-herdr-pi-stale-registration-live-e2e.test.sh|\ fm-opencode-primary-live-e2e.test.sh|fm-pi-branch-live-e2e.test.sh|\ fm-pi-branch-responsiveness-live-e2e.test.sh|\ fm-pi-primary-live-e2e.test.sh|fm-pi-codex-native.test.sh|fm-omp-primary-live-e2e.test.sh|\ @@ -1309,7 +1310,7 @@ families_for_changed_path() { # runner's logic is right, not that the suite it drives still runs. printf '%s\n' pure-contract-unit ;; - bin/backends/herdr*|bin/fm-herdr-lab.sh|tests/herdr-test-safety.sh) + bin/backends/herdr*|bin/fm-herdr-lab.sh|tests/herdr-test-safety.sh|tests/herdr-client-pair-fixture.sh) printf '%s\n' real-herdr-gated printf '%s\n' backend-dispatch printf '%s\n' pure-contract-unit @@ -1335,6 +1336,13 @@ families_for_changed_path() { printf '%s\n' backend-dispatch printf '%s\n' real-herdr-gated ;; + bin/fm-agent-process-lib.sh) + # The shared harness-process classifier feeds both the tmux and Herdr + # liveness verdicts, so a change to it is proven by both backends' suites. + printf '%s\n' backend-dispatch + printf '%s\n' real-herdr-gated + printf '%s\n' pure-contract-unit + ;; bin/fm-watch*|bin/fm-wake*|bin/fm-inactive-reconcile.sh|\ bin/fm-classify-lib.sh|bin/fm-daemon*|bin/fm-turnend-guard*|bin/fm-guard.sh) printf '%s\n' watcher-wake-lock diff --git a/docs/herdr-backend.md b/docs/herdr-backend.md index 7b79cd77c4b..316643fdda2 100644 --- a/docs/herdr-backend.md +++ b/docs/herdr-backend.md @@ -282,10 +282,19 @@ A restored same-labeled tab with a missing pane or no registered agent is a husk Create replaces only a confidently dead or no-agent husk, creates the replacement before closing the old tab, and refuses live or unknown states. This prevents closing the workspace's last tab before a replacement exists. -The generic Herdr agent-liveness probe reuses the same pane classifier, then applies one recovery-only exception. -A structurally gone pane or a pane read from a session positively reported as having no running server becomes `missing`, a restored agent-less shell becomes `dead`, a registered agent becomes `alive`, and every other unexpected read becomes `unreadable`. -The stopped-server exception does not widen husk detection or any close authority; those paths still refuse an unreadable pane. -Unlike tmux process-name inspection, native registration can classify Pi without guessing from a generic interpreter name. +A registration alone never proves an agent. +Herdr keeps a Pi registration (`agent get` still reports `agent=pi` with its last status) after the Pi process has exited to a plain shell whenever a nested interactive shell sits under the pane's top shell, which is the crew shape `treehouse get` leaves behind (measured on Herdr 0.9.0 - [verification](verification/runtime-backends.md) "Stale agent registration"; upstream issue #4115). +So before a registered agent counts as live, the pane classifier reads `pane process-info` and the real process table through the shared harness-process classifier in `bin/fm-agent-process-lib.sh`, the same rule the tmux adapter proves liveness with: a harness in the foreground process group, or still a descendant of the pane shell, keeps the registration live; a foreground that is nothing but shells with no harness descendant is a `stale-agent` pane, agent-free with that explicit reason; a foreground holding anything else keeps the registration live, but only after the same bounded settle window the idle-shell proof uses, because an idle shell transiently hosts prompt helpers such as starship in its foreground group and the first agent or shell sample in that window decides; an unreadable process view makes the pane `unknown`, trusting neither the registration nor its absence. +No registered status outranks the process view, because an agent killed mid-turn leaves `working` behind just as a quit one leaves `idle`, and the native busy verdict is verified the same way so a shell-only pane never reads busy. +The `pane process-info` subcommand that this process-level proof depends on is present in every supported release client from the 0.7.1 floor upward (measured 2026-09-10 on the pinned 0.7.1, 0.7.3, 0.7.4, and 0.7.5 release clients - [verification](verification/runtime-backends.md) "Stale agent registration"). +The response shape the adapter parses (`result.type` of `pane_process_info`, `process_info.shell_pid`, and `foreground_processes` entries carrying `name`, `argv0`, `argv`, and `cmdline`) is verified live only on Herdr 0.9.0, with the idle-shell proof's narrower parse previously verified on 0.7.5. +A server response below 0.9.0 has not been measured for this parse. +An unreadable or unparseable process view reads `unknown`, which refuses lifecycle verbs and recovery rather than trusting the registration. + +The generic Herdr agent-liveness probe reuses that pane classifier, then applies one recovery-only exception. +A structurally gone pane or a pane read from a session positively reported as having no running server becomes `missing`, a restored agent-less shell and a stale registration over a shell-only pane both become `dead`, a registered agent with a live process becomes `alive`, and every other unexpected read becomes `unreadable`. +Neither the stopped-server exception nor the stale-registration verdict widens husk detection or any close authority; those paths still refuse an unreadable pane, and a `stale-agent` pane is reused by recovery, never closed as a husk, because the shell it holds may be a nested worktree shell. +Native registration still identifies Pi by name where tmux would see a generic interpreter; the process-level proof only decides whether that registration is backed by a running process. `tests/fm-backend-herdr-agent-exit-shell-e2e.test.sh` pins the live-Pi versus leftover-shell distinction; [`verification/runtime-backends.md`](verification/runtime-backends.md#agent-lifecycle-control) owns the versioned evidence. The session-start sweep uses this probe. @@ -355,6 +364,7 @@ tests/fm-backend-herdr-workspace-per-home-e2e.test.sh tests/fm-backend-herdr-launcher-workspace-e2e.test.sh tests/fm-backend-herdr-presentation-e2e.test.sh tests/fm-backend-herdr-agent-exit-shell-e2e.test.sh +tests/fm-herdr-pi-stale-registration-live-e2e.test.sh tests/fm-backend-herdr-eventwait-smoke.test.sh tests/fm-control-herdr-smoke.test.sh tests/fm-herdr-session-cleanup.test.sh diff --git a/docs/scripts.md b/docs/scripts.md index e323c515eed..9b149e538c7 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -60,6 +60,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-backend.sh` | Runtime-backend selection, meta helpers, selector resolution, and operation dispatch | | `fm-backend-hometag-lib.sh` | Shared per-installation home-tag derivation for zellij tab and cmux workspace titles | | `fm-composer-lib.sh` | Single fleet-wide owner of composer shapes, capability-aware screen classification, and verdicts | +| `fm-agent-process-lib.sh` | Backend-neutral harness-process name classifier shared by the tmux and herdr adapters | | `backends/tmux.sh` | Verified tmux session-provider adapter | | `backends/herdr.sh` | Herdr session-provider adapter with its own required CI lane | | `backends/zellij.sh` | Experimental zellij session-provider adapter | diff --git a/docs/tmux-backend.md b/docs/tmux-backend.md index 3da0930635e..a5b4fd9c8ff 100644 --- a/docs/tmux-backend.md +++ b/docs/tmux-backend.md @@ -49,6 +49,7 @@ Verify setup by spawning a small task and confirming its `fm-` window appear A target-existence check proves only that the pane exists. The deeper tmux agent-liveness probe first verifies exact window membership, then reads process names to distinguish a running harness from a bare idle shell. It classifies recognized Claude, Codex, OpenCode, Pi, pi-signed, Grok, Kimi, Cursor, Muse, and Rovo process identities as `alive`, common shells as `dead`, an authoritatively absent window as `missing`, unreadable state as `unreadable`, and every other process as `ambiguous`. +The process-name vocabulary behind those verdicts is owned by `bin/fm-agent-process-lib.sh` and shared with the Herdr adapter, which proves a registered agent against the same names ([herdr-backend.md](herdr-backend.md) "Restart and liveness behavior"). Only `dead` and `missing` authorize recovery because a false dead result could launch a duplicate agent. For positive attribution, the probe combines two independent name sources rather than making either one load-bearing. diff --git a/docs/verification/rovo.md b/docs/verification/rovo.md index 8a588e3ad8f..2d6c722f1d6 100644 --- a/docs/verification/rovo.md +++ b/docs/verification/rovo.md @@ -184,7 +184,7 @@ $ ls "$LAB/outside/inbox/handled" ## Backend liveness: tmux verified live, herdr placement verified live with a herdr-side agent-detection gap tmux 3.6a is now installed and was exercised live in an isolated `tmux -L ` session, so tmux pane liveness is fully verified rather than pending. -`bin/backends/tmux.sh`'s `fm_backend_tmux_classify_process_name` matches `*rovo*` alongside the other globbed harness names, so a rovo pane classifies `agent` (not `other`). +`bin/fm-agent-process-lib.sh`'s `fm_agent_process_classify_name` (then still inside `bin/backends/tmux.sh`) matches `*rovo*` alongside the other globbed harness names, so a rovo pane classifies `agent` (not `other`). The two independent name sources behaved as designed: `#{pane_current_command}` reported the truncated on-disk binary name `atlassian_cli_r` - macOS's 15-char `comm` truncation cuts `atlassian_cli_rovodev` off just before the `rovo` substring begins, the same truncation-volatility class [`runtime-backends.md`](runtime-backends.md) already documents for codex/kimi's own patch-release name drift - while the foreground ps-based `comm` correctly reported `rovo`, and `fm_backend_tmux_agent_state` correctly returned `alive` through that primary source. The two-independent-name-sources design is exactly why the truncation quirk does not break the verdict. `tmux capture-pane` correctly rendered the box composer and the `Rovo is thinking...` busy line while a real `sleep`-based bash tool call ran; `fm_busy_rovo_tail_busy` classified it busy, then idle once the tool call completed and the reply landed. The Escape/`Agent cancelled` evidence in the interrupt section above was captured in this same live tmux session. `/exit` closed the tmux window cleanly, and `fm_backend_tmux_agent_state` reported `missing` immediately afterward - a clean, unambiguous exit verdict. diff --git a/docs/verification/runtime-backends.md b/docs/verification/runtime-backends.md index acaaf0708f7..200a5974131 100644 --- a/docs/verification/runtime-backends.md +++ b/docs/verification/runtime-backends.md @@ -63,7 +63,7 @@ The crewmate-only Muse Code 0.1.0-R708.1 adapter was verified separately on 2026 Its installed `muse-bin-0.1.0-R708.1` foreground identity classified `alive`, while `musescore`, `amuse`, `muse-binary`, and `muse-bind` remained ambiguous in the portable regression. [`muse.md`](muse.md#process-identity) owns the artifact identity and launcher evidence for that verification. -The crewmate/scout-only Rovo CLI 202609.1.2 adapter added `*rovo*` to the same glob family as `*grok*`/`*kimi*` in `fm_backend_tmux_classify_process_name`, and was relaunched live under tmux 3.6a in an isolated private socket. +The crewmate/scout-only Rovo CLI 202609.1.2 adapter added `*rovo*` to the same glob family as `*grok*`/`*kimi*` in the shared process-name classifier (now `fm_agent_process_classify_name` in `bin/fm-agent-process-lib.sh`), and was relaunched live under tmux 3.6a in an isolated private socket. `#{pane_current_command}` reported the truncated on-disk binary name `atlassian_cli_r` - macOS's 15-char `comm` truncation cuts `atlassian_cli_rovodev` off just before the `rovo` substring begins, the same truncation-volatility class codex/kimi's own patch-release name drift shows above - while the foreground ps-based `comm` correctly reported `rovo`, so `fm_backend_tmux_agent_state` returned `alive` through that primary source; the two-independent-name-sources design is exactly why the truncated title does not break the verdict. [`rovo.md`](rovo.md#backend-liveness-tmux-verified-live-herdr-placement-verified-live-with-a-herdr-side-agent-detection-gap) owns the fuller record, including the busy/interrupt/exit facts captured in that same live tmux session and the herdr agent-detection gap found when herdr placement was verified live in an isolated lab session. @@ -1012,17 +1012,22 @@ Herdr is one of the two backends whose recovery-grade agent-state classifier the tests/fm-control-herdr-smoke.test.sh ``` -Observed output: +Observed output, refreshed 2026-09-10 on Herdr 0.9.0 after the stale-registration fix (the two stale-registration lines are recorded under "Stale agent registration" below): ```text ok - real herdr: exit on a pane with no registered agent is idempotent success +ok - real herdr 0.9.0: a gone session reads recoverable while a live pane and a malformed target do not +ok - real herdr: a drifted agent-free shell returns to its worktree and reuses the same endpoint ok - real herdr: interrupt refuses when herdr's own agent registry reports no agent ok - real herdr: interrupt delivers the harness's key and proves the agent survived it ok - real herdr: no control verb removed the endpoint or the task's local copy +ok - real herdr 0.9.0: a registration Herdr keeps after its agent exits reads stale-agent and recovers as dead +ok - real herdr: exit on a pane with a stale registration is idempotent success +ok - real herdr: a stale registration no longer blocks relaunch, and the endpoint and local copy survive ok - real herdr: an agent that does not stop fails closed instead of being reported as stopped ``` -The registry read through `herdr pane report-agent` is the same source `fm_backend_herdr_agent_state` classifies, so registering and not registering an agent on a plain shell pane exercises exactly the gate every lifecycle verb depends on, with no real agent launched. +The registry read through `herdr pane report-agent` is the same source `fm_backend_herdr_agent_state` classifies, and since 2026-09-10 that registration counts as an agent only while `pane process-info` shows a harness process behind it, so the guard backs the registration with a real process named like a harness (a symlink to `sleep`) and then stops that process, with no real harness launched. That command is the guard that refreshes this record; run it after every Herdr upgrade rather than trusting the version above. For Pi on Herdr 0.9.0, `herdr agent get` reflects whether the agent process remains live; its registration does not persist merely because the pane and parent shell do. @@ -1089,6 +1094,74 @@ ok - real herdr: a drifted agent-free shell returns to its worktree and reuses t `tests/fm-control-relaunch.test.sh` drives a tmux stub and proves that tmux retains its prior refusal without sending `cd` or any other input to the pane. The Herdr refusal when a shell accepts the command but does not move is not exercised in this change. +### Stale agent registration + +Measured 2026-09-10 on macOS aarch64 against Herdr 0.9.0 (protocol 22) and Pi 0.85.1 in an isolated `fm-lab-` session (upstream issue #4115, duplicates #3639, #3487, #2908, #3545). + +Herdr keeps a Pi registration after the Pi process has exited to a shell when a nested interactive shell sits under the pane's top shell, which is the crew shape `treehouse get` leaves behind; a plain `/quit` directly under the top shell, and a `kill -9` of Pi, both released it on this version. +Reproduced in the lab with a nested `zsh` under the pane shell, then `pi` with no prompt, then `/quit`: + +```sh +herdr pane run w1:p1 zsh --session "$LAB"; herdr pane run w1:p1 pi --session "$LAB" +herdr agent get w1:p1 --session "$LAB" | jq -c '.result.agent | {agent, agent_status}' +herdr pane process-info --pane w1:p1 --session "$LAB" | jq -c '.result.process_info | {shell_pid, fg: .foreground_process_group_id, procs: [.foreground_processes[] | {pid, name, argv0}]}' +herdr pane send-text w1:p1 '/quit' --session "$LAB"; herdr pane send-keys w1:p1 Enter --session "$LAB" +herdr agent get w1:p1 --session "$LAB" | jq -c '.result.agent | {agent, agent_status}' +herdr pane process-info --pane w1:p1 --session "$LAB" | jq -c '.result.process_info | {shell_pid, fg: .foreground_process_group_id, procs: [.foreground_processes[] | {pid, name, argv0}]}' +``` + +```text +{"agent":"pi","agent_status":"idle"} +{"shell_pid":87754,"fg":35952,"procs":[{"pid":35952,"name":"node","argv0":"pi"}]} +{"agent":"pi","agent_status":"idle"} +{"shell_pid":87754,"fg":35834,"procs":[{"pid":35834,"name":"zsh","argv0":"zsh"}]} +``` + +Before the fix `fm_backend_agent_state herdr` read that second state as `alive`, so `bin/fm-control.sh relaunch` and `bin/fm-spawn.sh --relaunch` were refused for as long as the registration lived, which is hours. +The registration is still present after the wait, and Herdr's own `pane report-agent` leaves the same shape behind on any pane, which is what the lifecycle-control guard uses. + +Two vendor facts the fix rests on, both read from the outputs above and from `fm_backend_herdr_pane_process_state`'s `pane process-info` parse: + +- Pi's process presents with kernel name `node` and argv0 `pi` (its foreground group also carries Pi's child `node` helpers with argv0 such as `npm view ... version`), so a running Pi is attributed by argv[0] exactly as the tmux probe attributes it; a symlink named `claude` to `sleep` presents as name `sleep`, argv0 `claude`. +- Herdr creates the record with its own placeholder `agent_status` of `unknown` the moment it notices Pi, before Pi's extension reports `idle`; that transient reads `unknown` in the pane classifier as it always did, and only a lifecycle status is subject to the process-level proof. + +Subcommand presence below the 0.9.0 measurement, checked 2026-09-10 on macOS aarch64 against the pinned upstream release clients fetched from `https://github.com/ogulcancelik/herdr/releases/download/v/herdr-macos-aarch64`: + +| Release | sha256 | +|---------|--------| +| 0.7.1 | `16f4653f0491ea1e7d2b46b5b02542f18e1b82e88daaf9e2900572e5bb634df8` | +| 0.7.3 | `b31345392d004ec1f1b2c821e1ad601019fa8385fe1e4c6931321eb58a920773` | +| 0.7.4 | `24992e1625dbdcb18354a59e299e4b263c312400b31396cdc07cd46ed57f24a7` | +| 0.7.5 | `37350546b0012555943b92eaf962665de4e264395baeb44227b8015e8ff5b0d6` | + +The command run against each client was ` pane --help`, which is client-side, session-independent, and opens no socket, and each printed the line: + +```text +process-info Show pane process information +``` + +This proves subcommand presence in the client only, not the server response shape, which is measured only on 0.9.0 above. + +The live guard that refreshes this record runs by default wherever Herdr and Pi are installed, spends no model token, and fails naming both versions: + +```sh +tests/fm-herdr-pi-stale-registration-live-e2e.test.sh +``` + +Observed 2026-09-10: + +```text +# pi 0.85.1 under herdr 0.9.0: registered idle, foreground [{"name":"node","argv0":"node"},{"name":"node","argv0":"node"},{"name":"node","argv0":"rpiv-ask-user-question version"},{"name":"node","argv0":"npm view gentle-engram version"},{"name":"node","argv0":"pi"}] +ok - real herdr 0.9.0 + pi 0.85.1: a running registered pi classifies alive at process level +# herdr 0.9.0 kept the pi registration (idle) after /quit under a nested shell: the stale-registration branch is exercised +ok - real herdr 0.9.0 + pi 0.85.1: the registration left behind by a quit pi reads stale-agent and recovers as dead +``` + +`tests/fm-control-herdr-smoke.test.sh` proves the same shape through the control plane with no harness launched (the two `stale` lines under "Agent lifecycle control" above): a registration over a real agent-named process reads `alive`, stopping that process makes the pane read `stale-agent` and recover as `dead` while `agent get` still reports the record, `exit` then reports `already-stopped`, and `--relaunch` reuses the same endpoint with the local copy intact. +`tests/fm-backend-herdr.test.sh` pins the logic portably with canned `process-info` bodies over real processes, driving the signals apart: the identical shell-only foreground reads `stale-agent` for a childless shell and `live` when an agent-named process is still a descendant of that shell, a `working`, `done`, or `blocked` record over a shell-only pane reads the same as `idle`, an unreadable process view reads `unknown` and refuses husk closing, a transient prompt helper beside the shell settles into `stale-agent` on the next shell-only sample while a foreground that never settles within the bound still reads `live`, and `busy_state` verifies a `working` record before reporting busy. +`tests/fm-crew-state.test.sh` pins the recovery classifier: a stale registration over a shell-only pane reports agent gone rather than alive or unreachable, and a stale `working` record never reports the pane working. +A stale-registration pane is never a husk: create, reclaim, presentation recovery, and session cleanup keep refusing it, and only recovery reuses it. + ### Away-mode transport The away daemon is no longer launched on Pi; the away posture there is the record `bin/fm-afk-contract.sh` owns. diff --git a/tests/fm-backend-herdr.test.sh b/tests/fm-backend-herdr.test.sh index c1b2c3bc44b..8f02af73011 100755 --- a/tests/fm-backend-herdr.test.sh +++ b/tests/fm-backend-herdr.test.sh @@ -419,6 +419,286 @@ test_recovery_grade_read_widens_only_at_its_own_boundary() { pass "herdr recovery-grade read: a stopped server means missing there, and nowhere else" } +# --- stale agent registration over a shell-only pane (issue #4115) ----------- +# +# Herdr keeps a Pi registration (`agent get` -> agent=pi, agent_status=idle) +# after the Pi process has exited to a plain shell whenever a nested interactive +# shell sits under the pane's top shell (the `treehouse get` crew shape; +# reproduced on Herdr 0.9.0 - docs/verification/runtime-backends.md "Stale agent +# registration"). Trusting that registration alone classified the pane `live`, +# so every relaunch and recovery was refused forever. The classifier must now +# prove an agent at process level before reporting one, exactly as the tmux +# adapter does, and a registration with no live agent process is agent-free +# with an explicit reason. +# +# The fixture pairs a canned `pane process-info` body with REAL processes: +# the shell pid it names is a real process this test owns, so the descendant +# walk runs against the real operating-system process table. + +stale_registration_case() { # [process-info-exit] + local dir="$TMP_ROOT/stale-reg-$1" resp log fb n + mkdir -p "$dir/responses"; resp="$dir/responses"; log="$dir/log"; : > "$log" + # The probe below classifies the same pane three times (pane state, the + # recovery-grade read, the husk check), and the canned fake consumes + # responses in call order, so the same three-call script is laid down for + # each pass: + for n in 0 3 6; do + # +1: pane get -> the pane structurally exists + printf '{"result":{"pane":{"pane_id":"w1:p2"}}}\n' > "$resp/$((n + 1)).out" + # +2: agent get -> a registered agent with the given status + printf '{"result":{"agent":{"agent":"pi","agent_status":"%s"}}}\n' "$2" > "$resp/$((n + 2)).out" + # +3: pane process-info -> the pane's actual process view + [ "$3" = - ] || printf '%s\n' "$3" > "$resp/$((n + 3)).out" + [ -z "${4:-}" ] || printf '%s\n' "$4" > "$resp/$((n + 3)).exit" + done + fb=$(make_herdr_fakebin "$dir") + PATH="$fb:$PATH" FM_HERDR_LOG="$log" FM_HERDR_RESPONSES="$resp" \ + bash -c '. "$0/bin/backends/herdr.sh" + printf "%s %s " "$(fm_backend_herdr_pane_agent_state fmtest w1:p2)" "$(fm_backend_herdr_agent_state fmtest:w1:p2)" + fm_backend_herdr_tab_is_husk fmtest w1:p2 && printf husk || printf refused' "$ROOT" +} + +shell_only_process_info() { # + printf '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p2","shell_pid":%s,"foreground_process_group_id":%s,"foreground_processes":[{"pid":%s,"name":"zsh","argv0":"zsh","argv":["-zsh"],"cmdline":"-zsh"}]}}}' "$1" "$1" "$1" +} + +test_stale_registration_over_a_shell_only_pane_is_agent_free() { + local sleep_bin shell_pid out + sleep_bin=$(command -v sleep) || fail "sleep not found" + # A real, childless process stands in for the pane's shell. + "$sleep_bin" 300 & + shell_pid=$! + out=$(stale_registration_case shell-only idle "$(shell_only_process_info "$shell_pid")") + kill "$shell_pid" 2>/dev/null || true + [ "$out" = "stale-agent dead refused" ] \ + || fail "a registered idle agent over a shell-only pane must read stale-agent, recover as dead, and still refuse husk closing; got '$out'" + pass "herdr stale registration: a shell-only pane with a lingering Pi record is agent-free with an explicit reason" +} + +test_stale_registration_ignores_status_and_reads_the_process() { + local sleep_bin shell_pid out status + sleep_bin=$(command -v sleep) || fail "sleep not found" + "$sleep_bin" 300 & + shell_pid=$! + for status in working 'done' blocked; do + out=$(stale_registration_case "shell-only-$status" "$status" "$(shell_only_process_info "$shell_pid")") + [ "$out" = "stale-agent dead refused" ] \ + || { kill "$shell_pid" 2>/dev/null; fail "a lingering '$status' record over a shell-only pane must still read stale-agent/dead, got '$out'"; } + done + kill "$shell_pid" 2>/dev/null || true + pass "herdr stale registration: no registered status can outrank a shell-only process view" +} + +test_registered_agent_with_a_live_foreground_process_stays_alive() { + local out + # The real Pi shape on Herdr 0.9.0: the kernel name is the interpreter and + # only argv0 says pi. + out=$(stale_registration_case live-pi idle \ + '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p2","shell_pid":4242,"foreground_process_group_id":4243,"foreground_processes":[{"pid":4243,"name":"node","argv0":"pi","argv":["pi"],"cmdline":"pi"}]}}}') + [ "$out" = "live alive refused" ] \ + || fail "a registered agent whose foreground process is Pi must stay live/alive, got '$out'" + pass "herdr stale registration: a registered agent with a live Pi foreground process still reads alive" +} + +test_registered_agent_with_a_non_shell_foreground_process_stays_alive() { + local out + # A registered agent running a foreground tool in its own process group is + # not a shell-only pane, so the registration keeps its authority. + out=$(FM_BACKEND_HERDR_IDLE_SHELL_PROOF_POLLS=1 stale_registration_case live-tool working \ + '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p2","shell_pid":4242,"foreground_process_group_id":4250,"foreground_processes":[{"pid":4250,"name":"git","argv0":"git","argv":["git","status"],"cmdline":"git status"}]}}}') + [ "$out" = "live alive refused" ] \ + || fail "a registered agent with a non-shell foreground process must stay live/alive, got '$out'" + pass "herdr stale registration: only a shell-only pane demotes a registration" +} + +# settle_registration_case: one pane classification over a scripted sequence +# of `pane process-info` samples, so the settle window's resampling is +# observable in the fake CLI's call log. +settle_registration_case() { # ... + local dir="$TMP_ROOT/settle-reg-$1" polls=$2 resp log fb n + shift 2 + mkdir -p "$dir/responses"; resp="$dir/responses"; log="$dir/log"; : > "$log" + printf '{"result":{"pane":{"pane_id":"w1:p2"}}}\n' > "$resp/1.out" + printf '{"result":{"agent":{"agent":"pi","agent_status":"idle"}}}\n' > "$resp/2.out" + n=3 + for body in "$@"; do + printf '%s\n' "$body" > "$resp/$n.out" + n=$((n + 1)) + done + fb=$(make_herdr_fakebin "$dir") + PATH="$fb:$PATH" FM_HERDR_LOG="$log" FM_HERDR_RESPONSES="$resp" \ + FM_BACKEND_HERDR_IDLE_SHELL_PROOF_POLLS="$polls" \ + bash -c '. "$0/bin/backends/herdr.sh" + printf "%s %s" "$(fm_backend_herdr_pane_agent_state fmtest w1:p2)" "$(grep -c "process-info" "$1")"' "$ROOT" "$log" +} + +prompt_helper_process_info() { # + printf '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p2","shell_pid":%s,"foreground_process_group_id":%s,"foreground_processes":[{"pid":99998,"name":"starship","argv":["/usr/local/bin/starship","prompt","--continuation"]},{"pid":%s,"name":"zsh","argv0":"zsh","argv":["-zsh"],"cmdline":"-zsh"}]}}}' "$1" "$1" "$1" +} + +test_transient_prompt_helper_settles_into_stale_agent() { + local sleep_bin shell_pid out + sleep_bin=$(command -v sleep) || fail "sleep not found" + "$sleep_bin" 300 & + shell_pid=$! + # Sample 1: the shell is redrawing its prompt with starship beside it (the + # real 0.7.5 shape); sample 2: the helper is gone and the shell is alone. + out=$(settle_registration_case helper-settles 3 \ + "$(prompt_helper_process_info "$shell_pid")" "$(shell_only_process_info "$shell_pid")") + kill "$shell_pid" 2>/dev/null || true + [ "$out" = "stale-agent 2" ] \ + || fail "a transient prompt helper followed by a shell-only sample must settle into stale-agent after exactly two samples, got '$out'" + pass "herdr stale registration: a transient prompt helper settles into stale-agent instead of reading live" +} + +test_exhausted_settle_window_keeps_a_non_shell_foreground_live() { + local sleep_bin shell_pid out + sleep_bin=$(command -v sleep) || fail "sleep not found" + "$sleep_bin" 300 & + shell_pid=$! + out=$(settle_registration_case helper-persists 2 \ + "$(prompt_helper_process_info "$shell_pid")" "$(prompt_helper_process_info "$shell_pid")" \ + "$(shell_only_process_info "$shell_pid")") + kill "$shell_pid" 2>/dev/null || true + [ "$out" = "live 2" ] \ + || fail "a foreground that never settles within the bound must stay live after exactly the bounded sample count, got '$out'" + pass "herdr stale registration: an exhausted settle window still reads a non-shell foreground as live" +} + +test_registered_agent_with_an_agent_descendant_outside_the_foreground_stays_alive() { + local lab sleep_bin shell_pid out shell_verdict + sleep_bin=$(command -v sleep) || fail "sleep not found" + lab="$TMP_ROOT/stale-reg-descendant-bin"; mkdir -p "$lab" + # A symlink to a real long-running binary so the kernel records `pi` as the + # executable identity (a copied platform binary fails code signing on macOS). + ln -sf "$sleep_bin" "$lab/pi" + # A real shell whose child is that agent-named process, while the canned + # foreground view shows only the shell (a suspended or backgrounded agent). + sh -c "'$lab/pi' 300; :" & + shell_pid=$! + sleep 0.3 + out=$(stale_registration_case descendant idle "$(shell_only_process_info "$shell_pid")") + pkill -P "$shell_pid" 2>/dev/null || true + kill "$shell_pid" 2>/dev/null || true + [ "$out" = "live alive refused" ] \ + || fail "a registered agent with a live agent-named descendant must stay live/alive, got '$out'" + # The divergence itself: the identical canned foreground view reads + # stale-agent for a childless shell, so the descendant walk is what carried + # this verdict. + "$sleep_bin" 300 & + shell_pid=$! + shell_verdict=$(stale_registration_case descendant-childless idle "$(shell_only_process_info "$shell_pid")") + kill "$shell_pid" 2>/dev/null || true + [ "$shell_verdict" = "stale-agent dead refused" ] \ + || fail "the childless control must read stale-agent so the descendant case is not vacuous, got '$shell_verdict'" + pass "herdr stale registration: an agent process outside the foreground group still counts as alive" +} + +test_agent_descendant_under_a_spaced_install_path_stays_alive() { + local lab sleep_bin shell_pid out + sleep_bin=$(command -v sleep) || fail "sleep not found" + # The executable path the process table reports contains a space (the macOS + # `/Library/Application Support/...` shape), so a field-split read of the + # process table sees only a fragment of the name. + lab="$TMP_ROOT/stale-reg-spaced-bin/Application Support/Some Dir"; mkdir -p "$lab" + ln -sf "$sleep_bin" "$lab/pi" + sh -c "'$lab/pi' 300; :" & + shell_pid=$! + sleep 0.3 + out=$(stale_registration_case spaced-descendant idle "$(shell_only_process_info "$shell_pid")") + pkill -P "$shell_pid" 2>/dev/null || true + kill "$shell_pid" 2>/dev/null || true + [ "$out" = "live alive refused" ] \ + || fail "an agent-named descendant under a spaced install path must stay live/alive, got '$out'" + pass "herdr stale registration: the descendant walk reads a spaced executable path whole" +} + +test_registered_agent_with_an_unreadable_process_view_is_unknown() { + local out + out=$(stale_registration_case unreadable-exit idle 'Error: socket unavailable' 1) + [ "$out" = "unknown unreadable refused" ] \ + || fail "a failed process-info read must not demote OR trust the registration: expected unknown/unreadable, got '$out'" + out=$(stale_registration_case unreadable-empty idle -) + [ "$out" = "unknown unreadable refused" ] \ + || fail "an empty process-info read must read unknown/unreadable, got '$out'" + out=$(stale_registration_case unreadable-mismatch idle \ + '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w9:p9","shell_pid":4242,"foreground_process_group_id":4242,"foreground_processes":[{"pid":4242,"name":"zsh","argv0":"zsh"}]}}}') + [ "$out" = "unknown unreadable refused" ] \ + || fail "a process view for a different pane must read unknown/unreadable, got '$out'" + out=$(stale_registration_case unreadable-no-foreground idle \ + '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p2","shell_pid":4242,"foreground_process_group_id":4242,"foreground_processes":[]}}}') + [ "$out" = "unknown unreadable refused" ] \ + || fail "an empty foreground list must read unknown/unreadable, got '$out'" + pass "herdr stale registration: an unreadable process view refuses instead of guessing either way" +} + +test_registered_agent_with_an_empty_foreground_over_a_real_shell_settles_via_descendant_walk() { + local sleep_bin shell_pid out + sleep_bin=$(command -v sleep) || fail "sleep not found" + # A real, childless shell process stands in for the pane's shell, and the + # foreground list is empty - the exec-to-shell handoff shape the flake fix + # targets. Unlike unreadable-no-foreground above (a synthetic pid absent + # from `ps`), this shell_pid is real, so the descendant walk can run to + # completion and prove the empty array settles to stale-agent, not + # unreadable. + "$sleep_bin" 300 & + shell_pid=$! + out=$(stale_registration_case empty-foreground idle \ + "$(printf '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p2","shell_pid":%s,"foreground_process_group_id":%s,"foreground_processes":[]}}}' "$shell_pid" "$shell_pid")") + kill "$shell_pid" 2>/dev/null || true + [ "$out" = "stale-agent dead refused" ] \ + || fail "an empty foreground list over a real childless shell must settle to stale-agent via the descendant walk, not unreadable, got '$out'" + pass "herdr stale registration: an empty foreground list over a real shell is not unreadable, it settles via the descendant walk" +} + +test_projection_reclaim_rollback_refuses_a_stale_registration() { + local out + out=$(bash -c '. "$0/bin/backends/herdr.sh" + fm_backend_herdr_pane_agent_state() { printf stale-agent; } + fm_backend_herdr_projection_close_pane_focus_preserving() { printf "CLOSED %s\n" "$2" >&2; exit 99; } + fm_backend_herdr_projection_reclaim_rollback fmtest w1:p9; printf "rc=%s" "$?"' "$ROOT" 2>&1) + [ "$out" = "rc=1" ] \ + || fail "reclaim rollback must refuse (never close) a pane with a stale registration, got '$out'" + pass "herdr stale registration: presentation reclaim never closes a stale-registration pane" +} + +test_busy_state_never_reports_a_shell_only_pane_busy() { + local sleep_bin shell_pid dir resp log fb out + sleep_bin=$(command -v sleep) || fail "sleep not found" + "$sleep_bin" 300 & + shell_pid=$! + dir="$TMP_ROOT/busy-stale"; mkdir -p "$dir/responses"; resp="$dir/responses"; log="$dir/log"; : > "$log" + # 1: agent get -> a lingering working record; 2: process-info -> shell only + printf '{"result":{"agent":{"agent":"pi","agent_status":"working"}}}\n' > "$resp/1.out" + shell_only_process_info "$shell_pid" > "$resp/2.out" + fb=$(make_herdr_fakebin "$dir") + out=$(PATH="$fb:$PATH" FM_HERDR_LOG="$log" FM_HERDR_RESPONSES="$resp" \ + bash -c '. "$0/bin/backends/herdr.sh"; fm_backend_herdr_busy_state fmtest:w1:p2' "$ROOT") + kill "$shell_pid" 2>/dev/null || true + [ "$out" = unknown ] \ + || fail "a working record over a shell-only pane must not read busy, got '$out'" + assert_contains "$(cat "$log")" $'pane\x1fprocess-info' "busy_state did not verify the working record at process level" + + # The control: the same working record with a live Pi foreground reads busy. + dir="$TMP_ROOT/busy-live"; mkdir -p "$dir/responses"; resp="$dir/responses"; log="$dir/log"; : > "$log" + printf '{"result":{"agent":{"agent":"pi","agent_status":"working"}}}\n' > "$resp/1.out" + printf '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p2","shell_pid":4242,"foreground_process_group_id":4243,"foreground_processes":[{"pid":4243,"name":"node","argv0":"pi","argv":["pi"],"cmdline":"pi"}]}}}\n' > "$resp/2.out" + fb=$(make_herdr_fakebin "$dir") + out=$(PATH="$fb:$PATH" FM_HERDR_LOG="$log" FM_HERDR_RESPONSES="$resp" \ + bash -c '. "$0/bin/backends/herdr.sh"; fm_backend_herdr_busy_state fmtest:w1:p2' "$ROOT") + [ "$out" = busy ] || fail "a working record with a live Pi foreground must read busy, got '$out'" + + # An idle record needs no process read: idle is never trusted as busy anyway. + dir="$TMP_ROOT/busy-idle"; mkdir -p "$dir/responses"; resp="$dir/responses"; log="$dir/log"; : > "$log" + printf '{"result":{"agent":{"agent":"pi","agent_status":"idle"}}}\n' > "$resp/1.out" + fb=$(make_herdr_fakebin "$dir") + out=$(PATH="$fb:$PATH" FM_HERDR_LOG="$log" FM_HERDR_RESPONSES="$resp" \ + bash -c '. "$0/bin/backends/herdr.sh"; fm_backend_herdr_busy_state fmtest:w1:p2' "$ROOT") + [ "$out" = idle ] || fail "an idle record should read idle without a process read, got '$out'" + assert_not_contains "$(cat "$log")" $'pane\x1fprocess-info' "busy_state ran a process read for an idle record" + pass "herdr stale registration: busy_state proves a working record at process level before reporting busy" +} + test_agent_state_bypasses_a_stale_client_shadowing_a_compatible_one() { local dir out err dir="$TMP_ROOT/client-pair-bypass"; make_herdr_client_pair "$dir" @@ -876,6 +1156,8 @@ test_create_task_refuses_duplicate_label_when_agent_live() { printf '{"result":{"pane":{"pane_id":"w1:p2"}}}\n' > "$resp/3.out" # 4: agent get -> a genuinely registered, live agent (idle, not just working) printf '{"result":{"agent":{"agent_status":"idle"}}}\n' > "$resp/4.out" + # 5: pane process-info -> a live Pi process backs that registration (#4115) + printf '%s\n' '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p2","shell_pid":4242,"foreground_process_group_id":4243,"foreground_processes":[{"pid":4243,"name":"node","argv0":"pi"}]}}}' > "$resp/5.out" fb=$(make_herdr_fakebin "$dir") out=$( PATH="$fb:$PATH" FM_HERDR_LOG="$log" FM_HERDR_RESPONSES="$resp" \ bash -c '. "$0/bin/backends/herdr.sh"; fm_backend_herdr_create_task fmtest:w1 fm-dup1 /tmp/proj' "$ROOT" 2>&1 ) @@ -897,6 +1179,8 @@ test_create_task_refuses_when_any_duplicate_label_is_live() { printf '{"result":{"panes":[{"pane_id":"w1:p2","tab_id":"w1:t2"},{"pane_id":"w1:p3","tab_id":"w1:t3"}]}}\n' > "$resp/5.out" printf '{"result":{"pane":{"pane_id":"w1:p3"}}}\n' > "$resp/6.out" printf '{"result":{"agent":{"agent_status":"idle"}}}\n' > "$resp/7.out" + # 8: pane process-info -> a live Pi process backs that registration (#4115) + printf '%s\n' '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p3","shell_pid":4242,"foreground_process_group_id":4243,"foreground_processes":[{"pid":4243,"name":"node","argv0":"pi"}]}}}' > "$resp/8.out" fb=$(make_herdr_fakebin "$dir") out=$( PATH="$fb:$PATH" FM_HERDR_LOG="$log" FM_HERDR_RESPONSES="$resp" \ bash -c '. "$0/bin/backends/herdr.sh"; fm_backend_herdr_create_task fmtest:w1 fm-mixed1 /tmp/proj' "$ROOT" 2>&1 ) @@ -3233,6 +3517,8 @@ test_projection_recovery_is_read_only_and_refuses_live_duplicate_risk() { printf '{"result":{"panes":[{"pane_id":"w1:p1","tab_id":"w1:t1"}]}}\n' > "$resp/2.out" printf '{"result":{"pane":{"pane_id":"w1:p1"}}}\n' > "$resp/3.out" printf '{"result":{"agent":{"agent_status":"idle"}}}\n' > "$resp/4.out" + # 5: process-info -> a live harness backs the registration (issue #4115) + printf '{"result":{"type":"pane_process_info","process_info":{"pane_id":"w1:p1","shell_pid":4242,"foreground_process_group_id":4243,"foreground_processes":[{"pid":4243,"name":"node","argv0":"pi"}]}}}\n' > "$resp/5.out" out=$(PATH="$fb:$PATH" FM_HERDR_LOG="$log" FM_HERDR_RESPONSES="$resp" \ bash -c '. "$0/bin/backends/herdr.sh"; fm_backend_herdr_projection_recovery_allows_flat fmtest "$1" task-p3' "$ROOT" "$journal" 2>&1) status=$? @@ -4895,6 +5181,18 @@ test_workspace_label_different_secondmates_get_different_labels test_cli_helper_sets_env_and_appends_trailing_session_flag test_agent_state_bypasses_a_stale_client_shadowing_a_compatible_one test_recovery_grade_read_widens_only_at_its_own_boundary +test_stale_registration_over_a_shell_only_pane_is_agent_free +test_stale_registration_ignores_status_and_reads_the_process +test_registered_agent_with_a_live_foreground_process_stays_alive +test_registered_agent_with_a_non_shell_foreground_process_stays_alive +test_transient_prompt_helper_settles_into_stale_agent +test_exhausted_settle_window_keeps_a_non_shell_foreground_live +test_registered_agent_with_an_agent_descendant_outside_the_foreground_stays_alive +test_agent_descendant_under_a_spaced_install_path_stays_alive +test_registered_agent_with_an_unreadable_process_view_is_unknown +test_registered_agent_with_an_empty_foreground_over_a_real_shell_settles_via_descendant_walk +test_projection_reclaim_rollback_refuses_a_stale_registration +test_busy_state_never_reports_a_shell_only_pane_busy test_cli_caches_the_selected_client_within_a_process test_cli_scopes_the_selected_client_to_its_session test_cli_unrelated_failure_never_triggers_reselection diff --git a/tests/fm-control-herdr-smoke.test.sh b/tests/fm-control-herdr-smoke.test.sh index 427b6ae1077..8c86947bbc2 100755 --- a/tests/fm-control-herdr-smoke.test.sh +++ b/tests/fm-control-herdr-smoke.test.sh @@ -9,9 +9,12 @@ # an agent is running, and therefore whether a lifecycle verb may act at all, # comes from herdr's own agent registry. # -# No real agent is launched. herdr's `pane report-agent` is the same registry -# the adapter reads, so registering and not registering an agent on a plain -# shell pane exercises exactly the classification the control plane gates on. +# No real harness is launched. herdr's `pane report-agent` is the same registry +# the adapter reads, and a symlink named like a harness is the same process +# identity the adapter proves through `pane process-info`, so registering an +# agent over a real agent-named process, over a plain shell, and not at all +# exercises exactly the classification the control plane gates on - including +# the registration Herdr keeps after the agent process is gone (issue #4115). # # Always runs on a private, named, throwaway lab session, never the default # one (tests/herdr-test-safety.sh; the 2026-07-02 incident). Skips cleanly @@ -199,14 +202,44 @@ case "$OUT" in esac pass "real herdr: interrupt refuses when herdr's own agent registry reports no agent" -# --- a registered agent: classification flips, and the verbs follow --------- +# --- a registered agent WITH a live process: classification flips ------------ +# +# A registration alone no longer proves an agent (issue #4115): the adapter +# verifies the pane's processes through the real `pane process-info` view. So +# the registered agent is backed by a real agent-named foreground process - a +# symlink to a long-running system binary named `claude`, the same construction +# tests/fm-tmux-agent-liveness.test.sh uses (a copied platform binary fails code +# signing on macOS arm64; the symlink name is what the kernel records as argv[0]). +AGENT_BIN="$SCRATCH/agentbin" +mkdir -p "$AGENT_BIN" +SLEEP_BIN=$(command -v sleep) || fail "sleep not found" +ln -s "$SLEEP_BIN" "$AGENT_BIN/claude" +printf -v AGENT_Q '%q' "$AGENT_BIN/claude" + +wait_process_state() { # + local expected=$1 tries=$2 i=0 + while [ "$i" -lt "$tries" ]; do + [ "$(fm_backend_herdr_pane_process_state "$SESSION" "$PANE_ID")" != "$expected" ] || return 0 + sleep 0.1 + i=$((i + 1)) + done + return 1 +} + +start_agent_process() { + fm_backend_herdr_send_text_line "$SESSION:$PANE_ID" "$AGENT_Q 900" \ + || fail "could not start the agent-named foreground process in the task pane" + wait_process_state agent 50 \ + || version_fail "a real agent-named foreground process reads '$(fm_backend_herdr_pane_process_state "$SESSION" "$PANE_ID")' rather than 'agent' through pane process-info" +} +start_agent_process herdr pane report-agent "$PANE_ID" --source fm-control-smoke --agent fm-control-smoke-agent \ --state idle --session "$SESSION" >/dev/null 2>&1 \ || fail "could not register a live agent on the task pane" STATE=$(fm_backend_agent_state herdr "$SESSION:$PANE_ID") -[ "$STATE" = alive ] || fail "herdr should classify a registered agent as alive, got '$STATE'" +[ "$STATE" = alive ] || fail "herdr should classify a registered agent with a live process as alive, got '$STATE'" OUT=$(run_control hsmoke interrupt) || fail "interrupt against a registered agent should succeed: $OUT" case "$OUT" in @@ -220,9 +253,71 @@ herdr pane get "$PANE_ID" --session "$SESSION" >/dev/null 2>&1 \ [ -d "$WT" ] || fail "the control plane must never remove the task's local copy" pass "real herdr: no control verb removed the endpoint or the task's local copy" -# Last, because it deliberately types a harness command into a pane that hosts -# a plain shell: the registered agent cannot actually be stopped that way, and -# the control plane must say so rather than report a stop it did not achieve. +# --- the stale registration (issue #4115): the agent process is gone, the --- +# --- record is not, and recovery must proceed anyway ------------------------ +# +# Stopping the agent-named process leaves the pane a plain shell while Herdr +# keeps the registration, which is exactly the shape a Pi crew leaves behind +# when it exits under a nested shell. Before the fix this read `alive` forever: +# exit waited out its timeout and refused, and relaunch was refused for good. +# This runs BEFORE the fail-closed exit case below, whose typed exit command +# stays buffered in the pane's tty while the stand-in ignores it and would be +# replayed into the shell the moment the stand-in died. +AGENT_PID=$(herdr pane process-info --pane "$PANE_ID" --session "$SESSION" 2>/dev/null \ + | jq -r '.result.process_info.foreground_processes[0].pid // empty') +[ -n "$AGENT_PID" ] || fail "could not read the agent-named process pid from pane process-info" +kill "$AGENT_PID" 2>/dev/null || fail "could not stop the agent-named process" +wait_process_state shell 50 \ + || version_fail "after the agent process exited the pane reads '$(fm_backend_herdr_pane_process_state "$SESSION" "$PANE_ID")' rather than 'shell' through pane process-info. Raw process-info: $(herdr pane process-info --pane "$PANE_ID" --session "$SESSION" 2>&1 | tr -d '\n')" + +# The divergence that makes this case non-vacuous: Herdr's own registry still +# reports the agent, and only the process-level view disagrees. +REGISTERED=$(herdr agent get "$PANE_ID" --session "$SESSION" 2>/dev/null | jq -r '.result.agent.agent_status // empty') +[ -n "$REGISTERED" ] \ + || version_fail "Herdr released the registration when the agent process exited, so this run cannot prove the stale-registration path; the classifier still reads dead through agent_not_found" + +PANE_STATE=$(fm_backend_herdr_pane_agent_state "$SESSION" "$PANE_ID") +[ "$PANE_STATE" = stale-agent ] \ + || version_fail "a registration over a shell-only pane reads '$PANE_STATE' rather than 'stale-agent'" +STATE=$(fm_backend_agent_state herdr "$SESSION:$PANE_ID") +[ "$STATE" = dead ] \ + || version_fail "a registration over a shell-only pane recovers as '$STATE' rather than 'dead'; every relaunch would be refused" +pass "real herdr $HERDR_VERSION: a registration Herdr keeps after its agent exits reads stale-agent and recovers as dead" + +OUT=$(run_control hsmoke exit) || fail "exit against a stale-registration pane should be idempotent success: $OUT" +case "$OUT" in + "already-stopped hsmoke"*) : ;; + *) fail "a stale-registration pane should report already-stopped, got: $OUT" ;; +esac +pass "real herdr: exit on a pane with a stale registration is idempotent success" + +rm -f "$SCRATCH/codex-launched" +OUT=$(env FM_HOME="$HOME_DIR" HERDR_SESSION="$SESSION" FM_SPAWN_NO_GUARD=1 \ + "$ROOT/bin/fm-spawn.sh" hsmoke --relaunch --harness codex) \ + || fail "a stale-registration Herdr pane should be relaunched: $OUT" +for _ in $(seq 1 20); do + [ ! -e "$SCRATCH/codex-launched" ] || break + sleep 0.1 +done +[ -e "$SCRATCH/codex-launched" ] || fail "the replacement harness was not launched after the stale registration" +[ "$(sed -n 's/^window=//p' "$HOME_DIR/state/hsmoke.meta" | tail -1)" = "$SESSION:$PANE_ID" ] \ + || fail "the relaunch replaced its endpoint instead of reusing it" +herdr pane get "$PANE_ID" --session "$SESSION" >/dev/null 2>&1 \ + || fail "the relaunch removed the endpoint it was required to reuse" +[ -d "$WT" ] || fail "the relaunch must never remove the task's local copy" +awk -F= '$1 == "harness" {$0="harness=claude"} {print}' "$HOME_DIR/state/hsmoke.meta" \ + > "$HOME_DIR/state/hsmoke.meta.tmp" +mv "$HOME_DIR/state/hsmoke.meta.tmp" "$HOME_DIR/state/hsmoke.meta" +pass "real herdr: a stale registration no longer blocks relaunch, and the endpoint and local copy survive" + +# Last, because it deliberately types a harness command into a foreground +# process that ignores it: the registered agent cannot actually be stopped +# that way, and the control plane must say so rather than report a stop it +# did not achieve. +start_agent_process +herdr pane report-agent "$PANE_ID" --source fm-control-smoke --agent fm-control-smoke-agent \ + --state idle --session "$SESSION" >/dev/null 2>&1 \ + || fail "could not re-register the live agent on the task pane" if OUT=$(run_control hsmoke exit 2>&1); then fail "exit should fail closed when the agent does not stop: $OUT" fi diff --git a/tests/fm-crew-state.test.sh b/tests/fm-crew-state.test.sh index f08f372d14a..da93917667d 100755 --- a/tests/fm-crew-state.test.sh +++ b/tests/fm-crew-state.test.sh @@ -153,6 +153,18 @@ case "${1:-}" in fi printf '{"result":{"pane":{"pane_id":"%s"}}}\n' "${3:-}" exit 0 ;; + process-info) + # The process-level view a registration is verified against (#4115): + # `agent` puts a live claude in the foreground, `shell` a bare zsh whose + # pid is the test script itself (a real, long-lived process with no + # harness descendant, so the adapter's real process-table walk finds + # it), and anything else answers nothing (unreadable). + pane=""; args=("$@"); for ((i=0; i<${#args[@]}; i++)); do [ "${args[$i]}" = --pane ] && pane=${args[$((i+1))]:-}; done + case "${FM_FAKE_HERDR_PROCESS:-agent}" in + agent) printf '{"result":{"type":"pane_process_info","process_info":{"pane_id":"%s","shell_pid":%s,"foreground_process_group_id":424242,"foreground_processes":[{"pid":424242,"name":"claude","argv0":"claude"}]}}}\n' "$pane" "${FM_FAKE_HERDR_SHELL_PID:-$PPID}" ;; + shell) printf '{"result":{"type":"pane_process_info","process_info":{"pane_id":"%s","shell_pid":%s,"foreground_process_group_id":%s,"foreground_processes":[{"pid":%s,"name":"zsh","argv0":"zsh","argv":["-zsh"]}]}}}\n' "$pane" "${FM_FAKE_HERDR_SHELL_PID:-$PPID}" "${FM_FAKE_HERDR_SHELL_PID:-$PPID}" "${FM_FAKE_HERDR_SHELL_PID:-$PPID}" ;; + esac + exit 0 ;; esac ;; agent) case "${2:-}" in @@ -218,10 +230,12 @@ reset_fakes() { FM_FAKE_HERDR_READ_FAIL=0 FM_FAKE_HERDR_HUSK=0 FM_FAKE_HERDR_AGENT_STATUS="" + FM_FAKE_HERDR_PROCESS=agent + FM_FAKE_HERDR_SHELL_PID=$$ FM_FAKE_CI_LOGS="" FM_FAKE_DAEMON_DOWN=0 export FM_FAKE_AXI_STATUS FM_FAKE_AXI_STATUS_RUN FM_FAKE_RUNS_LIST FM_FAKE_BUSY FM_FAKE_BUSY_TEXT FM_FAKE_TMUX_MISSING FM_FAKE_TMUX_UNREADABLE - export FM_FAKE_HERDR_BUSY FM_FAKE_HERDR_MISSING FM_FAKE_HERDR_READ_FAIL FM_FAKE_HERDR_HUSK FM_FAKE_HERDR_AGENT_STATUS FM_FAKE_CI_LOGS + export FM_FAKE_HERDR_BUSY FM_FAKE_HERDR_MISSING FM_FAKE_HERDR_READ_FAIL FM_FAKE_HERDR_HUSK FM_FAKE_HERDR_AGENT_STATUS FM_FAKE_HERDR_PROCESS FM_FAKE_HERDR_SHELL_PID FM_FAKE_CI_LOGS export FM_FAKE_DAEMON_DOWN } @@ -1507,6 +1521,53 @@ test_no_run_herdr_alive_with_failed_read_stays_live() { pass "an alive endpoint whose scrollback read failed stays working" } +# Issue #4115: a registration Herdr kept after its Pi exited to a plain shell is +# not an agent. The recovery-grade read proves the process level, so the +# shell-only pane reads as positive agent-gone evidence, never as a live agent +# or as unreachable. +test_no_run_herdr_stale_registration_over_shell_reads_agent_gone() { + command -v jq >/dev/null 2>&1 || { pass "herdr stale-registration test skipped without jq"; return; } + reset_fakes + local d; d=$(new_case herdr-stale-reg) + make_repo_on_branch "$d/wt" fm/feat-herdr-stale + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-herdr-stale.meta" "window=default:w1:p2" "worktree=$d/wt" "kind=ship" \ + "backend=herdr" "harness=pi" + FM_FAKE_TMUX_MISSING=1 + FM_FAKE_HERDR_READ_FAIL=1 + FM_FAKE_HERDR_AGENT_STATUS=idle + FM_FAKE_HERDR_PROCESS=shell + local out; out=$(run_crew_state "$d" feat-herdr-stale) + assert_contains "$out" "state: unknown" "a stale registration over a shell-only pane is not a live state" + assert_contains "$out" "backend target gone" "a stale registration over a shell-only pane must read as positive agent-gone evidence" + assert_contains "$out" "agent gone, pane shell remains" "the agent-gone reason must name the remaining shell" + assert_not_contains "$out" "backend unreachable" "a readable shell-only pane is not unreachable" + pass "herdr stale registration over a shell-only pane reads agent gone, not alive" +} + +# The busy half of the same defect: a `working` record Herdr kept after the +# agent was killed mid-turn must never make a shell-only pane read as working. +test_no_run_herdr_stale_working_record_is_never_busy() { + command -v jq >/dev/null 2>&1 || { pass "herdr stale-working test skipped without jq"; return; } + reset_fakes + local d; d=$(new_case herdr-stale-working) + make_repo_on_branch "$d/wt" fm/feat-herdr-stale-working + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-herdr-stale-working.meta" "window=default:w1:p2" "worktree=$d/wt" "kind=ship" \ + "backend=herdr" "harness=pi" + FM_FAKE_TMUX_MISSING=1 + FM_FAKE_HERDR_AGENT_STATUS=working + FM_FAKE_HERDR_PROCESS=shell + local out; out=$(run_crew_state "$d" feat-herdr-stale-working) + assert_not_contains "$out" "state: working" "a stale working record over a shell-only pane must never read busy" + assert_not_contains "$out" "herdr-native" "the native busy verdict must not be trusted for a shell-only pane" + # The control: the same record with a live harness in the foreground is busy. + FM_FAKE_HERDR_PROCESS=agent + out=$(run_crew_state "$d" feat-herdr-stale-working) + assert_contains "$out" "state: working" "the same working record with a live harness process must still read working" + pass "herdr stale working record never reports a shell-only pane busy" +} + # Decision follow-up (2026-09-05 review): a husk pane (pane present, # agent_not_found) is authoritative death evidence - it keeps the gone-class # text so the stale sweep may still reclaim it, never unknown/unreachable. @@ -2506,5 +2567,7 @@ test_active_fix_round_unfetched_pipeline_head_reports_current test_unanchored_unfetched_active_row_does_not_match test_unresolved_terminal_row_is_history_not_current test_runs_list_continuation_found_when_axi_answers_other_branch +test_no_run_herdr_stale_registration_over_shell_reads_agent_gone +test_no_run_herdr_stale_working_record_is_never_busy echo "all fm-crew-state tests passed" diff --git a/tests/fm-cursor-harness.test.sh b/tests/fm-cursor-harness.test.sh index 23ecc74948b..cbdb3047a91 100755 --- a/tests/fm-cursor-harness.test.sh +++ b/tests/fm-cursor-harness.test.sh @@ -156,19 +156,19 @@ test_tmux_classifies_cursor_pane_without_inferring_dead() { tree="$TMP_ROOT/tree5"; bin=$(make_cursor_tree "$tree") # shellcheck source=bin/backends/tmux.sh ( FM_BACKEND_LIB_DIR="$ROOT/bin"; . "$ROOT/bin/backends/tmux.sh" - [ "$(fm_backend_tmux_classify_process_name node "$bin/cursor-agent")" = agent ] \ + [ "$(fm_agent_process_classify_name node "$bin/cursor-agent")" = agent ] \ || fail "a cursor pane reported as node must classify agent" - [ "$(fm_backend_tmux_classify_process_name '' "$bin/cursor-agent")" = agent ] \ + [ "$(fm_agent_process_classify_name '' "$bin/cursor-agent")" = agent ] \ || fail "the argv[0]-only call must classify a cursor pane agent" # The safety half: an unrelated node is `other`, and the callers turn # `other` into `ambiguous`, never `dead`. - [ "$(fm_backend_tmux_classify_process_name node /usr/bin/node)" = other ] \ + [ "$(fm_agent_process_classify_name node /usr/bin/node)" = other ] \ || fail "an unrelated node must stay 'other', never agent" - [ "$(fm_backend_tmux_classify_process_name agent /usr/local/bin/agent)" = other ] \ + [ "$(fm_agent_process_classify_name agent /usr/local/bin/agent)" = other ] \ || fail "an unrelated agent must stay 'other', never agent" # Neighbours must not regress. - [ "$(fm_backend_tmux_classify_process_name claude '')" = agent ] || fail "claude regressed" - [ "$(fm_backend_tmux_classify_process_name zsh '')" = shell ] || fail "zsh regressed" + [ "$(fm_agent_process_classify_name claude '')" = agent ] || fail "claude regressed" + [ "$(fm_agent_process_classify_name zsh '')" = shell ] || fail "zsh regressed" ) || exit 1 pass "tmux liveness: a cursor pane is agent; an unrelated node/agent is other, never dead" } diff --git a/tests/fm-harness-liveness-drift-live-e2e.test.sh b/tests/fm-harness-liveness-drift-live-e2e.test.sh index e71d59b7191..a885d7f975f 100755 --- a/tests/fm-harness-liveness-drift-live-e2e.test.sh +++ b/tests/fm-harness-liveness-drift-live-e2e.test.sh @@ -134,7 +134,7 @@ for harness in claude codex opencode pi pi-signed grok kimi cursor muse; do comms=$(fm_backend_tmux_foreground_comms "$target" | tr '\n' ' ') [ "$state" = alive ] || fail \ - "LIVENESS DRIFT: $harness $version is running but classifies '$state', not 'alive'. Supervision and lifecycle control treat this endpoint as unattributable. Observed process title '$title'; observed foreground process names [$comms]. Teach bin/backends/tmux.sh's fm_backend_tmux_classify_process_name the identity this release actually reports." + "LIVENESS DRIFT: $harness $version is running but classifies '$state', not 'alive'. Supervision and lifecycle control treat this endpoint as unattributable. Observed process title '$title'; observed foreground process names [$comms]. Teach bin/fm-agent-process-lib.sh's fm_agent_process_classify_name the identity this release actually reports." note "$harness $version: title='$title' foreground=[$comms]" diff --git a/tests/fm-herdr-pi-stale-registration-live-e2e.test.sh b/tests/fm-herdr-pi-stale-registration-live-e2e.test.sh new file mode 100755 index 00000000000..3df52274ba2 --- /dev/null +++ b/tests/fm-herdr-pi-stale-registration-live-e2e.test.sh @@ -0,0 +1,159 @@ +#!/usr/bin/env bash +# Default-on live guard for the Herdr stale-registration classifier (issue +# #4115) against the REAL Pi harness under the REAL Herdr binary. +# +# The defect: Herdr keeps a Pi registration (`agent get` -> agent=pi, +# agent_status=idle) after the Pi process has exited to a plain shell whenever +# a nested interactive shell sits under the pane's top shell - the crew shape, +# where `treehouse get` leaves a worktree shell under the pane's login shell. +# The adapter now proves an agent at process level before trusting a +# registration, and this guard measures the two vendor facts that proof rests +# on, which no fixture can prove: +# +# 1. how Pi presents in `pane process-info` (on Herdr 0.9.0 the kernel name +# is `node` and only argv0 says `pi`), so the shared process classifier +# must still attribute the running harness as `agent`; +# 2. whether this Herdr release still leaves the registration behind after +# Pi quits under a nested shell, so the stale-registration branch is +# exercised against the real record rather than a canned one. +# +# It fails naming the Herdr and Pi versions when either fact drifts. Pi is +# launched with no prompt and quit immediately, so no model token is spent and +# the shared live gate runs it by default wherever both tools are installed. +# Run it after every Herdr or Pi upgrade and before trusting a refreshed +# docs/verification/runtime-backends.md "Stale agent registration" entry. +# +# Always runs on a private, named, throwaway lab session, never the default +# one (tests/herdr-test-safety.sh; bin/fm-herdr-lab.sh owns the isolation). +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" + +fail() { printf 'not ok - %s\n' "$1" >&2; exit 1; } +pass() { printf 'ok - %s\n' "$1"; } +note() { printf '# %s\n' "$1"; } + +fm_live_gate default-on FM_HERDR_PI_STALE_REGISTRATION_LIVE_E2E herdr pi jq + +# shellcheck source=tests/herdr-test-safety.sh +. "$ROOT/tests/herdr-test-safety.sh" +herdr_forget_inherited_pane + +HERDR_VERSION=$(herdr --version 2>&1 | head -1) +HERDR_VERSION=${HERDR_VERSION#herdr } +PI_VERSION=$(pi --version 2>/dev/null | head -1 | tr -d '\r') +[ -n "$PI_VERSION" ] || PI_VERSION=unknown +version_fail() { # + fail "$1 [herdr $HERDR_VERSION, pi $PI_VERSION]" +} + +SESSION="fm-lab-pi-stale-$$" +export HERDR_SESSION="$SESSION" +SCRATCH= +cleanup_all() { + local status=$? + [ -n "$SCRATCH" ] && rm -rf "$SCRATCH" + herdr_safe_stop_and_delete "$SESSION" + exit "$status" +} +trap cleanup_all EXIT +fm_herdr_lab_prepare "$SESSION" || fail "could not prepare isolated Herdr lab session" + +SCRATCH=$(mktemp -d "${TMPDIR:-/tmp}/fm-pi-stale.XXXXXX") +SCRATCH=$(cd "$SCRATCH" && pwd) +mkdir -p "$SCRATCH/cwd" + +# shellcheck source=/dev/null +. "$ROOT/bin/fm-backend.sh" +fm_backend_source herdr || fail "fm_backend_source herdr failed" + +lab() { fm_herdr_lab_cli "$SESSION" "$@"; } + +# prepare only records the tripwire; the adapter's own server-ensure starts +# the lab session's server exactly as a spawn would. +fm_backend_herdr_server_ensure "$SESSION" || fail "could not start the isolated Herdr lab server" +WS=$(lab workspace create --label fm-pi-stale --cwd "$SCRATCH/cwd" 2>&1) \ + || fail "could not create the lab workspace: $WS" +PANE_ID=$(printf '%s' "$WS" | jq -r '.result.root_pane.pane_id // empty') +[ -n "$PANE_ID" ] || fail "workspace create did not return a root pane id" +TARGET="$SESSION:$PANE_ID" + +wait_process_state() { # + local expected=$1 tries=$2 got i=0 + while [ "$i" -lt "$tries" ]; do + got=$(fm_backend_herdr_pane_process_state "$SESSION" "$PANE_ID") + [ "$got" = "$expected" ] && return 0 + sleep 0.2 + i=$((i + 1)) + done + return 1 +} + +registered_status() { + herdr agent get "$PANE_ID" --session "$SESSION" 2>/dev/null | jq -r '.result.agent.agent_status // empty' +} + +# The crew shape: a nested interactive shell under the pane's top shell, then +# the real Pi TUI with no prompt. +lab pane run "$PANE_ID" zsh >/dev/null 2>&1 || fail "could not start the nested shell in the pane" +sleep 1 +lab pane run "$PANE_ID" pi >/dev/null 2>&1 || fail "could not start pi in the pane" + +# Herdr creates the record with its own placeholder status (`unknown`, verified +# 0.9.0) the moment it notices Pi, before Pi's extension reports a lifecycle +# state; only a lifecycle state is the registration this guard is about. +STATUS= +for _ in $(seq 1 300); do + STATUS=$(registered_status) + case "$STATUS" in working|idle|done|blocked) break ;; esac + sleep 0.2 +done +case "$STATUS" in + working|idle|done|blocked) ;; + *) version_fail \ + "pi never reported a lifecycle state to Herdr in this pane (agent get read '${STATUS:-agent_not_found}' for 60s); the herdr pi integration (~/.pi/agent/extensions/herdr-agent-state.ts) is what reports it" ;; +esac + +wait_process_state agent 100 || version_fail \ + "pi is running and registered ($STATUS) but pane process-info reads '$(fm_backend_herdr_pane_process_state "$SESSION" "$PANE_ID")', not 'agent'. Observed foreground: $(herdr pane process-info --pane "$PANE_ID" --session "$SESSION" 2>/dev/null | jq -c '.result.process_info.foreground_processes'). Teach bin/fm-agent-process-lib.sh's fm_agent_process_classify the identity this release actually reports" +FOREGROUND=$(herdr pane process-info --pane "$PANE_ID" --session "$SESSION" 2>/dev/null \ + | jq -c '[.result.process_info.foreground_processes[] | {name, argv0}]') +STATE=$(fm_backend_agent_state herdr "$TARGET") +[ "$STATE" = alive ] || version_fail "a running, registered pi reads '$STATE' rather than 'alive' (registration '$(registered_status)', pane state '$(fm_backend_herdr_pane_agent_state "$SESSION" "$PANE_ID")', process state '$(fm_backend_herdr_pane_process_state "$SESSION" "$PANE_ID")', agent get: $(herdr agent get "$PANE_ID" --session "$SESSION" 2>&1 | tr -d '\n'))" +note "pi $PI_VERSION under herdr $HERDR_VERSION: registered $STATUS, foreground $FOREGROUND" +pass "real herdr $HERDR_VERSION + pi $PI_VERSION: a running registered pi classifies alive at process level" + +# Quit Pi to the nested shell. A slash command can open a completion popup that +# swallows the first Enter, so one extra Enter is allowed before judging. +lab pane send-text "$PANE_ID" '/quit' >/dev/null 2>&1 || fail "could not type /quit" +sleep 0.5 +lab pane send-keys "$PANE_ID" Enter >/dev/null 2>&1 || fail "could not submit /quit" +if ! wait_process_state shell 50; then + lab pane send-keys "$PANE_ID" Enter >/dev/null 2>&1 || true + wait_process_state shell 150 || version_fail \ + "pi did not exit to a shell within 40s of /quit; pane process-info reads '$(fm_backend_herdr_pane_process_state "$SESSION" "$PANE_ID")'" +fi + +# Let Herdr settle whatever release it is going to do, then read the record. +sleep 2 +STATUS=$(registered_status) +PANE_STATE=$(fm_backend_herdr_pane_agent_state "$SESSION" "$PANE_ID") +STATE=$(fm_backend_agent_state herdr "$TARGET") +BUSY=$(fm_backend_herdr_busy_state "$TARGET") +[ "$STATE" = dead ] || version_fail \ + "after pi quit to a shell the endpoint recovers as '$STATE' (pane state '$PANE_STATE', registration '${STATUS:-none}') rather than 'dead'; every relaunch would be refused" +[ "$BUSY" != busy ] || version_fail "a shell-only pane after pi quit reads busy (registration '${STATUS:-none}')" +if [ -n "$STATUS" ]; then + [ "$PANE_STATE" = stale-agent ] || version_fail \ + "Herdr kept the registration ($STATUS) over the shell-only pane but the classifier reads '$PANE_STATE' rather than 'stale-agent'" + note "herdr $HERDR_VERSION kept the pi registration ($STATUS) after /quit under a nested shell: the stale-registration branch is exercised" + pass "real herdr $HERDR_VERSION + pi $PI_VERSION: the registration left behind by a quit pi reads stale-agent and recovers as dead" +else + [ "$PANE_STATE" = no-agent ] || version_fail \ + "Herdr released the registration but the pane reads '$PANE_STATE' rather than 'no-agent'" + note "herdr $HERDR_VERSION released the pi registration after /quit under a nested shell; the stale-registration branch was not exercised by this release, the agent-free verdict still held through agent_not_found" + pass "real herdr $HERDR_VERSION + pi $PI_VERSION: a quit pi under a nested shell recovers as dead" +fi diff --git a/tests/fm-omp-harness.test.sh b/tests/fm-omp-harness.test.sh index 0b3f597fedd..7c25848eedd 100755 --- a/tests/fm-omp-harness.test.sh +++ b/tests/fm-omp-harness.test.sh @@ -97,10 +97,10 @@ test_lock_identity_and_liveness_classification() { # shellcheck source=bin/fm-backend.sh . "$ROOT/bin/fm-backend.sh" fm_backend_source tmux || fail "fm_backend_source tmux failed" - [ "$(fm_backend_tmux_classify_process_name omp)" = agent ] || fail "tmux liveness must classify omp as an agent" - [ "$(fm_backend_tmux_classify_process_name /opt/omp/bin/omp)" = agent ] || fail "tmux liveness must classify an omp path as an agent" - [ "$(fm_backend_tmux_classify_process_name ompd)" != agent ] || fail "tmux liveness must not classify ompd as an agent" - [ "$(fm_backend_tmux_classify_process_name comp)" != agent ] || fail "tmux liveness must not classify comp as an agent" + [ "$(fm_agent_process_classify_name omp)" = agent ] || fail "tmux liveness must classify omp as an agent" + [ "$(fm_agent_process_classify_name /opt/omp/bin/omp)" = agent ] || fail "tmux liveness must classify an omp path as an agent" + [ "$(fm_agent_process_classify_name ompd)" != agent ] || fail "tmux liveness must not classify ompd as an agent" + [ "$(fm_agent_process_classify_name comp)" != agent ] || fail "tmux liveness must not classify comp as an agent" pass "session lock and tmux liveness: omp is anchored, decoys stay out" } diff --git a/tests/fm-tmux-agent-liveness.test.sh b/tests/fm-tmux-agent-liveness.test.sh index 321f51c68db..ce31e801e1d 100755 --- a/tests/fm-tmux-agent-liveness.test.sh +++ b/tests/fm-tmux-agent-liveness.test.sh @@ -117,7 +117,7 @@ wait_for_state() { # [tries] title_classifies_agent() { # local name name=$(fm_backend_tmux_current_command "$1" 2>/dev/null) - [ "$(fm_backend_tmux_classify_process_name "$name")" = agent ] + [ "$(fm_agent_process_classify_name "$name")" = agent ] } # Does the foreground-process-group identity, including argv[0], name one? @@ -125,13 +125,13 @@ comms_classify_agent() { # local name while IFS= read -r name; do [ -n "$name" ] || continue - [ "$(fm_backend_tmux_classify_process_name "$name")" = agent ] && return 0 + [ "$(fm_agent_process_classify_name "$name")" = agent ] && return 0 done <> "$LOG" jq_state() { jq "$@" "$STATE"; } save() { tmp="$STATE.tmp.$$"; cat > "$tmp" && mv "$tmp" "$STATE"; } -ws=""; label=""; cwd="" +ws=""; label=""; cwd=""; pane="" args=("$@") for ((i=0; i<${#args[@]}; i++)); do case "${args[$i]}" in --workspace) ws=${args[$((i+1))]:-} ;; --label) label=${args[$((i+1))]:-} ;; --cwd) cwd=${args[$((i+1))]:-} ;; + --pane) pane=${args[$((i+1))]:-} ;; esac done case "${1:-} ${2:-}" in @@ -97,7 +98,9 @@ case "${1:-} ${2:-}" in [ ! -f "$SEND_FAIL" ] || exit 1 jq_state --arg p "${3:-}" '.typed[$p] = true | .working[$p] = true' | save ;; "pane read") printf '\n' ;; - "pane process-info") printf '{"result":{"process":{"name":"codex"}}}\n' ;; + "pane process-info") + printf '{"result":{"type":"pane_process_info","process_info":{"pane_id":"%s","shell_pid":%s,"foreground_process_group_id":%s,"foreground_processes":[{"pid":%s,"name":"codex","argv0":"codex","argv":["codex"],"cmdline":"codex"}]}}}\n' \ + "$pane" "$$" "$$" "$$" ;; "agent get") pane=${3:-} if [ "$(jq_state -r --arg p "$pane" '.working[$p] // false')" = true ]; then From 31f062d43d25775f8010480cbcba56134aac1998 Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Fri, 11 Sep 2026 10:24:18 -0700 Subject: [PATCH 05/31] feat: add live-head merge gates and away task grants (#4199) * Bind GitHub merges to a live green head and require an away-task grant. A GitHub merge now re-reads the pull request and passes --match-head-commit, so a red or moved head cannot land the way GitLab already refused. While an away record exists, only yolo or a named grant may merge, so hold-for-return cannot ship an ungated PR. Co-authored-by: Cursor * no-mistakes(review): Harden away merge authorization and grant parsing * no-mistakes(review): Restrict fallback outcomes to proved GitHub merges * no-mistakes(document): Refresh merge safety documentation * no-mistakes(ci): Fixed all three CI failures by updating legacy GitHub merge fixtures for live-head verification/direct gh merges and removing a process-event runner cleanup race. Verified fm-pr-check-security, fm-captain-hold-lifecycle, and fm-watch-triage pass locally; shell syntax and git diff checks also pass * no-mistakes(document): Document attended red-check exception --------- Co-authored-by: Cursor --- .agents/skills/afk/SKILL.md | 8 +- AGENTS.md | 5 +- bin/fm-afk-contract.sh | 119 +++- bin/fm-afk-launch.sh | 3 + bin/fm-merge-outcome-lib.sh | 18 +- bin/fm-pr-merge.sh | 329 +++++++++- docs/architecture.md | 6 +- docs/gitlab-merge-watch.md | 2 +- tests/fm-afk-contract.test.sh | 98 +++ tests/fm-captain-hold-lifecycle.test.sh | 25 +- tests/fm-pr-check-security.test.sh | 14 +- tests/fm-pr-merge.test.sh | 774 +++++++++++++++--------- 12 files changed, 1043 insertions(+), 358 deletions(-) diff --git a/.agents/skills/afk/SKILL.md b/.agents/skills/afk/SKILL.md index d2375ac6ffc..573eed53c07 100644 --- a/.agents/skills/afk/SKILL.md +++ b/.agents/skills/afk/SKILL.md @@ -27,7 +27,10 @@ Hold-for-return is the default and the only reach profile this release records: Write only clauses the words actually support; a wish with no object or no stated precondition is not a clause. Plain `/afk` with no words has no clauses. 2. **Propose and read back.** - Run `bin/fm-afk-launch.sh propose --words-file [--action --object --when [--stop ]]... [--expected-return ] [--spend ]` (or `--words `), and relay its read-back to the captain in `AGENTS.md` section 9 language: the accepted clauses as a numbered list, every refused clause with the part it is missing, the expected return, the spend cap, and the one-sentence reach announcement. + Run `bin/fm-afk-launch.sh propose --words-file [--action --object --when [--stop ]]... [--expected-return ] [--spend ] [--grant ]...` (or `--words `), and relay its read-back to the captain in `AGENTS.md` section 9 language: the accepted clauses as a numbered list, every refused clause with the part it is missing, the expected return, the spend cap, any merge-when-green task ids, and the one-sentence reach announcement. + When the captain names task ids that may merge while green, pass `--grant ` for each named id. + Never infer task ids from clause prose, object text, or the away words. + Red-check exceptions stay in the words or clause `when` text and are not executed. A refused clause does not fail the proposal; the captain can restate it or leave it refused. Exit 3 only means a clause was refused; the proposal stands. 3. **Confirm on the captain's go.** @@ -79,6 +82,9 @@ Bias ambiguous cases toward exit: a present captain beats token savings, and a f afk changes how the captain is informed and what happens at a captain-owned decision point, **not who approves what**. "Away" never means "approves more" or "approves less." A PR ready for merge keeps the merge authority from `AGENTS.md` section 7, and a needs-decision finding keeps the `ask-user-authority` policy; anything requiring the captain still waits for the captain's explicit word. +While the away-posture record exists, a merge proceeds only when that task's recorded yolo posture is on or its id is in the record's merge-grant list; otherwise it is held for the captain's return. +A merge grant never releases a captain hold, and it expires when the away record is archived. +`--allow-red` remains attended-only and is refused while the record exists. A mandate clause is the captain's explicit instruction given before leaving, recorded with its named object and condition; a clause is never inferred, never applied by analogy, and expires at return. Forbidden, destructive, irreversible, and security-sensitive actions are never pre-authorizable regardless of clause text, and no recorded clause is authority by itself. This release records clauses and does not execute them. diff --git a/AGENTS.md b/AGENTS.md index 906195aef1e..167d14da4cc 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -346,8 +346,9 @@ The path's worker, automated gates, and captain approval remain authoritative: Delivery mode and `yolo` are orthogonal. `yolo` governs merge authority only: with it off, the captain approves every PR merge and every local-only landing; with it on, firstmate merges green, in-scope work itself. -Never merge a red PR under either setting; destructive, irreversible, and security-sensitive merges still escalate. -Without a current explicit captain instruction that states the concrete merge, that default stands, and standing `yolo` cannot authorize a red merge; section 1 owns when such an instruction overrides a Firstmate-written standing rule within its exact scope. +Never merge a red PR under either setting unless a current explicit captain instruction names the single GitHub check waived through `fm-pr-merge.sh --allow-red`; that attended-only waiver still requires every other check green. +Destructive, irreversible, and security-sensitive merges still escalate. +Without a current explicit captain instruction that states the concrete merge, the green default stands, and standing `yolo` cannot authorize a red merge; section 1 owns when such an instruction overrides a Firstmate-written standing rule within its exact scope. Load `ask-user-authority` before deciding any ask-user finding; the implementation worker never answers its own finding. Use `bin/fm-pr-merge.sh` for every task PR merge so merge metadata is recorded and an unproved merge is refused instead of reported as landed, and use `bin/fm-merge-local.sh` for approved local-only landing; never call a lower-level merge command around their guards. After an autonomous merge, give the captain a one-line full-URL or local-main outcome. diff --git a/bin/fm-afk-contract.sh b/bin/fm-afk-contract.sh index 593b65e3a01..272e810846d 100755 --- a/bin/fm-afk-contract.sh +++ b/bin/fm-afk-contract.sh @@ -21,6 +21,9 @@ # reach_channels: none # reach_announced: # spend_max_concurrent_workers: +# merge_grants: - | task ids that may merge while this record exists +# - (empty is `merge_grants: -`; a missing field on +# ... a pre-field v1 record reads as an empty list) # confirmed: # confirmed_epoch: # words: | or |- the captain's words, verbatim, never edited, @@ -82,12 +85,14 @@ # Usage: # fm-afk-contract.sh propose [--words-file | --words ] # [--action --object --when [--stop ]]... -# [--expected-return ] [--spend ] +# [--expected-return ] [--spend ] [--grant ]... # Compile and write the proposal, then print the read-back. Exit 0 with every # clause accepted, 3 when at least one clause was refused (the read-back names # the missing part), and 2 on a usage error. --words-file keeps the file's # bytes verbatim, trailing newlines included. A refused clause remains in the -# proposal so the captain can restate it before saying go. +# proposal so the captain can restate it before saying go. Repeatable --grant +# records captain-named task ids that may merge-when-green while the record +# exists; invalid or duplicate ids are a usage error, never a refused clause. # fm-afk-contract.sh confirm # Promote the proposal into the record with the confirmed timestamp and # print the entry announcement. A proposal is required when no confirmed @@ -103,6 +108,7 @@ # (`\\`, `\t`, `\r`, and `\n`) so every record remains one row per clause; # a literal `-` is `\x2d` to distinguish it from the empty-stop marker. # fm-afk-contract.sh refused [--proposal | --path ] TSV: id text missing +# fm-afk-contract.sh grants [--proposal | --path ] one task id per line # fm-afk-contract.sh archive move the record aside; print its path # fm-afk-contract.sh archived print that archived record's path # @@ -162,6 +168,15 @@ fm_afk_contract_blank() { # [ -z "$(printf '%s' "$1" | tr -d '[:space:]')" ] } +# Same alphabet as fm_pr_task_id_valid / fm_task_id_path_safe in bin/fm-pr-lib.sh. +# Kept local so sourcing this file cannot reset that library's parse globals. +fm_afk_contract_grant_id_valid() { # + local LC_ALL=C id=${1-} + case "$id" in + ''|.*|*[!A-Za-z0-9._-]*) return 1 ;; + esac +} + fm_afk_contract_escape() { # local value=$1 value=${value//\\/\\\\} @@ -266,9 +281,10 @@ fm_afk_contract_validate_iso() { # # Compile every input into a record body on stdout (everything except the # confirmed fields). Inputs: WORDS (verbatim), the parallel clause field arrays -# CLAUSE_ACTIONS CLAUSE_OBJECTS CLAUSE_WHENS CLAUSE_STOPS, EXPECTED_RETURN, SPEND. +# CLAUSE_ACTIONS CLAUSE_OBJECTS CLAUSE_WHENS CLAUSE_STOPS, EXPECTED_RETURN, +# SPEND, MERGE_GRANTS. fm_afk_contract_render_body() { # - local entered=$1 entered_epoch=$2 ordinal=0 i as_given + local entered=$1 entered_epoch=$2 ordinal=0 i as_given grant local accepted_block="" refused_block="" i=0 while [ "$i" -lt "${#CLAUSE_ACTIONS[@]}" ]; do @@ -303,6 +319,14 @@ fm_afk_contract_render_body() { # printf 'reach_channels: none\n' printf 'reach_announced: %s\n' "$FM_AFK_CONTRACT_REACH_ANNOUNCED" printf 'spend_max_concurrent_workers: %s\n' "${SPEND:-$FM_AFK_CONTRACT_SPEND_DEFAULT}" + if [ "${#MERGE_GRANTS[@]}" -eq 0 ]; then + printf 'merge_grants: -\n' + else + printf 'merge_grants:\n' + for grant in "${MERGE_GRANTS[@]}"; do + printf ' - %s\n' "$grant" + done + fi if [ -n "$WORDS" ]; then local words_body=$WORDS words_indicator='|-' case "$words_body" in @@ -370,6 +394,52 @@ fm_afk_contract_read_words() { # ' "$path" } +# One granted task id per line. A missing merge_grants field is an empty list +# so a pre-field v1 record fails closed for non-yolo merges instead of skipping +# the grant check. A present but unreadable field fails rather than guessing. +fm_afk_contract_read_grants() { # + local path=$1 + [ -f "$path" ] || return 1 + awk -v record="$path" ' + function die(reason) { + printf "fm-afk-contract: record %s has an invalid merge_grants field: %s\n", record, reason > "/dev/stderr" + bad = 1 + exit 2 + } + function valid_id(value) { + if (value == "" || substr(value, 1, 1) == ".") return 0 + return value ~ /^[A-Za-z0-9._-]+$/ + } + /^merge_grants:/ { + if (found) die("the field is defined more than once") + found = 1 + if ($0 == "merge_grants: -") { empty = 1; next } + if ($0 == "merge_grants:") { inlist = 1; next } + die("the empty form is merge_grants: -") + } + inlist && /^ - / { + id = substr($0, 5) + if (!valid_id(id)) die("task id \"" id "\" is not a valid task id") + if (seen[id]++) die("task id \"" id "\" is listed more than once") + print id + count++ + next + } + inlist && /^[^ ]/ { + if (count == 0) die("the list form has no stored ids") + inlist = 0 + next + } + empty && /^[^ ]/ { empty = 0; next } + inlist || empty { die("a stored grant line is malformed") } + END { + if (bad) exit 2 + if (!found) exit 0 + if (inlist && count == 0) die("the list form has no stored ids") + } + ' "$path" +} + # TSV rows for a list section:
is clauses or refused. fm_afk_contract_read_list() { #
local path=$1 section=$2 @@ -466,6 +536,10 @@ fm_afk_contract_validate() { # words_header=$(sed -n '/^words: /{p;q;}' "$path") case "$words_header" in 'words: -'|'words: |'|'words: |-') ;; *) fm_afk_contract_log "record $path has no valid words field"; return 1 ;; esac fm_afk_contract_read_words "$path" >/dev/null || return 1 + fm_afk_contract_read_grants "$path" >/dev/null || { + fm_afk_contract_log "record $path has no valid merge_grants field" + return 1 + } if [ "$require_confirmed" -eq 1 ]; then confirmed=$(fm_afk_contract_read_field "$path" confirmed) fm_afk_contract_validate_iso "$confirmed" || { fm_afk_contract_log "record $path has no valid confirmed time"; return 1; } @@ -526,13 +600,22 @@ EOF # --- rendering -------------------------------------------------------------- fm_afk_contract_render_readback() { # - local path=$1 title=$2 words count id action object when stop text missing expected spend flag + local path=$1 title=$2 words count id action object when stop text missing expected spend flag grants grant_list expected=$(fm_afk_contract_read_field "$path" expected_return) spend=$(fm_afk_contract_read_field "$path" spend_max_concurrent_workers) + grants=$(fm_afk_contract_read_grants "$path") || return 1 + grant_list= + while IFS= read -r id; do + [ -n "$id" ] || continue + grant_list="${grant_list:+$grant_list, }$id" + done <<EOF +$grants +EOF printf '%s\n' "$title" printf ' entered: %s\n' "$(fm_afk_contract_read_field "$path" entered)" printf ' expected return: %s\n' "$( [ "$expected" = - ] && printf 'not given' || printf '%s' "$expected")" printf ' spend cap: %s concurrent workers\n' "$spend" + printf ' merge when green (task ids): %s\n' "${grant_list:-(none)}" printf ' reach: hold-for-return only. %s\n' "$(fm_afk_contract_read_field "$path" reach_announced)" words=$(fm_afk_contract_read_words "$path"; printf x) words=${words%x} @@ -601,10 +684,11 @@ fm_afk_contract_render_announcement() { # <path> # --- subcommands ------------------------------------------------------------ -fm_afk_contract_parse_inputs() { # <args...>; sets WORDS, the CLAUSE_* arrays, EXPECTED_RETURN, SPEND - local words_file='' open=-1 +fm_afk_contract_parse_inputs() { # <args...>; sets WORDS, the CLAUSE_* arrays, EXPECTED_RETURN, SPEND, MERGE_GRANTS + local words_file='' open=-1 grant WORDS=; EXPECTED_RETURN=-; SPEND=$FM_AFK_CONTRACT_SPEND_DEFAULT CLAUSE_ACTIONS=(); CLAUSE_OBJECTS=(); CLAUSE_WHENS=(); CLAUSE_STOPS=(); CLAUSE_STOP_GIVENS=() + MERGE_GRANTS=() while [ "$#" -gt 0 ]; do case "$1" in --words-file) @@ -642,6 +726,23 @@ fm_afk_contract_parse_inputs() { # <args...>; sets WORDS, the CLAUSE_* arrays, case "$2" in ''|*[!0-9]*|0) fm_afk_contract_log "--spend must be a positive integer, got '$2'"; return 2 ;; esac SPEND=$2 shift 2 ;; + --grant) + [ "$#" -gt 1 ] || { fm_afk_contract_log '--grant requires a task id'; return 2; } + fm_afk_contract_grant_id_valid "$2" || { + fm_afk_contract_log "--grant must be a valid task id, got '$2'" + return 2 + } + for grant in "${MERGE_GRANTS[@]+"${MERGE_GRANTS[@]}"}"; do + [ "$grant" != "$2" ] || { + fm_afk_contract_log "--grant lists '$2' more than once" + return 2 + } + done + MERGE_GRANTS+=("$2") + shift 2 ;; + --grant=*) + fm_afk_contract_log '--grant takes a separate task-id argument' + return 2 ;; *) fm_afk_contract_log "unknown option '$1'" return 2 ;; @@ -812,6 +913,10 @@ fm_afk_contract_main() { refused) path=$(fm_afk_contract_select_path "$@") || { fm_afk_contract_usage >&2; return 2; } fm_afk_contract_read_list "$path" refused ;; + grants) + path=$(fm_afk_contract_select_path "$@") || { fm_afk_contract_usage >&2; return 2; } + [ -f "$path" ] || { fm_afk_contract_log "no record at $path"; return 1; } + fm_afk_contract_read_grants "$path" ;; archive) fm_afk_contract_cmd_archive ;; archived) [ "$#" -eq 1 ] || { fm_afk_contract_usage >&2; return 2; } diff --git a/bin/fm-afk-launch.sh b/bin/fm-afk-launch.sh index a821236e233..2466356f5bd 100755 --- a/bin/fm-afk-launch.sh +++ b/bin/fm-afk-launch.sh @@ -40,11 +40,14 @@ # fm-afk-launch.sh propose [--words-file <path> | --words <text>] # [--action <verb> --object <text> --when <text> [--stop <text>]]... # [--expected-return <UTC ISO 8601>] [--spend <n>] +# [--grant <task-id>]... # Record the captain's away words and mandate # clause fields into a proposal and print the # read-back. Exit 3 when a clause was refused (its # missing part is named in the read-back); the # proposal still records it as refused. +# Repeatable --grant records captain-named task +# ids that may merge-when-green while away. # fm-afk-launch.sh confirm Promote the required proposal and print the entry # announcement. On Pi this is the whole entry. # fm-afk-launch.sh start Capture the captain pane, then (unless the daemon diff --git a/bin/fm-merge-outcome-lib.sh b/bin/fm-merge-outcome-lib.sh index ab0b96a6778..bf1f26c9c17 100755 --- a/bin/fm-merge-outcome-lib.sh +++ b/bin/fm-merge-outcome-lib.sh @@ -35,27 +35,37 @@ _FM_MERGE_OUTCOME_LIB_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" # shellcheck disable=SC2034 # Public result consumed by sourcing callers. FM_MERGE_OUTCOME_ALREADY_RECORDED=false -# fm_merge_outcome_report <home> <state> <task-id> <pr-url> <origin> +# fm_merge_outcome_report <home> <state> <task-id> <pr-url> <origin> [authority] # # <origin> says who observed the merge, because that decides whether the # existing poll path also needs a local wake: # self - this home performed the merge. # poll - this home's merge poll detected the merge, so the canonical outcome # also wakes this home after any upward hop needed by a secondmate. +# Optional <authority> is yolo or away-grant when the merge ran while the +# away-posture record existed; it is appended to the ledger line. Known audit +# gap: queued merges and a poll that wins direct-merge deduplication publish an +# untagged row because the poll path does not persist merge authority. # # Returns 0 when the outcome is recorded (or already was), 2 on an invalid # request, 3 when this home's own role or parent binding cannot be read well # enough to say where the outcome belongs, and 1 on any other failure to # record. A caller that has already merged must report a non-zero return rather # than treat it as success: the merge landed and the record did not. -fm_merge_outcome_report() { # <home> <state> <task-id> <pr-url> <origin> +fm_merge_outcome_report() { # <home> <state> <task-id> <pr-url> <origin> [authority] local home=$1 state=$2 id=$3 url=$4 origin=$5 + local authority=${6-} suffix= local self_rc=0 destination='' line lock status=0 local provider host path number # shellcheck disable=SC2034 # Sourced wake helpers consume these scoped globals. local STATE FM_WAKE_QUEUE FM_WAKE_QUEUE_LOCK FM_MERGE_OUTCOME_ALREADY_RECORDED=false case "$origin" in self|poll) ;; *) return 2 ;; esac + case "$authority" in + yolo|away-grant) suffix=" $authority" ;; + '') ;; + *) return 2 ;; + esac fm_pr_task_id_valid "$id" || return 2 fm_pr_url_parse "$url" || return 2 provider=$FM_PR_PROVIDER @@ -65,7 +75,7 @@ fm_merge_outcome_report() { # <home> <state> <task-id> <pr-url> <origin> [ -d "$state" ] && [ ! -L "$state" ] || return 1 if destination=$(fm_parent_channel_destination "$home" "$state"); then - line="done [key=merged-$id]: merged $id $FM_PR_URL" + line="done [key=merged-$id]: merged $id $FM_PR_URL$suffix" else self_rc=$? [ "$self_rc" -eq 1 ] || return 3 @@ -90,7 +100,7 @@ fm_merge_outcome_report() { # <home> <state> <task-id> <pr-url> <origin> fi if [ "$status" -eq 0 ] && { [ "$origin" = poll ] || [ -z "$destination" ]; }; then fm_wake_append check "merged-$id-$FM_PR_URL" \ - "check: merge landed: $id $FM_PR_URL" || status=1 + "check: merge landed: $id $FM_PR_URL$suffix" || status=1 fi if [ "$status" -eq 0 ]; then fm_pr_poll_merge_mark_notified "$state" "$id" \ diff --git a/bin/fm-pr-merge.sh b/bin/fm-pr-merge.sh index 4111f1e2c50..174c3188dae 100755 --- a/bin/fm-pr-merge.sh +++ b/bin/fm-pr-merge.sh @@ -2,23 +2,34 @@ # Merge a task's PR or MR after recording pr= and any available pr_head= through # bin/fm-pr-check.sh, so teardown can verify landed work after squash merges. # The full canonical URL is parsed by bin/fm-pr-lib.sh. A GitHub pull request is -# addressed through gh-axi by the derived owner and repository; a GitLab merge +# addressed through gh by the derived owner and repository; a GitLab merge # request is addressed through glab by the project URL rebuilt from the parsed # host and path, so any instance works and no host is hardcoded. # # Merge method on GitHub defaults to --squash when the caller passes none of # --squash, --merge, --rebase, or --method after the optional -- separator. -# The gh-axi merge abstraction always performs the merge; the outcome read that -# follows it never becomes a prerequisite for reaching that abstraction. After -# gh-axi returns success, GitHub's live state is read back and accepted only -# when the pull request is merged or in the merge queue. gh's GraphQL API -# supplies that queue-aware read when gh is on PATH; when gh is absent or its -# read fails, gh-axi's own view still proves a landed merge, and every outcome -# it cannot prove refuses, reporting the single failed read when gh is absent -# and naming both failed reads when gh is present and its own read failed. +# A GitHub merge is refused unless every pre-merge condition holds, each read +# live at merge time rather than taken from recorded metadata: the pull request +# is open, not a draft, mergeable, free of conflicts, and every unwaived check +# is green at the exact current head commit. Every failing condition is reported, not +# just the first. The verified head is then passed to gh as +# --match-head-commit, so a push that lands between that read and the merge +# fails the merge instead of landing commits nothing verified. Reading that +# state needs gh and jq, and either one absent stops the merge before any +# state is recorded. An attended --allow-red <check-name> may be passed once, +# with the name as a separate argument; it waives only checks with that exact +# name, still requires every other check green, and still binds the head. It is +# refused while the away-posture record exists, and it never +# applies on GitLab, where a merge already requires the head pipeline to have +# succeeded. After gh returns success, GitHub's live state is read back and +# accepted only when the pull request is merged or in the merge queue. gh's +# GraphQL API supplies that queue-aware read; when that read fails, gh-axi's +# own view still proves a landed merge, and every outcome it cannot prove +# refuses, reporting the failed gh read and naming both failed reads when the +# gh-axi view could not prove the outcome either. # If the pull request remains open and the base branch has an effective # merge_queue rule, the refusal names the queue's configured merge method and -# the exact -- --auto --<method> retry flags, unless the caller already passed +# the exact --attended-override -- --auto --<method> retry flags, unless the caller already passed # that method with --auto to a merge command that returned success, in which # case it reports instead that the accepted request has not entered the queue # and the queue state has to be re-checked. @@ -30,9 +41,7 @@ # queued is refused the same way and says auto-merge was armed with nothing # landed or queued yet, or, when the merge command itself failed, that auto-merge # was only requested; both are read from the caller's own arguments rather than -# from the forge's prose. The observed state is judged the same way whichever -# read produced it, and a refusal built on the gh-axi view says the merge queue -# could not be observed at all rather than implying an unqueued pull request. +# from the forge's prose. # Every refusal that follows a merge command which returned success quotes that # command's own output, marked as the forge's text and kept apart from this # script's verdict, including the refusal for an outcome that cannot be read; @@ -56,13 +65,28 @@ # Before either forge merge, the task's existing per-task control lock # serializes the captain-hold check through the forge command. A still-held or # unreadable row refuses before that command, so a captain approval must be -# recorded as an `answer --release` before this entrypoint is invoked. The lock -# ends when the local forge command returns; docs/captain-hold-lifecycle.md owns +# recorded as an `answer --release` before this entrypoint is invoked. While +# state/.afk-contract exists, a merge for this task also proceeds only if its +# meta yolo=on or its id is in that record's merge-grant list; otherwise it is +# held for the captain return. An unreadable record refuses rather than being +# skipped. Neither posture releases a captain hold, and the grant lapses when +# the record is archived. +# The lock ends when the local forge command returns; docs/captain-hold-lifecycle.md owns # the accepted asynchronous-landing and merge-to-cleanup residuals. # # Extra args must not include --repo or -R in any form, including a bundled # short-option cluster such as -yR, because the repository comes only from the -# URL, nor --sha on GitLab because the head comes only from the live read. +# URL, nor --sha or --match-head-commit because the head comes only from the +# live read. An existing task-meta pr= must equal the requested canonical URL; +# a task cannot be rebound here. Auto-merge (--auto), a protection bypass +# (--admin), and branch +# deletion (--delete-branch, -d and short-flag clusters, and GitLab's +# --remove-source-branch) are refused by default; --attended-override, parsed +# before the optional -- separator, re-enables those forge flags for an +# explicit captain instruction and never skips the live green check, the +# away-grant check, or a captain hold. +# +# Usage: fm-pr-merge.sh <task-id> <pr-url> [--attended-override] [--allow-red <check-name>] [-- <extra forge merge args>] # # On GitLab, this script confirms the MR is actually merged before reporting it; # an auto-merge-queued or unconfirmed request leaves the poll armed and records @@ -70,7 +94,6 @@ # destination, normal-case deduplication, and at-least-once recovery. # A landed merge whose outcome cannot be written is reported loudly rather than # misreported as a failed merge. -# Usage: fm-pr-merge.sh <task-id> <pr-url> [-- <extra forge merge args>] set -eu SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -84,6 +107,8 @@ STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" . "$SCRIPT_DIR/fm-backlog-transition-lib.sh" # shellcheck source=bin/fm-merge-outcome-lib.sh . "$SCRIPT_DIR/fm-merge-outcome-lib.sh" +# shellcheck source=bin/fm-afk-contract.sh +. "$SCRIPT_DIR/fm-afk-contract.sh" if [ "$#" -lt 2 ]; then echo "error: invalid PR merge request" >&2 @@ -104,7 +129,36 @@ PR_NUMBER=$FM_PR_NUMBER # rebuilt from the parsed identity rather than read from any ambient default. PROJECT_URL="https://$FM_PR_HOST/$FM_PR_PATH" shift 2 -[ "${1:-}" = "--" ] && shift +ATTENDED_OVERRIDE=false +ALLOW_RED=() +while [ "$#" -gt 0 ]; do + case "$1" in + --attended-override) + ATTENDED_OVERRIDE=true + shift + ;; + --attended-override=*) + echo "error: --attended-override takes no value" >&2 + exit 2 + ;; + --allow-red) + [ -n "${2:-}" ] || { echo "error: --allow-red requires a check name" >&2; exit 2; } + [ "${#ALLOW_RED[@]}" -eq 0 ] || { echo "error: --allow-red may be specified only once" >&2; exit 2; } + ALLOW_RED+=("$2") + shift 2 + ;; + --allow-red=*) + echo "error: --allow-red requires a separate check name argument" >&2 + exit 2 + ;; + --) shift; break ;; + *) break ;; + esac +done +if [ "${#ALLOW_RED[@]}" -gt 0 ] && [ "$PROVIDER" = gitlab ]; then + echo "error: --allow-red does not apply to GitLab, where a merge already requires the head pipeline to have succeeded" >&2 + exit 2 +fi caller_has_merge_method() { local arg @@ -180,7 +234,7 @@ reject_head_overrides() { local arg for arg in "$@"; do case "$arg" in - --sha|--sha=*) + --sha|--sha=*|--match-head-commit|--match-head-commit=*) echo "error: extra merge arguments must not override the head commit" >&2 return 1 ;; @@ -188,8 +242,29 @@ reject_head_overrides() { done } +reject_protected_forge_args() { + local arg + [ "$ATTENDED_OVERRIDE" = true ] && return 0 + for arg in "$@"; do + case "$arg" in + --auto|--auto=*|--admin|--admin=*|--delete-branch|--delete-branch=*|--remove-source-branch|--remove-source-branch=*) + echo "error: extra merge arguments must not request auto-merge, a protection bypass, or branch deletion; pass --attended-override only for an explicit captain instruction" >&2 + return 1 + ;; + --*) ;; + # A single-dash argument is a short-option cluster. -d is gh's + # --delete-branch, and -yd carries it the same way -yR carries --repo. + -*d*) + echo "error: extra merge arguments must not request auto-merge, a protection bypass, or branch deletion; pass --attended-override only for an explicit captain instruction" >&2 + return 1 + ;; + esac + done +} + reject_repo_overrides "$@" || exit 1 -[ "$PROVIDER" != gitlab ] || reject_head_overrides "$@" || exit 1 +reject_head_overrides "$@" || exit 1 +reject_protected_forge_args "$@" || exit 1 fm_backlog_directory_present "$STATE" "state directory" || { echo "error: PR merge refused: $FM_BACKLOG_TRANSITION_ERROR" >&2 @@ -247,6 +322,17 @@ if [ "$PROVIDER" = gitlab ]; then exit 1 fi fi +GITHUB_MISSING= +if [ "$PROVIDER" = github ]; then + command -v gh >/dev/null 2>&1 || GITHUB_MISSING="gh" + if ! command -v jq >/dev/null 2>&1; then + GITHUB_MISSING="${GITHUB_MISSING:+$GITHUB_MISSING and }jq" + fi + if [ -n "$GITHUB_MISSING" ]; then + echo "error: merging a GitHub pull request requires $GITHUB_MISSING on PATH" >&2 + exit 1 + fi +fi # The recorded head is read before bin/fm-pr-check.sh rewrites the metadata, # because that script re-records pr= and drops a pr_head= it cannot resolve. @@ -356,11 +442,130 @@ FIELDS FM_PR_MERGE_HEAD=$live_head } -# Read one live GitHub pull request view after gh-axi returns. The selected +# Every GitHub check that is not green in the given live pull-request JSON, one +# name per line: a status context whose state is not SUCCESS, or a check run +# that has not completed with SUCCESS, NEUTRAL, or SKIPPED (so a pending +# check is not green either). Exits nonzero when the rollup cannot be read, so +# a malformed answer is a failed read and never an empty red set. +github_checks_not_green() { + local json=$1 + printf '%s' "$json" | jq -r ' + if (.statusCheckRollup | type) != "array" then error("no check rollup") else . end + | .statusCheckRollup[] + | if .__typename == "CheckRun" then + {name: (.name // ""), ok: (.status == "COMPLETED" and (.conclusion == "SUCCESS" or .conclusion == "NEUTRAL" or .conclusion == "SKIPPED"))} + else + {name: (.context // ""), ok: (.state == "SUCCESS")} + end + | select(.ok | not) + | if .name == "" then "(unnamed check)" else .name end + ' 2>/dev/null || return 1 +} + +# Pre-merge conditions for a GitHub pull request, read from one live view. +# Sets FM_PR_MERGE_HEAD to the verified head on success. +github_verify_mergeable() { + local json fields line red name covered + local total=0 named=0 refusals='' + local state='' draft='' mergeable='' merge_state='' live_head='' + + if ! json=$(gh pr view "$URL" --json state,isDraft,mergeable,mergeStateStatus,headRefOid,statusCheckRollup 2>/dev/null) \ + || [ -z "$json" ]; then + echo "error: could not read the GitHub pull request state before merging" >&2 + return 1 + fi + if ! fields=$(printf '%s' "$json" | jq -r ' + if type == "object" then + "state=" + ((.state // "") | tostring), + "draft=" + (if (.isDraft | type) == "boolean" then (.isDraft | tostring) else "" end), + "mergeable=" + ((.mergeable // "") | tostring), + "merge_state=" + ((.mergeStateStatus // "") | tostring), + "head=" + ((.headRefOid // "") | tostring) + else + error("pull request payload is not an object") + end' 2>/dev/null); then + echo "error: could not read the GitHub pull request state before merging" >&2 + return 1 + fi + while IFS= read -r line; do + total=$((total + 1)) + case "$line" in + state=*) state=${line#state=} ;; + draft=*) draft=${line#draft=} ;; + mergeable=*) mergeable=${line#mergeable=} ;; + merge_state=*) merge_state=${line#merge_state=} ;; + head=*) live_head=${line#head=} ;; + *) continue ;; + esac + named=$((named + 1)) + done <<FIELDS +$fields +FIELDS + if [ "$named" -ne 5 ] || [ "$total" -ne 5 ]; then + echo "error: could not read the GitHub pull request state before merging" >&2 + return 1 + fi + + if ! fm_pr_head_valid "$live_head"; then + echo "error: could not read the GitHub pull request head commit before merging" >&2 + return 1 + fi + if ! red=$(github_checks_not_green "$json"); then + echo "error: could not read the GitHub pull request state before merging" >&2 + return 1 + fi + + case "$state" in + [oO][pP][eE][nN]) ;; + *) + refusals="$refusals - state is \"${state:-unreadable}\", not open +" + ;; + esac + [ "$draft" = false ] \ + || refusals="$refusals - the pull request is a draft +" + [ "$mergeable" = MERGEABLE ] \ + || refusals="$refusals - mergeable is \"${mergeable:-unreadable}\", not MERGEABLE +" + [ "$merge_state" != DIRTY ] \ + || refusals="$refusals - mergeStateStatus is DIRTY (conflicts) +" + + uncovered='' + while IFS= read -r name; do + [ -n "$name" ] || continue + covered=0 + if [ "${#ALLOW_RED[@]}" -gt 0 ]; then + for check in "${ALLOW_RED[@]}"; do + [ "$check" = "$name" ] && covered=1 + done + fi + [ "$covered" -eq 1 ] || { + refusals="$refusals - check '$name' is not green +" + uncovered="${uncovered:+$uncovered, }$name" + } + done <<EOF +$red +EOF + + if [ -n "$refusals" ]; then + printf 'error: refusing to merge %s\n' "$URL" >&2 + printf '%s' "$refusals" >&2 + [ -z "$uncovered" ] || printf 'error: these checks are not green: %s\n' "$uncovered" >&2 + return 1 + fi + printf 'verified: %s is open and mergeable, with every required check green at head %s\n' \ + "$URL" "$live_head" >&2 + FM_PR_MERGE_HEAD=$live_head +} + +# Read one live GitHub pull request view after gh returns. The selected # fields distinguish a landed pull request from a merge-queue entry and retain # the concrete state needed for a refusal. gh supplies the complete queue-aware -# view when available; gh-axi remains the degradation path that can prove a -# landed merge without making gh a prerequisite for the merge abstraction. +# view; if that post-merge read becomes unavailable, gh-axi is the degradation +# path that can prove only a landed merge. gh remains a pre-merge prerequisite. FM_PR_GITHUB_STATE= FM_PR_GITHUB_MERGED= FM_PR_GITHUB_QUEUED= @@ -435,7 +640,9 @@ github_read_outcome_with_gh_axi() { github_read_outcome() { if ! command -v gh >/dev/null 2>&1; then - github_read_outcome_with_gh_axi && return 0 + if github_read_outcome_with_gh_axi && [ "$FM_PR_GITHUB_MERGED" = true ]; then + return 0 + fi echo "error: could not read the GitHub pull request outcome after the merge attempt; PR metadata and merge poll remain recorded" >&2 return 1 fi @@ -557,6 +764,54 @@ require_released_captain_hold() { esac } +FM_PR_MERGE_AUTHORITY= +require_away_merge_grant() { + local yolo grants grant + FM_PR_MERGE_AUTHORITY= + fm_afk_contract_present "$STATE" || return 0 + if ! FM_HOME="$FM_HOME" FM_STATE_OVERRIDE="$STATE" \ + "$SCRIPT_DIR/fm-afk-contract.sh" validate >/dev/null 2>&1; then + echo "error: PR merge refused - the away-posture record could not be read; nothing was merged" >&2 + return 1 + fi + yolo=$(grep '^yolo=' "$META" | tail -1 | cut -d= -f2- || true) + if [ "$yolo" = on ]; then + FM_PR_MERGE_AUTHORITY=yolo + return 0 + fi + grants=$(FM_HOME="$FM_HOME" FM_STATE_OVERRIDE="$STATE" \ + "$SCRIPT_DIR/fm-afk-contract.sh" grants 2>/dev/null) || { + echo "error: PR merge refused - the away-posture record's grants could not be read; nothing was merged" >&2 + return 1 + } + while IFS= read -r grant; do + [ "$grant" = "$ID" ] || continue + FM_PR_MERGE_AUTHORITY=away-grant + return 0 + done <<EOF +$grants +EOF + echo "error: task $ID is held for the captain return" >&2 + return 1 +} + +require_current_away_authority() { + require_away_merge_grant || return 1 + if fm_afk_contract_present "$STATE" && [ "${#ALLOW_RED[@]}" -gt 0 ]; then + echo "error: --allow-red is attended-only; while the away-posture record exists the green check is absolute" >&2 + return 2 + fi +} + +require_recorded_pr_identity() { + local existing + existing=$(grep '^pr=' "$META" | tail -1 | cut -d= -f2- || true) + [ -n "$existing" ] || return 0 + [ "$existing" = "$URL" ] && return 0 + echo "error: task $ID is bound to $existing, not $URL" >&2 + return 1 +} + FM_PR_GITHUB_AUTO_REQUESTED=false FM_PR_GITHUB_MERGE_ACCEPTED=false FM_PR_GITHUB_CALLER_METHOD= @@ -615,7 +870,7 @@ github_report_queue_rules() { printf 'error: this run refuses even though the request for %s was accepted with the exact flags base branch %s requires (--auto --%s): the pull request has still not entered the merge queue, so no landed or queued outcome is proven; re-check the pull request'"'"'s merge queue state before retrying\n' \ "$URL" "$FM_PR_GITHUB_BASE" "$queue_method" >&2 else - printf 'error: base branch %s requires the merge queue; retry with: %s %s %s -- --auto --%s\n' \ + printf 'error: base branch %s requires the merge queue; retry with: %s %s %s --attended-override -- --auto --%s\n' \ "$FM_PR_GITHUB_BASE" "$0" "$ID" "$URL" "$queue_method" >&2 fi ;; @@ -681,7 +936,12 @@ gitlab_confirm_merged() { # Record before either forge call. This arms the merge poll without claiming a # landed outcome, so even a provider read failure after a real merge cannot # leave teardown without the PR identity it needs to verify the result. +away_status=0 +require_current_away_authority || away_status=$? +[ "$away_status" -eq 0 ] || exit "$away_status" +require_recorded_pr_identity || exit 1 record_pr_metadata || exit 1 +require_released_captain_hold || exit 1 case "$PROVIDER" in github) @@ -694,9 +954,15 @@ case "$PROVIDER" in FM_PR_GITHUB_AUTO_REQUESTED=true fi FM_PR_GITHUB_CALLER_METHOD=$(caller_merge_method "$@") - require_released_captain_hold || exit 1 + github_verify_mergeable || exit 1 + # This last presence and authority read narrows the publication race to the + # forge handoff; without a shared lock, a residual sub-second race remains. + away_status=0 + require_current_away_authority || away_status=$? + [ "$away_status" -eq 0 ] || exit "$away_status" merge_status=0 - merge_output=$(gh-axi pr merge "$PR_NUMBER" --repo "$PR_OWNER/$PR_REPO" \ + merge_output=$(gh pr merge "$PR_NUMBER" --repo "$PR_OWNER/$PR_REPO" \ + --match-head-commit "$FM_PR_MERGE_HEAD" \ "${merge_args[@]+"${merge_args[@]}"}" "$@" 2>&1) || merge_status=$? fm_lock_release "$MERGE_CONTROL_LOCK" || true MERGE_CONTROL_LOCK= @@ -737,7 +1003,11 @@ case "$PROVIDER" in # in between is refused by GitLab instead of merged unverified. --yes only # skips the interactive confirmation, which no supervised run can answer; # the conditions above are what authorize the merge. - require_released_captain_hold || exit 1 + # This last presence and authority read narrows the publication race to the + # forge handoff; without a shared lock, a residual sub-second race remains. + away_status=0 + require_current_away_authority || away_status=$? + [ "$away_status" -eq 0 ] || exit "$away_status" merge_status=0 GITLAB_HOST="$FM_PR_HOST" glab mr merge "$PR_NUMBER" -R "$PROJECT_URL" \ --sha "$FM_PR_MERGE_HEAD" --yes "$@" || merge_status=$? @@ -758,7 +1028,8 @@ esac # refused or failed merge above, and a queued forge merge exits without an # outcome while its existing poll remains armed. outcome_rc=0 -fm_merge_outcome_report "$FM_HOME" "$STATE" "$ID" "$URL" self || outcome_rc=$? +fm_merge_outcome_report "$FM_HOME" "$STATE" "$ID" "$URL" self \ + "${FM_PR_MERGE_AUTHORITY:-}" || outcome_rc=$? case "$outcome_rc" in 0) ;; 3) diff --git a/docs/architecture.md b/docs/architecture.md index dcb378642f4..43b5272b35b 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -313,12 +313,14 @@ Where a no-mistakes pipeline stores evidence in the repo, it publishes that PR-v This repo uses that setting, and its own `.no-mistakes/` directory remains local state that stays gitignored and is rejected by CI if tracked; [`configuration.md`](configuration.md) owns the setting. PR-based task merges go through `bin/fm-pr-merge.sh`, which records `pr=` and any available `pr_head=` through `bin/fm-pr-check.sh` before calling the forge CLI. The helper requires a full canonical URL and rejects malformed URLs or repo override flags before recording merge state. -A `https://github.com/<owner>/<repo>/pull/<n>` URL invokes `gh-axi pr merge <n> --repo <owner>/<repo>`, defaults to `--squash`, and preserves explicit merge-method flags. +A `https://github.com/<owner>/<repo>/pull/<n>` URL requires `gh` and `jq`, is merged only after one live read confirms the pull request is open, not a draft, mergeable, conflict-free, and every unwaived check is green at the current head, then `gh pr merge` binds that verified head with `--match-head-commit`. +`--auto`, `--admin`, and branch-deletion flags are refused unless `--attended-override` is passed for an explicit captain instruction; that override never skips the live green check, the away-grant check, or a captain hold. +An attended `--allow-red <check-name>` may appear once, waives only GitHub checks with that exact name, and is refused while the away-posture record exists. A `https://<host>/<path>/-/merge_requests/<n>` URL (see [docs/gitlab-merge-watch.md](gitlab-merge-watch.md)) invokes `glab mr merge <n> -R https://<host>/<path>`, so the instance comes from the URL, and adds no merge-method flag because the project's own merge method applies. That path merges only after one live read of the merge request confirms it is open, mergeable, conflict-free, with blocking discussions resolved and a successful pipeline at the current head, and it binds the merge to that verified head; recorded metadata is never the authority for those conditions because a rebase leaves it stale. After either forge command returns, the script confirms the PR or MR actually landed, and only a confirmed landing records a landed outcome; a queued or unconfirmed request records none and leaves its poll armed. On GitLab an auto-merge-queued or unconfirmed request is reported without failing the run. -On GitHub an outcome that is neither merged nor queued is refused loudly and non-zero, naming the observed state, and a base branch that requires the merge queue is refused with the concrete retry flags its configured method requires rather than having a merge method chosen on the caller's behalf. +On GitHub an outcome that is neither merged nor queued is refused loudly and non-zero, naming the observed state, and a base branch that requires the merge queue is refused with the concrete `--attended-override -- --auto --<method>` retry flags its configured method requires rather than having a merge method chosen on the caller's behalf. When the forge already accepted exactly those flags and the pull request still has not entered the queue, that refusal points at the queue state to re-check instead of echoing back the flags the caller just ran. An auto-merge request is held to the same standard: `--auto` that leaves the pull request neither merged nor queued is refused rather than reported as success. Every GitHub refusal states what it could not observe as plainly as what it did, so an unreadable branch-rule response, an unrecognised queue method, and a merge queue no available read can see are each named rather than left to look like a base branch with no queue at all. diff --git a/docs/gitlab-merge-watch.md b/docs/gitlab-merge-watch.md index 0483b0e5557..215d75c0ab9 100644 --- a/docs/gitlab-merge-watch.md +++ b/docs/gitlab-merge-watch.md @@ -239,7 +239,7 @@ $ echo $? A project that runs no pipeline at all therefore cannot merge through this path. That is the intended reading of the requirement rather than an oversight: a successful pipeline at the head is a condition, and "there is no pipeline" does not satisfy it. -Both refusals came after `pr=` was recorded and the merge poll was armed, exactly as a failing `gh-axi pr merge` does on the GitHub side, so a refusal still leaves the audit trail and the watch in place. +Both refusals came after `pr=` was recorded and the merge poll was armed, as a failed live verification or `gh pr merge` does on the GitHub side, so a refusal still leaves the audit trail and the watch in place. A recorded `pr_head=` that no longer matches the live head is reported, and the live head is what gets verified. The stale value below was written into the task record by hand, because a GitLab task never records one on its own: diff --git a/tests/fm-afk-contract.test.sh b/tests/fm-afk-contract.test.sh index 1ad4a6cd05b..8a6ba9fb580 100755 --- a/tests/fm-afk-contract.test.sh +++ b/tests/fm-afk-contract.test.sh @@ -525,6 +525,98 @@ test_inputs_are_validated() { pass "malformed inputs and foreign record versions are refused rather than guessed" } +test_merge_grants_round_trip_and_read_back() { + local home out + home=$(make_home grants-roundtrip) + out=$(contract "$home" propose --grant task-x1 --grant task-y2 --words 'merge those two when green') || fail "grant proposal failed: $out" + assert_contains "$out" 'merge when green (task ids): task-x1, task-y2' 'read-back did not list the granted ids' + [ "$(contract "$home" grants --proposal)" = "$(printf 'task-x1\ntask-y2')" ] \ + || fail "proposal grants subcommand: $(contract "$home" grants --proposal)" + contract "$home" confirm >/dev/null || fail "grant confirm failed" + [ "$(contract "$home" grants)" = "$(printf 'task-x1\ntask-y2')" ] \ + || fail "confirmed grants subcommand: $(contract "$home" grants)" + grep -q '^merge_grants:$' "$home/state/.afk-contract" || fail "confirmed record lacks merge_grants list" + grep -q ' - task-x1' "$home/state/.afk-contract" || fail "confirmed record dropped task-x1" + pass "merge grants round-trip through propose, confirm, read-back, and grants" +} + +test_merge_grants_empty_form_and_usage_errors() { + local home out rc + home=$(make_home grants-empty) + contract "$home" propose >/dev/null || fail "empty grant proposal failed" + grep -qxF 'merge_grants: -' "$home/state/.afk-contract.proposed" \ + || fail "empty grants did not write merge_grants: -" + [ -z "$(contract "$home" grants --proposal)" ] || fail "empty grants subcommand was not empty" + set +e + out=$(contract "$home" propose --grant 'bad id' 2>&1) + rc=$? + set -e + [ "$rc" -eq 2 ] || fail "invalid grant id should be usage error (rc=$rc): $out" + set +e + out=$(contract "$home" propose --grant task-x1 --grant task-x1 2>&1) + rc=$? + set -e + [ "$rc" -eq 2 ] || fail "duplicate grant id should be usage error (rc=$rc): $out" + pass "empty grants write the scalar form, and invalid or duplicate ids are usage errors" +} + +test_legacy_record_without_merge_grants_reads_empty() { + local home record + home=$(make_home grants-legacy) + contract "$home" propose >/dev/null || fail "legacy proposal failed" + contract "$home" confirm >/dev/null || fail "legacy confirm failed" + record="$home/state/.afk-contract" + awk '!/^merge_grants/' "$record" > "$home/legacy" || fail "could not strip merge_grants" + mv "$home/legacy" "$record" + contract "$home" validate >/dev/null || fail "a pre-field v1 record must still validate" + [ -z "$(contract "$home" grants)" ] || fail "a missing merge_grants field must read as an empty list" + pass "a pre-field v1 record reads as empty grants rather than skipping the field" +} + +test_malformed_merge_grants_refuse_validation() { + local home record out rc + home=$(make_home grants-malformed-scalar) + contract "$home" propose >/dev/null || fail "malformed scalar proposal failed" + contract "$home" confirm >/dev/null || fail "malformed scalar confirm failed" + record="$home/state/.afk-contract" + awk '{ print; if ($0 == "merge_grants: -") print " - task-x1" }' "$record" > "$home/malformed" + mv "$home/malformed" "$record" + set +e + out=$(contract "$home" validate 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "indented data attached to scalar merge_grants validated" + assert_contains "$out" 'invalid merge_grants field' 'attached scalar data refusal wording' + + home=$(make_home grants-malformed-duplicate) + contract "$home" propose --grant task-x1 >/dev/null || fail "duplicate field proposal failed" + contract "$home" confirm >/dev/null || fail "duplicate field confirm failed" + record="$home/state/.afk-contract" + printf 'merge_grants: -\n' >> "$record" + set +e + out=$(contract "$home" validate 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "duplicate merge_grants fields validated" + assert_contains "$out" 'invalid merge_grants field' 'duplicate field refusal wording' + pass "malformed and duplicate merge-grant fields fail record validation" +} + +test_archive_drops_live_grants() { + local home rc + home=$(make_home grants-archive) + contract "$home" propose --grant task-x1 >/dev/null || fail "archive grant proposal failed" + contract "$home" confirm >/dev/null || fail "archive grant confirm failed" + contract "$home" archive >/dev/null || fail "archive failed" + [ ! -f "$home/state/.afk-contract" ] || fail "archive left the live record" + set +e + contract "$home" grants >/dev/null 2>&1 + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "grants on the live path succeeded after archive" + pass "archive removes live grants so archived copies are not consulted" +} + test_fields_refuse_each_missing_part_by_name test_omitted_stop_confirms_as_no_stop test_never_set_flags_without_refusing_and_never_over_matches @@ -544,3 +636,9 @@ test_validation_rejects_blank_stop_and_refused_text test_validation_rejects_damaged_words_blocks test_archive_moves_the_record_aside_and_is_idempotent test_inputs_are_validated +test_merge_grants_round_trip_and_read_back +test_merge_grants_empty_form_and_usage_errors +test_legacy_record_without_merge_grants_reads_empty +test_malformed_merge_grants_refuse_validation +test_archive_drops_live_grants + diff --git a/tests/fm-captain-hold-lifecycle.test.sh b/tests/fm-captain-hold-lifecycle.test.sh index fcc3b923e1d..dae536d8684 100755 --- a/tests/fm-captain-hold-lifecycle.test.sh +++ b/tests/fm-captain-hold-lifecycle.test.sh @@ -103,7 +103,15 @@ configure_merged_github() { # <home> #!/usr/bin/env bash printf '%s\n' "$*" >> "$FM_TEST_GH_LOG" case "${1:-} ${2:-}" in - "pr view") printf '%s\n' 1111111111111111111111111111111111111111 ;; + "pr view") + case " $* " in + *statusCheckRollup*) + printf '%s\n' '{"state":"OPEN","isDraft":false,"mergeable":"MERGEABLE","mergeStateStatus":"CLEAN","headRefOid":"1111111111111111111111111111111111111111","statusCheckRollup":[{"__typename":"CheckRun","name":"ci","status":"COMPLETED","conclusion":"SUCCESS"}]}' + ;; + *headRefOid*) printf '%s\n' 1111111111111111111111111111111111111111 ;; + esac + ;; + "pr merge") printf 'merged:\n number: %s\n status: ok\n' "${3:-}" ;; "api graphql") printf '%s\n' 'state=MERGED' 'merged=true' 'queued=false' 'base=main' ;; @@ -113,7 +121,6 @@ SH #!/usr/bin/env bash printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" case "${1:-} ${2:-}" in - "pr merge") printf 'merged:\n number: %s\n status: ok\n' "${3:-}" ;; "pr view") printf 'pull_request:\n number: %s\n state: merged\n' "${3:-}" ;; esac SH @@ -3197,14 +3204,14 @@ test_pr_merge_entrypoint_refuses_a_captain_held_task() { run_captain "$home" hold "$pr_id" --reason "captain merge approval pending" >/dev/null \ || fail "could not hold the PR entrypoint fixture" - # Without the entrypoint guard, this run reaches gh-axi and returns success - # even though the task is still held for the captain. + # Without the entrypoint guard, this run reaches gh and returns success even + # though the task is still held for the captain. set +e run_pr_merge "$home" "$pr_id" "$pr" > "$home/pr.out" 2> "$home/pr.err" rc=$? set -e [ "$rc" -ne 0 ] || fail "the PR merge entrypoint accepted a still-held task" - assert_no_grep 'pr merge 31 ' "$home/gh-axi.log" \ + assert_no_grep 'pr merge 31 ' "$home/gh.log" \ "the PR merge entrypoint reached the irreversible forge call for a held task" assert_grep "$pr_id is still held for the captain" "$home/pr.err" \ "the PR merge refusal did not name the held task" @@ -3272,7 +3279,7 @@ test_pr_merge_entrypoint_separates_an_unreadable_record_from_an_absent_one() { [ "$rc" -ne 0 ] || fail "the PR merge entrypoint accepted an unreadable captain-hold authority record" assert_grep "could not determine whether task $id is still held for the captain" "$home/missing-pr.err" \ "the PR merge refusal did not name its unreadable authority record" - assert_no_grep 'pr merge 43 ' "$home/gh-axi.log" \ + assert_no_grep 'pr merge 43 ' "$home/gh.log" \ "the PR merge entrypoint reached the forge without a readable authority record" # A home with no backlog at all records no captain calls, so nothing can be @@ -3280,7 +3287,7 @@ test_pr_merge_entrypoint_separates_an_unreadable_record_from_an_absent_one() { rm "$home/data/backlog.md" run_pr_merge "$home" "$id" "$pr" > "$home/absent-pr.out" 2> "$home/absent-pr.err" \ || fail "the PR merge entrypoint refused a home carrying no backlog" - merge_count=$(grep -c 'pr merge 43 ' "$home/gh-axi.log" || true) + merge_count=$(grep -c 'pr merge 43 ' "$home/gh.log" || true) [ "$merge_count" -eq 1 ] || fail "the absent backlog did not permit exactly one PR merge" pass "the PR merge entrypoint separates an unreadable authority record from an absent one" } @@ -3476,7 +3483,7 @@ test_merge_entrypoints_refuse_a_reused_task_incarnation() { # Without the pre-wait generation capture and locked comparison, the waiter # records and merges pull request 42 against the replacement task record. [ "$merge_rc" -ne 0 ] || fail "the PR merge accepted a replacement task incarnation" - assert_no_grep 'pr merge 42 ' "$home/gh-axi.log" \ + assert_no_grep 'pr merge 42 ' "$home/gh.log" \ "the PR merge reached the forge for a replacement task incarnation" assert_grep "changed incarnation while waiting to merge" "$home/reuse-merge.err" \ "the PR merge did not identify the replacement task incarnation" @@ -3666,7 +3673,7 @@ SH "PR cleanup was not refused by the merge's task control lock" [ "$merge_rc" -eq 0 ] || fail "the serialized PR merge failed after cleanup was refused" assert_present "$home/state/$id.meta" "the refused PR cleanup removed task metadata" - assert_grep 'pr merge 33 ' "$home/gh-axi.log" \ + assert_grep 'pr merge 33 ' "$home/gh.log" \ "the serialized PR merge did not reach the forge after cleanup was refused" local_home=$(make_home teardown-race-local-entrypoint) diff --git a/tests/fm-pr-check-security.test.sh b/tests/fm-pr-check-security.test.sh index 07d7a3c79e9..240f26f40ba 100755 --- a/tests/fm-pr-check-security.test.sh +++ b/tests/fm-pr-check-security.test.sh @@ -144,6 +144,14 @@ case "${1:-} ${2:-}" in 'base=main' exit 0 ;; + "pr view") + case " $* " in + *statusCheckRollup*) + printf '%s\n' "{\"state\":\"OPEN\",\"isDraft\":false,\"mergeable\":\"MERGEABLE\",\"mergeStateStatus\":\"CLEAN\",\"headRefOid\":\"${FM_TEST_GH_HEAD:-0123456789abcdef0123456789abcdef01234567}\",\"statusCheckRollup\":[{\"__typename\":\"CheckRun\",\"name\":\"ci\",\"status\":\"COMPLETED\",\"conclusion\":\"SUCCESS\"}]}" + exit 0 + ;; + esac + ;; esac case " $* " in *" headRefOid "*) printf '%s\n' "${FM_TEST_GH_HEAD:-0123456789abcdef0123456789abcdef01234567}" ;; @@ -515,11 +523,11 @@ test_valid_recording_and_merge_derivation() { count=$(grep -c '^pr_head=' "$dir/home/state/task-a.meta") [ "$count" -eq 1 ] || fail "duplicate pr_head metadata was appended" - : > "$dir/gh-axi.log" + : > "$dir/gh.log" run_merge_entry "$dir" task-a https://github.com/my-org/repo_name.with-dots/pull/37 -- --merge \ >/dev/null 2>/dev/null || fail "valid merge wrapper failed" - grep -qxF 'pr merge 37 --repo my-org/repo_name.with-dots --merge' "$dir/gh-axi.log" \ - || fail "merge wrapper did not preserve repository derivation and method" + grep -qxF "pr merge 37 --repo my-org/repo_name.with-dots --match-head-commit $expected --merge" "$dir/gh.log" \ + || fail "merge wrapper did not preserve repository derivation, live head, and method" # A merge this home performed leaves its own durable outcome, so the poll's # confirmation is no longer the first the captain hears of it. Acknowledge that # record before the watcher cycle below, which is what still retires the poll. diff --git a/tests/fm-pr-merge.test.sh b/tests/fm-pr-merge.test.sh index b1e840d92fa..7d29879a6d0 100755 --- a/tests/fm-pr-merge.test.sh +++ b/tests/fm-pr-merge.test.sh @@ -5,69 +5,8 @@ # repos with no PR CI where the usual "checks green" fm-pr-check.sh trigger # never fires. # -# Matrix: -# (a) a verified merge records pr= and pr_head= -# (b) merge is refused when gh-axi pr merge itself fails (no silent success) -# (c) extra gh-axi pr merge args are forwarded after number and --repo -# (d) merge is refused before gh-axi when task meta is missing -# (e) PR URL is parsed to number + --repo for gh-axi (defaults to --squash) -# (f) malformed PR URL fails fast without calling gh-axi -# (g) explicit merge method is not overridden by the default --squash -# (h) repo override args fail fast because the repo comes from the URL, -# including a bundled short-option cluster that carries -R -# (i) a GitLab MR URL resolves and merges through glab instead of erroring -# (j) glab is addressed by the host from the URL, never an assumed one -# (k) no merge method is imposed on GitLab, so the project's own one applies -# (l) each pre-merge condition refuses independently, and all of them report -# (m) a stale recorded pr_head= is reported and the live head is verified -# (n) an unreadable merge request state refuses rather than merging blind -# (o) glab or jq absent refuses before any state is recorded -# (p) --sha in extra GitLab args fails fast, and still forwards on GitHub -# (q) a GitLab refusal still leaves pr= recorded and the merge poll armed -# (r) GitHub success is accepted only after the PR is read back as merged -# (s) an open GitHub PR that is neither merged nor queued fails verification -# (t) a GitHub PR in the merge queue is reported as queued, not merged -# (u) a queue-required refusal names the exact compatible retry flags -# (v) a failed poll setup cannot be reported as a verified GitHub merge -# (w) a zero-exit queue-required refusal keeps merge semantics unchanged -# (x) an unreadable outcome after a successful merge call keeps the PR -# recorded and the merge poll armed -# (y) agreeing queue rules still produce exact retry flags -# (z) conflicting queue rules report ambiguous retry guidance -# (aa) gh-axi remains usable when gh is absent -# (ab) a landed merge whose fallback outcome read fails keeps its poll armed -# (ac) a successful merge in a secondmate home reports the landed PR upward -# once, on the route its parent binding names, and a repeat merge of the -# same PR does not duplicate that line -# (ad) a refused or failed merge reports nothing -# (ae) a successful merge in a main home leaves a durable wake naming the PR -# (af) a secondmate home with no usable parent binding says so loudly instead -# of merging in silence -# (ag) an accepted queued GitHub merge emits nothing and leaves its poll armed -# (ah) an accepted queued GitLab merge emits nothing and leaves its poll armed -# (ai) an uncommitted marker retry never loses the durable outcome -# (aj) distinct merged PRs for a reused task each survive queue deduplication -# (ak) pr= is already recorded when the forge call that can land the merge runs -# (al) a failed gh read falls back to the gh-axi view, which can prove a merge -# (am) a failed merge command still names an outcome read that proves a landed -# or queued pull request, without masking the forge failure -# (an) a refusal after a zero-exit merge quotes the forge's own output, marked -# apart from the wrapper's verdict and never leaked to stdout -# (ao) a caller-requested auto-merge on a queue-less base refuses and says -# auto-merge is armed with nothing merged or queued yet -# (ap) a caller-requested auto-merge whose merge command failed refuses -# without ever claiming auto-merge was armed -# (aq) an outcome read that fails after a zero-exit merge still quotes the -# forge's own output, the only evidence left -# (ar) auto-merge with the queue's own method that is still unqueued refuses -# without echoing back the flags just used, and names the next step -# (as) a caller method the queue does not use still gets exact retry flags -# (at) an unrecognised queue method still names the queue requirement and -# guesses no method -# (au) unreadable branch rules are reported apart from a queue-less base -# (av) a base branch with no queue rule says nothing about a merge queue -# (aw) a refusal built on the gh-axi view says the merge queue could not be -# observed, and judges that view's state like the queue-aware one +# The test_* functions below name the covered merge, refusal, live-head, +# away-authority, outcome-publication, and recovery behavior directly. set -u # shellcheck source=tests/lib.sh @@ -90,8 +29,8 @@ MR_STALE_HEAD=bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb JQ_BIN=$(command -v jq) || fail "these tests read glab's JSON with the real jq, which was not found" REAL_MV=$(command -v mv) || fail "these tests need mv to simulate a failed poll publish" -# Build a fresh sandbox for one test case: a state dir with a task meta and a -# fakebin with a gh-axi mock that records how it was invoked. Echoes the case dir. +# Build a fresh sandbox for one test case: a state dir with task metadata and a +# directory for its forge-command mocks. Echoes the case directory. make_case() { local name=$1 case_dir fakebin case_dir="$TMP_ROOT/$name" @@ -119,15 +58,42 @@ make_case() { printf '%s\n' "$case_dir" } -# gh-axi mock recording every invocation to a log file, and gh mock answering -# headRefOid for fm-pr-check.sh's pr_head lookup. Args: case_dir head_sha +# Live GitHub JSON for the pre-merge verify, plus gh-axi for the +# post-merge fallback view. Merge itself is `gh pr merge --match-head-commit`. +# Args: case_dir head_sha +write_github_live_json() { + local case_dir=$1 head=$2 + printf '%s\n' "$head" > "$case_dir/github-head" + cat > "$case_dir/github-view.json" <<JSON +{"state":"OPEN","isDraft":false,"mergeable":"MERGEABLE","mergeStateStatus":"CLEAN","headRefOid":"$head","statusCheckRollup":[{"__typename":"CheckRun","name":"ci","status":"COMPLETED","conclusion":"SUCCESS"}]} +JSON +} + +write_github_red_json() { + local case_dir=$1 head=$2 name=$3 + printf '%s\n' "$head" > "$case_dir/github-head" + cat > "$case_dir/github-view.json" <<JSON +{"state":"OPEN","isDraft":false,"mergeable":"MERGEABLE","mergeStateStatus":"CLEAN","headRefOid":"$head","statusCheckRollup":[{"__typename":"CheckRun","name":"$name","status":"COMPLETED","conclusion":"FAILURE"}]} +JSON +} + +assert_logged_gh_merge() { + local case_dir=$1 number=$2 repo=$3 head line extra= + shift 3 + head=$(cat "$case_dir/github-head") + [ "$#" -eq 0 ] || extra=" $*" + line="pr merge $number --repo $repo --match-head-commit $head$extra" + grep -qxF "$line" "$case_dir/gh.log" \ + || fail "expected gh merge line: $line"$'\n'"got: $(grep '^pr merge ' "$case_dir/gh.log" || true)" +} + add_gh_mocks() { local case_dir=$1 head=$2 + write_github_live_json "$case_dir" "$head" cat > "$case_dir/fakebin/gh-axi" <<'SH' #!/usr/bin/env bash printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" case "${1:-} ${2:-}" in - "pr merge") printf 'merged:\n number: %s\n status: ok\n' "${3:-}" ;; "pr view") [ "$#" -eq 5 ] && [ "${4:-}" = --repo ] || exit 2 printf 'pull_request:\n number: %s\n state: %s\n' "$3" "${FM_TEST_GH_MERGE_STATE:-merged}" @@ -135,21 +101,53 @@ case "${1:-} ${2:-}" in esac exit 0 SH - cat > "$case_dir/fakebin/gh" <<SH + cat > "$case_dir/fakebin/gh" <<'SH' #!/usr/bin/env bash -printf '%s\n' "\$*" >> "\$FM_TEST_GH_LOG" -case "\${1:-} \${2:-}" in +printf '%s\n' "$*" >> "$FM_TEST_GH_LOG" +case "${1:-} ${2:-}" in "pr view") - case " \$* " in - *headRefOid*) printf '%s\n' '$head' ; exit 0 ;; + case " $* " in + *statusCheckRollup*) + cat "$FM_TEST_GH_VIEW_JSON" + if [ -f "${FM_TEST_AWAY_RECORD_AFTER_VIEW:-}" ]; then + cp "$FM_TEST_AWAY_RECORD_AFTER_VIEW" "$FM_STATE_OVERRIDE/.afk-contract" + fi + exit 0 + ;; + *headRefOid*) + cat "$FM_TEST_GH_HEAD" + exit 0 + ;; esac ;; + "pr merge") + if [ -n "${FM_TEST_META_AT_MERGE:-}" ] && [ -f "${FM_STATE_OVERRIDE:-}/task-x1.meta" ]; then + cat "$FM_STATE_OVERRIDE/task-x1.meta" > "$FM_TEST_META_AT_MERGE" + fi + if [ -n "${FM_TEST_GH_MERGE_OUTPUT:-}" ]; then + printf '%s\n' "$FM_TEST_GH_MERGE_OUTPUT" + else + printf 'merged:\n number: %s\n status: ok\n' "${3:-}" + fi + merge_rc=0 + if [ -f "${FM_TEST_GH_MERGE_RC_FILE:-}" ]; then + merge_rc=$(cat "$FM_TEST_GH_MERGE_RC_FILE") + fi + exit "$merge_rc" + ;; "api graphql") - cat "\$FM_TEST_GH_OUTCOME" + if [ -f "${FM_TEST_GH_GRAPHQL_FAIL:-}" ]; then + echo 'error: could not reach the GitHub API' >&2 + exit 1 + fi + cat "$FM_TEST_GH_OUTCOME" exit 0 ;; api\ *) - cat "\$FM_TEST_GH_RULES" + if [ -f "${FM_TEST_GH_RULES_FAIL:-}" ]; then + exit 1 + fi + cat "$FM_TEST_GH_RULES" exit 0 ;; esac @@ -158,58 +156,21 @@ SH chmod +x "$case_dir/fakebin/gh-axi" "$case_dir/fakebin/gh" } -# gh-axi mock that fails the merge call but succeeds everything else, so a -# real merge failure is distinguishable from the recording step. +# gh mock that fails the merge call but succeeds live verify, so a real merge +# failure is distinguishable from the recording step. add_gh_mocks_merge_fails() { local case_dir=$1 - cat > "$case_dir/fakebin/gh-axi" <<'SH' -#!/usr/bin/env bash -printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" -case "${1:-} ${2:-}" in - "pr merge") echo "error: pr merge failed" >&2 ; exit 1 ;; - esac - exit 0 -SH - cat > "$case_dir/fakebin/gh" <<'SH' -#!/usr/bin/env bash -printf '%s\n' "$*" >> "$FM_TEST_GH_LOG" -case "${1:-} ${2:-}" in - "api graphql") - cat "$FM_TEST_GH_OUTCOME" - exit 0 - ;; - api\ *) - cat "$FM_TEST_GH_RULES" - exit 0 - ;; -esac -exit 0 -SH - chmod +x "$case_dir/fakebin/gh-axi" "$case_dir/fakebin/gh" + local head=${2:-bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb} + add_gh_mocks "$case_dir" "$head" + printf '1\n' > "$case_dir/github-merge-rc" + printf 'error: pr merge failed\n' > "$case_dir/github-merge-output" } -# gh mock that still answers fm-pr-check.sh's head lookup but cannot answer the -# outcome read, so a merge call that returned success is followed by a live -# state nothing can prove. Args: case_dir head_sha +# Flag the shared gh mock so GraphQL outcome reads fail while live verify and +# merge still succeed. Args: case_dir [head_sha ignored] add_gh_mock_outcome_read_fails() { - local case_dir=$1 head=$2 - cat > "$case_dir/fakebin/gh" <<SH -#!/usr/bin/env bash -printf '%s\n' "\$*" >> "\$FM_TEST_GH_LOG" -case "\${1:-} \${2:-}" in - "pr view") - case " \$* " in - *headRefOid*) printf '%s\n' '$head' ; exit 0 ;; - esac - ;; - "api graphql") - echo 'error: could not reach the GitHub API' >&2 - exit 1 - ;; -esac -exit 0 -SH - chmod +x "$case_dir/fakebin/gh" + local case_dir=$1 + : > "$case_dir/github-graphql-fail" } # gh-axi mock that merges but cannot answer its own view, so a case can prove @@ -363,7 +324,14 @@ run_pr_merge() { FM_TEST_GH_LOG="$case_dir/gh.log" \ FM_TEST_GH_OUTCOME="$case_dir/github-outcome" \ FM_TEST_GH_RULES="$case_dir/github-rules" \ + FM_TEST_GH_VIEW_JSON="$case_dir/github-view.json" \ + FM_TEST_GH_HEAD="$case_dir/github-head" \ + FM_TEST_GH_MERGE_RC_FILE="$case_dir/github-merge-rc" \ + FM_TEST_GH_MERGE_OUTPUT="$(cat "$case_dir/github-merge-output" 2>/dev/null || true)" \ + FM_TEST_GH_GRAPHQL_FAIL="$case_dir/github-graphql-fail" \ + FM_TEST_GH_RULES_FAIL="$case_dir/github-rules-fail" \ FM_TEST_META_AT_MERGE="$case_dir/meta-at-merge" \ + FM_TEST_AWAY_RECORD_AFTER_VIEW="$case_dir/away-record-after-view" \ FM_TEST_REAL_MV="$REAL_MV" \ FM_TEST_GLAB_LOG="$case_dir/glab.log" \ FM_TEST_GLAB_JSON="$case_dir/mr.json" \ @@ -387,6 +355,15 @@ write_github_outcome() { "base=$base" > "$case_dir/github-outcome" } +write_away_record() { + local case_dir=$1 + shift + FM_HOME="$case_dir/home" FM_STATE_OVERRIDE="$case_dir/state" \ + "$ROOT/bin/fm-afk-contract.sh" propose "$@" >/dev/null + FM_HOME="$case_dir/home" FM_STATE_OVERRIDE="$case_dir/state" \ + "$ROOT/bin/fm-afk-contract.sh" confirm >/dev/null +} + test_verified_merge_records_pr_and_head() { local case_dir rc case_dir=$(make_case records-before-merge) @@ -405,8 +382,7 @@ test_verified_merge_records_pr_and_head() { "records-before-merge: pr= was not recorded" assert_grep 'pr_head=deadbeefcafefeed0000000000000000deadbeef' "$case_dir/state/task-x1.meta" \ "records-before-merge: pr_head= was not recorded" - grep -qxF 'pr merge 9 --repo example/repo --squash' "$case_dir/gh-axi.log" \ - || fail "records-before-merge: gh-axi pr merge was not invoked with number, --repo, and default --squash" + assert_logged_gh_merge "$case_dir" 9 example/repo --squash pass "fm-pr-merge records pr= and pr_head= for a verified GitHub merge" } @@ -418,21 +394,6 @@ test_pr_metadata_is_recorded_before_the_forge_call() { case_dir=$(make_case records-ahead-of-forge-call) mkdir -p "$case_dir/wt" add_gh_mocks "$case_dir" 5151515151515151515151515151515151515151 - cat > "$case_dir/fakebin/gh-axi" <<'SH' -#!/usr/bin/env bash -printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" -case "${1:-} ${2:-}" in - "pr merge") - cat "$FM_STATE_OVERRIDE/task-x1.meta" > "$FM_TEST_META_AT_MERGE" - printf 'merged:\n number: %s\n status: ok\n' "${3:-}" - ;; - "pr view") - printf 'pull_request:\n number: %s\n state: merged\n' "$3" - ;; -esac -exit 0 -SH - chmod +x "$case_dir/fakebin/gh-axi" : > "$case_dir/gh-axi.log" : > "$case_dir/meta-at-merge" @@ -443,8 +404,7 @@ SH set -e expect_code 0 "$rc" "records-ahead-of-forge-call: fm-pr-merge should succeed" - assert_grep 'pr merge 62 --repo example/repo --squash' "$case_dir/gh-axi.log" \ - "records-ahead-of-forge-call: the merge abstraction was never invoked" + assert_logged_gh_merge "$case_dir" 62 example/repo --squash assert_grep 'pr=https://github.com/example/repo/pull/62' "$case_dir/meta-at-merge" \ "records-ahead-of-forge-call: the merge ran before pr= was recorded" pass "fm-pr-merge records pr= before the forge call can land the merge" @@ -578,15 +538,8 @@ test_github_refusal_quotes_the_forge_output() { case_dir=$(make_case github-refusal-quotes-forge) mkdir -p "$case_dir/wt" add_gh_mocks "$case_dir" 6161616161616161616161616161616161616161 - cat > "$case_dir/fakebin/gh-axi" <<'SH' -#!/usr/bin/env bash -printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" -case "${1:-} ${2:-}" in - "pr merge") echo "will be added to the merge queue when all requirements are met" ;; -esac -exit 0 -SH - chmod +x "$case_dir/fakebin/gh-axi" + printf '%s\n' 'will be added to the merge queue when all requirements are met' \ + > "$case_dir/github-merge-output" write_github_outcome "$case_dir" OPEN false false main : > "$case_dir/gh-axi.log" : > "$case_dir/gh.log" @@ -630,7 +583,7 @@ test_github_auto_merge_without_queue_refuses_legibly() { set +e run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/66 \ - -- "$spelling" --merge \ + --attended-override -- "$spelling" --merge \ > "$case_dir/stdout" 2> "$case_dir/stderr" rc=$? set -e @@ -642,9 +595,8 @@ test_github_auto_merge_without_queue_refuses_legibly() { "$case_dir/stderr" "github-auto-no-queue: the refusal never explained the armed auto-merge" assert_grep 'nothing is merged or in the merge queue yet' "$case_dir/stderr" \ "github-auto-no-queue: the refusal left the operator to infer the pending state" - grep -qxF "pr merge 66 --repo example/repo $spelling --merge" "$case_dir/gh-axi.log" \ - || fail "github-auto-no-queue: the attempted merge was changed unexpectedly" - [ "$(wc -l < "$case_dir/gh-axi.log" | tr -d '[:space:]')" = 1 ] \ + assert_logged_gh_merge "$case_dir" 66 example/repo "$spelling" --merge + [ "$(grep -c '^pr merge ' "$case_dir/gh.log")" -eq 1 ] \ || fail "github-auto-no-queue: the wrapper attempted more than one merge" assert_grep 'pr=https://github.com/example/repo/pull/66' "$case_dir/state/task-x1.meta" \ "github-auto-no-queue: the attempted merge lost its PR reference" @@ -665,7 +617,7 @@ test_github_failed_merge_never_claims_armed_auto_merge() { : > "$case_dir/gh.log" set +e - run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/67 -- --auto --merge \ + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/67 --attended-override -- --auto --merge \ > "$case_dir/stdout" 2> "$case_dir/stderr" rc=$? set -e @@ -696,7 +648,7 @@ test_github_failed_merge_with_queue_flags_never_claims_acceptance() { : > "$case_dir/gh.log" set +e - run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/74 -- --auto --merge \ + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/74 --attended-override -- --auto --merge \ > "$case_dir/stdout" 2> "$case_dir/stderr" rc=$? set -e @@ -712,7 +664,7 @@ test_github_failed_merge_with_queue_flags_never_claims_acceptance() { "github-failed-merge-queue-flags: a failed merge command was reported as an armed auto-merge" assert_grep 'base branch main requires the merge queue; retry with:' "$case_dir/stderr" \ "github-failed-merge-queue-flags: the failed merge command lost its concrete retry guidance" - assert_grep 'task-x1 https://github.com/example/repo/pull/74 -- --auto --merge' "$case_dir/stderr" \ + assert_grep 'task-x1 https://github.com/example/repo/pull/74 --attended-override -- --auto --merge' "$case_dir/stderr" \ "github-failed-merge-queue-flags: the retry guidance named no queue flags" assert_no_grep 'verified: ' "$case_dir/stdout" \ "github-failed-merge-queue-flags: a failed merge command was reported as verified" @@ -730,7 +682,7 @@ test_github_accepted_queue_flags_do_not_echo_back_the_same_command() { : > "$case_dir/gh.log" set +e - run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/68 -- --auto --merge \ + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/68 --attended-override -- --auto --merge \ > "$case_dir/stdout" 2> "$case_dir/stderr" rc=$? set -e @@ -761,7 +713,7 @@ test_github_mismatched_queue_flags_still_name_the_retry() { : > "$case_dir/gh.log" set +e - run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/69 -- --auto --merge \ + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/69 --attended-override -- --auto --merge \ > "$case_dir/stdout" 2> "$case_dir/stderr" rc=$? set -e @@ -769,7 +721,7 @@ test_github_mismatched_queue_flags_still_name_the_retry() { expect_code 1 "$rc" "github-mismatched-queue-flags: an unproved merge must still fail" assert_grep 'base branch main requires the merge queue; retry with:' "$case_dir/stderr" \ "github-mismatched-queue-flags: a caller method the queue does not use lost its retry guidance" - assert_grep '-- --auto --rebase' "$case_dir/stderr" \ + assert_grep '--attended-override -- --auto --rebase' "$case_dir/stderr" \ "github-mismatched-queue-flags: the exact compatible flags were not named" pass "fm-pr-merge still names retry flags when the caller used a different method" } @@ -807,24 +759,7 @@ test_github_unreadable_queue_rules_are_not_reported_as_no_queue() { mkdir -p "$case_dir/wt" add_gh_mocks "$case_dir" 8484848484848484848484848484848484848484 write_github_outcome "$case_dir" OPEN false false main - cat > "$case_dir/fakebin/gh" <<'SH' -#!/usr/bin/env bash -printf '%s\n' "$*" >> "$FM_TEST_GH_LOG" -case "${1:-} ${2:-}" in - "pr view") - case " $* " in - *headRefOid*) printf '%s\n' 8484848484848484848484848484848484848484 ; exit 0 ;; - esac - ;; - "api graphql") - cat "$FM_TEST_GH_OUTCOME" - exit 0 - ;; - api\ *) exit 1 ;; -esac -exit 0 -SH - chmod +x "$case_dir/fakebin/gh" + : > "$case_dir/github-rules-fail" : > "$case_dir/gh-axi.log" : > "$case_dir/gh.log" @@ -866,49 +801,40 @@ test_github_no_queue_rule_says_nothing_about_a_queue() { pass "fm-pr-merge says nothing about a merge queue when the base branch has no queue rule" } -test_github_fallback_view_refusal_says_the_queue_was_unobservable() { - local case_dir ghless_path rc - case_dir=$(make_case github-fallback-unobservable-queue) +test_github_unmerged_fallback_cannot_replace_queue_aware_read() { + local case_dir rc + case_dir=$(make_case github-unmerged-fallback) mkdir -p "$case_dir/wt" add_gh_mocks "$case_dir" 8686868686868686868686868686868686868686 + add_gh_mock_outcome_read_fails "$case_dir" cat > "$case_dir/fakebin/gh-axi" <<'SH' #!/usr/bin/env bash printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" case "${1:-} ${2:-}" in - "pr merge") printf 'merged:\n number: %s\n status: ok\n' "${3:-}" ;; "pr view") printf 'pull_request:\n number: %s\n state: open\n' "$3" ;; esac exit 0 SH chmod +x "$case_dir/fakebin/gh-axi" - rm "$case_dir/fakebin/gh" - ghless_path="$case_dir/path-without-gh" - mirror_path_without "$ghless_path" gh "$case_dir/fakebin" : > "$case_dir/gh-axi.log" set +e - PATH="$ghless_path" run_pr_merge "$case_dir" task-x1 \ - https://github.com/example/repo/pull/73 -- --auto --merge \ + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/73 \ > "$case_dir/stdout" 2> "$case_dir/stderr" rc=$? set -e - expect_code 1 "$rc" "github-fallback-unobservable-queue: an unproved merge must fail" - assert_grep 'isInMergeQueue=unknown' "$case_dir/stderr" \ - "github-fallback-unobservable-queue: refusal did not name the concrete observed state" - assert_grep 'the merge queue could not be observed for https://github.com/example/repo/pull/73' \ - "$case_dir/stderr" \ - "github-fallback-unobservable-queue: the refusal implied an unqueued PR it could not see" - assert_grep "re-check the pull request's merge queue state" "$case_dir/stderr" \ - "github-fallback-unobservable-queue: the refusal named no concrete next step" - # The lowercase state the fallback view reports must be judged the same way - # the queue-aware read's uppercase enum is, or every explanation is skipped. - assert_grep 'auto-merge was requested and armed for https://github.com/example/repo/pull/73' \ + expect_code 1 "$rc" "github-unmerged-fallback: an unproved merge must fail" + assert_grep 'pr view 73 --repo example/repo' "$case_dir/gh-axi.log" \ + "github-unmerged-fallback: the fallback view was not consulted" + assert_grep 'the gh read failed and the gh-axi view could not prove the outcome either' \ "$case_dir/stderr" \ - "github-fallback-unobservable-queue: the fallback view's state skipped the auto-merge explanation" + "github-unmerged-fallback: an unmerged fallback was treated as a readable outcome" + assert_no_grep 'GitHub merge outcome was not successful' "$case_dir/stderr" \ + "github-unmerged-fallback: an unmerged fallback reached detailed outcome handling" assert_no_grep 'verified: ' "$case_dir/stdout" \ - "github-fallback-unobservable-queue: an unproved merge was reported as verified" - pass "fm-pr-merge says the merge queue was unobservable when only the gh-axi view answered" + "github-unmerged-fallback: an unproved merge was reported as verified" + pass "fm-pr-merge accepts only a proved merge from the gh-axi fallback" } test_github_unreadable_outcome_refusal_quotes_the_forge_output() { @@ -916,16 +842,9 @@ test_github_unreadable_outcome_refusal_quotes_the_forge_output() { case_dir=$(make_case github-unreadable-outcome-quotes-forge) mkdir -p "$case_dir/wt" add_gh_mocks "$case_dir" 8787878787878787878787878787878787878787 - cat > "$case_dir/fakebin/gh-axi" <<'SH' -#!/usr/bin/env bash -printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" -case "${1:-} ${2:-}" in - "pr merge") echo "will be added to the merge queue when all requirements are met" ;; - "pr view") exit 1 ;; -esac -exit 0 -SH - chmod +x "$case_dir/fakebin/gh-axi" + printf '%s\n' 'will be added to the merge queue when all requirements are met' \ + > "$case_dir/github-merge-output" + add_gh_axi_mock_view_fails "$case_dir" add_gh_mock_outcome_read_fails "$case_dir" 8787878787878787878787878787878787878787 : > "$case_dir/gh-axi.log" : > "$case_dir/gh.log" @@ -1022,30 +941,22 @@ test_github_without_gh_still_uses_gh_axi_merge() { rc=$? set -e - expect_code 0 "$rc" "github-without-gh: gh-axi can prove a landed merge without gh" - assert_grep 'pr merge 60 --repo example/repo --squash' "$case_dir/gh-axi.log" \ - "github-without-gh: the configured merge abstraction was not invoked" - assert_grep 'pr view 60 --repo example/repo' "$case_dir/gh-axi.log" \ - "github-without-gh: the gh-axi fallback did not verify the landed state" - assert_grep 'verified: https://github.com/example/repo/pull/60 is merged' \ - "$case_dir/stdout" "github-without-gh: the fallback did not report the proven merge" - pass "fm-pr-merge reaches and verifies the gh-axi merge path without gh" + expect_code 1 "$rc" "github-without-gh: missing gh must refuse before recording" + assert_grep 'merging a GitHub pull request requires gh on PATH' "$case_dir/stderr" \ + "github-without-gh: missing gh was not named" + assert_no_grep 'pr=' "$case_dir/state/task-x1.meta" \ + "github-without-gh: pr= was recorded without gh" + assert_absent "$case_dir/state/task-x1.check.sh" \ + "github-without-gh: a merge poll was armed without gh" + pass "fm-pr-merge refuses a GitHub merge when gh is missing, before recording" } test_github_without_gh_failed_read_keeps_bookkeeping() { local case_dir ghless_path rc case_dir=$(make_case github-without-gh-read-fails) mkdir -p "$case_dir/wt" - cat > "$case_dir/fakebin/gh-axi" <<'SH' -#!/usr/bin/env bash -printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" -case "${1:-} ${2:-}" in - "pr merge") exit 0 ;; - "pr view") exit 1 ;; -esac -exit 0 -SH - chmod +x "$case_dir/fakebin/gh-axi" + add_gh_mocks "$case_dir" 4141414141414141414141414141414141414141 + rm "$case_dir/fakebin/gh" ghless_path="$case_dir/path-without-gh" mirror_path_without "$ghless_path" gh "$case_dir/fakebin" : > "$case_dir/gh-axi.log" @@ -1057,16 +968,14 @@ SH rc=$? set -e - expect_code 1 "$rc" "github-without-gh-read-fails: an unreadable outcome must fail" - assert_grep 'pr merge 61 --repo example/repo --squash' "$case_dir/gh-axi.log" \ - "github-without-gh-read-fails: the merge call did not happen before the failed read" - assert_grep 'could not read the GitHub pull request outcome after the merge attempt' \ - "$case_dir/stderr" "github-without-gh-read-fails: the failed read was not reported" - assert_grep 'pr=https://github.com/example/repo/pull/61' "$case_dir/state/task-x1.meta" \ - "github-without-gh-read-fails: a landed merge lost its PR metadata" - assert_present "$case_dir/state/task-x1.check.sh" \ - "github-without-gh-read-fails: a landed merge lost its merge poll" - pass "fm-pr-merge preserves bookkeeping when gh is absent and the fallback read fails" + expect_code 1 "$rc" "github-without-gh-read-fails: missing gh must refuse before recording" + assert_grep 'merging a GitHub pull request requires gh on PATH' "$case_dir/stderr" \ + "github-without-gh-read-fails: missing gh was not named" + assert_no_grep 'pr=' "$case_dir/state/task-x1.meta" \ + "github-without-gh-read-fails: pr= was recorded without gh" + assert_absent "$case_dir/state/task-x1.check.sh" \ + "github-without-gh-read-fails: a merge poll was armed without gh" + pass "fm-pr-merge refuses a GitHub merge when gh is missing rather than merging blind" } test_github_zero_exit_queue_required_refuses_with_exact_retry() { @@ -1090,15 +999,14 @@ test_github_zero_exit_queue_required_refuses_with_exact_retry() { "github-zero-exit-queue-required: refusal did not name the concrete observed state" assert_grep 'base branch release/2026 requires the merge queue' "$case_dir/stderr" \ "github-zero-exit-queue-required: refusal did not name the queue requirement" - assert_grep '-- --auto --rebase' "$case_dir/stderr" \ + assert_grep '--attended-override -- --auto --rebase' "$case_dir/stderr" \ "github-zero-exit-queue-required: refusal did not name the exact compatible flags" assert_grep 'api --paginate repos/example/repo/rules/branches/release%2F2026' "$case_dir/gh.log" \ "github-zero-exit-queue-required: queue rules were not read with pagination and encoded branch path" - grep -qxF 'pr merge 56 --repo example/repo --squash' "$case_dir/gh-axi.log" \ - || fail "github-zero-exit-queue-required: the attempted merge was changed unexpectedly" - [ "$(wc -l < "$case_dir/gh-axi.log" | tr -d '[:space:]')" = 1 ] \ + assert_logged_gh_merge "$case_dir" 56 example/repo --squash + [ "$(grep -c '^pr merge ' "$case_dir/gh.log")" -eq 1 ] \ || fail "github-zero-exit-queue-required: the wrapper attempted more than one merge" - assert_no_grep --auto "$case_dir/gh-axi.log" \ + assert_no_grep --auto "$case_dir/gh.log" \ "github-zero-exit-queue-required: queue flags were auto-applied to the attempted merge" assert_grep 'pr=https://github.com/example/repo/pull/56' "$case_dir/state/task-x1.meta" \ "github-zero-exit-queue-required: the attempted merge lost its PR reference" @@ -1128,7 +1036,7 @@ test_github_closed_unqueued_outcome_omits_retry_flags() { "github-closed-unqueued: refusal did not name the concrete observed state" assert_no_grep 'requires the merge queue' "$case_dir/stderr" \ "github-closed-unqueued: closed PR received unusable queue guidance" - assert_no_grep '-- --auto --merge' "$case_dir/stderr" \ + assert_no_grep '--attended-override -- --auto --merge' "$case_dir/stderr" \ "github-closed-unqueued: closed PR received retry flags" assert_grep 'pr=https://github.com/example/repo/pull/57' "$case_dir/state/task-x1.meta" \ "github-closed-unqueued: the attempted merge lost its PR reference" @@ -1147,7 +1055,7 @@ test_github_queued_outcome_is_verified() { : > "$case_dir/gh.log" set +e - run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/53 -- --auto --merge \ + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/53 --attended-override -- --auto --merge \ > "$case_dir/stdout" 2> "$case_dir/stderr" rc=$? set -e @@ -1183,10 +1091,9 @@ test_github_queue_required_refusal_names_retry_flags() { "github-queue-required: the original forge failure was not preserved" assert_grep 'base branch master requires the merge queue' "$case_dir/stderr" \ "github-queue-required: refusal did not name the queue requirement" - grep -F -- '-- --auto --merge' "$case_dir/stderr" >/dev/null \ + grep -F -- '--attended-override -- --auto --merge' "$case_dir/stderr" >/dev/null \ || fail "github-queue-required: refusal did not name the exact compatible flags" - grep -qxF 'pr merge 54 --repo example/repo --squash' "$case_dir/gh-axi.log" \ - || fail "github-queue-required: the wrapper silently changed the attempted merge semantics" + assert_logged_gh_merge "$case_dir" 54 example/repo --squash assert_present "$case_dir/state/task-x1.check.sh" \ "github-queue-required: the failed forge call did not leave the merge poll armed" pass "fm-pr-merge explains how to retry with the required GitHub merge queue method" @@ -1211,7 +1118,7 @@ test_github_agreeing_queue_rules_keep_retry_guidance() { expect_code 1 "$rc" "github-agreeing-queue-rules: an unproved merge must fail" assert_grep 'base branch main requires the merge queue' "$case_dir/stderr" \ "github-agreeing-queue-rules: refusal did not name the queue requirement" - assert_grep '-- --auto --rebase' "$case_dir/stderr" \ + assert_grep '--attended-override -- --auto --rebase' "$case_dir/stderr" \ "github-agreeing-queue-rules: agreeing rules omitted exact retry flags" assert_no_grep 'exact retry flags are ambiguous' "$case_dir/stderr" \ "github-agreeing-queue-rules: agreeing rules were reported as ambiguous" @@ -1239,9 +1146,9 @@ test_github_conflicting_queue_rules_report_ambiguity() { assert_grep 'base branch main has conflicting merge queue methods (MERGE, SQUASH)' \ "$case_dir/stderr" \ "github-conflicting-queue-rules: conflicting methods were not named" - assert_no_grep '-- --auto --merge' "$case_dir/stderr" \ + assert_no_grep '--attended-override -- --auto --merge' "$case_dir/stderr" \ "github-conflicting-queue-rules: an exact retry method was guessed" - assert_no_grep '-- --auto --squash' "$case_dir/stderr" \ + assert_no_grep '--attended-override -- --auto --squash' "$case_dir/stderr" \ "github-conflicting-queue-rules: an exact retry method was guessed" assert_no_grep 'SQUASH, SQUASH' "$case_dir/stderr" \ "github-conflicting-queue-rules: a repeated queue method was named twice" @@ -1255,12 +1162,25 @@ test_extra_merge_args_forwarded() { add_gh_mocks "$case_dir" 2222222222222222222222222222222222222222 : > "$case_dir/gh-axi.log" + set +e run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/15 -- --squash --delete-branch \ - > "$case_dir/stdout" 2> "$case_dir/stderr" || fail "extra-args: fm-pr-merge failed" + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "extra-args: branch deletion must be refused without --attended-override" + assert_grep 'pass --attended-override only for an explicit captain instruction' "$case_dir/stderr" \ + "extra-args: refusal did not name --attended-override" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "extra-args: gh pr merge ran despite the denylist" - grep -qxF 'pr merge 15 --repo example/repo --squash --delete-branch' "$case_dir/gh-axi.log" \ - || fail "extra-args: extra gh-axi pr merge flags were not forwarded" - pass "fm-pr-merge forwards extra flags to gh-axi pr merge after the -- separator" + case_dir=$(make_case extra-args-attended) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 2222222222222222222222222222222222222222 + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/15 \ + --attended-override -- --squash --delete-branch \ + > "$case_dir/stdout" 2> "$case_dir/stderr" || fail "extra-args-attended: attended override should merge" + assert_logged_gh_merge "$case_dir" 15 example/repo --squash --delete-branch + pass "fm-pr-merge refuses branch deletion unless --attended-override is passed" } test_missing_meta_refuses_before_merge() { @@ -1280,7 +1200,7 @@ test_missing_meta_refuses_before_merge() { expect_code 1 "$rc" "missing-meta: fm-pr-merge should refuse" assert_grep 'error: task metadata is unavailable' "$case_dir/stderr" \ "missing-meta: refusal did not explain missing meta" - [ ! -s "$case_dir/gh-axi.log" ] || fail "missing-meta: gh-axi pr merge was invoked" + [ ! -s "$case_dir/gh.log" ] || fail "missing-meta: gh pr merge was invoked" assert_absent "$case_dir/state/missing-x1.check.sh" \ "missing-meta: fm-pr-check should not arm a poll for an unknown task" pass "fm-pr-merge refuses before merging when task meta is missing" @@ -1309,7 +1229,7 @@ test_malformed_url_refuses_before_merge() { "malformed-url: malformed PR URL was recorded in meta" assert_absent "$case_dir/state/task-x1.check.sh" \ "malformed-url: malformed PR URL armed a merge poll" - assert_no_grep 'pr merge' "$case_dir/gh-axi.log" \ + assert_no_grep 'pr merge' "$case_dir/gh.log" \ "malformed-url: gh-axi pr merge was invoked for a malformed URL" pass "fm-pr-merge refuses malformed PR URLs before calling gh-axi" } @@ -1336,7 +1256,7 @@ test_rejects_unsafe_url_segments_before_recording() { "unsafe-url-segment: unsafe PR URL was recorded in meta" assert_absent "$case_dir/state/task-x1.check.sh" \ "unsafe-url-segment: unsafe PR URL armed a merge poll" - assert_no_grep 'pr merge' "$case_dir/gh-axi.log" \ + assert_no_grep 'pr merge' "$case_dir/gh.log" \ "unsafe-url-segment: gh-axi pr merge was invoked for an unsafe URL" pass "fm-pr-merge refuses unsafe PR URL segments before recording state" } @@ -1361,7 +1281,7 @@ test_repo_override_args_refuse_before_recording() { "repo-override: PR URL was recorded before rejecting repo override" assert_absent "$case_dir/state/task-x1.check.sh" \ "repo-override: repo override armed a merge poll" - assert_no_grep 'pr merge' "$case_dir/gh-axi.log" \ + assert_no_grep 'pr merge' "$case_dir/gh.log" \ "repo-override: gh-axi pr merge was invoked despite repo override" pass "fm-pr-merge refuses repo override args before recording state" } @@ -1390,7 +1310,7 @@ test_bundled_repo_override_args_refuse_before_recording() { "bundled-repo-override: PR URL was recorded before rejecting the bundled repo override" assert_absent "$case_dir/state/task-x1.check.sh" \ "bundled-repo-override: a bundled repo override armed a merge poll" - assert_no_grep 'pr merge' "$case_dir/gh-axi.log" \ + assert_no_grep 'pr merge' "$case_dir/gh.log" \ "bundled-repo-override: gh-axi pr merge was invoked despite the bundled repo override" case_dir=$(make_gitlab_case bundled-repo-override-gitlab) @@ -1418,13 +1338,23 @@ test_bundled_repo_override_args_refuse_before_recording() { add_gh_mocks "$case_dir" bcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbc : > "$case_dir/gh-axi.log" + set +e run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/8 -- -d \ - > "$case_dir/stdout" 2> "$case_dir/stderr" \ - || fail "bundled-non-repo-cluster: fm-pr-merge refused a short flag that overrides nothing" + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "bundled-non-repo-cluster: -d is branch deletion and must be refused" + assert_grep 'pass --attended-override only for an explicit captain instruction' "$case_dir/stderr" \ + "bundled-non-repo-cluster: refusal did not name --attended-override" - grep -qxF 'pr merge 8 --repo example/repo --squash -d' "$case_dir/gh-axi.log" \ - || fail "bundled-non-repo-cluster: a short flag carrying no repository override was not forwarded" - pass "fm-pr-merge refuses a bundled short-option repo override and forwards other short flags" + case_dir=$(make_case bundled-delete-attended) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" bcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbcbc + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/8 --attended-override -- -d \ + > "$case_dir/stdout" 2> "$case_dir/stderr" \ + || fail "bundled-delete-attended: attended override should merge" + assert_logged_gh_merge "$case_dir" 8 example/repo --squash -d + pass "fm-pr-merge refuses a bundled short-option repo override and refuses -d unless attended" } test_explicit_merge_method_not_overridden() { @@ -1437,8 +1367,7 @@ test_explicit_merge_method_not_overridden() { run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/22 -- --merge \ > "$case_dir/stdout" 2> "$case_dir/stderr" || fail "explicit-merge-method: fm-pr-merge failed" - grep -qxF 'pr merge 22 --repo example/repo --merge' "$case_dir/gh-axi.log" \ - || fail "explicit-merge-method: caller --merge was not forwarded without an extra default --squash" + assert_logged_gh_merge "$case_dir" 22 example/repo --merge pass "fm-pr-merge does not add default --squash when the caller passes an explicit merge method" } @@ -1452,8 +1381,7 @@ test_method_equals_merge_method_not_overridden() { run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/23 -- --method=merge \ > "$case_dir/stdout" 2> "$case_dir/stderr" || fail "method-equals-merge-method: fm-pr-merge failed" - grep -qxF 'pr merge 23 --repo example/repo --method=merge' "$case_dir/gh-axi.log" \ - || fail "method-equals-merge-method: caller --method=merge was not forwarded without an extra default --squash" + assert_logged_gh_merge "$case_dir" 23 example/repo --method=merge pass "fm-pr-merge respects --method=<value> as an explicit merge method" } @@ -1467,8 +1395,7 @@ test_parses_pr_url_for_gh_axi() { run_pr_merge "$case_dir" task-x1 https://github.com/my-org/my-repo/pull/126 \ > "$case_dir/stdout" 2> "$case_dir/stderr" || fail "url-parsing: fm-pr-merge failed" - grep -qxF 'pr merge 126 --repo my-org/my-repo --squash' "$case_dir/gh-axi.log" \ - || fail "url-parsing: gh-axi pr merge was not invoked as number + --repo + default --squash" + assert_logged_gh_merge "$case_dir" 126 my-org/my-repo --squash pass "fm-pr-merge parses a GitHub PR URL into gh-axi number and --repo arguments" } @@ -1551,12 +1478,22 @@ test_gitlab_extra_args_forwarded() { > "$case_dir/stdout" 2> "$case_dir/stderr" rc=$? set -e + expect_code 1 "$rc" "gitlab-extra-args: source-branch deletion must be refused without --attended-override" + assert_grep 'pass --attended-override only for an explicit captain instruction' "$case_dir/stderr" \ + "gitlab-extra-args: refusal did not name --attended-override" + [ ! -s "$case_dir/glab.log" ] || fail "gitlab-extra-args: glab ran despite the denylist" - expect_code 0 "$rc" "gitlab-extra-args: merge should succeed" + case_dir=$(make_gitlab_case gitlab-extra-args-attended) + set +e + run_pr_merge "$case_dir" task-x1 "$MR_URL" --attended-override -- --remove-source-branch \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 0 "$rc" "gitlab-extra-args-attended: attended override should merge" merge_line=$(glab_merge_line "$case_dir/glab.log") [ "$merge_line" = "GITLAB_HOST=$MR_HOST mr merge 7 -R $MR_PROJECT_URL --sha $MR_HEAD --yes --remove-source-branch" ] \ - || fail "gitlab-extra-args: extra glab flags were not forwarded: '$merge_line'" - pass "fm-pr-merge forwards extra flags to glab mr merge after the -- separator" + || fail "gitlab-extra-args-attended: extra glab flags were not forwarded: '$merge_line'" + pass "fm-pr-merge refuses GitLab source-branch deletion unless --attended-override is passed" } test_gitlab_merge_failure_propagates() { @@ -1579,7 +1516,7 @@ test_gitlab_merge_failure_propagates() { # Each pre-merge condition, driven one at a time, so no condition can be # carried by another. The refusal names that condition, no merge is attempted, # and pr= is still recorded and the poll still armed exactly as the GitHub path -# leaves them when gh-axi itself fails. +# leaves them when live verification or the gh merge fails. test_gitlab_each_condition_refuses_independently() { local case_dir rc name expected spec set -- \ @@ -1771,20 +1708,23 @@ test_gitlab_head_override_args_refuse_before_recording() { } test_github_still_forwards_sha_arg() { - local case_dir + local case_dir rc case_dir=$(make_case github-sha-arg) mkdir -p "$case_dir/wt" add_gh_mocks "$case_dir" dddddddddddddddddddddddddddddddddddddddd : > "$case_dir/gh-axi.log" - # --sha is rejected only where the head is firstmate's to determine. GitHub's - # extra args are the caller's business exactly as they were. + set +e run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/44 -- --sha abc123 \ - > "$case_dir/stdout" 2> "$case_dir/stderr" || fail "github-sha-arg: fm-pr-merge failed" - - grep -qxF 'pr merge 44 --repo example/repo --squash --sha abc123' "$case_dir/gh-axi.log" \ - || fail "github-sha-arg: the GitHub path stopped forwarding a caller --sha" - pass "fm-pr-merge leaves GitHub extra-arg handling unchanged, including --sha" + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "github-sha-arg: a caller --sha must be refused on GitHub too" + assert_grep 'extra merge arguments must not override the head commit' "$case_dir/stderr" \ + "github-sha-arg: refusal did not name the head override" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-sha-arg: gh pr merge ran despite the head override" + pass "fm-pr-merge refuses a caller --sha on GitHub because the head comes from the live read" } # --- durable merge outcome --------------------------------------------------- @@ -1997,6 +1937,11 @@ test_distinct_merged_prs_keep_distinct_wakes() { rm -f "$case_dir/state/task-x1.check.sh" \ "$case_dir/state/task-x1.pr-poll" \ "$case_dir/state/task-x1.pr-poll-registration" + # Reused tasks re-bind through fm-pr-check before the next merge. Merge + # refuses a URL that is not the recorded pr=, so drop the first PR identity. + grep -vE '^(pr|pr_head)=' "$case_dir/state/task-x1.meta" \ + > "$case_dir/state/task-x1.meta.rebind" + mv "$case_dir/state/task-x1.meta.rebind" "$case_dir/state/task-x1.meta" FM_TEST_HOME="$case_dir/home" run_pr_merge "$case_dir" task-x1 "$second_url" \ >"$case_dir/stdout-2" 2>"$case_dir/stderr-2" \ || fail "distinct-merge-wakes: second merge failed" @@ -2102,7 +2047,7 @@ test_github_mismatched_queue_flags_still_name_the_retry test_github_unrecognised_queue_method_still_names_the_queue test_github_unreadable_queue_rules_are_not_reported_as_no_queue test_github_no_queue_rule_says_nothing_about_a_queue -test_github_fallback_view_refusal_says_the_queue_was_unobservable +test_github_unmerged_fallback_cannot_replace_queue_aware_read test_github_auto_merge_without_queue_refuses_legibly test_github_failed_merge_never_claims_armed_auto_merge test_github_failed_merge_with_queue_flags_never_claims_acceptance @@ -2158,8 +2103,7 @@ test_absent_backlog_still_merges() { expect_code 0 "$rc" "absent-backlog-merges: a home with no backlog must still merge" assert_no_grep 'held for the captain' "$case_dir/stderr" \ "absent-backlog-merges: an absent backlog was read as a captain hold" - grep -qxF 'pr merge 61 --repo example/repo --squash' "$case_dir/gh-axi.log" \ - || fail "absent-backlog-merges: the merge was not attempted" + assert_logged_gh_merge "$case_dir" 61 example/repo --squash pass "fm-pr-merge proceeds when the home carries no backlog at all" } @@ -2181,8 +2125,8 @@ test_unreadable_backlog_refuses_the_merge() { expect_code 1 "$rc" "unreadable-backlog-refuses: an unreadable authority record must refuse" assert_grep 'refusing to merge' "$case_dir/stderr" \ "unreadable-backlog-refuses: the refusal did not say it refused to merge" - [ ! -s "$case_dir/gh-axi.log" ] \ - || fail "unreadable-backlog-refuses: the forge was called despite an unreadable record" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "unreadable-backlog-refuses: the forge merge ran despite an unreadable record" pass "fm-pr-merge refuses when the backlog exists but cannot be read" } @@ -2205,8 +2149,8 @@ test_unreadable_backend_config_refuses_the_merge() { expect_code 1 "$rc" "unreadable-backend-config-refuses: an unreadable authority route must refuse" assert_grep 'tasks-axi backend configuration cannot be read' "$case_dir/stderr" \ "unreadable-backend-config-refuses: the unreadable authority route was not named" - [ ! -s "$case_dir/gh-axi.log" ] \ - || fail "unreadable-backend-config-refuses: the forge was called despite an unreadable authority route" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "unreadable-backend-config-refuses: the forge merge ran despite an unreadable authority route" pass "fm-pr-merge refuses when its configured backend cannot be read" } @@ -2232,8 +2176,8 @@ test_unreadable_user_backend_config_refuses_the_merge() { expect_code 1 "$rc" "unreadable-user-backend-config-refuses: an unreadable authority route must refuse" assert_grep "tasks-axi backend configuration cannot be read at $user_config" "$case_dir/stderr" \ "unreadable-user-backend-config-refuses: the unreadable authority route was not named" - [ ! -s "$case_dir/gh-axi.log" ] \ - || fail "unreadable-user-backend-config-refuses: the forge was called despite an unreadable authority route" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "unreadable-user-backend-config-refuses: the forge merge ran despite an unreadable authority route" pass "fm-pr-merge refuses when its user backend configuration cannot be read" } @@ -2259,8 +2203,8 @@ test_untraversable_user_backend_config_directory_refuses_the_merge() { expect_code 1 "$rc" "untraversable-user-backend-config-directory-refuses: an unreadable authority route must refuse" assert_grep "tasks-axi backend configuration cannot be read at $user_config" "$case_dir/stderr" \ "untraversable-user-backend-config-directory-refuses: the unreadable authority route was not named" - [ ! -s "$case_dir/gh-axi.log" ] \ - || fail "untraversable-user-backend-config-directory-refuses: the forge was called despite an unreadable authority route" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "untraversable-user-backend-config-directory-refuses: the forge merge ran despite an unreadable authority route" pass "fm-pr-merge refuses when its user backend configuration directory cannot be traversed" } @@ -2280,10 +2224,9 @@ test_absent_user_backend_config_directory_and_backlog_still_merge() { set -e expect_code 0 "$rc" "absent-user-backend-config-directory-and-backlog-merges: sound defaults and no backlog must permit merging" - [ "$(grep -c '^pr merge ' "$case_dir/gh-axi.log")" -eq 1 ] \ + [ "$(grep -c '^pr merge ' "$case_dir/gh.log")" -eq 1 ] \ || fail "absent-user-backend-config-directory-and-backlog-merges: the forge must merge exactly once" - grep -qxF 'pr merge 67 --repo example/repo --squash' "$case_dir/gh-axi.log" \ - || fail "absent-user-backend-config-directory-and-backlog-merges: the expected merge was not attempted" + assert_logged_gh_merge "$case_dir" 67 example/repo --squash pass "fm-pr-merge proceeds once when its user configuration directory and backlog are genuinely absent" } @@ -2307,11 +2250,235 @@ test_backend_override_bypasses_unreadable_user_config() { chmod 644 "$user_config" expect_code 0 "$rc" "backend-override-bypasses-unreadable-user-config: an explicit backend must bypass config" - grep -qxF 'pr merge 65 --repo example/repo --squash' "$case_dir/gh-axi.log" \ - || fail "backend-override-bypasses-unreadable-user-config: the merge was not attempted" + assert_logged_gh_merge "$case_dir" 65 example/repo --squash pass "fm-pr-merge honors a backend override over an unreadable user configuration" } +test_github_red_checks_refuse_and_allow_red_waives_named() { + local case_dir rc head + head=aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa + case_dir=$(make_case github-red-checks) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_red_json "$case_dir" "$head" lint + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/80 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "github-red: a red check must refuse" + assert_grep "check 'lint' is not green" "$case_dir/stderr" \ + "github-red: the red check was not named" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-red: gh pr merge ran on a red PR" + + case_dir=$(make_case github-allow-red) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_red_json "$case_dir" "$head" lint + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/81 \ + --allow-red lint \ + > "$case_dir/stdout" 2> "$case_dir/stderr" || fail "github-allow-red: named waiver should merge" + assert_logged_gh_merge "$case_dir" 81 example/repo --squash + pass "fm-pr-merge refuses red GitHub checks and waives only a named --allow-red check" +} + +test_allow_red_is_refused_while_away() { + local case_dir rc head + head=abababababababababababababababababababab + case_dir=$(make_case github-allow-red-away) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_red_json "$case_dir" "$head" lint + write_away_record "$case_dir" --grant task-x1 + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/82 \ + --allow-red lint \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 2 "$rc" "github-allow-red-away: --allow-red must be refused while away" + assert_grep '--allow-red is attended-only' "$case_dir/stderr" \ + "github-allow-red-away: refusal did not name attended-only" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-allow-red-away: gh pr merge ran despite away --allow-red" + + case_dir=$(make_case github-allow-red-away-after-view) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_red_json "$case_dir" "$head" lint + write_away_record "$case_dir" --grant task-x1 + mv "$case_dir/state/.afk-contract" "$case_dir/away-record-after-view" + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/82 \ + --allow-red lint \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 2 "$rc" "github-allow-red-away-after-view: late away publication must refuse --allow-red" + assert_grep '--allow-red is attended-only' "$case_dir/stderr" \ + "github-allow-red-away-after-view: late refusal did not name attended-only" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-allow-red-away-after-view: gh pr merge ran after late away publication" + pass "fm-pr-merge rechecks away presence before an attended red merge" +} + +test_allow_red_requires_one_separate_name() { + local case_dir rc head + head=afafafafafafafafafafafafafafafafafafafaf + + case_dir=$(make_case github-allow-red-equals) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_red_json "$case_dir" "$head" lint + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/87 \ + --allow-red=lint > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 2 "$rc" "github-allow-red-equals: equals form must be refused" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-allow-red-equals: gh pr merge ran for the equals alias" + + case_dir=$(make_case github-allow-red-duplicate) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_red_json "$case_dir" "$head" lint + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/88 \ + --allow-red lint --allow-red unit > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 2 "$rc" "github-allow-red-duplicate: duplicate waiver must be refused" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-allow-red-duplicate: gh pr merge ran for duplicate waivers" + pass "fm-pr-merge accepts exactly one separately named red-check waiver" +} + +test_away_grant_and_yolo_and_hold_for_return() { + local case_dir rc url head + head=acacacacacacacacacacacacacacacacacacacac + url=https://github.com/example/repo/pull/83 + + case_dir=$(make_case away-held) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_away_record "$case_dir" + set +e + run_pr_merge "$case_dir" task-x1 "$url" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "away-held: ungranted merge must refuse" + assert_grep 'task task-x1 is held for the captain return' "$case_dir/stderr" \ + "away-held: refusal did not name hold-for-return" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "away-held: gh pr merge ran without a grant" + + case_dir=$(make_case away-held-attended-override) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_away_record "$case_dir" + set +e + run_pr_merge "$case_dir" task-x1 "$url" --attended-override \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "away-held-override: --attended-override must not skip the grant" + assert_grep 'task task-x1 is held for the captain return' "$case_dir/stderr" \ + "away-held-override: override skipped the grant" + + case_dir=$(make_case away-grant) + mkdir -p "$case_dir/wt" "$case_dir/home" + add_gh_mocks "$case_dir" "$head" + write_away_record "$case_dir" --grant task-x1 + FM_TEST_HOME="$case_dir/home" run_pr_merge "$case_dir" task-x1 "$url" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" || fail "away-grant: granted green merge should succeed" + assert_logged_gh_merge "$case_dir" 83 example/repo --squash + assert_grep "merge landed: task-x1 $url away-grant" "$case_dir/state/.wake-queue" \ + "away-grant: the durable outcome did not tag away-grant" + + case_dir=$(make_case away-yolo) + mkdir -p "$case_dir/wt" "$case_dir/home" + add_gh_mocks "$case_dir" "$head" + printf '\nyolo=on\n' >> "$case_dir/state/task-x1.meta" + write_away_record "$case_dir" + FM_TEST_HOME="$case_dir/home" run_pr_merge "$case_dir" task-x1 "$url" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" || fail "away-yolo: yolo green merge should succeed" + assert_grep "merge landed: task-x1 $url yolo" "$case_dir/state/.wake-queue" \ + "away-yolo: the durable outcome did not tag yolo" + pass "away merges require yolo or a grant, and --attended-override does not skip that" +} + +test_away_grant_does_not_bypass_red_or_identity() { + local case_dir rc head + head=adadadadadadadadadadadadadadadadadadadad + case_dir=$(make_case away-grant-red) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_red_json "$case_dir" "$head" lint + write_away_record "$case_dir" --grant task-x1 + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/84 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "away-grant-red: a grant must not waive red checks" + assert_grep "check 'lint' is not green" "$case_dir/stderr" \ + "away-grant-red: C1 did not refuse the red check" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "away-grant-red: gh pr merge ran on a granted red PR" + + case_dir=$(make_case pr-identity-mismatch) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + printf '\npr=https://github.com/example/repo/pull/99\n' >> "$case_dir/state/task-x1.meta" + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/85 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "pr-identity: a different recorded URL must refuse" + assert_grep 'is bound to https://github.com/example/repo/pull/99' "$case_dir/stderr" \ + "pr-identity: refusal did not name the recorded URL" + pass "a grant does not bypass red checks, and a recorded pr= must match the URL" +} + +test_unreadable_away_record_refuses_merge() { + local case_dir rc + case_dir=$(make_case away-unreadable) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" aeaeaeaeaeaeaeaeaeaeaeaeaeaeaeaeaeaeaeae + printf 'not-a-contract\n' > "$case_dir/state/.afk-contract" + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/86 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "away-unreadable: an unreadable away record must refuse" + assert_grep 'away-posture record could not be read' "$case_dir/stderr" \ + "away-unreadable: refusal did not fail closed" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "away-unreadable: gh pr merge ran despite an unreadable record" + pass "an unreadable away-posture record refuses the merge instead of skipping the grant" +} + +test_allow_red_refused_on_gitlab() { + local case_dir rc + case_dir=$(make_gitlab_case gitlab-allow-red) + set +e + run_pr_merge "$case_dir" task-x1 "$MR_URL" --allow-red lint \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 2 "$rc" "gitlab-allow-red: --allow-red must not apply on GitLab" + assert_grep '--allow-red does not apply to GitLab' "$case_dir/stderr" \ + "gitlab-allow-red: refusal did not name GitLab" + [ ! -s "$case_dir/glab.log" ] || fail "gitlab-allow-red: glab ran despite --allow-red" + pass "fm-pr-merge refuses --allow-red on GitLab" +} + test_gitlab_head_override_args_refuse_before_recording test_secondmate_merge_reports_upward_once test_secondmate_merge_reports_on_the_local_route @@ -2331,3 +2498,10 @@ test_unreadable_user_backend_config_refuses_the_merge test_untraversable_user_backend_config_directory_refuses_the_merge test_absent_user_backend_config_directory_and_backlog_still_merge test_backend_override_bypasses_unreadable_user_config +test_github_red_checks_refuse_and_allow_red_waives_named +test_allow_red_is_refused_while_away +test_allow_red_requires_one_separate_name +test_away_grant_and_yolo_and_hold_for_return +test_away_grant_does_not_bypass_red_or_identity +test_unreadable_away_record_refuses_merge +test_allow_red_refused_on_gitlab From 9074f9d20d3dd6b632051623f71797084a7aab36 Mon Sep 17 00:00:00 2001 From: Christoph Meise <christoph@scripe.io> Date: Fri, 11 Sep 2026 20:19:45 +0200 Subject: [PATCH 06/31] fix(bin): bound the Claude turn-end re-block against a frozen auto-arm epoch (#4221) The --claude guard's re-block budget charged the auto-arm ledger epoch, not the re-block: `budget_account_current_epoch` advanced the session count only when `state/.claude-autoarm-epoch` named a different generation than the previous accounting. The epoch advances only inside the auto-arm hook's generation claim, so a hook kept inert before that claim - a session lock held by a live harness outside its ancestry, a hook that never fires, or an identity or write failure ahead of `fm_autoarm_claim_next` - left the ledger frozen at its last outcome and the count frozen with it. Reproduced in a fixture: twelve consecutive Stops re-blocked with the count at 0 and the attended fail-open never fired, leaving only Claude's silent 8-block override, the blind end the bounded alarm exists to prevent. The budget now charges a re-block against an epoch the previous re-block already charged, while still charging each epoch at most once per Stop so the wait loop's repeated observations of one fresh terminal outcome and the same invocation's block decision cannot double count. The advancing-epoch progression is unchanged: three re-blocks, then one attended fail-open for a verified failure episode, and a frozen epoch now follows the same shape. Budget exhaustion without a verified failure still blocks, by the existing contract, and positive watcher recovery still clears the whole episode. Regression coverage drives the real auto-arm hook against a foreign session lock holder, asserts the ledger itself stays frozen, and fails before the fix in both the verified and unverified shapes; the existing unverified budget test now proves its budget actually ran out. --- bin/fm-turnend-guard.sh | 40 +++++++++-- docs/turnend-guard.md | 6 +- tests/fm-turnend-guard.test.sh | 124 +++++++++++++++++++++++++++++++++ 3 files changed, 163 insertions(+), 7 deletions(-) diff --git a/bin/fm-turnend-guard.sh b/bin/fm-turnend-guard.sh index 398fa68b7b7..ffceafaee51 100755 --- a/bin/fm-turnend-guard.sh +++ b/bin/fm-turnend-guard.sh @@ -79,7 +79,12 @@ # with the repair banner, bounded to FM_CLAUDE_TURNEND_BLOCK_BUDGET # (default 3) consecutive blocks per session - safely below Claude Code's # hard 8-consecutive-block override - then allow one loud attended -# fail-open only for an already verified failure episode. +# fail-open only for an already verified failure episode. The budget +# charges each event epoch once, and it also charges every re-block +# against an epoch the auto-arm never advanced past the previous +# re-block (budget_account_current_epoch owns that rule), so an inert +# hook that leaves the ledger frozen cannot hold the guard in an +# unbounded re-block loop below that override. set -u SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -249,12 +254,31 @@ fi # The Stop-owned auto-arm fires on the same Stop event. Give it a brief bounded # window to prove it owns recovery for this event epoch before consuming one of # Claude's bounded continuations. -budget_account_current_epoch() { - local current_epoch outcome old_session old_count old_epoch tmp initialized +# +# Budget accounting, under the budget lock. Sets COUNT (the session's +# consumed continuations, including this one) and BUDGET_INITIALIZED_FAILURE. +# The ledger's epoch identity is what is charged: a new epoch charges once, +# and an epoch this same invocation already charged is never charged again, +# because the wait loop above can observe one fresh terminal epoch many times +# before the block decision. Across Stops the two callers differ: +# - observe (the allow paths in autoarm_owns_recovery): seeing an +# already-charged epoch again is free - it is the same claim, seen again. +# - block (the re-block path): a re-block against the epoch the previous +# re-block already charged is a new consumed continuation, because the +# auto-arm advanced nothing between the two Stops - it did not participate +# at all, which is exactly the absence this budget bounds. Charging only +# epoch changes let an inert hook (identity-gated, never fired, or failing +# before its generation claim) freeze the ledger and the count together, +# so the guard re-blocked without limit and the attended fail-open below +# never became reachable. +BUDGET_CHARGED_EPOCH= +budget_account_current_epoch() { # [observe|block] + local mode=${1:-observe} current_epoch outcome old_session old_count old_epoch tmp initialized charged fm_lock_try_acquire "$BUDGET_LOCK" || return 1 current_epoch=$(sed -n '1s/^epoch=\([0-9][0-9]*\) .*/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) outcome=$(sed -n '1s/^.*outcome=\([a-z][a-z-]*\) .*$/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) initialized=0 + charged=0 COUNT=0 if [ -f "$BUDGET_FILE" ]; then old_session=$(sed -n '1s/^session=//p' "$BUDGET_FILE" 2>/dev/null || true) @@ -266,13 +290,18 @@ budget_account_current_epoch() { if [ "$old_session" = "$SESSION_ID" ]; then COUNT=$old_count if [ -n "$current_epoch" ] && [ "$old_epoch" = "$current_epoch" ]; then - : + if [ "$mode" = block ] && [ "$BUDGET_CHARGED_EPOCH" != "$current_epoch" ]; then + COUNT=$((COUNT + 1)) + charged=1 + fi else COUNT=$((COUNT + 1)) + charged=1 fi fi fi if [ ! -f "$BUDGET_FILE" ] || [ "${old_session:-}" != "$SESSION_ID" ]; then + charged=1 case "$outcome" in failed|failed-suppressed) if [ -e "$FAILURE_NOTICE" ]; then @@ -293,6 +322,7 @@ budget_account_current_epoch() { return 1 fi rm -f "$tmp" 2>/dev/null || true + [ "$charged" -eq 0 ] || BUDGET_CHARGED_EPOCH=$current_epoch BUDGET_INITIALIZED_FAILURE=$initialized fm_lock_release "$BUDGET_LOCK" return 0 @@ -456,7 +486,7 @@ fi # The auto-arm genuinely failed to establish: consume the bounded re-block # budget before considering the verified one-time attended fail-open. -budget_account_current_epoch || block_stop +budget_account_current_epoch block || block_stop terminal_fail_open terminal_status=$? if [ "$terminal_status" -eq 0 ]; then diff --git a/docs/turnend-guard.md b/docs/turnend-guard.md index 6c5d134dcd5..579225d2970 100644 --- a/docs/turnend-guard.md +++ b/docs/turnend-guard.md @@ -111,7 +111,9 @@ The first fresh exhausted-failure epoch preserves its handoff without consuming When none of those proofs appears, it re-blocks up to `FM_CLAUDE_TURNEND_BLOCK_BUDGET` times (default 3, below Claude's 8-block override). In Claude mode, positive watcher recovery clears the block budget, failure notice, and attended alarm together under the existing budget lock before either hook reports ordinary recovery. The one loud attended fail-open is available only when the auto-arm has recorded an exhausted failure, its one notice is already consumed, the block budget is exhausted, and a final check finds neither a healthy watcher nor an automatic continuation. -Each epoch identity is accounted at most once under the budget lock. +Each epoch identity is charged at most once per Stop under the budget lock, and a re-block against an epoch the auto-arm did not advance past the previous re-block is charged as well. +That second rule is what bounds an inert auto-arm: a hook kept silent by a session lock held by a live harness outside its ancestry, a hook that never fires, or a hook failing before its generation claim leaves the ledger frozen at its last outcome. +Charging only epoch changes let the count freeze with that ledger, so the guard re-blocked without limit and the attended fail-open was never reachable; `budget_account_current_epoch` in `bin/fm-turnend-guard.sh` owns the rule. Whenever both coordination locks are needed, positive auto-arm recovery and the terminal check acquire the auto-arm owner lock before the budget lock. After that alarm, the Stop auto-arm suppresses further exit-2 continuations until positive watcher recovery, so the final fail-open remains reachable. The alarm cannot repeat during that failure episode, and a later unhealthy stop blocks again. @@ -183,7 +185,7 @@ That warning uses `bin/fm-supervision-instructions.sh --repair-line`, so it alwa ## Regression coverage -`tests/fm-turnend-guard.test.sh` covers the predicate, main and secondmate primary scope, child-worktree exclusion, `FM_HOME` and `FM_STATE_OVERRIDE` precedence, the live-lock and fresh-beacon guard predicate, the cooperative `--claude` open-generation claim wait, monotonic failed-epoch progression, bounded attended fail-open, post-alarm continuation suppression, positive recovery reset, generation and legacy claim cases that must block or clear instead of allowing a blind stop, away-mode daemon ownership between watcher cycles and over a watcher lock left behind by an exited watcher, plus its dead, pid-reused, absent, stale-beacon, and away-mode-off negatives, the away-mode beacon's poll-derived grace widening for a live daemon still mid-cycle and its bound against a dead daemon, a beacon older than that wider grace, and FM_POLL's inapplicability with away mode off, Pi logical-run latching, missing-`jq` behavior, all five primary registrations, Grok native and legacy selection, typed field precedence, malformed input, and exactly-one-path safety. +`tests/fm-turnend-guard.test.sh` covers the predicate, main and secondmate primary scope, child-worktree exclusion, `FM_HOME` and `FM_STATE_OVERRIDE` precedence, the live-lock and fresh-beacon guard predicate, the cooperative `--claude` open-generation claim wait, monotonic failed-epoch progression, bounded attended fail-open, the same bound against a ledger frozen by an inert auto-arm with and without a verified failure episode, post-alarm continuation suppression, positive recovery reset, generation and legacy claim cases that must block or clear instead of allowing a blind stop, away-mode daemon ownership between watcher cycles and over a watcher lock left behind by an exited watcher, plus its dead, pid-reused, absent, stale-beacon, and away-mode-off negatives, the away-mode beacon's poll-derived grace widening for a live daemon still mid-cycle and its bound against a dead daemon, a beacon older than that wider grace, and FM_POLL's inapplicability with away mode off, Pi logical-run latching, missing-`jq` behavior, all five primary registrations, Grok native and legacy selection, typed field precedence, malformed input, and exactly-one-path safety. `tests/fm-guard-stale-banner.test.sh` covers the pull-guard predicate, including the persistent-model fresh-leftover-beacon negative control; the auto-arm model's healthy fresh-beacon-without-a-watcher case, session-and-recovery-bound long-turn rewake tolerance, independently broken tolerance signals, open-claim negative control, stale-beacon alarm, and isolation from other models; and the extension model's live-watcher path, ownership-qualified fresh hand-off, held-lock failures, independently broken ownership signals, stale-beacon alarm, queued-wake warning, and Pi and pi-signed harness routing. It also covers true-reason banner wording and reason-keyed episode dedup surviving a beacon mtime change. `tests/fm-cursor-primary.test.sh` covers the Cursor park end to end over real processes with no harness installed: each tracked Claude-shaped entrypoint standing down on a Cursor payload, both follow-up sources, the bounded repair nag and its reset, the nested loop bounds, supersession, away-mode and lock-ownership inertness, Pi-host stand-down without Cursor identity and continued parking when `PI_CODING_AGENT` leaks alongside `CURSOR_AGENT` or `CURSOR_INVOKED_AS`, child-worktree exclusion, and that the adapter never exits 2. diff --git a/tests/fm-turnend-guard.test.sh b/tests/fm-turnend-guard.test.sh index 7d656ef2b35..1fd43c8d6c3 100755 --- a/tests/fm-turnend-guard.test.sh +++ b/tests/fm-turnend-guard.test.sh @@ -1220,6 +1220,17 @@ run_integrated_autoarm() { ' 2>&1 } +# The same real hook, fired from a harness-named process that does NOT write +# state/.lock: whoever already holds that lock decides whether this firing is +# the owning session's or a competing one. +run_integrated_autoarm_unowned() { + local dir=$1 home + home=$(cd "$dir" && pwd) + # shellcheck disable=SC2016 # the fake harness expands FM_HOME inside its child shell. + printf '{"session_id":"sess-claude-mode","stop_hook_active":false}\n' \ + | FM_HOME="$home" "$dir/fake-claude" -c '"$FM_HOME/bin/fm-claude-stop-autoarm.sh"' 2>&1 +} + write_integrated_failed_arm() { local dir=$1 cat > "$dir/bin/fm-watch-arm.sh" <<'SH' @@ -1614,6 +1625,115 @@ test_hook_claude_mode_integrated_monotonic_fail_open() { pass "fm-turnend-guard --claude: integrated fresh failures reach one bounded fail-open, stop continuation, and reset on recovery" } +# The auto-arm's ledger epoch advances only when the hook reaches its +# generation claim. A live harness-named process outside the hook's ancestry +# holding state/.lock keeps the hook inert by its identity contract, so the +# ledger stays at the exhausted-failure epoch the hook wrote before it went +# quiet. The block budget used to advance only on an epoch change, so this +# shape re-blocked without limit and the attended fail-open never fired: the +# budget must count consecutive re-blocks against an unchanged epoch instead. +hold_session_lock_from_foreign_harness() { # sets FOREIGN_LOCK_HOLDER + local dir=$1 + # `bash -c` execs a single command in place, which would rename the process + # to sleep; the trailing no-op keeps the harness-named shell as the holder. + # Started in this shell, not a command substitution, so the caller can reap + # it and no inherited pipe keeps a substitution waiting on the sleeper. + "$dir/fake-claude" -c 'sleep 60; true' >/dev/null 2>&1 & + FOREIGN_LOCK_HOLDER=$! + printf '%s\n' "$FOREIGN_LOCK_HOLDER" > "$dir/state/.lock" +} + +test_hook_claude_mode_frozen_epoch_reaches_bounded_fail_open() { + local dir out status guard_out guard_status holder i pid identity count epoch_line + dir=$(make_primary_dir "$TMP_ROOT/hook-claude-frozen-epoch") + : > "$dir/state/task1.meta" + install_integrated_autoarm "$dir" + write_integrated_failed_arm "$dir" + + out=$(run_integrated_autoarm "$dir"); status=$? + expect_code 2 "$status" "the exhausted auto-arm cycle must emit its one failure notice before going quiet" + guard_out=$(FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=100 run_hook_claude "$dir" true); guard_status=$? + expect_code 0 "$guard_status" "the first failed epoch must own its Stop handoff" + epoch_line=$(sed -n '1p' "$dir/state/.claude-autoarm-epoch") + + hold_session_lock_from_foreign_harness "$dir" + holder=$FOREIGN_LOCK_HOLDER + for i in 1 2 3 4; do + out=$(run_integrated_autoarm_unowned "$dir"); status=$? + expect_code 0 "$status" "an auto-arm outside the lock owner's ancestry must stay inert at stop $i" + [ -z "$out" ] || fail "inert auto-arm produced output at stop $i: $out" + [ "$(sed -n '1p' "$dir/state/.claude-autoarm-epoch")" = "$epoch_line" ] \ + || fail "the ledger epoch advanced at stop $i, so this case no longer drives a frozen epoch" + guard_out=$(FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=100 run_hook_claude "$dir" true); guard_status=$? + if [ "$i" -lt 4 ]; then + expect_code 2 "$guard_status" "frozen-epoch stop $i must still re-block within the budget" + assert_contains "$guard_out" "TURN WOULD END BLIND" "frozen-epoch re-block $i lost the blind-turn banner" + assert_not_contains "$guard_out" 'FIRSTMATE SUPERVISION IS GENUINELY DOWN' "fail-open fired before the frozen-epoch budget was spent" + assert_absent "$dir/state/.claude-autoarm-failure-alarmed" "frozen-epoch re-block $i consumed the attended alarm early" + else + expect_code 0 "$guard_status" "the frozen-epoch progression must reach the attended fail-open" + assert_contains "$guard_out" 'FIRSTMATE SUPERVISION IS GENUINELY DOWN' "the frozen-epoch fail-open alarm is missing" + assert_present "$dir/state/.claude-autoarm-failure-alarmed" "the frozen-epoch fail-open did not consume its episode alarm" + fi + done + + guard_out=$(FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=100 run_hook_claude "$dir" true); guard_status=$? + expect_code 2 "$guard_status" "a later unhealthy stop after the frozen-epoch alarm must remain attended" + assert_not_contains "$guard_out" 'FIRSTMATE SUPERVISION IS GENUINELY DOWN' "the attended alarm repeated against the frozen epoch" + + # The other direction: the bound must not outlive the failure. A verified + # healthy watcher still lets the stop through and clears the whole episode. + sleep 60 & + pid=$! + identity=$(watcher_identity "$dir" "$pid") || { + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + kill "$holder" 2>/dev/null || true + wait "$holder" 2>/dev/null || true + fail "could not identify the frozen-epoch recovery watcher" + } + record_watcher_lock "$dir" "$pid" "$identity" + touch "$dir/state/.last-watcher-beat" + guard_out=$(run_hook_claude "$dir" true); guard_status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + kill "$holder" 2>/dev/null || true + wait "$holder" 2>/dev/null || true + rm -rf "$dir/state/.watch.lock" + expect_code 0 "$guard_status" "a healthy watcher must still allow the stop after a frozen-epoch alarm" + [ -z "$guard_out" ] || fail "healthy allow after the frozen-epoch alarm produced output: $guard_out" + assert_absent "$dir/state/.turnend-claude-blocks" "positive recovery left the frozen-epoch block budget" + assert_absent "$dir/state/.claude-autoarm-failure-notified" "positive recovery left the failure notice" + assert_absent "$dir/state/.claude-autoarm-failure-alarmed" "positive recovery left the attended alarm" + guard_out=$(FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=100 run_hook_claude "$dir" true); guard_status=$? + expect_code 2 "$guard_status" "a later unhealthy stop must re-block from a fresh budget" + count=$(sed -n '2s/^count=//p' "$dir/state/.turnend-claude-blocks") + [ "$count" = 1 ] || fail "the post-recovery episode must restart its budget at 1, got $count" + pass "fm-turnend-guard --claude: an inert auto-arm's frozen epoch reaches one bounded fail-open and resets on recovery" +} + +# The same frozen ledger without a verified failure episode: the budget must +# still provably run out, and the verified-failure gate - not a stuck counter - +# is what keeps the stop blocking after that. +test_hook_claude_mode_frozen_epoch_without_verified_failure_spends_budget_and_keeps_blocking() { + local dir out status i count + dir=$(make_primary_dir "$TMP_ROOT/hook-claude-frozen-unverified") + : > "$dir/state/task1.meta" + printf 'epoch=7 owner_pid=999 outcome=clean updated_at=1\n' > "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.claude-autoarm-epoch" + for i in 1 2 3 4 5; do + out=$(FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=100 run_hook_claude "$dir" false); status=$? + expect_code 2 "$status" "frozen unverified stop $i must keep blocking" + assert_not_contains "$out" 'systemMessage' "an unverified frozen epoch must never fail open" + [ "$(sed -n '1p' "$dir/state/.claude-autoarm-epoch")" = 'epoch=7 owner_pid=999 outcome=clean updated_at=1' ] \ + || fail "the guard rewrote the frozen ledger at stop $i" + done + count=$(sed -n '2s/^count=//p' "$dir/state/.turnend-claude-blocks") + [ "$count" -gt 3 ] 2>/dev/null || fail "the block budget must run out against a frozen epoch, but the recorded count is ${count:-absent}" + assert_absent "$dir/state/.claude-autoarm-failure-alarmed" "an unverified frozen epoch recorded an attended alarm" + pass "fm-turnend-guard --claude: a frozen unverified epoch spends the budget yet still blocks" +} + test_hook_claude_mode_recovery_contention_is_not_ordinary_allow() { local dir pid identity holder out status dir=$(make_primary_dir "$TMP_ROOT/hook-claude-recovery-contention") @@ -1703,6 +1823,8 @@ test_hook_claude_mode_budget_without_verified_failure_keeps_blocking() { out=$(FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=100 run_hook_claude "$dir" false); status=$? expect_code 2 "$status" "--claude block $i must exit 2 within the budget" done + count=$(sed -n '2s/^count=//p' "$dir/state/.turnend-claude-blocks") + [ "$count" -gt 3 ] 2>/dev/null || fail "four consecutive blocks must spend the budget, but the recorded count is ${count:-absent}" assert_not_contains "$out" 'systemMessage' "budget exhaustion without verified auto-arm failure must not fail open" assert_absent "$dir/state/.claude-autoarm-failure-alarmed" "unverified budget exhaustion recorded an attended alarm" pass "fm-turnend-guard --claude: budget exhaustion alone cannot permit a blind stop" @@ -2135,6 +2257,8 @@ test_hook_claude_mode_blocks_on_stuck_generation_claim test_hook_claude_mode_terminal_fail_open_clears_abandoned_claim test_hook_claude_mode_preserves_fresh_failed_progression test_hook_claude_mode_integrated_monotonic_fail_open +test_hook_claude_mode_frozen_epoch_reaches_bounded_fail_open +test_hook_claude_mode_frozen_epoch_without_verified_failure_spends_budget_and_keeps_blocking test_hook_claude_mode_recovery_contention_is_not_ordinary_allow test_hook_claude_mode_concurrent_recovery_resets_are_idempotent test_hook_claude_mode_stale_rewake_epoch_blocks From ff5c7af9d7653b14ee9b9a74ec522cd85a3e829c Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Fri, 11 Sep 2026 15:26:02 -0700 Subject: [PATCH 07/31] feat(herdr): add guarded foreground viewer for live validation (#4242) * feat(herdr): attach a real foreground viewer so the live-client teardown cases can be driven PR #4131 gated the Herdr active-tab close refusal on a live foreground client instead of the persisted `.focused` pointer, but only its two detached scenarios could be validated live. Every pseudo-terminal the runner built started at a zero-sized window grid, so Herdr registered no foreground client and `terminal title clear` kept answering `no_foreground_client`, leaving the four attached-client scenarios untested. That was a harness limit, not a product one. Add `fm-herdr-lab.sh viewer start|stop <session>`, backed by `bin/fm-herdr-lab-viewer.py`. The launcher sets the pty window size on the master fd BEFORE the fork, so the TUI cannot read the grid until it is already non-zero, and scrubs the inherited `HERDR_*` variables so Herdr's nested-viewer refusal does not fire when the helper runs inside one of its own panes. Attach and detach are both confirmed against the session's own foreground-client reason rather than assumed from a signal. The viewer inherits the lab's isolation contract: it attaches only to a session carrying this lab's ownership tripwire, never to `default`, and it signals only the processes it recorded, so a client someone else attached is never touched. Teardown now refuses while an owned viewer is still attached. Turn the reproduction into the regression with `tests/fm-herdr-attached-viewer-live-e2e.test.sh`, which drives #4131's scenarios 3, 4, 5, and 7 live against real Herdr and asserts the close refusal fires. Scenarios 4 and 5 need a focus change at one exact product boundary, so a PATH shim performs the real `tab focus` when the close helper issues its planning `pane get`. Removing either half of the recipe from the launcher makes the guard fail with the same `no_foreground_client` symptom #4131 reported. * test(herdr): fail loudly when an attached-viewer fixture cannot be created The fixture helpers run inside command substitutions, where fail() exits only the subshell and leaves the script running with empty ids. Return non-zero instead and carry the message at each call site. * fix(herdr): stop the viewer launcher's kill timer from raising on an exited child The SIGALRM escalation called os.kill unguarded, so a viewer that exited during the grace window turned an ordinary shutdown into a traceback inside the signal handler. * docs: list the lab viewer's pty engine in the bin toolbelt * no-mistakes(review): Harden Herdr viewer ownership and live CI coverage * no-mistakes(review): Validate viewer startup timeout and process ownership * no-mistakes(document): Document Herdr viewer safety contracts * no-mistakes(review): Fix viewer timeout to two seconds * no-mistakes(review): Cancel timed-out viewers and fix PTY grid * no-mistakes(review): Serialize viewer transitions and verify process parentage * no-mistakes(review): Harden viewer ownership locks and deduplicate CI * no-mistakes(review): Release interrupted locks and preserve viewer escalation * no-mistakes(review): Remove viewer locks and cancel interrupted launches * no-mistakes(review): Close viewer launch signal races * no-mistakes(document): Document attached Herdr viewer regression --- bin/fm-herdr-lab-viewer.py | 204 +++++++++++++ bin/fm-herdr-lab.sh | 241 ++++++++++++++- bin/fm-test-run.sh | 5 +- docs/herdr-backend.md | 2 + docs/scripts.md | 1 + docs/verification/runtime-backends.md | 32 ++ .../fm-herdr-attached-viewer-live-e2e.test.sh | 259 +++++++++++++++++ tests/fm-herdr-lab.test.sh | 274 ++++++++++++++++++ 8 files changed, 1015 insertions(+), 3 deletions(-) create mode 100755 bin/fm-herdr-lab-viewer.py create mode 100755 tests/fm-herdr-attached-viewer-live-e2e.test.sh diff --git a/bin/fm-herdr-lab-viewer.py b/bin/fm-herdr-lab-viewer.py new file mode 100755 index 00000000000..ce40e0b4a9e --- /dev/null +++ b/bin/fm-herdr-lab-viewer.py @@ -0,0 +1,204 @@ +#!/usr/bin/env python3 +"""Attach one real foreground Herdr viewer to a named lab session over a pty. + +bin/fm-herdr-lab.sh's ``viewer start`` is the only supported caller; run this +through that guard rather than directly, so the lab's ownership tripwire and +refuse-default checks still apply. + +Herdr registers a foreground client only when the attaching terminal reports a +usable window grid. A pty created by ``script`` or a bare ``pty.fork()`` from a +non-tty parent starts at 0x0, which makes Herdr report a zero-sized grid and +keeps ``client.window_title.clear`` answering ``no_foreground_client``. That is +why firstmate could not drive the attached-viewer teardown cases live before +this helper existed. The fix is ordering as much as sizing: the window size is +set on the master fd BEFORE the fork, so the TUI cannot read the grid until it +is already non-zero. + +The child also drops the inherited ``HERDR_*`` variables listed in +``SCRUBBED_ENV`` below. Herdr refuses to launch a nested viewer inside one of +its own panes, and this helper normally runs from exactly there. + +Usage: fm-herdr-lab-viewer.py <session> <pidfile> + +Exit status: + 0 the viewer ran and exited; + 2 the session or pidfile was invalid; + 3 the pty or the viewer process could not be created. +""" + +import errno +import fcntl +import os +import re +import signal +import struct +import subprocess +import sys +import termios + +# Herdr inherits these from the pane this helper runs in, and a nested viewer +# is refused outright. HERDR_SESSION is scrubbed with them so the explicit +# --session argument stays the viewer's only session source. +SCRUBBED_ENV = ( + "HERDR_ENV", + "HERDR_PANE_ID", + "HERDR_TAB_ID", + "HERDR_WORKSPACE_ID", + "HERDR_SOCKET_PATH", + "HERDR_BIN_PATH", + "HERDR_SESSION", +) + +SESSION_PATTERN = re.compile(r"\Afm-lab-[A-Za-z0-9][A-Za-z0-9_-]*\Z") +TERMINATE_GRACE_SECONDS = 5.0 +READ_CHUNK = 65536 +ROWS = 40 +COLS = 120 +TERMINATION_SIGNALS = (signal.SIGTERM, signal.SIGINT, signal.SIGHUP) + + +def _child(slave, master, session): + signal.pthread_sigmask(signal.SIG_UNBLOCK, TERMINATION_SIGNALS) + os.setsid() + try: + fcntl.ioctl(slave, termios.TIOCSCTTY, 0) + except OSError: + pass + for target in (0, 1, 2): + os.dup2(slave, target) + if slave > 2: + os.close(slave) + os.close(master) + env = {key: value for key, value in os.environ.items() if key not in SCRUBBED_ENV} + env.setdefault("TERM", "xterm-256color") + try: + os.execvpe("herdr", ["herdr", "--session", session], env) + except OSError: + pass + os._exit(127) + + +def _process_start(pid): + result = subprocess.run( + ["ps", "-p", str(pid), "-o", "lstart="], + check=True, + capture_output=True, + text=True, + env={**os.environ, "LC_ALL": "C"}, + ) + value = result.stdout.strip() + if not value: + raise RuntimeError("process start time unavailable") + return value + + +def _write_pidfile(path, launcher_pid, viewer_pid): + launcher_start = _process_start(launcher_pid) + viewer_start = _process_start(viewer_pid) + temporary = "%s.%d.tmp" % (path, launcher_pid) + with open(temporary, "w", encoding="utf-8") as handle: + handle.write("launcher_pid=%d\n" % launcher_pid) + handle.write("launcher_start=%s\n" % launcher_start) + handle.write("viewer_pid=%d\n" % viewer_pid) + handle.write("viewer_start=%s\n" % viewer_start) + os.rename(temporary, path) + + +def _drain(master): + while True: + try: + if not os.read(master, READ_CHUNK): + return + except OSError as error: + if error.errno == errno.EINTR: + continue + return + + +def main(argv): + if len(argv) != 3: + sys.stderr.write("fm-herdr-lab-viewer: usage: <session> <pidfile>\n") + return 2 + session, pidfile = argv[1:] + if session == "default" or not SESSION_PATTERN.match(session): + sys.stderr.write("fm-herdr-lab-viewer: refusing session %r\n" % session) + return 2 + if not os.path.isabs(pidfile): + sys.stderr.write("fm-herdr-lab-viewer: pidfile must be an absolute path\n") + return 2 + + try: + master, slave = os.openpty() + except OSError as error: + sys.stderr.write("fm-herdr-lab-viewer: could not create a pty: %s\n" % error) + return 3 + # Before the fork, so the TUI's first grid read already sees a real size. + fcntl.ioctl(master, termios.TIOCSWINSZ, struct.pack("HHHH", ROWS, COLS, 0, 0)) + + signal.pthread_sigmask(signal.SIG_BLOCK, TERMINATION_SIGNALS) + try: + viewer_pid = os.fork() + except OSError as error: + signal.pthread_sigmask(signal.SIG_UNBLOCK, TERMINATION_SIGNALS) + sys.stderr.write("fm-herdr-lab-viewer: could not fork the viewer: %s\n" % error) + return 3 + if viewer_pid == 0: + _child(slave, master, session) + + def _cancel_before_record(signum, _frame): + try: + os.kill(viewer_pid, signal.SIGKILL) + except OSError: + pass + os._exit(128 + signum) + + signal.signal(signal.SIGTERM, _cancel_before_record) + signal.signal(signal.SIGINT, _cancel_before_record) + signal.signal(signal.SIGHUP, _cancel_before_record) + signal.pthread_sigmask(signal.SIG_UNBLOCK, TERMINATION_SIGNALS) + + os.close(slave) + try: + _write_pidfile(pidfile, os.getpid(), viewer_pid) + except (OSError, RuntimeError, subprocess.SubprocessError) as error: + sys.stderr.write("fm-herdr-lab-viewer: could not record process identity: %s\n" % error) + try: + os.kill(viewer_pid, signal.SIGKILL) + except OSError: + pass + os.close(master) + try: + os.waitpid(viewer_pid, 0) + except OSError: + pass + return 3 + + def _signal_viewer(number): + # The viewer may already be gone; that is the outcome we wanted anyway. + try: + os.kill(viewer_pid, number) + except OSError: + pass + + def _terminate(_signum, _frame): + _signal_viewer(signal.SIGTERM) + signal.setitimer(signal.ITIMER_REAL, TERMINATE_GRACE_SECONDS) + + signal.signal(signal.SIGTERM, _terminate) + signal.signal(signal.SIGINT, _terminate) + signal.signal(signal.SIGHUP, _terminate) + signal.signal(signal.SIGALRM, lambda _s, _f: _signal_viewer(signal.SIGKILL)) + + _drain(master) + _terminate(None, None) + signal.setitimer(signal.ITIMER_REAL, TERMINATE_GRACE_SECONDS) + try: + _, status = os.waitpid(viewer_pid, 0) + except OSError: + status = 0 + signal.setitimer(signal.ITIMER_REAL, 0) + return 0 if os.WIFSIGNALED(status) else os.WEXITSTATUS(status) + + +if __name__ == "__main__": + sys.exit(main(sys.argv)) diff --git a/bin/fm-herdr-lab.sh b/bin/fm-herdr-lab.sh index f8ea014c6bc..d0aa633df55 100755 --- a/bin/fm-herdr-lab.sh +++ b/bin/fm-herdr-lab.sh @@ -7,6 +7,8 @@ # fm-herdr-lab.sh prepare <session> # fm-herdr-lab.sh provision <session> # fm-herdr-lab.sh run <session> <herdr arguments...> +# fm-herdr-lab.sh viewer start <session> +# fm-herdr-lab.sh viewer stop <session> # fm-herdr-lab.sh stop <session> # fm-herdr-lab.sh teardown <session> # @@ -23,6 +25,14 @@ # destructive call. # Provision records the running default session as a fleet-state tripwire and # teardown requires that record to be identical afterward. +# The viewer command attaches or detaches one real foreground Herdr client on +# an owned lab session over a fixed 40-row by 120-column pty; +# bin/fm-herdr-lab-viewer.py owns the pty mechanics. +# Start succeeds only when that session reports a foreground client and the +# recorded viewer process still matches its launch identity. +# Stop signals only identity-matched recorded processes and retains its +# ownership record until detach is confirmed or the session is stopped or +# absent; teardown refuses when that stop cannot be confirmed. set -u fm_herdr_lab_error() { @@ -153,6 +163,227 @@ fm_herdr_lab_cli() { # <session> <herdr arguments...> fm_herdr_lab_raw "$name" "$@" } +# --- foreground viewer ------------------------------------------------------ +# +# Herdr counts a client as the session's foreground viewer only once that +# client reports a usable window grid, so a zero-sized pty attaches nothing and +# leaves `terminal title clear` answering no_foreground_client. Attaching a +# real viewer is what lets a test drive the live-client teardown paths instead +# of only their detached halves. bin/fm-herdr-lab-viewer.py owns the pty and +# environment mechanics; the guards below own who may be attached to. +# Per-session locks are deliberately absent: generated fm-lab-<label>-$$-$RANDOM +# names have no caller that starts one viewer concurrently, so locks add risk. +# A subsecond interrupt window and SIGKILL residue are accepted in this isolated +# lab helper because teardown drops any stray viewer connection with the session. + +readonly fm_herdr_lab_viewer_timeout_seconds=5 +readonly fm_herdr_lab_viewer_launcher_grace_seconds=6 + +fm_herdr_lab_viewer_record_path() { # <session> + printf '%s/%s.viewer' "$(fm_herdr_lab_state_dir)" "$1" +} + +fm_herdr_lab_viewer_log_path() { # <session> + printf '%s/%s.viewer.log' "$(fm_herdr_lab_state_dir)" "$1" +} + +fm_herdr_lab_viewer_launcher_path() { + printf '%s/fm-herdr-lab-viewer.py' "$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +} + +# Prints the session's current foreground-client reason, or nothing when it +# cannot be read. +fm_herdr_lab_viewer_reason() { # <session> + local name=$1 out + out=$(fm_herdr_lab_cli "$name" terminal title clear 2>/dev/null) || return 1 + printf '%s' "$out" | jq -r '.result.reason // empty' 2>/dev/null +} + +fm_herdr_lab_process_start() { # <pid> + LC_ALL=C ps -p "$1" -o lstart= 2>/dev/null | sed 's/^[[:space:]]*//;s/[[:space:]]*$//' +} + +fm_herdr_lab_process_parent() { # <pid> + LC_ALL=C ps -p "$1" -o ppid= 2>/dev/null | sed 's/^[[:space:]]*//;s/[[:space:]]*$//' +} + +fm_herdr_lab_viewer_recorded_value() { # <session> <key> + local record value + record=$(fm_herdr_lab_viewer_record_path "$1") + [ -f "$record" ] || return 1 + value=$(sed -n "s/^$2=//p" "$record" | head -n 1) + [ -n "$value" ] || return 1 + printf '%s' "$value" +} + +fm_herdr_lab_viewer_owned_pair() { # <session> + local launcher_pid viewer_pid launcher_start viewer_start current_start parent_pid + launcher_pid=$(fm_herdr_lab_viewer_recorded_value "$1" launcher_pid) || return 1 + viewer_pid=$(fm_herdr_lab_viewer_recorded_value "$1" viewer_pid) || return 1 + case "$launcher_pid:$viewer_pid" in + *[!0-9:]*) return 1 ;; + esac + launcher_start=$(fm_herdr_lab_viewer_recorded_value "$1" launcher_start) || return 1 + viewer_start=$(fm_herdr_lab_viewer_recorded_value "$1" viewer_start) || return 1 + current_start=$(fm_herdr_lab_process_start "$launcher_pid") || return 1 + [ -n "$current_start" ] && [ "$current_start" = "$launcher_start" ] || return 1 + current_start=$(fm_herdr_lab_process_start "$viewer_pid") || return 1 + [ -n "$current_start" ] && [ "$current_start" = "$viewer_start" ] || return 1 + parent_pid=$(fm_herdr_lab_process_parent "$viewer_pid") || return 1 + [ "$parent_pid" = "$launcher_pid" ] || return 1 + printf '%s %s' "$launcher_pid" "$viewer_pid" +} + +fm_herdr_lab_viewer_owned_pid() { # <session> <launcher|viewer> + local pair + pair=$(fm_herdr_lab_viewer_owned_pair "$1") || return 1 + case "$2" in + launcher) printf '%s' "${pair%% *}" ;; + viewer) printf '%s' "${pair#* }" ;; + *) return 1 ;; + esac +} + +fm_herdr_lab_viewer_signal() { # <session> <launcher|viewer> <signal> + local pid + pid=$(fm_herdr_lab_viewer_owned_pid "$1" "$2") || return 0 + kill "-$3" "$pid" 2>/dev/null || true +} + +# True while this lab owns a viewer process that is still running. +fm_herdr_lab_viewer_owned_alive() { # <session> + fm_herdr_lab_viewer_owned_pair "$1" >/dev/null +} + +fm_herdr_lab_viewer_session_stopped_or_absent() { # <session> + local sessions running + sessions=$(fm_herdr_lab_session_list "$1" 2>/dev/null) || return 1 + running=$(printf '%s' "$sessions" | jq -r --arg name "$1" \ + '[.sessions[]? | select(.name == $name) | .running] | if length == 0 then "absent" elif length == 1 then .[0] else "ambiguous" end' \ + 2>/dev/null) || return 1 + [ "$running" = false ] || [ "$running" = absent ] +} + +fm_herdr_lab_viewer_start() { # <session> + local name=$1 record log launcher launcher_pid waited attempt reason pid interrupt_traps=0 timeout=$fm_herdr_lab_viewer_timeout_seconds + fm_herdr_lab_validate_name "$name" || return 1 + command -v herdr >/dev/null 2>&1 || { fm_herdr_lab_error "herdr is required"; return 1; } + command -v jq >/dev/null 2>&1 || { fm_herdr_lab_error "jq is required"; return 1; } + command -v python3 >/dev/null 2>&1 || { fm_herdr_lab_error "python3 is required for the lab viewer"; return 1; } + + [ -f "$(fm_herdr_lab_tripwire_path "$name")" ] || { + fm_herdr_lab_error "missing fleet-state tripwire for '$name'; refusing to attach a viewer to a session this lab does not own" + return 1 + } + fm_herdr_lab_refuse_if_default "$name" || return 1 + + record=$(fm_herdr_lab_viewer_record_path "$name") + if fm_herdr_lab_viewer_owned_alive "$name"; then + fm_herdr_lab_error "a lab viewer is already attached to '$name'; stop it before starting another" + return 1 + fi + rm -f "$record" + + launcher=$(fm_herdr_lab_viewer_launcher_path) + [ -f "$launcher" ] || { fm_herdr_lab_error "missing viewer launcher at $launcher"; return 1; } + log=$(fm_herdr_lab_viewer_log_path "$name") + mkdir -p "$(fm_herdr_lab_state_dir)" || return 1 + launcher_pid= + if [ "${BASH_SOURCE[0]}" = "$0" ]; then + interrupt_traps=1 + trap 'trap - INT TERM; [ -z "${launcher_pid:-}" ] || fm_herdr_lab_cancel_viewer_launcher "$launcher_pid"; exit 130' INT + trap 'trap - INT TERM; [ -z "${launcher_pid:-}" ] || fm_herdr_lab_cancel_viewer_launcher "$launcher_pid"; exit 143' TERM + fi + nohup python3 "$launcher" "$name" "$record" >"$log" 2>&1 & + launcher_pid=$! + + waited=0 + attempt=$((timeout * 5)) + while [ "$waited" -lt "$attempt" ]; do + reason=$(fm_herdr_lab_viewer_reason "$name") || reason= + if [ "$reason" = cleared ]; then + pid=$(fm_herdr_lab_viewer_owned_pid "$name" viewer) || pid= + if [ -n "$pid" ]; then + [ "$interrupt_traps" = 0 ] || trap - INT TERM + disown "$launcher_pid" 2>/dev/null || true + printf 'viewer attached to %s (pid %s)\n' "$name" "$pid" + return 0 + fi + fi + sleep 0.2 + waited=$((waited + 1)) + done + fm_herdr_lab_cancel_viewer_launcher "$launcher_pid" + [ "$interrupt_traps" = 0 ] || trap - INT TERM + fm_herdr_lab_error "lab viewer did not become the foreground client of '$name' within $timeout seconds (last reason: ${reason:-<unreadable>})" + [ ! -s "$log" ] || fm_herdr_lab_error "viewer log: $(tail -n 5 "$log" | tr '\n' ' ')" + fm_herdr_lab_viewer_stop "$name" >/dev/null 2>&1 || true + return 1 +} + +fm_herdr_lab_viewer_stop() { # <session> + local name=$1 record log role waited attempt reason timeout=$fm_herdr_lab_viewer_timeout_seconds + fm_herdr_lab_validate_name "$name" || return 1 + record=$(fm_herdr_lab_viewer_record_path "$name") + log=$(fm_herdr_lab_viewer_log_path "$name") + # An absent record means this lab owns no viewer. Any client attached in that + # case belongs to someone else and must never be signalled from here. + [ -f "$record" ] || return 0 + + for role in viewer launcher; do + fm_herdr_lab_viewer_signal "$name" "$role" TERM + done + waited=0 + while fm_herdr_lab_viewer_owned_alive "$name" && [ "$waited" -lt 50 ]; do + sleep 0.1 + waited=$((waited + 1)) + done + for role in viewer launcher; do + fm_herdr_lab_viewer_signal "$name" "$role" KILL + done + + waited=0 + attempt=$((timeout * 5)) + while [ "$waited" -lt "$attempt" ]; do + reason=$(fm_herdr_lab_viewer_reason "$name") || reason= + if [ "$reason" = no_foreground_client ] \ + || { [ -z "$reason" ] && fm_herdr_lab_viewer_session_stopped_or_absent "$name"; }; then + rm -f "$record" "$log" + return 0 + fi + sleep 0.2 + waited=$((waited + 1)) + done + fm_herdr_lab_error "lab viewer for '$name' did not detach within $timeout seconds (last reason: ${reason:-<unreadable>})" + return 1 +} + +fm_herdr_lab_viewer() { # <start|stop> <session> + case "${1:-}" in + start) fm_herdr_lab_viewer_start "$2" ;; + stop) fm_herdr_lab_viewer_stop "$2" ;; + *) + fm_herdr_lab_error "viewer takes 'start' or 'stop'" + return 2 + ;; + esac +} + +fm_herdr_lab_cancel_viewer_launcher() { # <pid> + local pid=$1 attempt=0 max_attempts=$((fm_herdr_lab_viewer_launcher_grace_seconds * 10)) + if kill -0 "$pid" 2>/dev/null; then + kill -TERM "$pid" 2>/dev/null || true + while kill -0 "$pid" 2>/dev/null && [ "$attempt" -lt "$max_attempts" ]; do + sleep 0.1 + attempt=$((attempt + 1)) + done + if kill -0 "$pid" 2>/dev/null; then + kill -KILL "$pid" 2>/dev/null || true + fi + fi + wait "$pid" 2>/dev/null || true +} + fm_herdr_lab_cancel_provision() { # <pid> local pid=$1 attempt=0 if kill -0 "$pid" 2>/dev/null; then @@ -261,6 +492,10 @@ fm_herdr_lab_teardown() { # <session> fm_herdr_lab_error "missing fleet-state tripwire for '$name'; refusing destructive calls" return 1 } + fm_herdr_lab_viewer_stop "$name" || { + fm_herdr_lab_error "refusing teardown of '$name' while this lab's viewer is still attached" + return 1 + } sessions=$(fm_herdr_lab_session_list "$name" 2>/dev/null) || { fm_herdr_lab_error "cannot list Herdr sessions before teardown" return 1 @@ -299,7 +534,7 @@ fm_herdr_lab_name() { # <label> } fm_herdr_lab_usage() { - sed -n '2,13p' "${BASH_SOURCE[0]}" | sed 's/^# \{0,1\}//' + sed -n '2,15p' "${BASH_SOURCE[0]}" | sed 's/^# \{0,1\}//' } fm_herdr_lab_main() { @@ -322,6 +557,10 @@ fm_herdr_lab_main() { shift fm_herdr_lab_cli "$@" ;; + viewer) + [ "$#" -eq 3 ] || { fm_herdr_lab_usage >&2; return 2; } + fm_herdr_lab_viewer "$2" "$3" + ;; stop) [ "$#" -eq 2 ] || { fm_herdr_lab_usage >&2; return 2; } fm_herdr_lab_stop "$2" diff --git a/bin/fm-test-run.sh b/bin/fm-test-run.sh index ab76526d8a2..2bdd8919a83 100755 --- a/bin/fm-test-run.sh +++ b/bin/fm-test-run.sh @@ -304,7 +304,7 @@ family_for_basename() { fm-backend-herdr-focus-flash-e2e.test.sh|\ fm-backend-herdr-stale-active-tab-e2e.test.sh|\ fm-backend-herdr-agent-exit-shell-e2e.test.sh|\ - fm-herdr-session-cleanup-e2e.test.sh|\ + fm-herdr-attached-viewer-live-e2e.test.sh|fm-herdr-session-cleanup-e2e.test.sh|\ fm-backend-herdr-smoke.test.sh|fm-backend-herdr-workspace-per-home-e2e.test.sh|\ fm-control-herdr-smoke.test.sh) printf '%s\n' real-herdr-gated @@ -491,7 +491,7 @@ tests/fm-composer-lib.test.sh 4798 tests/fm-crew-state.test.sh 11557 tests/fm-ensure-agents-md.test.sh 901 tests/fm-grok-harness.test.sh 6563 -tests/fm-herdr-lab.test.sh 6936 +tests/fm-herdr-lab.test.sh 9800 tests/fm-lint.test.sh 164262 tests/fm-pi-primary-types.test.sh 8624 tests/fm-pr-merge.test.sh 111145 @@ -696,6 +696,7 @@ tests/fm-guard-stale-banner.test.sh 32981 tests/fm-harness-adapter-instructions-live-e2e.test.sh 20 tests/fm-harness-adapter-references.test.sh 55 tests/fm-harness-liveness-drift-live-e2e.test.sh 21 +tests/fm-herdr-attached-viewer-live-e2e.test.sh 19000 tests/fm-herdr-session-cleanup.test.sh 6704 tests/fm-herdr-submit-confirm-live-e2e.test.sh 23 tests/fm-herdr-version-floor-live-e2e.test.sh 23 diff --git a/docs/herdr-backend.md b/docs/herdr-backend.md index 316643fdda2..97523457071 100644 --- a/docs/herdr-backend.md +++ b/docs/herdr-backend.md @@ -185,6 +185,7 @@ Operational compromises: `tests/fm-herdr-session-cleanup-e2e.test.sh` covers the restored-shell cleanup in a guarded non-default named lab. `tests/fm-backend-herdr-focus-flash-e2e.test.sh` reproduces the raw explicit-close focus steal on the installed release and proves the focus-safe emptying-close plan removes a doomed workspace with no wrong-focus interval; [`verification/runtime-backends.md`](verification/runtime-backends.md#workspace-removal-focus-safety) owns the active versioned evidence. `tests/fm-backend-herdr-stale-active-tab-e2e.test.sh` proves a persisted-focused tab still closes when no foreground client is attached. +`tests/fm-herdr-attached-viewer-live-e2e.test.sh` proves the other half against a real attached viewer, which `bin/fm-herdr-lab.sh viewer start` supplies over a pty sized before the fork; [`verification/runtime-backends.md`](verification/runtime-backends.md#attached-foreground-viewer) owns the active versioned evidence and the re-run trigger. ## Default-tab prune safety @@ -369,6 +370,7 @@ tests/fm-backend-herdr-eventwait-smoke.test.sh tests/fm-control-herdr-smoke.test.sh tests/fm-herdr-session-cleanup.test.sh tests/fm-herdr-session-cleanup-e2e.test.sh +tests/fm-herdr-attached-viewer-live-e2e.test.sh tests/fm-afk-inject-herdr-e2e.test.sh tests/fm-afk-pi-herdr-return-e2e.test.sh ``` diff --git a/docs/scripts.md b/docs/scripts.md index 9b149e538c7..4f84c42dcce 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -35,6 +35,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-brief.sh` | Scaffold ship (explicit `--mode`), scout, secondmate-charter, and Herdr-lab briefs, with Captain's intent and Firstmate spec subsections on ship/scout | | [`fm-dod-lib.sh`](../bin/fm-dod-lib.sh) | Own ship/scout worker role scope, ship definitions of done, and the no-mistakes `--intent` contract | | `fm-herdr-lab.sh` | Provision and guardedly operate an isolated, never-default Herdr lab session | +| `fm-herdr-lab-viewer.py` | The pty engine behind `fm-herdr-lab.sh viewer`: one real foreground Herdr client on a non-zero window grid | | `fm-install-herdr.sh` | Install CI's exact-version Herdr pin with official asset URL, SHA-256, and protocol checks | | `fm-install-treehouse.sh`| Install CI's exact-version Treehouse pin for real-Herdr E2E that needs spawn worktrees | | `fm-herdr-ci-cleanup.sh` | Snapshot and tear down only job-owned `fm-lab-*` sessions in the Herdr CI lane | diff --git a/docs/verification/runtime-backends.md b/docs/verification/runtime-backends.md index 200a5974131..9d4d8d2cd21 100644 --- a/docs/verification/runtime-backends.md +++ b/docs/verification/runtime-backends.md @@ -884,6 +884,38 @@ Part C is the case the suite could not reach before: a doomed pane whose shell h On 0.7.5 that fallback exposed a bounded four-sample wrong-focus window and restored the anchor exactly; on 0.8.0 the same fallback exposed none, which is why default-on projection is floored at 0.8.0 rather than mitigated further below it. The suite also cross-checks its own Part A measurement against the floor classifier on whatever release it runs, so a drifted protocol-to-release mapping fails there rather than silently gating on the wrong thing. +### Attached foreground viewer + +A pseudo-terminal registers as a Herdr foreground client only when its window grid is non-zero. +`script` and a bare `pty.fork()` from a non-tty parent both start at 0x0, which is why PR #4131 could validate only the detached half of the teardown focus guard and left its four attached-client scenarios untested. +The guarded `viewer start` path fixes the pty at the proven 40-row by 120-column grid, sets that size on the master fd before the fork, and scrubs inherited `HERDR_*` variables, which makes the attached scenarios reachable from a headless runner. + +Measured on 2026-09-11 against Herdr 0.9.0 protocol 22 on macOS 26.5.2 aarch64 with Python 3.14.6: + +```sh +HERDR_LAB_HELPER=bin/fm-herdr-lab.sh \ + tests/fm-herdr-attached-viewer-live-e2e.test.sh +``` + +```text +ok - attached viewer: a pty sized before the fork registers as a real Herdr foreground client +ok - attached viewer: a live client on the target tab refuses the close and keeps the pane +ok - attached viewer: focus moving onto the target between planning and mutation still blocks the close +ok - attached viewer: a close preserves the fresh non-target focus the viewer moved to +ok - attached viewer: the projection seeded-tab prune refuses while a live client watches it +ok - attached viewer: detaching releases the refusal, so the guard tracks the client and not the pointer +``` + +Both halves of the recipe are load-bearing, and each was measured by removing it from the helper and re-running the guard on the same host and release. +Dropping the `TIOCSWINSZ` call and dropping the environment scrub each left startup reporting `no_foreground_client`, followed by the guard failure: + +```text +not ok - could not attach a real foreground Herdr viewer over a sized pty +``` + +Re-run this guard after every Herdr upgrade. +A release that changed the foreground-client contract, the window-grid requirement, or the nested-viewer refusal would fail here first, and the detached regressions would keep passing while saying nothing about it. + ### Presentation version floor Default-on presentation projection is floored at Herdr 0.8.0. diff --git a/tests/fm-herdr-attached-viewer-live-e2e.test.sh b/tests/fm-herdr-attached-viewer-live-e2e.test.sh new file mode 100755 index 00000000000..061b227de50 --- /dev/null +++ b/tests/fm-herdr-attached-viewer-live-e2e.test.sh @@ -0,0 +1,259 @@ +#!/usr/bin/env bash +# Live attached-viewer regression for the Herdr teardown focus guard. +# +# PR #4131 gated the active-tab close refusal on a LIVE foreground client +# instead of the persisted `.focused` pointer, but only its two detached +# scenarios could be driven live: every pseudo-terminal the runner built +# started at a zero-sized window grid, so Herdr registered no foreground client +# and `terminal title clear` kept answering `no_foreground_client`. That was a +# harness limit, not a product one. `fm-herdr-lab.sh viewer start` now attaches +# a real Herdr TUI over a pty sized before the fork, which turns those +# untestable cases into this regression: +# +# 3. a viewer sitting on the target tab blocks the close; +# 4. a viewer that moves ONTO the target between planning and the mutation +# boundary still blocks it, because the guard re-reads focus there; +# 5. a viewer that moves OFF the target instead keeps its fresh non-target +# focus after the close, rather than being dragged back to a stale +# pre-planning pointer; +# 7. the projection's seeded-tab prune inherits the same refusal. +# +# Scenarios 4 and 5 need a focus change at one exact product boundary, so a +# PATH shim performs the real `tab focus` when the close helper issues its +# planning `pane get`. Every Herdr call, the shim's included, still routes +# through the guarded lab helper against a named non-default session. +# +# The guard submits no model prompts, so the shared live gate runs it wherever +# herdr, jq, and python3 exist. Re-run it after every Herdr upgrade: a release +# that changed the foreground-client contract would surface here first. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +LAB_HELPER=${HERDR_LAB_HELPER:-$ROOT/bin/fm-herdr-lab.sh} + +fm_live_gate default-on FM_HERDR_ATTACHED_VIEWER_LIVE_E2E herdr jq python3 + +[ -x "$LAB_HELPER" ] || { echo "skip: Herdr lab helper not executable at $LAB_HELPER"; exit 0; } + +TMP_ROOT=$(fm_test_tmproot fm-herdr-attached-viewer) +FAKEBIN=$(fm_fakebin "$TMP_ROOT") +FOCUS_SWITCH_CONTROL="$TMP_ROOT/focus-switch" +ORIGINAL_PATH=$PATH +LAB_SESSION=$("$LAB_HELPER" name fm-herdr-attached-viewer) +export LAB_HELPER LAB_SESSION ORIGINAL_PATH FOCUS_SWITCH_CONTROL + +cleanup() { + local status=$? + env PATH="$ORIGINAL_PATH" "$LAB_HELPER" viewer stop "$LAB_SESSION" >/dev/null 2>&1 || status=1 + env PATH="$ORIGINAL_PATH" "$LAB_HELPER" teardown "$LAB_SESSION" || status=1 + fm_test_cleanup + exit "$status" +} +trap cleanup EXIT +"$LAB_HELPER" provision "$LAB_SESSION" || fail "could not provision the isolated Herdr lab" + +lab() { env PATH="$ORIGINAL_PATH" "$LAB_HELPER" run "$LAB_SESSION" "$@"; } + +# The adapter under test calls `herdr` by name. This shim strips the trailing +# session flag the helper will re-append, refuses any caller-supplied one, and +# optionally performs one real focus change at the requested product boundary +# before forwarding the call. +cat > "$FAKEBIN/herdr" <<'SH' +#!/usr/bin/env bash +set -u +args=("$@") +last=$((${#args[@]} - 1)) +flag=$((last - 1)) +if [ "${#args[@]}" -ge 2 ] \ + && [ "${args[$flag]}" = --session ] \ + && [ "${args[$last]}" = "$LAB_SESSION" ]; then + unset "args[$last]" "args[$flag]" +fi +set -- "${args[@]}" +for arg in "$@"; do + case "$arg" in --session|--session=*) exit 9 ;; esac +done +if [ -f "$FOCUS_SWITCH_CONTROL/trigger" ] && [ "$*" = "$(cat "$FOCUS_SWITCH_CONTROL/trigger")" ]; then + rm -f "$FOCUS_SWITCH_CONTROL/trigger" + env PATH="$ORIGINAL_PATH" "$LAB_HELPER" run "$LAB_SESSION" \ + tab focus "$(cat "$FOCUS_SWITCH_CONTROL/tab")" >/dev/null 2>&1 + printf '%s\n' switched > "$FOCUS_SWITCH_CONTROL/done" +fi +exec env PATH="$ORIGINAL_PATH" "$LAB_HELPER" run "$LAB_SESSION" "$@" +SH +chmod +x "$FAKEBIN/herdr" + +mkdir -p "$FOCUS_SWITCH_CONTROL" + +# Arm the shim to run `tab focus <tab>` immediately before the adapter's own +# `pane get <pane>` planning read, which is the last product call before the +# close helper re-reads focus at its mutation boundary. +arm_focus_switch() { # <tab-id> <pane-id> + rm -f "$FOCUS_SWITCH_CONTROL/done" + printf '%s\n' "$1" > "$FOCUS_SWITCH_CONTROL/tab" + printf 'pane get %s\n' "$2" > "$FOCUS_SWITCH_CONTROL/trigger" +} + +assert_focus_switch_fired() { # <label> + [ -f "$FOCUS_SWITCH_CONTROL/done" ] \ + || fail "$1: the mid-close focus switch never ran, so the timing boundary was not exercised" + rm -f "$FOCUS_SWITCH_CONTROL/done" "$FOCUS_SWITCH_CONTROL/trigger" +} + +# Drive one real adapter entry point with the shim on PATH. +drive() { # <function> <argument...> + PATH="$FAKEBIN:$ORIGINAL_PATH" bash -c ' + . "$1/bin/backends/herdr.sh" + fm_backend_herdr_cli() { + local session=$1 + shift + HERDR_SESSION="$session" herdr "$@" --session "$session" + } + fn=$2 + shift 2 + "$fn" "$@" + ' _ "$ROOT" "$@" 2>&1 +} + +# These run inside command substitutions, where `fail` would exit only the +# subshell and let the script carry on with empty ids. They return non-zero +# instead, and every call site carries its own `|| fail`. +new_workspace() { # <label> -> "<workspace>\t<tab>\t<pane>" + local out + out=$(lab workspace create --cwd "$ROOT" --label "$1" --no-focus) || return 1 + printf '%s' "$out" | jq -er ' + [.result.workspace.workspace_id, .result.tab.tab_id, .result.root_pane.pane_id] | @tsv + ' +} + +new_tab() { # <workspace> <label> -> "<tab>\t<pane>" + local out + out=$(lab tab create --workspace "$1" --label "$2" --cwd "$ROOT") || return 1 + printf '%s' "$out" | jq -er '[.result.tab.tab_id, .result.root_pane.pane_id] | @tsv' +} + +focused_tab() { + lab workspace list | jq -er ' + [.result.workspaces[] | select(.focused == true)] | select(length == 1) | .[0].active_tab_id + ' +} + +pane_exists() { lab pane get "$1" >/dev/null 2>&1; } + +foreground_reason() { + lab terminal title clear | jq -er '.result.reason' +} + +# --- the attachment itself, which is what #4131 could not do ---------------- + +REASON=$(foreground_reason) || fail "could not probe the session's foreground client" +[ "$REASON" = no_foreground_client ] \ + || fail "the fresh lab already had a foreground client (reason=$REASON)" + +"$LAB_HELPER" viewer start "$LAB_SESSION" >/dev/null \ + || fail "could not attach a real foreground Herdr viewer over a sized pty" +REASON=$(foreground_reason) || fail "could not probe the session's foreground client" +[ "$REASON" = cleared ] \ + || fail "the attached pty viewer did not register as a foreground client (reason=$REASON)" +pass "attached viewer: a pty sized before the fork registers as a real Herdr foreground client" + +# --- scenario 3: a viewer on the target tab blocks the close --------------- + +FIXTURE=$(new_workspace viewer-active) || fail "could not create the scenario 3 workspace" +IFS=$'\t' read -r WS_THREE TAB_THREE_A PANE_THREE_A <<<"$FIXTURE" +# A second tab keeps the close a plain one rather than an emptying-workspace plan. +new_tab "$WS_THREE" viewer-active-b >/dev/null || fail "could not create the scenario 3 companion tab" +lab tab focus "$TAB_THREE_A" >/dev/null || fail "could not focus the scenario 3 target tab" +[ "$(focused_tab)" = "$TAB_THREE_A" ] || fail "scenario 3 did not start focused on the target tab" + +OUT=$(drive fm_backend_herdr_projection_close_pane_focus_preserving "$LAB_SESSION" "$PANE_THREE_A") +STATUS=$? +[ "$STATUS" -ne 0 ] || fail "a live viewer on the target tab did not block the close: $OUT" +assert_contains "$OUT" "target is the captain's active tab" \ + "the live-viewer refusal did not name the captain's active tab: $OUT" +pane_exists "$PANE_THREE_A" \ + || fail "the close proceeded and destroyed the tab the live viewer was watching" +pass "attached viewer: a live client on the target tab refuses the close and keeps the pane" + +# --- scenario 4: the viewer moves ONTO the target mid-close ---------------- + +FIXTURE=$(new_workspace viewer-late-on) || fail "could not create the scenario 4 workspace" +IFS=$'\t' read -r WS_FOUR TAB_FOUR_A PANE_FOUR_A <<<"$FIXTURE" +FIXTURE=$(new_tab "$WS_FOUR" viewer-late-on-b) || fail "could not create the scenario 4 companion tab" +IFS=$'\t' read -r TAB_FOUR_B _ <<<"$FIXTURE" +lab tab focus "$TAB_FOUR_B" >/dev/null || fail "could not focus away from the scenario 4 target" +[ "$(focused_tab)" = "$TAB_FOUR_B" ] || fail "scenario 4 did not start focused off the target tab" + +arm_focus_switch "$TAB_FOUR_A" "$PANE_FOUR_A" +OUT=$(drive fm_backend_herdr_projection_close_pane_focus_preserving "$LAB_SESSION" "$PANE_FOUR_A") +STATUS=$? +assert_focus_switch_fired "scenario 4" +[ "$STATUS" -ne 0 ] \ + || fail "a viewer that moved onto the target after planning did not block the close: $OUT" +assert_contains "$OUT" "target is the captain's active tab" \ + "the late-switch refusal did not come from the fresh active-tab check: $OUT" +pane_exists "$PANE_FOUR_A" \ + || fail "the close destroyed a tab the viewer had moved onto before the mutation boundary" +pass "attached viewer: focus moving onto the target between planning and mutation still blocks the close" + +# --- scenario 5: the viewer moves OFF the target mid-close ----------------- + +FIXTURE=$(new_workspace viewer-late-off) || fail "could not create the scenario 5 workspace" +IFS=$'\t' read -r WS_FIVE TAB_FIVE_A PANE_FIVE_A <<<"$FIXTURE" +FIXTURE=$(new_tab "$WS_FIVE" viewer-late-off-b) || fail "could not create the scenario 5 companion tab" +IFS=$'\t' read -r TAB_FIVE_B _ <<<"$FIXTURE" +lab tab focus "$TAB_FIVE_A" >/dev/null || fail "could not focus the scenario 5 target tab" +[ "$(focused_tab)" = "$TAB_FIVE_A" ] || fail "scenario 5 did not start focused on the target tab" + +arm_focus_switch "$TAB_FIVE_B" "$PANE_FIVE_A" +OUT=$(drive fm_backend_herdr_projection_close_pane_focus_preserving "$LAB_SESSION" "$PANE_FIVE_A") +STATUS=$? +assert_focus_switch_fired "scenario 5" +[ "$STATUS" -eq 0 ] \ + || fail "the close was refused even though the live viewer had moved off the target: $OUT" +if pane_exists "$PANE_FIVE_A"; then + fail "the close reported success but left the target pane behind" +fi +[ "$(focused_tab)" = "$TAB_FIVE_B" ] \ + || fail "the close did not preserve the viewer's fresh non-target focus (focus is $(focused_tab), expected $TAB_FIVE_B)" +pass "attached viewer: a close preserves the fresh non-target focus the viewer moved to" + +# --- scenario 7: the projection's seeded-tab prune inherits the refusal ---- + +FIXTURE=$(new_workspace viewer-seeded) || fail "could not create the scenario 7 workspace" +IFS=$'\t' read -r WS_SEVEN TAB_SEVEN_SEEDED PANE_SEVEN_SEEDED <<<"$FIXTURE" +FIXTURE=$(new_tab "$WS_SEVEN" fm-viewer-seeded-task) || fail "could not create the scenario 7 task tab" +IFS=$'\t' read -r _ PANE_SEVEN_TASK <<<"$FIXTURE" +lab tab list --workspace "$WS_SEVEN" \ + | jq -e --arg tab "$TAB_SEVEN_SEEDED" '.result.tabs[] | select(.tab_id == $tab) | .label == "1"' >/dev/null \ + || fail "the seeded tab is not the label-1 default tab the prune identifies" +lab tab focus "$TAB_SEVEN_SEEDED" >/dev/null || fail "could not focus the seeded tab" +[ "$(focused_tab)" = "$TAB_SEVEN_SEEDED" ] || fail "scenario 7 did not start focused on the seeded tab" + +OUT=$(drive fm_backend_herdr_workspace_prune_seeded_default_tab \ + "$LAB_SESSION" "$WS_SEVEN" "$TAB_SEVEN_SEEDED" focus-preserving) +STATUS=$? +[ "$STATUS" -ne 0 ] || fail "the seeded prune did not refuse the tab a live viewer was watching: $OUT" +assert_contains "$OUT" "target is the captain's active tab" \ + "the seeded prune refusal did not come from the live-viewer guard: $OUT" +pane_exists "$PANE_SEVEN_SEEDED" \ + || fail "the seeded prune closed the tab the live viewer was watching" +pane_exists "$PANE_SEVEN_TASK" || fail "the seeded prune disturbed the task pane" +pass "attached viewer: the projection seeded-tab prune refuses while a live client watches it" + +# --- detaching restores the no-client contract the detached tests rely on --- + +"$LAB_HELPER" viewer stop "$LAB_SESSION" >/dev/null \ + || fail "could not detach the lab viewer" +REASON=$(foreground_reason) || fail "could not probe the session's foreground client" +[ "$REASON" = no_foreground_client ] \ + || fail "the lab still reported a foreground client after the viewer stopped (reason=$REASON)" +OUT=$(drive fm_backend_herdr_projection_close_pane_focus_preserving "$LAB_SESSION" "$PANE_THREE_A") +STATUS=$? +[ "$STATUS" -eq 0 ] || fail "the same close was still refused after the viewer detached: $OUT" +if pane_exists "$PANE_THREE_A"; then + fail "the detached close reported success but left the pane behind" +fi +pass "attached viewer: detaching releases the refusal, so the guard tracks the client and not the pointer" diff --git a/tests/fm-herdr-lab.test.sh b/tests/fm-herdr-lab.test.sh index 14ab7497a09..474b3f3e87e 100755 --- a/tests/fm-herdr-lab.test.sh +++ b/tests/fm-herdr-lab.test.sh @@ -64,6 +64,12 @@ case "$1 ${2:-}" in [ "${FM_FAKE_HERDR_DELETE_FAIL:-}" != 1 ] || exit 93 printf '%s\n' deleted > "$state/$session" ;; + "terminal title") + [ "${FM_FAKE_HERDR_TITLE_FAIL:-}" != 1 ] || exit 94 + reason=no_foreground_client + [ ! -f "$state/$session.foreground" ] || reason=$(cat "$state/$session.foreground") + jq -nc --arg reason "$reason" '{result:{reason:$reason,type:"client_window_title"}}' + ;; *) printf '%s\n' '{"ok":true}' ;; @@ -82,6 +88,7 @@ run_with_fake() { FM_FAKE_HERDR_SERVER_DELAY="${FM_FAKE_HERDR_SERVER_DELAY:-0}" \ FM_FAKE_HERDR_FAST_POLL="${FM_FAKE_HERDR_FAST_POLL:-}" \ FM_FAKE_HERDR_DELETE_FAIL="${FM_FAKE_HERDR_DELETE_FAIL:-}" \ + FM_FAKE_HERDR_TITLE_FAIL="${FM_FAKE_HERDR_TITLE_FAIL:-}" \ FM_HERDR_LAB_STATE_DIR="$TRIPWIRES" \ "$@" } @@ -213,6 +220,9 @@ test_timed_out_provision_cancels_late_launch() { cat > "$FAKEBIN/sleep" <<'SH' #!/usr/bin/env bash if [ "${FM_FAKE_HERDR_FAST_POLL:-}" = 1 ]; then + while [ -n "${FM_FAKE_HERDR_WAIT_MARKER:-}" ] && [ ! -f "$FM_FAKE_HERDR_WAIT_MARKER" ]; do + "$FM_FAKE_HERDR_REAL_SLEEP" 0.01 + done exit 0 fi exec "$FM_FAKE_HERDR_REAL_SLEEP" "$@" @@ -234,6 +244,260 @@ SH pass "fm-herdr-lab: timed-out provisioning cancels the launch before teardown" } + +# The pty attachment itself needs a real Herdr client, so the live guard +# tests/fm-herdr-attached-viewer-live-e2e.test.sh owns that proof. What is +# portable is who the helper will ever attach to, and who it will signal. +test_viewer_refuses_unowned_sessions() { + local name="fm-lab-viewer-guard-$$" status=0 out + : > "$FAKE_LOG" + out=$(run_with_fake fm_herdr_lab_viewer_start "$name" 2>&1) || status=$? + expect_code 1 "$status" "a session without an ownership tripwire must not be attached to" + assert_contains "$out" "does not own" \ + "the viewer refusal did not name the missing ownership record" + [ ! -s "$FAKE_LOG" ] \ + || fail "the unowned-session refusal reached Herdr instead of refusing first" + + status=0 + run_with_fake fm_herdr_lab_viewer_start default >/dev/null 2>&1 || status=$? + expect_code 1 "$status" "the default session must never be attached to" + pass "fm-herdr-lab: the viewer attaches only to a session this lab owns" +} + +start_viewer_fixture() { + local pair=$1 + ( + "$REAL_SLEEP" 20 & + printf '%s\n' "$!" > "$pair" + wait + ) & + FIXTURE_LAUNCHER_PID=$! + while [ ! -s "$pair" ]; do + "$REAL_SLEEP" 0.01 + done + FIXTURE_VIEWER_PID=$(cat "$pair") +} + +write_viewer_record() { + local record=$1 launcher_pid=$2 viewer_pid=$3 launcher_start viewer_start + launcher_start=$(fm_herdr_lab_process_start "$launcher_pid") || fail "could not identify launcher fixture process" + viewer_start=$(fm_herdr_lab_process_start "$viewer_pid") || fail "could not identify viewer fixture process" + printf 'launcher_pid=%s\nlauncher_start=%s\nviewer_pid=%s\nviewer_start=%s\n' \ + "$launcher_pid" "$launcher_start" "$viewer_pid" "$viewer_start" > "$record" +} + +test_viewer_start_cancels_an_unrecorded_launcher() { + local name="fm-lab-viewer-late-$$" out status=0 launcher_pid + local started="$TMP_ROOT/viewer-launcher-started" + run_with_fake fm_herdr_lab_provision "$name" || fail "viewer-late fixture provision failed" + cat > "$FAKEBIN/python3" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$$" > "$FM_FAKE_VIEWER_STARTED" +exec "$FM_FAKE_HERDR_REAL_SLEEP" 20 +SH + chmod +x "$FAKEBIN/python3" + out=$(FM_FAKE_HERDR_FAST_POLL=1 FM_FAKE_HERDR_WAIT_MARKER="$started" \ + FM_FAKE_VIEWER_STARTED="$started" run_with_fake fm_herdr_lab_viewer_start "$name" 2>&1) || status=$? + rm -f "$FAKEBIN/python3" + expect_code 1 "$status" "an unrecorded launcher must not outlive viewer start" + assert_present "$started" "delayed viewer launcher did not start" + launcher_pid=$(cat "$started") + kill -0 "$launcher_pid" 2>/dev/null && fail "timed-out viewer launcher remained alive" + assert_contains "$out" "did not become the foreground client" "launcher timeout was unclear" + run_with_fake fm_herdr_lab_teardown "$name" || fail "viewer-late fixture teardown failed" + pass "fm-herdr-lab: timed-out viewer startup cancels its exact launcher" +} + +test_viewer_timeout_allows_launcher_escalation() { + local launcher_pid started="$TMP_ROOT/viewer-grace-started" + local terminating="$TMP_ROOT/viewer-grace-terminating" completed="$TMP_ROOT/viewer-grace-completed" + cat > "$FAKEBIN/viewer-launcher" <<'SH' +#!/usr/bin/env bash +trap 'printf "" > "$FM_FAKE_VIEWER_TERMINATING"; "$FM_FAKE_HERDR_REAL_SLEEP" 1.2; printf "" > "$FM_FAKE_VIEWER_COMPLETED"; exit 0' TERM +printf '' > "$FM_FAKE_VIEWER_STARTED" +while :; do + "$FM_FAKE_HERDR_REAL_SLEEP" 0.1 +done +SH + chmod +x "$FAKEBIN/viewer-launcher" + FM_FAKE_HERDR_REAL_SLEEP="$REAL_SLEEP" FM_FAKE_VIEWER_STARTED="$started" \ + FM_FAKE_VIEWER_TERMINATING="$terminating" FM_FAKE_VIEWER_COMPLETED="$completed" \ + "$FAKEBIN/viewer-launcher" & + launcher_pid=$! + while [ ! -f "$started" ]; do + "$REAL_SLEEP" 0.01 + done + run_with_fake fm_herdr_lab_cancel_viewer_launcher "$launcher_pid" + assert_present "$terminating" "timed-out viewer launcher did not receive TERM" + assert_present "$completed" "viewer launcher was killed before completing child escalation" + pass "fm-herdr-lab: startup timeout allows launcher child escalation" +} + +test_viewer_start_requires_its_owned_process() { + local name="fm-lab-viewer-ownership-$$" out status=0 marker="$TMP_ROOT/viewer-launched" + run_with_fake fm_herdr_lab_provision "$name" || fail "viewer-ownership fixture provision failed" + printf '%s\n' cleared > "$FAKE_STATE/$name.foreground" + cat > "$FAKEBIN/python3" <<'SH' +#!/usr/bin/env bash +: > "$FM_FAKE_VIEWER_MARKER" +exit 0 +SH + chmod +x "$FAKEBIN/python3" + out=$(FM_FAKE_HERDR_FAST_POLL=1 FM_FAKE_VIEWER_MARKER="$marker" \ + run_with_fake fm_herdr_lab_viewer_start "$name" 2>&1) || status=$? + rm -f "$FAKEBIN/python3" + expect_code 1 "$status" "a foreign foreground client must not satisfy viewer start" + assert_present "$marker" "viewer ownership fixture did not launch" + assert_contains "$out" "did not become the foreground client" "ownership failure did not time out clearly" + assert_not_contains "$out" "viewer attached" "start claimed a foreign foreground client as its own" + printf '%s\n' no_foreground_client > "$FAKE_STATE/$name.foreground" + run_with_fake fm_herdr_lab_teardown "$name" || fail "viewer-ownership fixture teardown failed" + pass "fm-herdr-lab: viewer start requires an identity-matched owned process" +} + +test_viewer_stop_only_signals_owned_processes() { + local name="fm-lab-viewer-stop-$$" record status=0 holder_pid pair="$TMP_ROOT/viewer-stop-pair" + run_with_fake fm_herdr_lab_provision "$name" || fail "viewer-stop fixture provision failed" + record=$(run_with_fake fm_herdr_lab_viewer_record_path "$name") + + # No record: a client attached by someone else is not ours to kill. + printf '%s\n' cleared > "$FAKE_STATE/$name.foreground" + run_with_fake fm_herdr_lab_viewer_stop "$name" \ + || fail "stopping with no recorded viewer must succeed without touching a foreign client" + [ "$(cat "$FAKE_STATE/$name.foreground")" = cleared ] \ + || fail "an unrecorded foreground client was detached by the lab helper" + + # A recorded viewer is signalled until it exits and the session reports no + # foreground client again. + start_viewer_fixture "$pair" + write_viewer_record "$record" "$FIXTURE_LAUNCHER_PID" "$FIXTURE_VIEWER_PID" + status=0 + FM_FAKE_HERDR_FAST_POLL=1 run_with_fake fm_herdr_lab_viewer_stop "$name" \ + >/dev/null 2>&1 || status=$? + expect_code 1 "$status" "stop must fail while the session still reports a foreground client" + wait "$FIXTURE_LAUNCHER_PID" 2>/dev/null || true + kill -0 "$FIXTURE_VIEWER_PID" 2>/dev/null && fail "stop left the recorded viewer process running" + assert_present "$record" "a failed detach discarded the viewer record it still needs" + + sleep 20 & + holder_pid=$! + printf 'launcher_pid=%s\nlauncher_start=not-this-process\nviewer_pid=%s\nviewer_start=not-this-process\n' \ + "$holder_pid" "$holder_pid" > "$record" + printf '%s\n' no_foreground_client > "$FAKE_STATE/$name.foreground" + run_with_fake fm_herdr_lab_viewer_stop "$name" || fail "stop rejected a stale process record" + kill -0 "$holder_pid" 2>/dev/null || fail "stop signalled a PID whose recorded identity did not match" + kill "$holder_pid" 2>/dev/null || true + wait "$holder_pid" 2>/dev/null || true + + run_with_fake fm_herdr_lab_viewer_stop "$name" || fail "stop failed once the client had detached" + assert_absent "$record" "a confirmed detach left the viewer record behind" + run_with_fake fm_herdr_lab_teardown "$name" || fail "teardown after viewer stop failed" + pass "fm-herdr-lab: viewer stop signals only recorded processes and confirms the detach" +} + +test_viewer_stop_requires_the_recorded_parent() { + local name="fm-lab-viewer-parent-$$" record launcher_pid viewer_pid + run_with_fake fm_herdr_lab_provision "$name" || fail "viewer-parent fixture provision failed" + record=$(run_with_fake fm_herdr_lab_viewer_record_path "$name") + sleep 20 & + launcher_pid=$! + sleep 20 & + viewer_pid=$! + write_viewer_record "$record" "$launcher_pid" "$viewer_pid" + printf '%s\n' no_foreground_client > "$FAKE_STATE/$name.foreground" + run_with_fake fm_herdr_lab_viewer_stop "$name" || fail "parent-mismatch stop failed" + kill -0 "$launcher_pid" 2>/dev/null || fail "stop signalled a launcher without its recorded child" + kill -0 "$viewer_pid" 2>/dev/null || fail "stop signalled a viewer outside the recorded launcher" + kill "$launcher_pid" "$viewer_pid" 2>/dev/null || true + wait "$launcher_pid" 2>/dev/null || true + wait "$viewer_pid" 2>/dev/null || true + run_with_fake fm_herdr_lab_teardown "$name" || fail "viewer-parent fixture teardown failed" + pass "fm-herdr-lab: viewer ownership requires the recorded parent" +} + +test_interrupted_viewer_start_cancels_launcher() { + local name="fm-lab-viewer-interrupt-$$" command_pid launcher_pid status=0 + local started="$TMP_ROOT/viewer-interrupt-started" attached="$TMP_ROOT/viewer-interrupt-attached" + run_with_fake fm_herdr_lab_provision "$name" || fail "viewer-interrupt fixture provision failed" + cat > "$FAKEBIN/python3" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$$" > "$FM_FAKE_VIEWER_STARTED" +"$FM_FAKE_HERDR_REAL_SLEEP" 0.5 +: > "$FM_FAKE_VIEWER_ATTACHED" +printf '%s\n' cleared > "$FM_FAKE_HERDR_STATE/$FM_FAKE_VIEWER_SESSION.foreground" +exec "$FM_FAKE_HERDR_REAL_SLEEP" 20 +SH + chmod +x "$FAKEBIN/python3" + FM_FAKE_VIEWER_STARTED="$started" FM_FAKE_VIEWER_ATTACHED="$attached" \ + FM_FAKE_VIEWER_SESSION="$name" run_with_fake exec "$ROOT/bin/fm-herdr-lab.sh" \ + viewer start "$name" >/dev/null 2>&1 & + command_pid=$! + while [ ! -f "$started" ]; do + "$REAL_SLEEP" 0.01 + done + launcher_pid=$(cat "$started") + kill -TERM "$command_pid" + wait "$command_pid" || status=$? + rm -f "$FAKEBIN/python3" + [ "$status" -ne 0 ] || fail "interrupted viewer start unexpectedly succeeded" + "$REAL_SLEEP" 0.6 + kill -0 "$launcher_pid" 2>/dev/null && fail "interrupted viewer start left its launcher running" + assert_absent "$attached" "interrupted viewer start attached after its command exited" + [ ! -f "$FAKE_STATE/$name.foreground" ] || fail "interrupted viewer start left a foreground client" + run_with_fake fm_herdr_lab_teardown "$name" || fail "viewer-interrupt fixture teardown failed" + pass "fm-herdr-lab: interrupted viewer start cancels its launcher" +} + +test_teardown_refuses_while_viewer_attached() { + local name="fm-lab-viewer-teardown-$$" record status=0 pair="$TMP_ROOT/viewer-teardown-pair" + run_with_fake fm_herdr_lab_provision "$name" || fail "viewer-teardown fixture provision failed" + record=$(run_with_fake fm_herdr_lab_viewer_record_path "$name") + printf '%s\n' cleared > "$FAKE_STATE/$name.foreground" + start_viewer_fixture "$pair" + write_viewer_record "$record" "$FIXTURE_LAUNCHER_PID" "$FIXTURE_VIEWER_PID" + : > "$FAKE_LOG" + FM_FAKE_HERDR_FAST_POLL=1 run_with_fake fm_herdr_lab_teardown "$name" \ + >/dev/null 2>&1 || status=$? + expect_code 1 "$status" "teardown must refuse while an owned viewer is still attached" + [ "$(cat "$FAKE_STATE/$name")" = running ] \ + || fail "the refused teardown stopped the lab session anyway" + assert_no_grep "session delete $name" "$FAKE_LOG" \ + "the refused teardown still reached the destructive delete" + + printf '%s\n' no_foreground_client > "$FAKE_STATE/$name.foreground" + run_with_fake fm_herdr_lab_teardown "$name" || fail "teardown after the viewer detached failed" + pass "fm-herdr-lab: teardown refuses to destroy a session an attached viewer still holds" +} + +test_viewer_stop_retains_record_when_detach_is_unreadable() { + local name="fm-lab-viewer-unreadable-$$" record status=0 + run_with_fake fm_herdr_lab_provision "$name" || fail "unreadable-detach fixture provision failed" + record=$(run_with_fake fm_herdr_lab_viewer_record_path "$name") + printf 'launcher_pid=99999999\nlauncher_start=stale\nviewer_pid=99999999\nviewer_start=stale\n' > "$record" + FM_FAKE_HERDR_FAST_POLL=1 FM_FAKE_HERDR_TITLE_FAIL=1 \ + run_with_fake fm_herdr_lab_viewer_stop "$name" >/dev/null 2>&1 || status=$? + expect_code 1 "$status" "an unreadable detach result on a running session must fail closed" + assert_present "$record" "an unreadable detach result discarded the ownership record" + printf '%s\n' no_foreground_client > "$FAKE_STATE/$name.foreground" + run_with_fake fm_herdr_lab_teardown "$name" || fail "teardown after a confirmed detach failed" + pass "fm-herdr-lab: unreadable detach results fail closed on running sessions" +} + +test_viewer_launcher_refuses_unsafe_arguments() { + local launcher="$ROOT/bin/fm-herdr-lab-viewer.py" status=0 + command -v python3 >/dev/null 2>&1 || { pass "fm-herdr-lab: viewer launcher argument guard (skipped, no python3)"; return; } + python3 "$launcher" default "$TMP_ROOT/pid" >/dev/null 2>&1 || status=$? + expect_code 2 "$status" "the launcher must refuse the default session" + status=0 + python3 "$launcher" arbitrary-session "$TMP_ROOT/pid" >/dev/null 2>&1 || status=$? + expect_code 2 "$status" "the launcher must refuse a non-lab session name" + status=0 + python3 "$launcher" fm-lab-args relative-pidfile >/dev/null 2>&1 || status=$? + expect_code 2 "$status" "the launcher must refuse a relative pidfile path" + assert_absent "$TMP_ROOT/pid" "a refused launch still wrote a pid record" + pass "fm-herdr-lab: the viewer launcher refuses unsafe sessions and pidfiles" +} + test_refuses_unsafe_names test_provision_run_and_guarded_teardown test_missing_tripwire_blocks_destruction @@ -241,3 +505,13 @@ test_changed_default_trips_after_teardown test_stopped_owned_lab_can_reprovision test_failed_delete_retains_tripwire test_timed_out_provision_cancels_late_launch +test_viewer_refuses_unowned_sessions +test_viewer_start_cancels_an_unrecorded_launcher +test_viewer_timeout_allows_launcher_escalation +test_viewer_start_requires_its_owned_process +test_viewer_stop_only_signals_owned_processes +test_viewer_stop_requires_the_recorded_parent +test_interrupted_viewer_start_cancels_launcher +test_teardown_refuses_while_viewer_attached +test_viewer_stop_retains_record_when_detach_is_unreadable +test_viewer_launcher_refuses_unsafe_arguments From dee156fe394a3a48e75ea798153716194f011531 Mon Sep 17 00:00:00 2001 From: Tiago <tiagop@hey.com> Date: Fri, 11 Sep 2026 19:32:38 -0300 Subject: [PATCH 08/31] fix(bin): let nonvisual work proceed when lavish-axi is unavailable (#3766) * fix(bootstrap): allow nonvisual work without Lavish * no-mistakes(review): Gate scout brief Lavish line on bootstrap version floor --- .agents/skills/bootstrap-diagnostics/SKILL.md | 7 +++- AGENTS.md | 6 +-- bin/fm-bootstrap.sh | 25 ++++++++---- bin/fm-brief.sh | 9 ++++- bin/fm-test-run.sh | 6 ++- docs/configuration.md | 9 +++-- tests/fm-bootstrap.test.sh | 24 +++++++----- tests/fm-brief.test.sh | 38 ++++++++++++++++++- 8 files changed, 94 insertions(+), 30 deletions(-) diff --git a/.agents/skills/bootstrap-diagnostics/SKILL.md b/.agents/skills/bootstrap-diagnostics/SKILL.md index cb10d536e98..4e5ab2efa04 100644 --- a/.agents/skills/bootstrap-diagnostics/SKILL.md +++ b/.agents/skills/bootstrap-diagnostics/SKILL.md @@ -2,7 +2,7 @@ name: bootstrap-diagnostics description: >- Agent-only handling playbook for session-start bootstrap diagnostics. - Use whenever the session-start digest's bootstrap or network-checks section prints an actionable diagnostic line - MISSING, MISSING_MANUAL, BACKEND_INVALID, NEEDS_GH_AUTH, TANGLE, STARTUP_MEMORY_BUDGET, CREW_DISPATCH invalid, FLEET_SYNC, NETWORK_CHECKS, HOME_SUMMARY, BACKLOG_RECONCILE, SECONDMATE_SYNC, SECONDMATE_LIVENESS, SECONDMATE_HANDOFF, NUDGE_SECONDMATES, or FMX - or reports that an interrupted backlog cleanup may have left an endpoint or local copy, or when a standalone bin/fm-bootstrap.sh or bin/fm-startup-network.sh run prints one of those lines. + Use whenever the session-start digest's bootstrap or network-checks section prints an actionable diagnostic line - MISSING, MISSING_MANUAL, PRESENTATION_UNAVAILABLE, BACKEND_INVALID, NEEDS_GH_AUTH, TANGLE, STARTUP_MEMORY_BUDGET, CREW_DISPATCH invalid, FLEET_SYNC, NETWORK_CHECKS, HOME_SUMMARY, BACKLOG_RECONCILE, SECONDMATE_SYNC, SECONDMATE_LIVENESS, SECONDMATE_HANDOFF, NUDGE_SECONDMATES, or FMX - or reports that an interrupted backlog cleanup may have left an endpoint or local copy, or when a standalone bin/fm-bootstrap.sh or bin/fm-startup-network.sh run prints one of those lines. A silent bootstrap section, or any other BOOTSTRAP_INFO fact, means no skill load. user-invocable: false metadata: @@ -19,9 +19,12 @@ When any diagnostic needs captain attention, report the plain consequence and re - `MISSING: <tool> (install: <command>)` - list the missing tools to the captain with a one-line purpose each plus the printed install commands, wait for consent (one approval may cover the list), then run `bin/fm-bootstrap.sh install <approved tools...>`. For `treehouse`, this also covers an installed version whose `treehouse get` lacks `--lease`; treat it as an upgrade request. For `no-mistakes`, this also covers an installed version older than 1.46.0, because this repo's PR gate requires structured pipeline attestation that older builds do not write. - For any axi-family tool - `gh-axi`, `lavish-axi`, `tasks-axi`, `quota-axi` - an installed version below its floor is a plain upgrade request; [`bin/fm-bootstrap.sh`](../../../bin/fm-bootstrap.sh) owns the floor policy, and never argue the floor down to whatever the home happens to have installed. + For essential axi-family tools - `gh-axi`, `tasks-axi`, `quota-axi` - an installed version below its floor is a plain upgrade request; [`bin/fm-bootstrap.sh`](../../../bin/fm-bootstrap.sh) owns the floor policy, and never argue the floor down to whatever the home happens to have installed. For `tasks-axi`, this additionally covers an installed build that fails the separate feature probe (`bin/fm-tasks-axi-lib.sh` owns the definition); `config/backlog-backend=manual` only suppresses the verbose `BOOTSTRAP_INFO: tasks-axi available` fact, not this missing-tool report. For `quota-axi`, bootstrap requires it because firstmate reads its current output directly before resolving every crew-dispatch profile array; without it, report the missing requirement and do not choose around an unexamined candidate. +- `PRESENTATION_UNAVAILABLE: lavish-axi ...` - explain that visual presentation is unavailable and continue nonvisual work with plain-text decisions and reports; do not hold unrelated dispatch for installation consent. + Do not use Lavish until it satisfies the floor owned by `bin/fm-bootstrap.sh`; when visual work needs it, request consent for the printed install or upgrade command, then rerun bootstrap to confirm compatibility before using it. + Scout briefs check the same floor when scaffolded and ask for a text report instead of a Lavish loop, so scaffold a visual scout only after that rerun confirms compatibility. - `MISSING_MANUAL: <tool> (instructions: <url>)` - tell the captain why the tool is required and give them the printed instructions URL, but do not pass the tool to `bin/fm-bootstrap.sh install`; wait for the captain to complete the manual installation, then rerun session start to confirm the dependency is present. - `BACKEND_INVALID: <name> (known: <names>)` - the resolved runtime backend has no verified dependency or lifecycle contract, so do not dispatch work until the invalid `FM_BACKEND` or `config/backend` value is corrected to one of the listed backends. - `NEEDS_GH_AUTH` - ask the captain to run `! gh auth login` (interactive; you cannot run it for them). diff --git a/AGENTS.md b/AGENTS.md index 167d14da4cc..a64288e8884 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -198,8 +198,8 @@ When that section reports its checks still in progress it names exactly what is The closing reminder points back to the emitted supervision block and preserves only the lock, afk, Relay, and read-once reminders. Bootstrap detects first, asks for consent, and installs only after the captain approves in the current session. -Do not dispatch until the required tools are present and GitHub authentication is good. -Use `gh-axi` for GitHub, `chrome-devtools-axi` for browser work, and `lavish-axi` for structured decisions or reports; consult current help rather than memorizing flags. +Do not dispatch until the essential launch tools are present and GitHub authentication is good; presentation availability follows `bootstrap-diagnostics` and does not block nonvisual work. +Use `gh-axi` for GitHub, `chrome-devtools-axi` for browser work, and compatible `lavish-axi` for visual decisions or reports; consult current help rather than memorizing flags. A silent bootstrap section needs no action; for any printed actionable diagnostic line, load `bootstrap-diagnostics` and follow its owner procedure. `BOOTSTRAP_INFO:` lines are completed no-action facts and do not require loading a skill. `secondmate-provisioning` owns startup secondmate sync, liveness, and inherited local-material convergence. @@ -556,7 +556,7 @@ The skill owns the guarded fleet update and restart procedure; it never touches These skills are not captain-invocable; load them only at their precise triggers. -- `bootstrap-diagnostics` - load whenever the session-start digest's bootstrap or network-checks section prints an actionable diagnostic line (`MISSING:`, `MISSING_MANUAL:`, `BACKEND_INVALID:`, `NEEDS_GH_AUTH`, `TANGLE:`, `STARTUP_MEMORY_BUDGET:`, `CREW_DISPATCH: invalid`, `FLEET_SYNC:`, `NETWORK_CHECKS:`, `HOME_SUMMARY:`, `BACKLOG_RECONCILE:`, `SECONDMATE_SYNC:`, `SECONDMATE_LIVENESS:`, `SECONDMATE_HANDOFF:`, `NUDGE_SECONDMATES:`, or `FMX:`), or when `BOOTSTRAP_INFO:` says an interrupted backlog cleanup may have left an endpoint or local copy; silence and other `BOOTSTRAP_INFO:` facts need no load. +- `bootstrap-diagnostics` - load whenever the session-start digest's bootstrap or network-checks section prints an actionable diagnostic line (`MISSING:`, `MISSING_MANUAL:`, `PRESENTATION_UNAVAILABLE:`, `BACKEND_INVALID:`, `NEEDS_GH_AUTH`, `TANGLE:`, `STARTUP_MEMORY_BUDGET:`, `CREW_DISPATCH: invalid`, `FLEET_SYNC:`, `NETWORK_CHECKS:`, `HOME_SUMMARY:`, `BACKLOG_RECONCILE:`, `SECONDMATE_SYNC:`, `SECONDMATE_LIVENESS:`, `SECONDMATE_HANDOFF:`, `NUDGE_SECONDMATES:`, or `FMX:`), or when `BOOTSTRAP_INFO:` says an interrupted backlog cleanup may have left an endpoint or local copy; silence and other `BOOTSTRAP_INFO:` facts need no load. - `diagnostic-reasoning` - load before scoping a reported bug and before acting on a diagnostic report. - `ask-user-authority` - load before deciding any ask-user finding. - `quota-array-dispatch` - load before choosing among a matched crew-dispatch profile array from current quota-axi default TOON. diff --git a/bin/fm-bootstrap.sh b/bin/fm-bootstrap.sh index 230b5327a19..8441dadea89 100755 --- a/bin/fm-bootstrap.sh +++ b/bin/fm-bootstrap.sh @@ -6,6 +6,7 @@ # exits 0. # Silent = all good. # Lines: "MISSING: <tool> (install: <command>)", +# "PRESENTATION_UNAVAILABLE: lavish-axi (requires >=<floor>; install: <command>) - nonvisual work may proceed with plain-text decisions and reports; install or upgrade before using Lavish", # "MISSING_MANUAL: <tool> (instructions: <url>)", "NEEDS_GH_AUTH", # "BACKEND_INVALID: <name> (known: <names>)", # "STARTUP_MEMORY_BUDGET: invalid config/startup-memory-budget - <reason>", @@ -57,11 +58,13 @@ # 1.46.0 (structured pipeline attestation floor; see CONTRIBUTING.md). # The AXI-family floor policy is owned beside GH_AXI_MIN and # LAVISH_AXI_MIN below; the per-tool owners point there. An installed -# build below its floor reports MISSING like no-mistakes, so the operator -# is asked to upgrade rather than silently running an older tool. +# essential build below its floor reports MISSING like no-mistakes. +# Missing or incompatible lavish-axi reports PRESENTATION_UNAVAILABLE: +# nonvisual dispatch continues with plain-text decisions and reports, +# but Lavish use still requires a compatible build at or above its floor. # tasks-axi feature probes remain a separate defense-in-depth check. -# tasks-axi and quota-axi are required bootstrap tools (same class as -# lavish-axi). A compatible tasks-axi default backend is silent. +# tasks-axi and quota-axi are essential bootstrap tools. +# A compatible tasks-axi default backend is silent. # quota-axi is required for the agent-owned dispatch-profile array # procedure in AGENTS.md section 4 and # .agents/skills/quota-array-dispatch/SKILL.md. @@ -144,6 +147,9 @@ # keeps detect-only meaning unlocked, exactly as before. # fm-bootstrap.sh install <tool>... # Install the named tools (only ones the captain approved). +# fm-bootstrap.sh lavish-compatible +# Exit 0 when lavish-axi meets LAVISH_AXI_MIN, 1 otherwise, printing +# nothing; bin/fm-brief.sh uses it to gate scout Lavish hosting. set -u SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -888,7 +894,7 @@ missing_tool_diagnostic() { # fm_backend_required_tools (bin/fm-backend.sh). So a herdr/zellij/cmux home is # never told tmux is missing, and only orca drops treehouse. A backend value with # no verified dependency set is reported before the universal checks continue. -COMMON_TOOLS="node git gh no-mistakes gh-axi chrome-devtools-axi lavish-axi tasks-axi quota-axi" +COMMON_TOOLS="node git gh no-mistakes gh-axi chrome-devtools-axi tasks-axi quota-axi" BACKEND=$(fm_backend_name) BACKEND_VALID=1 if ! BACKEND_TOOLS=$(fm_backend_required_tools "$BACKEND"); then @@ -1329,6 +1335,11 @@ startup_memory_budget_setup() { fi } +if [ "${1:-}" = "lavish-compatible" ]; then + tool_version_at_least lavish-axi "$LAVISH_AXI_MIN" + exit +fi + if [ "${1:-}" = "install" ]; then shift [ $# -gt 0 ] || { echo "usage: fm-bootstrap.sh install <tool>..." >&2; exit 1; } @@ -1425,8 +1436,8 @@ detect_local_tools() { if command -v gh-axi >/dev/null 2>&1 && ! tool_version_at_least gh-axi "$GH_AXI_MIN"; then echo "MISSING: gh-axi (install: $(install_cmd gh-axi))" fi - if command -v lavish-axi >/dev/null 2>&1 && ! tool_version_at_least lavish-axi "$LAVISH_AXI_MIN"; then - echo "MISSING: lavish-axi (install: $(install_cmd lavish-axi))" + if ! tool_version_at_least lavish-axi "$LAVISH_AXI_MIN"; then + echo "PRESENTATION_UNAVAILABLE: lavish-axi (requires >=$LAVISH_AXI_MIN; install: $(install_cmd lavish-axi)) - nonvisual work may proceed with plain-text decisions and reports; install or upgrade before using Lavish" fi if command -v quota-axi >/dev/null 2>&1 && ! fm_quota_axi_compatible; then echo "MISSING: quota-axi (install: $(install_cmd quota-axi))" diff --git a/bin/fm-brief.sh b/bin/fm-brief.sh index c8d50ce09f8..750f19feba0 100755 --- a/bin/fm-brief.sh +++ b/bin/fm-brief.sh @@ -17,6 +17,8 @@ # fm-brief.sh <task-id> --secondmate {<project>...|--no-projects} # --scout writes the scout contract instead: the deliverable is a report at # data/<task-id>/report.md (no branch, no push, no PR) and the worktree is scratch. +# It offers the Lavish review loop only when `fm-bootstrap.sh lavish-compatible` +# confirms the supported lavish-axi floor; otherwise it asks for a text report. # --secondmate writes a persistent secondmate charter. The project list # is cloned into the secondmate home, while the natural-language scope # tells the main firstmate when to route work there; routine churn stays in its own home; @@ -354,6 +356,11 @@ EOF TASK_SECTION=${TASK_SECTION%$'\n'} if [ "$KIND" = scout ]; then +if "$SCRIPT_DIR/fm-bootstrap.sh" lavish-compatible >/dev/null 2>&1; then + LAVISH_LINE='If your deliverable is a visual artifact the captain will review and iterate on, you may host the Lavish review loop yourself (poll, revise, re-serve, staying alive) instead of handing it back to firstmate.' +else + LAVISH_LINE='Lavish is unavailable (lavish-axi is missing or below its supported version floor), so deliver your findings as a text report without Lavish, even for a visual deliverable.' +fi cat > "$BRIEF" <<EOF You are a crewmate: an autonomous worker agent managed by firstmate. Work on your own; do not wait for a human. @@ -409,7 +416,7 @@ $INBOX_SECTION # Definition of done Write your findings to \`$DATA/$ID/report.md\`. The report must stand alone: what you did, what you found, the evidence (commands run, output, file:line references), and what you recommend. -If your deliverable is a visual artifact the captain will review and iterate on, you may host the Lavish review loop yourself (poll, revise, re-serve, staying alive) instead of handing it back to firstmate. +$LAVISH_LINE Before reporting done, read and follow \`$FM_ROOT/.agents/skills/captain-hold-lifecycle/SKILL.md\` and pass its shared completion gate for the report and any visual review. When the report is complete, append \`done: {one-line conclusion}\` to the status file and stop. If your findings reveal work that should ship (e.g. you reproduced a bug and the fix is clear), say so in the report; firstmate may promote this task in place, and you would then receive mode-specific ship instructions as a follow-up message. diff --git a/bin/fm-test-run.sh b/bin/fm-test-run.sh index 2bdd8919a83..3633959d7cf 100755 --- a/bin/fm-test-run.sh +++ b/bin/fm-test-run.sh @@ -1368,11 +1368,15 @@ families_for_changed_path() { bin/fm-stow-cascade.sh) printf '%s\n' secondmate ;; - bin/fm-session-start.sh|bin/fm-bootstrap.sh|bin/fm-fleet-sync.sh|\ + bin/fm-session-start.sh|bin/fm-fleet-sync.sh|\ bin/fm-sessionstart-nudge.sh|bin/fm-startup-network.sh|bin/fm-tangle*|bin/fm-update.sh|\ bin/fm-gate-refuse*|bin/fm-lock*) printf '%s\n' session-bootstrap ;; + bin/fm-bootstrap.sh) + printf '%s\n' session-bootstrap + printf '%s\n' "__script__:fm-brief.test.sh" + ;; bin/fm-quota-axi-lib.sh) printf '%s\n' session-bootstrap printf '%s\n' "__script__:fm-procevent-quota.test.sh" diff --git a/docs/configuration.md b/docs/configuration.md index 8bcc89448b1..63ac6534ff7 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -441,10 +441,11 @@ Secondmate homes inherit this file from the primary, so a secondmate's own crewm On session start the first mate detects what its required toolchain is missing or too old and lists each problem with either an exact install command or manual instructions. It installs automatically supported tools only after you say go; manual-only tools remain for you to install from the printed instructions. Required tools come in two parts: a universal toolchain every home needs regardless of backend, and a per-backend delta that follows the runtime backend actually resolved for this home. -The universal toolchain is node, git, gh with GitHub auth via `gh auth login`, no-mistakes v1.46.0 or newer, compatible gh-axi, chrome-devtools-axi, compatible lavish-axi, compatible tasks-axi per "Backlog backend" above, and compatible quota-axi. +The essential universal toolchain is node, git, gh with GitHub auth via `gh auth login`, no-mistakes v1.46.0 or newer, compatible gh-axi, chrome-devtools-axi, compatible tasks-axi per "Backlog backend" above, and compatible quota-axi. [`bin/fm-bootstrap.sh`](../bin/fm-bootstrap.sh) owns the axi-family floor policy and the gh-axi and lavish-axi floors, while [`bin/fm-tasks-axi-lib.sh`](../bin/fm-tasks-axi-lib.sh) and [`bin/fm-quota-axi-lib.sh`](../bin/fm-quota-axi-lib.sh) hold their own tools' floor constants. This section is the single owner of that universal toolchain list; backend guides' prerequisites point here and add only their backend-specific tools. -In that list, no-mistakes runs the validation pipeline, gh-axi, chrome-devtools-axi, and lavish-axi cover GitHub, browser, and rich-review operations, and tasks-axi plus quota-axi back backlog mutations and quota-aware array dispatch. +In that list, no-mistakes runs the validation pipeline, gh-axi and chrome-devtools-axi cover GitHub and browser operations, and tasks-axi plus quota-axi back backlog mutations and quota-aware array dispatch. +Lavish is a presentation-only dependency for visual decisions and reports; nonvisual work can proceed with plain text when it is unavailable. The per-backend delta is required only for the backend resolved from `FM_BACKEND`, then `config/backend`, then runtime auto-detection, then default `tmux`, so a home is never told to install a tool an inactive backend or feature would need. That delta is owned in code by `fm_backend_required_tools` in `bin/fm-backend.sh`: the resolved backend's own session-provider CLI (`tmux`, `herdr`, `zellij`, `orca`, or `cmux`), `jq` for the JSON-emitting adapters (`herdr`, `zellij`, `cmux`) whose spawn and liveness paths parse the backend's JSON output, and the `treehouse` worktree provider for every session-provider-only backend (`tmux`, `herdr`, `zellij`, `cmux`). Backend tool availability uses the adapter's own executable resolver, so bootstrap and spawn agree on supported non-`PATH` locations such as cmux's bundled CLI. @@ -453,10 +454,10 @@ Orca provides both the task worktree and terminal endpoint (see "Runtime backend A herdr, zellij, or cmux home is therefore never told `tmux` is missing, and the `treehouse` durable-lease upgrade check runs only for the backends that actually use treehouse. When `config/crew-dispatch.json` exists, bootstrap also requires `jq` for dispatch profile validation. When Relay is opted in, bootstrap also requires `curl` and `jq` before arming the relay poll shim. -`tasks-axi` and `quota-axi` are required bootstrap tools in every profile, the same class as `lavish-axi`. +`tasks-axi` and `quota-axi` are essential bootstrap tools in every profile. An absent or incompatible `tasks-axi` reports `MISSING: tasks-axi (install: npm install -g tasks-axi)`; when `config/backlog-backend` is not `manual`, a home with a configured non-markdown adapter or a markdown backlog refuses lifecycle mutation until compatible `tasks-axi` is on `PATH`, while a manual-backend home keeps its backlog hand-edited. An absent or incompatible `gh-axi` reports `MISSING: gh-axi (install: npm install -g gh-axi && gh-axi setup hooks)`. -An absent or incompatible `lavish-axi` reports `MISSING: lavish-axi (install: npm install -g lavish-axi && lavish-axi setup hooks)`. +An absent or incompatible `lavish-axi` reports `PRESENTATION_UNAVAILABLE` with its required floor, install command, and explicit text fallback; [`bootstrap-diagnostics`](../.agents/skills/bootstrap-diagnostics/SKILL.md) owns the response and compatibility check before visual use. An absent or too-old `quota-axi` reports `MISSING: quota-axi (install: npm install -g quota-axi)`; firstmate cannot resolve a profile array without a compatible binary. Bootstrap also reports a `TANGLE:` line when `FM_ROOT` is on a named non-default branch; follow the printed checkout remediation rather than treating it as an installable tool problem. In a read-only session that did not get the fleet lock, the same line is advisory and omits the checkout command. diff --git a/tests/fm-bootstrap.test.sh b/tests/fm-bootstrap.test.sh index b4b9a1a67aa..23eac5b67f5 100755 --- a/tests/fm-bootstrap.test.sh +++ b/tests/fm-bootstrap.test.sh @@ -5,7 +5,7 @@ # BOOTSTRAP_INFO fact, or completed bootstrap no-action fact and is silent when # all is well. firstmate consumes the exact 'MISSING: treehouse (install: ...)', # 'MISSING: tasks-axi (install: ...)', 'MISSING: quota-axi (install: ...)', -# 'MISSING: gh-axi (install: ...)', 'MISSING: lavish-axi (install: ...)', and +# 'MISSING: gh-axi (install: ...)', 'PRESENTATION_UNAVAILABLE: lavish-axi ...', and # 'BOOTSTRAP_INFO: ...' lines, so those contracts are pinned verbatim. The cases # are table-driven over the inputs that vary: whether `treehouse get --help` # advertises --lease, which (if any) tasks-axi version is on PATH, whether @@ -373,8 +373,8 @@ ROWS } test_lavish_axi_min_version() { - local label version mode case_dir fakebin out missing n - missing='MISSING: lavish-axi (install: npm install -g lavish-axi && lavish-axi setup hooks)' + local label version mode case_dir fakebin out unavailable n + unavailable='PRESENTATION_UNAVAILABLE: lavish-axi (requires >=0.1.46; install: npm install -g lavish-axi && lavish-axi setup hooks) - nonvisual work may proceed with plain-text decisions and reports; install or upgrade before using Lavish' n=0 while IFS='^' read -r label version mode; do [ -n "$label" ] || continue @@ -383,24 +383,28 @@ test_lavish_axi_min_version() { mkdir -p "$case_dir/home/config" printf '%s\n' manual > "$case_dir/home/config/backlog-backend" fakebin=$(make_fake_toolchain "$case_dir") + [ "$version" != absent ] || rm -f "$fakebin/lavish-axi" out=$(PATH="$fakebin:$BASE_PATH" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ - FM_FAKE_TREEHOUSE_LEASE_HELP=1 FM_FAKE_LAVISH_AXI_VERSION="$version" "$ROOT/bin/fm-bootstrap.sh") + FM_FAKE_TREEHOUSE_LEASE_HELP=1 FM_FAKE_LAVISH_AXI_VERSION="$version" "$ROOT/bin/fm-bootstrap.sh") \ + || fail "$label: optional presentation must not fail bootstrap" + assert_not_contains "$out" 'MISSING:' "$label: optional presentation must not block nonvisual dispatch" case "$mode" in empty) [ -z "$out" ] || fail "$label: expected silence, got: $out" ;; - missing) - [ "$out" = "$missing" ] || fail "$label: expected '$missing', got: $out" ;; + unavailable) + [ "$out" = "$unavailable" ] || fail "$label: expected '$unavailable', got: $out" ;; esac done <<'ROWS' +absent lavish-axi permits text fallback^absent^unavailable minimum lavish-axi version is accepted^0.1.46^empty newer lavish-axi patch is accepted^0.1.47^empty newer lavish-axi minor is accepted^0.2.0^empty newer lavish-axi major is accepted^1.0.0^empty -the patch just below the floor reports an upgrade^0.1.45^missing -much older lavish-axi minor reports an upgrade^0.0.9^missing -unparseable lavish-axi version reports an upgrade^lavish-axi development build^missing +the patch just below the floor permits text fallback^0.1.45^unavailable +much older lavish-axi minor permits text fallback^0.0.9^unavailable +unparseable lavish-axi version permits text fallback^lavish-axi development build^unavailable ROWS - pass "bootstrap enforces lavish-axi minimum version" + pass "bootstrap permits nonvisual work without compatible lavish-axi and retains its presentation floor" } test_tasks_axi_min_version() { diff --git a/tests/fm-brief.test.sh b/tests/fm-brief.test.sh index 5467b5cbdca..72140fae459 100755 --- a/tests/fm-brief.test.sh +++ b/tests/fm-brief.test.sh @@ -817,6 +817,41 @@ test_scout_and_secondmate_load_decision_hold_policy() { pass "fm-brief.sh: investigation and visual-review completions load the shared decision policy" } +# A scout brief offers the Lavish review loop only when bootstrap confirms the +# supported lavish-axi floor at scaffold time; a missing or older build gets a +# text-report instruction instead, so a scout never drives a below-floor Lavish. +test_scout_lavish_line_follows_presentation_floor() { + local base label version expect case_dir fakebin brief n=0 + local hosting='you may host the Lavish review loop yourself' + local text_only='deliver your findings as a text report without Lavish' + base=$(fm_test_base_path_sans "${FM_TEST_BASE_PATH:-/usr/bin:/bin:/usr/sbin:/sbin}" lavish-axi) + while IFS='^' read -r label version expect; do + [ -n "$label" ] || continue + n=$((n + 1)) + case_dir="$TMP_ROOT/scout-lavish-$n" + mkdir -p "$case_dir/home/data" + fakebin=$(fm_fakebin "$case_dir") + [ "$version" = absent ] || fm_fake_version_tool "$fakebin" lavish-axi FM_FAKE_LAVISH_AXI_VERSION "$version" + PATH="$fakebin:$base" FM_HOME="$case_dir/home" \ + "$ROOT/bin/fm-brief.sh" scout-lavish alpha --scout >/dev/null \ + || fail "$label: scout scaffold failed" + brief="$case_dir/home/data/scout-lavish/brief.md" + if [ "$expect" = hosting ]; then + assert_grep "$hosting" "$brief" "$label: scout brief did not offer the Lavish review loop" + assert_no_grep "$text_only" "$brief" "$label: scout brief withheld Lavish from a compatible build" + else + assert_grep "$text_only" "$brief" "$label: scout brief did not ask for a text report" + assert_no_grep "$hosting" "$brief" "$label: scout brief offered a below-floor Lavish" + fi + done <<'ROWS' +lavish-axi at the floor^0.1.46^hosting +lavish-axi above the floor^0.2.0^hosting +lavish-axi just below the floor^0.1.45^text +absent lavish-axi^absent^text +ROWS + pass "fm-brief.sh: scout Lavish hosting follows the bootstrap lavish-axi floor" +} + # Scout and secondmate paths still scaffold well-formed briefs. test_scout_and_secondmate_scaffold() { local brief @@ -826,8 +861,6 @@ test_scout_and_secondmate_scaffold() { assert_present "$brief" "scout brief was not scaffolded" assert_grep "SCOUT task" "$brief" "scout brief must declare itself a scout task" assert_grep "report.md" "$brief" "scout brief must point at the report deliverable" - assert_grep "you may host the Lavish review loop yourself" "$brief" \ - "scout brief must mention the option to host a Lavish review loop" assert_grep "## Captain's intent" "$brief" "scout brief missing Captain's intent subsection" assert_grep "## Firstmate spec" "$brief" "scout brief missing Firstmate spec subsection" assert_grep "{FIRSTMATE_SPEC}" "$brief" "scout brief missing the spec placeholder" @@ -891,3 +924,4 @@ test_secondmate_directory_paths_are_absolute_and_output_is_stable test_pause_verb_override_renders_all_brief_scaffolds test_scout_and_secondmate_load_decision_hold_policy test_scout_and_secondmate_scaffold +test_scout_lavish_line_follows_presentation_floor From eb0ea3ab862c5121fd435486431911845015cb90 Mon Sep 17 00:00:00 2001 From: Tiago <tiagop@hey.com> Date: Fri, 11 Sep 2026 19:32:43 -0300 Subject: [PATCH 09/31] test: isolate fixture Git config from host global and system settings (#3825) * fix(tests): isolate fixture Git configuration from host preferences Ignore global and system Git configuration in the shared test library, which all four fixture helper entry points source. Keep local config, command-line overrides and explicitly supplied test config usable without changing the caller's environment or real project signing preferences. Exercise global and system signing inputs through all four helpers, real fixture and child commits, explicit signing overrides, unchanged input files, and signing refusal outside fixture subprocesses. Verification evidence for issue #3770: On pristine upstream f09de8a3, all 12 reported suites failed and each logged "No secret key" using a private GIT_CONFIG_GLOBAL containing commit.gpgsign=true and gpg.format=openpgp, GIT_CONFIG_NOSYSTEM=1, and an empty private GNUPGHOME (GIT_CONFIG_COUNT and GIT_CONFIG_PARAMETERS unset). With this change, all 12 pass in the identical environment through bin/fm-test-run.sh --per-script-timeout-secs 900: fm-backlog-atomicity, fm-bootstrap-network-parallel, fm-bootstrap, fm-crew-state, fm-fleet-sync, fm-gate-refuse, fm-grok-harness, fm-session-start, fm-sessionstart-nudge, fm-tangle-guard, fm-test-run, and fm-update (all tests/<name>.test.sh). The new fm-test-fixtures regression failed before the library change and passes after it. Canonical bin/fm-lint.sh passes. Additional verification exposed fm-teardown's herdr-preflight-missing-adapter assertion on both this branch and an unchanged f09de8a3 archive with signing neutralized. That pre-existing failure needs separate disposition; it is not repaired or skipped here. The separately owned Muse and composer fixture defects remain untouched. Fixes #3770 * no-mistakes(review): Complete fixture Git isolation and scope config assertions * no-mistakes(review): Share Git isolation across standalone fixture entry points * no-mistakes(review): Map git-config helper changes to lib.sh dependents * no-mistakes(review): Select fixture-isolation regression on runner change; halve config matrix * no-mistakes(review): Scope fixture-isolation regression selection to the runner alone * no-mistakes(document): Give fixture Git isolation helper its owning header * no-mistakes(document): Record fixture Git-isolation coverage in fixtures suite header * no-mistakes(review): Fix linked-worktree fixtures and remove redundant Git isolation * no-mistakes(document): Correct stale runner-selection documentation * no-mistakes(document): Clarify family antecedent in isolation-proof runner evidence --- CONTRIBUTING.md | 2 +- bin/fm-test-run.sh | 35 +++++++- docs/fm-test-isolation-proof.md | 6 +- tests/fm-gitignore-config.test.sh | 2 + tests/fm-test-fixtures.test.sh | 141 ++++++++++++++++++++++++++++++ tests/fm-test-run.test.sh | 23 +++++ tests/git-config-helpers.sh | 26 ++++++ tests/herdr-test-safety.sh | 3 + tests/lib.sh | 5 ++ 9 files changed, 235 insertions(+), 8 deletions(-) create mode 100644 tests/git-config-helpers.sh diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 628cead8ab2..c1c3d15ee91 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -113,7 +113,7 @@ Those sleeps look like recoverable overhead - `fm-watch-triage.test.sh` alone is Sampling less often does not remove that wait, it only delays detection: raising the interval to 0.5s and charging each sample proportionally measured `fm-watch-triage.test.sh` at 435s and 440s against 390s and 393s for the unchanged script, back to back on 2026-09-03, because each of its ~40 poll-cycle waits and ~73 process-exit waits paid up to half a second more. Some of those loops are also catching a transient rather than waiting for a settled condition, so a coarser sample can step over the state they assert on. Discover tests by listing `tests/*.test.sh`: each is a self-contained bash script named `<subject>.test.sh`, and its header comment describes what it covers, so pass one to `bin/fm-test-run.sh` to focus on a subject with canonical timing output. -Shared test helpers live in `tests/lib.sh` (reporters, temp roots, git fixtures), `tests/fixtures.sh` (fake toolchain and spawn-world builders), `tests/wake-helpers.sh`, and `tests/secondmate-helpers.sh`. +Shared test helpers live in `tests/lib.sh` (reporters, temp roots, git fixtures), `tests/fixtures.sh` (fake toolchain and spawn-world builders), `tests/wake-helpers.sh`, `tests/secondmate-helpers.sh`, and `tests/git-config-helpers.sh` (fixture Git isolation from the host's global and system configuration, already sourced by `tests/lib.sh` and `tests/herdr-test-safety.sh`; a suite that sources neither must source it itself before its first Git operation so a direct invocation stays isolated). Source those instead of copying a fake toolchain into a new suite. A fixture may shorten a production timeout to keep a failure path prompt, but never below what the real work inside that window costs on a loaded machine: a fork, an exec, a lock acquisition, a beacon publication, or a first-poll check. Where a case's assertion is not about the timeout itself, give that window headroom over the measured loaded cost, and bound the test's own waiting with iteration-counted poll loops, which stretch under load where a wall-clock budget does not. diff --git a/bin/fm-test-run.sh b/bin/fm-test-run.sh index 3633959d7cf..43f3bff4d06 100755 --- a/bin/fm-test-run.sh +++ b/bin/fm-test-run.sh @@ -116,6 +116,10 @@ # live-capability (a live-harness guard governed by fm_live_gate, which records # unavailable tools and explicit policy skips; see tests/lib.sh), or none. # +# Every selected script runs isolated from the host's global and system Git +# configuration, including one that sources no test helper of its own; +# tests/git-config-helpers.sh owns that contract and its limits. +# # Family labels, the changed-file map, and production portable-shard composition # live in this script only (one owner). The proven-isolated candidate set remains # owned by bin/fm-test-isolation-proof.sh; portable parallel shards are a @@ -1222,12 +1226,14 @@ select_family() { [ "$found" -eq 1 ] || die "no tests mapped to family '$want'" } -families_for_test_reference() { - local needle=$1 s +families_for_test_reference() { # <needle>... + local s needle local found=0 + local -a needles=() + for needle in "$@"; do needles+=(-e "$needle"); done while IFS= read -r s; do [ -n "$s" ] || continue - if grep -Fq "$needle" "$s"; then + if grep -Fq "${needles[@]}" "$s"; then family_for_basename "$(basename "$s")" found=1 fi @@ -1304,12 +1310,21 @@ families_for_changed_path() { # resolution in the caller; emit a marker family of __script__ printf '%s\n' "__script__:$(basename "$path")" ;; - bin/fm-test-run.sh|bin/fm-test-isolation-proof.sh) + bin/fm-test-run.sh) # Deliberately the WHOLE family, not just the two contract tests. This # runner executes every pure-contract-unit script, so a change to it is # only proven by running them: its own contract test passing says the # runner's logic is right, not that the suite it drives still runs. printf '%s\n' pure-contract-unit + # Only this script wraps each suite in run_script_bounded's fixture Git + # isolation, and only a standalone-family script proves it. + printf '%s\n' "__script__:fm-test-fixtures.test.sh" + ;; + bin/fm-test-isolation-proof.sh) + # Same reason as the runner above: the proof drives every + # pure-contract-unit script. It runs each candidate directly, never + # through run_script_bounded, so it cannot regress fixture Git isolation. + printf '%s\n' pure-contract-unit ;; bin/backends/herdr*|bin/fm-herdr-lab.sh|tests/herdr-test-safety.sh|tests/herdr-client-pair-fixture.sh) printf '%s\n' real-herdr-gated @@ -1526,6 +1541,12 @@ families_for_changed_path() { docs/configuration.md|docs/supervision-protocols/*) printf '%s\n' pure-contract-unit ;; + tests/git-config-helpers.sh) + # The reference scan is not transitive, so match the two helpers that + # source this one as well: most suites inherit it only through them. + families_for_test_reference git-config-helpers.sh lib.sh herdr-test-safety.sh \ + || printf '%s\n' "__unmapped__:$path" + ;; tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh) families_for_test_reference "$(basename "$path")" \ || printf '%s\n' "__unmapped__:$path" @@ -2280,6 +2301,12 @@ record_script_result() { # because an unbounded suite is what silently outruns its caller's budget. run_script_bounded() { # <script> <out> <stream> <id> local script=$1 out=$2 stream=$3 id=$4 + # Declaring the variables local first keeps the helper's export scoped to this + # call and its child script, so the runner's own environment is left as the + # caller had it. + local GIT_CONFIG_GLOBAL GIT_CONFIG_NOSYSTEM + # shellcheck source=tests/git-config-helpers.sh + . "$ROOT/tests/git-config-helpers.sh" || return local rc : "$id" set +e diff --git a/docs/fm-test-isolation-proof.md b/docs/fm-test-isolation-proof.md index e8df1878041..6f0766a16bc 100644 --- a/docs/fm-test-isolation-proof.md +++ b/docs/fm-test-isolation-proof.md @@ -119,10 +119,10 @@ Both `bin/fm-test-run.sh` and the current proof harness therefore order concurre | 1 | `FM_ISOLATION_SUMMARY total=32 failed=0 concurrency=4 duration_ms=161837` | | 2 | `FM_ISOLATION_SUMMARY total=32 failed=0 concurrency=4 duration_ms=156462` | -This family is what a change to `bin/fm-test-run.sh` itself selects, so it decides that selection's wall clock. -Before admission, 14 of its scripts fell to the serial tail and the 33-script selection measured 327.3s against a 300s budget: the concurrent group was 19 scripts totalling 273.4s while the tail alone was 215.7s, dominated by `fm-calm-pi-extension` (77.5s), `fm-vendor-auth-probe` (51.0s), and `fm-muse-harness` (39.7s). +The current runner-change selection is owned by [`bin/fm-test-run.sh`](../bin/fm-test-run.sh)'s changed-file map. +Before admission, 14 of the family's scripts fell to the serial tail and the 33-script selection measured 327.3s against a 300s budget: the concurrent group was 19 scripts totalling 273.4s while the tail alone was 215.7s, dominated by `fm-calm-pi-extension` (77.5s), `fm-vendor-auth-probe` (51.0s), and `fm-muse-harness` (39.7s). Admitting the family moves that tail into the bounded concurrent group. -Current runner-file selection was verified on 2026-08-28 with the runner and its tests bound to each measured Bash version. +The then-current runner-file selection was verified on 2026-08-28 with the runner and its tests bound to each measured Bash version. Because the runner uses `#!/usr/bin/env bash` and invokes each test with `bash` from `PATH`, the stock macOS measurement used `PATH=/bin:$PATH bin/fm-test-run.sh --changed --max-wall-ms 300000` so both resolved to `/bin/bash` 3.2.57. Two runs selected all 33 scripts, passed the five-minute result check in 153.5s and 166.8s, and reported the same two failures as `main`: `tests/fm-muse-harness.test.sh` and `tests/fm-composer-lib.test.sh`. With Bash 5.3.9 on `PATH`, three runs of `bin/fm-test-run.sh --changed --max-wall-ms 300000` selected the same 33 scripts, completed with 0 failures, and reported 163.8s, 172.0s, and 166.9s. diff --git a/tests/fm-gitignore-config.test.sh b/tests/fm-gitignore-config.test.sh index a221b4176d2..9c6c4c7ad6f 100755 --- a/tests/fm-gitignore-config.test.sh +++ b/tests/fm-gitignore-config.test.sh @@ -8,6 +8,8 @@ set -u ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +# shellcheck source=tests/git-config-helpers.sh +. "$ROOT/tests/git-config-helpers.sh" fail() { printf 'not ok - %s\n' "$1" >&2 diff --git a/tests/fm-test-fixtures.test.sh b/tests/fm-test-fixtures.test.sh index 9176771dd5d..5d9c67c8b1a 100755 --- a/tests/fm-test-fixtures.test.sh +++ b/tests/fm-test-fixtures.test.sh @@ -6,6 +6,12 @@ # filesystem effects - never on helper source text. Migrated spawn suites cover # fm_test_run_spawn through the real fm-spawn.sh; this file pins the shared # primitives and stubs those suites use. +# +# It is also the fixture Git-config isolation regression, with host signing +# armed on a scratch config file: it drives every entry point that must reach +# tests/git-config-helpers.sh - the shared helpers, bin/fm-test-run.sh's +# per-suite wrapper, and the standalone scripts runnable without a live vendor. +# That helper's header owns the contract and the layers it leaves in force. set -u # shellcheck source=tests/fixtures.sh @@ -13,6 +19,140 @@ set -u TMP_ROOT=$(fm_test_tmproot fm-test-fixtures) +test_git_config_isolation() ( + local dir="$TMP_ROOT/git-config" helper jobs timeout fakebin rc + mkdir -p "$dir/runner/bin" "$dir/runner/tests" + git init -q "$dir/caller" + git -C "$dir/caller" config commit.gpgsign false + cd "$dir/caller" || exit 1 + cp "$ROOT/bin/fm-test-run.sh" "$ROOT/bin/fm-timeout-lib.sh" "$dir/runner/bin/" + cp "$ROOT/tests/git-config-helpers.sh" "$dir/runner/tests/" + fakebin=$(fm_fakebin "$dir/standalone") + fm_fake_exit0 "$fakebin" pi + cat > "$fakebin/tmux" <<'SH' +#!/usr/bin/env bash +set -eu +while [ "$#" -gt 0 ]; do + if [ "$1" = -c ]; then + git -C "$2" log -1 --format=%s > "${FM_TEST_STANDALONE_COMMIT:?}" + exit 1 + fi + shift +done +SH + chmod +x "$fakebin/tmux" + cat > "$dir/runner/tests/fm-test-run.test.sh" <<'SH' +#!/usr/bin/env bash +set -eu +repo=$(mktemp -d "${TMPDIR:-/tmp}/fm-git-runner.XXXXXX") +trap 'rm -rf "$repo"' EXIT +git init -q "$repo" +git -C "$repo" config user.name 'Runner Fixture' +git -C "$repo" config user.email runner@example.invalid +git -C "$repo" commit -q --allow-empty -m initial +[ "$(git -C "$repo" log -1 --format='%s:%an:%ae')" = 'initial:Runner Fixture:runner@example.invalid' ] +[ "$(git -C "$repo" config --get fixture.input)" = preserved ] +[ "$(GIT_CONFIG_GLOBAL="$FM_TEST_GIT_CONFIG" git config --global --get commit.gpgsign)" = true ] +SH + chmod +x "$dir/runner/tests/fm-test-run.test.sh" + export GIT_CONFIG_GLOBAL="$dir/global" GIT_CONFIG_SYSTEM="$dir/system" + export GIT_CONFIG_NOSYSTEM=0 + unset GIT_CONFIG_COUNT GIT_CONFIG_PARAMETERS + + # A failing signer exposes inherited config without requiring GPG or keys. + arm_host_signing() { # <scope>: only this layer carries the failing signer + : > "$dir/global" + : > "$dir/system" + git config --file "$dir/$1" commit.gpgsign true + git config --file "$dir/$1" gpg.format openpgp + git config --file "$dir/$1" gpg.program /usr/bin/false + cp "$dir/$1" "$dir/expected" + } + + assert_helper_isolates() { # <helper> <scope> + bash -eus -- "$ROOT/tests/$1.sh" "$dir/$2-$1" "$dir/$2" <<'SH' || exit 1 +. "$1" +fm_git_init_commit "$2" +[ "$(git -C "$2" log -1 --format=%s)" = initial ] || fail "fixture has no initial commit" +fm_git_identity +# Child Git processes and direct commits inherit the same isolation. +bash -eu -c 'git -C "$1" commit -q --allow-empty -m child' _ "$2" +# Repository-local config and explicit command inputs remain authoritative. +git -C "$2" config commit.gpgsign true +git -C "$2" config gpg.program /usr/bin/false +if git -C "$2" commit -q --allow-empty -m signed > "$2/signing.log" 2>&1; then + fail "repository-local signing config was ignored" +fi +assert_grep 'gpg failed to sign' "$2/signing.log" "local signing was not attempted" +git -C "$2" -c commit.gpgsign=false commit -q --allow-empty -m explicit +GIT_CONFIG_COUNT=1 GIT_CONFIG_KEY_0=commit.gpgsign GIT_CONFIG_VALUE_0=false \ + git -C "$2" commit -q --allow-empty -m environment +# A config test can deliberately supply its own global file after sourcing. +[ "$(GIT_CONFIG_GLOBAL="$3" git config --global --get commit.gpgsign)" = true ] || fail "explicit global config was ignored" +SH + } + + assert_host_config_still_governs() { # <scope> + # Sourcing in test subprocesses cannot change the caller or its config files. + [ "$(git config --"$1" --get commit.gpgsign)" = true ] || fail "caller lost signing preference" + cmp -s "$dir/$1" "$dir/expected" || fail "host config file was changed" + git init -q "$dir/$1-outside" + if git -C "$dir/$1-outside" -c user.name=test -c user.email=test@example.invalid \ + commit -q --allow-empty -m outside > "$dir/outside.log" 2>&1; then + fail "commit outside fixtures bypassed signing" + fi + assert_grep 'gpg failed to sign' "$dir/outside.log" "outside commit did not attempt signing" + } + + # Every fixture entry point, once. Each only has to reach the shared helper; + # which layers that helper neutralizes is the helper's own property, settled + # by the system-layer case below. + arm_host_signing global + for helper in lib fixtures secondmate-helpers wake-helpers; do + assert_helper_isolates "$helper" global + done + bash -eus -- "$ROOT/tests/herdr-test-safety.sh" "$dir/global-herdr" <<'SH' || exit 1 +. "$1" +git init -q "$2" +git -C "$2" -c user.name=test -c user.email=test@example.invalid \ + commit -q --allow-empty -m initial +[ "$(git -C "$2" log -1 --format=%s)" = initial ] +SH + for jobs in 1 2; do + for timeout in 0 30; do + GIT_CONFIG_COUNT=1 GIT_CONFIG_KEY_0=fixture.input GIT_CONFIG_VALUE_0=preserved \ + FM_TEST_GIT_CONFIG="$dir/global" \ + "$dir/runner/bin/fm-test-run.sh" --jobs "$jobs" --per-script-timeout-secs "$timeout" \ + tests/fm-test-run.test.sh > "$dir/runner.log" 2>&1 \ + || fail "runner inherited global config (jobs=$jobs, timeout=$timeout): $(cat "$dir/runner.log")" + assert_grep 'FM_TEST_SUMMARY total=1 failed=0 skipped_gate=0' "$dir/runner.log" \ + "runner did not execute the Git fixture" + done + done + rc=0 + FM_SESSIONSTART_INSTRUCTION_REFRESH_LIVE_E2E=1 FM_SESSIONSTART_INSTRUCTION_REFRESH_REF=HEAD \ + FM_SESSIONSTART_INSTRUCTION_REFRESH_EXPECT=updated \ + FM_TEST_STANDALONE_COMMIT="$dir/global-standalone-commit" PATH="$fakebin:$PATH" \ + bash "$ROOT/tests/fm-sessionstart-instruction-refresh-live-e2e.test.sh" \ + > "$dir/standalone.log" 2>&1 || rc=$? + [ "$rc" = 1 ] || fail "standalone fixture did not stop at the tmux launch" + assert_grep 'could not start isolated Pi session' "$dir/standalone.log" \ + "standalone fixture failed before the tmux launch: $(cat "$dir/standalone.log")" + [ "$(cat "$dir/global-standalone-commit")" = 'test: initial instruction contract' ] \ + || fail "standalone fixture did not create its initial commit" + bash "$ROOT/tests/fm-gitignore-config.test.sh" > "$dir/gitignore.log" 2>&1 \ + || fail "standalone gitignore fixture inherited global config: $(cat "$dir/gitignore.log")" + assert_host_config_still_governs global + + # The system layer is the shared helper's other half: one entry point settles + # it, and the caller still signing proves the layer was genuinely armed. + arm_host_signing system + assert_helper_isolates lib system + assert_host_config_still_governs system + + pass "runner and shared helpers isolate host Git config and preserve explicit config and outside commits" +) + test_touch_epoch_preserves_repeated_dst_hour() { local TZ=Europe/Paris epoch path actual export TZ @@ -139,6 +279,7 @@ test_spawn_home_layout() { pass "spawn-home layout writes harness pin, beat, and brief" } +test_git_config_isolation || fail "Git fixture config isolation" test_touch_epoch_preserves_repeated_dst_hour test_no_mistakes_version_constant test_no_mistakes_init_doctor_markers diff --git a/tests/fm-test-run.test.sh b/tests/fm-test-run.test.sh index 35f3a8d4980..b1f18390745 100755 --- a/tests/fm-test-run.test.sh +++ b/tests/fm-test-run.test.sh @@ -92,6 +92,7 @@ init_changed_fixture_repo() { local repo=$1 script mkdir -p "$repo/bin" "$repo/tests" cp "$RUNNER" "$repo/bin/fm-test-run.sh" + cp "$ROOT/tests/git-config-helpers.sh" "$repo/tests/" chmod +x "$repo/bin/fm-test-run.sh" for script in \ fm-brief.test.sh \ @@ -99,6 +100,7 @@ init_changed_fixture_repo() { fm-documentation-audiences.test.sh \ fm-test-isolation-proof.test.sh \ fm-test-run.test.sh \ + fm-test-fixtures.test.sh \ fm-cd-pretool-check.test.sh \ fm-daemon.test.sh \ fm-harness-adapter-instructions-live-e2e.test.sh \ @@ -174,6 +176,7 @@ init_primary_and_linked_worktree() { for tree in "$repo" "$linked"; do mkdir -p "$tree/bin" "$tree/tests" cp "$RUNNER" "$tree/bin/fm-test-run.sh" + cp "$ROOT/tests/git-config-helpers.sh" "$tree/tests/" chmod +x "$tree/bin/fm-test-run.sh" cat >"$tree/tests/probe.test.sh" <<PROBE #!/usr/bin/env bash @@ -252,6 +255,12 @@ test_changed_runner_surfaces_select_their_family() { *tests/fm-ask-user-authority.test.sh*) ;; *) fail "runner change did not select its pure-contract-unit family: $listed" ;; esac + # The suite that proves the runner's per-suite fixture Git isolation lives in + # the standalone family, which pure-contract-unit never reaches. + case "$listed" in + *tests/fm-test-fixtures.test.sh*) ;; + *) fail "runner change did not select its fixture-isolation regression: $listed" ;; + esac git -C "$repo" add bin/fm-test-run.sh git -C "$repo" -c user.name=test -c user.email=test@example.invalid commit -qm runner-change @@ -299,6 +308,14 @@ test_changed_dependency_selection_and_unmapped_failure() { git -C "$repo" add tests/lib.sh git -C "$repo" -c user.name=test -c user.email=test@example.invalid commit -qm helper-change + printf '\n' >>"$repo/tests/git-config-helpers.sh" + listed=$(cd "$repo" && bin/fm-test-run.sh --list --changed --base HEAD) + assert_contains "$listed" "tests/fm-pr-merge.test.sh" "git-config helper selects lib.sh dependents" + assert_contains "$listed" "tests/fm-secondmate-safety.test.sh" "git-config helper selects secondmate dependents" + assert_contains "$listed" "tests/fm-bearings-snapshot.test.sh" "git-config helper selects snapshot dependents" + git -C "$repo" add tests/git-config-helpers.sh + git -C "$repo" -c user.name=test -c user.email=test@example.invalid commit -qm git-config-helper-change + printf '\n' >>"$repo/tests/fm-backend-herdr-eventwait.test.py" listed=$(cd "$repo" && bin/fm-test-run.sh --list --changed --base HEAD) assert_contains "$listed" "tests/fm-backend-herdr-smoke.test.sh" "eventwait test selects Herdr coverage" @@ -484,6 +501,7 @@ PY timeout_script=tests/fm-calm-pi-extension.test.sh mkdir -p "$timeout_repo/bin" "$timeout_repo/tests" cp "$RUNNER" "$timeout_repo/bin/fm-test-run.sh" + cp "$ROOT/tests/git-config-helpers.sh" "$timeout_repo/tests/" cat >"$timeout_repo/bin/fm-timeout-lib.sh" <<'SH' fm_run_timed() { [ "$1" -eq 900 ] || return 99 @@ -635,6 +653,7 @@ test_family_proofs_run_in_separate_concurrent_phases() { repo="$tmp/repo" mkdir -p "$repo/bin" "$repo/tests" cp "$RUNNER" "$repo/bin/fm-test-run.sh" + cp "$ROOT/tests/git-config-helpers.sh" "$repo/tests/" cp "$ROOT/bin/fm-timeout-lib.sh" "$repo/bin/fm-timeout-lib.sh" chmod +x "$repo/bin/fm-test-run.sh" for script in \ @@ -1262,6 +1281,7 @@ test_unmapped_new_test_never_inherits_family_concurrency() { repo="$tmp/repo" mkdir -p "$repo/bin" "$repo/tests" cp "$RUNNER" "$repo/bin/fm-test-run.sh" + cp "$ROOT/tests/git-config-helpers.sh" "$repo/tests/" chmod +x "$repo/bin/fm-test-run.sh" # Two members of the proven residual family, plus a test basename the family # map has never seen - the shape of any test added tomorrow. @@ -1339,6 +1359,7 @@ test_per_script_timeout_bounds_a_hang() { hang=tests/fm-hang-fixture.test.sh mkdir -p "$repo/bin" "$repo/tests" cp "$RUNNER" "$runner" + cp "$ROOT/tests/git-config-helpers.sh" "$repo/tests/" cp "$ROOT/bin/fm-timeout-lib.sh" "$repo/bin/fm-timeout-lib.sh" grandchild_pid="$tmp/grandchild.pid" cat >"$repo/$hang" <<'SH' @@ -1402,6 +1423,7 @@ test_max_wall_ms_is_a_result_not_advice() { fast=tests/fm-budget-fixture.test.sh mkdir -p "$repo/bin" "$repo/tests" cp "$RUNNER" "$runner" + cp "$ROOT/tests/git-config-helpers.sh" "$repo/tests/" cat >"$repo/$fast" <<'SH' #!/usr/bin/env bash sleep 1 @@ -1466,6 +1488,7 @@ test_jobs_parallel_scheduler_and_failure_propagation() { d=tests/fm-supervision-instructions.test.sh mkdir -p "$repo/bin" "$repo/tests" "$evidence" "$fake_bin" cp "$RUNNER" "$runner" + cp "$ROOT/tests/git-config-helpers.sh" "$repo/tests/" cat >"$fake_bin/stat" <<'SH' #!/usr/bin/env bash if [ "$1" = "-c" ] && [ "$2" = "%a" ]; then diff --git a/tests/git-config-helpers.sh b/tests/git-config-helpers.sh new file mode 100644 index 00000000000..8fde3e7f85e --- /dev/null +++ b/tests/git-config-helpers.sh @@ -0,0 +1,26 @@ +#!/usr/bin/env bash +# tests/git-config-helpers.sh - fixture Git isolation from the host's global and +# system configuration. +# +# Source this before a fixture's first Git operation: +# # shellcheck source=tests/git-config-helpers.sh +# . "$(dirname "${BASH_SOURCE[0]}")/git-config-helpers.sh" +# +# Fixture Git processes must not inherit host signing, hooks, or other global and +# system preferences: with commit.gpgsign=true set globally and no secret key for +# the fixture identities, every fixture commit fails before its assertion. +# Isolation is limited to those two layers on purpose, so repository-local config, +# inline `git -c`, GIT_CONFIG_COUNT, and a GIT_CONFIG_GLOBAL the caller supplies +# after sourcing all stay authoritative - the Git-config suites assert on them. +# The export reaches only the sourcing shell and its children, so the developer's +# own config files are never written and real project commits made outside the +# fixtures keep their configuration and signing. +# +# tests/lib.sh and tests/herdr-test-safety.sh source this for every suite that +# uses them, bin/fm-test-run.sh sources it per suite in run_script_bounded, and a +# suite reaching none of those sources it directly so a hand-run invocation is +# isolated too. tests/fm-test-fixtures.test.sh is the regression - it drives the +# shared helpers, the runner, and the standalone entry points that run without a +# live vendor - and the changed-file map selects it for a change to this file. + +export GIT_CONFIG_GLOBAL=/dev/null GIT_CONFIG_NOSYSTEM=1 diff --git a/tests/herdr-test-safety.sh b/tests/herdr-test-safety.sh index 59a2bb46cc2..810b4a518fe 100644 --- a/tests/herdr-test-safety.sh +++ b/tests/herdr-test-safety.sh @@ -4,6 +4,9 @@ # fleet-state tripwire contract is bin/fm-herdr-lab.sh. set -u +# shellcheck source=tests/git-config-helpers.sh +. "$(dirname "${BASH_SOURCE[0]}")/git-config-helpers.sh" + # Herdr backend tests drive the real fm-spawn/fm-teardown but do not source # tests/lib.sh, so exempt them from the gate-lifecycle refusal here too (see # tests/lib.sh and bin/fm-gate-refuse-lib.sh for why firstmate's own suite, diff --git a/tests/lib.sh b/tests/lib.sh index 5942e711623..c10f87fe9b1 100644 --- a/tests/lib.sh +++ b/tests/lib.sh @@ -33,6 +33,11 @@ FM_TEST_LIB_SOURCED=1 # suite's fixtures were written against. umask 022 +# Fixture Git isolation for every suite that reaches this library; the helper's +# header owns the invariant and the layers it deliberately leaves in force. +# shellcheck source=tests/git-config-helpers.sh +. "$(dirname "${BASH_SOURCE[0]}")/git-config-helpers.sh" + # Exempt firstmate's own test suite from the gate-lifecycle refusal # (bin/fm-gate-refuse-lib.sh). The no-mistakes gate runs this suite FROM a gate # worktree - the exact environment that guard refuses - so without this every From 2c1017e5257fd4a4ef5625496bfd5c806549cc05 Mon Sep 17 00:00:00 2001 From: sree <sreekaran@harvey.ai> Date: Fri, 11 Sep 2026 18:48:08 -0400 Subject: [PATCH 10/31] feat(bin): add config/claude-permission-mode to launch Claude workers in auto mode (#4239) * feat(spawn): add config/claude-permission-mode to launch Claude workers in auto mode Every Claude worker launched with --dangerously-skip-permissions, and a captain who refuses bypass mode had no way to select Claude Code's classifier-reviewed auto mode instead. A new one-token local config, config/claude-permission-mode, selects the permission flag for every Claude launch: absent or `bypass` keeps today's launch byte-for-byte, `auto` swaps in --permission-mode auto, and any other value refuses the spawn before any endpoint, worktree, or record exists and names the accepted values. fm-spawn resolves the file on every spawn and relaunch, threads the flag through the Claude launch template for crewmates, scouts, and secondmates alike, and records claude_permission_mode=auto in the task meta only under auto so the default meta stays unchanged; a relaunch re-resolves rather than preserving the line. The file is a captain-wide safety preference, so it joins the inherited local material pushed into secondmate homes. The Claude adapter reference records the verified auto launch shape on Claude Code 2.1.269 and that it never meets the once-per-machine bypass confirmation dialog; docs/configuration.md owns the schema. * no-mistakes(review): drop unread claude_permission_mode meta line and its assertions --- .../references/harness/claude.md | 3 + AGENTS.md | 1 + bin/fm-config-inherit-lib.sh | 5 +- bin/fm-spawn.sh | 43 +++++++- docs/configuration.md | 11 +++ tests/fm-secondmate-harness.test.sh | 49 ++++++++++ tests/fm-spawn-dispatch-profile.test.sh | 98 +++++++++++++++++++ 7 files changed, 208 insertions(+), 2 deletions(-) diff --git a/.agents/skills/harness-adapters/references/harness/claude.md b/.agents/skills/harness-adapters/references/harness/claude.md index c9b9834f33b..24591bc65e9 100644 --- a/.agents/skills/harness-adapters/references/harness/claude.md +++ b/.agents/skills/harness-adapters/references/harness/claude.md @@ -12,6 +12,7 @@ Busy hooks verified 2026-07-28 on Claude Code 2.1.220. | Skill | `/<skill>`, for example `/no-mistakes`. | | Model | `--model <model>`; discover through the interactive `/model` picker, with alias or full-name shape documented by `claude --help`. | | Effort | `--effort <low\|medium\|high\|xhigh\|max>`, verified on 2.1.196. | +| Permissions | `--dangerously-skip-permissions` by default, or `--permission-mode auto` when `config/claude-permission-mode` is `auto`; the `auto` shape verified on 2.1.269, and `../../../../../docs/configuration.md` "Claude permission mode" owns the file. | ## Workspace trust @@ -28,6 +29,8 @@ The once-per-machine bypass-permissions confirmation is a separate dialog, scope Never send Enter to that one either: it was observed rendering in the same shape as the trust dialog, with the selection on `No, exit` and the footer `Enter to confirm . Esc to cancel`, so Enter ends the session rather than accepting. Firstmate cannot move a selection with Enter, Escape, and C-c alone, so it cannot accept this dialog at all, and an operator accepts it once per machine instead. Inspect the pane to identify which dialog is on screen, and report it rather than answering it. +A launch under `config/claude-permission-mode=auto` never meets the bypass confirmation, because it does not request bypass mode: on 2.1.269 `claude --permission-mode auto` reached the composer directly with the footer `⏵⏵ auto mode on (shift+tab to cycle)`, so a captain who refuses the bypass dialog selects `auto` there instead of accepting it. +The workspace-trust dialog is unaffected by the permission mode and still needs the pre-registration above. ## Composer ghost diff --git a/AGENTS.md b/AGENTS.md index a64288e8884..16ec3a472b9 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -68,6 +68,7 @@ skills/ standalone public installer-facing skills, committed; not l bin/ helper scripts, committed; read each script's header before first use .env optional Relay pairing token (presence-gates section 14) and mail-plane credentials (schema: docs/configuration.md "Mail plane"); LOCAL, gitignored config/crew-harness crewmate harness override; LOCAL, gitignored; absent or "default" = same as firstmate. Inherited as the literal file: a concrete primary adapter value also controls a secondmate home's own crewmates (section 4) +config/claude-permission-mode optional one-token permission posture for every Claude worker launch: absent or "bypass" keeps --dangerously-skip-permissions, "auto" launches with --permission-mode auto; LOCAL, gitignored; inherited by secondmate homes; see docs/configuration.md "Claude permission mode" config/crew-dispatch.json optional crewmate dispatch profiles; LOCAL, gitignored; firstmate-maintained but human-editable natural-language rules that choose a per-task harness/model/effort profile (section 4). Inherited by secondmate homes config/secondmate-harness harness the PRIMARY uses to launch SECONDMATE agents, optionally followed by a model and effort token on the same line ("<harness> [<model>] [<effort>]"; section 4); LOCAL, gitignored; absent or "default" harness falls back to config/crew-harness then firstmate's own. The primary's own setting; NOT inherited into secondmate homes (secondmates do not spawn secondmates) config/backlog-backend backlog backend override; LOCAL, gitignored; absent or "tasks-axi" = the configured tasks-axi backend, "manual" = force routine backlog updates to hand-editing; inherited by secondmate homes (section 10) diff --git a/bin/fm-config-inherit-lib.sh b/bin/fm-config-inherit-lib.sh index de54245ad22..79ff10605c2 100644 --- a/bin/fm-config-inherit-lib.sh +++ b/bin/fm-config-inherit-lib.sh @@ -17,6 +17,9 @@ # default-off W3C trace-context setup, while live convergence leaves it unchanged. # The primary passes its frozen home-session decision into a newly launched # Secondmate; see docs/trace-context.md. +# Primary config/claude-permission-mode is a captain-wide safety preference +# (bypass or auto for every claude launch), so it flows down too and a +# secondmate's own claude crewmates launch on the same permission posture. # It also pushes # the one primary-authoritative shared captain-preference file, # data/captain-shared.md, into each secondmate home's data/ as a read-only copy. @@ -63,7 +66,7 @@ FM_SHARED_CAPTAIN_MODE="444" # The declared inheritable set (space-separated, config-dir-relative item paths). # Extend here to inherit more of the primary's local config; override via the # environment only in tests. Items must not contain whitespace. -FM_INHERITABLE_CONFIG="${FM_INHERITABLE_CONFIG:-crew-dispatch.json crew-harness backlog-backend backend herdr-presentation-spaces startup-memory-budget trace-context launch-env-allowlist}" +FM_INHERITABLE_CONFIG="${FM_INHERITABLE_CONFIG:-crew-dispatch.json crew-harness backlog-backend backend herdr-presentation-spaces startup-memory-budget trace-context launch-env-allowlist claude-permission-mode}" # Items whose value is a home-SESSION enablement decision rather than durable # local configuration. They are inherited at the launch convergence point, where diff --git a/bin/fm-spawn.sh b/bin/fm-spawn.sh index 35a701c7041..bdf04ca7a6f 100755 --- a/bin/fm-spawn.sh +++ b/bin/fm-spawn.sh @@ -241,8 +241,20 @@ # This is an exec environment boundary, not a sandbox for the pane's startup # shell, credential files, same-user processes, or later shell initialization. # See docs/configuration.md for provider/Git setup and supported limits. +# Claude permission mode (config/claude-permission-mode): +# One token selecting the permission flag every claude launch (ship, scout, +# secondmate, and relaunch) carries. Absent or `bypass` keeps today's +# `--dangerously-skip-permissions`; `auto` launches with `--permission-mode +# auto` instead, Claude Code's classifier-reviewed mode, for a captain who +# refuses to run workers in bypass mode. Every other part of the claude launch +# is unchanged. The token is the file's whitespace-trimmed content; any other +# value, or an unreadable file, refuses the spawn before any endpoint, +# worktree, or record exists and names the accepted values. The file is read +# on every spawn and relaunch, so a change reaches the next launch without a +# restart, and it is inherited into secondmate homes (bin/fm-config-inherit-lib.sh). # Launch templates live in launch_template() below; placeholders replaced before launch: # __BRIEF__ absolute path to data/<task-id>/brief.md +# __CLAUDEPERMFLAG__ the claude permission flag selected by config/claude-permission-mode # __PIBIN__ quoted concrete Pi-family executable path resolved from PATH # __PITUIMODE__ optional --tui-mode regular when that executable advertises it # __TURNEND__ absolute path to state/<task-id>.turn-ended (for harnesses whose @@ -401,6 +413,31 @@ if [ "$LAUNCH_ENV_ENABLED" = 1 ]; then exit 1 fi fi +# config/claude-permission-mode (header above): resolved once per spawn or +# relaunch, before any mutation, so a malformed file refuses instead of +# launching a worker on a permission posture the captain did not choose. +if ! CLAUDE_PERM_PRESENT=$(fm_config_source_present "$CONFIG/claude-permission-mode"); then + exit 1 +fi +CLAUDE_PERMISSION_MODE=bypass +if [ "$CLAUDE_PERM_PRESENT" = 1 ]; then + if [ ! -f "$CONFIG/claude-permission-mode" ] || [ ! -r "$CONFIG/claude-permission-mode" ]; then + echo "error: config/claude-permission-mode must be a readable regular file holding one of: bypass, auto" >&2 + exit 1 + fi + CLAUDE_PERMISSION_MODE=$(tr -d '[:space:]' < "$CONFIG/claude-permission-mode" || true) + case "$CLAUDE_PERMISSION_MODE" in + bypass|auto) ;; + *) + echo "error: config/claude-permission-mode holds '$CLAUDE_PERMISSION_MODE'; accepted values are: bypass (--dangerously-skip-permissions, the default when the file is absent), auto (--permission-mode auto)" >&2 + exit 1 + ;; + esac +fi +case "$CLAUDE_PERMISSION_MODE" in + auto) CLAUDE_PERM_FLAG='--permission-mode auto' ;; + *) CLAUDE_PERM_FLAG='--dangerously-skip-permissions' ;; +esac SUB_HOME_MARKER=".fm-secondmate-home" if [ -e "$STATE" ] || [ -L "$STATE" ]; then fm_backlog_directory_present "$STATE" "state directory" || { @@ -1437,7 +1474,10 @@ launch_template() { # sources are not guaranteed to load that scope, so a worker would # otherwise run with attribution back on; carrying it per launch keeps the # policy in force regardless of which settings scopes end up loaded. - claude) printf '%s' 'CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '\''{"feedbackDrafts":"off","attribution":{"commit":"","pr":"","sessionUrl":false}}'\'' __MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' ;; + # __CLAUDEPERMFLAG__ is the permission flag config/claude-permission-mode + # selects (header above): --dangerously-skip-permissions by default, or + # --permission-mode auto for a captain who refuses bypass mode. + claude) printf '%s' 'CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude __CLAUDEPERMFLAG__ --settings '\''{"feedbackDrafts":"off","attribution":{"commit":"","pr":"","sessionUrl":false}}'\'' __MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' ;; codex) if [ "$kind" = secondmate ]; then printf '%s' 'codex __MODELFLAG____EFFORTFLAG__--dangerously-bypass-approvals-and-sandbox "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' @@ -3776,6 +3816,7 @@ MODELFLAG=$(model_flag_for_harness "$HARNESS" "$MODEL") EFFORTFLAG=$(effort_flag_for_harness "$HARNESS" "$EFFORT" "$MODEL") || exit 1 LAUNCH=${LAUNCH//__MODELFLAG__/$MODELFLAG} LAUNCH=${LAUNCH//__EFFORTFLAG__/$EFFORTFLAG} +LAUNCH=${LAUNCH//__CLAUDEPERMFLAG__/$CLAUDE_PERM_FLAG} if [ "$HARNESS" = rovo ]; then ROVOCONFIGOVERRIDE=$(rovo_config_override_flag "$EFFORT" "$DATA" "$STATE" "$ID") || { echo "error: could not resolve this task's home paths for rovo's allowedExternalPaths grant" >&2 diff --git a/docs/configuration.md b/docs/configuration.md index 63ac6534ff7..59421eab7fb 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -344,6 +344,17 @@ Its `remove` action excises only the marker-delimited Firstmate region and remov For Pi and pi-signed secondmate launches, `fm-spawn.sh` starts the selected executable with `-e` pointed at the secondmate home's own tracked `.pi/extensions/fm-primary-pi-watch.ts` and `.pi/extensions/fm-primary-turnend-guard.ts`, both already present from the secondmate home's git worktree. For omp secondmate launches, `fm-spawn.sh` passes no `-e` at all: omp auto-discovers the home's tracked `.omp/extensions/` with no trust gate, and naming a discovered file with `-e` as well loads it twice; every omp launch instead carries the tracked `.omp/fm-worker-overlay.yml` posture overlay through `--config`, which [`fm-spawn.sh --help`](../bin/fm-spawn.sh) owns. +## Claude permission mode (config/claude-permission-mode) + +The optional local, gitignored `config/claude-permission-mode` holds one token selecting the permission flag every Claude worker launch carries: crewmates, scouts, Claude secondmates, and control-plane relaunches alike. +The token is the file's whitespace-trimmed content. +`bypass` keeps today's launch, `claude --dangerously-skip-permissions`, and is also the default when the file is absent, so an unconfigured home launches byte-for-byte as before. +`auto` replaces that flag with `--permission-mode auto`, Claude Code's classifier-reviewed permission mode, for a captain who refuses to run workers in bypass mode; every other part of the Claude launch, including its environment prefix, inline settings, model, and effort flags, is unchanged. +Any other value, or an unreadable file, refuses every spawn from that home, whichever harness it would launch, before any endpoint, worktree, or task record exists, and names the accepted values; Firstmate never falls back to a permission posture the captain did not choose. +`bin/fm-spawn.sh` reads the file on every spawn and relaunch, so a change takes effect at the next launch without a restart. +The file is a captain-wide safety preference, so it is inherited into secondmate homes under the [`secondmate-provisioning`](../.agents/skills/secondmate-provisioning/SKILL.md) inherited-local-material contract; a secondmate's own Claude crewmates then launch on the same posture. +The [Claude adapter reference](../.agents/skills/harness-adapters/references/harness/claude.md) records the verified shape of both launches and which once-per-machine dialog each one can meet. + ## Worker launch environment (config/launch-env-allowlist) The optional local, gitignored `config/launch-env-allowlist` limits the ambient environment passed to newly launched workers, scouts, and secondmates, including relaunches. diff --git a/tests/fm-secondmate-harness.test.sh b/tests/fm-secondmate-harness.test.sh index 602b2ff567c..83668e21593 100755 --- a/tests/fm-secondmate-harness.test.sh +++ b/tests/fm-secondmate-harness.test.sh @@ -1007,6 +1007,7 @@ new_world() { [ "$dispatch_ignore" = no ] || printf 'config/crew-dispatch.json\n' printf 'config/crew-harness\nconfig/secondmate-harness\nconfig/backlog-backend\n' printf 'config/backend\nconfig/herdr-presentation-spaces\nconfig/startup-memory-budget\n' + printf 'config/claude-permission-mode\n' } > "$w/main/.gitignore" printf 'v1\n' > "$w/main/AGENTS.md" printf 'r1\n' > "$w/main/README.md" @@ -1391,6 +1392,52 @@ test_bootstrap_sweep_materializes_and_inherits_memory_default() { } # config/backend: present and absent primary state converges exactly. +# config/claude-permission-mode=auto reaches a Claude SECONDMATE launch too: the +# same template swap as a crewmate, with model/effort untouched. +test_spawn_secondmate_claude_permission_mode_auto() { + local w sm meta launchlog launch out status + w="$TMP_ROOT/spawn-claude-permmode" + sm="$w/sm" + launchlog="$w/launch.log" + mkdir -p "$w/home/config" + printf 'claude opus\n' > "$w/home/config/secondmate-harness" + printf 'auto\n' > "$w/home/config/claude-permission-mode" + make_seeded_home "$sm" sm + + out=$(spawn_secondmate_capture "$w" sm "$sm" "$launchlog" 2>&1); status=$? + expect_code 0 "$status" "claude secondmate spawn under claude-permission-mode=auto should succeed" + + meta="$w/home/state/sm.meta" + [ "$(meta_field "$meta" harness)" = claude ] || fail "permmode: meta harness not claude" + launch=$(cat "$launchlog") + assert_contains "$launch" "claude --permission-mode auto --settings '{\"feedbackDrafts\":\"off\",\"attribution\":{\"commit\":\"\",\"pr\":\"\",\"sessionUrl\":false}}' --model 'opus'" \ + "permmode: secondmate launch did not swap the permission flag while keeping --model" + assert_not_contains "$launch" "--dangerously-skip-permissions" "permmode: secondmate launch must not request bypass mode" + pass "C2b spawn: config/claude-permission-mode=auto reaches a Claude secondmate launch" +} + +# The file is a captain-wide safety preference, so it inherits like +# config/backend: present values converge exactly and primary absence mirrors. +test_claude_permission_mode_inheritance_present_and_absent() { + local w head out err status + w=$(new_world permmode-inherit) + head=$(git -C "$w/main" rev-parse HEAD) + add_sm_worktree "$w" sm "$head" + + printf 'auto\n' > "$w/home/config/claude-permission-mode" + err="$w/permmode-inherit.err" + out=$(run_config_push "$w" 2>"$err"); status=$? + expect_code 0 "$status" "claude-permission-mode present push should succeed" + assert_contains "$out" "claude-permission-mode: pushed" "present value should report pushed" + [ "$(cat "$w/sm/config/claude-permission-mode")" = auto ] || fail "claude-permission-mode present value not pushed" + + rm -f "$w/home/config/claude-permission-mode" + out=$(run_config_push "$w" 2>"$err"); status=$? + expect_code 0 "$status" "claude-permission-mode absence push should succeed" + [ -e "$w/sm/config/claude-permission-mode" ] && fail "claude-permission-mode not removed on primary absence" + pass "B12c claude-permission-mode inheritance: present values and primary absence converge exactly" +} + test_backend_inheritance_present_and_absent() { local w head out err status instruction w=$(new_world backend-inherit) @@ -2588,6 +2635,8 @@ test_bootstrap_sweep_propagates_when_tracked_current test_bootstrap_sweep_defers_dispatch_on_stale_unignored_home test_bootstrap_sweep_materializes_and_inherits_memory_default test_backend_inheritance_present_and_absent +test_spawn_secondmate_claude_permission_mode_auto +test_claude_permission_mode_inheritance_present_and_absent test_presentation_inheritance_default_on_and_opt_out test_bootstrap_sweep_surfaces_config_propagation_failure test_bootstrap_rereads_after_partial_propagation diff --git a/tests/fm-spawn-dispatch-profile.test.sh b/tests/fm-spawn-dispatch-profile.test.sh index 10279b7f188..188e6890bf3 100755 --- a/tests/fm-spawn-dispatch-profile.test.sh +++ b/tests/fm-spawn-dispatch-profile.test.sh @@ -1203,6 +1203,99 @@ SH pass "fm-spawn: actual ship/scout launch commands deliver the worker role contract" } +# config/claude-permission-mode (bin/fm-spawn.sh header): absent and `bypass` +# must both produce today's launch byte-for-byte, `auto` swaps only the +# permission flag, and any other token refuses before endpoint or metadata. +claude_expected_launch() { # <home> <id> <permission-flag> + local home=$1 id=$2 flag=$3 + printf '%s' "env -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u GEMINI_CLI CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude $flag --settings '{\"feedbackDrafts\":\"off\",\"attribution\":{\"commit\":\"\",\"pr\":\"\",\"sessionUrl\":false}}' \"\$('${ROOT}/bin/fm-operational-input.sh' encode launch-brief < '$home/data/$id/launch-brief.md')\"" +} + +test_claude_permission_mode_bypass_matches_absent_launch() { + local rec id out status launch expected + id=permmode-bypass-z19 + rec=$(make_spawn_case permmode-bypass claude "$id") + read_case_record "$rec" + printf 'bypass\n' > "$HOME_DIR/config/claude-permission-mode" + + out=$(run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR") + status=$? + expect_code 0 "$status" "claude spawn with claude-permission-mode=bypass should succeed" + launch=$(cat "$LAUNCH_LOG") + expected=$(claude_expected_launch "$HOME_DIR" "$id" --dangerously-skip-permissions) + [ "$launch" = "$expected" ] || fail "explicit bypass did not reproduce the absent-file launch"$'\n'"expected: $expected"$'\n'"actual: $launch" + pass "config/claude-permission-mode=bypass launches exactly as an absent file does" +} + +test_claude_permission_mode_auto_swaps_only_the_permission_flag() { + local rec id out status launch expected + id=permmode-auto-z20 + rec=$(make_spawn_case permmode-auto claude "$id") + read_case_record "$rec" + # Surrounding whitespace is trimmed, so an editor's trailing newline or indent is fine. + printf ' auto\n' > "$HOME_DIR/config/claude-permission-mode" + + out=$(run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR") + status=$? + expect_code 0 "$status" "claude spawn with claude-permission-mode=auto should succeed" + assert_contains "$out" "spawned $id harness=claude" "auto spawn did not report claude" + launch=$(cat "$LAUNCH_LOG") + expected=$(claude_expected_launch "$HOME_DIR" "$id" '--permission-mode auto') + [ "$launch" = "$expected" ] || fail "auto changed more than the permission flag"$'\n'"expected: $expected"$'\n'"actual: $launch" + assert_not_contains "$launch" "--dangerously-skip-permissions" "auto launch must not request bypass mode" + pass "config/claude-permission-mode=auto replaces --dangerously-skip-permissions with --permission-mode auto" +} + +test_claude_permission_mode_auto_reaches_scout_launch() { + local rec id out status launch + id=permmode-scout-z21 + rec=$(make_spawn_case permmode-scout claude "$id") + read_case_record "$rec" + printf 'auto\n' > "$HOME_DIR/config/claude-permission-mode" + + out=$(run_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --scout) + status=$? + expect_code 0 "$status" "claude scout spawn with claude-permission-mode=auto should succeed" + launch=$(cat "$LAUNCH_LOG") + assert_contains "$launch" "claude --permission-mode auto --settings" "scout launch did not carry --permission-mode auto" + assert_not_contains "$launch" "--dangerously-skip-permissions" "scout launch must not request bypass mode" + pass "config/claude-permission-mode=auto reaches scout launches too" +} + +test_claude_permission_mode_invalid_refuses_before_endpoint_or_metadata() { + local rec id out status + id=permmode-invalid-z22 + rec=$(make_spawn_case permmode-invalid claude "$id") + read_case_record "$rec" + printf 'yolo\n' > "$HOME_DIR/config/claude-permission-mode" + + out=$(run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR") + status=$? + expect_code 1 "$status" "an unrecognized claude-permission-mode token must refuse the spawn" + assert_contains "$out" "config/claude-permission-mode holds 'yolo'" "refusal must name the file and the offending token" + assert_contains "$out" "bypass" "refusal must list bypass as an accepted value" + assert_contains "$out" "--permission-mode auto" "refusal must list auto as an accepted value" + [ ! -s "$LAUNCH_LOG" ] || fail "an invalid permission mode must launch nothing (got: $(cat "$LAUNCH_LOG"))" + assert_absent "$HOME_DIR/state/$id.meta" "refusal must happen before meta is written" + pass "an unrecognized config/claude-permission-mode token refuses before any endpoint or metadata" +} + +test_non_claude_harness_ignores_claude_permission_mode() { + local rec id out status launch + id=permmode-codex-z23 + rec=$(make_spawn_case permmode-codex codex "$id") + read_case_record "$rec" + printf 'auto\n' > "$HOME_DIR/config/claude-permission-mode" + + out=$(run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness codex) + status=$? + expect_code 0 "$status" "codex spawn under claude-permission-mode=auto should succeed" + launch=$(cat "$LAUNCH_LOG") + assert_contains "$launch" "codex " "codex launch did not run codex" + assert_not_contains "$launch" "--permission-mode" "the claude permission flag must not leak into a codex launch" + pass "config/claude-permission-mode changes claude launches only" +} + test_worker_launch_delivers_role_scope test_no_profile_keeps_claude_profile_defaults test_non_cursor_launch_clears_inherited_cursor_markers @@ -1236,6 +1329,11 @@ test_pi_signed_persistent_secondmate_uses_pi_extensions_and_identity test_batch_forwards_shared_profile_flags test_claude_forwards_firstmate_config_dir_when_set test_claude_omits_config_dir_prefix_when_unset +test_claude_permission_mode_bypass_matches_absent_launch +test_claude_permission_mode_auto_swaps_only_the_permission_flag +test_claude_permission_mode_auto_reaches_scout_launch +test_claude_permission_mode_invalid_refuses_before_endpoint_or_metadata +test_non_claude_harness_ignores_claude_permission_mode test_non_claude_harness_ignores_config_dir test_claude_crewmate_launch_carries_the_attribution_policy test_claude_secondmate_launch_carries_the_attribution_policy From 7d14fc126aa8965611f1958a00542d3adc4abdb5 Mon Sep 17 00:00:00 2001 From: Christoph Meise <christoph@scripe.io> Date: Sat, 12 Sep 2026 00:57:18 +0200 Subject: [PATCH 11/31] fix(teardown): leave a Treehouse pool slot reassigned to another task untouched (#4243) * fix(teardown): refuse to return a Treehouse pool slot reassigned to another task A pool slot is reused across tasks, so a finished task's worktree= line can name a slot a different, live task now holds. Teardown already refused when a second task record named the same live path, but that scan cannot prove the record it is tearing down is the current owner: the task that took the slot next may leave no record the scan can reach - its own worker may have exited and its record been cleaned up, or it may live in a home this machine does not register. Teardown then killed every process under the path, hard-reset it and returned it, and its unlanded-work refusal never fired because it was inspecting a directory that no longer belonged to the task being torn down (observed 2026-09-07). Treehouse's own state file cannot answer the ownership question. It records a slot's owner as a live process lease (owner_pid plus owner_started_at, with `treehouse status` reporting in-use from the processes actually running under the path), which names no task and is released by the very event that makes a record stale - the worker exiting. An unleased slot therefore reads identical whether it is still this task's or has since been handed on, and a slot whose new holder has also exited but left uncommitted work reads as free. So the identity source is Firstmate's own claim, not Treehouse's lease. fm-spawn writes that claim - the task id - into the slot at the moment it takes it, under the same project lock that allocates the slot, and fm-teardown drops it only after the slot is genuinely returned. It lives at <pool>/<slot>/.fm-slot-owner, a sibling of the repo checkout rather than a file inside it, so claiming a slot can never dirty the copy the landed-work checks inspect. A claim naming another task, or one that cannot be read, refuses; --force does not lift either refusal, because --force authorizes discarding this task's unlanded work, never another task's live work. A slot that cannot be claimed refuses the spawn instead. An absent claim proceeds on exactly the record-scan protection it had before: slots taken before claims existed, and slots already returned, carry none, and refusing those would strand every task in flight across this change on no evidence at all. The refusal is deliberately all-or-nothing rather than partially completing the task's own cleanup. state/<id>.meta is the only durable record naming the worktree and endpoint, so removing it would destroy the evidence needed to reconcile which record is wrong, and its removal is one step with the backlog transition. Nothing is stranded: clearing the stale worktree= line leaves a record with no slot to release, which then tears down normally, and the refusal names that remedy. Repairing the previous claimant's stale worktree= line at spawn time is left for separate work. It would have the new owner write another task's record - the same class of cross-task mutation this bug is - and would need that record's own meta lock; with the claim in place teardown refuses on evidence rather than depending on the stale pointer having been scrubbed. For the same reason the relaunch path writes no claim: it holds no allocation lock, and a record whose worktree= is already stale would stamp the wrong task's claim onto a live sibling's slot. The regression reproduces the reuse sequence with only one discoverable record, including a clean, fully landed ship copy torn down without --force - the shape of the real incident, which the previous code returned to the pool - and fails against the previous code; the existing two-record, cross-home, own-slot and no-claim cases still pass unchanged. This builds ON upstream b028e8b1 (#3837), which is already in this branch's base (origin/main 40c50ea8) and owns the record-exclusivity scan. Nothing here replaces that scan; the claim is the positive proof it cannot supply. Claude-Session: https://claude.ai/code/session_01JTBmuqKugaPUj7k9TXQwFS * no-mistakes(review): teardown leaves reassigned slot; spawn abort drops claim * no-mistakes(review): narrow Treehouse lease evidence; gate abort claim release on lock * no-mistakes(review): pin spawn-side slot claim; narrow abort-release header * no-mistakes(document): docs: point slot-claim rationale at fm-wake-lib owner --- bin/fm-spawn.sh | 48 ++++- bin/fm-teardown.sh | 202 +++++++++++++++++----- bin/fm-wake-lib.sh | 115 ++++++++++++ docs/architecture.md | 3 +- tests/fm-spawn-pool-base-freshen.test.sh | 68 ++++++++ tests/fm-teardown-endpoint-safety.test.sh | 149 ++++++++++++++++ 6 files changed, 537 insertions(+), 48 deletions(-) diff --git a/bin/fm-spawn.sh b/bin/fm-spawn.sh index bdf04ca7a6f..eefc4abbf51 100755 --- a/bin/fm-spawn.sh +++ b/bin/fm-spawn.sh @@ -115,7 +115,15 @@ # root Firstmate home's state directory before slot allocation and holds it through # task metadata publication. Teardown holds that same lock while proving and # returning a slot, so allocation cannot reuse a slot before its owner record -# is published. The local root is whatever bin/fm-wake-lib.sh's +# is published. Under that same lock it writes the slot's owner claim, which is +# what lets teardown leave a slot reassigned since untouched; bin/fm-wake-lib.sh +# owns the claim and bin/fm-teardown.sh owns what it protects. A slot that +# cannot be claimed refuses the spawn rather than launching a worker whose slot +# could later be released out from under its successor. A spawn that aborts +# while it still holds the allocation lock drops its own claim; an abort after +# metadata publication has released that lock leaves the claim in place, and +# the next spawn's claim replaces it. +# The local root is whatever bin/fm-wake-lib.sh's # fm_firstmate_root_home resolves, so a home seeded from another machine anchors # that lock itself rather than failing to resolve one; # contention refuses rather than waits. @@ -912,6 +920,7 @@ SPAWN_TASK_SET_LOCK= SPAWN_TASK_SET_LOCK_HELD=0 SPAWN_TREEHOUSE_PROJECT_LOCK= SPAWN_TREEHOUSE_PROJECT_LOCK_HELD=0 +SPAWN_SLOT_CLAIMED=0 RELAUNCH_REPLACEMENT_PENDING=0 RELAUNCH_REPLACEMENT_BUSY_GEN= RELAUNCH_REPLACEMENT_HARNESS= @@ -1043,6 +1052,23 @@ spawn_abort_cleanup() { SPAWN_META_LOCK_HELD=0 fm_lock_release "$SPAWN_META_LOCK" || true fi + # A spawn that aborts after claiming its slot but before its record survives + # must not leave a claim naming a task no record describes. The release is a + # read-then-remove, so it runs only while the project lock that wrote the + # claim is still held (aborts before metadata publication); a later abort has + # already released that lock and leaves the claim for the next spawn's + # atomic replacement rather than racing it. The release itself never removes + # another task's claim. + if [ "$SPAWN_SLOT_CLAIMED" = 1 ] && [ -n "${WT:-}" ] \ + && [ ! -e "$STATE/$ID.meta" ] && [ ! -L "$STATE/$ID.meta" ] \ + && fm_treehouse_pool_slot "$PROJ_ABS" "$WT"; then + SPAWN_SLOT_CLAIMED=0 + if [ "$SPAWN_TREEHOUSE_PROJECT_LOCK_HELD" = 1 ]; then + fm_treehouse_slot_owner_release "$WT" "$ID" || true + else + echo "warning: leaving task $ID's slot claim on $WT in place; the Treehouse project lock is no longer held, so the next spawn's claim replaces it" >&2 + fi + fi if [ "$SPAWN_TREEHOUSE_PROJECT_LOCK_HELD" = 1 ]; then SPAWN_TREEHOUSE_PROJECT_LOCK_HELD=0 fm_lock_release "$SPAWN_TREEHOUSE_PROJECT_LOCK" || true @@ -3166,6 +3192,26 @@ elif [ "$KIND" != secondmate ] && [ "$BACKEND" != orca ]; then fi validate_spawn_worktree "treehouse get" "$T" + + # Claim the pool slot for this task. The interactive `treehouse get` sent to + # the pane above records only a process lease (Treehouse's durable + # `get --lease --lease-holder`, which bin/fm-home-seed.sh uses for secondmate + # homes, is not this path), so Treehouse cannot say which task a slot belongs + # to once that task's worker exits - and that is exactly when the slot is + # handed on and this task's worktree= line goes stale. The claim is what lets + # bin/fm-teardown.sh leave a slot that has since been reassigned untouched, so + # a slot that cannot be claimed is refused here, at the cheapest point, rather + # than launching a worker whose slot teardown could later release out from + # under its successor. + # Written under the Treehouse project lock held from before slot allocation + # through metadata publication, so no other spawn or return sees a half-claim. + if fm_treehouse_pool_slot "$PROJ_ABS" "$WT"; then + if ! fm_treehouse_slot_owner_claim "$WT" "$ID" "$FM_HOME"; then + echo "error: could not claim Treehouse pool slot $WT for task $ID; refusing to launch a worker whose slot cannot later be proved to be its own; inspect window $T" >&2 + exit 1 + fi + SPAWN_SLOT_CLAIMED=1 + fi fi if [ "$RELAUNCH" -eq 0 ] && [ "$KIND" != secondmate ]; then freshen_spawn_worktree_base "$WT" || exit 1 diff --git a/bin/fm-teardown.sh b/bin/fm-teardown.sh index 3a188227cde..e8a6a195ce8 100755 --- a/bin/fm-teardown.sh +++ b/bin/fm-teardown.sh @@ -79,8 +79,36 @@ # cleanup step, teardown verifies record exclusivity: no OTHER task record in # this home or any locally registered Firstmate home may name the same live path # in its worktree= or home=. One live path with two task records is the reuse -# collision itself, whichever record is stale. The recorded endpoint's exact -# task identity and the record's spawn incarnation are validated separately +# collision itself, whichever record is stale. +# That scan alone cannot prove THIS record is the current owner, because the task +# that took the slot next may leave no record it can reach - its own worker may +# have exited and its record been cleaned up, or it may live in a home this +# machine does not register - which is how a released-then-reassigned slot was +# returned out from under a live worker (observed 2026-09-07). So teardown also +# reads the slot's own owner claim, written by bin/fm-spawn.sh at the moment the +# slot is taken and dropped here once it is genuinely returned; bin/fm-wake-lib.sh +# owns the claim, its location, and its states. A claim naming another task is +# proof of reassignment: the slot is no longer this task's, so teardown warns, +# names the claimant, and then finishes only this task's own cleanup - endpoint, +# status, records, checks, backlog - while every step that would read or touch +# that slot is skipped: no process kill under it, no dirty or landed-work +# inspection of it, no branch or hook removal in it, no Treehouse return, and +# never the other task's claim. Skipping the inspection discards nothing of this +# task's: whatever unlanded work it had in that slot was already destroyed when +# the pool handed the slot on. Refusing instead would strand the record, because +# bin/fm-backend.sh's endpoint validation refuses an empty or missing worktree= +# unconditionally, so there is no line an operator could clear to get past it. +# A claim that cannot be read proves nothing either way and refuses; inspect or +# repair the claim file at the printed path and re-run - never remove it, since +# an absent claim proceeds and would return a slot that may be another task's. An +# absent claim - a slot taken before claims existed, or already returned - keeps +# exactly the record-scan protection it had before, because refusing it would +# strand every task in flight across that change on no evidence at all. +# Why Treehouse's own state cannot answer this for crewmate slots, and why the +# claim file sits on top of it, is owned by bin/fm-wake-lib.sh's slot-owner +# claim comment. +# The recorded endpoint's exact task identity and the record's spawn incarnation +# are validated separately # before cleanup. Its current working directory is only incidental process # state: the same worker remains the owner after changing directory, so cwd can # never veto teardown of that exact recorded endpoint. @@ -93,9 +121,10 @@ # through metadata publication, closing the publication # gap; forced secondmate teardown takes it and runs the same checks for every # descendant Treehouse slot before touching any child. -# This refusal is not relaxed by --force: --force authorizes discarding THIS -# task's unlanded work, never another task's live work. Reconcile whichever -# record is wrong and re-run. Orca is not a pool slot and proves its path through +# These refusals are not relaxed by --force: --force authorizes discarding THIS +# task's unlanded work, never another task's live work. Nothing of this task's +# own is removed by a refusal; reconcile whichever record is wrong and re-run. +# Orca is not a pool slot and proves its path through # require_orca_worktree_path_match instead. # Orca tasks use the same safety checks, then close the recorded terminal and # remove the recorded worktree through `orca worktree rm`; teardown never guesses @@ -297,23 +326,6 @@ if [ "$FORCE" = --force ] && [ "$(fm_lease_actor)" = branch ]; then fi fm_lease_guard "$ID" "teardown (fm-teardown)" -# A Treehouse slot has the managed pool's fixed <pool>/<slot>/<repo> layout. -# Require both its pool state and the same Git common directory as the recorded -# project; an ordinary linked worktree is not evidence that Treehouse owns it. -is_treehouse_pool_slot() { # <project> <worktree> - local project=$1 worktree=$2 slot pool state project_common slot_common - [ -d "$project" ] && [ -d "$worktree" ] || return 1 - slot=$(CDPATH='' cd -- "$worktree" 2>/dev/null && pwd -P) || return 1 - pool=$(dirname "$(dirname "$slot")") - state="$pool/treehouse-state.json" - [ -f "$state" ] && [ ! -L "$state" ] || return 1 - project_common=$(git -C "$project" rev-parse --path-format=absolute --git-common-dir 2>/dev/null) || return 1 - slot_common=$(git -C "$slot" rev-parse --path-format=absolute --git-common-dir 2>/dev/null) || return 1 - project_common=$(CDPATH='' cd -- "$project_common" 2>/dev/null && pwd -P) || return 1 - slot_common=$(CDPATH='' cd -- "$slot_common" 2>/dev/null && pwd -P) || return 1 - [ "$project_common" = "$slot_common" ] -} - META="$STATE/$ID.meta" TREEHOUSE_PROJECT_LOCK= TREEHOUSE_PROJECT_LOCK_HELD=0 @@ -327,7 +339,7 @@ if [ -f "$META" ] && [ ! -L "$META" ]; then TEARDOWN_LOCK_PROJECT=$(fm_meta_get "$META" project) if [ "$TEARDOWN_LOCK_KIND" != secondmate ] \ && [ "$TEARDOWN_LOCK_BACKEND" != orca ] \ - && is_treehouse_pool_slot "$TEARDOWN_LOCK_PROJECT" "$TEARDOWN_LOCK_WT"; then + && fm_treehouse_pool_slot "$TEARDOWN_LOCK_PROJECT" "$TEARDOWN_LOCK_WT"; then TREEHOUSE_SLOT_LOCK_REQUIRED=1 TREEHOUSE_PROJECT_LOCK=$(fm_treehouse_project_lock_path "$TEARDOWN_LOCK_PROJECT") || { echo "REFUSED: cannot resolve the shared Treehouse project lock for ${TEARDOWN_LOCK_PROJECT:-<missing>}; nothing was changed" >&2 @@ -926,7 +938,7 @@ CLEANUP_RECOVERY=$TEARDOWN_CLEANUP_RECOVERY KIND=$TEARDOWN_META_KIND EXPECTED_TREEHOUSE_PROJECT_LOCK= if [ "$KIND" != secondmate ] && [ "$BACKEND" != orca ] \ - && is_treehouse_pool_slot "$PROJ" "$WT"; then + && fm_treehouse_pool_slot "$PROJ" "$WT"; then EXPECTED_TREEHOUSE_PROJECT_LOCK=$(fm_treehouse_project_lock_path "$PROJ") || { echo "REFUSED: cannot resolve the shared Treehouse project lock for ${PROJ:-<missing>}; nothing was changed" >&2 exit 1 @@ -2083,7 +2095,7 @@ require_orca_worktree_path_match_if_present() { # record with nothing live to return skips them rather than refusing. teardown_live_slot_path() { [ "$KIND" != secondmate ] || return 1 - is_treehouse_pool_slot "$PROJ" "$WT" || return 1 + fm_treehouse_pool_slot "$PROJ" "$WT" || return 1 canonical_existing_dir "$WT" } @@ -2163,6 +2175,72 @@ require_exclusive_task_worktree_slot() { require_exclusive_worktree_slot_record "$META" "$ID" "$STATE" "$slot" } +# Positive slot ownership, read from the claim the task that took the slot wrote +# into the slot itself (bin/fm-wake-lib.sh owns the claim and its states). +# +# The record scan above proves that no OTHER task record names this slot. It +# cannot prove that THIS record is not the stale one, because the task that took +# the slot next may leave no record this scan can reach: its own worker may have +# exited and its record been cleaned up, or it may belong to a home this machine +# does not register. The claim closes that gap from the other side - it names the +# task that actually took the slot, and it is written under the same project lock +# that allocates it - so a claim naming another task is proof the slot was +# reassigned after this record was written. +# +# A claim naming another task does not refuse: it means the slot is no longer +# this task's, so the record's own cleanup proceeds and every slot step is +# skipped (see the script header for why refusing would strand the record and +# why skipping discards nothing). Returns TEARDOWN_SLOT_REASSIGNED_RC for that +# state so each caller gates its slot steps on one determination; the claimant +# stays in FM_TREEHOUSE_SLOT_OWNER_ID and FM_TREEHOUSE_SLOT_OWNER_HOME. +# +# An absent claim proceeds as the slot's owner: a slot taken before claims +# existed, or already returned to the pool, carries none, and refusing those +# would strand every task in flight across the change for no evidence at all. +# Those keep exactly the record-scan protection they had before. +TEARDOWN_SLOT_REASSIGNED_RC=3 +require_owned_worktree_slot_record() { # <task-id> <worktree> + local record_id=$1 worktree=$2 marker + fm_treehouse_slot_owner_state "$worktree" "$record_id" + case "$FM_TREEHOUSE_SLOT_OWNER" in + mine|absent) return 0 ;; + other) + echo "warning: task $record_id's recorded worktree $worktree was reassigned to task $FM_TREEHOUSE_SLOT_OWNER_ID${FM_TREEHOUSE_SLOT_OWNER_HOME:+ (home $FM_TREEHOUSE_SLOT_OWNER_HOME)}, which claimed that pool slot after this record was written; that slot is no longer $record_id's, so its processes, copy, and claim are left untouched and only $record_id's own cleanup runs." >&2 + return "$TEARDOWN_SLOT_REASSIGNED_RC" + ;; + esac + marker=$(fm_treehouse_slot_owner_marker "$worktree" 2>/dev/null) || marker="beside $worktree" + echo "REFUSED: task $record_id's recorded worktree $worktree carries a slot-owner claim that cannot be read, so the slot cannot be proved to still be this task's; nothing was changed - not even with --force." >&2 + echo "Inspect or repair the claim file at $marker (task= and home= lines), then re-run teardown." >&2 + return 1 +} + +# The one ownership determination for this task's recorded slot. Every later +# step that would read or touch $WT consults teardown_owns_worktree, so a +# reassigned slot is skipped consistently rather than by each step's own guess. +TEARDOWN_SLOT_REASSIGNED=0 +TEARDOWN_SLOT_REASSIGNED_TO= +TEARDOWN_SLOT_REASSIGNED_HOME= +require_owned_task_worktree_slot() { + local slot rc=0 + slot=$(teardown_live_slot_path) || return 0 + require_owned_worktree_slot_record "$ID" "$slot" || rc=$? + case "$rc" in + 0) return 0 ;; + "$TEARDOWN_SLOT_REASSIGNED_RC") + TEARDOWN_SLOT_REASSIGNED=1 + TEARDOWN_SLOT_REASSIGNED_TO=$FM_TREEHOUSE_SLOT_OWNER_ID + TEARDOWN_SLOT_REASSIGNED_HOME=$FM_TREEHOUSE_SLOT_OWNER_HOME + return 0 + ;; + esac + return 1 +} + +teardown_owns_worktree() { + [ "$TEARDOWN_SLOT_REASSIGNED" != 1 ] +} + firstmate_home_has_treehouse_slot() { local home=$1 worktree_registered_for_project "$FM_ROOT" "$home" @@ -2630,7 +2708,7 @@ preflight_descendant_task_locks() { } preflight_descendant_treehouse_slots() { - local i state task_id meta kind backend target worktree project lock_path held + local i state task_id meta kind backend target worktree project lock_path held owner_rc for ((i=0; i < ${#DESCENDANT_TASK_IDS[@]}; i++)); do state=${DESCENDANT_TASK_STATES[$i]} task_id=${DESCENDANT_TASK_IDS[$i]} @@ -2643,7 +2721,7 @@ preflight_descendant_treehouse_slots() { if [ "$kind" = secondmate ] || [ "$backend" = orca ]; then continue fi - if ! is_treehouse_pool_slot "$project" "$worktree"; then + if ! fm_treehouse_pool_slot "$project" "$worktree"; then continue fi lock_path=$(fm_treehouse_project_lock_path "$project") || { @@ -2652,7 +2730,7 @@ preflight_descendant_treehouse_slots() { } held=0 [ "$TREEHOUSE_PROJECT_LOCK_HELD" != 1 ] || [ "$TREEHOUSE_PROJECT_LOCK" != "$lock_path" ] || held=1 - for target in "${DESCENDANT_TREEHOUSE_LOCK_PATHS[@]}"; do + for target in "${DESCENDANT_TREEHOUSE_LOCK_PATHS[@]+"${DESCENDANT_TREEHOUSE_LOCK_PATHS[@]}"}"; do [ "$target" != "$lock_path" ] || held=1 done if [ "$held" = 0 ]; then @@ -2676,11 +2754,17 @@ preflight_descendant_treehouse_slots() { if [ "$kind" = secondmate ] || [ "$backend" = orca ]; then continue fi - if ! is_treehouse_pool_slot "$project" "$worktree"; then + if ! fm_treehouse_pool_slot "$project" "$worktree"; then continue fi fm_backend_validate_task_endpoint "$meta" "$task_id" || return 1 require_exclusive_worktree_slot_record "$meta" "$task_id" "$state" "$worktree" || return 1 + owner_rc=0 + require_owned_worktree_slot_record "$task_id" "$worktree" || owner_rc=$? + case "$owner_rc" in + 0|"$TEARDOWN_SLOT_REASSIGNED_RC") ;; + *) return 1 ;; + esac done } @@ -2852,7 +2936,7 @@ preflight_firstmate_home_herdr_children() { # <home> } cleanup_firstmate_home_children() { - local home=$1 sub_state child_meta child_id child_t child_wt child_proj child_kind child_home child_backend child_orca_worktree_id child_return_rc child_busy_gen + local home=$1 sub_state child_meta child_id child_t child_wt child_proj child_kind child_home child_backend child_orca_worktree_id child_return_rc child_busy_gen child_owner_rc sub_state="$home/state" [ -d "$sub_state" ] || return 0 for child_meta in "$sub_state"/*.meta; do @@ -2909,22 +2993,36 @@ cleanup_firstmate_home_children() { fi fm_backend_remove_worktree "$child_backend" "$child_orca_worktree_id" || return 1 elif [ -n "$child_wt" ] && [ -d "$child_wt" ]; then - validate_child_worktree_for_removal "$child_wt" "$child_proj" >/dev/null || return 1 - rm -f "$child_wt/.claude/settings.local.json" "$child_wt/.opencode/plugins/fm-turn-end.js" \ - "$child_wt/.opencode/plugins/fm-busy-state.js" \ - "$child_wt/.fm-grok-turnend" "$child_wt/.fm-kimi-turnend" - if [ -n "$child_proj" ] && [ -d "$child_proj" ] && command -v treehouse >/dev/null 2>&1; then - if teardown_treehouse_return "$child_wt" "$child_proj" "child worktree"; then - : - else - child_return_rc=$? - if [ "$child_return_rc" -eq "$TEARDOWN_TREEHOUSE_LOCK_REFUSED" ]; then - return "$child_return_rc" + # The same ownership determination as the parent's own slot: a child + # slot reassigned to another task is not this child's to kill, reset, + # or return, so only its records are cleaned up. The preflight above + # already named the reassignment on stderr under the same lock. + child_owner_rc=0 + if fm_treehouse_pool_slot "$child_proj" "$child_wt"; then + require_owned_worktree_slot_record "$child_id" "$child_wt" 2>/dev/null || child_owner_rc=$? + fi + if [ "$child_owner_rc" -eq "$TEARDOWN_SLOT_REASSIGNED_RC" ]; then + : + elif [ "$child_owner_rc" -ne 0 ]; then + require_owned_worktree_slot_record "$child_id" "$child_wt" || return 1 + else + validate_child_worktree_for_removal "$child_wt" "$child_proj" >/dev/null || return 1 + rm -f "$child_wt/.claude/settings.local.json" "$child_wt/.opencode/plugins/fm-turn-end.js" \ + "$child_wt/.opencode/plugins/fm-busy-state.js" \ + "$child_wt/.fm-grok-turnend" "$child_wt/.fm-kimi-turnend" + if [ -n "$child_proj" ] && [ -d "$child_proj" ] && command -v treehouse >/dev/null 2>&1; then + if teardown_treehouse_return "$child_wt" "$child_proj" "child worktree"; then + fm_treehouse_slot_owner_release "$child_wt" "$child_id" + else + child_return_rc=$? + if [ "$child_return_rc" -eq "$TEARDOWN_TREEHOUSE_LOCK_REFUSED" ]; then + return "$child_return_rc" + fi + safe_rm_rf_child_worktree "$child_wt" "$child_proj" fi + else safe_rm_rf_child_worktree "$child_wt" "$child_proj" fi - else - safe_rm_rf_child_worktree "$child_wt" "$child_proj" fi fi remove_grok_turnend_auth "$sub_state" "$child_id" || return 1 @@ -2962,6 +3060,7 @@ remove_secondmate_registry_entry() { } require_exclusive_task_worktree_slot || exit 1 +require_owned_task_worktree_slot || exit 1 validate_pr_poll_cleanup "$STATE" "$ID" || exit 1 @@ -3069,7 +3168,7 @@ if [ "$BACKEND" = orca ] && [ "$KIND" != scout ] && [ "$KIND" != secondmate ] && ORCA_PATH_MATCH_VERIFIED=1 fi -if [ -d "$WT" ] && [ "$FORCE" != "--force" ]; then +if teardown_owns_worktree && [ -d "$WT" ] && [ "$FORCE" != "--force" ]; then if validate_worktree_teardown_safety; then : else @@ -3184,9 +3283,11 @@ fi # kind=secondmate: a secondmate home's own runtime lifecycle is owned by the # dedicated process-event and firstmate-home removal machinery further below, # not by task-worktree cleanup. -if [ "$KIND" != secondmate ]; then +if [ "$KIND" != secondmate ] && teardown_owns_worktree; then conclude_task_no_mistakes_run "$WT" reap_task_worktree_processes worktree "$WT" "$TASK_TMP" +elif [ "$KIND" != secondmate ]; then + reap_task_worktree_processes tasktmp "$TASK_TMP" fi # Fix 3 (see script header): sweep remote job workers abandoned by an already @@ -3212,6 +3313,8 @@ if [ "$BACKEND" = orca ] && [ "$KIND" != secondmate ]; then fi [ -z "$T_ORCA" ] || fm_backend_kill "$BACKEND" "$T" "$(meta_value "$META" zellij_tab_id)" "fm-$ID" 2>/dev/null || true fm_backend_remove_worktree "$BACKEND" "$ORCA_WORKTREE_ID" +elif [ "$KIND" != secondmate ] && ! teardown_owns_worktree; then + : elif [ -d "$WT" ] && [ "$KIND" != secondmate ]; then branch=$(git -C "$WT" rev-parse --abbrev-ref HEAD 2>/dev/null || echo HEAD) if [ "$branch" != "HEAD" ]; then @@ -3234,6 +3337,11 @@ elif [ -d "$WT" ] && [ "$KIND" != secondmate ]; then echo "error: treehouse return failed for worktree $WT; teardown aborted" >&2 exit 1 } + # The slot is back in the pool, so this task's claim on it is spent. Dropping + # it here - and only after a return that succeeded - keeps a returned slot + # unclaimed until its next holder claims it, and leaves the claim in place + # whenever the return did not actually happen. + fm_treehouse_slot_owner_release "$WT" "$ID" fi HERDR_PRESENTATION_JOURNAL="$STATE/$ID.herdr-presentation" @@ -3397,7 +3505,9 @@ if [ -d "$STATE" ]; then fi if [ "$TEARDOWN_LEGACY_ACCEPTED" = 1 ]; then echo "teardown $ID complete (window $T, worktree $WT, legacy record accepted without spawn_gen: endpoint $TEARDOWN_LEGACY_ENDPOINT, incarnation $TEARDOWN_META_SPAWN_GEN)" -else +elif teardown_owns_worktree; then echo "teardown $ID complete (window $T, worktree $WT)" +else + echo "teardown $ID complete (window $T; pool slot $WT left to task $TEARDOWN_SLOT_REASSIGNED_TO${TEARDOWN_SLOT_REASSIGNED_HOME:+ (home $TEARDOWN_SLOT_REASSIGNED_HOME)}, which it was reassigned to)" fi backlog_refresh_reminder diff --git a/bin/fm-wake-lib.sh b/bin/fm-wake-lib.sh index 1ee40021360..54d3770ace3 100755 --- a/bin/fm-wake-lib.sh +++ b/bin/fm-wake-lib.sh @@ -1210,6 +1210,121 @@ fm_treehouse_project_lock_path() { # <project-dir> printf '%s/.treehouse-project-%s.lock\n' "$root/state" "$hash" } +# A Treehouse slot has the managed pool's fixed <pool>/<slot>/<repo> layout. +# Require both its pool state and the same Git common directory as the recorded +# project; an ordinary linked worktree is not evidence that Treehouse owns it. +fm_treehouse_pool_slot() { # <project-dir> <worktree> + local project=$1 worktree=$2 slot pool state project_common slot_common + [ -d "$project" ] && [ -d "$worktree" ] || return 1 + slot=$(CDPATH='' cd -- "$worktree" 2>/dev/null && pwd -P) || return 1 + pool=$(dirname "$(dirname "$slot")") + state="$pool/treehouse-state.json" + [ -f "$state" ] && [ ! -L "$state" ] || return 1 + project_common=$(git -C "$project" rev-parse --path-format=absolute --git-common-dir 2>/dev/null) || return 1 + slot_common=$(git -C "$slot" rev-parse --path-format=absolute --git-common-dir 2>/dev/null) || return 1 + project_common=$(CDPATH='' cd -- "$project_common" 2>/dev/null && pwd -P) || return 1 + slot_common=$(CDPATH='' cd -- "$slot_common" 2>/dev/null && pwd -P) || return 1 + [ "$project_common" = "$slot_common" ] +} + +# Slot-owner claim: which task a Treehouse pool slot currently belongs to. +# +# Treehouse can record ownership durably: `treehouse get --lease --lease-holder` +# reserves a slot under a label until `treehouse return --if-lease-holder` +# releases it, and Firstmate uses exactly that for secondmate homes +# (bin/fm-home-seed.sh). Crewmate spawns do not take that path: they acquire +# their slot through the interactive pane-driven `treehouse get`, whose state +# entry is a live process lease (owner_pid plus owner_started_at, and `treehouse +# status` reports in-use from the processes actually running under the path). +# That answers "is anything running here", never "which task owns this", and it +# is released by the very event that makes a task record stale - the worker +# exiting - so a slot whose lease has lapsed reads identical whether it is still +# this task's or has since been handed to another one. Firstmate therefore keeps +# its own claim on top: one file naming the task that took the slot, written by +# bin/fm-spawn.sh under the same project lock that allocates the slot and +# released by bin/fm-teardown.sh when the slot goes back to the pool. Moving +# crewmate spawns onto the durable lease is separate follow-up work. +# +# The claim lives at <pool>/<slot>/.fm-slot-owner - a sibling of the repo +# checkout rather than a file inside it - so claiming a slot can never dirty the +# copy teardown's landed-work checks inspect, and a returned slot carries no +# untracked leftover from it. +fm_treehouse_slot_owner_marker() { # <worktree> + local worktree=$1 slot + slot=$(CDPATH='' cd -- "$worktree" 2>/dev/null && pwd -P) || return 1 + printf '%s/.fm-slot-owner\n' "$(dirname "$slot")" +} + +# Claim a pool slot for a task, replacing whatever the previous holder left. +# The rename is atomic, so a reader either sees the old claim or the new one. +fm_treehouse_slot_owner_claim() { # <worktree> <task-id> <home> + local worktree=$1 id=$2 home=$3 marker tmp + [ -n "$id" ] || return 1 + marker=$(fm_treehouse_slot_owner_marker "$worktree") || return 1 + # Only a plain claim file may be replaced: renaming onto a directory would + # move the new claim inside it and leave the slot reading as unclaimable. + if { [ -e "$marker" ] || [ -L "$marker" ]; } \ + && { [ ! -f "$marker" ] || [ -L "$marker" ]; }; then + return 1 + fi + tmp="$marker.tmp.${BASHPID:-$$}" + rm -f "$tmp" || return 1 + { + printf 'task=%s\n' "$id" + printf 'home=%s\n' "$home" + } > "$tmp" 2>/dev/null || { rm -f "$tmp"; return 1; } + mv -f "$tmp" "$marker" 2>/dev/null || { rm -f "$tmp"; return 1; } +} + +# Read the claim on a pool slot and compare it with a task id. +# Sets FM_TREEHOUSE_SLOT_OWNER to one of: +# mine - the claim names this task +# other - the claim names a different task, so the slot was reassigned +# absent - no claim: the slot was taken before claims existed, or returned since +# unsafe - a claim file exists but cannot be read as a claim +# FM_TREEHOUSE_SLOT_OWNER_ID and FM_TREEHOUSE_SLOT_OWNER_HOME carry the recorded +# claimant as evidence. The home is reported, never matched: a home that moved +# must not turn a task's own slot into a refusal. +fm_treehouse_slot_owner_state() { # <worktree> <task-id> + local worktree=$1 id=$2 marker line owner_id='' owner_home='' + FM_TREEHOUSE_SLOT_OWNER=unsafe + FM_TREEHOUSE_SLOT_OWNER_ID= + FM_TREEHOUSE_SLOT_OWNER_HOME= + marker=$(fm_treehouse_slot_owner_marker "$worktree") || return 0 + if [ ! -e "$marker" ] && [ ! -L "$marker" ]; then + FM_TREEHOUSE_SLOT_OWNER=absent + return 0 + fi + [ -f "$marker" ] && [ ! -L "$marker" ] || return 0 + while IFS= read -r line || [ -n "$line" ]; do + case "$line" in + task=*) owner_id=${line#task=} ;; + home=*) owner_home=${line#home=} ;; + esac + done < "$marker" || return 0 + [ -n "$owner_id" ] || return 0 + # shellcheck disable=SC2034 # Output globals, read by the sourcing caller. + FM_TREEHOUSE_SLOT_OWNER_ID=$owner_id + # shellcheck disable=SC2034 # Output globals, read by the sourcing caller. + FM_TREEHOUSE_SLOT_OWNER_HOME=$owner_home + if [ "$owner_id" = "$id" ]; then + FM_TREEHOUSE_SLOT_OWNER=mine + else + FM_TREEHOUSE_SLOT_OWNER=other + fi +} + +# Drop a task's own claim once its slot is back in the pool. Never removes +# another task's claim, so a misdirected release cannot strip the evidence that +# protects the slot's real owner. +fm_treehouse_slot_owner_release() { # <worktree> <task-id> + local worktree=$1 id=$2 marker + fm_treehouse_slot_owner_state "$worktree" "$id" + [ "$FM_TREEHOUSE_SLOT_OWNER" = mine ] || return 0 + marker=$(fm_treehouse_slot_owner_marker "$worktree") || return 0 + rm -f "$marker" 2>/dev/null || true +} + fm_failure_episode_reset() { local state=$1 mode=${2:-acquire} lock current pid acquired=0 path lock="$state/.turnend-claude-blocks.lock" diff --git a/docs/architecture.md b/docs/architecture.md index 43b5272b35b..e1f8bb57be1 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -327,7 +327,8 @@ Every GitHub refusal states what it could not observe as plainly as what it did, A confirmed merge leaves a durable role-routed outcome instead of living only in the merging agent's memory, and [`bin/fm-merge-outcome-lib.sh`](../bin/fm-merge-outcome-lib.sh)'s header owns its destination, shape, identity, normal-case deduplication, and at-least-once recovery. The same emitter handles a merge firstmate performed and one its poll detected, while the watcher immediately delivers the emitter's local actionable poll row. Teardown is fail-closed for ship worktrees: dirty worktrees refuse, and committed work must be landed before the worktree is returned. -A pool worktree is only returned after teardown passes the slot-ownership proof: a contradictory task record or supported live endpoint refuses without touching either task, and no discard authority relaxes that. +A pool worktree is only returned after teardown passes the slot-ownership proof: a contradictory task record or a supported live endpoint refuses without touching either task, and no discard authority relaxes that. +A slot's own owner claim, written by the spawn that takes it under the allocation lock and owned by [`bin/fm-wake-lib.sh`](../bin/fm-wake-lib.sh), covers a slot reassigned to a task that left no record the scan could reach: a claim naming a different task releases nothing - teardown warns, names the claimant, and finishes only the task's own cleanup - because Treehouse's own live process lease cannot answer ownership once the worker's exit releases it. Allocation and return serialize on one project lock per machine-local Firstmate tree: every home reachable through local parent links shares that lock, and a home seeded from another machine anchors its own, because a lock taken on this filesystem is neither held nor observable across that boundary. Before the worktree is returned, teardown concludes the task's own no-mistakes run when it is parked at a gate, including a run whose head the task copy cannot resolve - the shared runs-ledger continuation proof is the only recognition for that case, so cleanup never orphans a parked run the pipeline advanced past the submitted head. [`bin/fm-teardown.sh`](../bin/fm-teardown.sh)'s header owns the landed-work proofs, slot-ownership proof, PR-discovery fallback, pre-teardown run conclusion, and stale-lock recovery procedure; [`tests/fm-teardown-endpoint-safety.test.sh`](../tests/fm-teardown-endpoint-safety.test.sh) and [`tests/fm-secondmate-safety.test.sh`](../tests/fm-secondmate-safety.test.sh) pin the slot-collision boundary. diff --git a/tests/fm-spawn-pool-base-freshen.test.sh b/tests/fm-spawn-pool-base-freshen.test.sh index aeb5a193149..2d39f3607ca 100755 --- a/tests/fm-spawn-pool-base-freshen.test.sh +++ b/tests/fm-spawn-pool-base-freshen.test.sh @@ -676,7 +676,75 @@ test_stale_pin_beside_other_dirt_reports_one_verdict() { pass "a stale pin beside other dirt yields the conservative refusal alone, with no stale-pin line" } +# Re-lay a case's pooled worktree as a managed Treehouse slot: <pool>/<slot>/<repo> +# with the pool's state file beside the slot, which is the shape fm-spawn claims +# for its task. Rewrites POOL_DIR to the relocated checkout. +lay_out_as_pool_slot() { + local slot_root="$CASE_DIR/slots" + mkdir -p "$slot_root/1" + git -C "$PROJECT_DIR" worktree move "$POOL_DIR" "$slot_root/1/project" + printf '{"worktrees":[{"name":"1","path":"%s"}]}\n' "$slot_root/1/project" \ + > "$slot_root/treehouse-state.json" + POOL_DIR="$slot_root/1/project" + SLOT_CLAIM="$slot_root/1/.fm-slot-owner" +} + +# The spawn side of the slot-owner claim that bin/fm-teardown.sh later reads: +# a launched task's claim names it, a slot that cannot be claimed refuses before +# anything is published, and an abort while the allocation lock is still held +# leaves no claim naming a task with no record. +test_pool_slot_claim_follows_the_spawn_outcome() { + local rec id out status before + + id='pool-slot-claim-r1' + rec=$(make_case slot-claim "$id") + read_case_record "$rec" + lay_out_as_pool_slot + out=$(run_spawn "$id" --scout) + status=$? + expect_code 0 "$status" "spawn from a Treehouse slot should launch"$'\n'"$out" + assert_grep "worktree=$POOL_DIR" "$HOME_DIR/state/$id.meta" \ + "spawn did not publish the relocated slot as its worktree" + [ -f "$SLOT_CLAIM" ] || fail "spawn left its Treehouse slot unclaimed: $out" + grep -Fxq -- "task=$id" "$SLOT_CLAIM" \ + || fail "the slot claim does not name the spawned task: $(cat "$SLOT_CLAIM")" + grep -Fxq -- "home=$HOME_DIR" "$SLOT_CLAIM" \ + || fail "the slot claim does not name the spawning home: $(cat "$SLOT_CLAIM")" + + id='pool-slot-unclaimable-r1' + rec=$(make_case slot-unclaimable "$id") + read_case_record "$rec" + lay_out_as_pool_slot + mkdir -p "$SLOT_CLAIM" + before=$(git -C "$POOL_DIR" rev-parse HEAD) + out=$(run_spawn "$id" --scout) + status=$? + [ "$status" -ne 0 ] || fail "spawn launched a worker on a slot it could not claim" + assert_contains "$out" "could not claim Treehouse pool slot" \ + "spawn did not name the unclaimable slot as the reason" + [ -d "$SLOT_CLAIM" ] || fail "spawn replaced the directory blocking its slot claim" + [ ! -e "$HOME_DIR/state/$id.meta" ] || fail "spawn published a record for an unclaimable slot" + [ "$(git -C "$POOL_DIR" rev-parse HEAD)" = "$before" ] \ + || fail "spawn moved the slot's HEAD after failing to claim it" + + id='pool-slot-claim-aborted-r1' + rec=$(make_originless_case slot-claim-aborted "$id") + read_case_record "$rec" + lay_out_as_pool_slot + git -C "$POOL_DIR" config remote.origin.fetch '+refs/heads/*:refs/remotes/origin/*' + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + [ "$status" -ne 0 ] || fail "spawn succeeded despite an unusable origin on the slot" + assert_contains "$out" "could not fetch origin" \ + "the aborted spawn did not refuse on its unusable origin" + [ ! -e "$HOME_DIR/state/$id.meta" ] || fail "the aborted spawn published task metadata" + [ ! -e "$SLOT_CLAIM" ] && [ ! -L "$SLOT_CLAIM" ] \ + || fail "the aborted spawn left a slot claim naming a task with no record: $(cat "$SLOT_CLAIM")" + pass "a Treehouse slot claim names the launched task, refuses when unclaimable, and is dropped by a locked abort" +} + test_remote_seeded_home_spawns_from_treehouse_pool +test_pool_slot_claim_follows_the_spawn_outcome test_linked_spawning_home_rejects_primary_before_refresh test_stale_pool_base_refreshes_before_branching test_non_main_default_branch_refreshes_before_branching diff --git a/tests/fm-teardown-endpoint-safety.test.sh b/tests/fm-teardown-endpoint-safety.test.sh index 2d185907abb..100f04e6785 100755 --- a/tests/fm-teardown-endpoint-safety.test.sh +++ b/tests/fm-teardown-endpoint-safety.test.sh @@ -48,6 +48,11 @@ mark_case_as_treehouse_pool() { # <case> : > "$dir/worktree/sentinel" } +claim_pool_slot() { # <case> <task-id> [home] + local dir=$1 id=$2 home=${3:-$1/home} + printf 'task=%s\nhome=%s\n' "$id" "$home" > "$dir/pool/1/.fm-slot-owner" +} + run_case() { # <case> <id> local dir=$1 id=$2 FM_HOME="$dir/home" FM_ROOT_OVERRIDE="$ROOT" \ @@ -825,6 +830,148 @@ test_remote_layout_homes_serialize_on_one_project_lock() { pass "Treehouse project locking still serializes two homes across the remote-seeded boundary" } +# The slot-reuse sequence with only ONE discoverable record: the finished task's +# worker exited, its slot was granted to another task, and that task leaves no +# record this home can enumerate. Nothing in the record scan contradicts the +# stale worktree= line, so the slot's own owner claim is the only evidence that +# it was reassigned. The slot is no longer this task's, so teardown finishes the +# task's own cleanup and leaves the slot - its worker, its copy, its claim - +# exactly as it found it. +assert_reassigned_slot_left_alone() { # <case> <id> <other> <description> + local dir=$1 id=$2 other=$3 description=$4 + assert_absent "$dir/home/state/$id.meta" "$description: the stale task's own record was not removed" + assert_present "$dir/pool/1/.fm-slot-owner" "$description: another task's slot claim was removed" + assert_contains "$(cat "$dir/pool/1/.fm-slot-owner")" "task=$other" \ + "$description: another task's slot claim was rewritten" + assert_present "$dir/pool/1/project/.git" "$description: the reassigned slot's checkout was removed" + ! grep -Fq "treehouse <return>" "$dir/runtime.log" \ + || fail "$description: the reassigned slot was returned to the pool: $(cat "$dir/runtime.log")" + assert_contains "$(cat "$dir/stderr")" "$other" \ + "$description: the warning should name the task the slot was reassigned to" + assert_contains "$(cat "$dir/stderr")" "reassigned" \ + "$description: the warning should name the reassignment as the cause" +} + +test_reassigned_pool_slot_finishes_own_cleanup_without_touching_the_slot() { + local dir id=stale-task other=reassigned-task worker rc + + # Dirty slot, --force, and a live worker inside it: --force authorizes + # discarding this task's unlanded work, which is already gone with the slot, + # never the other task's live work. + dir=$(make_case slot-reassigned) + mark_case_as_treehouse_pool "$dir" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + claim_pool_slot "$dir" "$other" "$dir/other-home" + # Staged in this shell, not a command substitution: a background child of a + # $(...) subshell does not outlive it, and the point of this worker is to be + # alive in the slot while teardown runs. + ( cd "$dir/worktree" && exec sleep 30 ) & + worker=$! + + set +e + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" + rc=$? + set -e + + [ "$rc" -eq 0 ] || fail "teardown of a task whose slot was reassigned failed: $(cat "$dir/stderr")" + kill -0 "$worker" 2>/dev/null || fail "teardown killed the worker holding the reassigned pool slot" + assert_present "$dir/worktree/sentinel" "teardown reset a pool slot another task had claimed" + assert_reassigned_slot_left_alone "$dir" "$id" "$other" "dirty reassigned slot with --force" + assert_contains "$(cat "$dir/stderr")" "$dir/other-home" \ + "the warning should name the claimant's home" + kill "$worker" 2>/dev/null || true + wait "$worker" 2>/dev/null || true + + # The same reassignment on a CLEAN slot: a landed ship task torn down without + # --force, which is the shape of the real incident. A clean, fully landed copy + # passes every unlanded-work check, so only the ownership determination can + # keep this slot out of the pool; a guard keyed off dirtiness would return it + # and destroy the live task's copy. + dir=$(make_case slot-reassigned-clean) + mark_case_as_treehouse_pool "$dir" + rm -f "$dir/worktree/sentinel" + [ -z "$(git -C "$dir/worktree" status --porcelain)" ] \ + || fail "clean-slot fixture is not clean: $(git -C "$dir/worktree" status --porcelain)" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=ship" + claim_pool_slot "$dir" "$other" "$dir/other-home" + ( cd "$dir/worktree" && exec sleep 30 ) & + worker=$! + + set +e + FM_HOME="$dir/home" FM_ROOT_OVERRIDE="$ROOT" \ + FM_RUNTIME_LOG="$dir/runtime.log" PATH="$dir/fakebin:$PATH" \ + "$TEARDOWN" "$id" > "$dir/stdout" 2> "$dir/stderr" + rc=$? + set -e + [ "$rc" -eq 0 ] || fail "teardown of a clean ship task whose slot was reassigned failed: $(cat "$dir/stderr")" + kill -0 "$worker" 2>/dev/null || fail "teardown killed the worker holding the clean reassigned pool slot" + assert_reassigned_slot_left_alone "$dir" "$id" "$other" "clean reassigned slot without --force" + kill "$worker" 2>/dev/null || true + wait "$worker" 2>/dev/null || true + + # A claim that exists but cannot be read as a claim proves nothing either way, + # so it refuses rather than guessing the slot is still this task's. + dir=$(make_case slot-claim-unreadable) + mark_case_as_treehouse_pool "$dir" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + printf 'not-a-claim\n' > "$dir/pool/1/.fm-slot-owner" + + set +e + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "teardown returned a pool slot whose claim could not be read" + assert_present "$dir/worktree/sentinel" "teardown reset a pool slot whose claim could not be read" + assert_present "$dir/pool/1/.fm-slot-owner" "teardown removed an unreadable slot claim" + assert_present "$dir/home/state/$id.meta" "teardown removed the task record on an unreadable claim" + [ ! -s "$dir/runtime.log" ] \ + || fail "teardown reached the runtime on an unreadable slot claim: $(cat "$dir/runtime.log")" + assert_contains "$(cat "$dir/stderr")" "$dir/pool/1/.fm-slot-owner" \ + "unreadable-claim refusal should name the claim file to inspect" + + pass "fm-teardown: a pool slot claimed by another task is left alone while the task's own cleanup finishes" +} + +# The two states that must never become a false refusal: the task's own claim, +# and no claim at all (a slot taken before claims existed, or already returned). +test_own_and_absent_slot_claims_still_tear_down() { + local dir id=owned-task + + dir=$(make_case slot-claim-own) + mark_case_as_treehouse_pool "$dir" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + claim_pool_slot "$dir" "$id" + + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" \ + || fail "teardown of a task holding its own slot claim failed: $(cat "$dir/stderr")" + assert_absent "$dir/home/state/$id.meta" "own-claim teardown left the task record" + assert_absent "$dir/pool/1/.fm-slot-owner" "own-claim teardown left its spent slot claim behind" + grep -Fq "treehouse <return>" "$dir/runtime.log" \ + || fail "own-claim teardown did not return its own pool slot: $(cat "$dir/runtime.log")" + + dir=$(make_case slot-claim-absent) + mark_case_as_treehouse_pool "$dir" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" \ + || fail "teardown of an unclaimed slot failed: $(cat "$dir/stderr")" + assert_absent "$dir/home/state/$id.meta" "unclaimed-slot teardown left the task record" + grep -Fq "treehouse <return>" "$dir/runtime.log" \ + || fail "unclaimed-slot teardown did not return its pool slot: $(cat "$dir/runtime.log")" + + pass "fm-teardown: a task's own slot claim, and an unclaimed slot, both still tear down" +} + test_invalid_endpoint_records_refuse_before_mutation test_control_lock_contention_refuses_before_mutation test_non_pool_teardown_ignores_task_set_lock @@ -837,6 +984,8 @@ test_bare_relative_origin_shares_project_lock_with_clone test_reused_pool_slot_refuses_before_touching_the_other_task test_cross_home_pool_slot_collision_refuses test_sole_slot_record_still_tears_down +test_reassigned_pool_slot_finishes_own_cleanup_without_touching_the_slot +test_own_and_absent_slot_claims_still_tear_down test_recorded_endpoint_that_changed_directory_still_tears_down test_project_lock_anchors_at_the_local_root_across_home_layouts test_remote_seeded_home_returns_its_uncontested_slot From 9e1e85e2fa15663d2dcbc9e160a8642d26f9f23e Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Fri, 11 Sep 2026 16:41:55 -0700 Subject: [PATCH 12/31] docs(AGENTS): keep brief-fill from widening the captain's ask (#4247) The reviewer treats Captain's intent as acceptance criteria, so a widened ask there drives over-built work; the spec should carry only what the ask requires. --- AGENTS.md | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/AGENTS.md b/AGENTS.md index 16ec3a472b9..b709a2b9c84 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -533,7 +533,8 @@ Preserve durable structured identifiers, dependencies, and completion artifact l ## 11. Crewmate briefs `bin/fm-brief.sh` and its help own scaffold syntax, generated variants, status protocol, delivery-mode definitions of done, and exact safety mechanics. -Use its scaffold as the contract, then fill `## Captain's intent` (`{TASK}`) with the captain's own ask plus the context needed to read it, including the substance of any report, decision, or PR the ask refers to, and fill `## Firstmate spec` (`{FIRSTMATE_SPEC}`) with Firstmate's build instructions. +Use its scaffold as the contract, then fill `## Captain's intent` (`{TASK}`) with the captain's own ask and any boundary the captain stated, plus the context needed to read it, including the substance of any report, decision, or PR the ask refers to; never widen the ask there into a general goal or an enumerated coverage list, because the reviewer treats that subsection as acceptance criteria. +Fill `## Firstmate spec` (`{FIRSTMATE_SPEC}`) with only the build instructions that ask requires, naming what stays out of scope when the ask is narrow; a generalization, consistency sweep, or extra hardening the captain did not ask for is follow-up work to note, not scope to add. `bin/fm-dod-lib.sh` owns what a no-mistakes worker may pass as `--intent` and its rule that the string must be self-sufficient. Keep additions task-specific rather than repeating lifecycle instructions, and alter generated sections only when the task genuinely differs from the standard shape. From 46ff99ccf9c7528f3c6906d3c427c1af83f99f5b Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Fri, 11 Sep 2026 16:42:37 -0700 Subject: [PATCH 13/31] fix: identify underway tasks and sort charted work (#4245) * feat(bearings): name the Underway rows and order Charted Next newest filed first The fleet board's Underway rows led with the run status alone, so a scan told the captain where a pipeline stood but never which task the row was, and Charted Next rendered in backlog order rather than by when work was filed. The snapshot now projects the durable task name onto every in_flight row - from this home's backlog title, and from a secondmate home's own ledger for an active child - and the durable filed date onto every gate. The board's Underway row leads with that name and keeps the run status on its second line, and Charted Next renders newest filed first, with rows carrying no comparable date keeping their payload order after every dated row. The payload validator requires an explicit name marker on every Underway row and refuses a filed value that is not an ISO date, so the board can never sort on garbage or invent a label. * no-mistakes(review): Fix Bearings labels, bounds, and filed validation * no-mistakes(review): Fix Bearings identifiers and eligible queue bounds * no-mistakes(document): Document Bearings labels and newest-first bounds * no-mistakes(ci): Updated the stock macOS Bash CI expectation from 56 to 59 Bearings tests. Verified the suite under /bin/bash 3.2: all 59 tests pass. git diff --check also passes --- .agents/skills/bearings/SKILL.md | 5 + .../bearings/assets/board-template.html | 20 ++- .github/workflows/ci.yml | 4 +- bin/fm-bearings-board.sh | 21 ++- bin/fm-bearings-snapshot.sh | 41 +++++- bin/fm-fleet-snapshot.sh | 21 ++- tests/assets/board-render-harness.mjs | 35 +++-- tests/fm-bearings-board-render.test.sh | 84 ++++++++++- tests/fm-bearings-board.test.sh | 14 ++ tests/fm-bearings-snapshot.test.sh | 136 ++++++++++++++++++ 10 files changed, 349 insertions(+), 32 deletions(-) diff --git a/.agents/skills/bearings/SKILL.md b/.agents/skills/bearings/SKILL.md index 811c638bffc..03d22ee708d 100644 --- a/.agents/skills/bearings/SKILL.md +++ b/.agents/skills/bearings/SKILL.md @@ -101,6 +101,11 @@ Compose the payload from the same snapshot with the same ranking judgment as the - When the card's task is a captain-gated WORK item (the answer should free it to proceed rather than complete it), set the card's `close: "release"` so the answer lifts the hold instead of closing the task; question-shaped items omit it. - A Charted Next row's optional `kind` separates work from alarms: omit it (or set `"queued"`) for real queued work, and set `"warning"` on every action-free fleet-integrity notice - the `(main-inventory)` gate, an unavailable secondmate home, and an inventory-mismatch repair notice. The board badges a warning row `needs repair` instead of `waiting` and leaves it out of the Charted Next count, so those rows never read as dispatchable queued work. - `charted_more` counts omitted queued rows only, while `charted_warning_more` counts omitted warning rows only; keep both counts separate whenever the board payload truncates Charted Next. +- Every Underway row copies the task-identifying `in_flight.name` from the snapshot into an explicit `name` field, which the board leads with while keeping the run status on its second line. + The snapshot command's header owns its durable-title-or-id normalization; never replace the projected label with run status or invent another label. +- Every Charted Next row copies the snapshot gate's durable filed date into `filed`, and the board orders the section by it, newest filed first. + Follow `bin/fm-bearings-board.sh`'s payload contract for the accepted format. + Omit it or pass null for a row with no durable filed date - the main-inventory warning, an unavailable secondmate home, or a queued row filed before dates were recorded - and the board keeps those rows in payload order after every dated row. - Every Captain's Call item and every Underway, Recently Landed, and Charted Next row carries an explicit `repo` field. Fill it from the snapshot and task records wherever known; use null or an empty string only as the deliberate genuinely-no-repo marker, in which case the template may show the internal id. Ids otherwise stay in the payload only as the routing channel, and composed reasons name blockers in plain words. Run `build` once after composing the payload. diff --git a/.agents/skills/bearings/assets/board-template.html b/.agents/skills/bearings/assets/board-template.html index 614bef7426b..987db2d8c03 100644 --- a/.agents/skills/bearings/assets/board-template.html +++ b/.agents/skills/bearings/assets/board-template.html @@ -439,6 +439,17 @@ so every count of queued work excludes them. */ function isWarning(t) { return t && t.kind === "warning"; } function chartedQueued(rows) { return (rows || []).filter(function (t) { return !isWarning(t); }); } + /* Charted Next reads newest filed first, so the most recently filed upcoming + work is at the top. ISO filed dates compare as text; a row with no + comparable date keeps its payload order after every dated row. */ + function chartedOrder(rows) { + var dated = [], undated = []; + (rows || []).forEach(function (t) { + if (t && typeof t.filed === "string" && t.filed) dated.push(t); else undated.push(t); + }); + dated.sort(function (a, b) { return a.filed < b.filed ? 1 : (a.filed > b.filed ? -1 : 0); }); + return dated.concat(undated); + } var chartedMoreQueued = data.charted_more || 0; var chartedMoreWarnings = data.charted_warning_more || 0; function utf8ByteLength(text) { return new TextEncoder().encode(text).length; } @@ -617,10 +628,13 @@ var row = el("div", "bb-row"); row.appendChild(badge(t.state === "working" ? "online" : "info", t.state)); var main = el("div", "bb-row__main"); - main.appendChild(el("div", "bb-row__title", t.doing)); + /* the snapshot's durable name-or-id label leads the row so a scan says + WHICH task this is; the run status keeps its place on the second line */ + main.appendChild(el("div", "bb-row__title", t.name)); /* captain-facing rows name the repo; the internal task id shows only when no repo is known */ - main.appendChild(el("div", "bb-row__sub", t.kind + " · " + (t.repo || t.id))); + main.appendChild(el("div", "bb-row__sub", + t.doing + " · " + t.kind + " · " + (t.repo || t.id))); row.appendChild(main); uw.appendChild(row); }); @@ -664,7 +678,7 @@ if (!chartedQueued(data.charted).length && !chartedMoreQueued) { ch.appendChild(el("div", "bb-empty", "Nothing is queued.")); } - data.charted.forEach(function (t) { + chartedOrder(data.charted).forEach(function (t) { var row = el("div", "bb-row"); if (t.dispatchable && !isWarning(t)) { anyPickable = true; diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index dc681e1e0d0..5dcac6dff5a 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -414,8 +414,8 @@ jobs: bearings_output=$(/bin/bash tests/fm-bearings-snapshot.test.sh) printf '%s\n' "$bearings_output" bearings_count=$(printf '%s\n' "$bearings_output" | grep -c '^ok - ') - [ "$bearings_count" -eq 56 ] || { - echo "::error::expected 56 Bearings tests, got $bearings_count" + [ "$bearings_count" -eq 59 ] || { + echo "::error::expected 59 Bearings tests, got $bearings_count" exit 1 } diff --git a/bin/fm-bearings-board.sh b/bin/fm-bearings-board.sh index b25ad5e9c10..2cb9506d721 100755 --- a/bin/fm-bearings-board.sh +++ b/bin/fm-bearings-board.sh @@ -72,6 +72,13 @@ # the template may display the routing id. Anything else refuses before the # existing board is touched. # +# Every Underway row likewise carries a non-empty `name`: the durable task name +# when known, otherwise its durable identifier. +# A Charted Next row MAY carry `filed`, the durable filed date (YYYY-MM-DD, or +# that date with a UTC timestamp) the template orders the section by, newest +# first; a row with no comparable date keeps its payload order after every dated +# row. Anything else in that field refuses rather than sorting on garbage. +# # The board path is stable - $FM_HOME/.lavish/bearings-board.html - so a # re-invocation rebuilds the same file in place, which keeps the same Lavish # session URL and the same canonical process-event source id. Injection escapes @@ -109,6 +116,17 @@ validate_payload() { # <data.json> def nonempty_string: type == "string" and length > 0; def slug($max): type == "string" and test("^[A-Za-z0-9._-]{1," + ($max | tostring) + "}$"); def repo_marker: has("repo") and (.repo == null or (.repo | type == "string")); + def name_marker: has("name") and (.name | nonempty_string); + def valid_filed: + . as $filed + | type == "string" + and test("^[0-9]{4}-[0-9]{2}-[0-9]{2}(T[0-9]{2}:[0-9]{2}:[0-9]{2}Z)?$") + and (if test("T") + then try ((fromdateiso8601 | strftime("%Y-%m-%dT%H:%M:%SZ")) == $filed) catch false + else try (((. + "T00:00:00Z") | fromdateiso8601 | strftime("%Y-%m-%d")) == $filed) catch false + end); + def optional_filed: + (has("filed") | not) or (.filed == null) or (.filed | valid_filed); def optional_string($name): (has($name) | not) or (.[$name] | type == "string"); def optional_https_url($name): (has($name) | not) @@ -152,7 +170,7 @@ validate_payload() { # <data.json> and ([.options[].value] | index("reconcile") == null) and (if .type == "merge" then (.risk | nonempty_string) else true end); def underway_item: - type == "object" and repo_marker and (.id | nonempty_string) + type == "object" and repo_marker and name_marker and (.id | nonempty_string) and (.state | nonempty_string) and (.doing | nonempty_string) and (.kind | nonempty_string); def landed_item: type == "object" and repo_marker and (.id | nonempty_string) @@ -164,6 +182,7 @@ validate_payload() { # <data.json> and (.title | nonempty_string) and (.reason | type == "string") and (.dispatchable | type == "boolean") and ((has("kind") | not) or (.kind == "queued" or .kind == "warning")) + and optional_filed and (if .kind == "warning" then .dispatchable == false else true end); type == "object" and (.schema == $schema) diff --git a/bin/fm-bearings-snapshot.sh b/bin/fm-bearings-snapshot.sh index 33e64835f16..154019977c1 100755 --- a/bin/fm-bearings-snapshot.sh +++ b/bin/fm-bearings-snapshot.sh @@ -25,7 +25,10 @@ # decisions from report or visual-review prose or reimplements snapshot semantics. # Underway (in_flight) projects every main live worker plus every active child # from every readable secondmate ledger, independently of that home's -# bearings_state. A home classified captain_decision because it has an open +# bearings_state. Each row's name is the durable task title when nonblank and +# its durable task id otherwise, so renderers always receive a task-identifying +# label instead of having to substitute run status. A home classified +# captain_decision because it has an open # captain hold still contributes each working child as its own Underway row; # the home row on secondmates[] keeps the decision and gate classification. # Captain-hold placement follows the canonical snapshot's hold_bucket and @@ -41,6 +44,10 @@ # Aging is a projection safety net only; the durable # deferral remains re-holding with --until. # +# Charted Next gates are ordered by durable filed date, newest first, before the +# FM_BEARINGS_GATES bound is applied. Gates without a comparable filed date keep +# their input order after dated gates. +# # Main-home inventory validity comes from the canonical snapshot's main_inventory # object (orphan structured in-flight without meta, unstructured current rows). # Bearings never invents Underway rows from backlog-only ids; it discloses those @@ -129,12 +136,14 @@ Default collection performs bounded concurrent remote-ledger reads for registere remote homes under one shared snapshot budget and may refresh the parent-side cache. --include-prs additionally performs live GitHub discovery and checks. -Default fields: schema, home, generated, prs, in_flight{id,kind,state,repo,doing}, +Default fields: schema, home, generated, prs, in_flight{id,kind,state,repo,name,doing}, secondmates{id,state,doing,provenance,freshness,age_seconds,contradiction,reason}, secondmate_reconcile{id,spawn_gen,host,kind,ids}, decisions_open{id,key,verb,summary,owner}, landed{id,what,artifact,owner}, - gates{id,title,blocked_by,reason,owner}, reports{id,path}, recorded_prs{id,url}, + gates{id,title,blocked_by,reason,owner,filed}, reports{id,path}, recorded_prs{id,url}, unhealthy_endpoints{...} (only when non-empty), omitted{surface,reveal}. +Default gates are selected newest filed first before their bound; undated gates + retain input order after dated gates. landed merges this home's Done with registered secondmate homes' Done, bounded by a per-home cap (FM_BEARINGS_LANDED_PER_HOME) and an overall cap (FM_BEARINGS_LANDED), with omitted[] disclosure. Default selection is balanced across deterministic home @@ -381,7 +390,8 @@ MODEL=$(printf '%s' "$SNAP" | jq \ def as_gate($owner): {id, title:(.title | trunc(60)), blocked_by:((.unresolved_blocker_ids // []) | if length > 0 then join(",") else "-" end | trunc(120)), - reason:(hold_gate_reason | trunc(40)), owner:$owner}; + reason:(hold_gate_reason | trunc(40)), owner:$owner, + filed:((.since // null) | trunc(40))}; def round_robin_landed($n): . as $groups | [range(0; (($groups | map(length) | max) // 0)) as $i @@ -457,6 +467,8 @@ MODEL=$(printf '%s' "$SNAP" | jq \ | {id, kind, state: .current_state.state, repo:(.backlog.repo // .project // null), + name:((.backlog.title // "") as $name + | (if ($name | test("[^[:space:]]")) then $name else .id end) | trunc(70)), doing: ((.current_state.detail // "") as $d | (if $d != "" then $d else (.hints.last_event_text // "") end) | trunc(90)) } ] @@ -466,6 +478,9 @@ MODEL=$(printf '%s' "$SNAP" | jq \ kind:(.kind // "secondmate"), state:(.state // "working"), repo:(.repo // null), + name:((.name // "") as $name + | (if (($name | type) == "string" and ($name | test("[^[:space:]]"))) + then $name else ($m.id + "/" + .id) end) | trunc(70)), doing:((.doing // .state) | trunc(90))} ]) as $in_flight_all | ([ .backlog.records[] | . as $record @@ -501,7 +516,8 @@ MODEL=$(printf '%s' "$SNAP" | jq \ title:((.main_inventory.reason // "main inventory invalid") | trunc(60)), blocked_by:"-", reason:"main inventory", - owner:"(main)"}] + owner:"(main)", + filed:null}] else [] end) + [ .backlog.records[] | . as $record @@ -522,7 +538,17 @@ MODEL=$(printf '%s' "$SNAP" | jq \ | select(($all_reports == 1) or (($rel_ids | index($r.id)) != null)) | {id, path} ]) as $reports_all | ([ .tasks[] | select(.kind != "secondmate" and .pr.url != null and .pr.source == "meta") | {id, url:.pr.url} ]) as $recorded_prs_all - | . as $snap + | def filed_epoch: + (.filed // null) as $filed + | if ($filed | type) != "string" then null + elif ($filed | test("T")) then try ($filed | fromdateiso8601) catch null + else try (($filed + "T00:00:00Z") | fromdateiso8601) catch null end; + def newest_filed_first: + to_entries + | sort_by((.value | filed_epoch) as $epoch + | if $epoch == null then [1, 0, .key] else [0, -$epoch, .key] end) + | map(.value); + . as $snap | { schema: "fm-bearings.v1", home: $home, @@ -536,7 +562,8 @@ MODEL=$(printf '%s' "$SNAP" | jq \ decisions_open: (if $all_decisions == 1 then $decisions_all else $decisions_all[:$decisions_n] end), landed: ($done | map({id, what:(.title | trunc(70)), artifact:(landed_artifact // "-"),owner:.home_id})), - gates: (if $all_queued == 1 then $gates_all else $gates_all[:$gates_n] end), + gates: ($gates_all | newest_filed_first + | if $all_queued == 1 then . else .[:$gates_n] end), reports: (if $all_reports == 1 then $reports_all else $reports_all[:$reports_n] end), recorded_prs: (if $all_recorded_prs == 1 then $recorded_prs_all else $recorded_prs_all[:$recorded_prs_n] end) } diff --git a/bin/fm-fleet-snapshot.sh b/bin/fm-fleet-snapshot.sh index 5ae389d047d..4f67b9be00f 100755 --- a/bin/fm-fleet-snapshot.sh +++ b/bin/fm-fleet-snapshot.sh @@ -85,6 +85,10 @@ # reconcile_inventory independently of projection trust. # Actionable captain holds appear in decisions_open; every captain hold remains # in the bounded queued inventory with its structured classification metadata. +# Before that queued bound is applied, non-captain-actionable rows are selected +# ahead of captain-actionable rows so separately projected live decisions cannot +# crowd Charted-Next-eligible work out of the summary. Each group is ordered by +# filed date newest first, with undated rows stable at the end. # Structured-home input must declare the current home-summary and hold-classifier # schemas; a live ledger or cached copy missing either declaration or declaring # an unsupported version is unavailable even when it contains no captain holds. @@ -953,6 +957,16 @@ secondmate_home_summary_json() { # <backlog-json-file> <tasks-json-file> | def trunc($n): tostring | gsub("\\s+"; " ") | if length > $n then .[:$n] + "…" else . end; + def filed_epoch: + (.since // null) as $filed + | if ($filed | type) != "string" then null + elif ($filed | test("T")) then try ($filed | fromdateiso8601) catch null + else try (($filed + "T00:00:00Z") | fromdateiso8601) catch null end; + def newest_filed_first: + to_entries + | sort_by((.value | filed_epoch) as $epoch + | if $epoch == null then [1, 0, .key] else [0, -$epoch, .key] end) + | map(.value); ([ $backlog.records[]? | select((.state == "in_flight" or .state == "queued") and (.structured | not)) ]) as $unstructured_current | ([ $backlog.records[]? | select(.state == "in_flight" and .structured) ]) as $owned_in_flight @@ -1016,6 +1030,7 @@ secondmate_home_summary_json() { # <backlog-json-file> <tasks-json-file> | select(.id == $work.id and .current_state.state == "working") | {id,kind,state:.current_state.state, repo:(($work.repo // .project // null) | if . == null then null else trunc(120) end), + name:(($work.title // null) | if . == null then null else trunc(70) end), source:.current_state.source, doing:((.current_state.detail // "") | trunc(120))} ]) as $active_all | ($captain_holds_all @@ -1082,7 +1097,11 @@ secondmate_home_summary_json() { # <backlog-json-file> <tasks-json-file> hold_age_days:(.hold_age_days // null), captain_actionable:(.captain_actionable // false), repo:((.repo // null) | if . == null then null else trunc(120) end), - kind:((.kind // null) | if . == null then null else trunc(40) end)}][:$queued_n]), + kind:((.kind // null) | if . == null then null else trunc(40) end), + since:((.since // null) | if . == null then null else trunc(40) end)}] + | ((map(select(.captain_actionable != true)) | newest_filed_first) + + (map(select(.captain_actionable == true)) | newest_filed_first)) + | .[:$queued_n]), landed:(if $landed_n == 0 then $landed_all else $landed_all[:$landed_n] end), endpoints:([$tasks[] | {id,state:.current_state.state,source:.current_state.source, endpoint:(.endpoint + {target:((.endpoint.target // null) | if . == null then null else trunc(240) end)})}][:$child_n]), diff --git a/tests/assets/board-render-harness.mjs b/tests/assets/board-render-harness.mjs index e21a8d2dd5d..c181aed5c88 100644 --- a/tests/assets/board-render-harness.mjs +++ b/tests/assets/board-render-harness.mjs @@ -3,7 +3,9 @@ // asserted through the real template rather than by reading its source. // // Usage: node board-render-harness.mjs <built-board.html> -// Prints one JSON document: { stats:[{n,label}], charted:[{title,sub,badges,pickable}] } +// Prints one JSON document: +// { stats:[{n,label}], underway:[{title,sub,badges}], +// charted:[{title,sub,badges,pickable}], empty, more, error } import { readFileSync } from "node:fs"; const html = readFileSync(process.argv[2], "utf8"); @@ -93,18 +95,24 @@ const stats = strip.children.map((t) => ({ label: t.children.find((c) => c.className.includes("bb-stat__label"))?.textContent, })); +const rowsOf = (container) => + container.children + .filter((r) => r.className.split(/\s+/).includes("bb-row")) + .map((row) => { + const main = row.children.find((c) => c.className.includes("bb-row__main")); + return { + title: main?.children.find((c) => c.className.includes("bb-row__title"))?.textContent ?? "", + sub: main?.children.find((c) => c.className.includes("bb-row__sub"))?.textContent ?? "", + badges: badgesOf(row), + pickable: row.children.some((c) => c.className.includes("bb-pick") && !c.className.includes("spacer")), + }; + }); + +const uw = byId.get("bb-underway") || new Node("div"); +const underway = rowsOf(uw); + const ch = byId.get("bb-charted") || new Node("div"); -const charted = ch.children - .filter((r) => r.className.split(/\s+/).includes("bb-row")) - .map((row) => { - const main = row.children.find((c) => c.className.includes("bb-row__main")); - return { - title: main?.children.find((c) => c.className.includes("bb-row__title"))?.textContent ?? "", - sub: main?.children.find((c) => c.className.includes("bb-row__sub"))?.textContent ?? "", - badges: badgesOf(row), - pickable: row.children.some((c) => c.className.includes("bb-pick") && !c.className.includes("spacer")), - }; - }); +const charted = rowsOf(ch); // A fail-closed render replaces the page body instead of the board sections, so // surface it rather than reporting an empty board as a successful render. const errorText = [...byId.entries()] @@ -114,4 +122,5 @@ const errorText = [...byId.entries()] const empty = ch.children.filter((c) => c.className.includes("bb-empty")).map((c) => c.textContent); const more = ch.children.filter((c) => c.className.includes("bb-morechip")).map((c) => c.textContent); -process.stdout.write(JSON.stringify({ stats, charted, empty, more, error: errorText }) + "\n"); +process.stdout.write( + JSON.stringify({ stats, underway, charted, empty, more, error: errorText }) + "\n"); diff --git a/tests/fm-bearings-board-render.test.sh b/tests/fm-bearings-board-render.test.sh index cf26fd31428..d32d0e9dd79 100755 --- a/tests/fm-bearings-board-render.test.sh +++ b/tests/fm-bearings-board-render.test.sh @@ -56,12 +56,14 @@ SH printf '%s\n' "$home" } -# Build the board from <charted-json> and return what the renderer produced. -render() { # <home> <charted-json> [charted_more] [charted_warning_more] - local home=$1 charted=$2 more=${3:-0} warning_more=${4:-0} data="$1/payload.json" - jq -n --argjson charted "$charted" --argjson more "$more" --argjson warning_more "$warning_more" '{ +# Build the board from <underway-json> plus <charted-json> and return what the +# renderer produced. +render_board() { # <home> <underway-json> <charted-json> [charted_more] [charted_warning_more] + local home=$1 underway=$2 charted=$3 more=${4:-0} warning_more=${5:-0} data="$1/payload.json" + jq -n --argjson underway "$underway" --argjson charted "$charted" \ + --argjson more "$more" --argjson warning_more "$warning_more" '{ schema:"fm-bearings-board.v1", home:"render-home", generated:"2026-08-26T00:00Z", - prs_live:false, captains_call:[], underway:[], landed:[], + prs_live:false, captains_call:[], underway:$underway, landed:[], charted:$charted, charted_more:$more, charted_warning_more:$warning_more}' > "$data" PATH="$home/fakebin:$PATH" FM_HOME="$home" \ FM_STATE_OVERRIDE="$home/state" FM_DATA_OVERRIDE="$home/data" \ @@ -71,6 +73,11 @@ render() { # <home> <charted-json> [charted_more] [charted_warning_more] || fail "the built board could not be rendered" } +# Build the board from <charted-json> alone and return what the renderer produced. +render() { # <home> <charted-json> [charted_more] [charted_warning_more] + render_board "$1" '[]' "$2" "${3:-0}" "${4:-0}" +} + charted_next_count() { # <render-json> printf '%s' "$1" | jq -r '.stats[] | select(.label == "charted next") | .n' } @@ -158,6 +165,73 @@ test_an_omitted_kind_keeps_the_existing_queued_rendering() { pass "an omitted kind renders exactly as queued work always did" } +test_an_underway_row_leads_with_the_task_name_and_keeps_its_run_status() { + local home out + home=$(make_home underway-name) + out=$(render_board "$home" '[ + {"id":"fm-board-name-r1","repo":"firstmate","name":"Show task names on the board", + "state":"working","kind":"ship","doing":"no-mistakes: review round 2"} + ]' '[]') + printf '%s' "$out" | jq -e ' + (.underway | length) == 1 + and (.underway[0] + | .title == "Show task names on the board" + and (.sub | test("no-mistakes: review round 2")) + and (.sub | test("ship")) and (.sub | test("firstmate")) + and [.badges[] | .text] == ["working"]) + ' >/dev/null || fail "an underway row did not lead with the task name: $out" + pass "an underway row leads with the task name and still reports its run status" +} + +test_an_underway_identifier_label_is_not_replaced_by_run_status() { + local home out + home=$(make_home underway-identifier) + out=$(render_board "$home" '[ + {"id":"mate/child-1","repo":null,"name":"mate/child-1", + "state":"working","kind":"secondmate","doing":"fixing the failing check"} + ]' '[]') + printf '%s' "$out" | jq -e ' + (.underway | length) == 1 + and (.underway[0] + | .title == "mate/child-1" + and (.sub | startswith("fixing the failing check · ")) + and (.title != "fixing the failing check")) + ' >/dev/null || fail "an identifier-labelled underway row rendered as status-only: $out" + pass "an underway identifier label is not replaced by run status" +} + +test_charted_next_reads_newest_filed_first() { + local home out + home=$(make_home charted-order) + out=$(render_board "$home" '[]' '[ + {"id":"oldest","repo":"sample","title":"Filed in June","reason":"queued","dispatchable":true,"filed":"2026-06-01"}, + {"id":"newest","repo":"sample","title":"Filed in August","reason":"queued","dispatchable":true,"filed":"2026-08-14T09:30:00Z"}, + {"id":"middle","repo":"sample","title":"Filed in July","reason":"queued","dispatchable":true,"filed":"2026-07-22"} + ]') + printf '%s' "$out" | jq -e ' + [.charted[] | .title] == ["Filed in August", "Filed in July", "Filed in June"] + ' >/dev/null || fail "charted next was not ordered newest filed first: $out" + pass "charted next renders the most recently filed work first" +} + +test_charted_rows_without_a_filed_date_follow_the_dated_rows_in_payload_order() { + local home out + home=$(make_home charted-undated) + out=$(render_board "$home" '[]' '[ + {"id":"undated-first","repo":"sample","title":"Undated one","reason":"queued","dispatchable":true}, + {"id":"dated","repo":"sample","title":"Dated","reason":"queued","dispatchable":true,"filed":"2026-07-22"}, + {"id":"undated-second","repo":"sample","title":"Undated two","reason":"queued","dispatchable":true,"filed":null} + ]') + printf '%s' "$out" | jq -e ' + [.charted[] | .title] == ["Dated", "Undated one", "Undated two"] + ' >/dev/null || fail "undated charted rows did not keep a stable trailing order: $out" + pass "charted rows with no filed date follow the dated rows in payload order" +} + +test_an_underway_row_leads_with_the_task_name_and_keeps_its_run_status +test_an_underway_identifier_label_is_not_replaced_by_run_status +test_charted_next_reads_newest_filed_first +test_charted_rows_without_a_filed_date_follow_the_dated_rows_in_payload_order test_a_warning_row_reads_as_a_repair_not_as_queued_work test_warnings_are_excluded_from_the_charted_next_count test_a_board_of_only_warnings_still_reports_nothing_queued diff --git a/tests/fm-bearings-board.test.sh b/tests/fm-bearings-board.test.sh index b59010036e9..b5254d42bfa 100644 --- a/tests/fm-bearings-board.test.sh +++ b/tests/fm-bearings-board.test.sh @@ -260,6 +260,20 @@ test_build_refuses_malformed_payloads_before_touching_the_board() { set +e; out=$(run_board "$home" build "$data" 2>&1); rc=$?; set -e [ "$rc" -ne 0 ] || fail "a fleet row without an explicit repo marker was accepted" + write_valid_payload "$data" + jq '.underway = [{"id":"sample-task","repo":"sample","state":"working", + "kind":"ship","doing":"implementing"}]' "$data" > "$data.tmp" && mv "$data.tmp" "$data" + set +e; out=$(run_board "$home" build "$data" 2>&1); rc=$?; set -e + [ "$rc" -ne 0 ] || fail "an underway row without an explicit name marker was accepted" + + for invalid_filed in "last Tuesday" "2026-13-01" "2026-08-14T99:30:00Z" "2026-02-29"; do + write_valid_payload "$data" + jq --arg filed "$invalid_filed" '.charted[0].filed = $filed' "$data" > "$data.tmp" \ + && mv "$data.tmp" "$data" + set +e; out=$(run_board "$home" build "$data" 2>&1); rc=$?; set -e + [ "$rc" -ne 0 ] || fail "an invalid filed date was accepted: $invalid_filed" + done + write_valid_payload "$data" jq '.captains_call[0].allow_freeform = "yes"' "$data" > "$data.tmp" && mv "$data.tmp" "$data" set +e; out=$(run_board "$home" build "$data" 2>&1); rc=$?; set -e diff --git a/tests/fm-bearings-snapshot.test.sh b/tests/fm-bearings-snapshot.test.sh index bc385251e8b..ecdde88a82c 100755 --- a/tests/fm-bearings-snapshot.test.sh +++ b/tests/fm-bearings-snapshot.test.sh @@ -2402,6 +2402,139 @@ EOF pass "active children reach Underway independently of a home captain hold" } +test_nameless_legacy_summary_uses_its_durable_identifier() { + local parent remote_home fakebin json + parent=$(make_home nameless-legacy-summary) + make_remote_ledger_fleet "$parent" 1 + remote_home="$TMP_ROOT/remote-ledger-home-1" + fakebin=$(make_remote_ledger_ssh "$parent/remote-ssh") + jq ' + .active_children = [ + {id:"legacy-child",kind:"ship",state:"working",repo:null, + source:"remote-ledger",doing:"running review"}, + {id:"blank-name-child",kind:"ship",state:"working",repo:null,name:" \t ", + source:"remote-ledger",doing:"running tests"} + ] + | .counts.active_children = 2 + | .state = "active_child_work" + ' "$remote_home/state/home-summary.json" > "$remote_home/state/legacy-summary.json" + mv "$remote_home/state/legacy-summary.json" "$remote_home/state/home-summary.json" + + json=$(run_remote_ledger_bearings "$parent" "$fakebin" 1100) \ + || fail "nameless legacy summary bearings failed" + printf '%s' "$json" | jq -e ' + (.in_flight | any(.id == "ledger-1/legacy-child" + and .name == "ledger-1/legacy-child" + and .doing == "running review" + and .name != .doing)) + and (.in_flight | any(.id == "ledger-1/blank-name-child" + and .name == "ledger-1/blank-name-child" + and .doing == "running tests" + and .name != .doing)) + ' >/dev/null || fail "a blank legacy child name was not replaced by its id: $json" + pass "blank legacy summary names use their durable identifier" +} + +test_newest_filed_gates_are_selected_before_snapshot_bounds() { + local home mate fakebin json i + home=$(make_home newest-before-bounds) + : > "$home/data/secondmates.md" + printf '## In flight\n\n## Queued\n' > "$home/data/backlog.md" + i=1 + while [ "$i" -le 20 ]; do + printf -- '- [ ] old-%02d - Older gate %02d (repo: sample) (kind: ship) (since 2026-06-%02d)\n' \ + "$i" "$i" "$i" >> "$home/data/backlog.md" + i=$((i + 1)) + done + printf -- '- [ ] newest - Newest gate (repo: sample) (kind: ship) (since 2026-07-01)\n\n## Done\n' \ + >> "$home/data/backlog.md" + fakebin=$(make_fakebin "$home") + json=$(run "$home" "$fakebin" --json) + printf '%s' "$json" | jq -e ' + (.gates | length) == 20 and .gates[0].id == "newest" + and (.gates | any(.id == "old-01") | not) + ' >/dev/null || fail "the bearings gate bound dropped the newest filed row: $json" + + mate="$TMP_ROOT/newest-before-bounds-mate" + make_valid_secondmate_home bounded-mate "$mate" + : > "$home/data/backlog.md" + append_secondmate_registry "$home" bounded-mate "$mate" + cat > "$mate/data/backlog.md" <<'EOF' +## In flight + +## Queued +- [ ] mate-eligible - Eligible remote gate (repo: sample) (kind: ship) (since 2026-07-08) +- [ ] mate-call-one - Newer captain call (repo: sample) (kind: captain) (hold: choose one) (hold-kind: captain) (since 2026-07-10) +- [ ] mate-call-two - Newest captain call (repo: sample) (kind: captain) (hold: choose two) (hold-kind: captain) (since 2026-07-11) + +## Done +EOF + json=$(FM_SNAPSHOT_SECONDMATE_QUEUED=2 run "$home" "$fakebin" --json) + printf '%s' "$json" | jq -e ' + [.gates[].id] == ["mate-eligible"] + and (.decisions_open | any(.id == "bounded-mate/mate-call-one")) + and (.decisions_open | any(.id == "bounded-mate/mate-call-two")) + ' >/dev/null || fail "captain calls crowded eligible Charted work out of the bound: $json" + pass "newest filed gates are selected before snapshot bounds" +} + +# A captain scanning Underway must be able to tell WHICH task a row is, and the +# board orders Charted Next by the durable filed date, so both facts have to come +# out of fleet state rather than being invented at render time. +test_underway_and_gate_rows_carry_the_durable_name_and_filed_date() { + local home mate fakebin json + home=$(make_home durable-name-filed) + : > "$home/data/secondmates.md" + mate="$TMP_ROOT/durable-name-home" + make_valid_secondmate_home named-mate "$mate" + append_secondmate_registry "$home" named-mate "$mate" + mkdir -p "$home/projects/main-wt" + cat > "$home/data/backlog.md" <<'EOF' +## In flight +- [ ] main-ship - Rename the fleet board rows (repo: firstmate) (kind: ship) (since 2026-07-09) + +## Queued +- [ ] newer-gate - Filed later (repo: firstmate) (kind: ship) (since 2026-07-10) +- [ ] older-gate - Filed earlier (repo: firstmate) (kind: ship) (since 2026-07-01) +- [ ] undated-gate - Filed before dates were recorded (repo: firstmate) (kind: ship) + +## Done +EOF + fm_write_meta "$home/state/main-ship.meta" \ + "window=firstmate:fm-main-ship" "worktree=$home/projects/main-wt" "project=firstmate" \ + "harness=claude" "kind=ship" "mode=no-mistakes" + record_claude_state "$home/state" main-ship busy + printf 'working: no-mistakes review round 2\n' > "$home/state/main-ship.status" + + printf '## In flight\n' > "$mate/data/backlog.md" + printf -- '- [ ] mate-child - Tighten the ledger contract (repo: sample) (kind: ship) (since 2026-07-08)\n' \ + >> "$mate/data/backlog.md" + printf '\n## Queued\n\n## Done\n' >> "$mate/data/backlog.md" + mkdir -p "$mate/projects/mate-child" + fm_write_meta "$mate/state/mate-child.meta" \ + "window=firstmate:fm-mate-child" "worktree=$mate/projects/mate-child" "project=sample" \ + "harness=claude" "kind=ship" "mode=no-mistakes" + record_claude_state "$mate/state" mate-child busy + printf 'working: waiting on the pipeline\n' > "$mate/state/mate-child.status" + + fakebin=$(make_fakebin "$home") + json=$(run "$home" "$fakebin" --json) + printf '%s' "$json" | jq -e ' + (.in_flight | any(.id == "main-ship" + and .name == "Rename the fleet board rows" + and (.doing | type == "string") and (.doing | length) > 0 + and .doing != .name)) + and (.in_flight | any(.id == "named-mate/mate-child" + and .name == "Tighten the ledger contract" + and (.doing | type == "string") and (.doing | length) > 0 + and .doing != .name)) + and (.gates | any(.id == "newer-gate" and .filed == "2026-07-10")) + and (.gates | any(.id == "older-gate" and .filed == "2026-07-01")) + and (.gates | any(.id == "undated-gate" and .filed == null)) + ' >/dev/null || fail "durable Underway names or gate filed dates are missing: $json" + pass "Underway rows carry the durable task name and gates carry their filed date" +} + test_mixed_secondmate_roles_partial_state_and_captain_readiness() { local home fakebin hibit wheel sshhip ha canonical json home=$(make_home mixed-domain-regressions) @@ -3214,6 +3347,9 @@ test_main_unstructured_current_is_disclosed_with_structured_sibling test_main_orphan_counterfactual_meta_clears_inventory_warning test_working_captain_holds_keep_their_bucket_surfaces test_active_children_project_independent_of_home_captain_hold +test_nameless_legacy_summary_uses_its_durable_identifier +test_newest_filed_gates_are_selected_before_snapshot_bounds +test_underway_and_gate_rows_carry_the_durable_name_and_filed_date test_mixed_secondmate_roles_partial_state_and_captain_readiness test_main_captain_readiness_matches_secondmate_projection test_completed_scout_report_not_pending From ad14a8db06fd7a6655c2c4cc351d7add92448ecf Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Fri, 11 Sep 2026 17:34:37 -0700 Subject: [PATCH 14/31] fix(bearings): surface return catch-up without blocking snapshots (#4248) * fix(bearings): report the away-return catch-up instead of refusing A captain returning from away and asking for bearings got zero bytes and an error: fm-bearings-snapshot.sh ran the away-return guard with `|| exit $?` before reading any fleet state, so the mere existence of the catch-up gate killed every bearings mode (and /ahoy with them). Bearings now consults that guard rather than obeying it. fm-afk-return.sh separates its two refusal branches by exit status, so an ACTIVE away window still refuses exactly as before - the right answer there is to run the return first - while return catch-up (exit 4) lets collection and projection proceed and is disclosed as one action-free `(return-catchup)` gate row, following the existing `(main-inventory)` precedent. It stays out of decisions_open: these blockers are firstmate-actionable, not the captain's own call, and the per-task blockers already project as their own Underway rows. The guard's refusal text also stops promising a blocker list it cannot produce: a gate retained for a lifecycle reason alone now names that retention reason, and bearings carries the same reason in the gate row's title. Reporting is not ordinary work. AGENTS.md already scopes the return hold to work rather than reporting, so only the /afk and bearings skills needed the correction. * no-mistakes(document): Refresh away-return Bearings verification * no-mistakes(review): Reserve catch-up gate outside Bearings truncation * no-mistakes(review): Preserve filed dates in catch-up gate output * no-mistakes(document): Document reserved catch-up gate projection --- .agents/skills/afk/SKILL.md | 3 +- .agents/skills/bearings/SKILL.md | 6 +- bin/fm-afk-return.sh | 65 ++++++++++++++++--- bin/fm-bearings-snapshot.sh | 59 ++++++++++++++--- docs/verification/runtime-backends.md | 11 ++-- tests/fm-afk-pi-herdr-return-e2e.test.sh | 21 ++++--- tests/fm-afk-return.test.sh | 80 ++++++++++++++++++++---- 7 files changed, 200 insertions(+), 45 deletions(-) diff --git a/.agents/skills/afk/SKILL.md b/.agents/skills/afk/SKILL.md index 573eed53c07..5b2479e2997 100644 --- a/.agents/skills/afk/SKILL.md +++ b/.agents/skills/afk/SKILL.md @@ -71,7 +71,8 @@ No `/back` is needed. The first genuine message is the return signal: The gate keeps every open `blocked:` event until that blocker's own resolution is proven: remediate each immediately through the normal lifecycle, or explicitly reclassify it with a durable reason and close its decision key with `resolved [key=...]`, then run `bin/fm-afk-return.sh check`. Captain-verdict outcomes are listed under "waiting on you", but do not exempt open blockers because per-blocker provenance is deferred to phase 4. Once the record is archived, resume full per-wake responsiveness through the emitted primary-harness supervision protocol while blocker handling proceeds, so the gate never creates a blind wait. - Do not answer a Bearings request or perform any other ordinary captain work until the check exits successfully. + A Bearings request may be answered while the gate is open, and the digest surfaces the catch-up state as a Charted Next `(return-catchup)` warning row naming what still holds it. + Acting on the fleet - dispatching, steering, merging, or any other ordinary captain work - still waits until the check exits successfully. - A message **with** the current operational prefix (`FM_OPERATIONAL_PREFIX`, U+2063 INVISIBLE SEPARATOR followed by `FIRSTMATE_OP: `), or a legacy bare `FM_INJECT_MARK` daemon escalation -> stay away and process it. - Re-invoking `/afk` while already away -> stay away (refresh); this does **not** trigger an exit. diff --git a/.agents/skills/bearings/SKILL.md b/.agents/skills/bearings/SKILL.md index 03d22ee708d..7de9c9c4a67 100644 --- a/.agents/skills/bearings/SKILL.md +++ b/.agents/skills/bearings/SKILL.md @@ -54,6 +54,8 @@ Board answers are acted on later under the normal authority rules; this skill's The `(main-inventory)` gate is an action-free integrity warning rather than queued work. Render it under Charted Next with the related `omitted` disclosure, never invent an Underway row from backlog-only state, and never move it into Captain's Call. The same holds for a secondmate home whose current state is unavailable, and for a readable home whose `invalidity` reports a backlog-vs-metadata mismatch: the mismatch is a repair notice about that home's own books, not a reason to drop its separately projected decisions, queued, landed, or live work. + The `(return-catchup)` gate is the same shape: an action-free notice that an away-return catch-up is still open, naming the blockers left to clear or the reason the catch-up was retained. + Render it under Charted Next like any other warning row: reporting is not ordinary work, while acting on the fleet still waits for `bin/fm-afk-return.sh check` (`/afk`). 2. **Record a later reconcile notification for any home whose own books disagree.** When the snapshot reports a secondmate home whose `invalidity` is `orphan_in_flight`, `unowned_current`, or `terminal_in_flight`, that home's backlog and its own task metadata disagree and only that home may fix it. @@ -99,13 +101,13 @@ Compose the payload from the same snapshot with the same ranking judgment as the - Decision cards carry agent-authored copy: a short noun-phrase title, one-line `about` and `decide` context rows, and option labels with hints, with the recommended option marked. - Card `type` (decision, merge, credential) is your composing judgment from the row's content; no backlog field types a card for you. - When the card's task is a captain-gated WORK item (the answer should free it to proceed rather than complete it), set the card's `close: "release"` so the answer lifts the hold instead of closing the task; question-shaped items omit it. -- A Charted Next row's optional `kind` separates work from alarms: omit it (or set `"queued"`) for real queued work, and set `"warning"` on every action-free fleet-integrity notice - the `(main-inventory)` gate, an unavailable secondmate home, and an inventory-mismatch repair notice. The board badges a warning row `needs repair` instead of `waiting` and leaves it out of the Charted Next count, so those rows never read as dispatchable queued work. +- A Charted Next row's optional `kind` separates work from alarms: omit it (or set `"queued"`) for real queued work, and set `"warning"` on every action-free fleet-integrity notice - the `(main-inventory)` gate, the `(return-catchup)` gate, an unavailable secondmate home, and an inventory-mismatch repair notice. The board badges a warning row `needs repair` instead of `waiting` and leaves it out of the Charted Next count, so those rows never read as dispatchable queued work. - `charted_more` counts omitted queued rows only, while `charted_warning_more` counts omitted warning rows only; keep both counts separate whenever the board payload truncates Charted Next. - Every Underway row copies the task-identifying `in_flight.name` from the snapshot into an explicit `name` field, which the board leads with while keeping the run status on its second line. The snapshot command's header owns its durable-title-or-id normalization; never replace the projected label with run status or invent another label. - Every Charted Next row copies the snapshot gate's durable filed date into `filed`, and the board orders the section by it, newest filed first. Follow `bin/fm-bearings-board.sh`'s payload contract for the accepted format. - Omit it or pass null for a row with no durable filed date - the main-inventory warning, an unavailable secondmate home, or a queued row filed before dates were recorded - and the board keeps those rows in payload order after every dated row. + Omit it or pass null for a row with no durable filed date - the main-inventory or return-catchup warning, an unavailable secondmate home, or a queued row filed before dates were recorded - and the board keeps those rows in payload order after every dated row. - Every Captain's Call item and every Underway, Recently Landed, and Charted Next row carries an explicit `repo` field. Fill it from the snapshot and task records wherever known; use null or an empty string only as the deliberate genuinely-no-repo marker, in which case the template may show the internal id. Ids otherwise stay in the payload only as the routing channel, and composed reasons name blockers in plain words. Run `build` once after composing the payload. diff --git a/bin/fm-afk-return.sh b/bin/fm-afk-return.sh index 923f1259eee..7773c7d325c 100755 --- a/bin/fm-afk-return.sh +++ b/bin/fm-afk-return.sh @@ -6,7 +6,9 @@ # fm-afk-return.sh Stop away mode, render the return brief, and open/check the gate. # fm-afk-return.sh begin Same as the default command. # fm-afk-return.sh check Re-render the brief and close the gate only after blockers resolve. -# fm-afk-return.sh guard Read-only refusal while away or catch-up is pending. +# fm-afk-return.sh guard Read-only consult: exit 3 while away mode is still +# active, exit 4 while return catch-up is pending. +# fm-afk-return.sh catchup-summary Read-only catch-up projection for a reporting surface. # # THE RETURN BRIEF (stdout, on begin and on every check) is rendered from durable # records, never from conversation memory: the archived away-posture record @@ -37,9 +39,12 @@ # so a crash between stopping, wake presentation, and blocker handling fails # closed. It retains the presented wake, buffered-escalation, wedge-marker, # health, and posture-record evidence until every live open blocker is closed -# and `check` succeeds. Repeated begin/check calls are idempotent. `guard` -# never mutates state and is suitable for ordinary read entrypoints such as -# fm-bearings-snapshot.sh. +# and `check` succeeds. Repeated begin/check calls are idempotent. `guard` and +# `catchup-summary` never mutate state and are suitable for ordinary read +# entrypoints such as fm-bearings-snapshot.sh. `guard` separates its two +# refusal branches by exit status so a reporting surface can keep refusing +# during an active away window while still rendering the catch-up posture as +# content; this file owns the gate format both branches read. set -u SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -59,7 +64,7 @@ RETURN_GRACE=${FM_GUARD_GRACE:-300} CONTRACT="$SCRIPT_DIR/fm-afk-contract.sh" usage() { - sed -n '2,9p' "${BASH_SOURCE[0]}" | sed 's/^# \{0,1\}//' + sed -n '2,11p' "${BASH_SOURCE[0]}" | sed 's/^# \{0,1\}//' } clean_field() { @@ -245,15 +250,58 @@ clear_delivery_artifacts() { "$STATE/.subsuper-inject-wedged" } +# The lifecycle retention reasons the gate kept, one per line, empty when the +# gate was retained for open blockers alone. +gate_retention_reasons() { # <file> + local file=$1 tag kind text + while IFS="$(printf '\t')" read -r tag kind text; do + [ "$tag" = evidence ] && [ "$kind" = lifecycle ] || continue + printf '%s\n' "$text" + done < "$file" +} + +gate_has_blockers() { # <file> + grep -q "^blocker$(printf '\t')" "$1" 2>/dev/null +} + +# Read-only catch-up projection for a reporting surface such as +# fm-bearings-snapshot.sh: one tab-separated line +# `<open-blocker-count><TAB><first-retention-reason>`, and exit 1 when no gate +# is open. The reason field is empty when open blockers alone hold the gate. +catchup_summary() { + local count reason + [ -e "$GATE" ] || return 1 + count=$(grep -c "^blocker$(printf '\t')" "$GATE" 2>/dev/null || true) + case "$count" in ''|*[!0-9]*) count=0 ;; esac + reason=$(gate_retention_reasons "$GATE" | head -1) + printf '%s\t%s\n' "$count" "$reason" +} + return_guard() { + local reasons if [ -e "$STATE/.afk" ] || fm_afk_contract_present "$STATE"; then printf 'fm-afk-return: away mode is still active; run bin/fm-afk-return.sh before ordinary captain work\n' >&2 return 3 fi if [ -e "$GATE" ]; then - printf 'fm-afk-return: return catch-up is pending; remediate or durably reclassify every listed blocker, then run bin/fm-afk-return.sh check\n' >&2 - print_blockers "$GATE" >&2 - return 3 + if gate_has_blockers "$GATE"; then + printf 'fm-afk-return: return catch-up is pending; remediate or durably reclassify every listed blocker, then run bin/fm-afk-return.sh check\n' >&2 + print_blockers "$GATE" >&2 + else + # No blocker row exists, so naming "every listed blocker" would ask for + # something the gate does not list. Name the lifecycle retention reason + # that actually holds it instead. + printf 'fm-afk-return: return catch-up is pending with no open blocker; clear the retention reason below, then run bin/fm-afk-return.sh check\n' >&2 + reasons=$(gate_retention_reasons "$GATE") + if [ -n "$reasons" ]; then + printf '%s\n' "$reasons" | while IFS= read -r text; do + printf 'catch-up retained: %s\n' "$text" >&2 + done + else + printf 'catch-up retained: the durable gate recorded no retention reason\n' >&2 + fi + fi + return 4 fi return 0 } @@ -650,6 +698,7 @@ main() { case "$mode" in begin|check) ;; guard) return_guard; return ;; + catchup-summary) catchup_summary; return ;; -h|--help|help) usage; return 0 ;; *) usage >&2; return 2 ;; esac diff --git a/bin/fm-bearings-snapshot.sh b/bin/fm-bearings-snapshot.sh index 154019977c1..8f7bda840db 100755 --- a/bin/fm-bearings-snapshot.sh +++ b/bin/fm-bearings-snapshot.sh @@ -44,9 +44,10 @@ # Aging is a projection safety net only; the durable # deferral remains re-holding with --until. # -# Charted Next gates are ordered by durable filed date, newest first, before the -# FM_BEARINGS_GATES bound is applied. Gates without a comparable filed date keep -# their input order after dated gates. +# Ordinary Charted Next gates are ordered by durable filed date, newest first, +# before the FM_BEARINGS_GATES bound is applied. Gates without a comparable filed +# date keep their input order after dated gates. The synthetic (return-catchup) +# posture row is reserved ahead of that ordering and bound so it always surfaces. # # Main-home inventory validity comes from the canonical snapshot's main_inventory # object (orphan structured in-flight without meta, unstructured current rows). @@ -54,6 +55,12 @@ # gaps in omitted[] and, when invalid, a Charted Next gate line so the four-section # chat cannot claim an empty fleet while main current state is broken. # +# An open away-return catch-up is disclosed the same way, as a single action-free +# (return-catchup) gate row naming the blockers left to clear or the reason the +# catch-up was retained. Reporting is not ordinary captain work, so the gate never +# suppresses the digest; an ACTIVE away window still refuses, because the right +# answer there is to run the return first. bin/fm-afk-return.sh owns the gate. +# # The landed section merges this home's Done with the canonical snapshot's # secondmate_landed roll-up (fm-fleet-snapshot.sh), so merges a secondmate managed - # recorded in ITS OWN backlog, never the main one - are visible. It stays bounded by @@ -197,10 +204,29 @@ done command -v jq >/dev/null 2>&1 || { echo "fm-bearings-snapshot: jq not found" >&2; exit 1; } -# The deterministic return-catch-up owner must clear before this or any other -# ordinary captain request proceeds. Bearings does not reproduce that policy; -# it only consults the shared read-only gate. -"$SCRIPT_DIR/fm-afk-return.sh" guard || exit $? +# The shared read-only away-return owner is consulted, not obeyed. An active +# away window still refuses here: the correct answer to a bearings request then +# is to run the return first. Return CATCH-UP is different - the captain is +# back and asking for the picture, so the catch-up posture is reported as +# content (a Charted Next gate row) and collection continues. bin/fm-afk-return.sh +# owns both the gate format and the branch distinction; bearings reproduces +# neither. Acting on the fleet still waits for its `check`. +RETURN_CATCHUP=null +GUARD_RC=0 +GUARD_ERR=$("$SCRIPT_DIR/fm-afk-return.sh" guard 2>&1 >/dev/null) || GUARD_RC=$? +if [ "$GUARD_RC" -ne 0 ] && [ "$GUARD_RC" -ne 4 ]; then + [ -z "$GUARD_ERR" ] || printf '%s\n' "$GUARD_ERR" >&2 + exit "$GUARD_RC" +fi +if [ "$GUARD_RC" -eq 4 ]; then + CATCHUP_LINE=$("$SCRIPT_DIR/fm-afk-return.sh" catchup-summary) || CATCHUP_LINE="" + CATCHUP_BLOCKERS=${CATCHUP_LINE%%$'\t'*} + case "$CATCHUP_BLOCKERS" in ''|*[!0-9]*) CATCHUP_BLOCKERS=0 ;; esac + CATCHUP_REASON="" + case "$CATCHUP_LINE" in *"$(printf '\t')"*) CATCHUP_REASON=${CATCHUP_LINE#*$'\t'} ;; esac + RETURN_CATCHUP=$(jq -n --argjson blockers "$CATCHUP_BLOCKERS" --arg reason "$CATCHUP_REASON" \ + '{pending:true,blockers:$blockers,reason:$reason}') +fi NOW=${FM_BEARINGS_NOW:-$(date -u +%Y-%m-%dT%H:%M:%SZ)} if [ "$ALL_LANDED" = 1 ] || [ "$ALL_SECONDMATES" = 1 ]; then @@ -340,6 +366,7 @@ MODEL=$(printf '%s' "$SNAP" | jq \ --argjson pr_repos_shown "$PR_REPOS_SHOWN" \ --argjson pr_rows_capped "$PR_ROWS_CAPPED" \ --argjson pr_rows_min_total "$PR_ROWS_MIN_TOTAL" \ + --argjson return_catchup "$RETURN_CATCHUP" \ --argjson candidate_prs "$CANDIDATE_PRS" "$FM_LANDED_JQ_DEFS"' def trunc($n): if . == null then null else (tostring | gsub("\\s+"; " ") | if (length > $n) then (.[:$n] + "…") else . end) end; @@ -511,6 +538,19 @@ MODEL=$(printf '%s' "$SNAP" | jq \ + [ (.secondmate_current.records // [])[] | .queued[]? | select(.hold_kind == "captain" and projected_deferred_hold) ] | length) as $decisions_marked_deferred + | (if ($return_catchup.pending // false) then + [{id:"(return-catchup)", + title:((if ($return_catchup.blockers // 0) > 0 then + "\($return_catchup.blockers) blocker(s) to clear before ordinary work" + elif (($return_catchup.reason // "") != "") then + ("catch-up retained: " + + ($return_catchup.reason | sub("[,;] *catch-up stays gated$"; ""))) + else "away-return catch-up is still open" end) | trunc(60)), + blocked_by:"-", + reason:"away-return catch-up", + owner:"(main)", + filed:null}] + else [] end) as $return_catchup_gate | ((if (.main_inventory.valid == false) then [{id:"(main-inventory)", title:((.main_inventory.reason // "main inventory invalid") | trunc(60)), @@ -562,8 +602,9 @@ MODEL=$(printf '%s' "$SNAP" | jq \ decisions_open: (if $all_decisions == 1 then $decisions_all else $decisions_all[:$decisions_n] end), landed: ($done | map({id, what:(.title | trunc(70)), artifact:(landed_artifact // "-"),owner:.home_id})), - gates: ($gates_all | newest_filed_first - | if $all_queued == 1 then . else .[:$gates_n] end), + gates: ($return_catchup_gate + + ($gates_all | newest_filed_first + | if $all_queued == 1 then . else .[:$gates_n] end)), reports: (if $all_reports == 1 then $reports_all else $reports_all[:$reports_n] end), recorded_prs: (if $all_recorded_prs == 1 then $recorded_prs_all else $recorded_prs_all[:$recorded_prs_n] end) } diff --git a/docs/verification/runtime-backends.md b/docs/verification/runtime-backends.md index 9d4d8d2cd21..1ec8b925be0 100644 --- a/docs/verification/runtime-backends.md +++ b/docs/verification/runtime-backends.md @@ -1197,22 +1197,23 @@ A stale-registration pane is never a husk: create, reclaim, presentation recover ### Away-mode transport The away daemon is no longer launched on Pi; the away posture there is the record `bin/fm-afk-contract.sh` owns. -The Pi/Herdr away posture and return path was verified on 2026-09-08 against a real Pi primary in an isolated Herdr lab session, Herdr 0.9.0 and Pi 0.82.0: +The Pi/Herdr away posture and return transport was verified on 2026-09-08 against a real Pi primary in an isolated Herdr lab session, Herdr 0.9.0 and Pi 0.82.0: ```sh FM_AFK_PI_HERDR_E2E=1 HERDR_LAB_HELPER=bin/fm-herdr-lab.sh \ tests/fm-afk-pi-herdr-return-e2e.test.sh ``` -``` +Relevant transport output: + +```text ok - real Pi primary: the away posture is recorded with no daemon launched ok - real Pi/Herdr: nothing injects into the captain pane under the away posture -ok - real unmarked Pi return renders the brief, opens catch-up, and blocks Bearings before the unresolved blocker can be deferred -ok - resolved return catch-up allows Bearings and a clean idempotent away re-entry evidence: herdr=herdr 0.9.0 pi=0.82.0 target=fm-lab-fm-afk-pi-return-37189-7133:w1:p1 archived-records=2 ``` -Observed guarantees: `fm-afk-launch.sh start` refused on the Pi primary and `confirm` recorded the posture with no daemon pid, flag, or terminal; a pending real Pi draft was left untouched with nothing submitted into the captain pane; the unmarked return request was recognized as the return, rendered the brief health first, opened the catch-up gate on the live blocker, and refused Bearings; resolving the blocker cleared the gate, and a clean re-entry and return left exactly one archived record per away window. +Observed guarantees: `fm-afk-launch.sh start` refused on the Pi primary and `confirm` recorded the posture with no daemon pid, flag, or terminal; a pending real Pi draft was left untouched with nothing submitted into the captain pane; the unmarked return request was recognized as the return, rendered the brief health first, and opened the catch-up gate on the live blocker; resolving the blocker cleared the gate, and a clean re-entry and return left exactly one archived record per away window. +The current catch-up reporting boundary is pinned by `tests/fm-afk-return.test.sh` and the same live entry point: Bearings continues through a pending return catch-up, projects its posture as an action-free warning outside Captain's Call, and drops that warning after the gate clears, while an active away window still refuses. The fixture captures submitted input through Pi's `input` extension hook, so the lab agent directory needs no provider credentials. The daemon injection transport into a live composer keeps its coverage in `tests/fm-afk-inject-herdr-e2e.test.sh` for the harnesses that still run the daemon, and the dedicated Herdr daemon workspace topology is covered by `tests/fm-afk-launch.test.sh` and preserves the captain tab's pane count. diff --git a/tests/fm-afk-pi-herdr-return-e2e.test.sh b/tests/fm-afk-pi-herdr-return-e2e.test.sh index 0daea78b06e..0b94a29d9b1 100755 --- a/tests/fm-afk-pi-herdr-return-e2e.test.sh +++ b/tests/fm-afk-pi-herdr-return-e2e.test.sh @@ -10,7 +10,8 @@ # records the posture with no daemon terminal, pid, or flag; # - a real Pi draft is never touched by away mode (nothing injects on Pi); # - an unmarked return request is recognized as the return, opens the -# catch-up gate before Bearings on the live blocker, and renders the brief; +# catch-up gate on the live blocker, renders the brief, and still lets +# Bearings report that catch-up posture as content; # - remediation/resolution clears the gate, and re-entry is idempotent. # The 2026-07-14 two-owner incident's daemon-injection assertions retired with # the daemon on Pi; the daemon transport keeps its coverage in @@ -252,19 +253,23 @@ assert_contains "$RETURN_OUT" 'firstmate-actionable blocker: repair-task [key=sy assert_contains "$RETURN_OUT" '=== Return brief (away ' "the return did not render the brief" assert_contains "$RETURN_OUT" 'Supervisor health:' "the brief did not lead with supervisor health" [ ! -f "$STATE/.afk-contract" ] || fail "the return did not archive the away-posture record" -set +e BEARINGS_OUT=$(PATH="$FAKEBIN:$ORIGINAL_PATH" HERDR_SESSION="$SESSION" FM_ROOT_OVERRIDE="$PROJECT" FM_HOME="$HOME_DIR" FM_STATE_OVERRIDE="$STATE" \ - "$ROOT/bin/fm-bearings-snapshot.sh" --json 2>&1) -BEARINGS_RC=$? -set -e -[ "$BEARINGS_RC" -eq 3 ] || fail "Bearings bypassed the return gate (rc=$BEARINGS_RC): $BEARINGS_OUT" -pass "real unmarked Pi return renders the brief, opens catch-up, and blocks Bearings before the unresolved blocker can be deferred" + "$ROOT/bin/fm-bearings-snapshot.sh" --json 2>&1) \ + || fail "Bearings refused behind the return gate instead of reporting it: $BEARINGS_OUT" +printf '%s' "$BEARINGS_OUT" | jq -e ' + (.in_flight | any(.id == "repair-task")) + and (.gates | any(.id == "(return-catchup)" and .reason == "away-return catch-up")) + and ([.decisions_open[].id] | index("(return-catchup)") | not)' >/dev/null \ + || fail "Bearings did not surface the catch-up posture as content: $BEARINGS_OUT" +pass "real unmarked Pi return renders the brief, opens catch-up, and reports that posture through Bearings while the blocker stays Firstmate's to remediate" printf 'resolved [key=synthetic-dependency]: refreshed the synthetic token and resumed the task\n' >> "$STATE/repair-task.status" PATH="$FAKEBIN:$ORIGINAL_PATH" HERDR_SESSION="$SESSION" FM_ROOT_OVERRIDE="$PROJECT" FM_HOME="$HOME_DIR" FM_STATE_OVERRIDE="$STATE" \ "$ROOT/bin/fm-afk-return.sh" check >/dev/null || fail "remediated blocker did not clear return catch-up" PATH="$FAKEBIN:$ORIGINAL_PATH" HERDR_SESSION="$SESSION" FM_ROOT_OVERRIDE="$PROJECT" FM_HOME="$HOME_DIR" FM_STATE_OVERRIDE="$STATE" \ - "$ROOT/bin/fm-bearings-snapshot.sh" --json >/dev/null || fail "Bearings remained gated after blocker remediation" + "$ROOT/bin/fm-bearings-snapshot.sh" --json \ + | jq -e '[.gates[].id] | index("(return-catchup)") | not' >/dev/null \ + || fail "Bearings kept the catch-up posture row after the gate cleared" # A clean re-entry records a fresh posture, and an immediate return is # idempotently clear because the keyed blocker is resolved. diff --git a/tests/fm-afk-return.test.sh b/tests/fm-afk-return.test.sh index 153f9c94e5a..4b76e7f9947 100755 --- a/tests/fm-afk-return.test.sh +++ b/tests/fm-afk-return.test.sh @@ -4,8 +4,10 @@ # Covers the second half of the 2026-07-14 incident: an away-mode blocked event # survived in durable state, but the ordinary return request could proceed to # Bearings before Firstmate owned remediation. The shared script now stops, -# drains, preserves evidence, and refuses ordinary work until every live open -# `blocked:` event is resolved or durably reclassified. +# drains, preserves evidence, and holds ordinary WORK until every live open +# `blocked:` event is resolved or durably reclassified. Reporting is not work: +# Bearings renders behind the catch-up gate and surfaces the catch-up posture +# as content, so a returning captain still gets the picture. # The brief cases pin the away-posture redesign's return: the brief is composed # from the archived posture record, the outcome store, the held set, and the # status logs, health first, and the gate shrinks to what the away session could @@ -94,11 +96,21 @@ EOF printf 'blocked [key=%s]: firstmate can refresh the synthetic token\n' "$key" > "$dir/home/state/repair-task.status" } -test_return_gate_orders_catchup_before_bearings() { - local dir out rc gate wake_count +test_return_gate_owns_remediation_and_reports_catchup_to_bearings() { + local dir out rc gate wake_count i toon gate_header dir="$TMP_ROOT/ordering" install_runner "$dir" seed_live_blocker "$dir" herdr synthetic-dependency + { + printf '## In flight\n\n## Queued\n' + i=1 + while [ "$i" -le 20 ]; do + printf -- '- [ ] queued-%02d - Queued gate %02d (repo: sample) (kind: ship) (since 2026-06-%02d)\n' \ + "$i" "$i" "$i" + i=$((i + 1)) + done + printf '\n## Done\n' + } > "$dir/home/data/backlog.md" date +%s > "$dir/home/state/.afk" printf 'repair-task.status: blocked synthetic dependency\n' > "$dir/home/state/.subsuper-escalations" printf 'fm away-mode inject WEDGED: 4555s undelivered\n' > "$dir/home/state/.subsuper-inject-wedged" @@ -124,14 +136,40 @@ test_return_gate_orders_catchup_before_bearings() { [ -s "$dir/home/state/.fake-drain" ] || fail "blocked return acknowledged its emitted wake before handling completed" [ ! -e "$dir/home/state/.fake-drain-acks" ] || fail "blocked return crossed the post-handling acknowledgement boundary" - # The exact incident regression: Bearings is an ordinary request and must - # refuse before reading/rendering while this shared gate remains open. + # The captain is back and asking for the picture: Bearings reports the + # catch-up posture as content rather than refusing. The blocked worker still + # projects as its own Underway row, and the catch-up posture is a separate + # action-free Charted Next gate row that never becomes a Captain's Call entry. + out=$(FM_HOME="$dir/home" FM_STATE_OVERRIDE="$dir/home/state" "$ROOT/bin/fm-bearings-snapshot.sh" --json 2>&1) \ + || fail "Bearings should render behind the return catch-up gate: $out" + # The live projected state of the blocked worker follows its endpoint, which + # this fixture deliberately does not stand up; what the gate must no longer + # do is stop the fleet read, so the worker has to reach Underway at all. + printf '%s' "$out" | jq -e ' + (.in_flight | any(.id == "repair-task")) + and (.gates[0].id == "(return-catchup)" and .gates[0].filed == null) + and (.gates | length == 21) + and ([.gates[] | select(.id | startswith("queued-"))] | length == 20) + and (.gates | any(.id == "(return-catchup)" + and .owner == "(main)" + and .reason == "away-return catch-up" + and (.title | test("^1 blocker")))) + and ([.decisions_open[].id] | index("(return-catchup)") | not)' >/dev/null \ + || fail "Bearings did not reserve the catch-up posture outside bounded action-free gate rows: $out" + toon=$(FM_HOME="$dir/home" FM_STATE_OVERRIDE="$dir/home/state" "$ROOT/bin/fm-bearings-snapshot.sh" 2>&1) \ + || fail "default Bearings should render behind the return catch-up gate: $toon" + gate_header=$(printf '%s\n' "$toon" | awk '/^gates\[[0-9]+\]\{/ { print; exit }') + assert_contains "$gate_header" '{id,title,blocked_by,reason,owner,filed}' "catch-up removed filed from the TOON gate schema" + assert_contains "$toon" '2026-06-20' "catch-up removed durable gate dates from default Bearings output" + + # The guard itself still separates its two branches by exit status, so an + # active away window keeps refusing while catch-up reports. set +e - out=$(FM_HOME="$dir/home" FM_STATE_OVERRIDE="$dir/home/state" "$ROOT/bin/fm-bearings-snapshot.sh" --json 2>&1) + out=$(FM_HOME="$dir/home" FM_STATE_OVERRIDE="$dir/home/state" "$dir/bin/fm-afk-return.sh" guard 2>&1) rc=$? set -e - [ "$rc" -eq 3 ] || fail "Bearings should refuse behind the return gate (rc=$rc): $out" - assert_contains "$out" 'return catch-up is pending' "Bearings refusal did not point to the shared return owner" + [ "$rc" -eq 4 ] || fail "the catch-up branch should be distinguishable by exit status (rc=$rc): $out" + assert_contains "$out" 'return catch-up is pending' "the catch-up refusal did not point to the shared return owner" # Restart/re-entry is idempotent: no second stop, no duplicate catch-up line, # and the same unresolved blocker remains authoritative. @@ -148,6 +186,9 @@ test_return_gate_orders_catchup_before_bearings() { printf 'resolved [key=synthetic-dependency]: refreshed the synthetic token and resumed the task\n' >> "$dir/home/state/repair-task.status" out=$(run_return "$dir" check) || fail "resolved blocker did not clear return catch-up: $out" + FM_HOME="$dir/home" FM_STATE_OVERRIDE="$dir/home/state" "$ROOT/bin/fm-bearings-snapshot.sh" --json \ + | jq -e '[.gates[].id] | index("(return-catchup)") | not' >/dev/null \ + || fail "the cleared gate left the catch-up posture row in Bearings" assert_contains "$out" 'catch-up clear' "successful check did not announce that ordinary work may proceed" [ ! -e "$gate" ] || fail "successful check left the return gate behind" [ ! -e "$dir/home/state/.subsuper-escalations" ] || fail "successful check left delivered escalation state behind" @@ -162,7 +203,7 @@ test_return_gate_orders_catchup_before_bearings() { out=$(run_return "$dir" check) || fail "an already-clear repeated check should be idempotent: $out" [ ! -e "$gate" ] || fail "idempotent clear check recreated a gate" - pass "return catch-up precedes Bearings, owns live blocker remediation, preserves evidence once, and clears idempotently" + pass "return catch-up owns live blocker remediation, reports itself to Bearings as content, preserves evidence once, and clears idempotently" } test_explicit_reclassification_requires_durable_reason() { @@ -508,7 +549,7 @@ test_unreadable_outcome_store_keeps_catchup_gated() { } test_failed_held_listing_keeps_catchup_gated() { - local dir out waiting rc gate + local dir out waiting rc gate guard_out guard_rc dir="$TMP_ROOT/held-list-failure" install_runner "$dir" mkdir -p "$dir/fakebin" @@ -528,6 +569,21 @@ SH set -e [ "$rc" -eq 3 ] || fail "a failed held-set read should keep catch-up gated (rc=$rc): $out" [ -f "$gate" ] || fail "a failed held-set read did not retain the return gate" + + # A gate retained for a lifecycle reason lists no blocker at all, so the + # refusal must name what actually holds it instead of promising a blocker + # list it cannot produce, and Bearings must carry that same reason. + set +e + guard_out=$(FM_HOME="$dir/home" FM_STATE_OVERRIDE="$dir/home/state" "$dir/bin/fm-afk-return.sh" guard 2>&1) + guard_rc=$? + set -e + [ "$guard_rc" -eq 4 ] || fail "a blockerless catch-up gate should use the catch-up branch (rc=$guard_rc): $guard_out" + assert_contains "$guard_out" 'no open blocker' "the blockerless refusal did not say the gate lists no blocker" + assert_contains "$guard_out" 'catch-up retained: held set unreadable' "the blockerless refusal did not name the retention reason" + assert_not_contains "$guard_out" 'every listed blocker' "the blockerless refusal still demanded an empty blocker list" + FM_HOME="$dir/home" FM_STATE_OVERRIDE="$dir/home/state" "$ROOT/bin/fm-bearings-snapshot.sh" --json \ + | jq -e '.gates | any(.id == "(return-catchup)" and (.title | startswith("catch-up retained:")))' >/dev/null \ + || fail "Bearings did not carry the blockerless catch-up retention reason" waiting=$(printf '%s\n' "$out" | awk '/^Waiting on you:/{show=1} /^Tried and failed, or could not be fixed:/{show=0} show') assert_contains "$waiting" "held listing unavailable: $dir/home/data/backlog.md: synthetic held backlog failure; catch-up stays gated" "the failed held listing was not disclosed" assert_not_contains "$waiting" '(nothing)' "an unavailable held set was also reported as empty" @@ -693,7 +749,7 @@ test_missing_final_archive_keeps_retained_contract_gated() { pass "the retained contract epoch requires its final archive on every check" } -test_return_gate_orders_catchup_before_bearings +test_return_gate_owns_remediation_and_reports_catchup_to_bearings test_explicit_reclassification_requires_durable_reason test_captain_decision_does_not_masquerade_as_firstmate_blocker test_evidence_publication_failure_preserves_wake_for_redrain From a27646c4eae5d807027c3ebcb783234e0d212958 Mon Sep 17 00:00:00 2001 From: ShaDev <shazellb@gmail.com> Date: Fri, 11 Sep 2026 21:23:52 -0500 Subject: [PATCH 15/31] fix(bin): address the home's backlog from any directory and detect a forked code-root copy (#4223) * fix(backlog): address the home's backlog from any directory and detect a forked code-root copy A home outside the code root forks its queue: the tracked .tasks.toml names data/backlog.md relative to tasks-axi's working directory, so a bare tasks-axi call from the code root writes the code root's data/ while session start, spawn, and teardown use $FM_HOME/data. Linking the code-root copy into the home does not hold, because tasks-axi 0.2.4 writes by renaming a temp file over its target and rename(2) replaces a symlink: add, start, hold, and done from the code root each turn the link back into a regular file. The archive path is resolved against the working directory too, even with --file. bin/fm-tasks-axi.sh runs tasks-axi against this home's backlog from any directory, using the lifecycle transitions' existing addressing (run from the data directory's parent, pin <data>/backlog.md through TASKS_AXI_FILE). It keeps relative --to/--*-file arguments meaning the caller's paths, and refuses a caller --file, an unresolvable home, and a symlinked home backlog. The fm-send hold lookup, fm-public-followup, and the fm-decision-hold shim, which relied on cwd discovery, now go through it with an explicit FM_HOME and a cleared data override, so they keep addressing exactly $FM_HOME/data and an ambient TASKS_AXI_FILE cannot divert them; every agent-facing backlog command names it instead of bare tasks-axi. Bootstrap gains a detect-only BACKLOG_RECONCILE check, also run read-only: when the home's data directory is not the code root's, a code-root data/backlog.md or data/done-archive.md that is not the home's own file is reported as a fork, with the merge procedure in bootstrap-diagnostics. * test(teardown): assert the completion hint names bin/fm-tasks-axi.sh ready The completion hint now points at the home-addressed command instead of a bare tasks-axi call, so the dependency-cleared follow-up assertion checks for that command. * no-mistakes(test): clear ambient tasks-axi env in tests/lib.sh * no-mistakes(document): drop bare tasks-axi example from cd-guard doc * no-mistakes(lint): replace ls -A decoy listing with find for SC2012 * no-mistakes: apply CI fixes * revert: keep the compliance gate unchanged; the synchronize race is filed separately --- .agents/skills/bootstrap-diagnostics/SKILL.md | 3 + .agents/skills/fmx-respond/SKILL.md | 2 +- .agents/skills/stow/SKILL.md | 4 +- AGENTS.md | 4 +- bin/fm-backlog-handoff.sh | 2 +- bin/fm-bootstrap.sh | 22 ++ bin/fm-branch-prompt.sh | 4 +- bin/fm-decision-hold.sh | 4 +- bin/fm-public-followup.sh | 12 +- bin/fm-send.sh | 2 +- bin/fm-session-start.sh | 14 +- bin/fm-spawn.sh | 2 +- bin/fm-tasks-axi.sh | 127 ++++++++++ bin/fm-teardown.sh | 4 +- bin/fm-x-link.sh | 2 +- docs/architecture.md | 2 +- docs/cd-guard.md | 2 +- docs/configuration.md | 4 + docs/scripts.md | 1 + tests/fm-public-followup.test.sh | 50 ++++ tests/fm-session-start.test.sh | 4 +- tests/fm-tasks-axi.test.sh | 230 ++++++++++++++++++ tests/fm-teardown.test.sh | 2 +- tests/lib.sh | 10 + 24 files changed, 481 insertions(+), 32 deletions(-) create mode 100755 bin/fm-tasks-axi.sh create mode 100755 tests/fm-tasks-axi.test.sh diff --git a/.agents/skills/bootstrap-diagnostics/SKILL.md b/.agents/skills/bootstrap-diagnostics/SKILL.md index 4e5ab2efa04..1d49f4b8312 100644 --- a/.agents/skills/bootstrap-diagnostics/SKILL.md +++ b/.agents/skills/bootstrap-diagnostics/SKILL.md @@ -61,6 +61,9 @@ When any diagnostic needs captain attention, report the plain consequence and re Resolve the named backlog read problem and rerun session start; never guess by starting or closing an unreadable item. - `BACKLOG_RECONCILE: <id>: worker record exists but its backlog item could not be moved to In flight: <reason>` - this home owns a worker whose backlog item is still queued, and the reconciliation could not correct it. Until it is corrected, the fleet view reads that worker as work no backlog item owns; resolve the named backlog problem and rerun session start. +- `BACKLOG_RECONCILE: code-root <file> is not this home's <file>; ...` - a tasks-axi write addressed the code root instead of this home, so the queue has already forked and either copy may hold rows the other lacks; [`docs/configuration.md`](../../../docs/configuration.md) ("Backlog backend") owns why. + Neither copy is a safe winner: union-merge them into this home's file by task id, resolve each conflicting id to its most recent real transition, check this home's archive before treating a missing Done row as lost, and verify the merged id set equals the union of both inputs before installing it. + Then move the code-root file aside rather than deleting it, tell the captain which rows were recovered, and run every later backlog command through `bin/fm-tasks-axi.sh`; re-linking the code-root copy is never the fix, because the next cwd-relative tasks-axi write replaces the link again. - `SECONDMATE_SYNC: secondmate <id>: skipped: <reason>` - secondmate convergence left a live home on its existing checkout because the home was dirty, diverged, unsafe, on the wrong branch, missing its placement-specific target commit, unreachable, or otherwise not fast-forwardable, or because inherited local-material propagation failed; bootstrap continued, but inspect the reason because the secondmate's tracked instructions, inherited settings, or shared captain preferences may be stale after a primary update. - `SECONDMATE_LIVENESS: secondmate <id>: skipped: <reason>|respawn failed after <cause>: <reason>` - the session-start liveness sweep could not guarantee that the registered secondmate is running a real agent process. Investigate the reason because that secondmate is not guaranteed live. diff --git a/.agents/skills/fmx-respond/SKILL.md b/.agents/skills/fmx-respond/SKILL.md index 0c704b38494..39cafc2961f 100644 --- a/.agents/skills/fmx-respond/SKILL.md +++ b/.agents/skills/fmx-respond/SKILL.md @@ -257,7 +257,7 @@ So treat second-mate-routed Relay work as a promised final by construction: the **When you promise a final (including every Relay request whose work is routed to a second mate):** -1. Create the typed obligation with `tasks-axi public-followup add` and bind the work with `bind-work`, keeping the public-safe summary and the opaque thread binding in the obligation and the full request context where the poll already put it. +1. Create the typed obligation with `bin/fm-tasks-axi.sh public-followup add` and bind the work with its `bind-work`, keeping the public-safe summary and the opaque thread binding in the obligation and the full request context where the poll already put it. When the public ask plainly implies follow-on work ("look into X and fix it"), register the promised-final against the outcome and deliver any interim report as a separate `--purpose milestone` obligation on the same thread. An ask that genuinely terminates at a report stays `report-ready`; do not invent a ship commitment for work the captain has not authorized. 2. Register it with `bin/fm-public-followup.sh register <obligation-id> --relation <relation-id> --work-home <main|secondmate:<id>> --work-id <task-id> --generation <n>`. diff --git a/.agents/skills/stow/SKILL.md b/.agents/skills/stow/SKILL.md index ed32e020fee..8b86468011d 100644 --- a/.agents/skills/stow/SKILL.md +++ b/.agents/skills/stow/SKILL.md @@ -193,7 +193,7 @@ A local skill exists only in this home, so offloading an entry out of `data/capt Autonomously relocate it only by adding it to an already-existing allowed JIT note, or by routing it through a project's established delivery path to its existing owning `AGENTS.md`, then confirming that destination holds the quoted entry before removing the memory entry. A destination that needs creation, uncompleted project delivery, or any other future work is not live and cannot count as relief, so continue with the next archival or eviction rung instead of leaving an over-budget proposal pending. 2. Propose pinned relocation only. - For a pinned candidate, append a `proposed-offload` section with the same fields to the completion receipt, create or refresh one durable backlog item with `tasks-axi add`, `tasks-axi show <id> --full`, and `tasks-axi update <id> --body-file <path>` as appropriate, then hold it through `bin/fm-captain-hold.sh hold`. + For a pinned candidate, append a `proposed-offload` section with the same fields to the completion receipt, create or refresh one durable backlog item with `bin/fm-tasks-axi.sh add`, `bin/fm-tasks-axi.sh show <id> --full`, and `bin/fm-tasks-axi.sh update <id> --body-file <path>` as appropriate, then hold it through `bin/fm-captain-hold.sh hold`. Preserve each candidate's approval state in that item, and require explicit plain-chat approval for that named item before any migration. If the captain never answers, nothing migrates and the held item persists, but it is never treated as budget relief. 3. Migrate an approved pinned candidate outside this pass. @@ -223,7 +223,7 @@ A local skill exists only in this home, so offloading an entry out of `data/capt - Project-intrinsic knowledge never goes directly into a project's `AGENTS.md`. Route it through a normal ship task so a crewmate records it with `bin/fm-ensure-agents-md.sh` and the project's delivery path. - Knowledge general to every Firstmate user belongs in this repo's shared tracked material through the normal branch, no-mistakes, PR, and captain-merge path. - - For task-scoped notes, inspect the item with `tasks-axi show <id> --full`, classify the change as new, duplicate, superseding, or obsolete, then use a considered replacement body through `tasks-axi update <id> --body-file <path>`. + - For task-scoped notes, inspect the item with `bin/fm-tasks-axi.sh show <id> --full`, classify the change as new, duplicate, superseding, or obsolete, then use a considered replacement body through `bin/fm-tasks-axi.sh update <id> --body-file <path>`. Use `--archive-body` when recoverability matters. Never append. - File each undone next step as a queued backlog item with a genuine `blocked-by` dependency when applicable. diff --git a/AGENTS.md b/AGENTS.md index b709a2b9c84..3abd8348a28 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -515,14 +515,14 @@ Mention cost as a courtesy when unusually much work is running, but never block The configured `tasks-axi` backend is the durable queue; the tracked default is `data/backlog.md`. It tracks work items only, never agents; persistent secondmates never appear as backlog items. Work routed to a secondmate is recorded in that secondmate home's own backlog, not the main backlog. -A decision is simply a task held for the captain: create the task with `tasks-axi add` when needed, then always hold it through `bin/fm-captain-hold.sh hold <id> --reason "<reason>"`, with `--until <date>` when the captain defers it. +A decision is simply a task held for the captain: create the task with `bin/fm-tasks-axi.sh add` when needed, then always hold it through `bin/fm-captain-hold.sh hold <id> --reason "<reason>"`, with `--until <date>` when the captain defers it. When a main-side thread such as a pending captain decision or relay reminder is worth durable tracking, file it as its own work item and hold it through that wrapper. Captain calls discovered by investigations or visual reviews follow `captain-hold-lifecycle`, which owns their completion gate and recorded-answer rules. When the automatic transition gate applies, dispatch and completion move the item themselves - `bin/fm-spawn.sh` and `bin/fm-teardown.sh` own those transitions and refuse rather than report success without them - so what remains yours is filing the item before dispatch, recording decisions, and keeping notes current; `docs/configuration.md` owns gate applicability and the manual-backend exception. Re-evaluate queued work after every teardown and heartbeat, dispatching items only when dependencies and time gates have cleared. `.tasks.toml`, `docs/configuration.md`, and current `tasks-axi --help` own the backlog schema, compatibility, retention, and routine command syntax. -Use compatible `tasks-axi` when the configured backend selects it and the documented manual path otherwise; keep only the configured recent Done entries. +Use compatible `tasks-axi` when the configured backend selects it, always through `bin/fm-tasks-axi.sh` so the call reaches this home's backlog from any directory, and the documented manual path otherwise; keep only the configured recent Done entries. `secondmate-provisioning` and `bin/fm-backlog-handoff.sh` own cross-home handoff safety. Keep free-form notes free of temporary paths, moving versions, ephemeral identifiers, and copied state that will rot. diff --git a/bin/fm-backlog-handoff.sh b/bin/fm-backlog-handoff.sh index b40aae28758..dff23761c1f 100755 --- a/bin/fm-backlog-handoff.sh +++ b/bin/fm-backlog-handoff.sh @@ -317,7 +317,7 @@ warn_stale_public_commitments() { # <secondmate-id> <moved-key>... out=$("$SCRIPT_DIR/fm-public-followup.sh" guard-work main "$key" 2>/dev/null) || rc=$? [ "$rc" -ne 0 ] || continue [ -z "$out" ] || printf '%s\n' "$out" >&2 - printf 'warning: %s still owes a public reply bound to main/%s; rebind it to secondmate:%s (tasks-axi public-followup bind-work, then bin/fm-public-followup.sh register <obligation-id> --relation <relation-id> --work-home secondmate:%s --work-id %s --generation <n>) or the promised reply will be reconciled against work this home no longer owns.\n' \ + printf 'warning: %s still owes a public reply bound to main/%s; rebind it to secondmate:%s (bin/fm-tasks-axi.sh public-followup bind-work, then bin/fm-public-followup.sh register <obligation-id> --relation <relation-id> --work-home secondmate:%s --work-id %s --generation <n>) or the promised reply will be reconciled against work this home no longer owns.\n' \ "$key" "$key" "$id" "$id" "$key" >&2 done if fm_pf_relay_active "$FM_HOME" && fm_pf_has_delivered_open_loops "$STATE"; then diff --git a/bin/fm-bootstrap.sh b/bin/fm-bootstrap.sh index 8441dadea89..b5e010905cf 100755 --- a/bin/fm-bootstrap.sh +++ b/bin/fm-bootstrap.sh @@ -15,6 +15,7 @@ # "HOME_SUMMARY: <ledger never published|not republished since # <stamp>>; <n> failed attempt(s) ... last: <recorded failure>", # "BACKLOG_RECONCILE: <id>: <what this home could not reconcile>", +# "BACKLOG_RECONCILE: code-root <file> is not this home's <file>; ...", # "TANGLE: <remediation>", # "SECONDMATE_SYNC: secondmate <id>: skipped: <reason>", # "NUDGE_SECONDMATES: secondmate <id>: send failed: <reason>", @@ -97,6 +98,9 @@ # reads or writes another home; the fleet snapshot's classifier and # bin/fm-secondmate-reconcile.sh's nudge stay as backstops. Replayed # transitions and restored In-flight rows print BOOTSTRAP_INFO facts. +# The `code-root <file>` variant is a detect-only local check that runs +# even in a read-only session; detect_code_root_backlog_fork owns what +# it reports. # Set FM_BOOTSTRAP_DETECT_ONLY=1 to skip the six MUTATING sweeps # (backlog_record_reconcile, secondmate_sync, # secondmate_liveness_sweep, secondmate_handoff_resume, x_mode_setup, @@ -1478,9 +1482,27 @@ detect_local_config() { && ! fm_backlog_backend_manual "$CONFIG" && fm_tasks_axi_compatible; then echo "BOOTSTRAP_INFO: tasks-axi available" fi + detect_code_root_backlog_fork detect_home_summary_publication } +# Shadow-backlog check. When this home's data directory is not the code root's, +# a code-root data/backlog.md or data/done-archive.md that is not this home's +# own file is a queue a cwd-relative tasks-axi write has already forked; a link +# into the home does not survive such a write (docs/configuration.md "Backlog +# backend" owns why). Detect-only: neither copy is a safe winner, so nothing is +# merged here. +detect_code_root_backlog_fork() { + local name root_copy + [ "$FM_ROOT/data" -ef "$DATA" ] && return 0 + for name in backlog.md done-archive.md; do + root_copy="$FM_ROOT/data/$name" + [ -e "$root_copy" ] || [ -L "$root_copy" ] || continue + [ "$root_copy" -ef "$DATA/$name" ] && continue + echo "BACKLOG_RECONCILE: code-root $root_copy is not this home's $DATA/$name; tasks-axi wrote the code root instead of this home, so rows in it may be missing here - merge it into this home's copy and move it aside" + done +} + # This home's ledger publication is deliberately best-effort: every lifecycle # trigger calls it with --best-effort so a failure can never change the result # of a session start, a spawn, a teardown, or a watcher poll. That is correct, diff --git a/bin/fm-branch-prompt.sh b/bin/fm-branch-prompt.sh index c426f7ec8d2..0ed62dd0552 100755 --- a/bin/fm-branch-prompt.sh +++ b/bin/fm-branch-prompt.sh @@ -45,9 +45,9 @@ Handle it start to finish in one turn sequence: 1. Drain first: run `bin/fm-wake-drain.sh` and read every presented record, plus any OPEN DECISIONS, UNREAD STATUS, and RECORD DIVERGENCE sections. 2. For each task you are about to mutate, claim its lease first: `bin/fm-lease.sh claim <task>`. - Claim the reserved `backlog` lease around backlog writes (`bin/fm-lease.sh claim backlog`, then `tasks-axi ...`, then release). + Claim the reserved `backlog` lease around backlog writes (`bin/fm-lease.sh claim backlog`, then `bin/fm-tasks-axi.sh ...`, then release). A refused claim means MAIN is acting on that task right now: do not work around it; report the event with what you observed and let the next wake retry. -3. Handle with real tools: `bin/fm-crew-state.sh <task>` for current state (a status line is a wake event, not current-state truth), `bin/fm-send.sh` for a short steer, `bin/fm-control.sh <task> interrupt|exit|relaunch` for lifecycle, `bin/fm-pr-check.sh <task> <url>` when the task's ready status or `pr=` metadata names the PR's URL, `tasks-axi` for backlog moves. +3. Handle with real tools: `bin/fm-crew-state.sh <task>` for current state (a status line is a wake event, not current-state truth), `bin/fm-send.sh` for a short steer, `bin/fm-control.sh <task> interrupt|exit|relaunch` for lifecycle, `bin/fm-pr-check.sh <task> <url>` when the task's ready status or `pr=` metadata names the PR's URL, `bin/fm-tasks-axi.sh` for backlog moves. 4. Report: call the fm_branch_report tool exactly once per handled event, with the task id, the verdict, and a one-or-two-sentence summary; set silent true only for a fleet-wide heartbeat review that found literally nothing worth reporting. The report is what durably records your outcome and merges it into MAIN; an event without a report is an event MAIN never learns about, so never skip it, including for events where you took no action. 5. Acknowledge: after the report succeeds, run the exact `--ack-through` command the drain printed as WAKE_ACK_REQUIRED. diff --git a/bin/fm-decision-hold.sh b/bin/fm-decision-hold.sh index c1a7a6c9f03..f5538deed16 100755 --- a/bin/fm-decision-hold.sh +++ b/bin/fm-decision-hold.sh @@ -59,7 +59,7 @@ compose() { # <origin> <key> } task_show() { - (cd "$FM_HOME" && tasks-axi show "$1" --full) 2>/dev/null + FM_HOME="$FM_HOME" FM_DATA_OVERRIDE='' "$SCRIPT_DIR/fm-tasks-axi.sh" show "$1" --full 2>/dev/null } show_field() { @@ -169,7 +169,7 @@ command_resolve() { for dep in $routed; do show=$(task_show "$dep") || fail "routed task $dep disappeared before routing" if list_has_key "$(normalized_blocked_by "$show")" "$id"; then - (cd "$FM_HOME" && tasks-axi unblock "$dep" --by "$id" >/dev/null) \ + FM_HOME="$FM_HOME" FM_DATA_OVERRIDE='' "$SCRIPT_DIR/fm-tasks-axi.sh" unblock "$dep" --by "$id" >/dev/null \ || fail "could not route the recorded decision to $dep" fi done diff --git a/bin/fm-public-followup.sh b/bin/fm-public-followup.sh index ea93e801dcc..ea5173902d5 100755 --- a/bin/fm-public-followup.sh +++ b/bin/fm-public-followup.sh @@ -208,9 +208,11 @@ require_tools() { command -v tasks-axi >/dev/null 2>&1 || die "tasks-axi is required" 1 } -# Every tasks-axi call runs from the home whose backlog owns the obligation, the -# same convention bin/fm-captain-hold.sh uses for typed backlog state. -tx() { (cd "$FM_HOME" && tasks-axi "$@"); } +# Every tasks-axi call addresses $FM_HOME/data, the home whose backlog owns the +# obligation, through bin/fm-tasks-axi.sh. An inherited FM_DATA_OVERRIDE is +# cleared because a caller such as a secondmate teardown names the parent home +# in FM_HOME while its own data override is still in the environment. +tx() { FM_HOME="$FM_HOME" FM_DATA_OVERRIDE='' "$SCRIPT_DIR/fm-tasks-axi.sh" "$@"; } # obligation_json <id>: the complete typed obligation payload on stdout, empty # when the backlog simply has no such public-followup item, and a non-zero exit @@ -285,7 +287,7 @@ cmd_register() { payload=$(obligation_json "$id") \ || die "could not read the backlog through tasks-axi" 1 [ -n "$payload" ] \ - || die "no public-followup obligation '$id' in this home's backlog; create it with tasks-axi public-followup add before registering" 1 + || die "no public-followup obligation '$id' in this home's backlog; create it with bin/fm-tasks-axi.sh public-followup add before registering" 1 # The relation must already be bound, so a registration can never describe a # binding tasks-axi does not have. @@ -293,7 +295,7 @@ cmd_register() { '(.public_followup.work_relations // []) | map(select(.relation_id == $r and .work_ref.home_id == $h and .work_ref.task_id == $w)) | length > 0' >/dev/null 2>&1 \ - || die "obligation '$id' has no bound relation '$relation' for $work_home/$work_id; run tasks-axi public-followup bind-work first" 1 + || die "obligation '$id' has no bound relation '$relation' for $work_home/$work_id; run bin/fm-tasks-axi.sh public-followup bind-work first" 1 [ -n "$platform" ] || platform=$(pf_field "$payload" '.public_followup.request.platform') [ -n "$request" ] || request=$(pf_field "$payload" '.public_followup.request.request_id') diff --git a/bin/fm-send.sh b/bin/fm-send.sh index b4e61796959..aa74940c08e 100755 --- a/bin/fm-send.sh +++ b/bin/fm-send.sh @@ -545,7 +545,7 @@ fm_send_hold_resolved_id() { # <task-id> <decision-key> local show id state hold_kind command -v tasks-axi >/dev/null 2>&1 || return 1 for id in "$2" "$1-decision-$2"; do - show=$( (cd "$FM_HOME" && tasks-axi show "$id" --full) 2>/dev/null ) || continue + show=$(FM_HOME="$FM_HOME" FM_DATA_OVERRIDE='' "$SCRIPT_DIR/fm-tasks-axi.sh" show "$id" --full 2>/dev/null) || continue state=$(printf '%s\n' "$show" | sed -n 's/^ state: //p' | head -1) hold_kind=$(printf '%s\n' "$show" | sed -n 's/^ hold_kind: //p' | head -1) [ "$state" != "done" ] || continue diff --git a/bin/fm-session-start.sh b/bin/fm-session-start.sh index a5042957c38..59b9044ac58 100755 --- a/bin/fm-session-start.sh +++ b/bin/fm-session-start.sh @@ -154,7 +154,7 @@ # stay out of the startup digest; the same never-bound-a-held-or-blocked-row # rule applies, recognized there from the title line's own hold/blocked-by # markers. -# Full bodies are targeted follow-up only: `tasks-axi show <id> --full` when +# Full bodies are targeted follow-up only: `bin/fm-tasks-axi.sh show <id> --full` when # compatible tasks-axi is available, or `data/backlog.md` when the file body is # truly needed. # @@ -386,7 +386,7 @@ print_file_or_absent() { } print_backlog_pointer() { - printf 'Full task bodies remain available on demand: tasks-axi show <id> --full when compatible tasks-axi is available, or data/backlog.md.\n' + printf 'Full task bodies remain available on demand: bin/fm-tasks-axi.sh show <id> --full when compatible tasks-axi is available, or data/backlog.md.\n' } # A queued title line whose own text already marks it held or blocked. The @@ -453,8 +453,8 @@ strip_axi_help() { # and every other line it prints (its count, its public-followup line) passes # through untouched. Whatever is cut is disclosed exactly. print_ready_queued_bounded() { - local ready=$1 path=$2 - printf '%s\n' "$ready" | awk -v max="$QUEUED_LIMIT" -v path="$path" ' + local ready=$1 + printf '%s\n' "$ready" | awk -v max="$QUEUED_LIMIT" ' /^help\[/ { exit } /^ready\[/ { rows = 1; print; next } rows && /^[[:space:]]/ { @@ -467,7 +467,7 @@ print_ready_queued_bounded() { if (total > 0) { printf "(shown %d of %d ready queued item(s))\n", shown, total if (total > shown) { - printf "(%d more queued - tasks-axi ready --file %s)\n", total - shown, path + printf "(%d more queued - bin/fm-tasks-axi.sh ready)\n", total - shown } } } @@ -494,7 +494,7 @@ print_backlog_tasks_axi_compact() { printf '\nblocked queued:\n' printf '%s\n' "$blocked" | strip_axi_help printf '\nready queued (dispatchable now):\n' - print_ready_queued_bounded "$ready" "$path" + print_ready_queued_bounded "$ready" return 0 fi printf 'tasks-axi compact listing failed; falling back to title-line rendering.\n' @@ -814,7 +814,7 @@ Go to a source directly only when: - an individual full status log is needed for older wake-event history, or a status line was capped and its tail matters (each task's full log path is printed with its tail), - - a full task body is needed (tasks-axi show <id> --full, or data/backlog.md), + - a full task body is needed (bin/fm-tasks-axi.sh show <id> --full, or data/backlog.md), - the backlog listing disclosed omitted queued items and this turn needs them, - the NETWORK CHECKS section reported its checks still IN PROGRESS and this turn needs their verdict (bin/fm-startup-network.sh report), diff --git a/bin/fm-spawn.sh b/bin/fm-spawn.sh index eefc4abbf51..df1181dd561 100755 --- a/bin/fm-spawn.sh +++ b/bin/fm-spawn.sh @@ -2639,7 +2639,7 @@ if fm_backlog_transition_applies "$CONFIG" "$DATA" "$KIND"; then if fm_backlog_row_probe "$DATA" "$ID"; then BACKLOG_ROW_STATE=$FM_BACKLOG_ROW_STATE elif [ "$FM_BACKLOG_ROW_RESULT" = not_found ]; then - echo "error: task $ID has no backlog item in this home, so dispatching it would leave a worker no record owns; add it first (tasks-axi add $ID '<title>' --kind $KIND) and re-run" >&2 + echo "error: task $ID has no backlog item in this home, so dispatching it would leave a worker no record owns; add it first (bin/fm-tasks-axi.sh add $ID '<title>' --kind $KIND) and re-run" >&2 exit 1 else echo "error: task $ID's backlog item could not be read before dispatch ($FM_BACKLOG_ROW_ERROR)" >&2 diff --git a/bin/fm-tasks-axi.sh b/bin/fm-tasks-axi.sh new file mode 100755 index 00000000000..b8e2844c0e5 --- /dev/null +++ b/bin/fm-tasks-axi.sh @@ -0,0 +1,127 @@ +#!/usr/bin/env bash +# fm-tasks-axi.sh - run tasks-axi against THIS home's backlog from any working directory. +# +# Usage: fm-tasks-axi.sh [<tasks-axi command> [args...]] +# fm-tasks-axi.sh --help +# +# Every routine firstmate backlog read or mutation goes through this command +# rather than a bare `tasks-axi`; `fm-tasks-axi.sh <command> --help` prints +# tasks-axi's own help. Arguments reach tasks-axi as given, apart from one +# rewrite that keeps file arguments meaning what the caller meant: a relative +# value of `--to` or any `--*-file` flag (`--body-file`, `--relation-file`, ...) +# is made absolute against the caller's working directory, because tasks-axi +# starts from the backlog root instead. `--report` stays as given: tasks-axi +# stores it verbatim as a link, which lifecycle transitions record relative to +# that same root. +# +# Why it exists: a bare `tasks-axi` resolves the tracked `.tasks.toml` paths +# against its working directory, so from the code root it forks the queue +# whenever the home lives elsewhere; docs/configuration.md ("Backlog backend") +# owns that rationale. +# +# Addressing is bin/fm-backlog-transition-lib.sh's fm_backlog_tasks_axi_addressing, +# the same resolution the lifecycle transitions use: tasks-axi runs from the +# configured data directory's parent, so that home's own `.tasks.toml` (or +# tasks-axi's built-in defaults, which keep the archive beside the backlog) +# supplies the adapter, done_keep, and the archive path; a markdown backlog is +# additionally pinned to `<data>/backlog.md` through TASKS_AXI_FILE. The +# environment carries the pin rather than a trailing --file so the no-command +# dashboard works too. A configured non-markdown adapter is addressed by that +# root alone, so an inherited TASKS_AXI_FILE is cleared for it. +# +# The data directory is FM_DATA_OVERRIDE, else $FM_HOME/data, else the code +# root's data/ (FM_HOME unset keeps the single-home layout unchanged). +# +# Refusals (exit 2, nothing run): +# - tasks-axi missing from PATH; +# - a caller-supplied --file, because this command owns the addressing and +# tasks-axi would silently let the last --file win; +# - a data directory that cannot be resolved, or whose backend configuration +# cannot be read (bin/fm-tasks-axi-lib.sh owns that diagnostic); +# - a markdown `<data>/backlog.md` that is itself a symlink, because the +# first write would replace the link with a private copy, exactly the fork +# this command exists to prevent. Lifecycle transitions refuse the same file. +# Otherwise the exit status is tasks-axi's own. +set -u + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" +FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" +DATA="${FM_DATA_OVERRIDE:-$FM_HOME/data}" +# shellcheck source=bin/fm-tasks-axi-lib.sh disable=SC1091 +. "$SCRIPT_DIR/fm-tasks-axi-lib.sh" +# shellcheck source=bin/fm-backlog-transition-lib.sh disable=SC1091 +. "$SCRIPT_DIR/fm-backlog-transition-lib.sh" + +usage() { + awk ' + NR == 1 { next } + /^#/ { sub(/^# ?/, ""); print; next } + { exit } + ' "$0" +} + +fail() { + printf 'fm-tasks-axi: %s\n' "$*" >&2 + exit 2 +} + +case "${1:-}" in + -h|--help) + usage + exit 0 + ;; +esac + +CALLER_DIR=$(pwd) + +absolute_from_caller() { # <path-value> + case "$1" in + ''|-|/*) printf '%s' "$1" ;; + *) printf '%s/%s' "$CALLER_DIR" "$1" ;; + esac +} + +ARGS=() +path_value_next=0 +for arg in "$@"; do + if [ "$path_value_next" = 1 ]; then + ARGS+=("$(absolute_from_caller "$arg")") + path_value_next=0 + continue + fi + case "$arg" in + --file|--file=*) + fail "this command always addresses this home's backlog at $DATA; drop --file, or run tasks-axi directly for another backlog" + ;; + --to|--*-file) + ARGS+=("$arg") + path_value_next=1 + ;; + --to=*|--*-file=*) + ARGS+=("${arg%%=*}=$(absolute_from_caller "${arg#*=}")") + ;; + *) + ARGS+=("$arg") + ;; + esac +done + +command -v tasks-axi >/dev/null 2>&1 || fail "tasks-axi is not on PATH; run bin/fm-bootstrap.sh for the install command" + +FM_BACKLOG_TRANSITION_ERROR= +if ! fm_backlog_tasks_axi_addressing "$DATA"; then + fail "${FM_BACKLOG_TRANSITION_ERROR:-data directory cannot be resolved: $DATA}" +fi + +if [ -n "$FM_BACKLOG_AXI_FILE" ]; then + if [ -L "$FM_BACKLOG_AXI_FILE" ]; then + fail "$FM_BACKLOG_AXI_FILE is a symlink; a tasks-axi write would replace it with a regular file and fork the backlog - make it this home's real file" + fi + export TASKS_AXI_FILE="$FM_BACKLOG_AXI_FILE" +else + unset TASKS_AXI_FILE +fi + +cd "$FM_BACKLOG_AXI_ROOT" || fail "cannot enter the backlog root $FM_BACKLOG_AXI_ROOT" +exec tasks-axi ${ARGS[@]+"${ARGS[@]}"} diff --git a/bin/fm-teardown.sh b/bin/fm-teardown.sh index e8a6a195ce8..7eaabe22d5a 100755 --- a/bin/fm-teardown.sh +++ b/bin/fm-teardown.sh @@ -1433,7 +1433,7 @@ backlog_refresh_reminder() { if [ "$BACKLOG_CLOSED" = 1 ] && [ "$BACKLOG_TRANSITION" = retain ]; then printf '%s\n' "Backlog: $ID stays open in $backlog_display, still held for the captain with its deliverable recorded. Relay the question and close it only with bin/fm-captain-hold.sh answer." elif [ "$BACKLOG_CLOSED" = 1 ]; then - printf '%s\n' "Backlog: $ID is closed in $backlog_display. Run tasks-axi ready for dependency-cleared candidates, check date gates, and dispatch only work whose blockers are gone and date is due." + printf '%s\n' "Backlog: $ID is closed in $backlog_display. Run bin/fm-tasks-axi.sh ready for dependency-cleared candidates, check date gates, and dispatch only work whose blockers are gone and date is due." else printf '%s\n' "Backlog: $ID just finished ($BACKLOG_SKIP_REASON). Update $backlog_display - move $ID to Done, keep Done to the 10 most recent, then re-scan Queued and dispatch only work whose blockers are gone and date is due." fi @@ -3137,7 +3137,7 @@ if [ "$FORCE" != "--force" ] \ "$SCRIPT_DIR/fm-public-followup.sh" guard-work "$PUBLIC_FOLLOWUP_WORK_HOME" "$ID" 2>/dev/null); then echo "REFUSED: task $ID still owes a public reply through the myfirstmate relay." >&2 printf '%s\n' "$PUBLIC_FOLLOWUP_BLOCKING" >&2 - echo "Deliver it with bin/fm-public-followup.sh deliver <obligation-id>, waive it with tasks-axi public-followup waive, or use --force after explicit discard approval." >&2 + echo "Deliver it with bin/fm-public-followup.sh deliver <obligation-id>, waive it with bin/fm-tasks-axi.sh public-followup waive, or use --force after explicit discard approval." >&2 exit 1 fi fi diff --git a/bin/fm-x-link.sh b/bin/fm-x-link.sh index 13b881c0c7c..fd28ca5a11f 100755 --- a/bin/fm-x-link.sh +++ b/bin/fm-x-link.sh @@ -177,7 +177,7 @@ if [ ! -f "$META" ]; then ''|*' '*) ;; *) ROUTE_HOME_ARG="secondmate:$ROUTE_MATCHES" ;; esac - printf 'fm-x-link: bind the public promise through the promised-final path instead: tasks-axi public-followup add + bind-work, then bin/fm-public-followup.sh register <obligation-id> --relation <relation-id> --work-home %s --work-id %s --generation <n>, and put the bin/fm-public-followup.sh brief <obligation-id> command into the routed worker instructions.\n' \ + printf 'fm-x-link: bind the public promise through the promised-final path instead: bin/fm-tasks-axi.sh public-followup add + bind-work, then bin/fm-public-followup.sh register <obligation-id> --relation <relation-id> --work-home %s --work-id %s --generation <n>, and put the bin/fm-public-followup.sh brief <obligation-id> command into the routed worker instructions.\n' \ "$ROUTE_HOME_ARG" "$ID" >&2 fi exit 1 diff --git a/docs/architecture.md b/docs/architecture.md index e1f8bb57be1..7ab4e496f7d 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -392,7 +392,7 @@ Home-domain captain preferences go to `data/captain.md`, cross-domain shared cap Memory writes use inspect-then-update rather than blind append; the internal [`stow` skill](../.agents/skills/stow/SKILL.md) owns tier markers, decay, cold archival, and offload. The same pass also persists open-work record state the session is holding - filing a thread that was never recorded and correcting one the session knows went stale - bounded to the open work that session is actually holding. It is deliberately not a reconciliation of durable records against repository or PR reality: its input is the volatile context, so it can only preserve what the session still knows, and no reconciliation that outlives a session exists today. -Task-scoped notes use `tasks-axi show <id> --full` followed by `tasks-axi update <id> --body-file <path>`, adding `--archive-body` when the prior body should remain recoverable. +Task-scoped notes use `bin/fm-tasks-axi.sh show <id> --full` followed by `bin/fm-tasks-axi.sh update <id> --body-file <path>`, adding `--archive-body` when the prior body should remain recoverable. The stow pass never writes a skill, but a separately executed, captain-approved migration may move conditional knowledge into a user-owned local skill excluded from the Firstmate clone; changes to Firstmate's tracked skills remain deliberate repository work through the normal PR pipeline. Invoked in a primary home, `/stow` then cascades the same sweep to every registered secondmate, enumerated through `bin/fm-stow-cascade.sh`: each home is accounted and curated against its own startup-memory allowance, a live secondmate sweeps its own session, and a slow or unreachable home is reported as an exception rather than blocking the primary. diff --git a/docs/cd-guard.md b/docs/cd-guard.md index 2814912180c..ae4dae95eba 100644 --- a/docs/cd-guard.md +++ b/docs/cd-guard.md @@ -11,7 +11,7 @@ the watcher-arm PreToolUse seatbelt (`bin/fm-arm-pretool-check.sh`, `docs/arm-pr ## Purpose and boundary The primary firstmate shell persists its working directory across tool calls. -A stray persistent top-level `cd projects/<clone>` therefore silently relocates the shell, so the next firstmate-owned command - a backlog write, an `fm-*` lifecycle call, `tasks-axi` - runs inside a project clone instead of the home. +A stray persistent top-level `cd projects/<clone>` therefore silently relocates the shell, so the next firstmate-owned command - a backlog write, an `fm-*` lifecycle call - runs inside a project clone instead of the home. That has actually happened: a persistent top-level `cd` caused a firstmate-owned backlog write to execute inside a project clone rather than the home. The seatbelt denies exactly that command shape - a cwd change that persists to the primary shell - before it runs. diff --git a/docs/configuration.md b/docs/configuration.md index 59421eab7fb..3a9ef6e0b43 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -117,6 +117,10 @@ A `manual` home owns its backlog file outright: the lifecycle transitions above Absent or `tasks-axi` selects the tasks-axi path. On the default markdown adapter, tasks-axi and manual edits produce the same `## In flight`, `## Queued`, and `## Done` sections. +The tracked `.tasks.toml` paths resolve against the directory tasks-axi runs in, not `FM_HOME`, so a bare `tasks-axi` run from the code root addresses the code root's `data/` whenever the home lives elsewhere. +tasks-axi writes by renaming a temp file over its target, which replaces a symlink with a regular file, so linking the code-root copy into the home forks the queue on the first such write rather than keeping the two in step. +Every routine firstmate backlog command therefore runs through [`bin/fm-tasks-axi.sh`](../bin/fm-tasks-axi.sh), which addresses this home's backlog and archive from any working directory exactly as lifecycle transitions do, and bootstrap reports a code-root `data/backlog.md` or `data/done-archive.md` that is not this home's own file as a `BACKLOG_RECONCILE: code-root ...` line even in a read-only session. + ## Runtime backend (config/backend / FM_BACKEND) For spawn-capable adapters, the runtime session-provider backend controls where task windows/endpoints are created, captured, sent to, watched, and killed. diff --git a/docs/scripts.md b/docs/scripts.md index 4f84c42dcce..3f0694bac8d 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -102,6 +102,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-ff-lib.sh` | Shared guarded fast-forward helper for origin pulls and secondmate syncs | | `fm-lock-lib.sh` | Shared "is this git lock provably abandoned?" proof used by teardown and fleet-sync | | `fm-config-inherit-lib.sh` | Shared primary-to-secondmate inherited local-material propagation and config-reread delivery | +| `fm-tasks-axi.sh` | Run `tasks-axi` against this home's backlog from any working directory | | `fm-tasks-axi-lib.sh` | Shared backlog-backend selector and `tasks-axi` compatibility probe | | `fm-backlog-transition-lib.sh` | Pair task-record changes with their backlog transitions and replay interrupted closes | | `fm-quota-axi-lib.sh` | Shared `quota-axi` compatibility floor and quota snapshot schema validation | diff --git a/tests/fm-public-followup.test.sh b/tests/fm-public-followup.test.sh index c009bdcbbe9..c5551d99a0b 100755 --- a/tests/fm-public-followup.test.sh +++ b/tests/fm-public-followup.test.sh @@ -268,6 +268,55 @@ expect_failure() { fi } +# --- 0. the suite's seeding never reaches a real backlog ------------------------ + +# An operator shell exports TASKS_AXI_FILE at its live home's backlog, and +# tasks-axi resolves that env ahead of the fixture's .tasks.toml. This suite +# seeds obligations with bare `tasks-axi` from the fixture home, so before +# tests/lib.sh cleared the ambient overrides, every seed landed in the operator's +# real backlog while the consumer read the empty fixture. Re-enter the suite's +# seeding exactly as a test process begins (source tests/lib.sh, then seed from +# the fixture) under a decoy "live" backlog and a backend tasks-axi refuses: the +# decoy must stay byte-identical and the fixture must hold the obligation. +test_ambient_tasks_axi_env_never_reaches_a_real_backlog() { + local home decoy_dir decoy fixture_state + home=$(make_home ambient-env) + decoy_dir="$TMP_ROOT/ambient-live/data" + decoy="$decoy_dir/backlog.md" + mkdir -p "$decoy_dir" + printf '## In flight\n\n## Queued\n\n## Done\n' > "$decoy" + cp "$decoy" "$decoy.expected" + jq -n '{request_id:"req-amb", platform:"discord", + context_binding:{version:"ctx1", value:"ctx1_req-amb"}, + public_safe_summary:"seeded under an ambient tasks-axi override", + received_at:"2026-07-30T10:00:00Z", + followup_expires_at:"2026-08-06T10:00:00Z", + reservation_expires_at:"2026-08-06T10:00:00Z"}' > "$home/request.json" + jq -n '{type:"pr-merged", project:"firstmate", + required_deliverables:["pr_url"], completion_policy:"all-required"}' \ + > "$home/expected.json" + + TASKS_AXI_FILE="$decoy" TASKS_AXI_BACKEND=no-such-backend bash -c ' + set -u + . "$1/tests/lib.sh" + [ -z "${TASKS_AXI_FILE+x}" ] || { echo "TASKS_AXI_FILE survived tests/lib.sh"; exit 1; } + [ -z "${TASKS_AXI_BACKEND+x}" ] || { echo "TASKS_AXI_BACKEND survived tests/lib.sh"; exit 1; } + cd "$2" && tasks-axi public-followup add pf-ambient \ + --request-context-file "$2/request.json" --purpose promised-final \ + --expected-final-file "$2/expected.json" --expires-at 2026-10-01T00:00:00Z >/dev/null + ' _ "$ROOT" "$home" \ + || fail "seeding under an ambient tasks-axi override did not reach the fixture backlog" + + cmp -s "$decoy" "$decoy.expected" \ + || fail "the suite's seeding wrote the ambient TASKS_AXI_FILE backlog instead of the fixture" + [ "$(find "$decoy_dir" -mindepth 1 -maxdepth 1 -exec basename {} \; | LC_ALL=C sort | tr '\n' ' ')" = "backlog.md backlog.md.expected " ] \ + || fail "the suite's seeding left an artifact beside the ambient TASKS_AXI_FILE backlog: $(find "$decoy_dir" -mindepth 1 -maxdepth 1 -exec basename {} \; | LC_ALL=C sort | tr '\n' ' ')" + fixture_state=$(task_state "$home" pf-ambient) + [ "$fixture_state" != absent ] \ + || fail "the fixture backlog does not hold the obligation seeded under the ambient override" + pass "the suite's seeding never reaches a backlog named by ambient TASKS_AXI_FILE/BACKEND" +} + # --- 0. bounded, single-line, character-safe outcome text ----------------------- # The outcome sentence becomes a public reply, so bounding it must not mangle @@ -3095,6 +3144,7 @@ if [ -n "${FM_TEST_ONLY:-}" ]; then exit 0 fi +test_ambient_tasks_axi_env_never_reaches_a_real_backlog test_outcome_text_is_bounded_without_corrupting_characters test_restart_e2e_delivers_exactly_once test_duplicate_event_and_replay_are_noops diff --git a/tests/fm-session-start.test.sh b/tests/fm-session-start.test.sh index 223e2404882..cee8c0e6091 100755 --- a/tests/fm-session-start.test.sh +++ b/tests/fm-session-start.test.sh @@ -1753,7 +1753,7 @@ EOF assert_not_contains "$out" "DONE-ROW-LINE" "tasks-axi compact digest listed a done row at startup" assert_contains "$out" "--- compact-startup ---" "in-flight meta identity disappeared from startup recovery digest" assert_contains "$out" "worktree=$home/projects/firstmate" "in-flight recovery worktree identity disappeared from startup digest" - assert_contains "$out" "Full task bodies remain available on demand: tasks-axi show <id> --full" \ + assert_contains "$out" "Full task bodies remain available on demand: bin/fm-tasks-axi.sh show <id> --full" \ "compact digest omitted the full-body lookup pointer" assert_contains "$out" "ready_public_followups: 0 delivery-ready obligations" \ "the composed listing dropped a real signal from the dispatchable set" @@ -1797,7 +1797,7 @@ EOF assert_not_contains "$out" "ready-4,queued" "the queued bound did not actually bound the ready listing" assert_contains "$out" "(shown 3 of 7 ready queued item(s))" \ "the bounded queued listing did not report what it showed" - assert_contains "$out" "(4 more queued - tasks-axi ready --file $home/data/backlog.md)" \ + assert_contains "$out" "(4 more queued - bin/fm-tasks-axi.sh ready)" \ "the bounded queued listing did not disclose an exact remainder and how to see it" # The bound is for dispatchable work only: held and blocked rows stay whole. diff --git a/tests/fm-tasks-axi.test.sh b/tests/fm-tasks-axi.test.sh new file mode 100755 index 00000000000..5ceeabea836 --- /dev/null +++ b/tests/fm-tasks-axi.test.sh @@ -0,0 +1,230 @@ +#!/usr/bin/env bash +# Behavior tests for bin/fm-tasks-axi.sh home addressing and bootstrap's +# shadow-backlog check, over the split layout where the operational home lives +# outside the code root that carries the tracked .tasks.toml. +# +# The fork these guard against: .tasks.toml names data/backlog.md relative to +# the caller's working directory, and tasks-axi writes by renaming a temp file +# over its target, so a bare tasks-axi run from the code root turns a code-root +# symlink into the home's backlog into a private regular copy. The suite proves +# that every write through bin/fm-tasks-axi.sh lands in $FM_HOME/data from the +# code root (including archiving and relative --body-file arguments), +# that the command refuses addressing it cannot keep correct, and that bootstrap +# reports any code-root copy that is not this home's own file while staying +# silent for a link into the home, an absent copy, and the single-home layout. +set -u + +# shellcheck source=tests/lib.sh disable=SC1091 +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +WRAPPER="$ROOT/bin/fm-tasks-axi.sh" +BOOTSTRAP="$ROOT/bin/fm-bootstrap.sh" +TMP_ROOT=$(fm_test_tmproot fm-tasks-axi) +BASE_PATH=${FM_TEST_BASE_PATH:-/usr/bin:/bin:/usr/sbin:/sbin} + +# The developer shell may pin any of these; each case states its own layout. +unset TASKS_AXI_FILE TASKS_AXI_BACKEND FM_HOME FM_ROOT_OVERRIDE \ + FM_DATA_OVERRIDE FM_STATE_OVERRIDE FM_CONFIG_OVERRIDE FM_PROJECTS_OVERRIDE + +HAVE_TASKS_AXI=0 +command -v tasks-axi >/dev/null 2>&1 && HAVE_TASKS_AXI=1 + +empty_backlog() { # <path> + printf '## In flight\n\n## Queued\n\n## Done\n' > "$1" +} + +# A code root carrying the tracked .tasks.toml and an operational home beside +# it, with the code-root backlog linked into the home the way an operator +# would try to keep the two in sync. +make_split() { # <name>; prints the case directory + local dir="$TMP_ROOT/$1" + mkdir -p "$dir/code/data" "$dir/home/data" "$dir/home/state" "$dir/home/config" + cp "$ROOT/.tasks.toml" "$dir/code/.tasks.toml" + empty_backlog "$dir/home/data/backlog.md" + ln -s "$dir/home/data/backlog.md" "$dir/code/data/backlog.md" + printf '%s\n' "$dir" +} + +# Run the wrapper from the code root, as firstmate does. +wrapper_from_code() { # <case-dir> <tasks-axi args...> + local dir=$1 + shift + (cd "$dir/code" && FM_HOME="$dir/home" FM_ROOT_OVERRIDE="$dir/code" "$WRAPPER" "$@") +} + +# Only the shadow-backlog lines matter here; the rest of a detect-only local +# bootstrap pass reports this host's toolchain, which is not under test, so it +# runs on the bare base PATH where every tool probe is a fast miss. +bootstrap_backlog_lines() { # <code-root> [<home>] + local code=$1 home=${2:-} + if [ -n "$home" ]; then + PATH="$BASE_PATH" FM_HOME="$home" FM_ROOT_OVERRIDE="$code" FM_BOOTSTRAP_DETECT_ONLY=1 \ + FM_BOOTSTRAP_NETWORK=skip "$BOOTSTRAP" 2>&1 | grep '^BACKLOG_RECONCILE: code-root' || true + else + PATH="$BASE_PATH" FM_ROOT_OVERRIDE="$code" FM_BOOTSTRAP_DETECT_ONLY=1 \ + FM_BOOTSTRAP_NETWORK=skip "$BOOTSTRAP" 2>&1 | grep '^BACKLOG_RECONCILE: code-root' || true + fi +} + +test_guard_reports_regular_code_root_backlog() { + local dir out + dir=$(make_split guard-regular) + out=$(bootstrap_backlog_lines "$dir/code" "$dir/home") + assert_equals "" "$out" "a code-root link into this home must stay silent" + + rm "$dir/code/data/backlog.md" + out=$(bootstrap_backlog_lines "$dir/code" "$dir/home") + assert_equals "" "$out" "an absent code-root backlog must stay silent" + + printf '## In flight\n\n## Queued\n\n- [ ] stray: written from the code root\n\n## Done\n' \ + > "$dir/code/data/backlog.md" + out=$(bootstrap_backlog_lines "$dir/code" "$dir/home") + assert_contains "$out" "BACKLOG_RECONCILE: code-root $dir/code/data/backlog.md is not this home's $dir/home/data/backlog.md" \ + "a regular code-root backlog beside a separate home was not reported" + assert_not_contains "$out" "done-archive.md" "an absent code-root archive was reported" + pass "bootstrap reports a regular code-root backlog and stays silent for a link into the home or no copy" +} + +test_guard_reports_foreign_link_and_archive() { + local dir out + dir=$(make_split guard-foreign) + empty_backlog "$dir/elsewhere.md" + rm "$dir/code/data/backlog.md" + ln -s "$dir/elsewhere.md" "$dir/code/data/backlog.md" + printf '## Done\n' > "$dir/code/data/done-archive.md" + out=$(bootstrap_backlog_lines "$dir/code" "$dir/home") + assert_contains "$out" "code-root $dir/code/data/backlog.md is not this home's" \ + "a code-root backlog linked outside this home was not reported" + assert_contains "$out" "code-root $dir/code/data/done-archive.md is not this home's $dir/home/data/done-archive.md" \ + "a regular code-root archive beside a separate home was not reported" + pass "bootstrap reports a code-root backlog linked elsewhere and a forked archive" +} + +test_guard_silent_for_single_home() { + local dir out + dir="$TMP_ROOT/single-guard" + mkdir -p "$dir/data" + cp "$ROOT/.tasks.toml" "$dir/.tasks.toml" + empty_backlog "$dir/data/backlog.md" + printf '## Done\n' > "$dir/data/done-archive.md" + out=$(bootstrap_backlog_lines "$dir") + assert_equals "" "$out" "the single-home layout's own backlog was reported as a fork" + out=$(bootstrap_backlog_lines "$dir" "$dir") + assert_equals "" "$out" "FM_HOME naming the code root was reported as a fork" + pass "bootstrap stays silent when the code root is the home" +} + +# The end-to-end fork: a bare tasks-axi write from the code root. Whatever the +# installed tasks-axi does to the link, bootstrap must agree with the result: +# a replaced link is reported, a written-through link is not. +test_bare_tasks_axi_fork_is_detected() { + local dir out + dir=$(make_split bare-fork) + (cd "$dir/code" && tasks-axi add bare-1 "written from the code root" >/dev/null 2>&1) \ + || fail "bare tasks-axi add failed in the code root" + out=$(bootstrap_backlog_lines "$dir/code" "$dir/home") + if [ -L "$dir/code/data/backlog.md" ]; then + assert_grep "bare-1" "$dir/home/data/backlog.md" "a written-through link lost the row" + assert_equals "" "$out" "a written-through link was reported as a fork" + pass "bare tasks-axi wrote through the code-root link and bootstrap stayed silent" + else + assert_no_grep "bare-1" "$dir/home/data/backlog.md" "the replaced link still reached the home" + assert_contains "$out" "code-root $dir/code/data/backlog.md is not this home's" \ + "bootstrap missed the fork a bare tasks-axi write left behind" + pass "bare tasks-axi replaced the code-root link and bootstrap reported the fork" + fi +} + +test_wrapper_writes_through_to_home() { + local dir i + dir=$(make_split wrapper-home) + for i in 1 2; do + wrapper_from_code "$dir" add "ship-$i" "ship $i" >/dev/null || fail "add ship-$i failed" + wrapper_from_code "$dir" start "ship-$i" >/dev/null || fail "start ship-$i failed" + wrapper_from_code "$dir" "done" "ship-$i" >/dev/null || fail "done ship-$i failed" + done + wrapper_from_code "$dir" add call-1 "captain call" >/dev/null || fail "add call-1 failed" + wrapper_from_code "$dir" hold call-1 --reason "awaiting the captain" --kind captain >/dev/null \ + || fail "hold call-1 failed" + printf 'RELATIVE-BODY-MARKER\n' > "$dir/code/body.md" + wrapper_from_code "$dir" update call-1 --body-file body.md >/dev/null \ + || fail "update with a caller-relative --body-file failed" + wrapper_from_code "$dir" prune --keep 1 >/dev/null || fail "prune failed" + + [ -L "$dir/code/data/backlog.md" ] || fail "a wrapper write replaced the code-root link" + [ "$dir/code/data/backlog.md" -ef "$dir/home/data/backlog.md" ] \ + || fail "the code-root link no longer names the home's backlog" + assert_grep "call-1" "$dir/home/data/backlog.md" "the held row did not land in the home" + assert_grep "RELATIVE-BODY-MARKER" "$dir/home/data/backlog.md" \ + "a caller-relative --body-file was not read from the caller's directory" + assert_present "$dir/home/data/done-archive.md" "archiving did not reach the home" + assert_grep "ship-1" "$dir/home/data/done-archive.md" "the oldest closed row was not archived in the home" + assert_absent "$dir/code/data/done-archive.md" "archiving wrote a code-root archive" + assert_equals "" "$(bootstrap_backlog_lines "$dir/code" "$dir/home")" \ + "bootstrap reported a fork after only wrapper writes" + pass "fm-tasks-axi.sh writes, holds, archives, and reads relative body files through to the home from the code root" +} + +test_wrapper_overrides_ambient_file() { + local dir + dir=$(make_split wrapper-ambient) + empty_backlog "$dir/decoy.md" + (cd "$dir/code" && TASKS_AXI_FILE="$dir/decoy.md" FM_HOME="$dir/home" "$WRAPPER" add amb-1 "ambient" >/dev/null) \ + || fail "add under an ambient TASKS_AXI_FILE failed" + assert_grep "amb-1" "$dir/home/data/backlog.md" "an ambient TASKS_AXI_FILE diverted the write from the home" + assert_no_grep "amb-1" "$dir/decoy.md" "an ambient TASKS_AXI_FILE received the write" + wrapper_from_code "$dir" >/dev/null || fail "the no-command dashboard failed" + pass "fm-tasks-axi.sh pins the home's backlog over an ambient TASKS_AXI_FILE and serves the dashboard" +} + +test_wrapper_refusals() { + local dir out rc before + dir=$(make_split wrapper-refuse) + before=$(cat "$dir/home/data/backlog.md") + out=$(wrapper_from_code "$dir" add r-1 "explicit" --file "$dir/home/data/backlog.md" 2>&1) + rc=$? + expect_code 2 "$rc" "--file" + assert_contains "$out" "drop --file" "--file refusal did not explain itself" + out=$(wrapper_from_code "$dir" list --file="$dir/home/data/backlog.md" 2>&1) + rc=$? + expect_code 2 "$rc" "--file=" + + mv "$dir/home/data/backlog.md" "$dir/home/real-backlog.md" + ln -s "$dir/home/real-backlog.md" "$dir/home/data/backlog.md" + out=$(wrapper_from_code "$dir" add r-2 "through a link" 2>&1) + rc=$? + expect_code 2 "$rc" "symlinked home backlog" + assert_contains "$out" "is a symlink" "the symlinked home backlog refusal did not name the link" + [ -L "$dir/home/data/backlog.md" ] || fail "a refused call still replaced the home link" + assert_equals "$before" "$(cat "$dir/home/real-backlog.md")" "a refused call changed the backlog" + + out=$(cd "$dir/code" && FM_HOME="$dir/missing-home" "$WRAPPER" list 2>&1) + rc=$? + expect_code 2 "$rc" "missing data directory" + pass "fm-tasks-axi.sh refuses caller --file, a symlinked home backlog, and an unresolvable home" +} + +test_wrapper_single_home() { + local dir + dir="$TMP_ROOT/single-wrapper" + mkdir -p "$dir/data" + cp "$ROOT/.tasks.toml" "$dir/.tasks.toml" + empty_backlog "$dir/data/backlog.md" + (cd "$dir" && FM_ROOT_OVERRIDE="$dir" "$WRAPPER" add solo-1 "single home" >/dev/null) \ + || fail "add in the single-home layout failed" + assert_grep "solo-1" "$dir/data/backlog.md" "the single-home layout lost its own backlog write" + pass "fm-tasks-axi.sh keeps the single-home layout addressing its own code-root backlog" +} + +test_guard_reports_regular_code_root_backlog +test_guard_reports_foreign_link_and_archive +test_guard_silent_for_single_home +if [ "$HAVE_TASKS_AXI" = 1 ]; then + test_bare_tasks_axi_fork_is_detected + test_wrapper_writes_through_to_home + test_wrapper_overrides_ambient_file + test_wrapper_refusals + test_wrapper_single_home +else + echo "skip: tasks-axi not found; home-addressing cases not run" +fi diff --git a/tests/fm-teardown.test.sh b/tests/fm-teardown.test.sh index c0b3b24cf85..43fa7df5543 100755 --- a/tests/fm-teardown.test.sh +++ b/tests/fm-teardown.test.sh @@ -718,7 +718,7 @@ test_teardown_closes_the_backlog_item_itself() { "closed backlog item did not record the task's PR" assert_absent "$case_dir/state/task-x1.backlog-close" \ "a landed close left its pending-close record behind" - printf '%s\n' "$out" | grep -F 'tasks-axi ready' >/dev/null \ + printf '%s\n' "$out" | grep -F 'bin/fm-tasks-axi.sh ready' >/dev/null \ || fail "teardown dropped the dependency-cleared follow-up: $out" printf '%s\n' "$out" | grep -F 'check date gates' >/dev/null \ || fail "teardown did not preserve date-gate check: $out" diff --git a/tests/lib.sh b/tests/lib.sh index c10f87fe9b1..65edb6defd9 100644 --- a/tests/lib.sh +++ b/tests/lib.sh @@ -53,6 +53,16 @@ export FM_GATE_REFUSE_BYPASS=1 # under the marker. A case that verifies the refusal sets FM_TASK_ID itself. unset FM_TASK_ID +# Clear the tasks-axi env overrides. An operator shell exports TASKS_AXI_FILE +# (and may export TASKS_AXI_BACKEND) at its real home's backlog, and tasks-axi +# resolves that env AHEAD of the .tasks.toml a fixture copies, so a suite that +# seeds a temp home with bare `tasks-axi` would silently write the operator's +# live backlog instead - tests/fm-public-followup.test.sh did exactly that. Every +# fixture addresses its own data/backlog.md through its copied .tasks.toml, an +# explicit --file, or bin/fm-tasks-axi.sh; a case that verifies the wrapper +# against an ambient override sets TASKS_AXI_FILE itself. +unset TASKS_AXI_FILE TASKS_AXI_BACKEND + # Resolve the repo root from this library's own location. Consumed by sourcing # test files, not by this library, so it reads as "unused" here. # shellcheck disable=SC2034 From 92856fd4ffeb996a6a2bd068fd0a811599466153 Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Sat, 12 Sep 2026 00:02:20 -0700 Subject: [PATCH 16/31] fix: pre-register Claude trust for secondmate homes (#4262) * fix(spawn): pre-register Claude workspace trust for secondmate homes A claude --secondmate launch skipped workspace-trust registration entirely, so a standalone-clone secondmate home (an explicit ~/fm-homes/<id> path) had no store entry and its pane wedged on the "Is this a project you trust?" dialog before it read its charter. The step was gated on the task kind rather than on the harness, so the spawn's fail-closed guard had nothing to run against and reported a launch that could never start work. fm-claude-trust.sh gains a secondmate-home mode. A secondmate home is a whole firstmate instance, produced either as a leased worktree or as a standalone clone, so the linked-worktree test cannot decide it and the seed is the evidence instead: the .fm-secondmate-home marker must be a regular file this user owns naming exactly the id being spawned, the home must hold AGENTS.md and bin/, and each operational directory must resolve inside the home. That is the set fm-home-seed.sh writes and fm-spawn.sh's own home validation re-checks, so nothing wider than a home a secondmate spawn would launch into can earn home-level trust. The worktree path is unchanged, and still refuses a home. fm-spawn.sh now runs the registration for every claude launch and keeps refusing the spawn when it fails, rather than launching an agent that would wedge. * no-mistakes(document): Correct Claude secondmate trust guidance --- .../references/common/control-and-recovery.md | 6 +- .../references/harness/claude.md | 8 +- bin/fm-claude-trust.sh | 183 +++++++++++---- bin/fm-spawn.sh | 64 +++--- docs/verification/runtime-backends.md | 43 +++- tests/fm-claude-trust.test.sh | 215 +++++++++++++++++- tests/fm-secondmate-harness.test.sh | 24 +- tests/fm-trace-context-spawn.test.sh | 12 +- 8 files changed, 464 insertions(+), 91 deletions(-) diff --git a/.agents/skills/harness-adapters/references/common/control-and-recovery.md b/.agents/skills/harness-adapters/references/common/control-and-recovery.md index c16a78bb8a8..78115e47174 100644 --- a/.agents/skills/harness-adapters/references/common/control-and-recovery.md +++ b/.agents/skills/harness-adapters/references/common/control-and-recovery.md @@ -17,15 +17,11 @@ Select only its documented trust choice from the active Firstmate home, binding No observed dialog proves only that launch. Each supported harness handles its folder-trust gate differently, and the tool reference owns the detail. -Claude gates a fresh worktree and cannot be answered by key, so the spawn pre-registers the path in Claude's own store. +For Claude, load `references/harness/claude.md`; its workspace-trust section owns the non-key-answerable gate and spawn-time pre-registration for every spawn kind. Cursor suppresses its dialog with launch-time `--trust`, and Muse suppresses its own with `--yolo`. Grok dodges its gate instead of granting trust, because its project picker appears only outside a project and the spawn starts in the isolated git root. Pi gates the fresh-worktree case too, but unlike Claude its dialog is answered with Enter, and `references/harness/pi.md` owns that recipe and where the decision persists. Codex shows a directory-trust dialog on the first run for a repository root. -A Claude secondmate is deliberately not pre-registered, because `../../../bin/fm-spawn.sh` runs its per-harness pre-launch setup only for non-secondmate kinds, so the registration is never invoked for one. -That kind guard is the whole exclusion, because a treehouse-leased secondmate home is itself a linked worktree that the scope test would accept, and only a plain-clone home would be refused as a primary checkout. -The consequence is that a claude secondmate whose home Claude has never trusted meets the workspace-trust dialog itself, and firstmate cannot answer it any more than it can for a crewmate. -This is rarely seen because a secondmate home is persistent and reused, so its trust decision is made once and survives, unlike a per-task worktree that is new every time. Use the tool's exact skill form, or natural language only when no separate command is verified or the form remains uncertain. A successful send or key return is not proof of submission; require the tool-specific postcondition. diff --git a/.agents/skills/harness-adapters/references/harness/claude.md b/.agents/skills/harness-adapters/references/harness/claude.md index 24591bc65e9..437eb77c072 100644 --- a/.agents/skills/harness-adapters/references/harness/claude.md +++ b/.agents/skills/harness-adapters/references/harness/claude.md @@ -16,10 +16,10 @@ Busy hooks verified 2026-07-28 on Claude Code 2.1.220. ## Workspace trust -Claude gates a folder it has never seen behind an interactive workspace-trust dialog, so every fresh task worktree would hit it. -`--dangerously-skip-permissions` does not cover that gate: `claude --help` records that the dialog is skipped only in non-interactive mode, through `-p` or a non-TTY stdout, and a crewmate pane is interactive. -A ship or scout spawn therefore pre-registers the worktree before launch, and the dialog does not appear. -`../../../bin/fm-claude-trust.sh` records `hasTrustDialogAccepted` for that worktree path in `${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json`, and `../../../bin/fm-spawn.sh` refuses the spawn when the write fails rather than launching a worker that would wedge. +Claude gates a folder it has never seen behind an interactive workspace-trust dialog, so every fresh task worktree would hit it, and so would every secondmate home no operator has opened by hand. +`--dangerously-skip-permissions` does not cover that gate: `claude --help` records that the dialog is skipped only in non-interactive mode, through `-p` or a non-TTY stdout, and a spawned pane is interactive. +Every claude spawn therefore pre-registers the directory its pane starts in before launch, and the dialog does not appear: the task worktree for a ship or scout, and the home itself for a `--secondmate` spawn, in either seeded shape (a leased worktree or a standalone clone). +`../../../bin/fm-claude-trust.sh` records `hasTrustDialogAccepted` for that path in `${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json` and owns the structural scope test each shape must pass, and `../../../bin/fm-spawn.sh` refuses the spawn when the registration fails rather than launching an agent that would wedge. Never try to answer the trust dialog with a key. Firstmate's key plane carries only Enter, Escape, and C-c with no arrow navigation, so it cannot move a dialog's selection at all, and the observed rendering starts on `No, exit`, which means a sent Enter ends the session instead of accepting. diff --git a/bin/fm-claude-trust.sh b/bin/fm-claude-trust.sh index 732a6b7c5c4..770dd79cfb5 100755 --- a/bin/fm-claude-trust.sh +++ b/bin/fm-claude-trust.sh @@ -1,26 +1,34 @@ #!/usr/bin/env bash -# Pre-register Claude Code's workspace trust for the isolated task worktree a -# ship/scout spawn is about to launch a claude crewmate into, so the worker -# reaches its brief instead of wedging on the trust dialog. +# Pre-register Claude Code's workspace trust for the directory a claude spawn is +# about to launch into - the isolated task worktree of a ship or scout crewmate, +# or the seeded home of a secondmate - so the agent reaches its brief or charter +# instead of wedging on the trust dialog. # # Usage: fm-claude-trust.sh <worktree> <project> +# fm-claude-trust.sh --secondmate-home <home> <id> # <worktree> the isolated task worktree this spawn launches into # <project> the primary checkout that worktree belongs to +# <home> the seeded secondmate home this spawn launches into +# <id> the secondmate id that home must already be marked for # Prints one line naming what it registered; refuses loudly on anything else. # # WHY THIS EXISTS. Claude Code gates a folder it has never seen behind an # interactive workspace-trust dialog, and --dangerously-skip-permissions does # NOT cover it: `claude --help` records that the dialog is skipped only in -# non-interactive mode (-p, or a non-TTY stdout), and a crewmate pane is -# interactive. Every fresh task worktree therefore hits it. The dialog renders +# non-interactive mode (-p, or a non-TTY stdout), and a spawned pane is +# interactive. Every fresh task worktree therefore hits it, and so does every +# secondmate home the operator has not opened by hand. The dialog renders # with the cursor on "No, exit" and firstmate's steering plane carries only # Enter, Escape and C-c with no arrow navigation, so firstmate cannot answer it -# and must not try - pressing Enter would select exit. The worker wedges before +# and must not try - pressing Enter would select exit. The agent wedges before # it ever reads the brief. Registering the trust before launch is the only # control that reaches an interactive pane. # # THE SCOPE TEST IS THE SAFETY PROPERTY, and it is STRUCTURAL rather than a -# path policy. <worktree> must be a LINKED git worktree - its own git dir, +# path policy. Each mode has its own, because the two directories have entirely +# different shapes on disk. +# +# WORKTREE MODE. <worktree> must be a LINKED git worktree - its own git dir, # sharing <project>'s common dir - whose top level is exactly the resolved # argument. Git is the ground truth, so the argument is never trusted on its # own word: a primary checkout (git dir == common dir), a worktree of an @@ -45,12 +53,36 @@ # opt-in guard family (FM_*_LIVE_E2E=1) and record the result in # docs/verification/runtime-backends.md, rather than assuming the shape here. # +# SECONDMATE-HOME MODE. A secondmate home is a whole firstmate instance rather +# than a task worktree, and bin/fm-home-seed.sh produces it in two shapes: a +# leased treehouse worktree (linked) and a standalone clone of the firstmate +# repo (a primary checkout). The worktree test above therefore cannot decide +# this case at all - it refuses the standalone clone as a primary checkout, +# which is why a claude secondmate in an explicit ~/fm-homes/<id> home met the +# dialog with nothing registered. Git shape is not the evidence here; THE SEED +# IS. The home must carry a .fm-secondmate-home marker that is a regular file +# this user owns, never a symlink, naming exactly the <id> passed; it must hold +# the firstmate instance files AGENTS.md and bin/; and each of its data, state, +# config and projects paths must resolve inside the home. That is the set +# bin/fm-home-seed.sh writes and bin/fm-spawn.sh's validate_firstmate_home_for_spawn +# re-checks before launch, so this accepts exactly the homes a secondmate spawn +# will launch into and nothing wider: a plain directory, a project checkout, an +# ordinary firstmate checkout, a home marked for a different secondmate, and a +# home whose operational directory escapes it are each refused. An ABSENT +# operational directory is accepted for the same reason the spawn accepts one - +# a test stricter than the spawn's own would move the wedge from the dialog to +# a refusal without making any unseeded directory less trusted. +# +# Home-level trust is broader than worktree trust, since the pane starts in the +# home and the secondmate works across it, so it is granted on that seed +# evidence alone and never on a caller's word about what a path is. +# # Only the launching user's own store is written: the projects entry for the -# worktree path in ${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json, which must be a +# registered path in ${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json, which must be a # regular file this uid owns. Every unrelated key and project entry is # preserved, and the replacement is atomic. fm-spawn.sh forwards CLAUDE_CONFIG_DIR -# onto the claude launch verbatim rather than resolving it, and the worker's pane -# starts in the task worktree, so only an absolute value names the same store on +# onto the claude launch verbatim rather than resolving it, and the pane starts +# in the registered directory, so only an absolute value names the same store on # both sides; a relative one is refused below rather than guessed at. set -u # Path resolution here must answer from the filesystem, never from the caller's @@ -69,9 +101,36 @@ unset CDPATH \ GIT_DISCOVERY_ACROSS_FILESYSTEM GIT_CONFIG GIT_CONFIG_GLOBAL \ GIT_CONFIG_SYSTEM GIT_CONFIG_NOSYSTEM GIT_CONFIG_COUNT -[ "$#" -eq 2 ] || { echo "usage: fm-claude-trust.sh <worktree> <project>" >&2; exit 2; } -WT_ARG=$1 -PROJ_ARG=$2 +usage() { + echo "usage: fm-claude-trust.sh <worktree> <project>" >&2 + echo " fm-claude-trust.sh --secondmate-home <home> <id>" >&2 + exit 2 +} + +# MODE selects which structural scope test decides the argument, and SCOPE_NOUN +# names what the argument was expected to be so every shared refusal below reads +# correctly in both modes. +case "${1:-}" in + --secondmate-home) + [ "$#" -eq 3 ] || usage + MODE=secondmate-home + TARGET_ARG=$2 + SUB_ID=$3 + PROJ_ARG= + SCOPE_NOUN="secondmate home" + ;; + '' | -h | --help) + usage + ;; + *) + [ "$#" -eq 2 ] || usage + MODE=worktree + TARGET_ARG=$1 + SUB_ID= + PROJ_ARG=$2 + SCOPE_NOUN="task worktree" + ;; +esac refuse() { echo "error: refusing to pre-register Claude trust: $1" >&2; exit 1; } @@ -90,10 +149,12 @@ common_dir_of() { (cd -P -- "$dir" && real_dir "$common") } -WT_REAL=$(real_dir "$WT_ARG") || true -[ -n "$WT_REAL" ] || refuse "worktree '$WT_ARG' is not an accessible directory" -PROJ_REAL=$(real_dir "$PROJ_ARG") || true -[ -n "$PROJ_REAL" ] || refuse "project '$PROJ_ARG' is not an accessible directory" +TARGET_REAL=$(real_dir "$TARGET_ARG") || true +[ -n "$TARGET_REAL" ] || refuse "$SCOPE_NOUN '$TARGET_ARG' is not an accessible directory" +if [ "$MODE" = worktree ]; then + PROJ_REAL=$(real_dir "$PROJ_ARG") || true + [ -n "$PROJ_REAL" ] || refuse "project '$PROJ_ARG' is not an accessible directory" +fi CONFIG_DIR=${CLAUDE_CONFIG_DIR:-${HOME:-}} [ -n "$CONFIG_DIR" ] || refuse "neither CLAUDE_CONFIG_DIR nor HOME is set, so the store cannot be located" @@ -116,30 +177,66 @@ if [ -z "$CONFIG_DIR_REAL" ]; then fi [ -n "$CONFIG_DIR_REAL" ] || refuse "Claude config directory '$CONFIG_DIR' does not exist and could not be created" -# A home or config directory is never a task worktree. Checked explicitly so -# the refusal names the real reason instead of the git verdict behind it. -[ "$WT_REAL" != "$CONFIG_DIR_REAL" ] || refuse "'$WT_REAL' is the Claude config directory, not a task worktree" +# The filesystem root, a home directory, and the config directory are never +# something this registers, in either mode. Checked explicitly so the refusal +# names the real reason instead of the scope verdict behind it. +[ "$TARGET_REAL" != / ] || refuse "'/' is the filesystem root, not a $SCOPE_NOUN" +[ "$TARGET_REAL" != "$CONFIG_DIR_REAL" ] || refuse "'$TARGET_REAL' is the Claude config directory, not a $SCOPE_NOUN" if [ -n "${HOME:-}" ]; then HOME_REAL=$(real_dir "$HOME") || true - [ "$WT_REAL" != "${HOME_REAL:-}" ] || refuse "'$WT_REAL' is the home directory, not a task worktree" + [ "$TARGET_REAL" != "${HOME_REAL:-}" ] || refuse "'$TARGET_REAL' is the home directory, not a $SCOPE_NOUN" fi -WT_TOP=$(git -C "$WT_REAL" rev-parse --show-toplevel 2>/dev/null) || true -[ -n "$WT_TOP" ] || refuse "'$WT_REAL' is not inside a git repository" -WT_TOP_REAL=$(real_dir "$WT_TOP") || true -[ "$WT_TOP_REAL" = "$WT_REAL" ] || refuse "'$WT_REAL' is not a worktree root (its root is '${WT_TOP_REAL:-unresolvable}')" +if [ "$MODE" = worktree ]; then + WT_TOP=$(git -C "$TARGET_REAL" rev-parse --show-toplevel 2>/dev/null) || true + [ -n "$WT_TOP" ] || refuse "'$TARGET_REAL' is not inside a git repository" + WT_TOP_REAL=$(real_dir "$WT_TOP") || true + [ "$WT_TOP_REAL" = "$TARGET_REAL" ] || refuse "'$TARGET_REAL' is not a worktree root (its root is '${WT_TOP_REAL:-unresolvable}')" -WT_GIT_DIR=$(git -C "$WT_REAL" rev-parse --absolute-git-dir 2>/dev/null) || true -[ -n "$WT_GIT_DIR" ] || refuse "'$WT_REAL' has no resolvable git directory" -WT_GIT_DIR=$(real_dir "$WT_GIT_DIR") || true -[ -n "$WT_GIT_DIR" ] || refuse "'$WT_REAL' has an unresolvable git directory" -WT_COMMON=$(common_dir_of "$WT_REAL") || true -[ -n "$WT_COMMON" ] || refuse "'$WT_REAL' has no resolvable git common directory" -[ "$WT_GIT_DIR" != "$WT_COMMON" ] || refuse "'$WT_REAL' is a primary checkout, not an isolated worktree" + WT_GIT_DIR=$(git -C "$TARGET_REAL" rev-parse --absolute-git-dir 2>/dev/null) || true + [ -n "$WT_GIT_DIR" ] || refuse "'$TARGET_REAL' has no resolvable git directory" + WT_GIT_DIR=$(real_dir "$WT_GIT_DIR") || true + [ -n "$WT_GIT_DIR" ] || refuse "'$TARGET_REAL' has an unresolvable git directory" + WT_COMMON=$(common_dir_of "$TARGET_REAL") || true + [ -n "$WT_COMMON" ] || refuse "'$TARGET_REAL' has no resolvable git common directory" + [ "$WT_GIT_DIR" != "$WT_COMMON" ] || refuse "'$TARGET_REAL' is a primary checkout, not an isolated worktree" -PROJ_COMMON=$(common_dir_of "$PROJ_REAL") || true -[ -n "$PROJ_COMMON" ] || refuse "project '$PROJ_REAL' is not inside a git repository" -[ "$WT_COMMON" = "$PROJ_COMMON" ] || refuse "'$WT_REAL' is not a worktree of project '$PROJ_REAL'" + PROJ_COMMON=$(common_dir_of "$PROJ_REAL") || true + [ -n "$PROJ_COMMON" ] || refuse "project '$PROJ_REAL' is not inside a git repository" + [ "$WT_COMMON" = "$PROJ_COMMON" ] || refuse "'$TARGET_REAL' is not a worktree of project '$PROJ_REAL'" +else + # The seed evidence, in the order that names the most useful reason first: the + # marker decides whether this is a secondmate home at all, the id decides + # whose, and the instance files and operational directories decide whether it + # is the shape bin/fm-home-seed.sh leaves behind. The marker is the token the + # whole boundary rests on, so it is judged as a file rather than as a value: a + # symlink is refused outright rather than followed, because a link is a way to + # make some other file's bytes stand in for the seed, and a marker this user + # does not own was planted by someone else. + [ -n "$SUB_ID" ] || refuse "no secondmate id was supplied, so '$TARGET_REAL' cannot be matched against its seed marker" + SUB_MARKER="$TARGET_REAL/.fm-secondmate-home" + [ ! -L "$SUB_MARKER" ] || refuse "'$SUB_MARKER' is a symlink; a seeded secondmate home carries the marker as a regular file" + [ -f "$SUB_MARKER" ] || refuse "'$TARGET_REAL' carries no .fm-secondmate-home marker, so it is not a seeded secondmate home" + [ -O "$SUB_MARKER" ] || refuse "'$SUB_MARKER' is not owned by this user" + SUB_MARKER_ID=$(cat "$SUB_MARKER" 2>/dev/null) || true + [ "$SUB_MARKER_ID" = "$SUB_ID" ] || refuse "'$TARGET_REAL' is marked for secondmate '${SUB_MARKER_ID:-unknown}', not '$SUB_ID'" + [ -f "$TARGET_REAL/AGENTS.md" ] || refuse "'$TARGET_REAL' has no AGENTS.md, so it is not a firstmate home" + [ -d "$TARGET_REAL/bin" ] || refuse "'$TARGET_REAL' has no bin/, so it is not a firstmate home" + for sub_dir_name in data state config projects; do + sub_dir="$TARGET_REAL/$sub_dir_name" + if [ -L "$sub_dir" ] && [ ! -e "$sub_dir" ]; then + refuse "'$sub_dir' is a broken symlink, so this home's $sub_dir_name directory cannot be shown to stay inside it" + fi + [ -e "$sub_dir" ] || continue + [ -d "$sub_dir" ] || refuse "'$sub_dir' is not a directory, so '$TARGET_REAL' is not a seeded secondmate home" + sub_dir_real=$(real_dir "$sub_dir") || true + [ -n "$sub_dir_real" ] || refuse "'$sub_dir' cannot be resolved" + case "$sub_dir_real" in + "$TARGET_REAL"/*) ;; + *) refuse "'$sub_dir' resolves to '$sub_dir_real', outside the home, so '$TARGET_REAL' is not a safe secondmate home" ;; + esac + done +fi # The store write needs node, and a missing interpreter refuses like every other # failure here. Degrading instead would launch a worker straight into the dialog @@ -190,11 +287,11 @@ fi # attempts, and it must fail loudly rather than report a trust it did not leave. # ponytail: fingerprint-and-refuse, not a lock; flock is absent on macOS and # cannot stop a vendor session's own rewrite anyway. -if ! node - "$STORE" "$WT_REAL" <<'NODE' +if ! node - "$STORE" "$TARGET_REAL" <<'NODE' const fs = require("node:fs"); const path = require("node:path"); const crypto = require("node:crypto"); -const [store, worktree] = process.argv.slice(2); +const [store, target] = process.argv.slice(2); const readStore = () => { try { return fs.readFileSync(store); @@ -223,12 +320,12 @@ const attempt = () => { if (projects === null || typeof projects !== "object" || Array.isArray(projects)) { throw new Error(`${store} has a non-object "projects" value`); } - let entry = projects[worktree]; + let entry = projects[target]; if (entry === undefined || entry === null || typeof entry !== "object" || Array.isArray(entry)) { entry = {}; } entry.hasTrustDialogAccepted = true; - projects[worktree] = entry; + projects[target] = entry; // Unpredictable name plus an exclusive create: the config directory may be // writable by another local account, and a predictable path could be // pre-created there as a symlink that a plain write would follow into some @@ -250,7 +347,7 @@ const attempt = () => { if (!renamed) fs.rmSync(tmp, { force: true }); } const back = JSON.parse(fs.readFileSync(store, "utf8")); - return back.projects?.[worktree]?.hasTrustDialogAccepted === true ? "recorded" : "dropped"; + return back.projects?.[target]?.hasTrustDialogAccepted === true ? "recorded" : "dropped"; }; try { for (let i = 0; i < 3; i += 1) { @@ -265,11 +362,11 @@ try { console.error(`error: ${err.message}`); process.exit(1); } -console.error(`error: ${store} did not retain trust for ${worktree} after 3 attempts`); +console.error(`error: ${store} did not retain trust for ${target} after 3 attempts`); process.exit(1); NODE then - refuse "could not record trust for '$WT_REAL' in '$STORE'" + refuse "could not record trust for '$TARGET_REAL' in '$STORE'" fi -echo "trusted: $WT_REAL" +echo "trusted: $TARGET_REAL" diff --git a/bin/fm-spawn.sh b/bin/fm-spawn.sh index df1181dd561..13aaad7df50 100755 --- a/bin/fm-spawn.sh +++ b/bin/fm-spawn.sh @@ -305,13 +305,13 @@ # park owns that home's supervision (docs/supervision-protocols/cursor.md). # claude is the one harness whose pre-launch setup can REFUSE the spawn: before # any per-task state exists, and before its worktree .claude/settings.local.json -# hooks are written, a non-secondmate claude launch pre-registers the worktree in -# the launching user's own Claude trust store through bin/fm-claude-trust.sh, -# because Claude's interactive workspace-trust dialog gates a fresh worktree and -# firstmate cannot answer it. That helper's header owns the structural scope test -# and every refusal; a failed registration stops this spawn rather than launching -# a worker that would wedge on the dialog. A --secondmate launch never runs it, -# so a claude secondmate home keeps its own one-time trust decision. +# hooks are written, every claude launch pre-registers the directory the pane +# starts in - the task worktree, or the secondmate home for a --secondmate spawn - +# in the launching user's own Claude trust store through bin/fm-claude-trust.sh, +# because Claude's interactive workspace-trust dialog gates a folder it has never +# seen and firstmate cannot answer it. That helper's header owns the structural +# scope test for both shapes and every refusal; a failed registration stops this +# spawn rather than launching a worker that would wedge on the dialog. # Every claude launch also carries the attribution-off policy in its per-launch # --settings JSON, so a spawned worker never writes a Co-Authored-By trailer, # Claude-Session link, or generated-with line into a commit or PR body; @@ -3217,27 +3217,35 @@ if [ "$RELAUNCH" -eq 0 ] && [ "$KIND" != secondmate ]; then freshen_spawn_worktree_base "$WT" || exit 1 fi -# Pre-register Claude's workspace trust for the worktree, at the first point the -# worktree is known and before any per-task state is created below. The dialog -# gates the pane before the brief is ever read, and it also gates loading the -# project settings written further down, so nothing armed below takes effect -# without it. bin/fm-claude-trust.sh owns the structural scope test and refuses -# any path that is not this project's own isolated worktree; a refusal blocks the -# spawn rather than launching a worker that would wedge on a dialog firstmate -# cannot answer. Refusing here rather than beside the arm keeps this in the same -# class as the two worktree refusals just above: no temp root, no retired -# relaunch wiring and no busy record exists yet to strand, so the refusal names -# the endpoint the same way they do and leaves nothing else behind. -if [ "$KIND" != secondmate ]; then - case "$HARNESS" in - claude*) - if ! "$FM_ROOT/bin/fm-claude-trust.sh" "$WT" "$PROJ_ABS" >/dev/null; then - echo "error: could not pre-register Claude workspace trust for $WT; refusing to launch a claude worker that would wedge on the trust dialog; inspect window $T" >&2 - exit 1 - fi - ;; - esac -fi +# Pre-register Claude's workspace trust for the directory this launch starts in, +# at the first point that directory is known and before any per-task state is +# created below. The dialog gates the pane before the brief is ever read, and it +# also gates loading the project settings written further down, so nothing armed +# below takes effect without it. EVERY claude launch needs it, a secondmate's +# included: its home is just as unseen by Claude as a fresh worktree, and +# skipping the step for that kind left a standalone-clone secondmate home with +# nothing registered and a pane wedged on a dialog firstmate cannot answer. +# bin/fm-claude-trust.sh owns the structural scope test for both shapes and +# refuses anything that is neither this project's own isolated worktree nor a +# seeded secondmate home marked for this id; a refusal blocks the spawn rather +# than launching a worker that would wedge. Refusing here rather than beside the +# arm keeps this in the same class as the two worktree refusals just above: no +# temp root, no retired relaunch wiring and no busy record exists yet to strand, +# so the refusal names the endpoint the same way they do and leaves nothing else +# behind. +case "$HARNESS" in + claude*) + if [ "$KIND" = secondmate ]; then + spawn_trust_args=(--secondmate-home "$PROJ_ABS" "$ID") + else + spawn_trust_args=("$WT" "$PROJ_ABS") + fi + if ! "$FM_ROOT/bin/fm-claude-trust.sh" "${spawn_trust_args[@]}" >/dev/null; then + echo "error: could not pre-register Claude workspace trust for $WT; refusing to launch a claude worker that would wedge on the trust dialog; inspect window $T" >&2 + exit 1 + fi + ;; +esac # Per-task temp root: /tmp/fm-<id>/ with Go's build temp nested at gotmp/. Go won't # create GOTMPDIR, so mkdir before it is used; fm-teardown removes the whole root. diff --git a/docs/verification/runtime-backends.md b/docs/verification/runtime-backends.md index 1ec8b925be0..80e032d706a 100644 --- a/docs/verification/runtime-backends.md +++ b/docs/verification/runtime-backends.md @@ -300,7 +300,48 @@ That warning rendered in the same shape as the trust dialog, with the selection That gate is not a production blocker, because a normal environment has already accepted it and the treatment arm above ran against the real config and saw neither dialog. This change does not address that warning and does not claim to. -`bin/fm-spawn.sh` therefore pre-registers the task worktree through `bin/fm-claude-trust.sh` before launch, and `tests/fm-claude-trust.test.sh` pins both halves of the scope contract: a fresh worktree is trusted, and an out-of-scope path is refused. +### Secondmate homes + +Verified 2026-09-11 on Claude Code 2.1.269. +A secondmate launches in its own firstmate home rather than a task worktree, and that home meets the same gate. +The control arm launched a standalone-clone secondmate home that the store had no entry for, the way `bin/fm-spawn.sh --secondmate` launches one. + +```sh +tmux -L <sock> new-session -d -s ctrl -x 180 -y 44 -c <home> \ + "CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false claude --dangerously-skip-permissions" +``` + +``` + Accessing workspace: + /private/tmp/fm-sm-trust-live-69759/fm-homes/livemate-n1 + Quick safety check: Is this a project you created or one you trust? ... + ❯ No, exit + Yes, I trust this folder +``` + +The treatment arm pre-registered that same home through the secondmate-home mode and launched it identically against the operator's real config. + +```sh +bin/fm-claude-trust.sh --secondmate-home <home> livemate-n1 +``` + +``` +trusted: /private/tmp/fm-sm-trust-live-69759/fm-homes/livemate-n1 +``` + +``` + ▐▛███▛█ Claude Code v2.1.269 +▝▜██████▀ Opus 4.8 with high effort · Claude Max + ▝▝ ▝▝ /private/tmp/fm-sm-trust-live-69759/fm-homes/livemate-n1 +... +❯ + ⏵⏵ bypass permissions on (shift+tab to cycle) · ← for agents +``` + +No dialog appeared, the composer was reached, and neither did the machine-scoped bypass warning, because this ran against the real config. +The lab home was deleted and the test entry was removed from the store and verified absent, with the same point-in-time caveat as the worktree arms above. + +`bin/fm-spawn.sh` therefore pre-registers the directory every claude launch starts in through `bin/fm-claude-trust.sh` before launch, and `tests/fm-claude-trust.test.sh` pins both halves of the scope contract for both shapes: a fresh worktree and a seeded secondmate home are trusted, and an out-of-scope path is refused. That automated spawn case runs against a fake claude, so it asserts the store entry and the launch command and nothing more; the live arms above are what establish that the entry actually suppresses the dialog. The composer-classification record below observes the same gate from the other side, where an untrusted worktree left Claude, Grok, and Muse unverified because the guard reads a first-launch trust dialog as an unreadable composer. diff --git a/tests/fm-claude-trust.test.sh b/tests/fm-claude-trust.test.sh index 94211e0e4b4..893840c1901 100755 --- a/tests/fm-claude-trust.test.sh +++ b/tests/fm-claude-trust.test.sh @@ -1,9 +1,10 @@ #!/usr/bin/env bash # Behavior tests for bin/fm-claude-trust.sh and the claude spawn that calls it. # -# Both halves of the contract are load-bearing and both are proven here: a -# legitimate fresh task worktree is trusted so a claude worker reaches its -# brief with no human, and every out-of-scope path is REFUSED rather than +# Both halves of the contract are load-bearing and both are proven here, for +# each directory a claude launch can start in: a legitimate fresh task worktree +# and a seeded secondmate home are trusted so the agent reaches its brief or +# charter with no human, and every out-of-scope path is REFUSED rather than # warned about or quietly skipped. set -u @@ -78,6 +79,53 @@ node_free_path() { # <case-dir> -> a bin dir holding the script's own tools but printf '%s\n' "$dir" } +# --- secondmate homes ------------------------------------------------------- + +# seed_secondmate_home <home> <id> [shape]: the on-disk shape bin/fm-home-seed.sh +# leaves behind - the identity marker, the firstmate instance files, the four +# operational directories, and a charter for the launch to carry. "clone" (the +# default) is the standalone-clone home an explicit ~/fm-homes/<id> path +# produces, a primary checkout of the firstmate repo; "worktree" is the linked +# worktree a treehouse lease produces. Both shapes are real homes, so both must +# be trusted. +seed_secondmate_home() { + local home=$1 id=$2 shape=${3:-clone} src + case "$shape" in + worktree) + src="$home.src" + fm_git_worktree "$src" "$home" "sm-$id" + ;; + *) + mkdir -p "$home" + fm_git_init_commit "$home" + ;; + esac + mkdir -p "$home/bin" "$home/data" "$home/state" "$home/config" "$home/projects" + printf '# Firstmate\n' > "$home/AGENTS.md" + printf 'charter\n' > "$home/data/charter.md" + printf '%s\n' "$id" > "$home/.fm-secondmate-home" +} + +# run_home_trust <config> <home> <id> [user-home]: invoke the secondmate-home +# mode against an isolated store. +run_home_trust() { + local config=$1 home=$2 id=$3 user_home=${4:-$1} + CLAUDE_CONFIG_DIR="$config" HOME="$user_home" "$TRUST" --secondmate-home "$home" "$id" 2>&1 +} + +# spawn_secondmate_claude <case-dir> <home> <id>: run a real --secondmate claude +# spawn against the isolated store at <case-dir>/claude-config, logging the +# launch to <case-dir>/launch.log. Echoes the spawn output. +spawn_secondmate_claude() { + local case_dir=$1 home=$2 id=$3 primary fakebin + primary="$case_dir/primary" + mkdir -p "$case_dir/claude-config" + fakebin=$(make_spawn_fakebin "$case_dir/fake" claude) + fm_test_spawn_home "$primary" claude + FM_TEST_CLAUDE_CONFIG_DIR="$case_dir/claude-config" FM_FAKE_LAUNCH_LOG="$case_dir/launch.log" \ + fm_test_run_spawn "$primary" "$home" "$fakebin" "$id" "$home" claude --secondmate +} + test_fresh_worktree_is_trusted() { local rec out rec=$(make_case fresh) @@ -440,6 +488,162 @@ test_claude_spawn_pretrusts_its_worktree_and_reaches_the_brief() { pass "fm-spawn.sh: a claude spawn pre-trusts its worktree and launches with the brief" } +# A secondmate home is the second directory a claude launch starts in, and it is +# as unseen by Claude as a fresh worktree. The standalone-clone shape is the one +# that wedged in production: the trust step was skipped for every secondmate, so +# nothing was registered and the pane stopped on the dialog before it read its +# charter. +test_secondmate_standalone_clone_home_is_trusted() { + local case_dir home out + case_dir="$TMP_ROOT/sm-clone-spawn" + home="$case_dir/fm-homes/nomistakes-n1" + seed_secondmate_home "$home" nomistakes-n1 clone + out=$(spawn_secondmate_claude "$case_dir" "$home" nomistakes-n1) + expect_code 0 $? "a claude secondmate spawn into a standalone-clone home must succeed: $out" + assert_trusted "$case_dir/claude-config/.claude.json" "$home" \ + "the claude secondmate spawn did not pre-register trust for its standalone-clone home" + assert_present "$case_dir/launch.log" "the claude secondmate spawn sent no launch command" + assert_grep 'claude --dangerously-skip-permissions' "$case_dir/launch.log" \ + "the launch command was not the claude secondmate launch" + assert_grep "$home/data/charter.md" "$case_dir/launch.log" \ + "the launch command did not carry the charter the secondmate must read" + # The pane must read the SAME store the registration wrote, or the trust would + # land somewhere it never looks and the dialog would appear anyway. + assert_grep "CLAUDE_CONFIG_DIR='$case_dir/claude-config'" "$case_dir/launch.log" \ + "the launch command did not point the secondmate at the store that was trusted" + pass "fm-spawn.sh: a claude secondmate spawn pre-trusts a standalone-clone home" +} + +# The other seeded shape, a treehouse-leased linked worktree. It must be trusted +# through the same seed evidence rather than incidentally, so the registration +# does not depend on which shape the home happens to have. +test_secondmate_leased_worktree_home_is_trusted() { + local case_dir home out + case_dir="$TMP_ROOT/sm-leased-spawn" + home="$case_dir/leased/home" + mkdir -p "$case_dir/leased" + seed_secondmate_home "$home" leased-n1 worktree + out=$(spawn_secondmate_claude "$case_dir" "$home" leased-n1) + expect_code 0 $? "a claude secondmate spawn into a leased worktree home must succeed: $out" + assert_trusted "$case_dir/claude-config/.claude.json" "$home" \ + "the claude secondmate spawn did not pre-register trust for its leased worktree home" + pass "fm-spawn.sh: a claude secondmate spawn pre-trusts a leased worktree home" +} + +# The seed is the whole security boundary for home-level trust, so every path +# that is not a home seeded for THIS secondmate is refused and left untrusted. +# Each row drives one structural property apart from a genuine home. +test_secondmate_home_trust_refuses_everything_unseeded() { + local case_dir config home target out + case_dir="$TMP_ROOT/sm-refusals" + config="$case_dir/claude-config" + mkdir -p "$config" + + # A plain directory: no marker at all. + target="$case_dir/plain" + mkdir -p "$target" + out=$(run_home_trust "$config" "$target" plain-n1) + expect_code 1 $? "a plain directory must be refused: $out" + assert_contains "$out" "no .fm-secondmate-home marker" "the refusal did not name the missing marker" + assert_not_trusted "$config/.claude.json" "$target" "a plain directory was trusted" + + # A firstmate checkout that was never seeded as a secondmate home: every other + # structural signal matches and only the marker is missing. + target="$case_dir/checkout" + seed_secondmate_home "$target" checkout-n1 clone + rm -f "$target/.fm-secondmate-home" + out=$(run_home_trust "$config" "$target" checkout-n1) + expect_code 1 $? "an unseeded firstmate checkout must be refused: $out" + assert_contains "$out" "no .fm-secondmate-home marker" "the refusal did not name the missing marker" + assert_not_trusted "$config/.claude.json" "$target" "an unseeded firstmate checkout was trusted" + + # A home seeded for a DIFFERENT secondmate: one home's trust must not be + # granted while spawning another id. + target="$case_dir/other-mate" + seed_secondmate_home "$target" other-n1 clone + out=$(run_home_trust "$config" "$target" wanted-n1) + expect_code 1 $? "a home marked for another secondmate must be refused: $out" + assert_contains "$out" "other-n1" "the refusal did not name the id the home is marked for" + assert_not_trusted "$config/.claude.json" "$target" "a home marked for another secondmate was trusted" + + # A marker that is a symlink: another file's bytes must not stand in for the + # seed, even when they read as the right id. + target="$case_dir/linked-marker" + seed_secondmate_home "$target" linked-n1 clone + printf 'linked-n1\n' > "$case_dir/planted-id" + ln -sf "$case_dir/planted-id" "$target/.fm-secondmate-home" + out=$(run_home_trust "$config" "$target" linked-n1) + expect_code 1 $? "a symlinked marker must be refused: $out" + assert_contains "$out" "symlink" "the refusal did not name the symlinked marker" + assert_not_trusted "$config/.claude.json" "$target" "a home whose marker is a symlink was trusted" + + # An operational directory that escapes the home: the home's own working + # surface must stay inside it. + target="$case_dir/escaping" + seed_secondmate_home "$target" escaping-n1 clone + rm -rf "$target/projects" + mkdir -p "$case_dir/elsewhere" + ln -s "$case_dir/elsewhere" "$target/projects" + out=$(run_home_trust "$config" "$target" escaping-n1) + expect_code 1 $? "a home whose operational directory escapes it must be refused: $out" + assert_contains "$out" "outside the home" "the refusal did not name the escaping directory" + assert_not_trusted "$config/.claude.json" "$target" "a home whose projects/ escapes it was trusted" + + # The user's own home directory, seeded to prove the marker alone cannot carry + # it: HOME is refused in this mode exactly as it is for a worktree. + target="$case_dir/user-home" + seed_secondmate_home "$target" userhome-n1 clone + out=$(run_home_trust "$config" "$target" userhome-n1 "$target") + expect_code 1 $? "the user's home directory must be refused: $out" + assert_contains "$out" "home directory" "the refusal did not name the home directory" + assert_not_trusted "$config/.claude.json" "$target" "the user's home directory was trusted" + # Prove the seed really would have been accepted, so the guard above is what + # refused rather than an unrelated failure. + out=$(run_home_trust "$config" "$target" userhome-n1 "$case_dir/elsewhere-home") + expect_code 0 $? "the same seeded home must be accepted once it is not HOME: $out" + + pass "fm-claude-trust.sh: home-level trust is refused for everything but a home seeded for this secondmate" +} + +# A secondmate home is not a linked worktree, so worktree mode must keep +# refusing it rather than quietly widening to cover the new case. +test_worktree_mode_still_refuses_a_secondmate_home() { + local case_dir config home out + case_dir="$TMP_ROOT/sm-wrong-mode" + config="$case_dir/claude-config" + home="$case_dir/home" + mkdir -p "$config" + seed_secondmate_home "$home" mode-n1 clone + out=$(run_trust "$config" "$home" "$home") + expect_code 1 $? "worktree mode must still refuse a standalone-clone home: $out" + assert_contains "$out" "primary checkout" "the refusal did not name the primary checkout" + assert_not_trusted "$config/.claude.json" "$home" "worktree mode trusted a standalone-clone home" + pass "fm-claude-trust.sh: worktree mode still refuses a secondmate home" +} + +# The fail-closed half for secondmates: when the home's trust genuinely cannot be +# recorded, the spawn must refuse rather than launch a pane that would wedge on +# the dialog. This is the guard that never fired while the step was skipped. +test_secondmate_spawn_fails_closed_when_home_trust_cannot_be_recorded() { + local case_dir home out + case_dir="$TMP_ROOT/sm-failclosed" + home="$case_dir/fm-homes/failclosed-n1" + # Root owns /etc/passwd, so a store resolving to it is refused as another + # user's file. Running as root would own it and make the refusal vacuous. + if [ "$(id -u)" = 0 ]; then + pass "fm-spawn.sh: a claude secondmate spawn refuses when home trust cannot be recorded (skipped as root)" + return 0 + fi + seed_secondmate_home "$home" failclosed-n1 clone + mkdir -p "$case_dir/claude-config" + ln -s /etc/passwd "$case_dir/claude-config/.claude.json" + out=$(spawn_secondmate_claude "$case_dir" "$home" failclosed-n1) + expect_code 1 $? "a secondmate spawn whose trust registration is refused must fail: $out" + assert_contains "$out" "workspace trust" "the spawn did not report the trust refusal" + assert_absent "$case_dir/launch.log" "a secondmate was launched into a home whose trust could not be recorded" + pass "fm-spawn.sh: a claude secondmate spawn refuses when home trust cannot be recorded" +} + test_fresh_worktree_is_trusted test_registration_is_idempotent test_primary_checkout_is_refused @@ -460,3 +664,8 @@ test_missing_node_is_refused test_scope_refusal_stays_fail_closed_without_node test_claude_spawn_pretrusts_its_worktree_and_reaches_the_brief test_refused_spawn_leaves_no_task_state +test_secondmate_standalone_clone_home_is_trusted +test_secondmate_leased_worktree_home_is_trusted +test_secondmate_home_trust_refuses_everything_unseeded +test_worktree_mode_still_refuses_a_secondmate_home +test_secondmate_spawn_fails_closed_when_home_trust_cannot_be_recorded diff --git a/tests/fm-secondmate-harness.test.sh b/tests/fm-secondmate-harness.test.sh index 83668e21593..9a6ebe2df1f 100755 --- a/tests/fm-secondmate-harness.test.sh +++ b/tests/fm-secondmate-harness.test.sh @@ -64,6 +64,12 @@ fm_git_identity fmtest fmtest@example.com TMP_ROOT=$(fm_test_tmproot fm-secondmate-harness) export FM_BACKEND=tmux +# Every claude launch pre-registers workspace trust for the directory it starts +# in, and for a secondmate that directory is the home (bin/fm-claude-trust.sh). +# Several cases here resolve claude, so every spawn below pins a throwaway HOME +# with an empty CLAUDE_CONFIG_DIR and puts node on the spawn's PATH; without the +# first, this suite would write the developer's real ~/.claude.json. + # =========================================================================== # A) fm-harness.sh secondmate resolution + fallback (deterministic detect_own) # =========================================================================== @@ -426,6 +432,10 @@ make_noop_tmux() { exit 0 SH chmod +x "$fakebin/tmux" + # BASE_PATH deliberately omits the developer's node, which the trust + # registration below needs, so link the real one in rather than presenting a + # node-less spawn host no real fleet member looks like. + ln -sf "$(command -v node)" "$fakebin/node" printf '%s\n' "$fakebin" } @@ -455,7 +465,7 @@ spawn_secondmate() { [ -n "$harness" ] && spawn_args+=("$harness") spawn_args+=(--secondmate) PATH="$fakebin:$BASE_PATH" TMUX='' CLAUDECODE=1 \ - FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$world/home" \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$world/home" HOME="$world/home/user-home" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$world/home/state" FM_DATA_OVERRIDE="$world/home/data" \ FM_PROJECTS_OVERRIDE="$world/home/projects" FM_CONFIG_OVERRIDE="$world/home/config" \ FM_SPAWN_NO_GUARD=1 \ @@ -568,7 +578,7 @@ test_spawn_unverified_secondmate_harness_refused() { err="$w/spawn.err" rc=0 PATH="$fakebin:$BASE_PATH" TMUX='' CLAUDECODE=1 \ - FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$w/home" \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$w/home" HOME="$w/home/user-home" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$w/home/state" FM_DATA_OVERRIDE="$w/home/data" \ FM_PROJECTS_OVERRIDE="$w/home/projects" FM_CONFIG_OVERRIDE="$w/home/config" \ FM_SPAWN_NO_GUARD=1 \ @@ -595,7 +605,7 @@ test_spawn_cursor_secondmate_launches_with_its_primary_contract() { : > "$launchlog" rc=0 PATH="$fakebin:$BASE_PATH" TMUX='' CLAUDECODE=1 \ - FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$w/home" \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$w/home" HOME="$w/home/user-home" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$w/home/state" FM_DATA_OVERRIDE="$w/home/data" \ FM_PROJECTS_OVERRIDE="$w/home/projects" FM_CONFIG_OVERRIDE="$w/home/config" \ FM_SPAWN_NO_GUARD=1 FM_FAKE_LAUNCH_LOG="$launchlog" FM_FAKE_PANE_PATH="$sm" \ @@ -661,6 +671,10 @@ exit 0 SH chmod +x "$fakebin/tmux" fm_fake_exit0 "$fakebin" pi + # BASE_PATH deliberately omits the developer's node, which the trust + # registration below needs, so link the real one in rather than presenting a + # node-less spawn host no real fleet member looks like. + ln -sf "$(command -v node)" "$fakebin/node" printf '%s\n' "$fakebin" } @@ -674,7 +688,7 @@ spawn_secondmate_capture() { fakebin=$(make_launch_capturing_tmux "$world/tmux-$id") : > "$launchlog" PATH="$fakebin:$BASE_PATH" TMUX='' CLAUDECODE=1 \ - FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$world/home" \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$world/home" HOME="$world/home/user-home" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$world/home/state" FM_DATA_OVERRIDE="$world/home/data" \ FM_PROJECTS_OVERRIDE="$world/home/projects" FM_CONFIG_OVERRIDE="$world/home/config" \ FM_SPAWN_NO_GUARD=1 FM_FAKE_LAUNCH_LOG="$launchlog" \ @@ -2572,7 +2586,7 @@ SH chmod +x "$fakebin/rm" launchlog="$w/spawn-quarantine.launch.log" out=$(PATH="$fakebin:$BASE_PATH" TMUX='' CLAUDECODE=1 \ - FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$w/home" \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$w/home" HOME="$w/home/user-home" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$w/home/state" FM_DATA_OVERRIDE="$w/home/data" \ FM_PROJECTS_OVERRIDE="$w/home/projects" FM_CONFIG_OVERRIDE="$w/home/config" \ FM_SPAWN_NO_GUARD=1 FM_FAKE_LAUNCH_LOG="$launchlog" \ diff --git a/tests/fm-trace-context-spawn.test.sh b/tests/fm-trace-context-spawn.test.sh index d38fbf787d1..b9a61736246 100755 --- a/tests/fm-trace-context-spawn.test.sh +++ b/tests/fm-trace-context-spawn.test.sh @@ -214,8 +214,12 @@ run_two_level() { smlog="$base/sm-launch.log" smfake=$(make_spawn_fakebin "$base/sm-fake") : > "$smlog" + # A claude secondmate spawn pre-registers workspace trust for the HOME it + # launches into (bin/fm-claude-trust.sh), so this runs against a throwaway + # HOME; without it this suite would write the developer's real ~/.claude.json. + mkdir -p "$base/user-home" env FM_TRACE_CONTEXT="$penv" \ - FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$prim" \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$prim" HOME="$base/user-home" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$prim/state" FM_DATA_OVERRIDE="$prim/data" \ FM_PROJECTS_OVERRIDE="$prim/projects" FM_CONFIG_OVERRIDE="$prim/config" \ FM_SPAWN_NO_GUARD=1 CLAUDECODE=1 TMUX="fake,1,0" \ @@ -387,8 +391,12 @@ test_duplicate_secondmate_spawn_does_not_converge_trace_context() { printf 'charter\n' > "$sm/data/charter.md" fake=$(make_spawn_fakebin "$base/fake") + # A claude secondmate spawn pre-registers workspace trust for the HOME it + # launches into (bin/fm-claude-trust.sh), so this runs against a throwaway + # HOME; without it this suite would write the developer's real ~/.claude.json. + mkdir -p "$base/user-home" out=$(env -u FM_TRACE_CONTEXT \ - FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$prim" \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$prim" HOME="$base/user-home" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$prim/state" FM_DATA_OVERRIDE="$prim/data" \ FM_PROJECTS_OVERRIDE="$prim/projects" FM_CONFIG_OVERRIDE="$prim/config" \ FM_SPAWN_NO_GUARD=1 CLAUDECODE=1 TMUX="fake,1,0" \ From c191eacd36a9a44db0fb1888c67fe3dbfe534919 Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Sat, 12 Sep 2026 00:06:14 -0700 Subject: [PATCH 17/31] fix: ignore superseded failed GitHub check runs (#4258) * fix(pr-merge): judge each required check by its current run When the base branch advances, GitHub cancels a pull request's in-flight run and re-triggers it. The cancelled run stays in statusCheckRollup beside the passing re-run, so the rollup can hold several runs of one check name at the same head while GitHub itself reports the pull request CLEAN. github_checks_not_green judged every run independently, so that superseded failure refused a genuinely mergeable pull request and pushed the operator toward a needless --allow-red. Group the rollup by the reported name and judge each check by its current run. Supersession is proven, never assumed: a name leaves the red set only when every one of its non-green runs is strictly older than one of its green runs, dated by the forge's own settled timestamp - a check run's completedAt once its status is COMPLETED, or a status context's createdAt - and only in the whole-second UTC form GitHub emits, which is the one spelling that orders correctly as plain text. A run with no such timestamp is never superseded, so a still-running, queued or undated run keeps its check red, and a name with no green run at all stays red. An unnamed entry is grouped alone so two unrelated unnamed checks are never treated as one. Every comparison is one-directional: it can only clear a failure a later success provably replaced, and never clears a check whose current run failed, is pending, or is missing. No other guard moves - the pull request must still be open, undrafted, mergeable, conflict-free and head-bound, and --allow-red still waives exactly its named check with every other check green. Live reproduction: PR #4224 read CLEAN with an old FAILURE and a newer SUCCESS for one check name and was refused; it now verifies, while #4208 and #4210, whose latest runs failed, still refuse. * no-mistakes(review): Use check-run start times for safe supersession * no-mistakes(document): Clarify GitHub check-rollup documentation --- bin/fm-pr-merge.sh | 86 ++++++++++-- docs/architecture.md | 1 + tests/fm-pr-merge.test.sh | 284 ++++++++++++++++++++++++++++++++++++++ 3 files changed, 358 insertions(+), 13 deletions(-) diff --git a/bin/fm-pr-merge.sh b/bin/fm-pr-merge.sh index 174c3188dae..0b8aa79aeb2 100755 --- a/bin/fm-pr-merge.sh +++ b/bin/fm-pr-merge.sh @@ -11,7 +11,9 @@ # A GitHub merge is refused unless every pre-merge condition holds, each read # live at merge time rather than taken from recorded metadata: the pull request # is open, not a draft, mergeable, free of conflicts, and every unwaived check -# is green at the exact current head commit. Every failing condition is reported, not +# is green at the exact current head commit, where github_checks_not_green below +# owns what makes a check green and judges each one by its current run. +# Every failing condition is reported, not # just the first. The verified head is then passed to gh as # --match-head-commit, so a push that lands between that read and the merge # fails the merge instead of landing commits nothing verified. Reading that @@ -443,22 +445,80 @@ FIELDS } # Every GitHub check that is not green in the given live pull-request JSON, one -# name per line: a status context whose state is not SUCCESS, or a check run -# that has not completed with SUCCESS, NEUTRAL, or SKIPPED (so a pending -# check is not green either). Exits nonzero when the rollup cannot be read, so -# a malformed answer is a failed read and never an empty red set. +# name per line. An entry is green when it is a status context whose state is +# SUCCESS, or a check run that completed with SUCCESS, NEUTRAL, or SKIPPED (so +# a pending check is not green either). Exits nonzero when the rollup cannot be +# read, so a malformed answer is a failed read and never an empty red set. +# +# The rollup can hold several runs of one check name at the same head, because +# GitHub cancels a pull request's in-flight run when the base branch advances +# and re-triggers it; the cancelled run stays in the rollup beside the passing +# re-run. A check is therefore judged by its current run rather than by any run +# that a later one superseded, which is what makes this agree with GitHub's own +# CLEAN mergeStateStatus instead of refusing a pull request GitHub considers +# mergeable. +# +# Supersession applies only among check runs with the same reported name. A +# name is dropped from the red set only when every non-green run is COMPLETED, +# has a whole-second UTC startedAt, and started strictly before a green run. +# Status contexts are never grouped or superseded, and every non-green one is +# reported independently. A still-running, queued, undated, or tied check run +# stays red. A name whose runs are all green needs no timestamp, while a name +# with no green run stays red. +# +# The reported name is also what --allow-red matches. An unnamed check run is +# grouped alone and can neither supersede nor be superseded, because unrelated +# unnamed checks must not be treated as one. github_checks_not_green() { local json=$1 printf '%s' "$json" | jq -r ' + def settled_at: + if type == "string" and test("^[0-9]{4}-[0-9]{2}-[0-9]{2}T[0-9]{2}:[0-9]{2}:[0-9]{2}Z$") + then . else null end; if (.statusCheckRollup | type) != "array" then error("no check rollup") else . end - | .statusCheckRollup[] - | if .__typename == "CheckRun" then - {name: (.name // ""), ok: (.status == "COMPLETED" and (.conclusion == "SUCCESS" or .conclusion == "NEUTRAL" or .conclusion == "SKIPPED"))} - else - {name: (.context // ""), ok: (.state == "SUCCESS")} - end - | select(.ok | not) - | if .name == "" then "(unnamed check)" else .name end + | [ .statusCheckRollup + | to_entries[] + | .key as $i + | .value + | if .__typename == "CheckRun" then + { + kind: "check_run", + name: (.name // ""), + completed: (.status == "COMPLETED"), + ok: (.status == "COMPLETED" and (.conclusion == "SUCCESS" or .conclusion == "NEUTRAL" or .conclusion == "SKIPPED")), + at: (.startedAt | settled_at) + } + | . + {group: (if .name == "" then ["", $i] else [.name, -1] end)} + else + {kind: "status_context", name: (.context // ""), ok: (.state == "SUCCESS")} + end + ] + | . as $entries + | ( + ($entries[] + | select(.kind == "status_context" and (.ok | not)) + | .name + ), + ($entries + | [.[] | select(.kind == "check_run")] + | group_by(.group)[] + | { + name: .[0].name, + reds: [.[] | select(.ok | not)], + newest_green: ([.[] | select(.ok) | .at | select(. != null)] | max) + } + | select( + (.reds | length) > 0 + and ( + .newest_green == null + or any(.reds[]; (.completed | not) or .at == null) + or ([.reds[] | .at] | max) >= .newest_green + ) + ) + | .name + ) + ) + | if . == "" then "(unnamed check)" else . end ' 2>/dev/null || return 1 } diff --git a/docs/architecture.md b/docs/architecture.md index 7ab4e496f7d..a02e4b9a1e5 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -314,6 +314,7 @@ This repo uses that setting, and its own `.no-mistakes/` directory remains local PR-based task merges go through `bin/fm-pr-merge.sh`, which records `pr=` and any available `pr_head=` through `bin/fm-pr-check.sh` before calling the forge CLI. The helper requires a full canonical URL and rejects malformed URLs or repo override flags before recording merge state. A `https://github.com/<owner>/<repo>/pull/<n>` URL requires `gh` and `jq`, is merged only after one live read confirms the pull request is open, not a draft, mergeable, conflict-free, and every unwaived check is green at the current head, then `gh pr merge` binds that verified head with `--match-head-commit`. +A check run is green when its current run is green, because GitHub leaves a cancelled run in the rollup beside the passing re-run it triggered when the base branch advanced; `bin/fm-pr-merge.sh`'s `github_checks_not_green` owns the rule, which uses `startedAt` to clear only an older completed check run that a passing run with the same name provably replaced, while unfinished check runs and non-green status contexts stay red. `--auto`, `--admin`, and branch-deletion flags are refused unless `--attended-override` is passed for an explicit captain instruction; that override never skips the live green check, the away-grant check, or a captain hold. An attended `--allow-red <check-name>` may appear once, waives only GitHub checks with that exact name, and is refused while the away-posture record exists. A `https://<host>/<path>/-/merge_requests/<n>` URL (see [docs/gitlab-merge-watch.md](gitlab-merge-watch.md)) invokes `glab mr merge <n> -R https://<host>/<path>`, so the instance comes from the URL, and adds no merge-method flag because the project's own merge method applies. diff --git a/tests/fm-pr-merge.test.sh b/tests/fm-pr-merge.test.sh index 7d29879a6d0..5f4122b001e 100755 --- a/tests/fm-pr-merge.test.sh +++ b/tests/fm-pr-merge.test.sh @@ -77,6 +77,41 @@ write_github_red_json() { JSON } +# One CheckRun rollup entry the way GitHub reports it. A conclusion or timestamp +# of "-" is emitted as JSON null. Args: name status conclusion [startedAt] +# [completedAt] +check_run() { + local name=$1 status=$2 conclusion=$3 started=${4:--} completed=${5:-${4:--}} + local conclusion_json='null' started_json='null' completed_json='null' + [ "$conclusion" = - ] || conclusion_json="\"$conclusion\"" + [ "$started" = - ] || started_json="\"$started\"" + [ "$completed" = - ] || completed_json="\"$completed\"" + printf '{"__typename":"CheckRun","name":"%s","status":"%s","conclusion":%s,"startedAt":%s,"completedAt":%s}' \ + "$name" "$status" "$conclusion_json" "$started_json" "$completed_json" +} + +status_context() { + local name=$1 state=$2 + printf '{"__typename":"StatusContext","context":"%s","state":"%s"}' "$name" "$state" +} + +# Live GitHub JSON whose rollup holds the given entries verbatim, so a test can +# put several runs of one check name at the same head the way GitHub does after +# it cancels a pull request's in-flight run and re-triggers it. mergeStateStatus +# stays CLEAN because that is what GitHub reports for exactly this case. +# Args: case_dir head_sha <rollup-entry-json>... +write_github_rollup_json() { + local case_dir=$1 head=$2 entry rollup='' + shift 2 + for entry in "$@"; do + rollup="${rollup:+$rollup,}$entry" + done + printf '%s\n' "$head" > "$case_dir/github-head" + cat > "$case_dir/github-view.json" <<JSON +{"state":"OPEN","isDraft":false,"mergeable":"MERGEABLE","mergeStateStatus":"CLEAN","headRefOid":"$head","statusCheckRollup":[$rollup]} +JSON +} + assert_logged_gh_merge() { local case_dir=$1 number=$2 repo=$3 head line extra= shift 3 @@ -2284,6 +2319,246 @@ test_github_red_checks_refuse_and_allow_red_waives_named() { pass "fm-pr-merge refuses red GitHub checks and waives only a named --allow-red check" } +# When the base branch advances, GitHub cancels a pull request's in-flight run +# and re-triggers it, leaving the cancelled run in the rollup beside the passing +# re-run while reporting the pull request itself CLEAN. The merge must follow the +# current run rather than the one that re-run replaced. +test_superseded_failed_check_run_no_longer_refuses() { + local case_dir head + head=cccccccccccccccccccccccccccccccccccccccc + case_dir=$(make_case github-superseded-red) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(check_run ci COMPLETED CANCELLED 2026-01-01T00:00:01Z)" \ + "$(check_run ci COMPLETED SUCCESS 2026-01-01T00:00:09Z)" + + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/90 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" \ + || fail "github-superseded-red: a failed run replaced by a passing re-run must merge"$'\n'"$(cat "$case_dir/stderr")" + assert_logged_gh_merge "$case_dir" 90 example/repo --squash + pass "fm-pr-merge merges when a failed check run was replaced by a passing re-run" +} + +# Legacy status contexts remain independent from check runs, even when their +# reported names match. +test_check_runs_never_supersede_status_contexts() { + local case_dir rc head + head=cdcdcdcdcdcdcdcdcdcdcdcdcdcdcdcdcdcdcdcd + case_dir=$(make_case github-cross-check-kind) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(status_context ci FAILURE)" \ + "$(check_run ci COMPLETED SUCCESS 2026-01-01T00:00:09Z)" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/97 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "github-cross-check-kind: a failing status context must refuse" + assert_grep "check 'ci' is not green" "$case_dir/stderr" \ + "github-cross-check-kind: the status context was not named" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-cross-check-kind: a passing check run hid a failing status context" + pass "fm-pr-merge never lets a check run supersede a legacy status context" +} + +# The inverse, and the one that matters most: a check whose current run failed is +# still red however many earlier runs of it passed. +test_current_failed_check_run_still_refuses() { + local case_dir rc head + head=dddddddddddddddddddddddddddddddddddddddd + case_dir=$(make_case github-current-red) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(check_run ci COMPLETED SUCCESS 2026-01-01T00:00:01Z)" \ + "$(check_run ci COMPLETED FAILURE 2026-01-01T00:00:09Z)" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/91 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "github-current-red: a currently failing check must refuse" + assert_grep "check 'ci' is not green" "$case_dir/stderr" \ + "github-current-red: the red check was not named" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-current-red: gh pr merge ran on a currently failing check" + pass "fm-pr-merge still refuses when a check's current run failed after an earlier pass" +} + +# Run generation follows startedAt rather than the order overlapping runs finish. +test_late_finishing_old_success_does_not_hide_current_failure() { + local case_dir rc head + head=dededededededededededededededededededede + case_dir=$(make_case github-old-success-finishes-last) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(check_run ci COMPLETED SUCCESS 2026-01-01T00:00:01Z 2026-01-01T00:00:10Z)" \ + "$(check_run ci COMPLETED FAILURE 2026-01-01T00:00:09Z 2026-01-01T00:00:09Z)" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/98 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "github-old-success-finishes-last: the later-started failure must refuse" + assert_grep "check 'ci' is not green" "$case_dir/stderr" \ + "github-old-success-finishes-last: the current failure was not named" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-old-success-finishes-last: completion order hid the current failure" + pass "fm-pr-merge uses start order when the old success finishes last" +} + +# A cancelled old run may settle after the passing re-run that superseded it. +test_late_finishing_old_cancellation_is_superseded() { + local case_dir head + head=dfdfdfdfdfdfdfdfdfdfdfdfdfdfdfdfdfdfdfdf + case_dir=$(make_case github-old-cancellation-finishes-last) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(check_run ci COMPLETED CANCELLED 2026-01-01T00:00:01Z 2026-01-01T00:00:10Z)" \ + "$(check_run ci COMPLETED SUCCESS 2026-01-01T00:00:09Z 2026-01-01T00:00:09Z)" + + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/99 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" \ + || fail "github-old-cancellation-finishes-last: the passing re-run must merge"$'\n'"$(cat "$case_dir/stderr")" + assert_logged_gh_merge "$case_dir" 99 example/repo --squash + pass "fm-pr-merge supersedes an old cancellation that finishes last" +} + +# A re-run that has not finished proves nothing, so it can neither be superseded +# nor supersede: the check stays red whether the run it replaces passed or failed. +test_unfinished_rerun_keeps_a_check_red() { + local case_dir rc head prior + head=eeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeeee + for prior in FAILURE SUCCESS; do + case_dir=$(make_case "github-pending-rerun-$prior") + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(check_run ci COMPLETED "$prior" 2026-01-01T00:00:01Z)" \ + "$(check_run ci IN_PROGRESS - -)" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/92 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "github-pending-rerun-$prior: an unfinished re-run must refuse" + assert_grep "check 'ci' is not green" "$case_dir/stderr" \ + "github-pending-rerun-$prior: the pending check was not named" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-pending-rerun-$prior: gh pr merge ran with a re-run still in flight" + done + pass "fm-pr-merge keeps a check red while its re-run is still in flight" +} + +# Supersession is scoped to one check name, which is also the name --allow-red +# matches, so a newer passing check never clears a different check's failure. +test_supersession_never_crosses_check_names() { + local case_dir rc head + head=ffffffffffffffffffffffffffffffffffffffff + case_dir=$(make_case github-cross-name) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(check_run lint COMPLETED FAILURE 2026-01-01T00:00:01Z)" \ + "$(check_run ci COMPLETED SUCCESS 2026-01-01T00:00:09Z)" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/93 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "github-cross-name: another check passing must not clear this failure" + assert_grep "check 'lint' is not green" "$case_dir/stderr" \ + "github-cross-name: the red check was not named" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-cross-name: gh pr merge ran on a red check of a different name" + pass "fm-pr-merge never lets one check's pass clear another check's failure" +} + +# Supersession has to be proven from the forge's own start timestamps, so a run +# GitHub dated in any other way is treated as undated and clears nothing. +test_undated_runs_never_supersede() { + local case_dir rc spec label older newer + local head=0a0a0a0a0a0a0a0a0a0a0a0a0a0a0a0a0a0a0a0a + set -- \ + 'undated-failure|-|2026-01-01T00:00:09Z' \ + 'undated-pass|2026-01-01T00:00:01Z|-' \ + 'fractional-pass|2026-01-01T00:00:01Z|2026-01-01T00:00:09.500Z' \ + 'offset-pass|2026-01-01T00:00:01Z|2026-01-01T00:00:09+00:00' + for spec in "$@"; do + label=${spec%%|*} + older=${spec#*|} + older=${older%%|*} + newer=${spec##*|} + case_dir=$(make_case "github-undated-$label") + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(check_run ci COMPLETED FAILURE "$older")" \ + "$(check_run ci COMPLETED SUCCESS "$newer")" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/94 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "github-undated-$label: an unproven supersession must refuse" + assert_grep "check 'ci' is not green" "$case_dir/stderr" \ + "github-undated-$label: the red check was not named" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "github-undated-$label: gh pr merge ran on an unproven supersession" + done + pass "fm-pr-merge clears a failure only on a proven later pass of the same check" +} + +# A superseded failure changes nothing about the waiver: --allow-red still covers +# exactly the named check, still needs every other check green, and the merge is +# still bound to the verified head. +test_allow_red_still_waives_only_the_current_failure() { + local case_dir rc head + head=0b0b0b0b0b0b0b0b0b0b0b0b0b0b0b0b0b0b0b0b + case_dir=$(make_case github-superseded-allow-red-wrong-name) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(check_run ci COMPLETED FAILURE 2026-01-01T00:00:01Z)" \ + "$(check_run ci COMPLETED SUCCESS 2026-01-01T00:00:09Z)" \ + "$(check_run lint COMPLETED FAILURE 2026-01-01T00:00:09Z)" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/95 \ + --allow-red ci > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" "superseded-allow-red-wrong-name: waiving the green check must not merge" + assert_grep "check 'lint' is not green" "$case_dir/stderr" \ + "superseded-allow-red-wrong-name: the unwaived red check was not named" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "superseded-allow-red-wrong-name: gh pr merge ran with an unwaived red check" + + case_dir=$(make_case github-superseded-allow-red-named) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_github_rollup_json "$case_dir" "$head" \ + "$(check_run ci COMPLETED FAILURE 2026-01-01T00:00:01Z)" \ + "$(check_run ci COMPLETED SUCCESS 2026-01-01T00:00:09Z)" \ + "$(check_run lint COMPLETED FAILURE 2026-01-01T00:00:09Z)" + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/96 \ + --allow-red lint > "$case_dir/stdout" 2> "$case_dir/stderr" \ + || fail "superseded-allow-red-named: the named waiver should merge"$'\n'"$(cat "$case_dir/stderr")" + assert_logged_gh_merge "$case_dir" 96 example/repo --squash + pass "fm-pr-merge keeps --allow-red scoped to its named check beside a superseded failure" +} + test_allow_red_is_refused_while_away() { local case_dir rc head head=abababababababababababababababababababab @@ -2499,6 +2774,15 @@ test_untraversable_user_backend_config_directory_refuses_the_merge test_absent_user_backend_config_directory_and_backlog_still_merge test_backend_override_bypasses_unreadable_user_config test_github_red_checks_refuse_and_allow_red_waives_named +test_superseded_failed_check_run_no_longer_refuses +test_check_runs_never_supersede_status_contexts +test_current_failed_check_run_still_refuses +test_late_finishing_old_success_does_not_hide_current_failure +test_late_finishing_old_cancellation_is_superseded +test_unfinished_rerun_keeps_a_check_red +test_supersession_never_crosses_check_names +test_undated_runs_never_supersede +test_allow_red_still_waives_only_the_current_failure test_allow_red_is_refused_while_away test_allow_red_requires_one_separate_name test_away_grant_and_yolo_and_hold_for_return From de00521b9e671675aaf0a7481c51eac5e74ec868 Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Sat, 12 Sep 2026 00:43:41 -0700 Subject: [PATCH 18/31] fix(bin): persist merge authority for poll-detected outcomes (#4266) * fix(merge): persist the merge authority on poll-detected merge outcomes The merge ledger tags a merge with the authority that permitted it while the away-posture record existed, but only the direct attended merge in bin/fm-pr-merge.sh recorded it. A merge the forge queued, or one the merge poll detected after the fact, published an untagged row, so exactly the merges no agent watched were the least auditable. bin/fm-merge-authority-lib.sh now owns that answer, read from the same structured sources the merge gate already used: the task's recorded yolo posture and the away-posture record's mechanical grant list, never prose. bin/fm-pr-merge.sh keeps its own refusal wording and gates on that answer; bin/fm-watch.sh only records it on the row its poll publishes, so reading the authority never becomes a second path to a merge. An unresolved answer records an untagged row rather than dropping the outcome or inventing an authority. * no-mistakes(review): Persist canonical merge authority for queued poll outcomes * no-mistakes(review): Harden merge authority persistence against lifecycle races * no-mistakes(review): Serialize poll authority publication with teardown * no-mistakes(document): Clarify persisted merge authority lifecycle * no-mistakes(ci): Added targeted SC2034 suppressions for the two public result assignments in bin/fm-merge-authority-lib.sh. Verified successfully with `CI=true bin/fm-lint.sh` --- AGENTS.md | 1 + bin/fm-merge-authority-lib.sh | 201 ++++++++++++++++++++ bin/fm-merge-outcome-lib.sh | 13 +- bin/fm-pr-merge.sh | 83 +++++--- bin/fm-teardown.sh | 10 +- bin/fm-watch.sh | 37 +++- docs/architecture.md | 3 + docs/scripts.md | 1 + tests/fm-pr-check-security.test.sh | 296 ++++++++++++++++++++++++++++- 9 files changed, 602 insertions(+), 43 deletions(-) create mode 100755 bin/fm-merge-authority-lib.sh diff --git a/AGENTS.md b/AGENTS.md index 3abd8348a28..0c7cf568b99 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -113,6 +113,7 @@ state/ runtime records and signals; gitignored <id>.pr-poll private validated data sidecar for the byte-static PR merge poll <id>.pr-poll-registration private transactional provenance record binding the task, canonical metadata identity, sidecar, and static poll publication <id>.pr-poll-retirement private identity-bound crash-recovery receipt for one exact validated merged result; removed after its poll artifacts retire + <id>.merge-authority private canonical-PR-bound authority persisted after firstmate's forge merge request is accepted and consumed by a later merged poll; bin/fm-merge-authority-lib.sh owns its format and lifecycle <id>.pr-poll-merge-notified canonical PR identity of the last merge outcome delivered for this task; bin/fm-pr-lib.sh owns the marker format and identity mechanics, while bin/fm-merge-outcome-lib.sh owns locked publication, duplicate suppression, and replacement branch-outcomes.jsonl .branch-outcomes-cursor .branch-outcomes-processed .<task>.branch-outcome-index .branch-outcome-index-ready Pi supervision-branch durable outcome store, its read cursor, main's processed marker, bounded latest per-task status-coverage caches, and their recovery marker; bin/fm-branch-outcome.sh owns the formats branch-session/ .branch-session .branch-mirror-cursor the branch's per-main-session conversations, the pointer to the current one, and the dialog-mirror cursor; extension-owned (docs/pi-supervision-branch.md) diff --git a/bin/fm-merge-authority-lib.sh b/bin/fm-merge-authority-lib.sh new file mode 100755 index 00000000000..9dbbadda2b1 --- /dev/null +++ b/bin/fm-merge-authority-lib.sh @@ -0,0 +1,201 @@ +#!/usr/bin/env bash +# Durable ownership of the authority under which a task's merge was accepted. +# +# The away-posture record (state/.afk-contract) and the task's recorded yolo +# posture are resolved only at the merge gate. After a forge accepts the merge, +# bin/fm-pr-merge.sh persists that answer as: +# state/<task-id>.merge-authority +# fm-merge-authority-v1 +# <provider> +# <host> +# <path> +# <number> +# <authority> yolo | away-grant | attended +# The identity comes from the merge run's immutable canonical URL parse; +# persistence revalidates the task's current pr= metadata under its metadata +# and lifecycle locks and refuses a mismatch. The file is atomically published, +# mode 0600, single-link, and on the state filesystem. A poll consumes it only +# when all identity fields match its own validated snapshot. Missing, malformed, +# or mismatched state means external; it is never resolved again from a later +# away-posture record. +# +# Resolution authorizes nothing by itself. bin/fm-pr-merge.sh owns the merge +# gate and persists only after a forge command succeeds, before releasing the +# task lifecycle lock. After observing a landed merge, bin/fm-watch.sh acquires +# that same lock, revalidates the poll, publishes its durable outcome, and +# retires only the exact authority record it read. Teardown uses the same lock, +# so it cannot interleave with that consumption transaction, and removes any +# remaining record. +# +# Sourced by those scripts and by tests. No side effects on source beyond its +# sourced libraries. + +_FM_MERGE_AUTHORITY_LIB_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +# shellcheck source=bin/fm-pr-lib.sh +. "$_FM_MERGE_AUTHORITY_LIB_DIR/fm-pr-lib.sh" +# shellcheck source=bin/fm-afk-contract.sh +. "$_FM_MERGE_AUTHORITY_LIB_DIR/fm-afk-contract.sh" + +# shellcheck disable=SC2034 # Public results consumed by sourcing callers. +FM_MERGE_AUTHORITY= +# shellcheck disable=SC2034 # Public results consumed by sourcing callers. +FM_MERGE_AUTHORITY_REASON= +# shellcheck disable=SC2034 # Public results consumed by sourcing callers. +FM_MERGE_AUTHORITY_RECORD_IDENTITY= + +fm_merge_authority_resolve() { # <home> <state> <meta> <task-id> + local home=${1-} state=${2-} meta=${3-} id=${4-} + local yolo='' grants grant + FM_MERGE_AUTHORITY= + FM_MERGE_AUTHORITY_REASON='invalid' + [ -n "$home" ] && [ -n "$state" ] && [ -n "$meta" ] && [ -n "$id" ] || return 1 + + if ! fm_afk_contract_present "$state"; then + FM_MERGE_AUTHORITY='attended' + FM_MERGE_AUTHORITY_REASON='attended' + return 0 + fi + if ! FM_HOME="$home" FM_STATE_OVERRIDE="$state" \ + "$_FM_MERGE_AUTHORITY_LIB_DIR/fm-afk-contract.sh" validate >/dev/null 2>&1; then + FM_MERGE_AUTHORITY_REASON='record-unreadable' + return 1 + fi + if [ -f "$meta" ]; then + yolo=$(grep '^yolo=' "$meta" | tail -1 | cut -d= -f2- || true) + fi + if [ "$yolo" = on ]; then + FM_MERGE_AUTHORITY='yolo' + FM_MERGE_AUTHORITY_REASON='granted' + return 0 + fi + grants=$(FM_HOME="$home" FM_STATE_OVERRIDE="$state" \ + "$_FM_MERGE_AUTHORITY_LIB_DIR/fm-afk-contract.sh" grants 2>/dev/null) || { + FM_MERGE_AUTHORITY_REASON='grants-unreadable' + return 1 + } + while IFS= read -r grant; do + [ "$grant" = "$id" ] || continue + FM_MERGE_AUTHORITY='away-grant' + FM_MERGE_AUTHORITY_REASON='granted' + return 0 + done <<EOF +$grants +EOF + # shellcheck disable=SC2034 # Public results consumed by sourcing callers. + FM_MERGE_AUTHORITY_REASON='not-granted' + return 1 +} + +fm_merge_authority_record_matches() { # <record> <device> <provider> <host> <path> <number> + local record=$1 device=$2 expected_provider=$3 expected_host=$4 expected_path=$5 expected_number=$6 + local version provider host path number authority + fm_pr_private_file_valid "$record" 600 "$device" || return 1 + exec 8< "$record" || return 1 + IFS= read -r version <&8 || { exec 8<&-; return 1; } + IFS= read -r provider <&8 || { exec 8<&-; return 1; } + IFS= read -r host <&8 || { exec 8<&-; return 1; } + IFS= read -r path <&8 || { exec 8<&-; return 1; } + IFS= read -r number <&8 || { exec 8<&-; return 1; } + IFS= read -r authority <&8 || { exec 8<&-; return 1; } + if IFS= read -r _extra <&8; then + exec 8<&- + return 1 + fi + exec 8<&- + case "$authority" in yolo|away-grant|attended) ;; *) return 1 ;; esac + [ "$version" = fm-merge-authority-v1 ] \ + && [ "$provider" = "$expected_provider" ] \ + && [ "$host" = "$expected_host" ] \ + && [ "$path" = "$expected_path" ] \ + && [ "$number" = "$expected_number" ] || return 1 + FM_MERGE_AUTHORITY=$authority +} + +fm_merge_authority_persist() { # <state> <task-id> <meta> <provider> <host> <path> <number> <authority> + local state=$1 id=$2 meta=$3 provider=$4 host=$5 path=$6 number=$7 authority=$8 + local record tmp='' state_device lock status=0 + fm_pr_task_id_valid "$id" || return 1 + case "$authority" in yolo|away-grant|attended) ;; *) return 1 ;; esac + [ -d "$state" ] && [ ! -L "$state" ] || return 1 + state_device=$(fm_pr_file_device "$state") || return 1 + fm_pr_metadata_identity_parse "$meta" || return 1 + [ "$FM_PR_META_PROVIDER" = "$provider" ] \ + && [ "$FM_PR_META_HOST" = "$host" ] \ + && [ "$FM_PR_META_PATH" = "$path" ] \ + && [ "$FM_PR_META_NUMBER" = "$number" ] || return 1 + record="$state/$id.merge-authority" + lock="$record.lock" + fm_lock_acquire_wait "$lock" || return 1 + fm_pr_regular_destination_on_device_or_absent "$record" "$state_device" || status=1 + if [ "$status" -eq 0 ]; then + umask 077 + tmp=$(mktemp "$state/.fm-merge-authority.XXXXXX") || status=1 + fi + if [ "$status" -eq 0 ]; then + printf '%s\n%s\n%s\n%s\n%s\n%s\n' \ + fm-merge-authority-v1 "$provider" "$host" "$path" "$number" "$authority" > "$tmp" \ + || status=1 + fi + if [ "$status" -eq 0 ]; then + chmod 0600 "$tmp" \ + && fm_merge_authority_record_matches "$tmp" "$state_device" \ + "$provider" "$host" "$path" "$number" \ + && fm_pr_regular_destination_on_device_or_absent "$record" "$state_device" \ + && mv -f -- "$tmp" "$record" \ + && fm_merge_authority_record_matches "$record" "$state_device" \ + "$provider" "$host" "$path" "$number" \ + || status=1 + fi + [ "$status" -eq 0 ] || rm -f -- "$tmp" + fm_lock_release "$lock" || status=1 + return "$status" +} + +fm_merge_authority_read() { # <state> <task-id> <provider> <host> <path> <number> + local state=$1 id=$2 provider=$3 host=$4 path=$5 number=$6 + local record state_device lock status=0 + FM_MERGE_AUTHORITY='external' + FM_MERGE_AUTHORITY_RECORD_IDENTITY= + fm_pr_task_id_valid "$id" || return 1 + [ -d "$state" ] && [ ! -L "$state" ] || return 1 + state_device=$(fm_pr_file_device "$state") || return 1 + record="$state/$id.merge-authority" + lock="$record.lock" + fm_lock_acquire_wait "$lock" || return 1 + if fm_merge_authority_record_matches "$record" "$state_device" \ + "$provider" "$host" "$path" "$number"; then + # shellcheck disable=SC2034 # Public results consumed by sourcing callers. + FM_MERGE_AUTHORITY_RECORD_IDENTITY=$(fm_pr_file_identity "$record") || status=1 + else + FM_MERGE_AUTHORITY='external' + status=1 + fi + fm_lock_release "$lock" || status=1 + return "$status" +} + +fm_merge_authority_remove_if_matches() { # <state> <task-id> <provider> <host> <path> <number> <authority> <file-identity> + local state=$1 id=$2 provider=$3 host=$4 path=$5 number=$6 + local authority=$7 expected_file_identity=$8 record state_device lock current_file_identity status=0 + fm_pr_task_id_valid "$id" || return 1 + [ -d "$state" ] && [ ! -L "$state" ] || return 1 + state_device=$(fm_pr_file_device "$state") || return 1 + record="$state/$id.merge-authority" + lock="$record.lock" + fm_lock_acquire_wait "$lock" || return 1 + if [ -e "$record" ] || [ -L "$record" ]; then + if fm_merge_authority_record_matches "$record" "$state_device" \ + "$provider" "$host" "$path" "$number"; then + current_file_identity=$(fm_pr_file_identity "$record") || status=1 + if [ "$status" -eq 0 ] \ + && [ "$FM_MERGE_AUTHORITY" = "$authority" ] \ + && [ "$current_file_identity" = "$expected_file_identity" ]; then + rm -f -- "$record" || status=1 + fi + elif ! fm_pr_private_file_valid "$record" 600 "$state_device"; then + status=1 + fi + fi + fm_lock_release "$lock" || status=1 + return "$status" +} diff --git a/bin/fm-merge-outcome-lib.sh b/bin/fm-merge-outcome-lib.sh index bf1f26c9c17..db279351145 100755 --- a/bin/fm-merge-outcome-lib.sh +++ b/bin/fm-merge-outcome-lib.sh @@ -42,10 +42,11 @@ FM_MERGE_OUTCOME_ALREADY_RECORDED=false # self - this home performed the merge. # poll - this home's merge poll detected the merge, so the canonical outcome # also wakes this home after any upward hop needed by a secondmate. -# Optional <authority> is yolo or away-grant when the merge ran while the -# away-posture record existed; it is appended to the ledger line. Known audit -# gap: queued merges and a poll that wins direct-merge deduplication publish an -# untagged row because the poll path does not persist merge authority. +# Optional <authority> is yolo, away-grant, attended, or external. Yolo, +# away-grant, and external are appended to the ledger line; attended remains +# untagged. The merge entrypoint supplies its authority after forge acceptance, +# while the poll supplies the persisted identity-bound value or external when +# no matching record proves that this home authorized the merge. # # Returns 0 when the outcome is recorded (or already was), 2 on an invalid # request, 3 when this home's own role or parent binding cannot be read well @@ -62,8 +63,8 @@ fm_merge_outcome_report() { # <home> <state> <task-id> <pr-url> <origin> [autho FM_MERGE_OUTCOME_ALREADY_RECORDED=false case "$origin" in self|poll) ;; *) return 2 ;; esac case "$authority" in - yolo|away-grant) suffix=" $authority" ;; - '') ;; + yolo|away-grant|external) suffix=" $authority" ;; + attended|'') ;; *) return 2 ;; esac fm_pr_task_id_valid "$id" || return 2 diff --git a/bin/fm-pr-merge.sh b/bin/fm-pr-merge.sh index 0b8aa79aeb2..85319a13777 100755 --- a/bin/fm-pr-merge.sh +++ b/bin/fm-pr-merge.sh @@ -73,8 +73,10 @@ # held for the captain return. An unreadable record refuses rather than being # skipped. Neither posture releases a captain hold, and the grant lapses when # the record is archived. -# The lock ends when the local forge command returns; docs/captain-hold-lifecycle.md owns -# the accepted asynchronous-landing and merge-to-cleanup residuals. +# A failed forge command releases the lock after it returns. A successful one +# retains the lock until the accepted merge authority is persisted against the +# still-matching task metadata; docs/captain-hold-lifecycle.md owns the accepted +# asynchronous-landing and merge-to-cleanup residuals. # # Extra args must not include --repo or -R in any form, including a bundled # short-option cluster such as -yR, because the repository comes only from the @@ -109,6 +111,8 @@ STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" . "$SCRIPT_DIR/fm-backlog-transition-lib.sh" # shellcheck source=bin/fm-merge-outcome-lib.sh . "$SCRIPT_DIR/fm-merge-outcome-lib.sh" +# shellcheck source=bin/fm-merge-authority-lib.sh +. "$SCRIPT_DIR/fm-merge-authority-lib.sh" # shellcheck source=bin/fm-afk-contract.sh . "$SCRIPT_DIR/fm-afk-contract.sh" @@ -124,6 +128,8 @@ if ! fm_pr_task_id_valid "$ID" || ! fm_pr_url_parse "$RAW_URL"; then fi URL=$FM_PR_URL PROVIDER=$FM_PR_PROVIDER +PR_HOST=$FM_PR_HOST +PR_PATH=$FM_PR_PATH PR_OWNER=$FM_PR_OWNER PR_REPO=$FM_PR_REPO PR_NUMBER=$FM_PR_NUMBER @@ -295,7 +301,9 @@ fi MERGE_EXPECTED_SPAWN_GEN=$FM_BACKLOG_META_SPAWN_GEN MERGE_CONTROL_LOCK= +MERGE_META_LOCK= merge_control_cleanup() { + [ -z "$MERGE_META_LOCK" ] || fm_lock_release "$MERGE_META_LOCK" || true [ -z "$MERGE_CONTROL_LOCK" ] || fm_lock_release "$MERGE_CONTROL_LOCK" || true } trap merge_control_cleanup EXIT @@ -825,33 +833,27 @@ require_released_captain_hold() { } FM_PR_MERGE_AUTHORITY= +# The gate on top of the shared authority read. bin/fm-merge-authority-lib.sh +# owns what the away-posture record and the task's recorded yolo posture say; +# this function owns what a merge run may do about it, so the answer the merge +# poll later tags its ledger row with is the same answer gated here. require_away_merge_grant() { - local yolo grants grant FM_PR_MERGE_AUTHORITY= - fm_afk_contract_present "$STATE" || return 0 - if ! FM_HOME="$FM_HOME" FM_STATE_OVERRIDE="$STATE" \ - "$SCRIPT_DIR/fm-afk-contract.sh" validate >/dev/null 2>&1; then - echo "error: PR merge refused - the away-posture record could not be read; nothing was merged" >&2 - return 1 - fi - yolo=$(grep '^yolo=' "$META" | tail -1 | cut -d= -f2- || true) - if [ "$yolo" = on ]; then - FM_PR_MERGE_AUTHORITY=yolo + if fm_merge_authority_resolve "$FM_HOME" "$STATE" "$META" "$ID"; then + FM_PR_MERGE_AUTHORITY=$FM_MERGE_AUTHORITY return 0 fi - grants=$(FM_HOME="$FM_HOME" FM_STATE_OVERRIDE="$STATE" \ - "$SCRIPT_DIR/fm-afk-contract.sh" grants 2>/dev/null) || { - echo "error: PR merge refused - the away-posture record's grants could not be read; nothing was merged" >&2 - return 1 - } - while IFS= read -r grant; do - [ "$grant" = "$ID" ] || continue - FM_PR_MERGE_AUTHORITY=away-grant - return 0 - done <<EOF -$grants -EOF - echo "error: task $ID is held for the captain return" >&2 + case "$FM_MERGE_AUTHORITY_REASON" in + record-unreadable) + echo "error: PR merge refused - the away-posture record could not be read; nothing was merged" >&2 + ;; + grants-unreadable) + echo "error: PR merge refused - the away-posture record's grants could not be read; nothing was merged" >&2 + ;; + *) + echo "error: task $ID is held for the captain return" >&2 + ;; + esac return 1 } @@ -863,6 +865,23 @@ require_current_away_authority() { fi } +persist_accepted_merge_authority() { + local status=0 + MERGE_META_LOCK=$(fm_meta_lock_path "$META") || return 1 + fm_lock_acquire_wait "$MERGE_META_LOCK" || return 1 + fm_merge_authority_persist "$STATE" "$ID" "$META" \ + "$PROVIDER" "$PR_HOST" "$PR_PATH" "$PR_NUMBER" "$FM_PR_MERGE_AUTHORITY" \ + || status=1 + fm_lock_release "$MERGE_META_LOCK" || status=1 + MERGE_META_LOCK= + if [ "$status" -eq 0 ]; then + return 0 + fi + printf 'actionable: the forge accepted the merge request for %s but its merge authority could not be persisted; the merge poll remains armed\n' \ + "$URL" >&2 + return 1 +} + require_recorded_pr_identity() { local existing existing=$(grep '^pr=' "$META" | tail -1 | cut -d= -f2- || true) @@ -1024,11 +1043,14 @@ case "$PROVIDER" in merge_output=$(gh pr merge "$PR_NUMBER" --repo "$PR_OWNER/$PR_REPO" \ --match-head-commit "$FM_PR_MERGE_HEAD" \ "${merge_args[@]+"${merge_args[@]}"}" "$@" 2>&1) || merge_status=$? - fm_lock_release "$MERGE_CONTROL_LOCK" || true - MERGE_CONTROL_LOCK= if [ "$merge_status" -eq 0 ]; then FM_PR_GITHUB_MERGE_ACCEPTED=true + persist_accepted_merge_authority || exit 1 + fm_lock_release "$MERGE_CONTROL_LOCK" || true + MERGE_CONTROL_LOCK= else + fm_lock_release "$MERGE_CONTROL_LOCK" || true + MERGE_CONTROL_LOCK= [ -z "$merge_output" ] || printf '%s\n' "$merge_output" >&2 if github_read_outcome; then if [ "$FM_PR_GITHUB_MERGED" != true ] && [ "$FM_PR_GITHUB_QUEUED" != true ]; then @@ -1071,9 +1093,14 @@ case "$PROVIDER" in merge_status=0 GITLAB_HOST="$FM_PR_HOST" glab mr merge "$PR_NUMBER" -R "$PROJECT_URL" \ --sha "$FM_PR_MERGE_HEAD" --yes "$@" || merge_status=$? + if [ "$merge_status" -ne 0 ]; then + fm_lock_release "$MERGE_CONTROL_LOCK" || true + MERGE_CONTROL_LOCK= + exit "$merge_status" + fi + persist_accepted_merge_authority || exit 1 fm_lock_release "$MERGE_CONTROL_LOCK" || true MERGE_CONTROL_LOCK= - [ "$merge_status" -eq 0 ] || exit "$merge_status" gitlab_confirm_rc=0 gitlab_confirm_merged || gitlab_confirm_rc=$? [ "$gitlab_confirm_rc" -eq 0 ] || exit 0 diff --git a/bin/fm-teardown.sh b/bin/fm-teardown.sh index 7eaabe22d5a..1c94623f8dc 100755 --- a/bin/fm-teardown.sh +++ b/bin/fm-teardown.sh @@ -1207,7 +1207,7 @@ validate_pr_poll_cleanup() { fm_task_id_path_safe "$id" || return 0 for artifact in "$state_dir/$id.check.sh" "$state_dir/$id.pr-poll" \ "$state_dir/$id.pr-poll-registration" "$state_dir/$id.pr-poll-retirement" \ - "$state_dir/$id.check-trust"; do + "$state_dir/$id.merge-authority" "$state_dir/$id.check-trust"; do [ -e "$artifact" ] || [ -L "$artifact" ] || continue has_artifact=1 done @@ -1216,11 +1216,13 @@ validate_pr_poll_cleanup() { state_device=$(fm_pr_file_device "$state_dir") || return 1 for artifact in "$state_dir/$id.check.sh" "$state_dir/$id.pr-poll" \ "$state_dir/$id.pr-poll-registration" "$state_dir/$id.pr-poll-retirement" \ - "$state_dir/$id.check-trust"; do + "$state_dir/$id.merge-authority" "$state_dir/$id.check-trust"; do [ -e "$artifact" ] || [ -L "$artifact" ] || continue if [ ! -f "$artifact" ] || [ -L "$artifact" ] \ || [ "$(fm_pr_file_device "$artifact")" != "$state_device" ] \ - || [ "$(fm_pr_file_link_count "$artifact")" != 1 ]; then + || [ "$(fm_pr_file_link_count "$artifact")" != 1 ] \ + || { [ "$artifact" = "$state_dir/$id.merge-authority" ] \ + && [ "$(fm_pr_file_mode "$artifact")" != 600 ]; }; then echo "REFUSED: unsafe task PR-check artifact; preserving task state." >&2 return 1 fi @@ -1241,7 +1243,7 @@ remove_pr_poll_artifacts() { fm_pr_poll_merge_notified_remove "$state_dir" "$id" || return 1 rm -f "$state_dir/$id.check.sh" "$state_dir/$id.pr-poll" \ "$state_dir/$id.pr-poll-registration" "$state_dir/$id.pr-poll-retirement" \ - "$state_dir/$id.check-trust" || return 1 + "$state_dir/$id.merge-authority" "$state_dir/$id.check-trust" || return 1 } # Resolve the PR number for a worktree branch via gh-axi. Echoes the number on a diff --git a/bin/fm-watch.sh b/bin/fm-watch.sh index 2a8e02b735a..bdd720a800d 100755 --- a/bin/fm-watch.sh +++ b/bin/fm-watch.sh @@ -147,6 +147,10 @@ mkdir -p "$STATE" # worker while adding no uncovered file. # shellcheck source=/dev/null . "$SCRIPT_DIR/fm-merge-outcome-lib.sh" +# The durable merge-authority owner is shared with bin/fm-pr-merge.sh. The +# watcher consumes only its identity-bound record after a poll observes landing. +# shellcheck source=/dev/null +. "$SCRIPT_DIR/fm-merge-authority-lib.sh" # shellcheck source=bin/fm-x-lib.sh . "$SCRIPT_DIR/fm-x-lib.sh" # shellcheck source=bin/fm-check-lib.sh @@ -1841,8 +1845,16 @@ reconcile_requests_detached() { RECONCILE_REQUEST_PID=$! } +PR_POLL_CONTROL_LOCK= + +pr_poll_control_release() { + [ -z "$PR_POLL_CONTROL_LOCK" ] || fm_lock_release "$PR_POLL_CONTROL_LOCK" || return 1 + PR_POLL_CONTROL_LOCK= +} + watcher_cleanup() { local cleanup_status=0 owns_lock=0 transition=release-lock + pr_poll_control_release || cleanup_status=1 if [ "$(cat "$WATCH_LOCK/pid" 2>/dev/null || true)" = "${WATCHER_PID:-}" ]; then owns_lock=1 if [ "${WATCHER_RECOVERY_PENDING:-0}" -eq 1 ] \ @@ -2012,6 +2024,13 @@ while :; do host=$FM_PR_POLL_SNAPSHOT_HOST path=$FM_PR_POLL_SNAPSHOT_PATH number=$FM_PR_POLL_SNAPSHOT_NUMBER + PR_POLL_CONTROL_LOCK="$STATE/.control-$id.lock" + fm_lock_acquire_wait "$PR_POLL_CONTROL_LOCK" || exit 1 + if ! fm_pr_poll_snapshot_matches "$STATE" "$id" "$SCRIPT_DIR/fm-pr-poll.sh"; then + pr_poll_control_release || exit 1 + triage_log "PR poll for $id changed before its validated check; skipping the stale snapshot" + continue + fi run_check_capture "$SCRIPT_DIR/fm-pr-poll.sh" --validated \ "$provider" "$url" "$host" "$path" "$number" || exit 1 out=$FM_CHECK_RESULT @@ -2029,14 +2048,28 @@ while :; do if [ -n "$out" ]; then reason="check: $c: $out" if [ "$is_pr_poll" -eq 1 ] && [ "$out" = merged ]; then + if ! fm_merge_authority_read "$STATE" "$id" \ + "$provider" "$host" "$path" "$number"; then + triage_log "no matching persisted merge authority for $id; recording an external merge outcome" + fi + merge_authority=$FM_MERGE_AUTHORITY + merge_authority_record_identity=$FM_MERGE_AUTHORITY_RECORD_IDENTITY merge_outcome_rc=0 fm_merge_outcome_report "$FM_HOME" "$STATE" "$id" "$url" poll \ - || merge_outcome_rc=$? + "$merge_authority" || merge_outcome_rc=$? if [ "$merge_outcome_rc" -ne 0 ]; then triage_log "merge outcome for $id could not be recorded (rc=$merge_outcome_rc)" exit 1 fi + if [ -n "$merge_authority_record_identity" ] \ + && ! fm_merge_authority_remove_if_matches "$STATE" "$id" \ + "$provider" "$host" "$path" "$number" "$merge_authority" \ + "$merge_authority_record_identity"; then + triage_log "published merge outcome for $id but could not retire its authority record" + exit 1 + fi retire_merged_pr_poll "$id" + pr_poll_control_release || exit 1 touch "$STATE/.last-check" if [ "$FM_MERGE_OUTCOME_ALREADY_RECORDED" = true ]; then triage_log "absorbed duplicate merged PR poll result for $id" @@ -2044,10 +2077,12 @@ while :; do fi wake "$reason" fi + pr_poll_control_release || exit 1 fm_wake_append check "$c" "$reason" || exit 1 touch "$STATE/.last-check" wake "$reason" fi + pr_poll_control_release || exit 1 done if [ -n "$rejected_checks" ]; then reason="check: rejected unauthenticated state checks:$rejected_checks" diff --git a/docs/architecture.md b/docs/architecture.md index a02e4b9a1e5..ef1b1519c54 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -327,6 +327,9 @@ An auto-merge request is held to the same standard: `--auto` that leaves the pul Every GitHub refusal states what it could not observe as plainly as what it did, so an unreadable branch-rule response, an unrecognised queue method, and a merge queue no available read can see are each named rather than left to look like a base branch with no queue at all. A confirmed merge leaves a durable role-routed outcome instead of living only in the merging agent's memory, and [`bin/fm-merge-outcome-lib.sh`](../bin/fm-merge-outcome-lib.sh)'s header owns its destination, shape, identity, normal-case deduplication, and at-least-once recovery. The same emitter handles a merge firstmate performed and one its poll detected, while the watcher immediately delivers the emitter's local actionable poll row. +After the forge accepts firstmate's merge request, the merge path persists the resolved yolo, away-grant, or attended authority bound to the task's canonical PR identity. +A later merged poll consumes only that matching persisted value; with no match it records the landing as external rather than consulting a live away-posture record that may have been archived or replaced. +[`bin/fm-merge-authority-lib.sh`](../bin/fm-merge-authority-lib.sh)'s header owns resolution, private atomic persistence, identity-checked consumption, and retirement, while only the merge path gates on the answer. Teardown is fail-closed for ship worktrees: dirty worktrees refuse, and committed work must be landed before the worktree is returned. A pool worktree is only returned after teardown passes the slot-ownership proof: a contradictory task record or a supported live endpoint refuses without touching either task, and no discard authority relaxes that. A slot's own owner claim, written by the spawn that takes it under the allocation lock and owned by [`bin/fm-wake-lib.sh`](../bin/fm-wake-lib.sh), covers a slot reassigned to a task that left no record the scan could reach: a claim naming a different task releases nothing - teardown warns, names the claimant, and finishes only the task's own cleanup - because Treehouse's own live process lease cannot answer ownership once the worker's exit releases it. diff --git a/docs/scripts.md b/docs/scripts.md index 3f0694bac8d..048a064aa3c 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -132,6 +132,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-pr-check.sh` | Record validated `pr=` and `pr_head=` values, then atomically arm a static merge poll | | `fm-pr-merge.sh` | Record PR metadata, merge a task's canonical full GitHub or GitLab URL, then refuse an outcome it cannot prove landed or queued | | `fm-merge-outcome-lib.sh` | Publish a confirmed merge's durable, role-routed supervision outcome | +| `fm-merge-authority-lib.sh` | Resolve merge authority at the gate, persist it against the accepted canonical PR, and identity-check its later poll consumption | | `fm-parent-channel-lib.sh` | Resolve a secondmate home's parent channel and append a captain-facing outcome line to it at most once | | `fm-promote.sh` | Promote a scout task in place to a protected ship task with an explicit delivery mode, and write the ship instructions carrying that mode's definition of done | | `fm-teardown.sh` | Fail-closed teardown: return landed ship worktrees, require completed scout deliverables, retire secondmate homes | diff --git a/tests/fm-pr-check-security.test.sh b/tests/fm-pr-check-security.test.sh index 240f26f40ba..40d8438fcb4 100755 --- a/tests/fm-pr-check-security.test.sh +++ b/tests/fm-pr-check-security.test.sh @@ -138,9 +138,9 @@ printf '%s\n' "$*" >> "$FM_TEST_GH_LOG" case "${1:-} ${2:-}" in "api graphql") printf '%s\n' \ - 'state=MERGED' \ - 'merged=true' \ - 'queued=false' \ + "state=${FM_TEST_GH_GRAPHQL_STATE:-MERGED}" \ + "merged=${FM_TEST_GH_GRAPHQL_MERGED:-true}" \ + "queued=${FM_TEST_GH_GRAPHQL_QUEUED:-false}" \ 'base=main' exit 0 ;; @@ -152,11 +152,16 @@ case "${1:-} ${2:-}" in ;; esac ;; + "pr merge") + [ -z "${FM_TEST_GH_MERGE_HOOK:-}" ] || "$FM_TEST_GH_MERGE_HOOK" + exit 0 + ;; esac case " $* " in *" headRefOid "*) printf '%s\n' "${FM_TEST_GH_HEAD:-0123456789abcdef0123456789abcdef01234567}" ;; *" state "*) [ "${FM_TEST_GH_FAIL:-0}" = 0 ] || exit 1 + [ -z "${FM_TEST_GH_STATE_STARTED:-}" ] || : > "$FM_TEST_GH_STATE_STARTED" [ "${FM_TEST_GH_SLEEP:-0}" = 0 ] || sleep "$FM_TEST_GH_SLEEP" printf '%s\n' "${FM_TEST_GH_STATE:-OPEN}" ;; @@ -201,10 +206,14 @@ write_task_meta() { "mode=no-mistakes" } +# Extra "field=value" arguments are written before pr=, because +# fm_pr_metadata_identity_parse rejects an unrecognised line after it. write_poll_meta() { local state=$1 id=$2 url=$3 + shift 3 fm_write_meta "$state/$id.meta" \ "window=fm-$id" \ + "$@" \ "pr=$url" } @@ -624,9 +633,10 @@ SH run_watcher_bounded() { local home=$1 fakebin=$2 check_interval=${FM_TEST_CHECK_INTERVAL:-0} watch_root=${FM_TEST_WATCH_ROOT:-$ROOT} + local check_timeout=${FM_TEST_CHECK_TIMEOUT:-1} shift 2 perl -e 'my $pid=fork; die unless defined $pid; if (!$pid) { exec @ARGV } local $SIG{ALRM}=sub { kill "TERM", $pid; waitpid $pid, 0; exit 124 }; alarm 10; waitpid $pid, 0; alarm 0; exit($? >> 8)' \ - env FM_HOME="$home" FM_ROOT_OVERRIDE="$watch_root" FM_CHECK_INTERVAL="$check_interval" FM_CHECK_TIMEOUT=1 \ + env FM_HOME="$home" FM_ROOT_OVERRIDE="$watch_root" FM_CHECK_INTERVAL="$check_interval" FM_CHECK_TIMEOUT="$check_timeout" \ FM_POLL=0.02 FM_HEARTBEAT=999999 FM_SIGNAL_GRACE=0 PATH="$fakebin:$BASE_PATH" "$WATCH" "$@" } @@ -2135,12 +2145,290 @@ test_gitlab_merged_poll_retires() { pass "GitHub and GitLab exact merged results share one retirement path" } +# --- poll-path merge authority ---------------------------------------------- + +write_away_record() { # <dir> [<fm-afk-contract.sh propose args>...] + local dir=$1 + shift + FM_HOME="$dir/home" FM_STATE_OVERRIDE="$dir/home/state" \ + "$ROOT/bin/fm-afk-contract.sh" propose "$@" >/dev/null \ + || fail "could not propose an away-posture record" + FM_HOME="$dir/home" FM_STATE_OVERRIDE="$dir/home/state" \ + "$ROOT/bin/fm-afk-contract.sh" confirm >/dev/null \ + || fail "could not confirm an away-posture record" +} + +archive_away_record() { # <dir> + FM_HOME="$1/home" FM_STATE_OVERRIDE="$1/home/state" \ + "$ROOT/bin/fm-afk-contract.sh" archive >/dev/null \ + || fail "could not archive the away-posture record" +} + +# The durable queue is TSV (epoch, sequence, kind, key, payload). +merged_ledger_row() { # <state> <task-id> + awk -F'\t' -v prefix="check: merge landed: $2 " \ + 'index($5, prefix) == 1 { print $5 }' "$1/.wake-queue" +} + +run_merged_poll_cycle() { # <dir> + local dir=$1 rc=0 + add_stop_custom_check "$dir" + set +e + FM_TEST_GH_STATE=MERGED run_watcher_bounded "$dir/home" "$dir/fakebin" \ + > "$dir/watch.out" 2> "$dir/watch.err" + rc=$? + set -e + [ "$rc" -eq 0 ] || fail "merged poll watcher failed: $(cat "$dir/watch.err")" +} + +queue_merge() { # <dir> <url> + local dir=$1 url=$2 rc=0 + set +e + FM_TEST_GH_GRAPHQL_STATE=OPEN FM_TEST_GH_GRAPHQL_MERGED=false \ + FM_TEST_GH_GRAPHQL_QUEUED=true \ + run_merge_entry "$dir" task-a "$url" > "$dir/merge.out" 2> "$dir/merge.err" + rc=$? + set -e + [ "$rc" -eq 0 ] || fail "queued merge failed: $(cat "$dir/merge.err")" + assert_grep "is queued" "$dir/merge.out" "the forge did not queue the merge" + [ -f "$dir/home/state/task-a.merge-authority" ] \ + || fail "the accepted queued merge did not persist its authority" +} + +test_merged_poll_row_carries_the_merge_authority() { + local dir state url expected posture + url=https://github.com/o/r/pull/1 + + for posture in yolo grant; do + dir=$(make_case "queued-merge-authority-$posture") + state="$dir/home/state" + write_task_meta "$dir" task-a + if [ "$posture" = yolo ]; then + printf 'yolo=on\n' >> "$state/task-a.meta" + write_away_record "$dir" + expected=yolo + else + write_away_record "$dir" --grant task-a + expected=away-grant + fi + run_check_entry "$dir" task-a "$url" >/dev/null 2> "$dir/seed.err" \ + || fail "$posture: could not arm the merge poll" + queue_merge "$dir" "$url" + archive_away_record "$dir" + run_merged_poll_cycle "$dir" + [ "$(merged_ledger_row "$state" task-a)" = "check: merge landed: task-a $url $expected" ] \ + || fail "$posture: archived posture lost persisted authority: $(merged_ledger_row "$state" task-a)" + [ ! -e "$state/task-a.merge-authority" ] \ + || fail "$posture: published merge left its authority record behind" + done + + pass "queued merges retain yolo and away-grant after captain return" +} + +test_merged_poll_row_names_no_authority_when_no_record_grants_one() { + local dir state url + url=https://github.com/o/r/pull/1 + + dir=$(make_case queued-merge-authority-attended) + state="$dir/home/state" + write_task_meta "$dir" task-a + run_check_entry "$dir" task-a "$url" >/dev/null 2> "$dir/seed.err" \ + || fail "attended: could not arm the merge poll" + queue_merge "$dir" "$url" + run_merged_poll_cycle "$dir" + [ "$(merged_ledger_row "$state" task-a)" = "check: merge landed: task-a $url" ] \ + || fail "attended queued merge was tagged: $(merged_ledger_row "$state" task-a)" + + dir=$(make_case merged-poll-authority-external) + state="$dir/home/state" + write_poll_meta "$state" task-a "$url" yolo=on + write_away_record "$dir" + seed_canonical_poll "$dir" task-a "$url" + run_merged_poll_cycle "$dir" + [ "$(merged_ledger_row "$state" task-a)" = "check: merge landed: task-a $url external" ] \ + || fail "external merge was attributed from live away posture: $(merged_ledger_row "$state" task-a)" + assert_poll_absent "$state" task-a + + pass "poll distinguishes attended authorization from external landing" +} + +test_authority_persistence_refuses_rebound_metadata() { + local dir state url_a url_b rc + url_a=https://github.com/o/r/pull/1 + url_b=https://github.com/o/r/pull/2 + dir=$(make_case merge-authority-rebound-metadata) + state="$dir/home/state" + write_task_meta "$dir" task-a + run_check_entry "$dir" task-a "$url_a" >/dev/null 2> "$dir/seed.err" \ + || fail "rebind: could not arm the original poll" + cat > "$dir/rebind.sh" <<SH +#!/usr/bin/env bash +"$PR_CHECK" task-a "$url_b" >/dev/null +SH + chmod +x "$dir/rebind.sh" + set +e + FM_TEST_GH_MERGE_HOOK="$dir/rebind.sh" \ + FM_TEST_GH_GRAPHQL_STATE=OPEN FM_TEST_GH_GRAPHQL_MERGED=false \ + FM_TEST_GH_GRAPHQL_QUEUED=true \ + run_merge_entry "$dir" task-a "$url_a" > "$dir/merge.out" 2> "$dir/merge.err" + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "rebind: accepted merge persisted against rebound metadata" + grep -qxF "pr=$url_b" "$state/task-a.meta" \ + || fail "rebind: merge hook did not replace the canonical identity" + [ ! -e "$state/task-a.merge-authority" ] \ + || fail "rebind: authority was published for the wrong canonical identity" + pass "accepted merge authority refuses rebound task metadata" +} + +test_authority_persists_before_control_unlock() { + local dir state url + url=https://github.com/o/r/pull/1 + dir=$(make_case merge-authority-control-lock) + state="$dir/home/state" + write_task_meta "$dir" task-a + run_check_entry "$dir" task-a "$url" >/dev/null 2> "$dir/seed.err" \ + || fail "control lock: could not arm the merge poll" + cat > "$dir/fakebin/mv" <<'SH' +#!/usr/bin/env bash +case " $* " in + *"task-a.merge-authority "*) + [ -d "$FM_TEST_CONTROL_LOCK" ] || exit 91 + ;; +esac +exec "$FM_TEST_REAL_MV" "$@" +SH + chmod +x "$dir/fakebin/mv" + FM_TEST_CONTROL_LOCK="$state/.control-task-a.lock" FM_TEST_REAL_MV="$REAL_MV" \ + queue_merge "$dir" "$url" + pass "accepted merge authority persists under the lifecycle lock" +} + +test_teardown_cannot_race_authority_consumption() { + local dir state url watcher_pid rc i + url=https://github.com/o/r/pull/1 + dir=$(make_case merge-authority-teardown-race) + state="$dir/home/state" + fm_write_meta "$state/task-a.meta" \ + 'window=firstmate:fm-task-a' \ + 'endpoint_task_id=task-a' \ + "worktree=$dir/wt" \ + "project=$dir/project" \ + 'kind=ship' \ + 'mode=local-only' \ + 'yolo=on' + write_away_record "$dir" + run_check_entry "$dir" task-a "$url" >/dev/null 2> "$dir/seed.err" \ + || fail "teardown race: could not arm the merge poll" + queue_merge "$dir" "$url" + archive_away_record "$dir" + FM_TEST_GH_STATE_STARTED="$dir/poll-started" FM_TEST_GH_STATE=MERGED \ + FM_TEST_GH_SLEEP=0.5 FM_TEST_CHECK_TIMEOUT=3 \ + run_watcher_bounded "$dir/home" "$dir/fakebin" \ + > "$dir/watch.out" 2> "$dir/watch.err" & + watcher_pid=$! + i=0 + while [ ! -e "$dir/poll-started" ]; do + sleep 0.01 + i=$((i + 1)) + if [ "$i" -ge 500 ]; then + kill "$watcher_pid" 2>/dev/null || true + wait "$watcher_pid" 2>/dev/null || true + fail "teardown race: watcher did not begin its validated poll" + fi + done + set +e + FM_HOME="$dir/home" FM_ROOT_OVERRIDE="$ROOT" PATH="$dir/fakebin:$BASE_PATH" \ + "$TEARDOWN" task-a --force > "$dir/teardown.out" 2> "$dir/teardown.err" + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "teardown race: cleanup crossed the active poll transaction" + [ -f "$state/task-a.merge-authority" ] \ + || fail "teardown race: refused cleanup removed persisted authority" + rc=0 + wait "$watcher_pid" || rc=$? + [ "$rc" -eq 0 ] || fail "teardown race: watcher failed with $rc: $(cat "$dir/watch.err")" + [ "$(merged_ledger_row "$state" task-a)" = "check: merge landed: task-a $url yolo" ] \ + || fail "teardown race: concurrent cleanup downgraded the merge authority" + pass "teardown cannot race merged-poll authority consumption" +} + +test_authority_retirement_preserves_replacement() { + local dir state url_a url_b rc i + url_a=https://github.com/o/r/pull/1 + url_b=https://github.com/o/r/pull/2 + dir=$(make_case merge-authority-retirement-replacement) + state="$dir/home/state" + write_task_meta "$dir" task-a + run_check_entry "$dir" task-a "$url_a" >/dev/null 2> "$dir/seed.err" \ + || fail "replacement: could not arm the original poll" + queue_merge "$dir" "$url_a" + cat > "$dir/replace-authority.sh" <<SH +#!/usr/bin/env bash +"$PR_CHECK" task-a "$url_b" >/dev/null +( + FM_TEST_GH_GRAPHQL_STATE=OPEN FM_TEST_GH_GRAPHQL_MERGED=false \\ + FM_TEST_GH_GRAPHQL_QUEUED=true \\ + "$PR_MERGE" task-a "$url_b" > "$dir/replacement-merge.out" 2> "$dir/replacement-merge.err" + printf '%s\n' \$? > "$dir/replacement-merge.rc" +) & +SH + chmod +x "$dir/replace-authority.sh" + cat > "$dir/fakebin/mv" <<'SH' +#!/usr/bin/env bash +"$FM_TEST_REAL_MV" "$@" || exit $? +case " $* " in + *"task-a.pr-poll-merge-notified "*) + if [ ! -e "$FM_TEST_REPLACEMENT_RAN" ]; then + : > "$FM_TEST_REPLACEMENT_RAN" + "$FM_TEST_REPLACEMENT_SCRIPT" + fi + ;; +esac +SH + chmod +x "$dir/fakebin/mv" + add_stop_custom_check "$dir" + set +e + FM_TEST_REAL_MV="$REAL_MV" FM_TEST_REPLACEMENT_RAN="$dir/replacement-ran" \ + FM_TEST_REPLACEMENT_SCRIPT="$dir/replace-authority.sh" \ + FM_TEST_GH_STATE=MERGED run_watcher_bounded "$dir/home" "$dir/fakebin" \ + > "$dir/watch-a.out" 2> "$dir/watch-a.err" + rc=$? + set -e + [ "$rc" -eq 0 ] || fail "replacement: original poll failed: $(cat "$dir/watch-a.err")" + i=0 + while [ ! -e "$dir/replacement-merge.rc" ]; do + sleep 0.01 + i=$((i + 1)) + [ "$i" -lt 200 ] || fail "replacement: serialized replacement merge did not finish" + done + [ "$(cat "$dir/replacement-merge.rc")" -eq 0 ] \ + || fail "replacement: serialized replacement merge failed: $(cat "$dir/replacement-merge.err")" + [ -f "$state/task-a.merge-authority" ] \ + || fail "replacement: original poll retirement deleted the replacement authority" + grep -qxF "pr=$url_b" "$state/task-a.meta" \ + || fail "replacement: replacement poll was not armed" + ack_watcher_cycle "$state" || fail "replacement: could not acknowledge the original wake" + rm -f "$dir/fakebin/mv" "$state/.last-check" + run_merged_poll_cycle "$dir" + awk -F'\t' -v expected="check: merge landed: task-a $url_b" \ + '$5 == expected { found=1 } END { exit !found }' "$state/.wake-queue" \ + || fail "replacement: replacement merge lost its attended authority" + pass "poll retirement preserves a replacement authority record" +} + test_parser_matrix test_gitlab_merge_watch test_merged_poll_retires_once test_merged_poll_reregistration_after_notification_is_absorbed test_merged_poll_retries_a_failed_upward_report test_self_merge_and_poll_publish_one_outcome +test_merged_poll_row_carries_the_merge_authority +test_merged_poll_row_names_no_authority_when_no_record_grants_one +test_authority_persistence_refuses_rebound_metadata +test_authority_persists_before_control_unlock +test_teardown_cannot_race_authority_consumption +test_authority_retirement_preserves_replacement test_merged_poll_reports_upward_from_a_secondmate_home_once test_different_merged_pr_for_same_task_is_not_absorbed test_persistent_secondmate_retirement_is_poll_only From 83c63cc6aa79d1cf913bb5ea988a75f867ffa322 Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Sat, 12 Sep 2026 01:05:15 -0700 Subject: [PATCH 19/31] ci: supersede superseded PR CI and bound unbounded jobs (#4281) The 2026-09-12 Actions starvation incident found firstmate CI with no concurrency deduplication, so every superseded PR head kept its full 13-job fan-out, and four jobs with no timeout at all. Add per-PR supersession keyed on the PR number for pull_request events and on the unique run id for push events, cancelling only pull_request runs, so a new PR head replaces its own in-flight CI while every main push keeps its own group and is never cancelled. Add hang tripwires to the four previously unbounded jobs: 25 minutes for lint (measured at 14-16 minutes) and 5 minutes each for the coverage guard, the timing aggregate, and the repo invariants. Measured lane bounds are unchanged. tests/fm-ci-workflow.test.sh resolves the workflow's concurrency expressions against simulated pull_request and push contexts and holds every job's finite timeout. --- .github/workflows/ci.yml | 22 +++++ CONTRIBUTING.md | 1 + tests/fm-ci-workflow.test.sh | 159 +++++++++++++++++++++++++++++++++++ 3 files changed, 182 insertions(+) create mode 100755 tests/fm-ci-workflow.test.sh diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 5dcac6dff5a..24ade79a846 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -9,10 +9,26 @@ on: permissions: contents: read +# Per-PR supersession: a new push to the same PR replaces that PR's in-flight +# CI instead of letting superseded heads keep 13 jobs of hosted-runner work. +# The group uses the PR number for pull_request events, so every run of one PR +# shares a group, and falls back to the unique run id for push events, so each +# main push gets its own group and is never cancelled. Cancellation is likewise +# limited to pull_request events. Evidence and rationale: the September 12 +# Actions starvation report, section 3 "Workflow mechanics to ship first". +# The compliance workflow deliberately keeps its own event-specific groups; do +# not collapse it onto this simpler shape. +concurrency: + group: ci-${{ github.workflow }}-${{ github.event_name }}-${{ github.event.pull_request.number || github.run_id }} + cancel-in-progress: ${{ github.event_name == 'pull_request' }} + jobs: lint: name: Lint runs-on: ubuntu-latest + # Hang tripwire only: lint executions measured at 14-16 minutes in the + # September 12 starvation report, so this leaves deliberate margin. + timeout-minutes: 25 steps: - uses: actions/checkout@v6 - name: Install pinned ShellCheck @@ -37,6 +53,8 @@ jobs: test-coverage: name: Test coverage guard runs-on: ubuntu-latest + # Hang tripwire: the coverage guard is a seconds-long local computation. + timeout-minutes: 5 steps: - uses: actions/checkout@v6 - name: Prove complete regression partition @@ -335,6 +353,8 @@ jobs: tests-timing-aggregate: name: Behavior timing aggregate runs-on: ubuntu-latest + # Hang tripwire: aggregation is seconds of work over lane artifacts. + timeout-minutes: 5 needs: - tests-portable-parallel-1 - tests-portable-parallel-2 @@ -433,6 +453,8 @@ jobs: invariants: name: Repo invariants runs-on: ubuntu-latest + # Hang tripwire: the invariant checks are seconds-long file comparisons. + timeout-minutes: 5 steps: - uses: actions/checkout@v6 - name: Compatibility pointers must stay intact diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index c1c3d15ee91..87317bd5c45 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -107,6 +107,7 @@ Its header and `--help` own the flags, family labels, lanes, and changed-file ma Portable shard balance evidence lives in `docs/fm-test-portable-shards.md`. Family selection is the ordinary local path; `--all` is deliberate full regression only. CI owns broad regression across required portable parallel shards, the portable serial lane's separate-runner shards, the Herdr lane, lint, invariants, the coverage guard, and stock macOS Bash compatibility in [`.github/workflows/ci.yml`](.github/workflows/ci.yml). +Pushing a new head to a pull request cancels that pull request's still-running CI so only the current head is validated; pushes to `main` are never cancelled, and the workflow owns that contract and its rationale. Use `bin/fm-test-run.sh --list-lanes` for exact lane names and `--help` for `--jobs` rules and required gate-skip flags when reproducing a lane locally. Leave the `sleep 0.1` cadence in the suites' bounded condition waits alone. Those sleeps look like recoverable overhead - `fm-watch-triage.test.sh` alone issues about 1,900 of them, each paying a flat ~100ms scheduler wake-up penalty on macOS - but they are not overhead added to the clock; they are how a test waits for a subject that only moves on `fm-watch.sh`'s own one-second `FM_POLL` cadence. diff --git a/tests/fm-ci-workflow.test.sh b/tests/fm-ci-workflow.test.sh new file mode 100755 index 00000000000..fd2f7918493 --- /dev/null +++ b/tests/fm-ci-workflow.test.sh @@ -0,0 +1,159 @@ +#!/usr/bin/env bash +# Contract tests for .github/workflows/ci.yml's runner-spend safeguards. +# +# Origin: the 2026-09-12 GitHub Actions starvation incident. firstmate CI had no +# concurrency deduplication, so every superseded PR head kept its full job +# fan-out, and four jobs carried no timeout at all. These tests hold both +# safeguards: PR runs supersede within one PR while main pushes are never +# cancelled, and every CI job carries a finite hang tripwire. +# +# The workflow is parsed as YAML and its concurrency expressions are resolved +# against simulated pull_request and push contexts, so the assertions describe +# what GitHub would do, not how the file happens to be spelled. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +CI_WORKFLOW="$ROOT/.github/workflows/ci.yml" + +assert_present "$CI_WORKFLOW" ".github/workflows/ci.yml is missing" +command -v ruby >/dev/null 2>&1 \ + || fail "ruby is required to parse .github/workflows/ci.yml as YAML" + +# Resolve the workflow's concurrency contract under one simulated event and +# print "<group><TAB><cancel-in-progress>". Only the two expression constructs +# this workflow uses are resolved: an `a || b` fallback and an `==` comparison. +resolve_concurrency() { + local event=$1 pr_number=$2 run_id=$3 + ruby -ryaml -e ' +doc = YAML.load_file(ARGV[0]) +concurrency = doc.fetch("concurrency") +context = { + "github.workflow" => doc.fetch("name"), + "github.event_name" => ARGV[1], + "github.event.pull_request.number" => ARGV[2], + "github.run_id" => ARGV[3], +} + +value = lambda do |token| + token = token.strip + next token[1..-2] if token.start_with?("\x27") && token.end_with?("\x27") + raise "unresolvable context reference: #{token}" unless context.key?(token) + context.fetch(token) +end + +evaluate = lambda do |expression| + expression = expression.strip + if expression.include?("==") + left, right = expression.split("==", 2) + next value.call(left) == value.call(right) ? "true" : "false" + end + resolved = expression.split("||").map { |token| value.call(token) }.find { |v| !v.empty? } + resolved.to_s +end + +interpolate = lambda do |raw| + raw.to_s.gsub(/\$\{\{(.+?)\}\}/) { evaluate.call(Regexp.last_match(1)) } +end + +puts [interpolate.call(concurrency.fetch("group")), + interpolate.call(concurrency.fetch("cancel-in-progress"))].join("\t") +' "$CI_WORKFLOW" "$event" "$pr_number" "$run_id" +} + +job_timeout() { + ruby -ryaml -e ' +puts YAML.load_file(ARGV[0]).fetch("jobs").fetch(ARGV[1]).fetch("timeout-minutes", "none") +' "$CI_WORKFLOW" "$1" +} + +group_of() { printf '%s\n' "$1" | cut -f1; } +cancel_of() { printf '%s\n' "$1" | cut -f2; } + +test_pr_pushes_supersede_within_one_pr() { + local first second + first=$(resolve_concurrency pull_request 108 900001) || fail "could not resolve PR concurrency" + second=$(resolve_concurrency pull_request 108 900002) || fail "could not resolve PR concurrency" + [ "$(group_of "$first")" = "$(group_of "$second")" ] \ + || fail "two runs of one PR must share a concurrency group, got $(group_of "$first") and $(group_of "$second")" + [ "$(cancel_of "$first")" = true ] \ + || fail "PR runs must cancel the in-progress run, got $(cancel_of "$first")" + pass "a newer push to one PR supersedes that PR's in-flight CI" +} + +test_separate_prs_do_not_cancel_each_other() { + local one two + one=$(resolve_concurrency pull_request 108 900001) || fail "could not resolve PR concurrency" + two=$(resolve_concurrency pull_request 109 900003) || fail "could not resolve PR concurrency" + [ "$(group_of "$one")" != "$(group_of "$two")" ] \ + || fail "distinct PRs must not share a concurrency group ($(group_of "$one"))" + pass "distinct PRs get distinct concurrency groups" +} + +test_main_pushes_are_never_cancelled() { + local first second + first=$(resolve_concurrency push '' 900010) || fail "could not resolve push concurrency" + second=$(resolve_concurrency push '' 900011) || fail "could not resolve push concurrency" + [ "$(group_of "$first")" != "$(group_of "$second")" ] \ + || fail "each main push must get its own concurrency group, got $(group_of "$first") twice" + [ "$(cancel_of "$first")" = false ] \ + || fail "push runs must never cancel an in-progress run, got $(cancel_of "$first")" + pass "every main push keeps its own group and is never cancelled" +} + +test_every_job_has_a_finite_timeout() { + local reported + reported=$(ruby -ryaml -e ' +YAML.load_file(ARGV[0]).fetch("jobs").each do |name, job| + timeout = job["timeout-minutes"] + next if timeout.is_a?(Integer) && timeout > 0 + puts "#{name}: #{timeout.inspect}" +end +' "$CI_WORKFLOW") || fail "could not read job timeouts from ci.yml" + [ -z "$reported" ] || fail "these CI jobs have no finite hang tripwire:"$'\n'"$reported" + pass "every ci.yml job carries a finite timeout" +} + +# The four jobs the incident found unbounded, at the report's recommended caps. +test_previously_unbounded_jobs_keep_their_caps() { + local job expected actual + while read -r job expected; do + [ -n "$job" ] || continue + actual=$(job_timeout "$job") || fail "could not read the $job timeout" + [ "$actual" = "$expected" ] \ + || fail "$job timeout must stay $expected minutes, got $actual" + done <<'CAPS' +lint 25 +test-coverage 5 +tests-timing-aggregate 5 +invariants 5 +CAPS + pass "the incident's unbounded jobs keep their recommended caps" +} + +# Cancellation makes an undersized cap costlier: a falsely tripped job now also +# discards a run nobody replaced. These bounds were measured, not guessed. +test_measured_lanes_keep_their_existing_bounds() { + local job expected actual + while read -r job expected; do + [ -n "$job" ] || continue + actual=$(job_timeout "$job") || fail "could not read the $job timeout" + [ "$actual" = "$expected" ] \ + || fail "$job timeout must stay $expected minutes, got $actual" + done <<'CAPS' +tests-portable-parallel-1 10 +tests-portable-parallel-2 10 +tests-portable-serial 30 +tests-herdr 75 +macos-stock-bash 10 +CAPS + pass "the already-measured lane bounds are unchanged" +} + +test_pr_pushes_supersede_within_one_pr +test_separate_prs_do_not_cancel_each_other +test_main_pushes_are_never_cancelled +test_every_job_has_a_finite_timeout +test_previously_unbounded_jobs_keep_their_caps +test_measured_lanes_keep_their_existing_bounds From fa65b5df10745515b4c1a897e8cc1dbfd95309b0 Mon Sep 17 00:00:00 2001 From: nateliuroberts <nate@cipherlab.ai> Date: Sat, 12 Sep 2026 03:41:12 -0700 Subject: [PATCH 20/31] test(watch): gate backlog-hold away-record fixture on tasks-axi (#4288) Every other make_hold_home caller in this file skips when tasks-axi is absent; this test was the one unguarded call, so hosts without tasks-axi hard-fail the fixture build instead of skipping. --- tests/fm-watch-triage.test.sh | 2 ++ 1 file changed, 2 insertions(+) diff --git a/tests/fm-watch-triage.test.sh b/tests/fm-watch-triage.test.sh index c4f3bd428d0..8093b733c7d 100755 --- a/tests/fm-watch-triage.test.sh +++ b/tests/fm-watch-triage.test.sh @@ -4665,6 +4665,8 @@ test_live_captain_held_first_sight_silenced_by_away_record() { test_backlog_hold_never_rechecked_while_away_record_exists() { local dir out capture wakes + command -v tasks-axi >/dev/null 2>&1 \ + || { echo "skip: tasks-axi not found (away-record backlog hold)"; return 0; } dir=$(make_hold_home away-record-backlog-hold 'done: PR https://example.test/pr/9 checks green' hold) \ || fail "could not build the backlog-hold fixture" out="$dir/watch.out"; capture="$dir/pane.txt" From fb19dd9a75f2a9ec0b4c6573e45a0172b4f20f68 Mon Sep 17 00:00:00 2001 From: Jon Roosevelt <jon@arcs.health> Date: Sat, 12 Sep 2026 14:21:09 -0400 Subject: [PATCH 21/31] fix(backlog): bound per-item backlog row reads so a wedged backend cannot blind a session start (#4027) * fix(bin): bound each backlog row read so one wedged backend cannot blind a session start bin/fm-bootstrap.sh's reconcile and close-replay sweeps read the backlog backend once per item through fm_backlog_row_show, and that read was unbounded. A single wedged `tasks-axi show` therefore consumed the whole FM_SESSION_START_TIMEOUT and truncated the digest before the wake queue, supervision instructions, fleet state, and context sections ever printed, leaving the fleet unsupervised with no live watcher. The harm was a blind startup, not a slow one. Bound the read with the existing shared timeout primitive (bin/fm-timeout-lib.sh), so a wedged backend degrades to a loud partial reconcile: the sweep's existing BACKLOG_RECONCILE diagnostic names the item it could not read and the loop continues to the next one. The first bound hit also latches FM_BACKLOG_ROW_SHOW_WEDGED, so a sweep over many items pays one bound rather than one per item and still names every item it skipped, which is what keeps the digest whole on a home carrying a large fleet. The bound holds regardless of any particular tasks-axi install, so it does not depend on the 0.2.5 `show` hang being resolved separately. * fix(bin): set the wedged-backend latch where it survives, and prove it The latch added with the read bound was inert. fm_backlog_row_show runs inside a command substitution in both of its status-capturing callers, so the subshell read the inherited value correctly but its write died with the subshell. Every item still paid a full bound and reported `exceeded`, never `skipped`, which left the large-fleet case the latch existed to cover completely uncovered. Move the write to the two callers that capture the read's status and own the surviving shell, and leave fm_backlog_row_show reading the latch only. Correct the comments that claimed an ownership the function never had. The test that was supposed to cover this asserted only that the second read finished under a generous ceiling, which is true whether or not the latch works. Assert instead that a latched read is strictly faster than one bound and that it reports its own item as skipped, so an inert latch fails the test. * test: cover every item the wedged-backend latch skips The latch assertion exercised a single skipped item, so "every skipped item is still named" was inferred rather than tested. Probe three items instead and assert each skipped one names itself and costs less than a bound. Verified as a real guard by removing both latch writes: the suite then fails on the first skipped item instead of passing. * no-mistakes(review): distinguish backlog read-bound hits from absent rows * no-mistakes(review): preserve read-bound status through the captain verify gates * no-mistakes(review): Preserve backlog read-bound hits through resolve_entry and reconcile instead of spending them as absent rows * no-mistakes(review): Preserve backlog read-bound 124 through migrated-prefix scan and remaining task_show call sites * no-mistakes(document): Document bounded backlog row reads and FM_BACKLOG_ROW_TIMEOUT_SECS * no-mistakes(ci): Fixed all four failing CI checks with one root-cause fix plus one test-heredity fix. (1) bin/fm-captain-hold.sh: task_show carries the row in TASK_SHOW_OUTPUT and emits no stdout, but four call sites still used the stale command-substitution convention show=$(task_show ...), leaving show empty: task_show_or_fail (every captain hold failed with 'did not retain its hold-set stamp' - broke fm-captain-hold-lifecycle in parallel 1 and fm-bearings-board in serial 3), resolve_migrated_entry (migrated-prefix resolution could never match), reconcile-requests (existing rows were refused as absent), and command_open --identity (printed a constant '#0' identity, so fm-watch-triage's re-held captain call inherited the previous call's silence in serial 1). This is also the Greptile P1. Fixed by invoking task_show in the current shell and reading show=$TASK_SHOW_OUTPUT, the convention the other eight call sites already use; read-bound hits still stop loudly by name. (2) tests/fm-backlog-read-bound.test.sh (serial 4, unclassified family): the new e2e half implicitly relied on the author's process tree containing a harness process so fm-lock.sh would grant the fleet lock; on CI runners the lock is refused, the reconcile sweep is skipped, and the final BACKLOG_RECONCILE assertion fails. Reproduced by simulating a CI ancestry via a ps shim, fixed by pinning the lock evidence with the established fake-ps harness fixture pattern from tests/fm-session-start.test.sh. Verified: shellcheck clean; parallel-1, serial-3, and serial-4 lanes fully green locally (failed=0); serial-1 lane green except fm-gemini-harness, which fails only under local Node v26 (comm=node-MainThread); CI's default Node 22 reports comm=node, the branch that test passes on, so it is not a CI failure * no-mistakes(document): Verified bounded backlog read docs accurate across branch --- bin/fm-backlog-transition-lib.sh | 65 ++++- bin/fm-captain-hold.sh | 135 ++++++--- docs/configuration.md | 1 + docs/sessionstart-nudge.md | 3 +- tests/fm-backlog-read-bound.test.sh | 430 ++++++++++++++++++++++++++++ 5 files changed, 594 insertions(+), 40 deletions(-) create mode 100755 tests/fm-backlog-read-bound.test.sh diff --git a/bin/fm-backlog-transition-lib.sh b/bin/fm-backlog-transition-lib.sh index 4442ec3c64f..c116d016b0e 100644 --- a/bin/fm-backlog-transition-lib.sh +++ b/bin/fm-backlog-transition-lib.sh @@ -73,6 +73,17 @@ FM_BACKLOG_ROW_HOLD_KIND= # shellcheck disable=SC2034 # Output global, read by the sourcing caller. FM_BACKLOG_CLOSE_REPLAY_RESULT= +# Bounded execution is fm-timeout-lib.sh's alone; source it rather than +# re-deriving a deadline here. It is stateless, so the memoisation reason this +# library does not source fm-tasks-axi-lib.sh does not apply. +# shellcheck source=bin/fm-timeout-lib.sh disable=SC1091 +. "$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/fm-timeout-lib.sh" + +# Latched when a row read hits its bound. fm_backlog_row_show runs inside a +# command substitution, so the subshell can READ this latch but cannot set it; +# the callers that capture its status own the write. +FM_BACKLOG_ROW_SHOW_WEDGED=0 + # Emit each byte of a value as a decimal number, locale-independently. # Deliberately perl rather than od: the spawn and teardown lifecycle runs under a # curated PATH (tests/fm-teardown.test.sh make_path_without_lsof pins that set) @@ -382,22 +393,64 @@ fm_tasks_axi() { exit 127 } -# Print one row's `tasks-axi show` output (plus stderr); the exit status is -# tasks-axi's. Extra flags (such as --full) are passed through. +# Print one row's `tasks-axi show` output (plus stderr) from the addressing +# fm_backlog_tasks_axi_addressing resolved, with `--file` only for the markdown +# backend. Addressing or backend-resolution errors return before tasks-axi runs; +# otherwise its exit status is preserved. Extra flags (--full) pass through. +# +# Every read is bounded, because a wedged backend read here is what blinds a +# whole session start: bin/fm-bootstrap.sh's reconcile and close-replay sweeps +# call this once per item, and one unbounded read consumes the entire +# FM_SESSION_START_TIMEOUT and truncates the digest before the wake queue, +# supervision instructions, fleet state and context sections ever print. The +# bound turns that into a loud partial reconcile: the caller reports the item it +# could not read and moves to the next one. +# +# A per-item bound alone is not enough on a home carrying a large fleet, because +# N wedged items still cost N bounds and the digest is truncated anyway. So the +# first bound hit latches FM_BACKLOG_ROW_SHOW_WEDGED and every later read in the +# same sweep returns immediately, still naming its own item so nothing is +# silently skipped. This function only READS that latch: it runs inside a +# command substitution, and a write here would die with the subshell, so the +# callers that capture its status set it. The latch is deliberately +# process-wide because these scripts are short-lived and a backend that wedged +# once will wedge again within the same run. fm_backlog_row_show() { # <resolved-data-dir> <id> [flag...] - local data=$1 id=$2 addressing_status + local data=$1 id=$2 out status addressing_status secs=${FM_BACKLOG_ROW_TIMEOUT_SECS:-10} shift 2 + # A non-positive bound is not a bound (fm-timeout-lib.sh), and a padded zero + # such as 00 is still zero, so the digits test alone would let the very read + # this bound exists to prevent back in. Compare arithmetically, tolerating a + # value too large for the shell to compare at all. + case "$secs" in ''|*[!0-9]*) secs=10 ;; esac + [ "$secs" -gt 0 ] 2>/dev/null || secs=10 fm_backlog_tasks_axi_addressing "$data" addressing_status=$? if [ "$addressing_status" -ne 0 ]; then [ -z "${FM_BACKLOG_TRANSITION_ERROR:-}" ] || printf '%s\n' "$FM_BACKLOG_TRANSITION_ERROR" >&2 return "$addressing_status" fi + if [ "$FM_BACKLOG_ROW_SHOW_WEDGED" = 1 ]; then + printf 'tasks-axi show %s skipped: the backlog backend already exceeded its %ss read bound\n' "$id" "$secs" + return 124 + fi if [ -n "$FM_BACKLOG_AXI_FILE" ]; then - (cd "$FM_BACKLOG_AXI_ROOT" 2>/dev/null && fm_tasks_axi show "$id" "$@" --file "$FM_BACKLOG_AXI_FILE" 2>&1) + set -- "$@" --file "$FM_BACKLOG_AXI_FILE" + fi + # shellcheck disable=SC2016 # Expansion is deliberately deferred to the child shell. + out=$(fm_run_timed "$secs" bash -c 'cd "$1" 2>/dev/null || exit 1; shift; exec tasks-axi show "$@"' \ + _ "$FM_BACKLOG_AXI_ROOT" "$id" "$@" 2>&1) + status=$? + # A backend that wrote a header or a progress line before wedging leaves that + # fragment as the first output line, and every caller reads the first line as + # the failure reason. Whatever a timed-out read managed to emit is incomplete + # by definition, so the bound speaks for it instead. + if [ "$status" -eq 124 ]; then + printf 'tasks-axi show %s exceeded its %ss backlog read bound\n' "$id" "$secs" else - (cd "$FM_BACKLOG_AXI_ROOT" 2>/dev/null && fm_tasks_axi show "$id" "$@" 2>&1) + printf '%s\n' "$out" fi + return "$status" } fm_backlog_row_list() { # <resolved-data-dir> [flag...] @@ -436,6 +489,7 @@ fm_backlog_row_probe() { # <data-dir> <id> fi out=$(fm_backlog_row_show "$data" "$id") command_status=$? + [ "$command_status" -ne 124 ] || FM_BACKLOG_ROW_SHOW_WEDGED=1 if [ "$command_status" -ne 0 ]; then if printf '%s\n' "$out" | grep -q '^code: NOT_FOUND$'; then FM_BACKLOG_ROW_RESULT=not_found @@ -559,6 +613,7 @@ fm_backlog_retain() { # <data-dir> <id> [flag...] if [ -n "$deliverable" ]; then out=$(fm_backlog_row_show "$data" "$id" --full) command_status=$? + [ "$command_status" -ne 124 ] || FM_BACKLOG_ROW_SHOW_WEDGED=1 if [ "$command_status" -ne 0 ]; then FM_BACKLOG_TRANSITION_ERROR=$(printf '%s\n' "$out" | sed -n '1p') [ -n "$FM_BACKLOG_TRANSITION_ERROR" ] \ diff --git a/bin/fm-captain-hold.sh b/bin/fm-captain-hold.sh index a9ef07b132a..c3d3a98f67e 100755 --- a/bin/fm-captain-hold.sh +++ b/bin/fm-captain-hold.sh @@ -350,10 +350,38 @@ require_tasks_axi() { || fail "tasks-axi does not expose the captain-hold contract" } -task_show() { # <id> - local data +# Read one row into TASK_SHOW_OUTPUT; a non-zero return means the row is +# absent. A read that could not finish inside its bound is NOT absence, and +# every caller below would otherwise spend it as one - minting a duplicate task, +# skipping a keyed answer, or reporting a task that exists as missing. So the +# bound's own status stops the command instead, loudly and by name, and it +# leaves 124 intact rather than collapsing to fail's 1 so a caller running this +# inside a command substitution can still tell a wedged backend from a +# genuinely unknown id. +TASK_SHOW_OUTPUT= +task_show() { # <id>; sets TASK_SHOW_OUTPUT + local data status=0 reason data=$(fm_backlog_data_absolute "$DATA") || fail "data directory cannot be resolved: $DATA" - fm_backlog_row_show "$data" "$1" --full 2>/dev/null + TASK_SHOW_OUTPUT=$(fm_backlog_row_show "$data" "$1" --full 2>/dev/null) || status=$? + if [ "$status" -eq 124 ]; then + reason=${TASK_SHOW_OUTPUT%%$'\n'*} + printf 'fm-captain-hold: %s\n' \ + "${reason:-tasks-axi show $1 exceeded its backlog read bound}" >&2 + exit 124 + fi + return "$status" +} + +# Read one row into `show`, failing with <absence-message> only when the read +# genuinely failed; a read-bound hit (124) stops the command by name instead. +# task_show must be called in THIS shell, not inside a command substitution: +# it carries the row in TASK_SHOW_OUTPUT, which a subshell cannot hand back. +task_show_or_fail() { # <id> <absence-message>; sets show + task_show "$1" || { + [ "$?" -ne 124 ] || fail "the backlog backend exceeded its read bound reading $1" + fail "$2" + } + show=$TASK_SHOW_OUTPUT } show_field() { # <show-output> <field> @@ -388,7 +416,7 @@ show_field_value() { # <show-output> <field> origin_exists_here() { # <origin-id> [ -f "$STATE/$1.meta" ] && return 0 [ -f "$DATA/$1/report.md" ] && return 0 - task_show "$1" >/dev/null 2>&1 + task_show "$1" } list_has_key() { # <comma-list> <key> @@ -493,7 +521,8 @@ resolution_block() { # <mode> # surviving even when a date gate has expired) or a recorded captain answer. verify_hold_durable() { # <task-id> local id=$1 show state hold_kind body - show=$(task_show "$id") || fail "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" + task_show "$id" || fail "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" + show=$TASK_SHOW_OUTPUT state=$(show_field "$show" state) hold_kind=$(show_field_value "$show" hold_kind) body=$(show_field "$show" body) @@ -678,7 +707,13 @@ resolve_migrated_entry() { # <origin-or-empty> <entry> *-) prefixed="$prefix$candidate" ;; *) prefixed="$prefix-$candidate" ;; esac - show=$(task_show "$prefixed" 2>/dev/null) || continue + # Same shell rule as task_show_or_fail: the row is read out of + # TASK_SHOW_OUTPUT, so the read cannot sit inside a command substitution. + task_show "$prefixed" 2>/dev/null || { + [ "$?" -ne 124 ] || return 124 + continue + } + show=$TASK_SHOW_OUTPUT [ "$(show_field_value "$show" hold_kind)" = captain ] || continue prefixed_matches="${prefixed_matches}${prefixed_matches:+$NL_SEP}$prefixed" done @@ -699,13 +734,13 @@ resolve_migrated_entry() { # <origin-or-empty> <entry> # migrated-prefix, so a caller can record which evidence carried the attestation. resolve_entry() { # <origin-or-empty> <entry>; prints "<id> <how>" or fails local origin=$1 entry=$2 legacy migrated rc - if task_show "$entry" >/dev/null 2>&1; then + if task_show "$entry"; then printf '%s exact' "$entry" return 0 fi if [ -n "$origin" ] && [ "$origin" != "$BINDING_ANY" ]; then legacy=$(legacy_hold_id "$origin" "$entry") - if task_show "$legacy" >/dev/null 2>&1; then + if task_show "$legacy"; then printf '%s legacy' "$legacy" return 0 fi @@ -715,6 +750,7 @@ resolve_entry() { # <origin-or-empty> <entry>; prints "<id> <how>" or fails case "$rc" in 0) printf '%s' "$migrated"; return 0 ;; 2) return 2 ;; + 124) return 124 ;; esac if [ -n "$origin" ] && [ "$origin" != "$BINDING_ANY" ]; then legacy=$(legacy_hold_id "$origin" "$entry") @@ -763,6 +799,24 @@ write_hold_set_stamp() { # <task-id> <shown-body> <timestamp> <preserve-existin rm -f -- "$tmp" } +# Resolve one entry and verify the row it names is durably captain-held. A +# resolution failure that is not the read bound keeps resolve_entry's own +# status - its stderr already named the entry; 124 means the backend never +# answered, which is not the same as an unknown entry and must not be spent +# as absence. On success prints "<id> <how>" so the caller can keep the +# attestation evidence. +verify_entry_durable() { # <origin-or-empty> <entry>; prints "<id> <how>" + local origin=$1 entry=$2 resolved resolve_status=0 + resolved=$(resolve_entry "$origin" "$entry") || resolve_status=$? + if [ "$resolve_status" -ne 0 ]; then + [ "$resolve_status" -ne 124 ] \ + || fail "the backlog backend exceeded its read bound resolving $entry" + exit "$resolve_status" + fi + printf '%s\n' "$resolved" + verify_hold_durable "${resolved%% *}" +} + command_hold() { local id=${1:-} title='' reason='' repo='' origin='' until='' show state existing_title body='' hold_kind hold_set occurrence local existing_hold_kind='' existing_held='' preserve_hold_set=0 @@ -798,7 +852,8 @@ command_hold() { esac acquire_task_control_lock "$id" require_tasks_axi - if show=$(task_show "$id"); then + if task_show "$id"; then + show=$TASK_SHOW_OUTPUT state=$(show_field "$show" state) [ "$state" != "done" ] \ || fail "task $id is already closed; a new captain call needs its own task" @@ -833,9 +888,9 @@ command_hold() { # Publish the timestamp before the captain-hold annotation. A concurrent # snapshot may see the harmless stamp by itself, but can never see a newly # held task without the timestamp that defines this hold lifecycle's age. - show=$(task_show "$id") || fail "task $id disappeared before recording its hold-set stamp" + task_show_or_fail "$id" "task $id disappeared before recording its hold-set stamp" write_hold_set_stamp "$id" "$(show_field "$show" body)" "$hold_set" "$preserve_hold_set" - show=$(task_show "$id") || fail "task $id disappeared while recording its hold-set stamp" + task_show_or_fail "$id" "task $id disappeared while recording its hold-set stamp" [ -n "$(body_hold_set_timestamp "$(show_field_value "$show" body)")" ] \ || fail "task $id did not retain its hold-set stamp" if [ -n "$until" ]; then @@ -845,7 +900,8 @@ command_hold() { tasks_axi hold "$id" --reason "$reason" --kind captain >/dev/null \ || fail "could not hold task $id for the captain" fi - show=$(task_show "$id") || fail "task $id disappeared while holding it" + task_show "$id" || fail "task $id disappeared while holding it" + show=$TASK_SHOW_OUTPUT hold_kind=$(show_field_value "$show" hold_kind) [ "$hold_kind" = captain ] || fail "task $id did not retain its captain hold" occurrence=$(( $(resolution_record_count "$(show_field "$show" body)") + 1 )) @@ -922,7 +978,7 @@ close_answered() { # <task-id> <release-0-or-1> remove_interrupted_answer_stamp() { # <task-id> local id=$1 show body existing tmp - show=$(task_show "$id") || fail "task $id disappeared after closing" + task_show_or_fail "$id" "task $id disappeared after closing" body=$(decode_shown_value "$(show_field "$show" body)") \ || fail "could not decode the closed body for $id" existing=$(body_hold_set_timestamp "$body") @@ -958,7 +1014,8 @@ command_answer() { load_decision "$decision_file" acquire_task_control_lock "$id" require_tasks_axi - show=$(task_show "$id") || fail "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" + task_show "$id" || fail "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" + show=$TASK_SHOW_OUTPUT state=$(show_field "$show" state) hold_kind=$(show_field_value "$show" hold_kind) body=$(show_field "$show" body) @@ -994,7 +1051,8 @@ command_answer() { || fail "task $id was never held for the captain; nothing to record an answer on" write_resolution_record "$id" repaired "$body" remove_interrupted_answer_stamp "$id" - show=$(task_show "$id") || fail "task $id disappeared while recording the answer" + task_show "$id" || fail "task $id disappeared while recording the answer" + show=$TASK_SHOW_OUTPUT [ "$(show_field "$show" state)" = "done" ] || fail "recording the answer reopened closed task $id" body_has_resolution_record "$(show_field "$show" body)" \ || fail "captain-held task $id did not retain its durable resolution record" @@ -1031,7 +1089,8 @@ command_answer() { fail "could not close answered captain-held task $id" fi remove_interrupted_answer_stamp "$id" - show=$(task_show "$id") || fail "task $id disappeared after closing" + task_show "$id" || fail "task $id disappeared after closing" + show=$TASK_SHOW_OUTPUT body_has_resolution_record "$(show_field "$show" body)" \ || fail "captain-held task $id did not retain its durable resolution record" publish_parent_resolution_then_retire "$id" "$occurrence" "$outcome" @@ -1152,6 +1211,7 @@ sanitize_reconcile_provenance() { command_answers() { local origin='' source='' row rest key answer label mode id show state hold_kind body digest legacy_digest legacy_key local recorded_digest recorded_mode occurrence tmp err closed=0 skipped=0 reason release_flag tab=$'\t' + local resolve_rc while [ "$#" -gt 0 ]; do case "$1" in --source) shift; source=${1:-} ;; @@ -1212,6 +1272,12 @@ command_answers() { continue fi if [ "$resolve_rc" -ne 0 ]; then + # resolve_entry runs in a command substitution, so task_show's exit + # cannot stop this loop; only its status crosses back. 124 means the + # backend never answered, which is not the same as an unknown key and + # must not be spent as a skip. + [ "$resolve_rc" -ne 124 ] \ + || fail "the backlog backend exceeded its read bound resolving $key" printf 'skipped: %s (no captain-held task with that id)\n' "$key" skipped=$((skipped + 1)) continue @@ -1231,7 +1297,8 @@ command_answers() { if [ -n "$legacy_key" ]; then legacy_digest=$(sha256_text "$(legacy_keyed_decision_text "$source" "$legacy_key" "$answer" "$label")") fi - show=$(task_show "$id") || { printf 'skipped: %s (absent)\n' "$id"; skipped=$((skipped + 1)); continue; } + task_show "$id" || { printf 'skipped: %s (absent)\n' "$id"; skipped=$((skipped + 1)); continue; } + show=$TASK_SHOW_OUTPUT state=$(show_field "$show" state) hold_kind=$(show_field_value "$show" hold_kind) body=$(show_field "$show" body) @@ -1347,7 +1414,7 @@ publish_parent_resolution_then_retire() { # <task-id> <occurrence> <note> } command_reconcile_requests() { - local source_id='' source='' origin row id note provenance show created=0 skipped=0 tab=$'\t' + local source_id='' source='' origin row id note provenance show show_status=0 created=0 skipped=0 tab=$'\t' while [ "$#" -gt 0 ]; do case "$1" in --source-id) shift; source_id=${1:-} ;; @@ -1372,7 +1439,13 @@ command_reconcile_requests() { [ "${#id}" -le 128 ] \ || { printf 'refused: %s (task id is too long)\n' "$id"; skipped=$((skipped + 1)); continue; } acquire_task_control_lock "$id" - show=$(task_show "$id") || true + show_status=0 + show='' + task_show "$id" || show_status=$? + [ "$show_status" -ne 0 ] || show=$TASK_SHOW_OUTPUT + if [ "$show_status" -eq 124 ]; then + fail "the backlog backend exceeded its read bound reading $id" + fi if [ -z "$show" ]; then printf 'refused: %s (absent)\n' "$id" skipped=$((skipped + 1)) @@ -1447,7 +1520,7 @@ reconcile_close() { reconcile_request_read "$id" \ || fail "task $id has no pending board-created reconcile request" require_tasks_axi - show=$(task_show "$id") || fail "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" + task_show_or_fail "$id" "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" state=$(show_field "$show" state) hold_kind=$(show_field_value "$show" hold_kind) body=$(show_field "$show" body) @@ -1483,7 +1556,7 @@ reconcile_close() { fi close_answered "$id" 0 || fail "could not close reconciled captain-held task $id" remove_interrupted_answer_stamp "$id" - show=$(task_show "$id") || fail "task $id disappeared after closing" + task_show_or_fail "$id" "task $id disappeared after closing" body_has_resolution_record "$(show_field "$show" body)" \ || fail "captain-held task $id did not retain its durable resolution record" publish_parent_hold "$id" "$occurrence" resolved reconciled @@ -1519,7 +1592,7 @@ reconcile_note() { require_tasks_axi command_open "$id" \ || fail "task $id is not an open captain call; a note cannot keep a closed call open" - show=$(task_show "$id") || fail "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" + task_show_or_fail "$id" "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" body=$(decode_shown_value "$(show_field "$show" body)") \ || fail "could not decode the existing body for $id" note_digest=$(sha256_text "$note") @@ -1584,13 +1657,9 @@ command_complete() { if [ -n "$keys" ]; then while IFS= read -r entry; do [ -n "$entry" ] || continue - if ! resolved=$(resolve_entry "$origin" "$entry"); then - # resolve_entry has already refused on stderr naming the entry. - exit 1 - fi + resolved=$(verify_entry_durable "$origin" "$entry") || exit $? resolved_how=${resolved##* } resolved=${resolved%% *} - verify_hold_durable "$resolved" if [ "$resolved_how" = migrated-prefix ]; then attested_by_prefix="${attested_by_prefix}${attested_by_prefix:+ }$entry=$resolved" fi @@ -1648,11 +1717,7 @@ command_verify() { if [ -n "$keys" ]; then while IFS= read -r entry; do [ -n "$entry" ] || continue - if ! resolved=$(resolve_entry "$origin" "$entry"); then - # resolve_entry has already refused on stderr naming the entry. - exit 1 - fi - verify_hold_durable "${resolved%% *}" + verify_entry_durable "$origin" "$entry" >/dev/null done <<EOF $(printf '%s\n' "$keys" | tr ',' '\n') EOF @@ -1770,7 +1835,8 @@ command_diverged() { while IFS= read -r key; do list_has_line "$tokens" "$key" || continue [ "$(status_key_closing_verb "$f" "$key")" = "$resolve" ] || continue - show=$(task_show "$id") || continue + task_show "$id" || continue + show=$TASK_SHOW_OUTPUT [ "$(show_field "$show" state)" != "done" ] || continue [ "$(show_field_value "$show" hold_kind)" = captain ] || continue # The title is the only free-text field here, and the report is @@ -1838,10 +1904,11 @@ command_open() { # <task-id> [--identity] [--distinguish-absent] state=${FM_BACKLOG_ROW_STATE%% *} if [ "$state" != "done" ] && [ "$FM_BACKLOG_ROW_HOLD_KIND" = captain ]; then if [ "$identity" -eq 1 ]; then - show=$(task_show "$id") || { + task_show "$id" || { printf 'fm-captain-hold: captain call %s is open but its record could not be read\n' "$id" >&2 exit 2 } + show=$TASK_SHOW_OUTPUT shown_body=$(show_field "$show" body) printf '%s#%s\n' \ "$(body_hold_set_timestamp "$(decode_shown_value "$shown_body")")" \ diff --git a/docs/configuration.md b/docs/configuration.md index 3a9ef6e0b43..b6ad202c6fa 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -976,6 +976,7 @@ FM_ZELLIJ_SESSION=firstmate # zellij-only: named session for normal backend ops CMUX_SOCKET_PASSWORD= # cmux-only: socket password fallback when config/cmux-socket-password is absent (docs/cmux-backend.md) FM_SESSION_START_STATUS_TAIL=5 # state/*.status lines printed per task in the session-start digest; each line is capped by bin/fm-line-cap-lib.sh FM_SESSION_START_QUEUED_LIMIT=20 # plain queued backlog rows in the session-start digest; in-flight, held, and blocked rows are never bounded and done rows are never listed +FM_BACKLOG_ROW_TIMEOUT_SECS=10 # seconds bounding each backlog row read (bin/fm-backlog-transition-lib.sh); nonpositive or invalid values fall back to 10; the first bound hit latches the sweep so later reads return immediately, each still naming its own item FM_BOOTSTRAP_DETECT_ONLY=0 # internal/read-only session-start mode: skip bootstrap's mutating sweeps and print advisory TANGLE wording FM_BOOTSTRAP_NETWORK=all # internal session-start phase split: all, skip (local steps only), or only (network steps only); see bin/fm-bootstrap.sh FM_STARTUP_NETWORK_TIMEOUT=120 # seconds bounding the deferred inactive-outcome scan plus network checks; hitting it prints an actionable NETWORK_CHECKS line diff --git a/docs/sessionstart-nudge.md b/docs/sessionstart-nudge.md index 68f9b9c4ecb..17d44c93c44 100644 --- a/docs/sessionstart-nudge.md +++ b/docs/sessionstart-nudge.md @@ -42,7 +42,8 @@ On a run-tier harness the nudge cannot also fire: `resume`, `reload`, and `fork` The run tier blocks either hook-driven session initialization or Pi's first provider preflight while the digest runs, so `bin/fm-session-start.sh` bounds itself rather than betting on an unbounded prerequisite. The digest makes no external-network call at all: every one it owes runs off the blocking path in the separately bounded deferred stage owned by `bin/fm-startup-network.sh`, so an unreachable host can no longer consume this budget. -What remains is still not individually bounded - tool version probes, the backlog listing, and the per-task endpoint reads are all local but unbounded subprocesses - so the whole digest runs as one bounded child, default 120s via `FM_SESSION_START_TIMEOUT`. +Tool version probes, the backlog listing, and the per-task endpoint reads remain local but unbounded subprocesses, so the whole digest still runs as one bounded child, default 120s via `FM_SESSION_START_TIMEOUT`. +The per-item backlog row reads inside bootstrap's reconcile and close-replay sweeps are the exception: each is bounded by `FM_BACKLOG_ROW_TIMEOUT_SECS` (default 10s) through `bin/fm-backlog-transition-lib.sh`, and the first bound hit latches the sweep so later reads return immediately while still naming their own item. The shared timeout owner falls back to a pure-Bash process-group watchdog when timeout, gtimeout, and perl are unavailable, so no supported host runs the digest unbounded. Because the child streams into the native transport as it runs, everything emitted before the bound was hit is retained for delivery; the parent then prints a `STARTUP TRUNCATED` banner naming the stage that did not finish and the stages that were therefore never emitted, and still exits 0. The registered hook timeouts sit above that budget so the harness never preempts the banner. diff --git a/tests/fm-backlog-read-bound.test.sh b/tests/fm-backlog-read-bound.test.sh new file mode 100755 index 00000000000..726ca052455 --- /dev/null +++ b/tests/fm-backlog-read-bound.test.sh @@ -0,0 +1,430 @@ +#!/usr/bin/env bash +# tests/fm-backlog-read-bound.test.sh - behavior tests for the per-item bound on +# bin/fm-backlog-transition-lib.sh's backlog row read. +# +# The defect this pins: bin/fm-bootstrap.sh's reconcile and close-replay sweeps +# read the backlog backend once per item, and an unbounded read of a wedged +# backend consumed the whole FM_SESSION_START_TIMEOUT. The digest was then +# truncated before the wake queue, supervision instructions, fleet state, and +# context sections printed, leaving a whole fleet unsupervised. +# +# Both halves are proved here: +# - a deliberately hanging `tasks-axi show` cannot exceed the per-item bound, +# and the failure names the item it could not read +# - a session start against that same wedged backend still completes end to +# end, with every digest section present and a loud partial reconcile +# +# The bound must hold on its own, independent of any particular tasks-axi +# install, so the fake here simply never returns. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +BASE_PATH=${FM_TEST_BASE_PATH:-/usr/bin:/bin:/usr/sbin:/sbin} +TMP_ROOT=$(fm_test_tmproot fm-backlog-read-bound-tests) +trap fm_test_cleanup EXIT + +BOUND_SECS=2 +# Generous enough that a slow CI box never flakes, far below the unbounded hang +# (300s per read) and below the session-start budget the defect consumed. +BOUND_CEILING=30 + +# A backend whose `show` never returns. Everything the compatibility gate and the +# startup listing need still answers promptly, so the only thing under test is +# the read that hangs. +make_hanging_tasks_axi() { # <fakebin> + local fakebin=$1 + cat > "$fakebin/tasks-axi" <<'SH' +#!/usr/bin/env bash +set -u +case "${1:-}" in + --version) printf '%s\n' '0.2.5'; exit 0 ;; + update) + [ "${2:-}" = --help ] || exit 0 + printf '%s\n' 'usage: tasks-axi update <id> [flags]' ' --body-file <path>' ' --archive-body' + exit 0 + ;; + mv) + [ "${2:-}" = --help ] || exit 0 + printf '%s\n' 'usage: tasks-axi mv <id> [<id>...] --to <path-or-dir>' + exit 0 + ;; + show) + # A real backend rejects an unusable id promptly instead of wedging, which + # is what makes a dropped 124 surface as "absent" rather than as a bound. + if [ -z "${2:-}" ]; then + printf 'code: NOT_FOUND\n' >&2 + exit 1 + fi + # The wedge under test: a read that never returns. + sleep 300 + exit 0 + ;; + hold) + [ "${2:-}" = --help ] || exit 0 + printf '%s\n' 'usage: tasks-axi hold <id> [flags]' ' --kind captain' ' --until <date>' + exit 0 + ;; + add) + # Recorded, never silent: creating a row that already exists is the damage a + # timed-out read must never be spent on. + [ -z "${FM_TEST_TASKS_AXI_ADD_LOG:-}" ] || printf '%s\n' "$*" >> "$FM_TEST_TASKS_AXI_ADD_LOG" + exit 0 + ;; + list) + printf 'count: 0\n' + printf 'tasks[0]{id,state,kind,repo,title,blocked_by,hold_kind,hold_reason}:\n' + exit 0 + ;; +esac +exit 0 +SH + chmod +x "$fakebin/tasks-axi" +} + +elapsed_since() { # <start-epoch> + local now + now=$(date +%s) + printf '%s\n' "$((now - $1))" +} + +# --- half one: the per-item bound holds ------------------------------------- + +UNIT="$TMP_ROOT/unit" +UNIT_FAKEBIN=$(fm_fakebin "$UNIT") +mkdir -p "$UNIT/data" +make_hanging_tasks_axi "$UNIT_FAKEBIN" +printf '# Backlog\n' > "$UNIT/data/backlog.md" + +# Three items, so "every skipped item is still named" is actually exercised +# rather than inferred from a single skip. +PROBE_OUT="$UNIT/probe.out" +PATH="$UNIT_FAKEBIN:$BASE_PATH" FM_BACKLOG_ROW_TIMEOUT_SECS="$BOUND_SECS" \ + bash -c ' + set -u + . "$1/bin/fm-tasks-axi-lib.sh" + . "$1/bin/fm-backlog-transition-lib.sh" + for id in wedged-one wedged-two wedged-three; do + start=$(date +%s) + fm_backlog_row_probe "$2" "$id" && printf "unexpected-success\n" + printf "elapsed:%s=%s\n" "$id" "$(( $(date +%s) - start ))" + printf "error:%s=%s\n" "$id" "$FM_BACKLOG_ROW_ERROR" + done + ' _ "$ROOT" "$UNIT/data" > "$PROBE_OUT" 2>&1 + +probe_elapsed() { # <id> + sed -n "s/^elapsed:$1=//p" "$PROBE_OUT" +} + +probe_error() { # <id> + sed -n "s/^error:$1=//p" "$PROBE_OUT" +} + +grep -q '^unexpected-success$' "$PROBE_OUT" \ + && fail "a hanging tasks-axi show must not report a successful row read: $(cat "$PROBE_OUT")" + +FIRST_ELAPSED=$(probe_elapsed wedged-one) +[ -n "$FIRST_ELAPSED" ] || fail "probe produced no timing: $(cat "$PROBE_OUT")" +[ "$FIRST_ELAPSED" -lt "$BOUND_CEILING" ] \ + || fail "bounded row read took ${FIRST_ELAPSED}s, over the ${BOUND_CEILING}s ceiling: $(cat "$PROBE_OUT")" +pass "a hanging tasks-axi show returns within the per-item bound instead of running unbounded" + +FIRST_ERROR=$(probe_error wedged-one) +case "$FIRST_ERROR" in + *wedged-one*bound*) ;; + *) fail "the timed-out read must name the item and its bound, got: $FIRST_ERROR" ;; +esac +pass "a timed-out row read reports one error naming the item that timed out" + +# The latch is what keeps a home carrying a large fleet from paying N bounds and +# losing the digest anyway, so assert it strictly: a latched read must be +# FASTER than one bound, not merely under the ceiling. A ceiling-only assertion +# passes whether or not the latch works, and fm_backlog_row_show runs inside a +# command substitution whose writes die with the subshell - the exact way this +# latch can silently become inert. +for SKIPPED in wedged-two wedged-three; do + SKIPPED_ERROR=$(probe_error "$SKIPPED") + SKIPPED_ELAPSED=$(probe_elapsed "$SKIPPED") + case "$SKIPPED_ERROR" in + *"$SKIPPED"*skipped*) ;; + *) fail "every skipped item must still be named as skipped, $SKIPPED got: $SKIPPED_ERROR" ;; + esac + [ -n "$SKIPPED_ELAPSED" ] && [ "$SKIPPED_ELAPSED" -lt "$BOUND_SECS" ] \ + || fail "the latch is inert: $SKIPPED paid ${SKIPPED_ELAPSED}s against a known-wedged backend" +done +pass "after the first bound hit the sweep continues and names every remaining item without paying the bound again" + +# A padded zero is still zero, and `timeout 0` / `alarm 0` disable the deadline +# outright, so a bound that only rejects the literal 0 silently restores the +# unbounded read this whole change exists to prevent. +PADDED_OUT="$UNIT/padded.out" +PADDED_START=$(date +%s) +PATH="$UNIT_FAKEBIN:$BASE_PATH" FM_BACKLOG_ROW_TIMEOUT_SECS=00 \ + bash -c ' + set -u + . "$1/bin/fm-tasks-axi-lib.sh" + . "$1/bin/fm-backlog-transition-lib.sh" + fm_backlog_row_probe "$2" padded-zero && printf "unexpected-success\n" + printf "error=%s\n" "$FM_BACKLOG_ROW_ERROR" + ' _ "$ROOT" "$UNIT/data" > "$PADDED_OUT" 2>&1 +PADDED_ELAPSED=$(elapsed_since "$PADDED_START") + +[ "$PADDED_ELAPSED" -lt "$BOUND_CEILING" ] \ + || fail "a padded-zero bound disabled the deadline: the read ran ${PADDED_ELAPSED}s" +case "$(sed -n 's/^error=//p' "$PADDED_OUT")" in + *padded-zero*bound*) ;; + *) fail "a padded-zero bound must fall back to the default bound and report it: $(cat "$PADDED_OUT")" ;; +esac +pass "a padded-zero bound falls back to the default instead of disabling the deadline" + +# --- a bound hit is not absence --------------------------------------------- +# +# Turning a hang into a fast 124 reaches every caller that reads a non-zero row +# status as "this row does not exist". bin/fm-captain-hold.sh's hold path is the +# one where that misreading corrupts: it would create a task that already +# exists. The bound must stop the command instead. + +CAPTAIN="$TMP_ROOT/captain" +CAPTAIN_FAKEBIN=$(fm_fakebin "$CAPTAIN") +mkdir -p "$CAPTAIN/data" "$CAPTAIN/state" "$CAPTAIN/config" +make_hanging_tasks_axi "$CAPTAIN_FAKEBIN" +cp "$ROOT/.tasks.toml" "$CAPTAIN/.tasks.toml" +printf '# Backlog\n' > "$CAPTAIN/data/backlog.md" + +ADD_LOG="$CAPTAIN/add.log" +HOLD_OUT="$CAPTAIN/hold.out" +HOLD_STATUS=0 +PATH="$CAPTAIN_FAKEBIN:$BASE_PATH" FM_HOME="$CAPTAIN" \ + FM_STATE_OVERRIDE="$CAPTAIN/state" FM_DATA_OVERRIDE="$CAPTAIN/data" \ + FM_CONFIG_OVERRIDE="$CAPTAIN/config" FM_BACKLOG_ROW_TIMEOUT_SECS="$BOUND_SECS" \ + FM_TEST_TASKS_AXI_ADD_LOG="$ADD_LOG" \ + "$ROOT/bin/fm-captain-hold.sh" hold wedged-hold --title 'Wedged hold' --reason 'backend wedged' \ + > "$HOLD_OUT" 2>&1 || HOLD_STATUS=$? + +[ "$HOLD_STATUS" -ne 0 ] \ + || fail "holding a task against a wedged backend must not report success: $(cat "$HOLD_OUT")" +[ ! -s "$ADD_LOG" ] \ + || fail "a timed-out read was spent as absence: tasks-axi add ran anyway: $(cat "$ADD_LOG")" +case "$(cat "$HOLD_OUT")" in + *wedged-hold*bound*) ;; + *) fail "the refusal must name the item and the bound it hit, got: $(cat "$HOLD_OUT")" ;; +esac +pass "a bound hit stops a captain hold loudly instead of being read as a missing task" + +# The teardown gate reaches a row read through the same resolver, so the bound +# hit has to survive the command substitution that carries the resolved id. +fm_write_meta "$CAPTAIN/state/wedged-origin.meta" \ + 'window=firstmate:fm-wedged-origin' \ + 'worktree=/nonexistent/wedged-origin' \ + 'project=alpha' \ + 'harness=claude' \ + 'decisions_reviewed=1' \ + 'decision_keys=wedged-entry' + +VERIFY_OUT="$CAPTAIN/verify.out" +VERIFY_STATUS=0 +PATH="$CAPTAIN_FAKEBIN:$BASE_PATH" FM_HOME="$CAPTAIN" \ + FM_STATE_OVERRIDE="$CAPTAIN/state" FM_DATA_OVERRIDE="$CAPTAIN/data" \ + FM_CONFIG_OVERRIDE="$CAPTAIN/config" FM_BACKLOG_ROW_TIMEOUT_SECS="$BOUND_SECS" \ + "$ROOT/bin/fm-captain-hold.sh" verify wedged-origin > "$VERIFY_OUT" 2>&1 || VERIFY_STATUS=$? + +[ "$VERIFY_STATUS" -ne 0 ] \ + || fail "verify must not attest an inventory it could not read: $(cat "$VERIFY_OUT")" +case "$(cat "$VERIFY_OUT")" in + *absent*) fail "a bound hit was reported as an absent task: $(cat "$VERIFY_OUT")" ;; +esac +case "$(cat "$VERIFY_OUT")" in + *wedged-entry*bound*) ;; + *) fail "verify must name the entry it could not read and the bound it hit, got: $(cat "$VERIFY_OUT")" ;; +esac +pass "the teardown verify gate reports a bound hit by name instead of as an absent inventory entry" + +# The reconcile-requests intake reads each row with task_show in this shell and +# must stop on a bound hit by name; spending the 124 as 'refused: <id> +# (absent)' would let a wedged backend erase real rows from the reconcile +# sweep. +REQ="$TMP_ROOT/req" +REQ_FAKEBIN=$(fm_fakebin "$REQ") +mkdir -p "$REQ/data" "$REQ/state" "$REQ/config" "$REQ/state/decision-bindings" +make_hanging_tasks_axi "$REQ_FAKEBIN" +cp "$ROOT/.tasks.toml" "$REQ/.tasks.toml" +printf '# Backlog\n' > "$REQ/data/backlog.md" +printf 'schema=fm-decision-binding.v1\norigin=wedged-origin\n' \ + > "$REQ/state/decision-bindings/probe.origin" + +REQ_OUT="$REQ/req.out" +REQ_STATUS=0 +printf 'wedged-req\n' \ + | PATH="$REQ_FAKEBIN:$BASE_PATH" FM_HOME="$REQ" \ + FM_STATE_OVERRIDE="$REQ/state" FM_DATA_OVERRIDE="$REQ/data" \ + FM_CONFIG_OVERRIDE="$REQ/config" FM_BACKLOG_ROW_TIMEOUT_SECS="$BOUND_SECS" \ + "$ROOT/bin/fm-captain-hold.sh" reconcile-requests --source-id probe --source 'test capture' \ + > "$REQ_OUT" 2>&1 || REQ_STATUS=$? + +[ "$REQ_STATUS" -ne 0 ] \ + || fail "reconcile-requests must not report success against a wedged backend: $(cat "$REQ_OUT")" +case "$(cat "$REQ_OUT")" in + *absent*|*refused*) fail "the reconcile intake spent a bound hit as an absent row: $(cat "$REQ_OUT")" ;; +esac +case "$(cat "$REQ_OUT")" in + *wedged-req*bound*) ;; + *) fail "the reconcile intake must name the row and the bound it hit, got: $(cat "$REQ_OUT")" ;; +esac +pass "the reconcile-requests intake stops loudly on a bound hit instead of refusing the row as absent" + +# The migrated-prefix scan is the resolution path whose exact and legacy ids +# genuinely answer NOT_FOUND: only the prefixed migrated row wedges. A dropped +# 124 there falls through to 'no captain-held task $entry resolves to nothing' +# - the exact bound-hit-as-absence outcome the resolver's own 124 arm exists to +# prevent - so the bound must survive the prefixed scan to verify_entry_durable. +MIG="$TMP_ROOT/migrated" +MIG_FAKEBIN=$(fm_fakebin "$MIG") +mkdir -p "$MIG/data" "$MIG/state" "$MIG/config" +cat > "$MIG_FAKEBIN/tasks-axi" <<'SH' +#!/usr/bin/env bash +set -u +case "${1:-}" in + --version) printf '%s\n' '0.2.5'; exit 0 ;; + show) + [ -z "${2:-}" ] && { printf 'code: NOT_FOUND\n' >&2; exit 1; } + # Only the prefixed migrated candidates wedge; the exact and legacy ids + # answer NOT_FOUND promptly, the concrete path the prefix scan exists for. + case "$2" in $FM_TEST_PREFIXED_GLOB) sleep 300; exit 0 ;; esac + printf 'code: NOT_FOUND\n' >&2 + exit 1 + ;; + update) + [ "${2:-}" = --help ] || exit 0 + printf '%s\n' 'usage: tasks-axi update <id> [flags]' ' --body-file <path>' ' --archive-body' + exit 0 + ;; + mv) + [ "${2:-}" = --help ] || exit 0 + printf '%s\n' 'usage: tasks-axi mv <id> [<id>...] --to <path-or-dir>' + exit 0 + ;; + hold) + [ "${2:-}" = --help ] || exit 0 + printf '%s\n' 'usage: tasks-axi hold <id> [flags]' ' --kind captain' ' --until <date>' + exit 0 + ;; + list) + printf 'count: 0\n' + printf 'tasks[0]{id,state,kind,repo,title,blocked_by,hold_kind,hold_reason}:\n' + exit 0 + ;; +esac +exit 0 +SH +chmod +x "$MIG_FAKEBIN/tasks-axi" +cat > "$MIG_FAKEBIN/bd" <<'SH' +#!/usr/bin/env bash +[ "${1:-}" = list ] && { printf '[]\n'; exit 0; } +exit 1 +SH +chmod +x "$MIG_FAKEBIN/bd" +cat > "$MIG/.tasks.toml" <<'TOML' +backend = "beads" + +[beads] +prefix = "bd" +path = "graph" +binary = "bd" +TOML +printf '# Backlog\n' > "$MIG/data/backlog.md" +fm_write_meta "$MIG/state/wedged-origin.meta" \ + 'window=firstmate:fm-wedged-origin' \ + 'worktree=/nonexistent/wedged-origin' \ + 'project=alpha' \ + 'harness=claude' \ + 'decisions_reviewed=1' \ + 'decision_keys=mig-entry' + +VERIFY_MIG_OUT="$MIG/verify.out" +VERIFY_MIG_STATUS=0 +PATH="$MIG_FAKEBIN:$BASE_PATH" FM_HOME="$MIG" \ + FM_STATE_OVERRIDE="$MIG/state" FM_DATA_OVERRIDE="$MIG/data" \ + FM_CONFIG_OVERRIDE="$MIG/config" FM_BACKLOG_ROW_TIMEOUT_SECS="$BOUND_SECS" \ + FM_TEST_PREFIXED_GLOB='bd-*' \ + "$ROOT/bin/fm-captain-hold.sh" verify wedged-origin > "$VERIFY_MIG_OUT" 2>&1 || VERIFY_MIG_STATUS=$? + +[ "$VERIFY_MIG_STATUS" -ne 0 ] \ + || fail "verify must not attest an inventory whose migrated-prefix read wedged: $(cat "$VERIFY_MIG_OUT")" +case "$(cat "$VERIFY_MIG_OUT")" in + *'no captain-held task'*|*absent*) + fail "the migrated-prefix bound hit was spent as an unresolved key: $(cat "$VERIFY_MIG_OUT")" ;; +esac +case "$(cat "$VERIFY_MIG_OUT")" in + *'exceeded its read bound resolving mig-entry') ;; + *) fail "verify must name the entry it could not read and the bound it hit, got: $(cat "$VERIFY_MIG_OUT")" ;; +esac +pass "a bound hit in the migrated-prefix scan stops verify by name instead of resolving to nothing" + +# --- half two: the digest still completes end to end ------------------------ + +E2E="$TMP_ROOT/e2e" +E2E_ROOT="$E2E/root" +E2E_HOME="$E2E/home" +E2E_FAKEBIN="$E2E/fakebin" +mkdir -p "$E2E_HOME/state" "$E2E_HOME/data" "$E2E_HOME/config" "$E2E_FAKEBIN" +git init -q -b main "$E2E_ROOT" +git -C "$E2E_ROOT" commit -q --allow-empty -m init + +make_hanging_tasks_axi "$E2E_FAKEBIN" +# The reconcile sweep this half asserts on runs only under a verified fleet +# lock, and fm-lock.sh finds its holder by walking the invoking process tree +# through `ps`. A CI runner's ancestry carries no harness process, so the lock +# would be refused there and the sweep silently skipped. Pin the lock evidence +# the same way tests/fm-session-start.test.sh's make_fake_ps_harness does: +# every queried pid reports a live `claude` harness, independent of whatever +# process tree the test itself was launched from. +cat > "$E2E_FAKEBIN/ps" <<'SH' +#!/usr/bin/env bash +set -u +case "$*" in + *"comm="*) printf '%s\n' '/usr/local/bin/claude'; exit 0 ;; + *"args="*) printf '%s\n' 'claude'; exit 0 ;; + *"ppid="*) exit 1 ;; +esac +exit 1 +SH +chmod +x "$E2E_FAKEBIN/ps" +fm_fake_exit0 "$E2E_FAKEBIN" tmux node chrome-devtools-axi gh treehouse +fm_fake_version_tool "$E2E_FAKEBIN" lavish-axi FM_FAKE_LAVISH_AXI_VERSION 0.1.46 +fm_fake_version_tool "$E2E_FAKEBIN" gh-axi FM_FAKE_GH_AXI_VERSION 0.1.29 +fm_fake_version_tool "$E2E_FAKEBIN" no-mistakes FM_FAKE_NO_MISTAKES_VERSION \ + 'no-mistakes version v1.46.0 (fake) 2026-06-27T00:02:18Z' + +printf '# Backlog\n' > "$E2E_HOME/data/backlog.md" +# One owned record, so the reconcile sweep actually reads the wedged backend. +fm_write_meta "$E2E_HOME/state/wedged-task.meta" \ + 'window=firstmate:fm-wedged-task' \ + 'worktree=/nonexistent/wedged-task' \ + 'project=alpha' \ + 'harness=claude' \ + 'mode=no-mistakes' \ + 'yolo=off' + +DIGEST="$E2E/digest.out" +DIGEST_START=$(date +%s) +env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + FM_HOME="$E2E_HOME" FM_ROOT_OVERRIDE="$E2E_ROOT" PATH="$E2E_FAKEBIN:$BASE_PATH" \ + FM_BACKLOG_ROW_TIMEOUT_SECS="$BOUND_SECS" \ + "$ROOT/bin/fm-session-start.sh" > "$DIGEST" 2>&1 || true +DIGEST_ELAPSED=$(elapsed_since "$DIGEST_START") + +[ "$DIGEST_ELAPSED" -lt "$BOUND_CEILING" ] \ + || fail "session start took ${DIGEST_ELAPSED}s against a wedged backlog backend" + +for SECTION in 'WAKE QUEUE' 'SUPERVISION OPERATING INSTRUCTIONS' 'FLEET STATE' 'CONTEXT'; do + grep -q "$SECTION" "$DIGEST" \ + || fail "the digest lost its $SECTION section against a wedged backlog backend: $(cat "$DIGEST")" +done +pass "a wedged backlog backend still leaves a complete digest: wake queue, supervision instructions, fleet state, and context all print" + +grep -q '^BACKLOG_RECONCILE: wedged-task: ' "$DIGEST" \ + || fail "the wedged item must be reported by name as a partial reconcile: $(cat "$DIGEST")" +pass "an unreachable backlog backend degrades to a loud partial reconcile naming the item it could not read" + +echo "# fm-backlog-read-bound.test.sh: all assertions passed" From 76d44055d81060de6fb8fbe8e94d30291425c0d5 Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Sat, 12 Sep 2026 11:54:46 -0700 Subject: [PATCH 22/31] fix(merge): serialize away authority with synchronous merges (#4285) * fix(merge): serialize the away-authority check with a synchronous merge bin/fm-pr-merge.sh read the away-posture record for merge authority (the per-task merge grant and the yolo/away-grant decision) and handed the merge to the forge afterwards. An archive at the captain's return or a grant revoked by a replacement record could land in between, so a merge could proceed on away authority that no longer held. The away record now carries a cross-subsystem lock, built on the existing bounded lock primitive rather than a new lock format: the record-mutating subcommands hold it across their mutation, and the merge holds it across both its authority read and the forge command. Because a queued or auto merge returns before the pull request lands, and would therefore outlive the lock, an away merge is now refused whenever it could land asynchronously: a requested --auto, a base branch whose merge-queue state does not prove an immediate merge, and GitLab's asynchronous flags and configuration. What remains permitted while away is the synchronous merge that lands inside the lock. This closes the common away-record/merge race against a live lock owner. It does not make the merge atomic in every case, and two narrow races are accepted and documented at their sites rather than hidden, both confused-agent-grade in the sense bin/fm-lease-lib.sh already uses: - A merge-queue rule change or a PR base change in the window between the queue-free preflight and the forge call can still enqueue the merge, which can then land after its grant lapses. - Killing the lock-owning shell while its gh or glab child is still running lets stale-owner recovery reclaim the lock and the record be archived or replaced, after which the orphaned child can complete the merge on lapsed authority. Closing either one needs landing verification or an ownership handoff, which is deliberately out of scope here. No existing gate is relaxed. The lock is taken after the live green-at-head verify and the captain-hold check, the in-lock authority read is unchanged, and a lock that cannot be taken refuses the merge rather than proceeding unlocked. The away grant stays a structured field; no prose is parsed. * no-mistakes(review): Fix GitHub rollup fixture base branch * no-mistakes(document): Document atomic away-authority merge locking * no-mistakes(ci): Updated two executable GitHub API fixtures to include the required baseRefName. Both previously failing test suites now pass: fm-captain-hold-lifecycle.test.sh and fm-pr-check-security.test.sh. git diff --check also passes --- .agents/skills/afk/SKILL.md | 1 + AGENTS.md | 2 +- bin/fm-afk-contract.sh | 95 ++++++++- bin/fm-pr-merge.sh | 145 ++++++++++--- docs/architecture.md | 9 +- docs/captain-hold-lifecycle.md | 1 + docs/scripts.md | 2 +- tests/fm-afk-contract.test.sh | 60 ++++++ tests/fm-captain-hold-lifecycle.test.sh | 2 +- tests/fm-pr-check-security.test.sh | 2 +- tests/fm-pr-merge.test.sh | 264 +++++++++++++++++++++++- 11 files changed, 543 insertions(+), 40 deletions(-) diff --git a/.agents/skills/afk/SKILL.md b/.agents/skills/afk/SKILL.md index 5b2479e2997..1de10bfddec 100644 --- a/.agents/skills/afk/SKILL.md +++ b/.agents/skills/afk/SKILL.md @@ -86,6 +86,7 @@ A PR ready for merge keeps the merge authority from `AGENTS.md` section 7, and a While the away-posture record exists, a merge proceeds only when that task's recorded yolo posture is on or its id is in the record's merge-grant list; otherwise it is held for the captain's return. A merge grant never releases a captain hold, and it expires when the away record is archived. `--allow-red` remains attended-only and is refused while the record exists. +A merge under away authority must be synchronous; `fm-pr-merge.sh` refuses auto-merge and any GitHub queue state that cannot prove an immediate merge while the record exists. A mandate clause is the captain's explicit instruction given before leaving, recorded with its named object and condition; a clause is never inferred, never applied by analogy, and expires at return. Forbidden, destructive, irreversible, and security-sensitive actions are never pre-authorizable regardless of clause text, and no recorded clause is authority by itself. This release records clauses and does not execute them. diff --git a/AGENTS.md b/AGENTS.md index 0c7cf568b99..135831d62f5 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -140,7 +140,7 @@ state/ runtime records and signals; gitignored .watcher-down private generation-bound recovery state coupling watcher downtime, durable wake presentation, and post-handling acknowledgement; never touch .<id>.open-decisions-cursor per-task byte cursor and folded open-decision set bounding the OPEN DECISIONS scan's cost to new status-log appends; written only by fm-classify-lib.sh's status_open_decisions_incremental, removed by teardown, safe to delete (forces one full re-fold) .status-presentation-cursor .status-presentation-lock fleet-wide per-task status identity plus independent annotation and outcome-backstop byte offsets, with a serialization lock preventing already-presented lines from replaying while preserving delayed signal annotations; owned by fm-classify-lib.sh, with each task's row retired by teardown - .afk-contract the away-posture record: the captain's verbatim away words, expected return, reach profile, spend cap, and structured mandate clauses; written only by bin/fm-afk-contract.sh after the captain confirms the read-back, archived under afk-contracts/ at return; its presence IS the away posture in every harness + .afk-contract the away-posture record: the captain's verbatim away words, expected return, reach profile, spend cap, and structured mandate clauses; written only by bin/fm-afk-contract.sh after the captain confirms the read-back, archived under afk-contracts/ at return; its presence IS the away posture in every harness; its sibling .afk-contract.lock serializes actions authorized by the live record (contract: bin/fm-afk-contract.sh) afk-contracts/ archived away-posture records: one final record per away window keyed by entry time, plus any superseded mandates from that window .afk durable away-mode daemon flag on the harnesses that still launch the daemon (never on Pi); present = sub-supervisor may inject escalations (set by the daemon entry, cleared on user return) .watch.lock .wake-queue.lock watcher singleton and queue serialization locks diff --git a/bin/fm-afk-contract.sh b/bin/fm-afk-contract.sh index 272e810846d..04f8197f9a6 100755 --- a/bin/fm-afk-contract.sh +++ b/bin/fm-afk-contract.sh @@ -112,9 +112,27 @@ # fm-afk-contract.sh archive move the record aside; print its path # fm-afk-contract.sh archived <entered_epoch> print that archived record's path # -# Sourceable: with the BASH_SOURCE guard, other scripts get the path and -# presence helpers (fm_afk_contract_path, fm_afk_contract_present, -# fm_afk_contract_proposal_path, fm_afk_contract_archive_dir) without running main. +# CROSS-SUBSYSTEM LOCK (state/.afk-contract.lock; this script is its one owner). +# This record is authority another subsystem reads and then ACTS on outside this +# script: bin/fm-pr-merge.sh reads the merge grants and afterwards hands a merge +# to the forge. A publication, replacement, or archive landing between that read +# and the forge handoff would land a merge on authority that no longer holds, so +# the two subsystems share one lock instead of each locking its own records: the +# record-mutating subcommands (confirm, archive) hold it across their mutation, +# and a reader that acts on the record holds it across both its read and that +# action (fm_afk_contract_lock_hold / fm_afk_contract_lock_release). The +# read-only subcommands never take it, so a holder can still read the record it +# locked. Neither side ever proceeds without it: the acquire is bounded, and a +# bound that is hit refuses and names the live holder rather than racing. That +# fixed bound is 120 seconds, sized so only a genuinely wedged holder trips it. +# A lock left by a killed process is reclaimed +# by the ordinary stale-owner recovery in bin/fm-wake-lib.sh, which owns the lock +# primitive itself. +# +# Sourceable: with the BASH_SOURCE guard, other scripts get the path, presence, +# and lock helpers (fm_afk_contract_path, fm_afk_contract_present, +# fm_afk_contract_proposal_path, fm_afk_contract_archive_dir, +# fm_afk_contract_lock_hold, fm_afk_contract_lock_release) without running main. set -u FM_AFK_CONTRACT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -129,6 +147,10 @@ FM_AFK_CONTRACT_VERSION=1 FM_AFK_CONTRACT_VERBS="merge land prerelease install rerun dispatch abort-run answer discard wake-me" FM_AFK_CONTRACT_REACH_ANNOUNCED='No phone channel is configured; anything that needs you waits for your return.' FM_AFK_CONTRACT_SPEND_DEFAULT=4 +# Generous against the longest legitimate holder, a merge waiting on the forge, +# so the bound only ever trips on something genuinely wedged. +_FM_AFK_CONTRACT_LOCK_TIMEOUT=120 +FM_AFK_CONTRACT_LOCK_HELD= fm_afk_contract_path() { # [state-dir] printf '%s/.afk-contract' "${1:-$FM_AFK_CONTRACT_STATE}" @@ -146,6 +168,54 @@ fm_afk_contract_present() { # [state-dir] [ -f "$(fm_afk_contract_path "${1:-$FM_AFK_CONTRACT_STATE}")" ] } +fm_afk_contract_lock_path() { # [state-dir] + printf '%s/.afk-contract.lock' "${1:-$FM_AFK_CONTRACT_STATE}" +} + +# Lazily reach the lock primitive. bin/fm-wake-lib.sh is a canonical lint root +# in its own right, so keep this an analysis boundary for the same reason +# bin/fm-lease-lib.sh's fm_lease_lock_helpers does. +fm_afk_contract_lock_helpers() { + command -v fm_lock_acquire_wait_bounded >/dev/null 2>&1 && return 0 + # shellcheck source=/dev/null + . "$FM_AFK_CONTRACT_DIR/fm-wake-lib.sh" +} + +# fm_afk_contract_lock_hold [state-dir]: take the cross-subsystem lock described +# in the header. The acquire is bounded so a wedged holder is refused instead of +# blocking a merge or a captain return forever, and returns 1 WITHOUT the lock so +# every caller refuses rather than proceeding unlocked. +fm_afk_contract_lock_hold() { # [state-dir] + local lock rc=0 STATE timeout + STATE=${1:-$FM_AFK_CONTRACT_STATE} + lock=$(fm_afk_contract_lock_path "$STATE") + timeout=${FM_TEST_AFK_CONTRACT_LOCK_TIMEOUT:-$_FM_AFK_CONTRACT_LOCK_TIMEOUT} + fm_afk_contract_lock_helpers || { + fm_afk_contract_log "could not load the lock primitive for $lock" + return 1 + } + fm_lock_acquire_wait_bounded "$lock" "$timeout" || rc=$? + if [ "$rc" -ne 0 ]; then + if [ "$rc" -eq 124 ] && [ -n "${FM_LOCK_HELD_PID:-}" ]; then + fm_afk_contract_log "the away-posture record is locked by live process $FM_LOCK_HELD_PID (an in-flight merge, or another change to this record); nothing was changed" + else + fm_afk_contract_log "could not take the away-posture record lock at $lock; nothing was changed" + fi + return 1 + fi + FM_AFK_CONTRACT_LOCK_HELD=$lock +} + +# Release the lock taken by fm_afk_contract_lock_hold. Idempotent, so callers can +# invoke it unconditionally from their own cleanup. +fm_afk_contract_lock_release() { + local lock=$FM_AFK_CONTRACT_LOCK_HELD + [ -n "$lock" ] || return 0 + FM_AFK_CONTRACT_LOCK_HELD= + fm_afk_contract_lock_helpers || return 1 + fm_lock_release "$lock" +} + fm_afk_contract_log() { printf 'fm-afk-contract: %s\n' "$*" >&2; } fm_afk_contract_usage() { @@ -874,13 +944,28 @@ fm_afk_contract_select_path() { # <args...> -> prints the record path chosen by printf '%s' "$path" } +# The record-mutating subcommands run inside the cross-subsystem lock, so no +# publication, replacement, or archive can land between another subsystem's +# authority read and the action it takes on that authority. +fm_afk_contract_locked_cmd() { # <command> [args...] + local rc=0 + fm_afk_contract_lock_hold || return 1 + trap 'fm_afk_contract_lock_release || true' EXIT + "$@" || rc=$? + trap - EXIT + fm_afk_contract_lock_release || true + return "$rc" +} + fm_afk_contract_main() { local cmd=${1:-} path [ -n "$cmd" ] || { fm_afk_contract_usage >&2; return 2; } shift case "$cmd" in propose) fm_afk_contract_cmd_propose "$@" ;; - confirm) [ "$#" -eq 0 ] || { fm_afk_contract_usage >&2; return 2; }; fm_afk_contract_cmd_confirm ;; + confirm) + [ "$#" -eq 0 ] || { fm_afk_contract_usage >&2; return 2; } + fm_afk_contract_locked_cmd fm_afk_contract_cmd_confirm ;; readback) path=$(fm_afk_contract_select_path "$@") || { fm_afk_contract_usage >&2; return 2; } [ -f "$path" ] || { fm_afk_contract_log "no record at $path"; return 1; } @@ -917,7 +1002,7 @@ fm_afk_contract_main() { path=$(fm_afk_contract_select_path "$@") || { fm_afk_contract_usage >&2; return 2; } [ -f "$path" ] || { fm_afk_contract_log "no record at $path"; return 1; } fm_afk_contract_read_grants "$path" ;; - archive) fm_afk_contract_cmd_archive ;; + archive) fm_afk_contract_locked_cmd fm_afk_contract_cmd_archive ;; archived) [ "$#" -eq 1 ] || { fm_afk_contract_usage >&2; return 2; } path="$(fm_afk_contract_archive_dir)/$1.afk-contract" diff --git a/bin/fm-pr-merge.sh b/bin/fm-pr-merge.sh index 85319a13777..9ca7bfb2f9e 100755 --- a/bin/fm-pr-merge.sh +++ b/bin/fm-pr-merge.sh @@ -30,11 +30,13 @@ # refuses, reporting the failed gh read and naming both failed reads when the # gh-axi view could not prove the outcome either. # If the pull request remains open and the base branch has an effective -# merge_queue rule, the refusal names the queue's configured merge method and -# the exact --attended-override -- --auto --<method> retry flags, unless the caller already passed -# that method with --auto to a merge command that returned success, in which -# case it reports instead that the accepted request has not entered the queue -# and the queue state has to be re-checked. +# merge_queue rule, an attended refusal names the queue's configured merge +# method and exact --attended-override -- --auto --<method> retry flags. While +# the away-posture record exists, asynchronous merge requests are refused and +# queue retry flags are not offered because they would outlive away authority. +# An attended caller that already passed the configured method with --auto is +# told instead that the accepted request has not entered the queue and its queue +# state has to be re-checked. # No method is selected for the caller in any case. A rules response that names # no queue rule, one that could not be read, rules that disagree, and a method # this script does not recognise are four distinct outcomes and are reported @@ -73,10 +75,17 @@ # held for the captain return. An unreadable record refuses rather than being # skipped. Neither posture releases a captain hold, and the grant lapses when # the record is archived. +# The authority read and synchronous forge command share the away record's +# cross-subsystem lock, which bin/fm-afk-contract.sh owns, closing the common +# live-owner TOCTOU; failure to take it refuses before the forge call. Async and +# queued paths are refused while away. Two confused-agent-grade limitations are +# accepted rather than hidden: queue or base changes after GitHub's preflight can +# still enqueue, and killing this shell can orphan a forge child after stale-lock +# recovery. docs/architecture.md owns those away-merge limits, while +# docs/captain-hold-lifecycle.md owns the separate merge-to-cleanup residual. # A failed forge command releases the lock after it returns. A successful one # retains the lock until the accepted merge authority is persisted against the -# still-matching task metadata; docs/captain-hold-lifecycle.md owns the accepted -# asynchronous-landing and merge-to-cleanup residuals. +# still-matching task metadata. # # Extra args must not include --repo or -R in any form, including a bundled # short-option cluster such as -yR, because the repository comes only from the @@ -274,6 +283,26 @@ reject_repo_overrides "$@" || exit 1 reject_head_overrides "$@" || exit 1 reject_protected_forge_args "$@" || exit 1 +FM_PR_GITHUB_AUTO_REQUESTED=false +if [ "$PROVIDER" = github ] && caller_requested_auto_merge "$@"; then + FM_PR_GITHUB_AUTO_REQUESTED=true +fi +FM_PR_GITLAB_ASYNC_REQUESTED=false +if [ "$PROVIDER" = gitlab ]; then + for arg in "$@"; do + case "$arg" in + --auto-merge|--when-pipeline-succeeds) FM_PR_GITLAB_ASYNC_REQUESTED=true ;; + --auto-merge=*|--when-pipeline-succeeds=*) + case "${arg#*=}" in + [tT]|[tT][rR][uU][eE]|1) FM_PR_GITLAB_ASYNC_REQUESTED=true ;; + [fF]|[fF][aA][lL][sS][eE]|0) FM_PR_GITLAB_ASYNC_REQUESTED=false ;; + esac + ;; + esac + done +fi +FM_PR_AWAY_POSTURE=false + fm_backlog_directory_present "$STATE" "state directory" || { echo "error: PR merge refused: $FM_BACKLOG_TRANSITION_ERROR" >&2 exit 1 @@ -304,6 +333,7 @@ MERGE_CONTROL_LOCK= MERGE_META_LOCK= merge_control_cleanup() { [ -z "$MERGE_META_LOCK" ] || fm_lock_release "$MERGE_META_LOCK" || true + fm_afk_contract_lock_release || true [ -z "$MERGE_CONTROL_LOCK" ] || fm_lock_release "$MERGE_CONTROL_LOCK" || true } trap merge_control_cleanup EXIT @@ -355,11 +385,12 @@ fi # the merge request. Sets FM_PR_MERGE_HEAD to the verified head on success and # returns non-zero after reporting every condition that failed. FM_PR_MERGE_HEAD= +FM_PR_GITLAB_ASYNC_CONFIGURED=false gitlab_verify_mergeable() { local json fields line local total=0 named=0 refusals='' local state='' detail='' conflicts='' discussions='' - local live_head='' pipeline_sha='' pipeline_status='' + local live_head='' pipeline_sha='' pipeline_status='' async_configured='' # GITLAB_HOST is set to the same host the project URL already carries, so the # instance is taken from the parsed URL by both signals and never from the @@ -381,7 +412,8 @@ gitlab_verify_mergeable() { "discussions=" + (.blocking_discussions_resolved | tostring), "head=" + ((.sha // "") | tostring), "pipeline_sha=" + ((.head_pipeline.sha // "") | tostring), - "pipeline_status=" + ((.head_pipeline.status // "") | tostring) + "pipeline_status=" + ((.head_pipeline.status // "") | tostring), + "async_configured=" + (if .merge_when_pipeline_succeeds == true or (.merge_after != null) then "true" else "false" end) else error("merge request payload is not an object") end' 2>/dev/null); then @@ -398,6 +430,7 @@ gitlab_verify_mergeable() { head=*) live_head=${line#head=} ;; pipeline_sha=*) pipeline_sha=${line#pipeline_sha=} ;; pipeline_status=*) pipeline_status=${line#pipeline_status=} ;; + async_configured=*) async_configured=${line#async_configured=} ;; *) continue ;; esac named=$((named + 1)) @@ -407,7 +440,7 @@ FIELDS # Every field named exactly once and no unnamed line: a value carrying a # newline would split into a line no name matches, so it is refused here # rather than silently truncated into a value a check could accept. - if [ "$named" -ne 7 ] || [ "$total" -ne 7 ]; then + if [ "$named" -ne 8 ] || [ "$total" -ne 8 ]; then echo "error: could not read the GitLab merge request state before merging" >&2 return 1 fi @@ -450,6 +483,7 @@ FIELDS printf 'verified: %s is open and mergeable, with a successful pipeline at head %s\n' \ "$URL" "$live_head" >&2 FM_PR_MERGE_HEAD=$live_head + FM_PR_GITLAB_ASYNC_CONFIGURED=$async_configured } # Every GitHub check that is not green in the given live pull-request JSON, one @@ -535,9 +569,9 @@ github_checks_not_green() { github_verify_mergeable() { local json fields line red name covered local total=0 named=0 refusals='' - local state='' draft='' mergeable='' merge_state='' live_head='' + local state='' draft='' mergeable='' merge_state='' live_head='' base='' - if ! json=$(gh pr view "$URL" --json state,isDraft,mergeable,mergeStateStatus,headRefOid,statusCheckRollup 2>/dev/null) \ + if ! json=$(gh pr view "$URL" --json state,isDraft,mergeable,mergeStateStatus,headRefOid,baseRefName,statusCheckRollup 2>/dev/null) \ || [ -z "$json" ]; then echo "error: could not read the GitHub pull request state before merging" >&2 return 1 @@ -548,7 +582,8 @@ github_verify_mergeable() { "draft=" + (if (.isDraft | type) == "boolean" then (.isDraft | tostring) else "" end), "mergeable=" + ((.mergeable // "") | tostring), "merge_state=" + ((.mergeStateStatus // "") | tostring), - "head=" + ((.headRefOid // "") | tostring) + "head=" + ((.headRefOid // "") | tostring), + "base=" + ((.baseRefName // "") | tostring) else error("pull request payload is not an object") end' 2>/dev/null); then @@ -563,13 +598,14 @@ github_verify_mergeable() { mergeable=*) mergeable=${line#mergeable=} ;; merge_state=*) merge_state=${line#merge_state=} ;; head=*) live_head=${line#head=} ;; + base=*) base=${line#base=} ;; *) continue ;; esac named=$((named + 1)) done <<FIELDS $fields FIELDS - if [ "$named" -ne 5 ] || [ "$total" -ne 5 ]; then + if [ "$named" -ne 6 ] || [ "$total" -ne 6 ] || [ -z "$base" ]; then echo "error: could not read the GitHub pull request state before merging" >&2 return 1 fi @@ -627,6 +663,7 @@ EOF printf 'verified: %s is open and mergeable, with every required check green at head %s\n' \ "$URL" "$live_head" >&2 FM_PR_MERGE_HEAD=$live_head + FM_PR_GITHUB_BASE=$base } # Read one live GitHub pull request view after gh returns. The selected @@ -857,9 +894,35 @@ require_away_merge_grant() { return 1 } +# Take the away record's own lock (bin/fm-afk-contract.sh owns it) so that +# record cannot be published, replaced, or archived between the authority read +# below and the forge command that acts on it. Refuses without the lock: a merge +# on authority nothing is holding still is exactly what this closes. This is the +# only path that holds both the per-task control lock and the away-record lock, +# and it always takes them in that order; the away-record side takes only its own +# lock, so the pair cannot deadlock. +hold_away_record_for_merge() { + fm_afk_contract_lock_hold "$STATE" && return 0 + echo "error: PR merge refused - the away-posture record could not be locked for the merge; nothing was merged" >&2 + return 1 +} + require_current_away_authority() { + FM_PR_AWAY_POSTURE=false + if fm_afk_contract_present "$STATE"; then + FM_PR_AWAY_POSTURE=true + if [ "$PROVIDER" = github ] && [ "$FM_PR_GITHUB_AUTO_REQUESTED" = true ]; then + echo "error: --auto is attended-only; while the away-posture record exists only a synchronous merge may run under its authority lock" >&2 + return 2 + fi + if [ "$PROVIDER" = gitlab ] \ + && { [ "$FM_PR_GITLAB_ASYNC_REQUESTED" = true ] || [ "$FM_PR_GITLAB_ASYNC_CONFIGURED" = true ]; }; then + echo "error: GitLab auto-merge is attended-only; while the away-posture record exists only an immediate merge may run under its authority lock" >&2 + return 2 + fi + fi require_away_merge_grant || return 1 - if fm_afk_contract_present "$STATE" && [ "${#ALLOW_RED[@]}" -gt 0 ]; then + if [ "$FM_PR_AWAY_POSTURE" = true ] && [ "${#ALLOW_RED[@]}" -gt 0 ]; then echo "error: --allow-red is attended-only; while the away-posture record exists the green check is absolute" >&2 return 2 fi @@ -882,6 +945,17 @@ persist_accepted_merge_authority() { return 1 } +refuse_github_queue_while_away() { + [ "$FM_PR_AWAY_POSTURE" = true ] || return 0 + # Accepted confused-agent-grade limitation, as in bin/fm-lease-lib.sh, not an + # oversight: a queue rule or PR base change after this preflight can still + # enqueue the merge, which can land after its away grant lapses. + github_read_queue_method + [ "$FM_PR_GITHUB_QUEUE_STATUS" = none ] && return 0 + echo "error: GitHub merge refused while away because the base branch's merge-queue state does not prove an immediate merge; nothing was handed to the forge" >&2 + return 2 +} + require_recorded_pr_identity() { local existing existing=$(grep '^pr=' "$META" | tail -1 | cut -d= -f2- || true) @@ -891,7 +965,6 @@ require_recorded_pr_identity() { return 1 } -FM_PR_GITHUB_AUTO_REQUESTED=false FM_PR_GITHUB_MERGE_ACCEPTED=false FM_PR_GITHUB_CALLER_METHOD= @@ -935,6 +1008,10 @@ github_caller_method_is() { github_report_queue_rules() { local queue_method methods_display + if [ "$FM_PR_AWAY_POSTURE" = true ]; then + printf 'error: the direct merge did not land while the away-posture record exists; merge-queue retry flags are unavailable because a queued merge would outlive its authority\n' >&2 + return 0 + fi github_read_queue_method case "$FM_PR_GITHUB_QUEUE_STATUS" in single) @@ -987,8 +1064,12 @@ github_report_unmerged_outcome() { fi fi if [ "$FM_PR_GITHUB_QUEUE_OBSERVED" != true ]; then - printf 'error: the merge queue could not be observed for %s because the queue-aware read was unavailable, so a pull request already in the merge queue cannot be told apart from one that never entered it; re-check the pull request'"'"'s merge queue state before retrying\n' \ - "$URL" >&2 + if [ "$FM_PR_AWAY_POSTURE" = true ]; then + printf 'error: the synchronous merge did not land while the away-posture record exists; no asynchronous merge or queue retry is available under away authority\n' >&2 + else + printf 'error: the merge queue could not be observed for %s because the queue-aware read was unavailable, so a pull request already in the merge queue cannot be told apart from one that never entered it; re-check the pull request'"'"'s merge queue state before retrying\n' \ + "$URL" >&2 + fi return 0 fi github_report_queue_rules @@ -1022,6 +1103,10 @@ require_recorded_pr_identity || exit 1 record_pr_metadata || exit 1 require_released_captain_hold || exit 1 +# Accepted confused-agent-grade limitation, as in bin/fm-lease-lib.sh, not an +# oversight: if this lock-owning shell dies while its gh or glab child lives, +# stale-owner recovery can release the record for archive or replacement and +# the orphaned forge child can still merge on the lapsed away authority. case "$PROVIDER" in github) merge_output= @@ -1029,16 +1114,15 @@ case "$PROVIDER" in if ! caller_has_merge_method "$@"; then merge_args=(--squash) fi - if caller_requested_auto_merge "$@"; then - FM_PR_GITHUB_AUTO_REQUESTED=true - fi FM_PR_GITHUB_CALLER_METHOD=$(caller_merge_method "$@") github_verify_mergeable || exit 1 - # This last presence and authority read narrows the publication race to the - # forge handoff; without a shared lock, a residual sub-second race remains. + # The away record is locked first, so this last presence and authority read + # and the forge command below share one live-owner critical section. + hold_away_record_for_merge || exit 1 away_status=0 require_current_away_authority || away_status=$? [ "$away_status" -eq 0 ] || exit "$away_status" + refuse_github_queue_while_away || exit 2 merge_status=0 merge_output=$(gh pr merge "$PR_NUMBER" --repo "$PR_OWNER/$PR_REPO" \ --match-head-commit "$FM_PR_MERGE_HEAD" \ @@ -1046,9 +1130,11 @@ case "$PROVIDER" in if [ "$merge_status" -eq 0 ]; then FM_PR_GITHUB_MERGE_ACCEPTED=true persist_accepted_merge_authority || exit 1 + fm_afk_contract_lock_release || true fm_lock_release "$MERGE_CONTROL_LOCK" || true MERGE_CONTROL_LOCK= else + fm_afk_contract_lock_release || true fm_lock_release "$MERGE_CONTROL_LOCK" || true MERGE_CONTROL_LOCK= [ -z "$merge_output" ] || printf '%s\n' "$merge_output" >&2 @@ -1085,20 +1171,27 @@ case "$PROVIDER" in # in between is refused by GitLab instead of merged unverified. --yes only # skips the interactive confirmation, which no supervised run can answer; # the conditions above are what authorize the merge. - # This last presence and authority read narrows the publication race to the - # forge handoff; without a shared lock, a residual sub-second race remains. + # The away record is locked first, so this last presence and authority read + # and the forge command below share one live-owner critical section. + hold_away_record_for_merge || exit 1 away_status=0 require_current_away_authority || away_status=$? [ "$away_status" -eq 0 ] || exit "$away_status" merge_status=0 + gitlab_merge_args=() + if [ "$FM_PR_AWAY_POSTURE" = true ]; then + gitlab_merge_args=(--auto-merge=false) + fi GITLAB_HOST="$FM_PR_HOST" glab mr merge "$PR_NUMBER" -R "$PROJECT_URL" \ - --sha "$FM_PR_MERGE_HEAD" --yes "$@" || merge_status=$? + --sha "$FM_PR_MERGE_HEAD" --yes "$@" "${gitlab_merge_args[@]+"${gitlab_merge_args[@]}"}" || merge_status=$? if [ "$merge_status" -ne 0 ]; then + fm_afk_contract_lock_release || true fm_lock_release "$MERGE_CONTROL_LOCK" || true MERGE_CONTROL_LOCK= exit "$merge_status" fi persist_accepted_merge_authority || exit 1 + fm_afk_contract_lock_release || true fm_lock_release "$MERGE_CONTROL_LOCK" || true MERGE_CONTROL_LOCK= gitlab_confirm_rc=0 diff --git a/docs/architecture.md b/docs/architecture.md index ef1b1519c54..1ccf3aa7c45 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -317,11 +317,18 @@ A `https://github.com/<owner>/<repo>/pull/<n>` URL requires `gh` and `jq`, is me A check run is green when its current run is green, because GitHub leaves a cancelled run in the rollup beside the passing re-run it triggered when the base branch advanced; `bin/fm-pr-merge.sh`'s `github_checks_not_green` owns the rule, which uses `startedAt` to clear only an older completed check run that a passing run with the same name provably replaced, while unfinished check runs and non-green status contexts stay red. `--auto`, `--admin`, and branch-deletion flags are refused unless `--attended-override` is passed for an explicit captain instruction; that override never skips the live green check, the away-grant check, or a captain hold. An attended `--allow-red <check-name>` may appear once, waives only GitHub checks with that exact name, and is refused while the away-posture record exists. +Because away merge authority is read from that record and then acted on by the forge, the authority read and synchronous forge command share the record's cross-subsystem lock, closing the common live-owner TOCTOU. +A lock that cannot be taken refuses the merge. +While the record exists, GitHub auto-merge and any base whose rules cannot prove the absence of a merge queue are refused before submission, and GitLab auto-merge flags or scheduled state are refused while an immediate merge is forced with a final `--auto-merge=false`. +This is deliberately confused-agent-grade, as `bin/fm-lease-lib.sh` defines that grade, rather than fully atomic. +A GitHub queue-rule or PR-base change after the queue-free preflight can still enqueue a merge that lands after its away grant lapses, and killing the lock-owning shell while its forge child survives lets stale-owner recovery admit archive or replacement before that child completes. +These are accepted limitations, not oversights; durable authority, landing re-verification, and child-lock handoff are outside this boundary. +`bin/fm-afk-contract.sh` owns the lock contract, while `tests/fm-afk-contract.test.sh` and `tests/fm-pr-merge.test.sh` pin the serialization and fail-closed merge behavior. A `https://<host>/<path>/-/merge_requests/<n>` URL (see [docs/gitlab-merge-watch.md](gitlab-merge-watch.md)) invokes `glab mr merge <n> -R https://<host>/<path>`, so the instance comes from the URL, and adds no merge-method flag because the project's own merge method applies. That path merges only after one live read of the merge request confirms it is open, mergeable, conflict-free, with blocking discussions resolved and a successful pipeline at the current head, and it binds the merge to that verified head; recorded metadata is never the authority for those conditions because a rebase leaves it stale. After either forge command returns, the script confirms the PR or MR actually landed, and only a confirmed landing records a landed outcome; a queued or unconfirmed request records none and leaves its poll armed. On GitLab an auto-merge-queued or unconfirmed request is reported without failing the run. -On GitHub an outcome that is neither merged nor queued is refused loudly and non-zero, naming the observed state, and a base branch that requires the merge queue is refused with the concrete `--attended-override -- --auto --<method>` retry flags its configured method requires rather than having a merge method chosen on the caller's behalf. +On GitHub an outcome that is neither merged nor queued is refused loudly and non-zero, naming the observed state, and in attended posture a base branch that requires the merge queue is refused with the concrete `--attended-override -- --auto --<method>` retry flags its configured method requires rather than having a merge method chosen on the caller's behalf. When the forge already accepted exactly those flags and the pull request still has not entered the queue, that refusal points at the queue state to re-check instead of echoing back the flags the caller just ran. An auto-merge request is held to the same standard: `--auto` that leaves the pull request neither merged nor queued is refused rather than reported as success. Every GitHub refusal states what it could not observe as plainly as what it did, so an unreadable branch-rule response, an unrecognised queue method, and a merge queue no available read can see are each named rather than left to look like a base branch with no queue at all. diff --git a/docs/captain-hold-lifecycle.md b/docs/captain-hold-lifecycle.md index 2de7b260a89..88da8e585f0 100644 --- a/docs/captain-hold-lifecycle.md +++ b/docs/captain-hold-lifecycle.md @@ -154,6 +154,7 @@ The window between a merge landing and cleanup is an accepted structural residua That local window is normally only seconds wide and requires re-holding a task whose merge has just landed. A re-hold inside the window makes cleanup retain the row rather than publish it, so the delivery is omitted until the stale hold is cleared from that row. Queued forge merges cannot be covered locally because the forge performs the merge asynchronously after the local command has returned, when no lock this code could hold would still be held. +The away-posture restriction on queued merges and its residual limits are owned by [architecture.md](architecture.md#delivery-modes-are-explicit-per-task). ## Record divergence diff --git a/docs/scripts.md b/docs/scripts.md index 048a064aa3c..09a27089aae 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -87,7 +87,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-watch-checkpoint.sh` | Run one bounded foreground watcher checkpoint for Codex-style supervision | | `fm-watch.sh` | Singleton-safe watcher: absorb benign wakes, detect stalled local-secondmate wake queues, and exit on actionable ones | | `fm-inactive-reconcile.sh` | Reconcile long-inactive direct crewmate terminal outcomes without forge access | -| `fm-afk-contract.sh` | Own the away-posture record: schema, mandate-clause fields and never-set scan, refusal naming the missing part, read-back, entry announcement, archive | +| `fm-afk-contract.sh` | Own the away-posture record: schema, mandate-clause fields and never-set scan, refusal naming the missing part, read-back, entry announcement, archive, and cross-subsystem authority lock | | `fm-afk-start.sh` | Run the common sourceable away-mode daemon entry in the foreground | | `fm-afk-launch.sh` | Own away-mode entry (read-back, confirm, record), exit, rollback, and any backend terminal lifecycle | | `fm-afk-return.sh` | Own deterministic return shutdown, the return brief, catch-up evidence, and the firstmate-actionable blocker gate | diff --git a/tests/fm-afk-contract.test.sh b/tests/fm-afk-contract.test.sh index 8a6ba9fb580..5ccacb6b58e 100755 --- a/tests/fm-afk-contract.test.sh +++ b/tests/fm-afk-contract.test.sh @@ -617,6 +617,65 @@ test_archive_drops_live_grants() { pass "archive removes live grants so archived copies are not consulted" } +# The record-mutating commands share one lock with the subsystems that read this +# record's authority and then act on it (bin/fm-pr-merge.sh reads the grants and +# merges). While a reader holds that lock, confirm and archive must refuse and +# change nothing, so no publication, replacement, or archive can land inside the +# window between that read and the action it authorized. +test_record_changes_refuse_while_a_reader_holds_the_lock() { + local home lock holder_pid i rc out before + home=$(make_home lock-contended) + contract "$home" propose --grant task-x1 >/dev/null || fail "lock-contended: proposal failed" + contract "$home" confirm >/dev/null || fail "lock-contended: confirm failed" + before=$(cat "$home/state/.afk-contract") + lock="$home/state/.afk-contract.lock" + + FM_STATE_OVERRIDE="$home/state" bash -c ' + . "$1" + fm_lock_acquire_wait "$2" || exit 10 + printf "ready\n" > "$3" + while [ ! -e "$4" ]; do sleep 0.05; done + fm_lock_release "$2" + ' _ "$ROOT/bin/fm-wake-lib.sh" "$lock" "$home/holder.ready" "$home/release" & + holder_pid=$! + i=0 + while [ "$i" -lt 100 ] && [ ! -s "$home/holder.ready" ]; do + sleep 0.05 + i=$((i + 1)) + done + [ -s "$home/holder.ready" ] \ + || { kill "$holder_pid" 2>/dev/null || true; fail "lock-contended: the fixture never took the lock"; } + + set +e + out=$(FM_TEST_AFK_CONTRACT_LOCK_TIMEOUT=1 contract "$home" archive 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || { kill "$holder_pid" 2>/dev/null || true; fail "lock-contended: archive ran while the record was locked"; } + assert_contains "$out" 'locked by live process' "lock-contended: the archive refusal did not name the live holder" + [ -f "$home/state/.afk-contract" ] \ + || { kill "$holder_pid" 2>/dev/null || true; fail "lock-contended: the refused archive still moved the record"; } + + contract "$home" propose --grant task-other >/dev/null || fail "lock-contended: replacement proposal failed" + set +e + out=$(FM_TEST_AFK_CONTRACT_LOCK_TIMEOUT=1 contract "$home" confirm 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || { kill "$holder_pid" 2>/dev/null || true; fail "lock-contended: confirm replaced the record while it was locked"; } + assert_contains "$out" 'locked by live process' "lock-contended: the confirm refusal did not name the live holder" + [ "$(cat "$home/state/.afk-contract")" = "$before" ] \ + || { kill "$holder_pid" 2>/dev/null || true; fail "lock-contended: the refused confirm changed the standing record"; } + [ "$(contract "$home" grants)" = task-x1 ] \ + || { kill "$holder_pid" 2>/dev/null || true; fail "lock-contended: a read subcommand did not see the unchanged grants"; } + + : > "$home/release" + wait "$holder_pid" || fail "lock-contended: the fixture holder did not release cleanly" + contract "$home" confirm >/dev/null 2>&1 || fail "lock-contended: confirm failed once the lock cleared" + [ "$(contract "$home" grants)" = task-other ] \ + || fail "lock-contended: the released replacement did not take effect" + contract "$home" archive >/dev/null || fail "lock-contended: archive failed once the lock cleared" + pass "confirm and archive refuse while the record is locked, and proceed once it clears" +} + test_fields_refuse_each_missing_part_by_name test_omitted_stop_confirms_as_no_stop test_never_set_flags_without_refusing_and_never_over_matches @@ -641,4 +700,5 @@ test_merge_grants_empty_form_and_usage_errors test_legacy_record_without_merge_grants_reads_empty test_malformed_merge_grants_refuse_validation test_archive_drops_live_grants +test_record_changes_refuse_while_a_reader_holds_the_lock diff --git a/tests/fm-captain-hold-lifecycle.test.sh b/tests/fm-captain-hold-lifecycle.test.sh index dae536d8684..91ad77d9eca 100755 --- a/tests/fm-captain-hold-lifecycle.test.sh +++ b/tests/fm-captain-hold-lifecycle.test.sh @@ -106,7 +106,7 @@ case "${1:-} ${2:-}" in "pr view") case " $* " in *statusCheckRollup*) - printf '%s\n' '{"state":"OPEN","isDraft":false,"mergeable":"MERGEABLE","mergeStateStatus":"CLEAN","headRefOid":"1111111111111111111111111111111111111111","statusCheckRollup":[{"__typename":"CheckRun","name":"ci","status":"COMPLETED","conclusion":"SUCCESS"}]}' + printf '%s\n' '{"state":"OPEN","isDraft":false,"mergeable":"MERGEABLE","mergeStateStatus":"CLEAN","headRefOid":"1111111111111111111111111111111111111111","baseRefName":"main","statusCheckRollup":[{"__typename":"CheckRun","name":"ci","status":"COMPLETED","conclusion":"SUCCESS"}]}' ;; *headRefOid*) printf '%s\n' 1111111111111111111111111111111111111111 ;; esac diff --git a/tests/fm-pr-check-security.test.sh b/tests/fm-pr-check-security.test.sh index 40d8438fcb4..b38435781bc 100755 --- a/tests/fm-pr-check-security.test.sh +++ b/tests/fm-pr-check-security.test.sh @@ -147,7 +147,7 @@ case "${1:-} ${2:-}" in "pr view") case " $* " in *statusCheckRollup*) - printf '%s\n' "{\"state\":\"OPEN\",\"isDraft\":false,\"mergeable\":\"MERGEABLE\",\"mergeStateStatus\":\"CLEAN\",\"headRefOid\":\"${FM_TEST_GH_HEAD:-0123456789abcdef0123456789abcdef01234567}\",\"statusCheckRollup\":[{\"__typename\":\"CheckRun\",\"name\":\"ci\",\"status\":\"COMPLETED\",\"conclusion\":\"SUCCESS\"}]}" + printf '%s\n' "{\"state\":\"OPEN\",\"isDraft\":false,\"mergeable\":\"MERGEABLE\",\"mergeStateStatus\":\"CLEAN\",\"headRefOid\":\"${FM_TEST_GH_HEAD:-0123456789abcdef0123456789abcdef01234567}\",\"baseRefName\":\"main\",\"statusCheckRollup\":[{\"__typename\":\"CheckRun\",\"name\":\"ci\",\"status\":\"COMPLETED\",\"conclusion\":\"SUCCESS\"}]}" exit 0 ;; esac diff --git a/tests/fm-pr-merge.test.sh b/tests/fm-pr-merge.test.sh index 5f4122b001e..b90c0468117 100755 --- a/tests/fm-pr-merge.test.sh +++ b/tests/fm-pr-merge.test.sh @@ -65,7 +65,7 @@ write_github_live_json() { local case_dir=$1 head=$2 printf '%s\n' "$head" > "$case_dir/github-head" cat > "$case_dir/github-view.json" <<JSON -{"state":"OPEN","isDraft":false,"mergeable":"MERGEABLE","mergeStateStatus":"CLEAN","headRefOid":"$head","statusCheckRollup":[{"__typename":"CheckRun","name":"ci","status":"COMPLETED","conclusion":"SUCCESS"}]} +{"state":"OPEN","isDraft":false,"mergeable":"MERGEABLE","mergeStateStatus":"CLEAN","headRefOid":"$head","baseRefName":"main","statusCheckRollup":[{"__typename":"CheckRun","name":"ci","status":"COMPLETED","conclusion":"SUCCESS"}]} JSON } @@ -73,7 +73,7 @@ write_github_red_json() { local case_dir=$1 head=$2 name=$3 printf '%s\n' "$head" > "$case_dir/github-head" cat > "$case_dir/github-view.json" <<JSON -{"state":"OPEN","isDraft":false,"mergeable":"MERGEABLE","mergeStateStatus":"CLEAN","headRefOid":"$head","statusCheckRollup":[{"__typename":"CheckRun","name":"$name","status":"COMPLETED","conclusion":"FAILURE"}]} +{"state":"OPEN","isDraft":false,"mergeable":"MERGEABLE","mergeStateStatus":"CLEAN","headRefOid":"$head","baseRefName":"main","statusCheckRollup":[{"__typename":"CheckRun","name":"$name","status":"COMPLETED","conclusion":"FAILURE"}]} JSON } @@ -108,7 +108,7 @@ write_github_rollup_json() { done printf '%s\n' "$head" > "$case_dir/github-head" cat > "$case_dir/github-view.json" <<JSON -{"state":"OPEN","isDraft":false,"mergeable":"MERGEABLE","mergeStateStatus":"CLEAN","headRefOid":"$head","statusCheckRollup":[$rollup]} +{"state":"OPEN","isDraft":false,"mergeable":"MERGEABLE","mergeStateStatus":"CLEAN","headRefOid":"$head","baseRefName":"main","statusCheckRollup":[$rollup]} JSON } @@ -159,6 +159,17 @@ case "${1:-} ${2:-}" in if [ -n "${FM_TEST_META_AT_MERGE:-}" ] && [ -f "${FM_STATE_OVERRIDE:-}/task-x1.meta" ]; then cat "$FM_STATE_OVERRIDE/task-x1.meta" > "$FM_TEST_META_AT_MERGE" fi + # The forge call runs inside the merge's critical section, so a real + # away-record change attempted from here is the TOCTOU itself: whatever + # happens to it happens between the authority read and the merge. + if [ -x "${FM_TEST_AWAY_MUTATE_AT_MERGE:-}" ]; then + away_rc=0 + "$FM_TEST_AWAY_MUTATE_AT_MERGE" > "$FM_TEST_AWAY_MUTATE_OUT" 2>&1 || away_rc=$? + printf '%s\n' "$away_rc" > "$FM_TEST_AWAY_MUTATE_RC" + "$FM_TEST_ROOT/bin/fm-afk-contract.sh" grants \ + > "$FM_TEST_AWAY_GRANTS_AT_MERGE" 2>/dev/null \ + || printf 'no-live-record\n' > "$FM_TEST_AWAY_GRANTS_AT_MERGE" + fi if [ -n "${FM_TEST_GH_MERGE_OUTPUT:-}" ]; then printf '%s\n' "$FM_TEST_GH_MERGE_OUTPUT" else @@ -278,6 +289,7 @@ write_mr_json() { local file=$1 kv key value local state=opened detail=mergeable conflicts=false discussions=true local head=$MR_HEAD pipeline_sha=$MR_HEAD pipeline_status=success pipeline=present + local merge_when_pipeline_succeeds=false merge_after=null shift for kv in "$@"; do key=${kv%%=*} @@ -291,6 +303,8 @@ write_mr_json() { pipeline_sha) pipeline_sha=$value ;; pipeline_status) pipeline_status=$value ;; pipeline) pipeline=$value ;; + merge_when_pipeline_succeeds) merge_when_pipeline_succeeds=$value ;; + merge_after) merge_after=$value ;; *) fail "write_mr_json: unknown field '$key'" ;; esac done @@ -299,8 +313,10 @@ write_mr_json() { fi printf '{"iid":7,"state":"%s","detailed_merge_status":"%s","has_conflicts":%s,' \ "$state" "$detail" "$conflicts" > "$file" - printf '"blocking_discussions_resolved":%s,"sha":"%s","head_pipeline":%s}\n' \ + printf '"blocking_discussions_resolved":%s,"sha":"%s","head_pipeline":%s,' \ "$discussions" "$head" "$pipeline" >> "$file" + printf '"merge_when_pipeline_succeeds":%s,"merge_after":%s}\n' \ + "$merge_when_pipeline_succeeds" "$merge_after" >> "$file" } # make_gitlab_case <name> [<field>=<value> ...]: a case dir with both forge @@ -367,6 +383,11 @@ run_pr_merge() { FM_TEST_GH_RULES_FAIL="$case_dir/github-rules-fail" \ FM_TEST_META_AT_MERGE="$case_dir/meta-at-merge" \ FM_TEST_AWAY_RECORD_AFTER_VIEW="$case_dir/away-record-after-view" \ + FM_TEST_ROOT="$ROOT" \ + FM_TEST_AWAY_MUTATE_AT_MERGE="${FM_TEST_AWAY_MUTATE_AT_MERGE:-}" \ + FM_TEST_AWAY_MUTATE_OUT="$case_dir/away-mutate-output" \ + FM_TEST_AWAY_MUTATE_RC="$case_dir/away-mutate-rc" \ + FM_TEST_AWAY_GRANTS_AT_MERGE="$case_dir/away-grants-at-merge" \ FM_TEST_REAL_MV="$REAL_MV" \ FM_TEST_GLAB_LOG="$case_dir/glab.log" \ FM_TEST_GLAB_JSON="$case_dir/mr.json" \ @@ -2686,6 +2707,79 @@ test_away_grant_and_yolo_and_hold_for_return() { pass "away merges require yolo or a grant, and --attended-override does not skip that" } +test_away_posture_refuses_asynchronous_merge_paths() { + local case_dir rc url head merge_line + head=abababababababababababababababababababab + url=https://github.com/example/repo/pull/89 + + case_dir=$(make_case away-auto-refused) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + write_away_record "$case_dir" --grant task-x1 + set +e + run_pr_merge "$case_dir" task-x1 "$url" --attended-override -- --auto --merge \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 2 "$rc" "away-auto-refused: auto-merge must be attended-only" + assert_grep '--auto is attended-only' "$case_dir/stderr" \ + "away-auto-refused: refusal did not name the asynchronous flag" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "away-auto-refused: gh pr merge ran for an away auto-merge request" + + case_dir=$(make_case away-queue-refused) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" "$head" + printf 'merge_method=MERGE\n' > "$case_dir/github-rules" + write_away_record "$case_dir" --grant task-x1 + set +e + run_pr_merge "$case_dir" task-x1 "$url" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 2 "$rc" "away-queue-refused: a required merge queue must refuse before submission" + assert_grep 'merge-queue state does not prove an immediate merge' "$case_dir/stderr" \ + "away-queue-refused: refusal did not explain the away restriction" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "away-queue-refused: gh received a merge that could enter its queue" + + case_dir=$(make_gitlab_case away-gitlab-auto) + write_away_record "$case_dir" --grant task-x1 + set +e + run_pr_merge "$case_dir" task-x1 "$MR_URL" --attended-override -- --auto-merge \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 2 "$rc" "away-gitlab-auto: GitLab auto-merge must refuse" + assert_grep 'GitLab auto-merge is attended-only' "$case_dir/stderr" \ + "away-gitlab-auto: refusal did not name auto-merge" + [ -z "$(glab_merge_line "$case_dir/glab.log")" ] \ + || fail "away-gitlab-auto: glab received an asynchronous merge" + + case_dir=$(make_gitlab_case away-gitlab-configured merge_when_pipeline_succeeds=true) + write_away_record "$case_dir" --grant task-x1 + set +e + run_pr_merge "$case_dir" task-x1 "$MR_URL" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 2 "$rc" "away-gitlab-configured: configured auto-merge must refuse" + [ -z "$(glab_merge_line "$case_dir/glab.log")" ] \ + || fail "away-gitlab-configured: glab received a configured asynchronous merge" + + case_dir=$(make_gitlab_case away-gitlab-sync) + write_away_record "$case_dir" --grant task-x1 + run_pr_merge "$case_dir" task-x1 "$MR_URL" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" \ + || fail "away-gitlab-sync: an immediate granted merge should succeed" + merge_line=$(glab_merge_line "$case_dir/glab.log") + case "$merge_line" in + *" --auto-merge=false") ;; + *) fail "away-gitlab-sync: the final glab flag did not force an immediate merge: '$merge_line'" ;; + esac + pass "away posture permits immediate merges but refuses every asynchronous path" +} + test_away_grant_does_not_bypass_red_or_identity() { local case_dir rc head head=adadadadadadadadadadadadadadadadadadadad @@ -2739,6 +2833,164 @@ test_unreadable_away_record_refuses_merge() { pass "an unreadable away-posture record refuses the merge instead of skipping the grant" } +# The race this closes: the away record is read for merge authority and the +# forge is called afterwards, so an archive (the captain's return) or a grant +# revocation landing in between would merge on authority that no longer holds. +# away_change_script writes the change the gh mock attempts from inside the +# forge call, which IS that window. Its body drives the real away-record +# commands /afk and the return use, never a file edit, and takes a one-second +# lock bound so a contended case refuses quickly instead of waiting. +away_change_script() { # <case-dir> <name>; script body on stdin + local case_dir=$1 name=$2 path + path="$case_dir/$name" + { + printf '#!/usr/bin/env bash\n' + printf 'set -eu\n' + printf 'export FM_TEST_AFK_CONTRACT_LOCK_TIMEOUT=1\n' + printf 'CONTRACT="%s/bin/fm-afk-contract.sh"\n' "$ROOT" + cat + } > "$path" + chmod +x "$path" + printf '%s\n' "$path" +} + +# Two away-record changes, each attempted from inside the merge's critical +# section: the archive a captain return performs, and the replacement that +# revokes a grant. Neither may land there, and the merge must still complete on +# the authority it read. +test_away_record_cannot_change_between_the_authority_read_and_the_merge() { + local case_dir rc mutate + case_dir=$(make_case away-archive-at-merge) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 1b1b1b1b1b1b1b1b1b1b1b1b1b1b1b1b1b1b1b1b + write_away_record "$case_dir" --grant task-x1 + mutate=$(away_change_script "$case_dir" archive-at-merge <<'SH' +"$CONTRACT" archive +SH + ) + + export FM_TEST_AWAY_MUTATE_AT_MERGE="$mutate" + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/71 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + unset FM_TEST_AWAY_MUTATE_AT_MERGE + + expect_code 0 "$rc" "away-archive-at-merge: the granted green merge should still land" + [ -s "$case_dir/away-mutate-rc" ] \ + || fail "away-archive-at-merge: the archive was never attempted inside the merge" + [ "$(cat "$case_dir/away-mutate-rc")" != 0 ] \ + || fail "away-archive-at-merge: the archive landed inside the merge's critical section" + assert_grep 'locked by live process' "$case_dir/away-mutate-output" \ + "away-archive-at-merge: the refused archive did not name the live holder" + assert_equals task-x1 "$(cat "$case_dir/away-grants-at-merge" 2>/dev/null || true)" \ + "away-archive-at-merge: the grant this merge read was not still standing at the forge call" + assert_grep "merge landed: task-x1 https://github.com/example/repo/pull/71 away-grant" \ + "$case_dir/state/.wake-queue" \ + "away-archive-at-merge: the landed merge was not recorded under the grant it read" + # The lock goes with the merge rather than leaking: the captain's return + # archives the record on its first try once the merge is done. + FM_HOME="$case_dir/home" FM_STATE_OVERRIDE="$case_dir/state" \ + "$ROOT/bin/fm-afk-contract.sh" archive >/dev/null \ + || fail "away-archive-at-merge: the record stayed locked after the merge" + + case_dir=$(make_case away-revoke-at-merge) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 2c2c2c2c2c2c2c2c2c2c2c2c2c2c2c2c2c2c2c2c + write_away_record "$case_dir" --grant task-x1 + mutate=$(away_change_script "$case_dir" revoke-at-merge <<'SH' +"$CONTRACT" propose --grant task-other +"$CONTRACT" confirm +SH + ) + export FM_TEST_AWAY_MUTATE_AT_MERGE="$mutate" + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/72 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + unset FM_TEST_AWAY_MUTATE_AT_MERGE + + expect_code 0 "$rc" "away-revoke-at-merge: the granted green merge should still land" + [ "$(cat "$case_dir/away-mutate-rc" 2>/dev/null || true)" != 0 ] \ + || fail "away-revoke-at-merge: the replacement landed inside the critical section" + assert_equals task-x1 "$(cat "$case_dir/away-grants-at-merge" 2>/dev/null || true)" \ + "away-revoke-at-merge: the grant was revoked inside the merge's critical section" + pass "no away-record archive or grant revocation lands between the authority read and the merge" +} + +# The same serialization from the other side. A revocation that wins the race +# lands BEFORE the in-lock authority read, and the merge then refuses: the lock +# decides an order, it never lets a stale grant through. +test_a_grant_revoked_before_the_merge_refuses_it() { + local case_dir rc + case_dir=$(make_case away-revoked-before-merge) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 3d3d3d3d3d3d3d3d3d3d3d3d3d3d3d3d3d3d3d3d + write_away_record "$case_dir" + mv "$case_dir/state/.afk-contract" "$case_dir/away-record-after-view" + write_away_record "$case_dir" --grant task-x1 + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/73 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "away-revoked-before-merge: a revoked grant must refuse" + assert_grep 'held for the captain return' "$case_dir/stderr" \ + "away-revoked-before-merge: refusal did not name hold-for-return" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "away-revoked-before-merge: gh pr merge ran on a revoked grant" + pass "a grant revoked before the merge's own authority read refuses the merge" +} + +# Fail closed. The lock is what makes the authority read and the merge one +# action, so a merge that cannot take it has no locked window to merge in and +# refuses - including on this attended case, where the record is absent and +# there is no grant to check at all. +test_merge_refuses_when_the_away_record_cannot_be_locked() { + local case_dir rc holder_pid i lock + case_dir=$(make_case away-lock-unavailable) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 4e4e4e4e4e4e4e4e4e4e4e4e4e4e4e4e4e4e4e4e + lock="$case_dir/state/.afk-contract.lock" + + FM_STATE_OVERRIDE="$case_dir/state" bash -c ' + . "$1" + fm_lock_acquire_wait "$2" || exit 10 + printf "ready\n" > "$3" + while [ ! -e "$4" ]; do sleep 0.05; done + fm_lock_release "$2" + ' _ "$ROOT/bin/fm-wake-lib.sh" "$lock" "$case_dir/holder.ready" "$case_dir/release-holder" & + holder_pid=$! + i=0 + while [ "$i" -lt 100 ] && [ ! -s "$case_dir/holder.ready" ]; do + sleep 0.05 + i=$((i + 1)) + done + [ -s "$case_dir/holder.ready" ] \ + || { kill "$holder_pid" 2>/dev/null || true; fail "away-lock-unavailable: the fixture never took the record lock"; } + + export FM_TEST_AFK_CONTRACT_LOCK_TIMEOUT=1 + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/74 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + unset FM_TEST_AFK_CONTRACT_LOCK_TIMEOUT + : > "$case_dir/release-holder" + wait "$holder_pid" || fail "away-lock-unavailable: the fixture holder did not release cleanly" + + expect_code 1 "$rc" "away-lock-unavailable: an unlockable away record must refuse the merge" + assert_grep 'could not be locked for the merge' "$case_dir/stderr" \ + "away-lock-unavailable: refusal did not name the lock it could not take" + assert_no_grep 'pr merge' "$case_dir/gh.log" \ + "away-lock-unavailable: gh pr merge ran without the away-record lock" + pass "a merge that cannot lock the away record refuses instead of merging unlocked" +} + test_allow_red_refused_on_gitlab() { local case_dir rc case_dir=$(make_gitlab_case gitlab-allow-red) @@ -2786,6 +3038,10 @@ test_allow_red_still_waives_only_the_current_failure test_allow_red_is_refused_while_away test_allow_red_requires_one_separate_name test_away_grant_and_yolo_and_hold_for_return +test_away_posture_refuses_asynchronous_merge_paths test_away_grant_does_not_bypass_red_or_identity test_unreadable_away_record_refuses_merge +test_away_record_cannot_change_between_the_authority_read_and_the_merge +test_a_grant_revoked_before_the_merge_refuses_it +test_merge_refuses_when_the_away_record_cannot_be_locked test_allow_red_refused_on_gitlab From 35761484e9f0a692ebec86532ff4dd252b224cf9 Mon Sep 17 00:00:00 2001 From: AnPod <drejc83@gmail.com> Date: Sun, 13 Sep 2026 00:50:23 +0200 Subject: [PATCH 23/31] feat(bin): add Antigravity CLI (agy) as third worker/scout adapter (#4200) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * feat(agy): verify Antigravity CLI as third worker/scout adapter Detection by anchored ancestry in fm-harness.sh (no marker of its own); bootstrap harness and effort validation; launch template with model and effort mapping plus reachable-catalog model validation; rendered-tail busy fallback in fm-busy-lib.sh with delivery footer in fm-composer-lib.sh; control mechanics with crewmate/scout-only refusal; tmux liveness naming; router entry with concise adapter reference; dated verification record; portable regression plus opt-in live drift guard. Verified live on agy 1.2.0: supervised spawn, durable steering, same-copy relaunch, and exit, with Herdr-native busy agreement. * no-mistakes(review): bound agy model probe, gate trust dialog, narrow busy signature * no-mistakes(review): pre-register agy workspace trust, make readiness gate strict * no-mistakes(review): Close Orca terminal on gate failure; isolate live-guard HOME; tighten agy matching * no-mistakes(document): Document agy adapter in stale harness enumerations * no-mistakes(review): Clamp non-positive FM_AGY_MODELS_TIMEOUT to the default bound * no-mistakes(document): Fix stale test-shard snapshots after agy lane additions * no-mistakes(ci): Fixed ci-3 (tests/fm-agy-harness.test.sh:519). Root cause: the agy spawn fixture's default base PATH (/usr/bin:/bin:/usr/sbin:/sbin) omits node's directory, but the spawn drives the real bin/fm-agy-trust.sh (which hard-requires node to record trust) and the fixture's fake tmux trust lookup (node -e) under that PATH. On the ubuntu-latest CI runner node lives in the toolcache (/usr/local/bin), so trust pre-registration failed on portable serial 2; on typical Arch hosts node is in /usr/bin, masking the defect. Fix (smallest, following the existing tests/fm-kimi-harness.test.sh precedent of carrying the interpreter's resolved directory): resolve node from the invoking environment (failing the test with 'test needs node' if absent, as kimi does for python3) and prepend its directory to the fixture's default base PATH; the FM_TEST_BASE_PATH override contract is untouched. Verified locally: (1) pre-fix reproduction with a CI-shaped base PATH (system bins minus node) produced exactly the reported failure — 'node is required to record workspace trust and was not found on PATH' plus the fake tmux 'node: command not found'; (2) post-fix, all 29 tests in the file pass both with node available only via a leading non-standard dir in the base PATH (CI's shape) and with the default base PATH on this host. bash -n clean; ShellCheck is not installed in this worktree (previously recorded as environmental) * no-mistakes(test): Give agy typed sends a longer submit-confirm budget * no-mistakes(document): Document agy send budget, trust gate, and control coverage * no-mistakes(document): Document agy busy fallback inventory and send-timing evidence --- .agents/skills/harness-adapters/SKILL.md | 7 +- .../references/common/control-and-recovery.md | 1 + .../references/harness/agy.md | 55 ++ AGENTS.md | 2 +- CONTRIBUTING.md | 2 +- bin/fm-agent-process-lib.sh | 5 + bin/fm-agy-trust.sh | 185 ++++ bin/fm-bootstrap.sh | 3 +- bin/fm-busy-lib.sh | 55 +- bin/fm-composer-lib.sh | 17 +- bin/fm-control-lib.sh | 26 +- bin/fm-harness.sh | 11 +- bin/fm-send.sh | 25 +- bin/fm-spawn.sh | 203 +++- bin/fm-test-run.sh | 6 +- docs/agent-control.md | 6 +- docs/architecture.md | 4 +- docs/configuration.md | 3 +- docs/documentation-audiences.json | 8 + docs/fm-test-portable-shards.md | 18 +- docs/tmux-backend.md | 3 +- docs/trace-context.md | 2 +- docs/verification/agy.md | 170 ++++ tests/fm-agy-harness.test.sh | 905 ++++++++++++++++++ tests/fm-agy-signals-live-e2e.test.sh | 193 ++++ tests/fm-bootstrap.test.sh | 4 + tests/fm-send-agy-confirm.test.sh | 165 ++++ 27 files changed, 2012 insertions(+), 72 deletions(-) create mode 100644 .agents/skills/harness-adapters/references/harness/agy.md create mode 100755 bin/fm-agy-trust.sh create mode 100644 docs/verification/agy.md create mode 100755 tests/fm-agy-harness.test.sh create mode 100755 tests/fm-agy-signals-live-e2e.test.sh create mode 100755 tests/fm-send-agy-confirm.test.sh diff --git a/.agents/skills/harness-adapters/SKILL.md b/.agents/skills/harness-adapters/SKILL.md index be447a41d7a..0c3a0353411 100644 --- a/.agents/skills/harness-adapters/SKILL.md +++ b/.agents/skills/harness-adapters/SKILL.md @@ -3,7 +3,7 @@ name: harness-adapters description: >- Agent-only reference for firstmate harness operations. Use before spawning or recovering a crewmate or secondmate, handling a trust dialog, sending a harness-specific skill invocation, interrupting or exiting an agent, resuming an exited agent, or verifying a new harness adapter. - Contains verified facts for claude, codex, opencode, pi, pi-signed, grok, kimi, cursor, gemini, muse, rovo, and omp. + Contains verified facts for claude, codex, opencode, pi, pi-signed, grok, kimi, cursor, gemini, muse, rovo, omp, and agy. user-invocable: false metadata: internal: true @@ -35,7 +35,7 @@ For recovery and control, use the exact `harness=` in `state/<id>.meta`; never i Deliver lifecycle actions only through `../../../bin/fm-control.sh <task-id> interrupt|exit|relaunch`. Never type an interrupt key or exit command through `fm-send`, where routing-marked lifecycle text becomes chat. Trust handling is complete only when inspection proves the target started processing its instructions; delivery success alone is not proof. -Muse and Gemini are verified only for crewmate and scout work, never a secondmate or primary. +Muse, Gemini, and AGY are verified only for crewmate and scout work, never a secondmate or primary. ## Detection @@ -93,7 +93,8 @@ A new tool remains undispatchable until the `verify` plan, its harness entry, ev "gemini": "references/harness/gemini.md", "muse": "references/harness/muse.md", "rovo": "references/harness/rovo.md", - "omp": "references/harness/omp.md" + "omp": "references/harness/omp.md", + "agy": "references/harness/agy.md" } } ``` diff --git a/.agents/skills/harness-adapters/references/common/control-and-recovery.md b/.agents/skills/harness-adapters/references/common/control-and-recovery.md index 78115e47174..4223b63b895 100644 --- a/.agents/skills/harness-adapters/references/common/control-and-recovery.md +++ b/.agents/skills/harness-adapters/references/common/control-and-recovery.md @@ -18,6 +18,7 @@ No observed dialog proves only that launch. Each supported harness handles its folder-trust gate differently, and the tool reference owns the detail. For Claude, load `references/harness/claude.md`; its workspace-trust section owns the non-key-answerable gate and spawn-time pre-registration for every spawn kind. +agy gates every fresh worktree too; the spawn pre-registers it in agy's own store the same way, and a strict post-launch gate answers any dialog that still renders before the spawn reports success. Cursor suppresses its dialog with launch-time `--trust`, and Muse suppresses its own with `--yolo`. Grok dodges its gate instead of granting trust, because its project picker appears only outside a project and the spawn starts in the isolated git root. Pi gates the fresh-worktree case too, but unlike Claude its dialog is answered with Enter, and `references/harness/pi.md` owns that recipe and where the decision persists. diff --git a/.agents/skills/harness-adapters/references/harness/agy.md b/.agents/skills/harness-adapters/references/harness/agy.md new file mode 100644 index 00000000000..0f38ca16ba6 --- /dev/null +++ b/.agents/skills/harness-adapters/references/harness/agy.md @@ -0,0 +1,55 @@ +# Antigravity CLI + +Antigravity's `agy` TUI, verified end to end on 2026-09-10 with agy 1.2.0 on Linux through the Herdr backend. +Verified as a CREWMATE and SCOUT adapter only; `../../../../../bin/fm-spawn.sh` refuses a secondmate launch on it because `../../../../../docs/supervision-protocols/` carries no agy wake protocol. +`../../../../../docs/verification/agy.md` owns how every fact below was established and what is still unproven. + +## Operating facts + +| Fact | Value | +|---|---| +| Binary | Absolute `agy` from `PATH`, refused if absent; a Go-compiled single binary, so the live process name is exactly `agy` with `argv[0]=agy`. | +| Launch | `agy --prompt-interactive "<brief>" --model <id> --effort <level> --dangerously-skip-permissions`, with the resolved absolute binary; the brief auto-submits with no extra Enter. The spawn pre-registers the worktree in agy's trust store first, then waits for a busy turn (answering the folder-trust dialog if it renders anyway) before reporting success. | +| Busy state | No hook or plugin writer, so nothing is armed and no record is seeded; on Herdr the native `working` status classifies busy, and everywhere else the `agy-regex` rendered-tail fallback in `../../../../../bin/fm-busy-lib.sh` does. | +| Rendered tail | Busy status row carries `esc to cancel` on the left; the idle row shows `? for shortcuts` instead. The `Generating...` word beside the braille spinner is free-floating output and is not a signal. | +| Turn end | No turn-end hook or notification touch exists; completion arrives through the worker status protocol and, on Herdr, the native return to `idle`. | +| Exit | `/quit`, one Enter; the process exits. | +| Interrupt | Single `Escape`, which prints the Interrupted row and leaves an idle composer with no repollution, so no clear key follows. | +| Skill | No verified slash-skill form; use natural language. | +| Autonomy | `--dangerously-skip-permissions` auto-approves tool calls for the run. | +| Marker | None; a live TUI carries no `AGY_*` or `ANTIGRAVITY_*` variable. | +| Resume | `--continue` and `--conversation` exist but carry no verified pane-resume contract; use deterministic relaunch. | +| Model | `--model <id>` with the bare catalog id from `agy models` (for example `gemini-3.8-flash-high`); `bin/fm-spawn.sh` refuses a requested id a reachable listing omits. The listing is a remote fetch, so the probe runs stdin-detached under the shared hard bound and an unreachable or hung listing launches unvalidated with a notice. | +| Effort | `--effort low\|medium\|high`; `xhigh` and `max` stay in task metadata under the record-and-omit contract. | +| Composer | Borderless bare `>` row, which the shared classifier reads as `unknown` under the dead-shell rule, never `empty`; steering confirms delivery through native agent-state and the delivery footer instead, the cursor precedent. | + +## Trust, and where the decision persists + +Every task worktree is a path agy has never seen, so an unregistered launch stops on `Do you trust the contents of this project?` with the safe choice `Yes, I trust this folder` preselected, and an unanswered dialog sends the turn into agy's scratch directory instead of the worktree. +There is no launch flag that suppresses the dialog, but agy honours a `trustedWorkspaces` entry in the captain's own `~/.gemini/antigravity-cli/settings.json` written ahead of launch (verified live), so `../../../../../bin/fm-spawn.sh` pre-registers the worktree through `../../../../../bin/fm-agy-trust.sh` before launch, the claude shape: the helper refuses anything but a linked worktree of the spawning project, records both the logical pane path and its resolved form because agy compares the logical cwd, and preserves every other key in the store. +The post-launch readiness gate is the backstop: it answers a dialog that renders anyway with a single Enter, then requires a busy verdict (Herdr's native `working` status or the pinned `esc to cancel` row) before the spawn reports success, and on a path that was not pre-registered it never counts a busy verdict as ready until the dialog has been answered, because Herdr's native verdict can precede the dialog. +A pane whose brief cannot be confirmed to run in the worktree fails the spawn, records the failure in the task status, and closes the endpoint. +Never steer into a pane still showing the dialog; a spawn that reported success has already cleared it. + +## Credential precondition + +A verified agy worker ran under a signed-in Google account with no key export and no dialog. +The unauthenticated failure mode was not observed, so treat any auth prompt or refusal as a credential blocker under `../../../../../AGENTS.md` section 9, fix the environment, and retire the endpoint rather than typing into it. + +## Detection + +Detected by ancestry alone: `../../../../../bin/fm-harness.sh` matches the anchored process name `agy`, never `*agy*`. +No environment marker is promoted: `AGENT=1` observed on a live TUI is an inherited launcher value, not an agy identity, and agy does not clear an inherited `CLAUDECODE`, so the spawn clears foreign markers at the launch boundary and the ancestry arm decides. +agy is deliberately absent from the session-lock name vocabulary in `../../../../../bin/fm-session-lock-lib.sh`, where muse, gemini, and rovo are also absent: a crewmate-only adapter must never own a home session lock. + +## Worker busy state and turn end + +`../../../../../bin/fm-spawn.sh` arms no busy generation for agy and writes no sidecar, exactly because no writer could ever clear a seeded record. +`fm_busy_agy_tail_busy` matches the pinned `esc to cancel` status row alone, hardcoded with no environment override, and `fm_busy_classify` reports `unknown agy-regex` rather than idle when it is absent, because a long turn can scroll the marker out of the captured tail. +Teardown removes nothing agy-specific because the spawn leaves nothing behind. + +## Primary integration + +Unsupported and unverified. +`../../../../../docs/supervision-protocols/` carries no agy protocol, no turn-end guard adapter exists for it, and this adapter verified only the crewmate-side launch, busy state, interrupt, and exit. +`references/common/primary-hooks.md`'s unsupported-boundary rule applies: never invent a wake protocol from a similar TUI. diff --git a/AGENTS.md b/AGENTS.md index 135831d62f5..7d58297bb07 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -209,7 +209,7 @@ A silent bootstrap section needs no action; for any printed actionable diagnosti ## 4. Harness and runtime dispatch Load `harness-adapters` before every spawn or recovery and before trust handling, skill invocation, interrupt, exit, resume, or adapter verification. -The verified harnesses are `claude`, `codex`, `opencode`, `pi`, `pi-signed`, `grok`, `kimi`, `cursor`, and `omp`, plus `muse`, `gemini`, and `rovo` for crewmates and scouts only; never dispatch on an unverified adapter. +The verified harnesses are `claude`, `codex`, `opencode`, `pi`, `pi-signed`, `grok`, `kimi`, `cursor`, and `omp`, plus `muse`, `gemini`, `rovo`, and `agy` for crewmates and scouts only; never dispatch on an unverified adapter. If static `config/crew-harness` or `config/secondmate-harness` names an unverified adapter, report it and fall back only to a verified adapter rather than launching it. `docs/configuration.md` owns dispatch-profile and runtime-backend schemas, `bin/fm-harness.sh` owns static resolution, and `bin/fm-spawn.sh` owns launch flags and fail-closed validation. diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 87317bd5c45..de4cc759a35 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -52,7 +52,7 @@ See the [no-mistakes quick start](https://kunchenguid.github.io/no-mistakes/star It pins one exact shellcheck version and one exact actionlint version and refuses to run under any other. Print the shellcheck pin with `bin/fm-lint.sh --required-version` and the actionlint pin with `bin/fm-lint-workflows.sh --required-version`. Use `bin/fm-install-shellcheck.sh` and `bin/fm-install-actionlint.sh` to install those exact builds locally; each installer's header owns its destination usage and supported platforms. -- Harness-adapter ownership spans detection in `bin/fm-harness.sh`, launch and hook mechanics in `bin/fm-spawn.sh`, spawn-time Claude workspace-trust pre-registration in `bin/fm-claude-trust.sh`, semantic busy sources and trust gates in `bin/fm-busy-lib.sh`, delivery-only rendered guards in `bin/fm-composer-lib.sh`, cleanup in `bin/fm-teardown.sh`, and facts in the skill tree rooted at `.agents/skills/harness-adapters/SKILL.md`; the `firstmate-coding-guidelines` skill owns the validation policy for checks that depend on those harnesses. +- Harness-adapter ownership spans detection in `bin/fm-harness.sh`, launch and hook mechanics in `bin/fm-spawn.sh`, spawn-time workspace-trust pre-registration in `bin/fm-claude-trust.sh` and `bin/fm-agy-trust.sh`, semantic busy sources and trust gates in `bin/fm-busy-lib.sh`, delivery-only rendered guards in `bin/fm-composer-lib.sh`, cleanup in `bin/fm-teardown.sh`, and facts in the skill tree rooted at `.agents/skills/harness-adapters/SKILL.md`; the `firstmate-coding-guidelines` skill owns the validation policy for checks that depend on those harnesses. - Changes to runtime session backends (`bin/fm-backend.sh`, `bin/backends/`, and the scripts that dispatch through them) keep current setup and limits in the relevant backend guide and active empirical evidence in [`docs/verification/runtime-backends.md`](docs/verification/runtime-backends.md). - [`docs/documentation-audiences.md`](docs/documentation-audiences.md) and its machine-consumed inventory own prose classification; run `bin/fm-doc-audience-check.sh` after documentation changes. - In Markdown, put each full sentence on its own line. diff --git a/bin/fm-agent-process-lib.sh b/bin/fm-agent-process-lib.sh index 5943f1b2749..dcf4b59ff4e 100644 --- a/bin/fm-agent-process-lib.sh +++ b/bin/fm-agent-process-lib.sh @@ -41,6 +41,11 @@ fm_agent_process_classify_name() { # <path> [argv0] -> agent|shell|other # name is the bare word `omp` (verified, omp 18.1.11) and a glob would claim # unrelated commands such as ompd or comp. *claude*|*codex*|*opencode*|*grok*|*kimi*|*rovo*|pi|pi-signed|pi-launcher|Pi|omp) printf 'agent' ;; + # agy (Antigravity CLI) is anchored for the same reason as muse and omp: its + # live process name is the bare word `agy` (verified, agy 1.2.0: a Go-compiled + # single binary, comm=agy with argv[0]=agy), and a glob would claim + # unrelated commands containing that fragment. + agy) printf 'agent' ;; zsh|bash|sh|dash|ash|ksh|mksh|tcsh|csh|fish) printf 'shell' ;; *) if fm_harness_path_name "$path" >/dev/null || fm_harness_path_name "$argv0" >/dev/null; then diff --git a/bin/fm-agy-trust.sh b/bin/fm-agy-trust.sh new file mode 100755 index 00000000000..627d892da2f --- /dev/null +++ b/bin/fm-agy-trust.sh @@ -0,0 +1,185 @@ +#!/usr/bin/env bash +# Pre-register Antigravity CLI's workspace trust for the isolated task worktree +# a ship/scout spawn is about to launch an agy crewmate into, so the worker +# reaches its brief in the worktree instead of parking on the folder-trust +# dialog and running its turn in agy's own scratch directory. +# +# Usage: fm-agy-trust.sh <worktree> <project> +# <worktree> the isolated task worktree this spawn launches into +# <project> the primary checkout that worktree belongs to +# Prints one line naming what it registered; refuses loudly on anything else. +# +# WHY THIS EXISTS. agy 1.2.0 gates a folder it has never seen behind +# "Do you trust the contents of this project?" and no launch flag suppresses +# it (`agy --help` lists none). Answering appends the folder to the +# `trustedWorkspaces` array of ${HOME}/.gemini/antigravity-cli/settings.json, +# and agy honours an entry written there ahead of launch: verified live under a +# throwaway HOME, a pre-registered folder launched straight into its turn while +# an unregistered sibling parked on the dialog (docs/verification/agy.md). agy +# compares the pane's LOGICAL working directory, not its resolved path (a +# symlinked cwd with only the real path registered still parked), so both the +# logical path and its resolved form are recorded when they differ. +# +# bin/fm-spawn.sh keeps a post-launch gate as the backstop: it answers the +# dialog if one renders anyway and never counts a busy turn as ready on a path +# that was neither pre-registered here nor answered there. +# +# THE SCOPE TEST IS THE SAFETY PROPERTY and mirrors bin/fm-claude-trust.sh: +# <worktree> must be a LINKED git worktree - its own git dir, sharing +# <project>'s common dir - whose top level is exactly the resolved argument. A +# primary checkout, a worktree of an unrelated repo, a subdirectory of a +# worktree, a plain directory, and a home directory are each refused with a +# non-zero exit, never a warning and never a silent skip. Only the launching +# user's own store is written, it must be a regular file this uid owns, every +# unrelated key and entry is preserved, and the replacement is atomic. +set -u +unset CDPATH \ + GIT_DIR GIT_WORK_TREE GIT_COMMON_DIR GIT_OBJECT_DIRECTORY GIT_INDEX_FILE \ + GIT_ALTERNATE_OBJECT_DIRECTORIES GIT_CEILING_DIRECTORIES GIT_NAMESPACE \ + GIT_DISCOVERY_ACROSS_FILESYSTEM GIT_CONFIG GIT_CONFIG_GLOBAL \ + GIT_CONFIG_SYSTEM GIT_CONFIG_NOSYSTEM GIT_CONFIG_COUNT + +[ "$#" -eq 2 ] || { echo "usage: fm-agy-trust.sh <worktree> <project>" >&2; exit 2; } +WT_ARG=$1 +PROJ_ARG=$2 + +refuse() { echo "error: refusing to pre-register agy trust: $1" >&2; exit 1; } + +real_dir() { (cd -P -- "$1" 2>/dev/null && pwd -P); } +logical_dir() { (cd -- "$1" 2>/dev/null && pwd -L); } +real_file() { node -e 'process.stdout.write(require("node:fs").realpathSync(process.argv[1]))' "$1" 2>/dev/null; } + +common_dir_of() { + local dir=$1 common + common=$(git -C "$dir" rev-parse --git-common-dir 2>/dev/null) || return 1 + (cd -P -- "$dir" && real_dir "$common") +} + +WT_REAL=$(real_dir "$WT_ARG") || true +[ -n "$WT_REAL" ] || refuse "worktree '$WT_ARG' is not an accessible directory" +WT_LOGICAL=$(logical_dir "$WT_ARG") || true +[ -n "$WT_LOGICAL" ] || WT_LOGICAL=$WT_REAL +PROJ_REAL=$(real_dir "$PROJ_ARG") || true +[ -n "$PROJ_REAL" ] || refuse "project '$PROJ_ARG' is not an accessible directory" + +[ -n "${HOME:-}" ] || refuse "HOME is not set, so agy's settings store cannot be located" +HOME_REAL=$(real_dir "$HOME") || true +[ -n "$HOME_REAL" ] || refuse "HOME '$HOME' is not an accessible directory" +[ "$WT_REAL" != "$HOME_REAL" ] || refuse "'$WT_REAL' is the home directory, not a task worktree" + +WT_TOP=$(git -C "$WT_REAL" rev-parse --show-toplevel 2>/dev/null) || true +[ -n "$WT_TOP" ] || refuse "'$WT_REAL' is not inside a git repository" +WT_TOP_REAL=$(real_dir "$WT_TOP") || true +[ "$WT_TOP_REAL" = "$WT_REAL" ] || refuse "'$WT_REAL' is not a worktree root (its root is '${WT_TOP_REAL:-unresolvable}')" + +WT_GIT_DIR=$(git -C "$WT_REAL" rev-parse --absolute-git-dir 2>/dev/null) || true +[ -n "$WT_GIT_DIR" ] || refuse "'$WT_REAL' has no resolvable git directory" +WT_GIT_DIR=$(real_dir "$WT_GIT_DIR") || true +[ -n "$WT_GIT_DIR" ] || refuse "'$WT_REAL' has an unresolvable git directory" +WT_COMMON=$(common_dir_of "$WT_REAL") || true +[ -n "$WT_COMMON" ] || refuse "'$WT_REAL' has no resolvable git common directory" +[ "$WT_GIT_DIR" != "$WT_COMMON" ] || refuse "'$WT_REAL' is a primary checkout, not an isolated worktree" + +PROJ_COMMON=$(common_dir_of "$PROJ_REAL") || true +[ -n "$PROJ_COMMON" ] || refuse "project '$PROJ_REAL' is not inside a git repository" +[ "$WT_COMMON" = "$PROJ_COMMON" ] || refuse "'$WT_REAL' is not a worktree of project '$PROJ_REAL'" + +command -v node >/dev/null 2>&1 || refuse "node is required to record workspace trust and was not found on PATH" + +STORE_DIR="$HOME_REAL/.gemini/antigravity-cli" +mkdir -p "$STORE_DIR" 2>/dev/null || true +STORE_DIR_REAL=$(real_dir "$STORE_DIR") || true +[ -n "$STORE_DIR_REAL" ] || refuse "agy settings directory '$STORE_DIR' does not exist and could not be created" +STORE="$STORE_DIR_REAL/settings.json" +if [ -L "$STORE" ]; then + STORE_REAL=$(real_file "$STORE") || true + [ -n "$STORE_REAL" ] || refuse "'$STORE' is a symlink whose target cannot be resolved" + STORE=$STORE_REAL +fi +if [ -e "$STORE" ]; then + [ -f "$STORE" ] || refuse "'$STORE' is not a regular file" + [ -O "$STORE" ] || refuse "'$STORE' is not owned by this user" + [ -w "$STORE" ] || refuse "'$STORE' is not writable" +fi + +# Read-modify-write with a fingerprint check before the rename and a readback +# after it, the bin/fm-claude-trust.sh shape: agy itself rewrites this file +# when a worker answers a dialog or changes a setting, so a store that moved +# under us is retried once and then refused rather than clobbered. +if ! node - "$STORE" "$WT_LOGICAL" "$WT_REAL" <<'NODE' +const fs = require("node:fs"); +const path = require("node:path"); +const crypto = require("node:crypto"); +const [store, ...wanted] = process.argv.slice(2); +const paths = [...new Set(wanted)]; +const readStore = () => { + try { + return fs.readFileSync(store); + } catch (err) { + if (err.code === "ENOENT") return null; + throw err; + } +}; +const fingerprint = (buf) => + buf === null ? "absent" : crypto.createHash("sha256").update(buf).digest("hex"); +const listed = (root) => + Array.isArray(root.trustedWorkspaces) && paths.every((p) => root.trustedWorkspaces.includes(p)); +const attempt = () => { + const original = readStore(); + const before = fingerprint(original); + let root = {}; + if (original !== null) { + const raw = original.toString("utf8"); + if (raw.trim() !== "") { + root = JSON.parse(raw); + if (root === null || typeof root !== "object" || Array.isArray(root)) { + throw new Error(`${store} is not a JSON object`); + } + } + } + if (root.trustedWorkspaces === undefined || root.trustedWorkspaces === null) root.trustedWorkspaces = []; + if (!Array.isArray(root.trustedWorkspaces)) { + throw new Error(`${store} has a non-array "trustedWorkspaces" value`); + } + if (listed(root)) return "recorded"; + for (const p of paths) { + if (!root.trustedWorkspaces.includes(p)) root.trustedWorkspaces.push(p); + } + const unique = `${process.pid}.${crypto.randomBytes(8).toString("hex")}`; + const tmp = path.join(path.dirname(store), `.settings.json.fm-trust.${unique}`); + fs.writeFileSync(tmp, `${JSON.stringify(root, null, 2)}\n`, { mode: 0o600, flag: "wx" }); + let renamed = false; + try { + if (fingerprint(readStore()) !== before) return "moved"; + fs.renameSync(tmp, store); + renamed = true; + } finally { + if (!renamed) fs.rmSync(tmp, { force: true }); + } + return listed(JSON.parse(fs.readFileSync(store, "utf8"))) ? "recorded" : "dropped"; +}; +try { + for (let i = 0; i < 3; i += 1) { + const result = attempt(); + if (result === "recorded") process.exit(0); + if (result === "moved" && i >= 1) { + console.error(`error: ${store} was modified while trust was being recorded; refusing to overwrite it`); + process.exit(1); + } + } +} catch (err) { + console.error(`error: ${err.message}`); + process.exit(1); +} +console.error(`error: ${store} did not retain trust for ${paths.join(", ")} after 3 attempts`); +process.exit(1); +NODE +then + refuse "could not record trust for '$WT_LOGICAL' in '$STORE'" +fi + +if [ "$WT_LOGICAL" != "$WT_REAL" ]; then + echo "trusted: $WT_LOGICAL ($WT_REAL)" +else + echo "trusted: $WT_REAL" +fi diff --git a/bin/fm-bootstrap.sh b/bin/fm-bootstrap.sh index b5e010905cf..98b791e52f7 100755 --- a/bin/fm-bootstrap.sh +++ b/bin/fm-bootstrap.sh @@ -1114,7 +1114,7 @@ crew_dispatch_validate() { return 0 fi err=$(jq -r ' - def verified($h): ["claude","codex","opencode","pi","pi-signed","grok","kimi","cursor","muse","rovo","omp"] | index($h); + def verified($h): ["claude","codex","opencode","pi","pi-signed","grok","kimi","cursor","agy","muse","rovo","omp"] | index($h); def effort_ok($h; $m; $e): if $e == null then true elif ($e | type) != "string" then false @@ -1122,6 +1122,7 @@ crew_dispatch_validate() { elif $h == "claude" then (["low","medium","high","xhigh","max"] | index($e)) elif $h == "codex" then (["low","medium","high","xhigh"] | index($e)) elif $h == "grok" then (["low","medium","high"] | index($e)) + elif $h == "agy" then (["low","medium","high"] | index($e)) elif $h == "pi" or $h == "pi-signed" or $h == "omp" then (["low","medium","high","xhigh","max"] | index($e)) elif $h == "muse" then (["low","medium","high","xhigh","max"] | index($e)) elif $h == "rovo" then (["low","medium","high","max"] | index($e)) diff --git a/bin/fm-busy-lib.sh b/bin/fm-busy-lib.sh index dce79a17941..d8f7a0ee111 100755 --- a/bin/fm-busy-lib.sh +++ b/bin/fm-busy-lib.sh @@ -42,7 +42,7 @@ # fm-interrupt the legacy Claude fm-send --key Escape idle event # fm-recovery a documented recovery reset after relaunch # Classifier-only sources (never written into a record): -# endpoint-gone, herdr-native, grok-regex, rovo-regex, muse-session-log, +# endpoint-gone, herdr-native, grok-regex, rovo-regex, agy-regex, muse-session-log, # cursor-transcript, missing, malformed, gen-mismatch, source-mismatch, # kimi-unverified, codex-unverified, capture-failed, no-target # @@ -53,14 +53,15 @@ # 3. a valid, gen-matching, source-trusted record -> its state and source # 4. no record at all: herdr's native busy verdict is trusted as busy # (generation state is sufficient for busy, not for idle), then the -# muse session-log and cursor transcript pull sources, then the Grok/Rovo -# temporary regex fallbacks classify a grok or rovo task from its -# rendered tail, then unknown missing +# muse session-log and cursor transcript pull sources, then the +# Grok/Rovo/AGY temporary regex fallbacks classify a grok, rovo, or agy +# task from its rendered tail, then unknown missing # 5. malformed, stale, or untrusted records -> unknown, never a fallback -# Grok and Rovo are the ONLY rendered-text classifications that survive the -# redesign, because neither's structured lifecycle was credited-live-verified +# Grok, Rovo, and AGY are the ONLY rendered-text classifications that survive the +# redesign, because none of their structured lifecycles was credited-live-verified # in the approved audit (Rovo's clean ACP stopReason lives outside the TUI -# path firstmate drives, see references/harness/rovo.md); each is scoped to +# path firstmate drives, see references/harness/rovo.md; agy 1.2.0 exposes no +# hook surface at all, see references/harness/agy.md); each is scoped to # its own harness= and can never classify another adapter. The delivery # guards in bin/fm-composer-lib.sh match rendered footers for submit # acknowledgement and away-mode supervisor injection only; neither is a @@ -851,12 +852,27 @@ fm_busy_rovo_tail_busy() { | grep -qiE "${FM_BUSY_ROVO_REGEX:-Rovo is thinking}" } +# fm_busy_agy_tail_busy: the AGY-only temporary rendered-tail fallback. +# Consumes the tail on stdin; 0 when AGY's verified busy signature matches: +# the `esc to cancel` token in the status row the TUI pins to the bottom of +# the pane while a turn runs (verified live on agy 1.2.0; the idle status row +# shows `? for shortcuts` instead). The `Generating...` spinner word that +# renders beside it is deliberately NOT matched: it is a free-floating output +# line, so ordinary worker output echoing the word would classify an idle +# worker as busy. agy exposes no hook surface, so this fallback is the only +# pane-side source; it is never armed as a semantic writer +# (fm_busy_sources_for_harness trusts nothing for agy). +fm_busy_agy_tail_busy() { + grep -v '^[[:space:]]*$' | tail -12 \ + | grep -qiE 'esc[[:space:]]+to[[:space:]]+cancel' +} + # fm_busy_classify: semantic classification for a task whose endpoint the # caller has already established as present. Prints "<verdict> <source>": # busy|idle|unknown plus the producing source (see header). Never probes # process state. <tail40> is optional pre-captured plain output used only by -# the Grok arm; when absent the Grok arm captures through fm_backend_capture -# if available, else reports unknown capture-failed. +# the grok, rovo, and agy arms; when absent each captures through +# fm_backend_capture if available, else reports unknown capture-failed. fm_busy_classify() { # <backend> <target> <harness> <id> <state-dir> [tail40] local backend=$1 target=$2 harness=$3 id=$4 state=$5 tail40=${6-} local out rc r_state r_source native log @@ -979,6 +995,27 @@ fm_busy_classify() { # <backend> <target> <harness> <id> <state-dir> [tail40] fi return 0 ;; + agy) + if [ -z "$tail40" ]; then + if command -v fm_backend_capture >/dev/null 2>&1; then + tail40=$(fm_backend_capture "$backend" "$target" 40 2>/dev/null) || { + printf 'unknown capture-failed' + return 0 + } + else + printf 'unknown capture-failed' + return 0 + fi + fi + # Best-effort like rovo: a long turn can scroll the busy marker out of + # the captured tail, so its absence means "can't tell," never idle. + if printf '%s' "$tail40" | fm_busy_agy_tail_busy; then + printf 'busy agy-regex' + else + printf 'unknown agy-regex' + fi + return 0 + ;; esac printf 'unknown missing' } diff --git a/bin/fm-composer-lib.sh b/bin/fm-composer-lib.sh index cdea8d98abd..058dadc7293 100644 --- a/bin/fm-composer-lib.sh +++ b/bin/fm-composer-lib.sh @@ -290,7 +290,7 @@ fm_composer_strip_ghost() { # Matching a footer to confirm a keystroke landed is a different question from # asking what a worker is doing, and the two must not be conflated. # Delivery-only rendered busy footers per harness. claude/codex: "esc to -# interrupt"; opencode: "esc interrupt"; pi: "Working..."; omp: "Working…"; grok: "Ctrl+c:cancel". +# interrupt"; opencode: "esc interrupt"; pi: "Working..."; omp: "Working…"; grok: "Ctrl+c:cancel"; agy: "esc to cancel". # Claude's current spinner has a rotating glyph and word, but every active-turn # line has an ellipsis followed by a parenthesized elapsed duration. Keep this # signature separate from the shared default because that shape is not generic @@ -311,7 +311,11 @@ fm_composer_strip_ghost() { # part of that union for the same reason the others are: without it a cursor # submit could never be acknowledged, because cursor parks its terminal cursor # outside its composer and the composer verdict is therefore always `unknown`. -FM_DELIVERY_BUSY_REGEX_DEFAULT='esc (to )?interrupt|Working(\.\.\.|…)|Ctrl\+c:cancel|ctrl\+c to stop' +# agy's `esc to cancel` is part of the union for the same reason: an explicit +# tmux agy endpoint reaches the submit core with no recorded harness, and its +# bare `>` composer verdict is `unknown`, so the busy footer is the only +# turn-started acknowledgement that path can read. +FM_DELIVERY_BUSY_REGEX_DEFAULT='esc (to )?interrupt|Working(\.\.\.|…)|Ctrl\+c:cancel|ctrl\+c to stop|esc[[:space:]]+to[[:space:]]+cancel' FM_DELIVERY_CLAUDE_BUSY_REGEX_DEFAULT='esc to interrupt|…[[:space:]]+\([0-9]+[smh]' FM_DELIVERY_CODEX_BUSY_REGEX_DEFAULT='esc to interrupt' FM_DELIVERY_OPENCODE_BUSY_REGEX_DEFAULT='esc interrupt' @@ -342,6 +346,14 @@ FM_DELIVERY_GROK_BUSY_REGEX_DEFAULT='Ctrl\+c:cancel' # injection. Cursor's recorded worker state comes from its transcript fold in # bin/fm-busy-lib.sh, never from this row. FM_DELIVERY_CURSOR_BUSY_REGEX_DEFAULT='ctrl\+c to stop' +# agy (Antigravity CLI) renders a pinned status row while a turn runs: the +# `esc to cancel` token on the left and the model cell on the right (verified +# live, agy 1.2.0; the idle row shows `? for shortcuts` instead). The +# `Generating...` spinner word beside it is a free-floating output line and is +# deliberately not matched, so echoed worker output cannot fake an +# acknowledgement. Delivery guard only; recorded worker state comes from the +# agy-regex fold in bin/fm-busy-lib.sh. +FM_DELIVERY_AGY_BUSY_REGEX_DEFAULT='esc[[:space:]]+to[[:space:]]+cancel' FM_DELIVERY_KIMI_BUSY_REGEX_DEFAULT='^[[:space:]]*(🌑|🌒|🌓|🌔|🌕|🌖|🌗|🌘)[[:space:]]+·[[:space:]]+' fm_busy_lines_match() { # [harness] @@ -357,6 +369,7 @@ fm_busy_lines_match() { # [harness] pi|pi-signed) regex=$FM_DELIVERY_PI_BUSY_REGEX_DEFAULT ;; omp) regex=$FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT ;; grok) regex=$FM_DELIVERY_GROK_BUSY_REGEX_DEFAULT ;; + agy) regex=$FM_DELIVERY_AGY_BUSY_REGEX_DEFAULT ;; kimi) regex=$FM_DELIVERY_KIMI_BUSY_REGEX_DEFAULT ;; cursor) regex=$FM_DELIVERY_CURSOR_BUSY_REGEX_DEFAULT ;; '') regex=$FM_DELIVERY_BUSY_REGEX_DEFAULT ;; diff --git a/bin/fm-control-lib.sh b/bin/fm-control-lib.sh index 96a7abc4860..516a00b4364 100644 --- a/bin/fm-control-lib.sh +++ b/bin/fm-control-lib.sh @@ -63,7 +63,7 @@ fm_control_verb_allowed() { # <verb> # than guessed at, exactly as a spawn on it would be. fm_control_harness_supported() { # <harness> case "${1-}" in - claude|codex|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp) return 0 ;; + claude|codex|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp|agy) return 0 ;; esac return 1 } @@ -74,13 +74,15 @@ fm_control_harness_supported() { # <harness> # harness= that way), which is why the spawn adapters match `claude*`, `muse*`, # and friends. This is the one place that prefix rule is stated. `pi` and # `pi-signed` are exact because a `pi*` prefix would swallow the signed adapter, -# `omp` is exact because an `omp*` prefix would claim unrelated commands, and an +# `omp` is exact because an `omp*` prefix would claim unrelated commands, `agy` +# is exact for the same reason on an even shorter name, and an # unrecognized value returns nonzero rather than being guessed into a family. fm_control_harness_family() { # <recorded-harness> case "${1-}" in pi) printf 'pi' ;; pi-signed) printf 'pi-signed' ;; omp) printf 'omp' ;; + agy) printf 'agy' ;; claude*) printf 'claude' ;; codex*) printf 'codex' ;; opencode*) printf 'opencode' ;; @@ -94,8 +96,8 @@ fm_control_harness_family() { # <recorded-harness> esac } -# Which task kinds an adapter is verified to run. muse, gemini, and rovo are -# crewmate/scout adapters only: none has a primary supervision protocol, +# Which task kinds an adapter is verified to run. muse, gemini, rovo, and agy +# are crewmate/scout adapters only: none has a primary supervision protocol, # and bin/fm-spawn.sh refuses a --secondmate launch on any of them. The control # plane asks this BEFORE it stops anything, so an incompatible relaunch target is # refused while the current agent is still running rather than after it has @@ -104,7 +106,7 @@ fm_control_harness_supports_kind() { # <harness> <kind> local harness=${1-} kind=${2-} fm_control_harness_supported "$harness" || return 1 case "$harness" in - muse|gemini|rovo) [ "$kind" != secondmate ] || return 1 ;; + muse|gemini|rovo|agy) [ "$kind" != secondmate ] || return 1 ;; esac return 0 } @@ -114,12 +116,14 @@ fm_control_harness_supports_kind() { # <harness> <kind> # gemini names its own key in the running turn's status row # (`(esc to cancel, <n>s)`), and a single Escape was verified to cancel it. # rovo cancels on a single Escape too, printing "Agent cancelled" (verified, -# 202609.1.2). omp (Oh My Pi) shares Pi's single Escape, empty composer +# 202609.1.2). agy cancels on a single Escape, printing the Interrupted row +# with an idle composer and no repollution (verified live, agy 1.2.0 through +# Herdr). omp (Oh My Pi) shares Pi's single Escape, empty composer # afterwards, and /quit exit (verified omp 18.1.2 in a PTY, re-verified 18.1.11 # through Herdr). fm_control_interrupt_key() { # <harness> case "${1-}" in - claude|codex|opencode|pi|pi-signed|omp|kimi|cursor|gemini|muse|rovo) printf 'Escape' ;; + claude|codex|opencode|pi|pi-signed|omp|kimi|cursor|gemini|muse|rovo|agy) printf 'Escape' ;; grok) printf 'C-c' ;; *) return 1 ;; esac @@ -130,7 +134,7 @@ fm_control_interrupt_key() { # <harness> fm_control_interrupt_repeat() { # <harness> case "${1-}" in opencode) printf '2' ;; - claude|codex|pi|pi-signed|omp|grok|kimi|cursor|gemini|muse|rovo) printf '1' ;; + claude|codex|pi|pi-signed|omp|grok|kimi|cursor|gemini|muse|rovo|agy) printf '1' ;; *) return 1 ;; esac } @@ -151,7 +155,7 @@ fm_control_interrupt_repeat() { # <harness> fm_control_interrupt_clear_key() { # <harness> case "${1-}" in muse) printf 'C-u' ;; - claude|codex|opencode|pi|pi-signed|omp|grok|kimi|cursor|gemini|rovo) ;; + claude|codex|opencode|pi|pi-signed|omp|grok|kimi|cursor|gemini|rovo|agy) ;; *) return 1 ;; esac } @@ -166,7 +170,7 @@ fm_control_interrupt_ack_source() { # <harness> # rovo's TUI prints "Agent cancelled" on Escape, but for parity with # claude/cursor this stays 'none': the ack is a rendered string, not a # recorded state source, and rovo has no busy wiring to confirm against. - claude|codex|opencode|pi|pi-signed|omp|grok|kimi|cursor|gemini|rovo) printf 'none' ;; + claude|codex|opencode|pi|pi-signed|omp|grok|kimi|cursor|gemini|rovo|agy) printf 'none' ;; *) return 1 ;; esac } @@ -175,7 +179,7 @@ fm_control_interrupt_ack_source() { # <harness> fm_control_exit_command() { # <harness> case "${1-}" in claude|opencode|grok|kimi|cursor|muse|rovo) printf '/exit' ;; - codex|pi|pi-signed|omp|gemini) printf '/quit' ;; + codex|pi|pi-signed|omp|gemini|agy) printf '/quit' ;; *) return 1 ;; esac } diff --git a/bin/fm-harness.sh b/bin/fm-harness.sh index 96443cf60c9..ecc19935197 100755 --- a/bin/fm-harness.sh +++ b/bin/fm-harness.sh @@ -1,6 +1,6 @@ #!/usr/bin/env bash # Detect the agent harness this process tree runs on. -# Usage: fm-harness.sh print own harness: claude|codex|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp|unknown +# Usage: fm-harness.sh print own harness: claude|codex|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp|agy|unknown # fm-harness.sh crew print the effective CREWMATE harness # (config/crew-harness; "default" resolves to own) # fm-harness.sh secondmate print the harness the PRIMARY uses to launch @@ -167,6 +167,15 @@ detect_own() { # named `claude` with its own node child, and that fallback's *claude* # args glob would otherwise claim it if that subtree were ever walked. omp) echo omp; return ;; + # agy (Antigravity CLI) is a Go-compiled single binary whose process name + # is exactly `agy` (verified, agy 1.2.0: `ps -o comm=` reports agy and + # Herdr's process-info reports name agy with argv[0] agy). Anchored, never + # *agy*, so unrelated commands cannot be misread as this harness. agy + # publishes no harness-identity marker of its own (a live 1.2.0 TUI + # carries no AGY_* or ANTIGRAVITY_* variable; AGENT=1 seen there is an + # inherited launcher value, not an agy identity), so like muse it is + # detected by ancestry alone. + agy) echo agy; return ;; node*|python*) # Bare interpreter: match the harness name in its script path. args=$(ps -o args= -p "$pid" 2>/dev/null) diff --git a/bin/fm-send.sh b/bin/fm-send.sh index aa74940c08e..885efff7002 100755 --- a/bin/fm-send.sh +++ b/bin/fm-send.sh @@ -70,10 +70,11 @@ # failure); any other nonzero = the send failed and nothing may be assumed # delivered. Submission dispatches through the target's recorded backend; the # tmux adapter shares its composer/submit core with the away-mode daemon via -# bin/fm-tmux-lib.sh. Tune with FM_SEND_RETRIES (default 3) / FM_SEND_SLEEP -# (0.4). Slash commands, and codex `$...` skill invocations resolved through -# harness meta, get a longer pre-Enter settle so completion popups do not -# swallow Enter. A remote secondmate target has no typed text plane at all: +# bin/fm-tmux-lib.sh. Tune with FM_SEND_RETRIES (default 3; agy typed targets +# default to 20 for agy's late busy render) / FM_SEND_SLEEP (0.4). Slash +# commands, and codex `$...` skill invocations resolved through harness meta, +# get a longer pre-Enter settle so completion popups do not swallow Enter. +# A remote secondmate target has no typed text plane at all: # every remote text steer rides the inbox (a marked secondmate request already # reaches the harness as marker-prefixed chat rather than a parser command, so # routing a remote "/..." or "$..." through the record changes nothing the @@ -1030,7 +1031,21 @@ else ;; *) settle=0.3 ;; esac - retries=${FM_SEND_RETRIES:-3} + # Per-harness submit-confirm budget. agy's bare `>` composer verdict is + # `unknown`, so a landed submit is acknowledged only by the idle-to-busy + # transition poll, and agy renders its verified busy footer well after the + # shared budget expires: ~1.5s after Enter for a short steer, ~4-5s for a + # realistic longer brief (live-measured, agy 1.2.1), against the shared + # default's 3 x 0.4s. With the shared default a typed steer to an agy + # endpoint was reported exit-1 non-delivery for a message that landed and + # ran, inviting a duplicate resend. agy typed targets get a longer default + # budget (~8s at the default cadence, twice the worst measured render); an + # explicit FM_SEND_RETRIES still wins, and every other harness keeps the + # shared 3-retry default untouched. + case "$TARGET_HARNESS" in + agy) retries=${FM_SEND_RETRIES:-20} ;; + *) retries=${FM_SEND_RETRIES:-3} ;; + esac sleep_s=${FM_SEND_SLEEP:-0.4} # Type once, submit, verify. Only exact empty confirms delivery; every other # verdict preserves the loud refusal boundary. Only LOCAL targets reach this diff --git a/bin/fm-spawn.sh b/bin/fm-spawn.sh index 13aaad7df50..2188298b37e 100755 --- a/bin/fm-spawn.sh +++ b/bin/fm-spawn.sh @@ -133,7 +133,7 @@ # profile consultation. A --secondmate spawn is exempt and resolves the SECONDMATE # harness (config/secondmate-harness -> config/crew-harness -> own), so the # secondmate-vs-crewmate split is DURABLE across every respawn (recovery, -# /updatefirstmate, restart). A bare adapter name (claude|codex|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp) +# /updatefirstmate, restart). A bare adapter name (claude|codex|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp|agy) # overrides it for this spawn (either kind). A non-flag string containing # whitespace is treated as a RAW launch command - the escape hatch for verifying # new adapters. For pi and pi-signed, fm-spawn resolves the selected executable @@ -281,6 +281,7 @@ # __CURSORBIN__ resolved, cursor-verified executable for a cursor launch # __GEMINISETTINGS__ firstmate-owned per-task gemini settings file (busy-state hooks) # __ROVOBIN__ resolved, rovo-verified executable for a rovo launch +# __AGYBIN__ resolved, agy-verified executable for an agy launch # Verified per-harness turn-end hooks are installed automatically where enabled; some live outside the worktree. # Kimi uses one surgically installed Firstmate region in $HOME/.kimi-code/config.toml, # a firstmate-owned global hook and registry, and a gitignored per-task pointer. @@ -288,7 +289,7 @@ # plus a gitignored .fm-grok-turnend worktree pointer and a state token. # muse installs no hook at all - its plugin engine is off in the default build - so # it writes state/<id>.muse-session to bind the pane to muse's own session event -# log; muse and gemini are crewmate/scout only and are refused for --secondmate. +# log; muse, gemini, and agy are crewmate/scout only and are refused for --secondmate. # rovo installs no hook either - its eventHooks fire at tool granularity only, # never turn-end - so it carries no busy-source wiring at all and no turn-end # hook. A positional brief is dead-on-arrival (rovo loads, never works, and drops @@ -296,6 +297,14 @@ # only after a TUI readiness gate, then a delivery-confirmation gate - the same # launch-then-send shape as kimi. Its busy state is a screen-scrape fallback like # grok. rovo is crewmate/scout only and is refused for --secondmate, like muse. +# agy installs no hook either - it exposes no hook surface at all - so it +# carries no busy-source wiring and no turn-end hook. Its brief rides the launch +# command, but a fresh worktree would park it on a folder-trust dialog, so the +# spawn pre-registers the worktree in agy's own trust store through +# bin/fm-agy-trust.sh (the claude shape, but non-fatal) and then waits for a +# busy turn - answering the dialog first if it renders anyway - before +# reporting success (the rovo/kimi launch-then-confirm shape). Its busy state +# is a screen-scrape fallback like grok and rovo, and it is crewmate/scout only. # cursor installs no per-task hook either: it writes state/<id>.cursor-session to # bind the pane to cursor's own conversation transcript (projects root, the exact # workspace path cursor records in .workspace-trusted, and the conversations that @@ -481,6 +490,8 @@ fm_backlog_directory_present "$STATE" "state directory" || { . "$SCRIPT_DIR/fm-trace-context-lib.sh" # shellcheck source=bin/fm-remote-readiness-lib.sh . "$SCRIPT_DIR/fm-remote-readiness-lib.sh" +# shellcheck source=bin/fm-timeout-lib.sh +. "$SCRIPT_DIR/fm-timeout-lib.sh" # Fail closed before any fleet mutation: a no-mistakes gate agent must never spawn # a direct report (see bin/fm-gate-refuse-lib.sh). fm_refuse_if_gate_agent @@ -1389,7 +1400,7 @@ if [ "$RELAUNCH" -eq 1 ]; then } elif [ "$KIND" = secondmate ]; then case "${POS[1]:-}" in - ''|claude|codex|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp) + ''|claude|codex|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp|agy) ARG3=${POS[1]:-} ;; *' '*) @@ -1467,6 +1478,37 @@ omp_model_validate() { # <omp-bin> <model> return 1 } +# agy pre-launch model validation. `agy models` (agy 1.2.0) prints one model per +# line as "<id>\t<label>" for the account's catalog only; model ids are bare +# (gemini-3.8-flash-high), never provider-prefixed. A requested model absent +# from a reachable listing is concrete unsupported evidence and refuses the +# spawn, so a stale id (the unlisted bare gemini-3.8-flash) fails loudly here +# instead of wedging a worker pane. The listing is a remote fetch that needs +# network and a signed-in account, so the probe runs under the shared hard +# bound (bin/fm-timeout-lib.sh) with stdin detached: a stalled fetch or a +# sign-in prompt can never block the spawn before any pane exists. An +# unreachable listing establishes nothing (harness-adapters +# model-and-effort.md) and launches unvalidated with a notice. +agy_model_validate() { # <agy-bin> <model> + local bin=$1 model=$2 listing rc=0 bound=${FM_AGY_MODELS_TIMEOUT:-15} + case "$bound" in ''|*[!0-9]*|0*) bound=15 ;; esac + [ -n "$model" ] && [ "$model" != default ] || return 0 + listing=$(fm_run_timed "$bound" "$bin" models 2>/dev/null < /dev/null) || rc=$? + if [ "$rc" -ne 0 ] || [ -z "$listing" ]; then + if [ "$rc" -eq 124 ]; then + echo "notice: 'agy models' did not answer within ${bound}s; launching with --model '$model' unvalidated" >&2 + else + echo "notice: 'agy models' listing is unreachable (exit $rc); launching with --model '$model' unvalidated" >&2 + fi + return 0 + fi + if printf '%s\n' "$listing" | awk '{print $1}' | grep -qxF -- "$model"; then + return 0 + fi + echo "error: agy model '$model' is not listed by 'agy models'; choose a listed id or omit --model" >&2 + return 1 +} + # The verified launch command per adapter. The knowledge half of each adapter # (busy-state source, exit command, dialogs, quirks) lives in the harness-adapters skill. launch_template() { @@ -1540,6 +1582,30 @@ launch_template() { printf '%s' ' __MODELFLAG____EFFORTFLAG__-e __OMPEXT__ "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' fi ;; + # agy (Antigravity CLI): --prompt-interactive "<brief>" starts the supervised + # interactive session and auto-submits it, so the brief rides the launch + # command (verified: a multi-line brief submitted itself with no extra Enter, + # agy 1.2.0). --model takes the bare catalog id from `agy models` + # (gemini-3.8-flash-high, never the unlisted bare gemini-3.8-flash). + # --effort takes low|medium|high. --dangerously-skip-permissions + # auto-approves every tool call, which an unattended crewmate needs. + # Every task worktree is a fresh path, so agy would show a folder-trust + # dialog ("Do you trust the contents of this project?") and no launch flag + # suppresses it (agy 1.2.0 --help lists none). Left unanswered, the turn + # runs in agy's own scratch directory instead of the worktree, so the + # worktree is pre-registered in the captain's own + # ~/.gemini/antigravity-cli/settings.json trustedWorkspaces before launch + # (bin/fm-agy-trust.sh, the claude shape), and the post-launch gate + # (agy_wait_for_working) answers the preselected safe default ("Yes, I + # trust this folder") with a single Enter if the dialog renders anyway, + # then requires the busy signature before the spawn reports success. + # The foreign primary markers are cleared for the same + # reason cursor clears them: agy publishes no marker of its own and does not + # clear an inherited CLAUDECODE (verified in the /proc environ of a live 1.2.0 + # TUI), so bin/fm-harness.sh must not read an agy worker as its launcher. + # agy exposes no hook surface, so busy state is a rendered-tail fallback + # (bin/fm-busy-lib.sh) and nothing is armed below. + agy) printf '%s' 'env -u CLAUDECODE -u PI_CODING_AGENT -u GROK_AGENT -u FM_PI_HARNESS __AGYBIN__ --prompt-interactive "$(__OPINPUT__ encode launch-brief < __BRIEF__)" __MODELFLAG____EFFORTFLAG__--dangerously-skip-permissions' ;; # grok (Grok Build TUI): a positional prompt starts the supervised interactive # session. --always-approve auto-approves every tool execution (verified: the # crewmate runs fully autonomously, no permission gate), which an unattended @@ -1691,7 +1757,7 @@ case "$ARG3" in ;; esac -# muse and gemini are verified as CREWMATE/SCOUT adapters only. A secondmate is +# muse, gemini, and agy are verified as CREWMATE/SCOUT adapters only. A secondmate is # a firstmate instance, so it needs a primary supervision protocol. # gemini has none: docs/supervision-protocols/ carries no gemini wake protocol # and this task verified only crewmate-side launch, busy state, interrupt, and @@ -1701,7 +1767,9 @@ esac # asyncRewake handlers that firstmate's primary turn-end supervision is built on # (muse 0.1.0-R708.1). Refusing here keeps that gap loud instead of standing up a # secondmate whose supervision cycle could never be armed. -if [ "$KIND" = secondmate ] && { [ "$HARNESS" = muse ] || [ "$HARNESS" = gemini ]; }; then +# agy has none either: it exposes no hook surface for primary supervision and +# docs/supervision-protocols/ carries no agy wake protocol (agy 1.2.0). +if [ "$KIND" = secondmate ] && { [ "$HARNESS" = muse ] || [ "$HARNESS" = gemini ] || [ "$HARNESS" = agy ]; }; then echo "error: $HARNESS is a verified crewmate/scout adapter only and cannot run a secondmate; it has no primary supervision protocol. Select a harness verified for secondmates." >&2 exit 1 fi @@ -1755,6 +1823,12 @@ case "$HARNESS" in exit 1 } ;; + agy) + AGY_BIN=$(resolve_pi_executable agy) || { + echo "error: agy executable not found on PATH; install Antigravity CLI or select a different verified harness" >&2 + exit 1 + } + ;; esac # config/secondmate-harness may carry optional model/effort tokens alongside the @@ -1790,6 +1864,9 @@ fi if [ "$HARNESS" = omp ]; then omp_model_validate "$OMP_BIN" "$MODEL" || exit 1 fi +if [ "$HARNESS" = agy ]; then + agy_model_validate "$AGY_BIN" "$MODEL" || exit 1 +fi secondmate_registry_value() { secondmate_registry_field "$DATA/secondmates.md" "$1" "$2" @@ -1902,7 +1979,7 @@ model_flag_for_harness() { local harness=$1 model=$2 [ -n "$model" ] && [ "$model" != default ] || return 0 case "$harness" in - claude|codex|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp) + claude|codex|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp|agy) printf -- '--model %s ' "$(shell_quote "$model")" ;; esac @@ -1934,6 +2011,13 @@ effort_flag_for_harness() { low|medium|high) printf -- '--reasoning-effort %s ' "$(shell_quote "$effort")" ;; esac ;; + agy) + # agy 1.2.0 --effort accepts exactly low|medium|high, so xhigh and max are + # omitted rather than passed as known-bad values (record-and-omit). + case "$effort" in + low|medium|high) printf -- '--effort %s ' "$(shell_quote "$effort")" ;; + esac + ;; pi|pi-signed) # Pi 0.80.6 accepts the full shared effort vocabulary, including max, through # its --thinking flag. @@ -3085,19 +3169,79 @@ rovo_spawn_fail() { # <detail> rovo_endpoint_cleanup } -# No task record is ever published on this failure path, so nothing else -# (teardown, the watcher) will ever learn this endpoint exists to close it: -# without this, the already-launched --yolo rovo process keeps running as an -# orphaned autonomous agent outside task control. Mirrors fm-teardown.sh's own -# generic non-orca kill call; orca's worktree+terminal are owned by the -# separate ORCA_ABORT_CLEANUP trap path and are out of scope here. +# The launch-then-confirm gates run after the task record is published, when +# ORCA_ABORT_CLEANUP is already cleared and neither the abort trap nor a +# teardown owns this endpoint yet, so a gate failure must close the launched +# process here or it keeps running as an orphaned autonomous agent outside +# task control. Mirrors fm-teardown.sh's own generic kill call. On orca only +# the exact terminal is closed: that stops the CLI while its worktree stays +# for the record's own teardown, which owns worktree deletion. rovo_endpoint_cleanup() { - [ "$BACKEND" = orca ] && return 0 + if [ "$BACKEND" = orca ]; then + fm_backend_kill orca "$T" 2>/dev/null || true + return 0 + fi local tab_id= [ "$BACKEND" = zellij ] && tab_id=$ZELLIJ_TAB_ID fm_backend_kill "$BACKEND" "$T" "$tab_id" "fm-$ID" 2>/dev/null || true } +# agy carries its brief on the launch command, so it needs no delivery gate, +# but a worktree agy does not trust parks the TUI on the folder-trust dialog +# and an unanswered dialog sends the turn into agy's scratch directory instead +# of the worktree. The trust is pre-registered before launch +# (bin/fm-agy-trust.sh, verified to remove the dialog), and this gate is the +# backstop in the rovo/kimi launch-then-confirm shape: answer the dialog once +# with the preselected safe default if it renders anyway, then require +# positive proof that the brief is being processed - the same verdict the +# supervisor reads (Herdr's native working state or the pinned `esc to cancel` +# status row through fm_busy_classify) - before the spawn reports success. +# The gate is strict about ordering because on Herdr the native working +# verdict is known to coexist with an unanswered dialog: a busy verdict counts +# only when the path was pre-registered or the dialog has been seen and +# answered; on an unregistered path it keeps polling for the dialog instead. +AGY_TRUST_DIALOG='Do you trust the contents of this project?' +AGY_TRUST_ANSWERED=0 + +agy_capture() { + fm_backend_capture "$BACKEND" "$T" 120 "$W" 2>/dev/null || true +} + +agy_pane_shows_trust_dialog() { # <plain-pane-capture> + printf '%s\n' "$1" | grep -Fq "$AGY_TRUST_DIALOG" +} + +agy_pane_is_working() { # <plain-pane-capture> + case "$(fm_busy_classify "$BACKEND" "$T" agy "$ID" "$STATE" "$1")" in + busy*) return 0 ;; + esac + return 1 +} + +agy_wait_for_working() { + local pane i=0 max=${FM_AGY_READY_POLLS:-60} interval=${FM_AGY_POLL_INTERVAL:-0.5} + while [ "$i" -lt "$max" ]; do + pane=$(agy_capture) + if agy_pane_shows_trust_dialog "$pane"; then + if [ "$AGY_TRUST_ANSWERED" -eq 0 ]; then + spawn_send_key "$T" Enter + AGY_TRUST_ANSWERED=1 + fi + elif [ "$AGY_TRUST_PREREGISTERED" -eq 1 ] || [ "$AGY_TRUST_ANSWERED" -eq 1 ]; then + agy_pane_is_working "$pane" && return 0 + fi + i=$((i + 1)) + [ "$i" -ge "$max" ] || sleep "$interval" + done + return 1 +} + +agy_spawn_fail() { # <detail> + printf 'failed: %s\n' "$1" >> "$STATE/$ID.status" + echo "error: $1; inspect window $T" >&2 + rovo_endpoint_cleanup +} + if [ "$RELAUNCH" -eq 1 ]; then # No worktree is acquired: the recorded one is reused as-is. What must be # proven instead is that the adopted endpoint's shell is actually sitting in @@ -3233,6 +3377,15 @@ fi # temp root, no retired relaunch wiring and no busy record exists yet to strand, # so the refusal names the endpoint the same way they do and leaves nothing else # behind. +# agy gates a fresh worktree behind its own folder-trust dialog and honours a +# trustedWorkspaces entry written ahead of launch (bin/fm-agy-trust.sh), so the +# same pre-registration removes the dialog for it. Unlike claude's dialog, agy's +# preselects the safe answer, so a failed registration is not fatal here: the +# post-launch gate (agy_wait_for_working) answers the dialog itself and, on a +# path that was not pre-registered, refuses to count a busy turn as ready until +# it has done so. agy is crewmate/scout only (refused above for secondmate), so +# only the worktree shape applies. +AGY_TRUST_PREREGISTERED=0 case "$HARNESS" in claude*) if [ "$KIND" = secondmate ]; then @@ -3245,6 +3398,15 @@ case "$HARNESS" in exit 1 fi ;; + agy) + if [ "$KIND" != secondmate ]; then + if "$FM_ROOT/bin/fm-agy-trust.sh" "$WT" "$PROJ_ABS" >/dev/null; then + AGY_TRUST_PREREGISTERED=1 + else + echo "warning: could not pre-register agy workspace trust for $WT; the launch will answer the folder-trust dialog in window $T instead" >&2 + fi + fi + ;; esac # Per-task temp root: /tmp/fm-<id>/ with Go's build temp nested at gotmp/. Go won't @@ -3891,10 +4053,11 @@ case "$HARNESS" in cursor) LAUNCH=${LAUNCH//__CURSORBIN__/"$(shell_quote "$CURSOR_BIN")"} ;; gemini) LAUNCH=${LAUNCH//__GEMINISETTINGS__/"$(shell_quote "$STATE_REAL/$ID.gemini-settings.json")"} ;; omp) LAUNCH=${LAUNCH//__OMPBIN__/"$(shell_quote "$OMP_BIN")"} ;; + agy) LAUNCH=${LAUNCH//__AGYBIN__/"$(shell_quote "$AGY_BIN")"} ;; esac LAUNCH=${LAUNCH//__WORKTREE__/$sq_worktree} case "$HARNESS" in - claude|codex|opencode|pi|pi-signed|grok|kimi|gemini|muse|rovo) + claude|codex|opencode|pi|pi-signed|grok|kimi|gemini|muse|rovo|agy) LAUNCH="env -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u GEMINI_CLI $LAUNCH" ;; esac @@ -4066,6 +4229,18 @@ if [ "$HARNESS" = rovo ]; then exit 1 fi fi +if [ "$HARNESS" = agy ]; then + if ! agy_wait_for_working; then + if [ "$AGY_TRUST_ANSWERED" -eq 1 ]; then + agy_spawn_fail "agy did not start processing its brief after the folder-trust dialog was answered in window $T" + elif [ "$AGY_TRUST_PREREGISTERED" -eq 1 ]; then + agy_spawn_fail "agy did not start processing its brief in the pre-trusted worktree in window $T" + else + agy_spawn_fail "agy never showed its folder-trust dialog on an unregistered worktree in window $T, so the brief could not be confirmed to run there" + fi + exit 1 + fi +fi if [ "$KIND" = secondmate ] && [ "${FM_SKIP_SECONDMATE_INHERIT:-0}" != 1 ]; then if ! fm_config_reread_discard_pending "$PROJ_ABS" "$ID" "$FM_HOME"; then if fm_config_reread_quarantine_pending "$PROJ_ABS" "$ID" "$FM_HOME"; then diff --git a/bin/fm-test-run.sh b/bin/fm-test-run.sh index 43f3bff4d06..e08ea75b444 100755 --- a/bin/fm-test-run.sh +++ b/bin/fm-test-run.sh @@ -278,7 +278,7 @@ family_for_basename() { fm-composer-ghost.test.sh|fm-composer-lib.test.sh|\ fm-crew-state.test.sh|fm-captain-hold-lifecycle.test.sh|\ fm-documentation-audiences.test.sh|fm-ensure-agents-md.test.sh|fm-grok-harness.test.sh|\ - fm-kimi-harness.test.sh|fm-muse-harness.test.sh|fm-rovo-harness.test.sh|fm-omp-harness.test.sh|fm-herdr-lab.test.sh|fm-lint.test.sh|\ + fm-kimi-harness.test.sh|fm-muse-harness.test.sh|fm-rovo-harness.test.sh|fm-agy-harness.test.sh|fm-omp-harness.test.sh|fm-herdr-lab.test.sh|fm-lint.test.sh|\ fm-lint-workflows.test.sh|\ fm-operational-input.test.sh|fm-pi-primary-types.test.sh|\ fm-harness-adapter-references.test.sh|\ @@ -341,7 +341,7 @@ family_for_basename() { fm-cursor-primary-live-e2e.test.sh|\ fm-grok-stop-live-e2e.test.sh|fm-harness-adapter-instructions-live-e2e.test.sh|\ fm-harness-liveness-drift-live-e2e.test.sh|\ - fm-muse-signals-live-e2e.test.sh|fm-rovo-signals-live-e2e.test.sh|\ + fm-muse-signals-live-e2e.test.sh|fm-rovo-signals-live-e2e.test.sh|fm-agy-signals-live-e2e.test.sh|\ fm-herdr-version-floor-live-e2e.test.sh|\ fm-herdr-pi-stale-registration-live-e2e.test.sh|\ fm-opencode-primary-live-e2e.test.sh|fm-pi-branch-live-e2e.test.sh|\ @@ -650,6 +650,8 @@ list_portable_serial() { # balance rather than coverage. That doc owns the refresh procedure. portable_serial_weight_hints() { cat <<'EOF' +tests/fm-agy-harness.test.sh 11000 +tests/fm-agy-signals-live-e2e.test.sh 23 tests/fm-afk-contract.test.sh 3000 tests/fm-afk-inject-e2e.test.sh 35792 tests/fm-afk-pi-herdr-return-e2e.test.sh 100 diff --git a/docs/agent-control.md b/docs/agent-control.md index 7d7fffa4722..ef92c1b7de8 100644 --- a/docs/agent-control.md +++ b/docs/agent-control.md @@ -48,7 +48,7 @@ The clear is refused before anything is sent when the recorded backend cannot de Removing a worktree, closing an endpoint, or discarding work stays with [`bin/fm-teardown.sh`](../bin/fm-teardown.sh), which owns the landed-work test. **`resume` is not a verb.** -It is not deterministic across the verified adapters: codex, grok, and gemini resume only from a session id printed at exit, opencode continues the most recent session for the cwd, and claude, pi, pi-signed, omp, and kimi have no verified pane-resume contract. +It is not deterministic across the verified adapters: codex, grok, and gemini resume only from a session id printed at exit, opencode continues the most recent session for the cwd, and claude, pi, pi-signed, omp, kimi, and agy have no verified pane-resume contract. `relaunch` covers the same need on every adapter, because the brief on disk - not a harness-private session - is the durable instruction. ## Transactional relaunch @@ -114,11 +114,11 @@ Backend capability comes from each adapter's real surface, not from a policy cho | cmux | yes | yes | yes | yes | no | | orca | no | yes | yes | no | no | -Per-harness interrupt keys, repeat counts, composer clears, exit commands, and supported task kinds live in `bin/fm-control-lib.sh` and are exercised for every verified harness by `tests/fm-control.test.sh`. +Per-harness interrupt keys, repeat counts, composer clears, exit commands, and supported task kinds live in `bin/fm-control-lib.sh` and are exercised for every verified harness by `tests/fm-control.test.sh`, with adapters outside its lane pinning their control mechanics in their own harness suites. The empirical basis for each adapter's value is the `harness-adapters` skill's verification record for that adapter. ## Verification -- `tests/fm-control.test.sh` - the adapter contract for every verified harness, the backend capability matrix, exact-id scoping, the closed verb list, the busy, idle, dead, and idempotent lifecycle cases, and marker non-regression, all against a stubbed session provider. +- `tests/fm-control.test.sh` - the adapter contract for its verified-harness lane (adapters outside the lane pin their control mechanics in their own harness suites), the backend capability matrix, exact-id scoping, the closed verb list, the busy, idle, dead, and idempotent lifecycle cases, and marker non-regression, all against a stubbed session provider. - `tests/fm-control-relaunch.test.sh` - the relaunch transaction: identity preservation, harness switching, the progress note, checkpoint refusals, and rollback after a failed launch. - `tests/fm-control-herdr-smoke.test.sh` - the second state-verified backend against the real herdr binary, on an isolated throwaway lab session. diff --git a/docs/architecture.md b/docs/architecture.md index 1ccf3aa7c45..443dcaa99d4 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -178,7 +178,7 @@ Every classification returns a verdict of busy, idle, unknown, or dead together Each converted adapter reports its own turn lifecycle through a machine-readable contract the vendor already exposes, rather than through rendered footer text: Pi and pi-signed through the Firstmate-owned extension's `agent_start` and `agent_settled` confirmed by `ctx.isIdle()`, omp through its extension's `agent_start` and `agent_end` without `willContinue`, OpenCode through its plugin's semantic `session.status`, Claude through owned `UserPromptSubmit`, `Stop`, `StopFailure`, and `SessionEnd` hooks, Muse through its session log, and Cursor through its conversation transcript. Kimi behind Pi inherits Pi's lifecycle. -Codex and standalone Kimi classify unknown behind explicit probes until a semantic source is live-verified for them, and Grok keeps one clearly isolated rendered-tail fallback that can only ever classify a Grok task. +Codex and standalone Kimi classify unknown behind explicit probes until a semantic source is live-verified for them, and Grok, Rovo, and AGY each keep one clearly isolated rendered-tail fallback that can only ever classify their own task. Missing, malformed, stale, untrusted, or unverified semantic state is unknown, never idle, and unknown is never promoted to busy either. Ordinary task-state consumers act only on an exact busy verdict, so an unreadable worker surfaces for a closer look instead of being absorbed as still-working or written off as finished. @@ -254,7 +254,7 @@ The session-start bootstrap step keeps valid dispatch configuration silent unles When the file exists, `fm-spawn.sh` refuses crewmate and scout launches without an explicit harness, so `config/crew-harness` is only automatic when no dispatch profile file is active. Secondmate launches are exempt because they resolve the secondmate harness and any optional secondmate model or effort tokens instead. Unsupported effort values are still recorded in task meta when passed to `fm-spawn.sh`, but the launch template omits any effort flag that the selected harness does not accept. -That keeps spawn launch compatible across claude, codex, opencode, pi, pi-signed, grok, kimi, cursor, gemini, muse, rovo, and omp while preserving the requested profile for later audit. +That keeps spawn launch compatible across claude, codex, opencode, pi, pi-signed, grok, kimi, cursor, gemini, muse, rovo, omp, and agy while preserving the requested profile for later audit. ## Optional secondmates diff --git a/docs/configuration.md b/docs/configuration.md index b6ad202c6fa..d10624d2a77 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -314,6 +314,7 @@ muse is verified for crewmate and scout launches ONLY, and `fm-spawn.sh` refuses muse also needs a worker-reachable credential before spawning, and the portable fleet path is the `<config>/muse/auth.json` credential stored by `muse login`, because a caller-only `META_API_KEY` does not cross a long-lived backend daemon. gemini is likewise refused for secondmates because it has no primary supervision protocol; [its adapter reference](../.agents/skills/harness-adapters/references/harness/gemini.md) owns the credential precondition, canonical-launch wiring, and raw-launch limitations. rovo is likewise verified for crewmate and scout launches ONLY, refused for a secondmate for the same reason - no turn-end hook and no primary supervision protocol; [`docs/verification/rovo.md`](verification/rovo.md) owns that evidence, including the OAuth token's silent background refresh from a stored refresh token and both tmux and herdr pane liveness (herdr placement is verified live, with a Herdr-side agent-detection gap left open for recovery classification). +agy is likewise verified for crewmate and scout launches ONLY, refused for a secondmate for the same reason - no hook surface and no primary supervision protocol; [`docs/verification/agy.md`](verification/agy.md) owns that evidence, including the spawn-time worktree trust pre-registration through `bin/fm-agy-trust.sh` and Herdr's native agy pane recognition. New harnesses get verified through a supervised trial task before joining the set. The verified adapter evidence - each harness's busy-state source, interrupt and exit behavior, skill-invocation syntax, and per-harness quirks - lives in the skill tree rooted at [`.agents/skills/harness-adapters/SKILL.md`](../.agents/skills/harness-adapters/SKILL.md). The executable interrupt and exit mechanics live in [`bin/fm-control-lib.sh`](../bin/fm-control-lib.sh), and [`docs/agent-control.md`](agent-control.md) owns their lifecycle-control architecture. @@ -1083,7 +1084,7 @@ FM_COMPOSER_CAPTURE_LINES=20 # fleet-wide bound for tail-capture composer read FM_COMPOSER_PI_MAX_LINES=8 # fleet-wide: maximum rows admitted between Pi's identity-corroborated separator pair; taller or ambiguous candidates stay unknown FM_COMPOSER_GHOST_LUMA_MAX=128 # fleet-wide: max perceived luminance (0.299R+0.587G+0.114B, 0-255) for a TRUECOLOR foreground to count as de-emphasised ghost/placeholder text and be stripped; dim/faint (SGR 2) is stripped regardless. Assumes a dark terminal theme (bin/fm-composer-lib.sh's fm_composer_strip_ghost, used by styled tmux, herdr, and Zellij reads) GROK_HOME= # optional Grok config home for firstmate's global grok turn-end hook; defaults to ~/.grok -FM_SEND_RETRIES=3 # fm-send typed-plane Enter-retry attempts after typing the line once +FM_SEND_RETRIES=3 # fm-send typed-plane Enter-retry attempts after typing the line once; agy typed targets use a longer per-harness default owned by bin/fm-send.sh FM_SEND_SLEEP=0.4 # seconds between fm-send typed-plane submit checks FM_SEND_SETTLE=1 # seconds fm-send waits after a successful typed-plane submit; 0 disables FM_PENDING_REPLY_GRACE_SECS=120 # seconds after marked-request delivery before a completed turn without a correlated parent report is eligible for its one recovery repost diff --git a/docs/documentation-audiences.json b/docs/documentation-audiences.json index 03124cd1628..1e66c6dbc9c 100644 --- a/docs/documentation-audiences.json +++ b/docs/documentation-audiences.json @@ -176,6 +176,10 @@ "path": ".agents/skills/harness-adapters/references/common/primary-hooks.md", "audience": "agent-runtime" }, + { + "path": ".agents/skills/harness-adapters/references/harness/agy.md", + "audience": "agent-runtime" + }, { "path": ".agents/skills/harness-adapters/references/harness/claude.md", "audience": "agent-runtime" @@ -432,6 +436,10 @@ "path": "docs/turnend-guard.md", "audience": "operator-current" }, + { + "path": "docs/verification/agy.md", + "audience": "maintainer-verification" + }, { "path": "docs/verification/dispatch-auth.md", "audience": "maintainer-verification" diff --git a/docs/fm-test-portable-shards.md b/docs/fm-test-portable-shards.md index 45c3a5fd763..a327902b99c 100644 --- a/docs/fm-test-portable-shards.md +++ b/docs/fm-test-portable-shards.md @@ -57,31 +57,21 @@ Each shard is still strictly serial in itself, and separate runners mean no two `.github/workflows/ci.yml` derives the same `n` from `strategy.job-total` rather than a literal, so changing the shard count in either file without the other fails the lane loudly instead of leaving part of the required suite unrun. Assignment is longest-processing-time bin packing over per-script duration hints embedded in `bin/fm-test-run.sh`. -The 145 current hints include the slowest measurements retained from the `fm-test-timing-portable-serial-*` artifacts of three green CI runs on 2026-09-01, [33558082172](https://github.com/kunchenguid/firstmate/actions/runs/33558082172), [33523597838](https://github.com/kunchenguid/firstmate/actions/runs/33523597838), and [33463326167](https://github.com/kunchenguid/firstmate/actions/runs/33463326167), the completed-script measurements from [run 34342484144](https://github.com/kunchenguid/firstmate/actions/runs/34342484144), plus the 5121 ms native-Windows focused runner measurement for `tests/fm-pi-windows-shell-invocation.test.sh` from 2026-09-06T21:02Z. -Those per-script maxima total 4312606 ms of conservative balance weight. +The embedded hints include the slowest measurements retained from the `fm-test-timing-portable-serial-*` artifacts of three green CI runs on 2026-09-01, [33558082172](https://github.com/kunchenguid/firstmate/actions/runs/33558082172), [33523597838](https://github.com/kunchenguid/firstmate/actions/runs/33523597838), and [33463326167](https://github.com/kunchenguid/firstmate/actions/runs/33463326167), the completed-script measurements from [run 34342484144](https://github.com/kunchenguid/firstmate/actions/runs/34342484144), plus the 5121 ms native-Windows focused runner measurement for `tests/fm-pi-windows-shell-invocation.test.sh` from 2026-09-06T21:02Z. Taking the slowest of several CI runs rather than a single run keeps the balance honest on a slow runner. -A script with no hint gets the conservative `PORTABLE_SERIAL_DEFAULT_WEIGHT_MS` default; the current 154-script lane has nine such scripts, bringing its assignment weight to 4555606 ms. +A script with no hint gets the conservative `PORTABLE_SERIAL_DEFAULT_WEIGHT_MS` default. Hints only affect balance: the coverage guard keeps the partition complete and disjoint whatever they say, so a stale hint costs a slower shard rather than lost coverage. Balance is still worth keeping current, because enough unmeasured scripts let one shard carry more than twice another shard's real work and reach the job cap while another runner sits idle. That is not hypothetical: by 2026-09-01 the lane had grown from 116 to 139 scripts and from ~42 to ~63 minutes, 17 scripts were still unmeasured, and several hints were low by 2-5x, so shard 3 of 4 ran 17-20 minutes against its 20-minute cap while shard 1 ran 11.5 minutes and run [33574154856](https://github.com/kunchenguid/firstmate/actions/runs/33574154856) timed out seconds after a passing test. `bin/fm-test-run.sh --check-coverage` now reports the unmeasured share as `serial_unhinted=` and refuses past `PORTABLE_SERIAL_MAX_UNHINTED_PERCENT`, so hint drift fails the coverage guard instead of silently pushing one shard into its job cap. Refresh the hints whenever the serial lane gains scripts, rather than waiting for that bound to trip. -| Lane | Script count | Estimated duration | -|---|---:|---:| -| `portable-serial-1of5` | 30 | 911111 ms (~15.19 min) | -| `portable-serial-2of5` | 31 | 911128 ms (~15.19 min) | -| `portable-serial-3of5` | 32 | 911128 ms (~15.19 min) | -| `portable-serial-4of5` | 31 | 911128 ms (~15.19 min) | -| `portable-serial-5of5` | 30 | 911111 ms (~15.19 min) | -| imbalance | | 17 ms | - -The current table is generated from the runner's retained maxima plus its default for the nine unhinted scripts. +`bin/fm-test-run.sh` owns the per-shard packing, so its `--check-coverage` output is the current account of lane size, shard composition, and balance rather than a copied table. Run 34342484144 observed a shard reach about 20 minutes of passing work, so the 30-minute job cap keeps meaningful hang-tripwire margin for job setup and runner-speed spread. The single longest script, `tests/fm-watch-triage.test.sh` at 262626 ms, is the floor for any shard count. -Refresh the CI-derived hints by downloading the per-shard timing artifacts from several green CI runs, replacing the `portable_serial_weight_hints` table in `bin/fm-test-run.sh` with the slowest measured `duration_ms` per `path`, and updating the table above: +Refresh the CI-derived hints by downloading the per-shard timing artifacts from several green CI runs and replacing the `portable_serial_weight_hints` table in `bin/fm-test-run.sh` with the slowest measured `duration_ms` per `path`: ```sh for run in <run-id> <run-id> <run-id>; do diff --git a/docs/tmux-backend.md b/docs/tmux-backend.md index a5b4fd9c8ff..77b836c04fa 100644 --- a/docs/tmux-backend.md +++ b/docs/tmux-backend.md @@ -48,7 +48,7 @@ Verify setup by spawning a small task and confirming its `fm-<id>` window appear A target-existence check proves only that the pane exists. The deeper tmux agent-liveness probe first verifies exact window membership, then reads process names to distinguish a running harness from a bare idle shell. -It classifies recognized Claude, Codex, OpenCode, Pi, pi-signed, Grok, Kimi, Cursor, Muse, and Rovo process identities as `alive`, common shells as `dead`, an authoritatively absent window as `missing`, unreadable state as `unreadable`, and every other process as `ambiguous`. +It classifies recognized Claude, Codex, OpenCode, Pi, pi-signed, Grok, Kimi, Cursor, Muse, Rovo, and AGY process identities as `alive`, common shells as `dead`, an authoritatively absent window as `missing`, unreadable state as `unreadable`, and every other process as `ambiguous`. The process-name vocabulary behind those verdicts is owned by `bin/fm-agent-process-lib.sh` and shared with the Herdr adapter, which proves a registered agent against the same names ([herdr-backend.md](herdr-backend.md) "Restart and liveness behavior"). Only `dead` and `missing` authorize recovery because a false dead result could launch a duplicate agent. @@ -62,6 +62,7 @@ The same scoping covers multi-process launchers without a special case, so the P Direct executable identities `pi`, `pi-signed`, and `Pi` remain accepted exactly, and similar or prefixed process names are not accepted through those exact Pi-family entries. Muse is likewise anchored to the exact `muse` launcher identity or the installed `muse-bin-<version>` prefix, so unrelated names such as `musescore` and `amuse` remain ambiguous. omp is anchored to the exact `omp` identity for the same reason, so `ompd` and `comp` remain ambiguous. +AGY is anchored to the exact `agy` identity for the same reason, so unrelated names containing that fragment remain ambiguous. Cursor is identified from its exact `cursor-agent` identity or versioned install tree in the foreground process path or structured argv[0]; a bare `node` or unrelated `agent` remains ambiguous. The CI-enforced portable regression and opt-in real-harness drift guard follow the split owned by `.agents/skills/firstmate-coding-guidelines/SKILL.md`. diff --git a/docs/trace-context.md b/docs/trace-context.md index 5f273345a7e..6a9cb5e83b9 100644 --- a/docs/trace-context.md +++ b/docs/trace-context.md @@ -23,7 +23,7 @@ When enabled, for each spawn Firstmate resolves one W3C `traceparent` carrier fo This feature parents no SDK span by itself. Because the injected carrier and the recorded carrier are the same string, an observer that reads the metadata reconstructs exactly the identity the child received. -The injection sits at the unconditional pre-launch export site, so it covers ship and scout spawns across `claude`, `codex`, `opencode`, `pi`, `pi-signed`, `grok`, `kimi`, `cursor`, `gemini`, `muse`, and `rovo`, plus Secondmate spawns across that same set except the deliberately crewmate-only `gemini`, `muse`, and `rovo` adapters. +The injection sits at the unconditional pre-launch export site, so it covers ship and scout spawns across `claude`, `codex`, `opencode`, `pi`, `pi-signed`, `grok`, `kimi`, `cursor`, `gemini`, `muse`, `rovo`, and `agy`, plus Secondmate spawns across that same set except the deliberately crewmate-only `gemini`, `muse`, `rovo`, and `agy` adapters. This is the same coverage `GOTMPDIR` already has and requires no trace-specific `launch_template()` behavior. Ship and scout spawns reach that site on every spawn backend (`tmux`, `herdr`, `zellij`, `orca`, `cmux`); a Secondmate reaches it on every backend that accepts a Secondmate spawn (`tmux`, `herdr`, `zellij`), because `bin/fm-spawn.sh` rejects a Secondmate on `orca` and `cmux`. diff --git a/docs/verification/agy.md b/docs/verification/agy.md new file mode 100644 index 00000000000..c8bec59a9a7 --- /dev/null +++ b/docs/verification/agy.md @@ -0,0 +1,170 @@ +# Verification: the agy (Antigravity CLI) crewmate/scout adapter + +Active empirical facts for firstmate's agy adapter. +The skill tree rooted at [`.agents/skills/harness-adapters/SKILL.md`](../../.agents/skills/harness-adapters/SKILL.md) owns the operating facts through [`references/harness/agy.md`](../../.agents/skills/harness-adapters/references/harness/agy.md); this record owns how they were established and what is still unproven. + +## Subject + +| Field | Value | +|---|---| +| Version | `agy 1.2.0`; the send-confirmation timing below was re-measured on `agy 1.2.1` (2026-09-12) | +| Verified | 2026-09-10 | +| Binary | `/home/andpod/.local/bin/agy`, an ELF 64-bit Go-compiled single executable | +| Platform | Linux x64 (Arch, kernel 7.2.3) | +| Backend | Herdr, in an isolated non-`default` lab session (`fm-lab-firstmate-agy-ad-*` via `bin/fm-herdr-lab.sh`); the live `default` session was unchanged throughout | + +Every command below ran inside the disposable firstmate task worktree or the named Herdr lab session. +No captain fleet state was touched. + +## Detection: ancestry only, no marker + +``` +$ agy --version +1.2.0 +``` + +A live TUI's `/proc/<pid>/environ` carries no `AGY_*` or `ANTIGRAVITY_*` variable. +It does carry `AGENT=1` and `CLAUDECODE=1`, both inherited from the launching environment, so neither is an agy identity and neither is promoted to a marker. +Herdr's `pane process-info` for the same pane reports the foreground process as `name=agy` with `argv=["agy", ...]`, and `ps -o comm=` reports `agy`. +`bin/fm-harness.sh` therefore matches the anchored process name `agy` alone, and the spawn clears `CLAUDECODE`, `PI_CODING_AGENT`, `GROK_AGENT`, and `FM_PI_HARNESS` at the launch boundary. +`tests/fm-agy-harness.test.sh` pins the anchored match, the rejection of unrelated names containing the fragment, and that an inherited `CLAUDECODE` never outranks a real `agy` ancestor once the spawn clears it. + +## Launch: positional prompt-interactive with auto-submit + +``` +$ agy --prompt-interactive "Reply with exactly AGY_LIVE_PROBE_OK and nothing else" --model gemini-3.8-flash-low --effort low --dangerously-skip-permissions +``` + +The brief submitted itself with no extra Enter, the turn ran, and the reply rendered in the pane. +A second launch into the same directory answered a fresh prompt the same way, so the shape is repeatable, not a first-run accident. +The footer rendered `Gemini 3.8 Flash · low`, proving both flags were accepted together. + +## Trust dialog: pre-registered before launch, gated on a busy turn as the backstop + +A first launch in a fresh worktree shows this dialog: + +``` +Accessing workspace: + +/home/andpod/.treehouse/firstmate-7bab20/1/firstmate/agy-probe-tmp + +Do you trust the contents of this project? + +Antigravity CLI requires permission to read, edit, and execute files here. + +> Yes, I trust this folder + No, exit +``` +`agy --help` (1.2.0) lists no trust flag or pre-registration command, but agy honours a `trustedWorkspaces` entry written to `~/.gemini/antigravity-cli/settings.json` ahead of launch. +Verified under a throwaway `HOME` holding a copy of `~/.gemini` (the real settings file was never written): a folder appended to that array by hand launched `--prompt-interactive` straight into its turn and rendered the reply with no dialog, while an unregistered sibling folder launched the same way parked on the dialog. +agy compares the pane's logical working directory, not its resolved path: a symlinked cwd whose real path alone was registered still parked on the dialog, so `bin/fm-agy-trust.sh` records both the logical path and its resolved form when they differ. +`bin/fm-spawn.sh` runs that helper before launch at the same point it pre-registers claude trust; the helper applies the same structural scope test (a linked worktree of the spawning project, never a primary checkout, a subdirectory, a plain directory, or the home directory), preserves every other key in the store, and writes atomically with a fingerprint check. +A failed registration is a stderr warning rather than a refusal, because agy's dialog preselects the safe answer and the gate below can answer it. +Two supervised Herdr runs in treehouse worktrees completed file-writing turns while the dialog was still unanswered at observation time (worker file and `done:` status line both verified on disk before Enter was ever sent to those panes). +Isolated runs in untrusted `/tmp` directories never reached the workspace until Enter: the turn spun through exploratory tool calls in agy's own scratch directory instead, and only the queued prompt ran after the answer. +One run left unanswered for several minutes wrote its file to agy's scratch directory instead of the workspace once finally answered. +The mechanism behind the difference was not established; path, backend, and latency were all varied across runs without isolating a single cause. +The spawn therefore does not depend on it: after pre-registration, `bin/fm-spawn.sh` runs a post-launch readiness gate (`agy_wait_for_working`) in the rovo/kimi launch-then-confirm shape as the backstop. +It polls the pane capture, answers the dialog with a single Enter the first time the `Do you trust the contents of this project?` text renders, and reports success only once `fm_busy_classify` returns a busy verdict for the pane (Herdr's native `working` status or the pinned `esc to cancel` status row). +Because Herdr's native `working` verdict is known to coexist with an unanswered dialog, the gate is strict about order: a busy verdict counts as ready only when the worktree was pre-registered before launch or the dialog has already been seen and answered; on an unregistered path it keeps polling for the dialog instead of accepting the early busy verdict. +When the brief cannot be confirmed to run within the window (an answered dialog never turns busy, a pre-trusted pane never turns busy, or an unregistered pane never shows the dialog), the spawn fails, records `failed:` in the task status, and closes the endpoint so no orphan worker survives outside task control. +`tests/fm-agy-harness.test.sh` covers the helper's registration and scope refusals against a throwaway store, and drives a fake pane whose dialog decision reads the store the spawn just wrote: the pre-trusted launch with no dialog, a dialog that renders anyway answered exactly once, the premature busy verdict on an unregistered path waiting for the dialog, and both fail-and-close paths. + +## Model and effort + +``` +$ agy models +Fetching available models... +gemini-3.8-flash-high Gemini 3.8 Flash (High) +gemini-3.8-flash-medium Gemini 3.8 Flash (Medium) +gemini-3.8-flash-low Gemini 3.8 Flash (Low) +... +``` + +`agy --help` documents `--effort` as `low|medium|high` and `--model` as the model for the session. +The bare `gemini-3.8-flash` id from this home's previous config is not listed; only the suffixed `-high`, `-medium`, and `-low` variants are. +`bin/fm-spawn.sh`'s `agy_model_validate` refuses a requested id a reachable `agy models` listing omits, and launches unvalidated with a stderr notice when the listing is unreachable. +The listing is a remote fetch (`Fetching available models...`), so the probe runs with stdin detached under the shared hard bound from `bin/fm-timeout-lib.sh` (15 seconds by default, `FM_AGY_MODELS_TIMEOUT`; a non-positive or non-numeric value clamps back to that default, because a non-positive bound is not a bound); a stalled fetch or a sign-in prompt is cut off and falls through to the unvalidated launch instead of blocking the spawn before any pane exists. +Print mode (`agy -p "Reply with exactly: AGY_PRINT_PROBE_OK" --model gemini-3.8-flash-low`) returned the exact reply with exit 0 in about 8 seconds, proving the credential path without a pane. + +## Busy state: the pinned status row, unknown on absence + +Mid-turn the pane rendered the status row and a spinner line at once: + +``` +⣯ Generating... +└ Tip: When reviewing a file edit, press f to see the full diff. +... +esc to cancel Gemini 3.8 Flash · low +``` + +The completed turn showed the reply, then the idle composer: + +``` +> +────────────────────────────────────────────────────────────────────────────── +? for shortcuts Gemini 3.8 Flash · low +``` + +`fm_busy_agy_tail_busy` and the delivery guard in `bin/fm-composer-lib.sh` match the `esc to cancel` token alone: the TUI pins that status row to the bottom of the pane for the whole turn, and the idle row replaces it with `? for shortcuts`. +The `Generating...` spinner word is deliberately not a signal: it is a free-floating output line, so ordinary worker output such as `Generating report...` would otherwise classify an idle worker as busy or acknowledge a submit that did not land. +No busy phase without the status row was observed live; every captured mid-turn frame carried it. +`fm_busy_classify` reports `unknown agy-regex` when the token is absent, because a long turn can scroll the marker out of the captured tail. +The signature is hardcoded with no environment override, so a stray variable can never change worker-state classification. +Herdr's own registry agreed throughout: `agent get` reported `agent_status=working` mid-turn and `idle` after, so on Herdr the native verdict carries busy with no new code. + +## Interrupt and exit + +A single `Escape` sent mid-turn through `herdr pane send-keys` cancelled it and printed this row, with the composer back at idle and no repolluted text: + +``` + ⎿ Interrupted · What should Antigravity CLI do instead? +``` + +Sending `/quit` plus Enter exited the process; the pane closed under the `exec` launch, and Herdr reported the pane gone. +`bin/fm-control-lib.sh` records `Escape` once, no clear key, no ack source, and `/quit` for agy. + +## Backend liveness: Herdr recognizes agy, tmux names it + +``` +$ herdr agent get w2:p1 --session fm-lab-firstmate-agy-ad-1599574-8823 +{"result":{"agent":{"agent":"agy","agent_status":"idle",...,"agent_session":{"agent":"agy","kind":"id","source":"herdr:antigravity_cli",...}}}} +``` + +Herdr tracks agy natively (`antigravity-cli` integration, detected as `agent=agy`), so `fm_backend_herdr_pane_agent_state` returns `live` for every registered agy status and no exit-detection hardening was needed. +The tmux adapter classifies the anchored process name `agy` as `agent` through the shared name vocabulary in `bin/fm-agent-process-lib.sh`, the muse/omp precedent for short bare-word names. +agy stays out of the session-lock name vocabulary in `bin/fm-session-lock-lib.sh`, where the other crewmate-only adapters are also absent. + +## Composer: unknown by design + +Byte-level capture of the idle pane shows a bare unstyled `>` between two full-width `─` rules, with an unstyled `? for shortcuts` cell and a dim (`SGR 2`) model cell in the status row below. +The shared classifier reads that bare `>` as `unknown` under the dead-shell rule, never `empty`. +Steering still confirms delivery: the Herdr submit core leads with the native `idle`-to-`working` transition, which agy performs, and the delivery footer regex covers the tmux path. +agy renders the busy footer late for that confirm loop - about 1.5 s after Enter for a short steer and 4-5 s for a realistic longer brief, measured live on `agy 1.2.1` (2026-09-12) against the shared budget's 3 x 0.4 s - so `bin/fm-send.sh` gives agy typed targets a longer default submit-confirm budget (20 retries, about 8 s at the default cadence); an explicit `FM_SEND_RETRIES` still wins and every other harness keeps the shared 3-retry default. +`tests/fm-send-agy-confirm.test.sh` pins the raised default and `tests/fm-agy-harness.test.sh` pins the Herdr transition path. +This is the cursor precedent, not a gap to patch in shared code. + +## Supervised task: spawn, steer, relaunch, and exit through the new path + +A trivial scout ran end to end through `bin/fm-spawn.sh --harness agy` against the same isolated lab session: `spawned agy-e2e1 harness=agy kind=scout` with a treehouse-provisioned worktree, `--model gemini-3.8-flash-low`, and `--effort low` all recorded in task metadata. +The worker wrote its worktree file and appended `done: agy e2e turn complete` to its status file, which lives outside the worktree, proving prompt processing, tool execution, outside-workspace file access, and a new completion event. +Durable steering held: a `bin/fm-send.sh` message landed in the task inbox, the worker appended the steered lines to both files, and its inbox record moved to `handled/`. +Same-copy relaunch held: `bin/fm-control.sh relaunch --note` replaced the worker in place on the identical worktree, model, and effort, the replacement verified both prior lines intact and appended `relaunched: done`. +Exit held: `bin/fm-control.sh exit` stopped the worker, the registry returned `agent_not_found`, and the pane remained a lone shell in the worktree with all work intact. +No automatic quota failover was exercised or claimed; every handoff above was an explicit supervised relaunch. + +## What is still unproven + +The unauthenticated failure mode was never observed; this host's agy runs signed in, so any auth prompt is a fail-loud credential blocker, not a handled dialog. +No slash-skill invocation form was verified, so skill invocation stays natural language. +`--continue` and `--conversation` resume were never exercised; recovery uses deterministic relaunch from the brief on disk. +No primary or secondmate behavior was built or tested, and none is claimed. + +## Refreshing this record + +Run the portable suite and the live guard after any agy upgrade, because the process name, marker set, trust dialog text, and rendered busy/interrupt text are all vendor-controlled surfaces that the spawn gate and the busy fallback match verbatim: + +``` +bin/fm-test-run.sh tests/fm-agy-harness.test.sh +FM_AGY_SIGNALS_LIVE=1 bin/fm-test-run.sh tests/fm-agy-signals-live-e2e.test.sh +``` diff --git a/tests/fm-agy-harness.test.sh b/tests/fm-agy-harness.test.sh new file mode 100755 index 00000000000..6acaf588d79 --- /dev/null +++ b/tests/fm-agy-harness.test.sh @@ -0,0 +1,905 @@ +#!/usr/bin/env bash +# Behavior tests for the verified Antigravity CLI crewmate/scout adapter. +# +# The facts pinned here are the ones an agy release could silently change and +# the ones a wrong guess would make dangerous: +# 1. agy publishes no harness-identity marker of its own (a live 1.2.0 TUI +# carries no AGY_* variable; AGENT=1 there is inherited launcher state), +# so detection is ancestry alone on the anchored process name `agy`. +# 2. The anchored match must never claim unrelated commands containing the +# fragment, and an inherited CLAUDECODE still outranks ancestry until the +# spawn clears it - the clearing is load-bearing, not cosmetic. +# 3. The launch carries the brief via --prompt-interactive with --model, +# --effort, and --dangerously-skip-permissions; a requested model a +# reachable `agy models` omits refuses loudly instead of wedging a pane, +# while a hung or unreachable listing is cut off and never blocks. +# 4. A fresh worktree would park agy on its folder-trust dialog, so the spawn +# pre-registers the worktree in agy's own trustedWorkspaces store through +# bin/fm-agy-trust.sh (scope-refused for anything but a linked worktree +# of the project) and the post-launch gate is the backstop: it answers a +# dialog that renders anyway exactly once, never counts a busy turn as +# ready on an unregistered path until the dialog has been answered (the +# Herdr native-busy-before-dialog race), and fails the spawn with endpoint +# cleanup when the brief cannot be confirmed to run in the worktree. +# 5. agy is a crewmate/scout adapter only: a secondmate launch is refused, +# and nothing is armed as busy wiring because no writer could clear it. +# 6. The busy signature is the pinned `esc to cancel` status row alone; the +# free-floating `Generating...` word must never read busy on its own. +# 7. Herdr's registry already tracks agy, and exit detection proves the +# agent at process level before trusting any registration (the shared +# post-#4115 contract in bin/backends/herdr.sh): a registered status plus +# a process view naming agy is live and refuses replacement, a registered +# status over a proven shell-only pane is the explicit stale-agent state, +# and nothing short of that shared proof flips an agy pane to agent-free. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +# bin/fm-harness.sh checks verified ENV markers before ancestry. A suite run +# from inside another harness inherits those markers, which outrank the fake +# ancestry the detection cases set up. Drop the ambient markers so the asserted +# verdict does not depend on which harness launched the suite. +unset CLAUDECODE PI_CODING_AGENT FM_PI_HARNESS GROK_AGENT CURSOR_AGENT CURSOR_INVOKED_AS \ + ATLASSIAN_AGENT_TYPE ROVODEV_CLI GEMINI_CLI AGENT FM_OMP_HARNESS + +# shellcheck source=/dev/null +. "$ROOT/bin/fm-control-lib.sh" +# shellcheck source=/dev/null +. "$ROOT/bin/fm-busy-lib.sh" +# shellcheck source=/dev/null +. "$ROOT/bin/fm-composer-lib.sh" + +HARNESS="$ROOT/bin/fm-harness.sh" +SPAWN="$ROOT/bin/fm-spawn.sh" +TRUST="$ROOT/bin/fm-agy-trust.sh" +TMP_ROOT=$(fm_test_tmproot fm-agy-harness) + +# The store is agy's own persisted settings JSON, so trust is asserted against +# the parsed trustedWorkspaces array and preservation against parsed values. +agy_trusted_paths() { # <store> + node -e 'const fs=require("node:fs");const j=fs.existsSync(process.argv[1])?JSON.parse(fs.readFileSync(process.argv[1],"utf8")):{};for(const p of (j.trustedWorkspaces||[]))console.log(p);' "$1" +} + +agy_store_value() { # <store> <key> + node -e 'const j=JSON.parse(require("node:fs").readFileSync(process.argv[1],"utf8"));console.log(JSON.stringify(j[process.argv[2]]));' "$1" "$2" +} + +assert_agy_trusted() { # <store> <path> <msg> + agy_trusted_paths "$1" | grep -Fqx "$2" || fail "$3" +} + +assert_agy_not_trusted() { # <store> <path> <msg> + agy_trusted_paths "$1" | grep -Fqx "$2" && fail "$3" + return 0 +} + +test_agy_ancestry_detects_the_native_command_name() { + local fakebin out + fakebin=$(fm_fakebin "$TMP_ROOT/anc-native") + cat > "$fakebin/ps" <<'SH' +#!/usr/bin/env bash +case "$*" in + *"comm="*) printf '%s\n' '/usr/local/bin/agy'; exit 0 ;; + *"args="*) printf '%s\n' 'agy --prompt-interactive hello'; exit 0 ;; +esac +exit 1 +SH + chmod +x "$fakebin/ps" + out=$(PATH="$fakebin:$PATH" "$HARNESS") + [ "$out" = agy ] \ + || fail "a natively-named agy command must be detected by ancestry, got '$out'" + pass "fm-harness.sh: ancestry detects a natively-named agy command" +} + +test_agy_ancestry_rejects_unrelated_mentions() { + local fakebin out + fakebin=$(fm_fakebin "$TMP_ROOT/anc-negatives") + cat > "$fakebin/ps" <<'SH' +#!/usr/bin/env bash +case "$*" in + *"comm="*) printf '%s\n' "${FAKE_PS_COMM:?}"; exit 0 ;; + *"args="*) printf '%s\n' "${FAKE_PS_ARGS:?}"; exit 0 ;; +esac +exit 1 +SH + chmod +x "$fakebin/ps" + + out=$(FAKE_PS_COMM=magyk FAKE_PS_ARGS='magyk --serve' \ + PATH="$fakebin:$PATH" "$HARNESS") + [ "$out" != agy ] \ + || fail "an unrelated magyk command must not detect agy, got '$out'" + + out=$(FAKE_PS_COMM=bash FAKE_PS_ARGS='bash -c "echo agy --help"' \ + PATH="$fakebin:$PATH" "$HARNESS") + [ "$out" != agy ] \ + || fail "a later shell argument naming agy must not detect agy, got '$out'" + pass "fm-harness.sh: ancestry rejects unrelated agy mentions" +} + +test_agy_claims_no_inherited_launcher_marker() { + local out + # AGENT=1 was observed on a live agy TUI as inherited launcher state, so it + # must never promote to an agy identity the way GEMINI_CLI does for gemini. + out=$(AGENT=1 "$HARNESS") + [ "$out" != agy ] \ + || fail "an inherited AGENT=1 must never claim the agy identity, got '$out'" + # Drive the hazard the other way: agy does not clear an inherited CLAUDECODE, + # so the marker still wins over a real agy ancestor until the spawn clears it + # at the launch boundary. Pin both halves so neither can rot silently. + out=$(CLAUDECODE=1 FAKE_PS_COMM=agy FAKE_PS_ARGS='agy --prompt-interactive hi' \ + PATH="$(fm_fakebin "$TMP_ROOT/anc-claude"):$PATH" "$HARNESS") + [ "$out" = claude ] \ + || fail "an inherited CLAUDECODE must still outrank agy ancestry, got '$out'" + pass "fm-harness.sh: no inherited launcher marker claims the agy identity" +} + +test_agy_control_mechanics_are_the_verified_ones() { + fm_control_harness_supported agy || fail "agy must be a supported control harness" + [ "$(fm_control_harness_family agy)" = agy ] || fail "agy must map to its own family" + fm_control_harness_supports_kind agy scout || fail "agy must run scouts" + fm_control_harness_supports_kind agy ship || fail "agy must run ships" + fm_control_harness_supports_kind agy secondmate \ + && fail "agy must refuse secondmates" || true + [ "$(fm_control_interrupt_key agy)" = Escape ] || fail "agy must interrupt on Escape" + [ "$(fm_control_interrupt_repeat agy)" = 1 ] || fail "agy must interrupt on a single press" + [ -z "$(fm_control_interrupt_clear_key agy)" ] || fail "agy must need no clear key" + [ "$(fm_control_interrupt_ack_source agy)" = none ] || fail "agy must have no ack source" + [ "$(fm_control_exit_command agy)" = /quit ] || fail "agy must exit on /quit" + pass "fm-control-lib: agy mechanics are Escape once, no clear key, and /quit" +} + +test_agy_busy_tail_needs_the_pinned_status_row() { + printf 'working\nesc to cancel\n' | fm_busy_agy_tail_busy \ + || fail "the esc-to-cancel status row must read busy" + printf 'working\n Generating...\n' | fm_busy_agy_tail_busy \ + && fail "the free-floating Generating word alone must not read busy" || true + printf 'Generating report...\ndone\n? for shortcuts\n>\n' | fm_busy_agy_tail_busy \ + && fail "echoed worker output naming Generating must not read busy" || true + printf 'idle\n? for shortcuts\n>\n' | fm_busy_agy_tail_busy \ + && fail "an idle footer must not read busy" || true + printf 'Generating report...\ndone\n? for shortcuts\n>\n' | fm_busy_lines_match agy \ + && fail "the delivery guard must not acknowledge on echoed Generating output" || true + FM_BUSY_AGY_REGEX='idle' bash -c '. "$0/bin/fm-busy-lib.sh"; printf "idle\n" | fm_busy_agy_tail_busy' "$ROOT" \ + && fail "an environment override must not change the agy busy signature" || true + pass "fm-busy-lib: only the pinned esc-to-cancel row carries the agy busy verdict" +} + +test_agy_busy_signatures_are_harness_scoped() { + printf 'esc to cancel\n' | fm_busy_lines_match agy \ + || fail "harness=agy must match its own esc token" + printf 'esc to cancel\n' | fm_busy_lines_match grok \ + && fail "harness=grok must never borrow agy's esc token" || true + printf 'Ctrl+c:cancel\n' | fm_busy_lines_match agy \ + && fail "harness=agy must never borrow grok's token" || true + printf 'esc to cancel\n' | fm_busy_lines_match kimi \ + && fail "harness=kimi must never borrow agy's token" || true + printf 'esc to cancel\n' | fm_busy_lines_match spaceship \ + && fail "an unverified harness must match nothing" || true + pass "fm-composer-lib: agy delivery signatures never cross harnesses" +} + +test_agy_classify_reports_unknown_when_the_marker_scrolls_out() { + local statedir busy idle + statedir="$TMP_ROOT/classify"; mkdir -p "$statedir" + busy=$(fm_busy_classify tmux fake:win agy agy-case-1 "$statedir" 'turn running +esc to cancel Gemini 3.8 Flash · low') + [ "$busy" = "busy agy-regex" ] || fail "a busy tail must classify busy agy-regex, got '$busy'" + idle=$(fm_busy_classify tmux fake:win agy agy-case-2 "$statedir" 'reply landed +? for shortcuts Gemini 3.8 Flash · low') + [ "$idle" = "unknown agy-regex" ] || fail "a scrolled-out marker must classify unknown, got '$idle'" + pass "fm-busy-lib: agy classifies busy on its marker and unknown without it" +} + +test_agy_tmux_names_the_native_binary_an_agent() { + local got + # shellcheck source=/dev/null + . "$ROOT/bin/fm-backend.sh" + fm_backend_source tmux || fail "fm_backend_source tmux failed" + got=$(fm_agent_process_classify_name agy) + [ "$got" = agent ] || fail "tmux liveness must read the agy binary as an agent, got '$got'" + got=$(fm_agent_process_classify_name magyk) + [ "$got" = other ] || fail "tmux liveness must not read magyk as an agent, got '$got'" + got=$(fm_agent_process_classify_name bash) + [ "$got" = shell ] || fail "tmux liveness must still read bash as a shell, got '$got'" + pass "bin/fm-agent-process-lib.sh: agy is an agent, fragments are not" +} + +# Canned `pane process-info` bodies for the herdr fixtures. The shared +# exit-detection contract proves a registered agent at process level before +# trusting it (bin/backends/herdr.sh fm_backend_herdr_pane_process_state), so +# every registered-status fixture pairs its `agent get` body with a process +# view. The agy-shaped body names the foreground process exactly `agy`, which +# is the same identity surface the tmux liveness probe and the ancestry +# detector use - no real agy process is needed because the foreground branch +# answers before the descendant walk touches the process table. +agy_herdr_process_info_body() { # <shell-pid> <foreground-name> -> JSON + printf '%s\n' "{\"result\":{\"type\":\"pane_process_info\",\"process_info\":{\"pane_id\":\"w9:p1\",\"shell_pid\":$1,\"foreground_processes\":[{\"pid\":$(( $1 + 1 )),\"name\":\"$2\",\"argv\":[\"$2\",\"--prompt-interactive\"],\"argv0\":\"$2\",\"cmdline\":\"$2 --prompt-interactive\"}]}}}" +} + +agy_herdr_agent_state() { # <fixture-dir> -> verdict; logs every CLI call + local dir=$1 + : > "$dir/calls.log" + AGY_FIX_RESP="$dir/agent-get.json" AGY_FIX_PROC="$dir/process-info.json" \ + AGY_FIX_LOG="$dir/calls.log" bash -c ' + . "$0/bin/backends/herdr.sh" + fm_backend_herdr_pane_presence_state() { printf "present"; } + fm_backend_herdr_cli() { + printf "%s\n" "$*" >> "$AGY_FIX_LOG" + case "$*" in + *"agent get"*) cat "$AGY_FIX_RESP" ;; + *"pane process-info"*) cat "$AGY_FIX_PROC" ;; + *) exit 0 ;; + esac + } + fm_backend_herdr_pane_agent_state testsession w9:p1' "$ROOT" 2>&1 +} + +test_herdr_done_with_live_registry_stays_live() { + local dir out + dir="$TMP_ROOT/herdr-done"; mkdir -p "$dir" + printf '%s\n' '{"result":{"agent":{"agent":"agy","agent_status":"done","pane_id":"w9:p1"}}}' > "$dir/agent-get.json" + agy_herdr_process_info_body 424242 agy > "$dir/process-info.json" + out=$(agy_herdr_agent_state "$dir") + [ "$out" = live ] || fail "a registered done status with an agy process view must stay live, got '$out'" + grep -q "process-info" "$dir/calls.log" \ + || fail "the shared contract proves a registered agent at process level; the verdict trusted the registration alone" + out=$(AGY_FIX_RESP="$dir/agent-get.json" AGY_FIX_PROC="$dir/process-info.json" AGY_FIX_LOG="$dir/calls.log" bash -c ' + . "$0/bin/backends/herdr.sh" + fm_backend_herdr_pane_presence_state() { printf "present"; } + fm_backend_herdr_cli() { + case "$*" in + *"agent get"*) cat "$AGY_FIX_RESP" ;; + *"pane process-info"*) cat "$AGY_FIX_PROC" ;; + *) exit 0 ;; + esac + } + fm_backend_herdr_tab_is_husk testsession w9:p1 && printf husk || printf refused' "$ROOT" 2>&1) + [ "$out" = refused ] || fail "a live pane must refuse husk replacement, got '$out'" + pass "herdr exit detection: done with a live registry and an agy process view stays live and refuses replacement" +} + +test_herdr_registered_status_over_a_shell_only_pane_is_stale_not_live() { + local dir out shell_pid + dir="$TMP_ROOT/herdr-stale"; mkdir -p "$dir" + # The descendant walk reads the REAL process table, so the canned pane shell + # must be a process this test owns and can prove alive: a short-lived sleep. + sleep 30 & shell_pid=$! + printf '%s\n' '{"result":{"agent":{"agent":"agy","agent_status":"done","pane_id":"w9:p1"}}}' > "$dir/agent-get.json" + agy_herdr_process_info_body "$shell_pid" bash > "$dir/process-info.json" + out=$(agy_herdr_agent_state "$dir") + kill "$shell_pid" 2>/dev/null || true + [ "$out" = stale-agent ] || fail "a registered status over a proven shell-only pane must read stale-agent, got '$out'" + out=$(AGY_FIX_RESP="$dir/agent-get.json" AGY_FIX_PROC="$dir/process-info.json" AGY_FIX_LOG="$dir/calls.log" bash -c ' + . "$0/bin/backends/herdr.sh" + fm_backend_herdr_pane_presence_state() { printf "present"; } + fm_backend_herdr_cli() { + case "$*" in + *"agent get"*) cat "$AGY_FIX_RESP" ;; + *"pane process-info"*) cat "$AGY_FIX_PROC" ;; + *) exit 0 ;; + esac + } + fm_backend_herdr_tab_is_husk testsession w9:p1 && printf husk || printf refused' "$ROOT" 2>&1) + [ "$out" = refused ] || fail "a stale registration must still refuse husk replacement, got '$out'" + pass "herdr exit detection: a registered status over a shell-only pane is stale-agent and still refuses closing" +} + +test_herdr_shell_first_with_live_registry_stays_live() { + local dir out + dir="$TMP_ROOT/herdr-idle"; mkdir -p "$dir" + printf '%s\n' '{"result":{"agent":{"agent":"agy","agent_status":"idle","pane_id":"w9:p1"}}}' > "$dir/agent-get.json" + # The pane shell is present in the process view too (shell_pid), but the + # foreground names agy: the verified harness identity outranks shell-first + # ranking, and the shared contract's process proof is satisfied. + agy_herdr_process_info_body 424242 agy > "$dir/process-info.json" + out=$(agy_herdr_agent_state "$dir") + [ "$out" = live ] || fail "a registered idle status with an agy foreground must stay live, got '$out'" + grep -q "process-info" "$dir/calls.log" \ + || fail "the shared contract proves a registered agent at process level; the verdict trusted the registration alone" + pass "herdr exit detection: a registered pane with an agy foreground stays live however its shell ranks" +} + +test_herdr_lone_unregistered_pane_is_agent_free() { + local dir out + dir="$TMP_ROOT/herdr-gone"; mkdir -p "$dir" + printf '%s\n' '{"error":{"code":"agent_not_found","message":"agent target w9:p1 not found"}}' > "$dir/agent-get.json" + out=$(agy_herdr_agent_state "$dir") + [ "$out" = no-agent ] || fail "an unregistered pane must read no-agent, got '$out'" + out=$(AGY_FIX_RESP="$dir/agent-get.json" AGY_FIX_LOG="$dir/calls.log" bash -c ' + . "$0/bin/backends/herdr.sh" + fm_backend_herdr_pane_presence_state() { printf "present"; } + fm_backend_herdr_cli() { + case "$*" in *"agent get"*) cat "$AGY_FIX_RESP" ;; *) exit 0 ;; esac + } + fm_backend_herdr_tab_is_husk testsession w9:p1 && printf husk || printf refused' "$ROOT" 2>&1) + [ "$out" = husk ] || fail "an agent-free pane must allow husk replacement, got '$out'" + pass "herdr exit detection: only a positively unregistered pane is agent-free" +} + +test_herdr_malformed_and_failed_reads_stay_unknown() { + local dir out + dir="$TMP_ROOT/herdr-malformed"; mkdir -p "$dir" + printf '%s\n' '{not json at all' > "$dir/agent-get.json" + out=$(agy_herdr_agent_state "$dir") + [ "$out" = unknown ] || fail "a malformed registry response must read unknown, got '$out'" + dir="$TMP_ROOT/herdr-failed"; mkdir -p "$dir" + printf '%s\n' '{"result":{}}' > "$dir/agent-get.json" + export AGY_FIX_FAIL=1 + out=$(AGY_FIX_RESP="$dir/agent-get.json" AGY_FIX_LOG="$dir/calls.log" bash -c ' + . "$0/bin/backends/herdr.sh" + fm_backend_herdr_pane_presence_state() { printf "present"; } + fm_backend_herdr_cli() { + printf "%s\n" "$*" >> "$AGY_FIX_LOG" + case "$*" in *"agent get"*) [ "${AGY_FIX_FAIL:-0}" = 1 ] && exit 3; cat "$AGY_FIX_RESP" ;; *) exit 0 ;; esac + } + fm_backend_herdr_pane_agent_state testsession w9:p1' "$ROOT" 2>&1) + unset AGY_FIX_FAIL + [ "$out" = unknown ] || fail "a failed registry query must read unknown, got '$out'" + pass "herdr exit detection: malformed and failed reads stay unknown" +} + +make_agy_trust_case() { # <name> -> "<case>|<proj>|<wt>|<home>" + local name=$1 case_dir proj wt home + case_dir="$TMP_ROOT/trust-$name" + proj="$case_dir/project" + wt="$case_dir/wt" + home="$case_dir/home" + mkdir -p "$home" + fm_git_worktree "$proj" "$wt" "wt-trust-$name" + printf '%s|%s|%s|%s\n' "$case_dir" "$proj" "$wt" "$home" +} + +read_agy_trust_case() { + IFS='|' read -r CASE_DIR PROJ_DIR WT_DIR HOME_DIR <<EOF +$1 +EOF +} + +run_agy_trust() { # <home> <worktree> <project> + HOME="$1" "$TRUST" "$2" "$3" 2>&1 +} + +test_agy_trust_registers_the_logical_and_resolved_worktree_paths() { + local rec store out link + rec=$(make_agy_trust_case fresh) + read_agy_trust_case "$rec" + store="$HOME_DIR/.gemini/antigravity-cli/settings.json" + mkdir -p "$(dirname "$store")" + printf '%s\n' '{"model":"Gemini 3.8 Flash (High)","allowNonWorkspaceAccess":true,"trustedWorkspaces":["/home/someone/elsewhere"]}' > "$store" + link="$CASE_DIR/wt-link" + ln -s "$WT_DIR" "$link" + out=$(run_agy_trust "$HOME_DIR" "$link" "$PROJ_DIR") || fail "a fresh linked worktree must be trusted: $out" + assert_agy_trusted "$store" "$link" "the logical (symlinked) pane path agy compares against was not registered" + assert_agy_trusted "$store" "$WT_DIR" "the resolved worktree path was not registered alongside the logical one" + assert_agy_trusted "$store" "/home/someone/elsewhere" "registration dropped an existing trustedWorkspaces entry" + [ "$(agy_store_value "$store" model)" = '"Gemini 3.8 Flash (High)"' ] \ + || fail "registration did not preserve an unrelated store key" + [ "$(agy_store_value "$store" allowNonWorkspaceAccess)" = true ] \ + || fail "registration did not preserve an unrelated boolean key" + out=$(run_agy_trust "$HOME_DIR" "$link" "$PROJ_DIR") || fail "repeat registration must succeed: $out" + [ "$(agy_trusted_paths "$store" | grep -Fcx "$WT_DIR")" -eq 1 ] \ + || fail "repeat registration duplicated the worktree entry" + pass "fm-agy-trust.sh: registers the logical and resolved worktree paths and preserves the store" +} + +test_agy_trust_creates_a_missing_store() { + local rec store out + rec=$(make_agy_trust_case nostore) + read_agy_trust_case "$rec" + store="$HOME_DIR/.gemini/antigravity-cli/settings.json" + out=$(run_agy_trust "$HOME_DIR" "$WT_DIR" "$PROJ_DIR") || fail "a missing store must be created: $out" + [ -f "$store" ] || fail "no settings store was created at $store" + assert_agy_trusted "$store" "$WT_DIR" "the worktree was not registered in the created store" + pass "fm-agy-trust.sh: creates agy's settings store when none exists" +} + +test_agy_trust_refuses_out_of_scope_paths() { + local rec store out rc plain before after + rec=$(make_agy_trust_case scope) + read_agy_trust_case "$rec" + store="$HOME_DIR/.gemini/antigravity-cli/settings.json" + mkdir -p "$(dirname "$store")" + printf '%s\n' '{"trustedWorkspaces":[]}' > "$store" + rc=0; out=$(run_agy_trust "$HOME_DIR" "$PROJ_DIR" "$PROJ_DIR") || rc=$? + [ "$rc" -ne 0 ] || fail "the primary checkout must be refused" + assert_contains "$out" "primary checkout" "primary-checkout refusal lacked its reason" + assert_agy_not_trusted "$store" "$PROJ_DIR" "a refused primary checkout was still registered" + rc=0; out=$(run_agy_trust "$HOME_DIR" "$HOME_DIR" "$PROJ_DIR") || rc=$? + [ "$rc" -ne 0 ] || fail "the home directory must be refused" + assert_agy_not_trusted "$store" "$HOME_DIR" "a refused home directory was still registered" + plain="$CASE_DIR/plain"; mkdir -p "$plain" + rc=0; out=$(run_agy_trust "$HOME_DIR" "$plain" "$PROJ_DIR") || rc=$? + [ "$rc" -ne 0 ] || fail "a plain directory must be refused" + assert_agy_not_trusted "$store" "$plain" "a refused plain directory was still registered" + rc=0; out=$(run_agy_trust "$HOME_DIR" "$WT_DIR/.git" "$PROJ_DIR") || rc=$? + [ "$rc" -ne 0 ] || fail "a path below the worktree root must be refused" + printf '%s\n' '{not json' > "$store" + before=$(cat "$store") + rc=0; out=$(run_agy_trust "$HOME_DIR" "$WT_DIR" "$PROJ_DIR") || rc=$? + [ "$rc" -ne 0 ] || fail "an unparseable store must be refused" + after=$(cat "$store") + [ "$before" = "$after" ] || fail "an unparseable store was rewritten" + pass "fm-agy-trust.sh: refuses every out-of-scope path and never rewrites a broken store" +} + +# The fake tmux renders an agy-shaped screen that advances through +# launched -> (trust dialog ->) busy as the real spawn drives it, so the launch +# command, the pre-registration, the single Enter that answers a dialog, and +# the readiness gate are exercised through their real code paths. Whether the +# dialog renders is decided the way agy decides it: the pane path is looked up +# in the trustedWorkspaces array of the store the spawn just wrote. +# FM_FAKE_AGY_IGNORE_TRUST=1 models a vendor that stopped honouring the store; +# FM_FAKE_AGY_ASSUME_TRUSTED=1 models a pane that never shows the dialog even +# though firstmate could not register the path (a busy verdict with no proof +# of where the turn runs); +# FM_FAKE_AGY_RACE=1 models Herdr's native busy verdict rendering one capture +# before the dialog paints; FM_FAKE_AGY_ANSWER=stuck models a dialog whose +# answer never turns into a busy turn. +make_agy_fakebin() { + local dir=$1 fakebin + fakebin=$(fm_fakebin "$dir") + cat > "$fakebin/tmux" <<'SH' +#!/usr/bin/env bash +set -u +printf '%s\n' "$*" >> "$FM_FAKE_TMUX_CALL_LOG" +state=$(cat "$FM_FAKE_AGY_STATE" 2>/dev/null || true) +fake_screen() { + case "$state" in + dialog) + printf 'Accessing workspace:\n\n%s\n\nDo you trust the contents of this project?\n\nAntigravity CLI requires permission to read, edit, and execute files here.\n\n> Yes, I trust this folder\n No, exit\n' "$FM_FAKE_PANE_PATH" + ;; + busy) + printf 'Generating...\n└ Tip: press f to see the full diff.\n\nesc to cancel Gemini 3.8 Flash · low\n' + ;; + racing) + printf 'esc to cancel Gemini 3.8 Flash · low\n' + printf 'dialog\n' > "$FM_FAKE_AGY_STATE" + ;; + *) + printf 'shell starting\n$ \n' + ;; + esac +} +fake_path_trusted() { + [ "${FM_FAKE_AGY_ASSUME_TRUSTED:-0}" = 1 ] && return 0 + [ "${FM_FAKE_AGY_IGNORE_TRUST:-0}" = 1 ] && return 1 + node -e 'const fs=require("node:fs");let j={};try{j=JSON.parse(fs.readFileSync(process.argv[1],"utf8"));}catch(e){process.exit(1);}process.exit(Array.isArray(j.trustedWorkspaces)&&j.trustedWorkspaces.includes(process.argv[2])?0:1);' \ + "$FM_FAKE_AGY_SETTINGS" "$FM_FAKE_PANE_PATH" +} +case "$*" in + *"#{pane_current_path}"*) printf '%s\n' "$FM_FAKE_PANE_PATH"; exit 0 ;; + *"#{cursor_y}"*) printf '1\n'; exit 0 ;; +esac +case "${1:-}" in + display-message) printf 'firstmate\n'; exit 0 ;; + list-windows) exit 0 ;; + has-session|new-session|new-window|kill-window) exit 0 ;; + send-keys) + literal= + prev= + for arg in "$@"; do + if [ "$prev" = -l ]; then literal=$arg; break; fi + prev=$arg + done + if [ -n "$literal" ]; then + case "$literal" in + *--prompt-interactive*) + printf '%s\n' "$literal" >> "$FM_FAKE_LAUNCH_LOG" + printf 'launched\n' > "$FM_FAKE_AGY_STATE" + ;; + esac + exit 0 + fi + case " $* " in + *' Enter '*) + case "$state" in + launched) + if fake_path_trusted; then + printf 'busy\n' > "$FM_FAKE_AGY_STATE" + elif [ "${FM_FAKE_AGY_RACE:-0}" = 1 ]; then + printf 'racing\n' > "$FM_FAKE_AGY_STATE" + else + printf 'dialog\n' > "$FM_FAKE_AGY_STATE" + fi + ;; + dialog) + if [ "${FM_FAKE_AGY_ANSWER:-works}" = works ]; then + printf 'busy\n' > "$FM_FAKE_AGY_STATE" + fi + ;; + esac + ;; + esac + exit 0 + ;; + capture-pane) fake_screen; exit 0 ;; +esac +exit 0 +SH + chmod +x "$fakebin/tmux" + cat > "$fakebin/agy" <<'SH' +#!/usr/bin/env bash +set -u +if [ "${1:-}" = models ]; then + if [ "${FM_FAKE_AGY_MODELS_FAIL:-0}" = 1 ]; then exit 3; fi + if [ "${FM_FAKE_AGY_MODELS_HANG:-0}" = 1 ]; then cat > /dev/null; sleep 30; exit 0; fi + printf 'gemini-3.8-flash-high\tGemini 3.8 Flash (High)\n' + printf 'gemini-3.8-flash-medium\tGemini 3.8 Flash (Medium)\n' + printf 'gemini-3.8-flash-low\tGemini 3.8 Flash (Low)\n' + exit 0 +fi +echo "fake agy must never execute" >&2 +exit 9 +SH + chmod +x "$fakebin/agy" + fm_fake_exit0 "$fakebin" treehouse gh-axi gh + printf '%s\n' "$fakebin" +} + +make_agy_spawn_case() { + local name=$1 id=$2 case_dir home proj wt fakebin + case_dir="$TMP_ROOT/$name" + home="$case_dir/home" + proj="$case_dir/project" + wt="$case_dir/wt" + fakebin=$(make_agy_fakebin "$case_dir/fake") + mkdir -p "$home/data/$id" "$home/projects" "$home/state" "$home/config" + cat > "$home/data/$id/brief.md" <<'EOF' +# Task +## Captain's intent +Exercise Antigravity dispatch. + +## Firstmate spec +Verify launch and delivery behavior. +EOF + printf 'agy\n' > "$home/config/crew-harness" + mkdir -p "$home/.gemini/antigravity-cli" + printf '%s\n' '{"model":"Gemini 3.8 Flash (High)","trustedWorkspaces":["/home/someone/elsewhere"]}' \ + > "$home/.gemini/antigravity-cli/settings.json" + fm_git_worktree "$proj" "$wt" "wt-$name" + touch "$home/state/.last-watcher-beat" + : > "$case_dir/launch.log" + : > "$case_dir/tmux-calls.log" + : > "$case_dir/agy.state" + printf '%s\n' "$case_dir|$home|$proj|$wt|$fakebin" +} + +read_agy_spawn_record() { + IFS='|' read -r CASE_DIR HOME_DIR PROJ_DIR WT_DIR FAKEBIN_DIR <<EOF +$1 +EOF +} + +# The spawn drives the real bin/fm-agy-trust.sh and the fake tmux's trust +# lookup under this base PATH, and both read agy's settings store with node, +# which runners do not keep in the system bin dirs. Carry the directory the +# invoking environment resolves node from, the fm-kimi-harness shape. +NODE_BIN=$(command -v node) || fail "test needs node" +NODE_BIN_DIR=$(dirname "$NODE_BIN") +BASE_PATH=${FM_TEST_BASE_PATH:-$NODE_BIN_DIR:/usr/bin:/bin:/usr/sbin:/sbin} + +run_agy_spawn() { + local case_dir=$1 home=$2 proj=$3 wt=$4 fakebin=$5 id=$6 + shift 6 + HOME="$home" FM_ROOT_OVERRIDE='' FM_HOME="$home" \ + FM_STATE_OVERRIDE="$home/state" FM_DATA_OVERRIDE="$home/data" \ + FM_PROJECTS_OVERRIDE="$home/projects" FM_CONFIG_OVERRIDE="$home/config" \ + FM_SPAWN_NO_GUARD=1 FM_FAKE_PANE_PATH="$wt" TMUX="fake,1,0" \ + FM_FAKE_LAUNCH_LOG="$case_dir/launch.log" \ + FM_FAKE_TMUX_CALL_LOG="$case_dir/tmux-calls.log" \ + FM_FAKE_AGY_STATE="$case_dir/agy.state" \ + FM_FAKE_AGY_SETTINGS="$home/.gemini/antigravity-cli/settings.json" \ + FM_FAKE_AGY_MODELS_FAIL="${FM_FAKE_AGY_MODELS_FAIL:-0}" \ + FM_FAKE_AGY_MODELS_HANG="${FM_FAKE_AGY_MODELS_HANG:-0}" \ + FM_FAKE_AGY_IGNORE_TRUST="${FM_FAKE_AGY_IGNORE_TRUST:-0}" \ + FM_FAKE_AGY_ASSUME_TRUSTED="${FM_FAKE_AGY_ASSUME_TRUSTED:-0}" \ + FM_FAKE_AGY_RACE="${FM_FAKE_AGY_RACE:-0}" \ + FM_FAKE_AGY_ANSWER="${FM_FAKE_AGY_ANSWER:-works}" \ + FM_AGY_READY_POLLS=4 FM_AGY_POLL_INTERVAL=0 FM_AGY_MODELS_TIMEOUT=${FM_AGY_MODELS_TIMEOUT:-1} \ + PATH="$fakebin:$BASE_PATH" \ + "$SPAWN" "$id" "$proj" --harness agy --mode no-mistakes --yolo off "$@" 2>&1 +} + +test_agy_launch_carries_the_brief_with_model_effort_and_autonomy() { + local id rec out rc launch meta + id="agy-launch-z1-$$" + rec=$(make_agy_spawn_case launch "$id") + read_agy_spawn_record "$rec" + out=$(run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$id" \ + --model gemini-3.8-flash-low --effort low) + rc=$? + expect_code 0 "$rc" "agy spawn with a listed model should succeed" + launch=$(cat "$CASE_DIR/launch.log") + assert_contains "$launch" "$FAKEBIN_DIR/agy" "agy launch did not pin the resolved absolute binary" + assert_contains "$launch" "--prompt-interactive" "agy launch did not carry the brief via --prompt-interactive" + assert_contains "$launch" "--model 'gemini-3.8-flash-low'" "agy launch did not carry the requested model" + assert_contains "$launch" "--effort 'low'" "agy launch did not carry the requested effort" + assert_contains "$launch" "--dangerously-skip-permissions" "agy launch omitted unattended autonomy" + assert_contains "$launch" "env -u CLAUDECODE" "agy launch did not clear the inherited launcher marker" + assert_not_contains "$launch" "__AGYBIN__" "agy launch left its binary placeholder unsubstituted" + assert_not_contains "$launch" "__MODELFLAG__" "agy launch left its model placeholder unsubstituted" + assert_not_contains "$launch" "__BRIEF__" "agy launch left its brief placeholder unsubstituted" + meta="$HOME_DIR/state/$id.meta" + assert_grep 'harness=agy' "$meta" "agy meta did not record its harness" + assert_grep 'model=gemini-3.8-flash-low' "$meta" "agy meta did not record its model" + assert_grep 'effort=low' "$meta" "agy meta did not record its effort" + pass "fm-spawn: agy launch carries brief, model, effort, and autonomy with cleared markers" +} + +test_agy_effort_xhigh_is_recorded_but_omitted() { + local id rec out rc launch meta + id="agy-xhigh-z2-$$" + rec=$(make_agy_spawn_case xhigh "$id") + read_agy_spawn_record "$rec" + out=$(run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$id" \ + --model gemini-3.8-flash-low --effort xhigh) + rc=$? + expect_code 0 "$rc" "agy spawn with an unsupported effort should still succeed" + launch=$(cat "$CASE_DIR/launch.log") + assert_not_contains "$launch" "--effort" "agy launch passed a known-bad effort value" + meta="$HOME_DIR/state/$id.meta" + assert_grep 'effort=xhigh' "$meta" "agy meta did not retain the unsupported effort axis" + pass "fm-spawn: agy omits xhigh from the launch but records it in task metadata" +} + +test_agy_unlisted_model_refuses_before_pane_creation() { + local id rec out rc + id="agy-badmodel-z3-$$" + rec=$(make_agy_spawn_case badmodel "$id") + read_agy_spawn_record "$rec" + rc=0 + out=$(run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$id" \ + --model gemini-3.8-flash) || rc=$? + [ "$rc" -ne 0 ] || fail "an unlisted agy model should refuse the spawn" + assert_contains "$out" "not listed by 'agy models'" "unlisted model refusal lacked its concrete reason" + [ -s "$CASE_DIR/launch.log" ] && fail "an unlisted model created a launch command" || true + pass "fm-spawn: an unlisted agy model refuses before pane creation" +} + +test_agy_unreachable_listing_launches_unvalidated() { + local id rec out rc + id="agy-nolisting-z4-$$" + rec=$(make_agy_spawn_case nolisting "$id") + read_agy_spawn_record "$rec" + rc=0 + out=$(FM_FAKE_AGY_MODELS_FAIL=1 run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" \ + "$FAKEBIN_DIR" "$id" --model gemini-3.8-flash-low) || rc=$? + expect_code 0 "$rc" "an unreachable model listing must not block the spawn" + [ -s "$CASE_DIR/launch.log" ] || fail "an unreachable listing produced no launch command" + assert_contains "$out" "listing is unreachable" "an unreachable listing launched without its notice" + pass "fm-spawn: an unreachable agy listing establishes nothing and launches" +} + +test_agy_hung_listing_is_cut_off_and_launches() { + local id rec out rc started elapsed + id="agy-hanglisting-z8-$$" + rec=$(make_agy_spawn_case hanglisting "$id") + read_agy_spawn_record "$rec" + rc=0 + started=$(date +%s) + out=$(FM_FAKE_AGY_MODELS_HANG=1 run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" \ + "$FAKEBIN_DIR" "$id" --model gemini-3.8-flash-low) || rc=$? + elapsed=$(( $(date +%s) - started )) + expect_code 0 "$rc" "a hung model listing must not block the spawn" + [ "$elapsed" -lt 20 ] || fail "the model probe was not cut off by its bound (took ${elapsed}s)" + assert_contains "$out" "did not answer within 1s" "a hung listing launched without its timeout notice" + [ -s "$CASE_DIR/launch.log" ] || fail "a hung listing produced no launch command" + assert_contains "$(cat "$CASE_DIR/launch.log")" "--model 'gemini-3.8-flash-low'" \ + "a hung listing dropped the requested model instead of launching it unvalidated" + pass "fm-spawn: a hung agy listing is cut off by the shared bound and launches unvalidated" +} + +test_agy_zero_model_timeout_is_clamped_to_the_default_bound() { + local id rec out rc started elapsed + id="agy-zerobound-z14-$$" + rec=$(make_agy_spawn_case zerobound "$id") + read_agy_spawn_record "$rec" + rc=0 + started=$(date +%s) + out=$(FM_FAKE_AGY_MODELS_HANG=1 FM_AGY_MODELS_TIMEOUT=0 \ + run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" \ + "$FAKEBIN_DIR" "$id" --model gemini-3.8-flash-low) || rc=$? + elapsed=$(( $(date +%s) - started )) + expect_code 0 "$rc" "a hung listing with a zero bound must not block the spawn" + [ "$elapsed" -lt 25 ] || fail "a zero model bound disabled the deadline (took ${elapsed}s)" + assert_contains "$out" "did not answer within 15s" \ + "a zero model bound was not clamped to the documented default" + [ -s "$CASE_DIR/launch.log" ] || fail "a zero model bound produced no launch command" + pass "fm-spawn: a zero FM_AGY_MODELS_TIMEOUT is clamped to the default bound" +} + +# Bare Enter key presses only: shell setup rides its Enter on the typed text +# (`send-keys -t <target> export X=Y Enter`), while the launch submit and the +# trust-dialog answer are lone key sends (`send-keys -t <target> Enter`). +count_enter_sends() { # <tmux-call-log> + grep -c '^send-keys -t [^ ]* Enter$' "$1" || true +} + +test_agy_fresh_worktree_is_pre_trusted_and_launches_without_a_dialog() { + local id rec out rc enters store + id="agy-trust-z9-$$" + rec=$(make_agy_spawn_case trust "$id") + read_agy_spawn_record "$rec" + store="$HOME_DIR/.gemini/antigravity-cli/settings.json" + out=$(run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$id" \ + --model gemini-3.8-flash-low) + rc=$? + expect_code 0 "$rc" "an agy spawn into a fresh worktree should succeed" + assert_contains "$out" "spawned $id harness=agy" "agy spawn did not report success" + assert_not_contains "$out" "could not pre-register" "a legitimate worktree failed trust pre-registration" + assert_agy_trusted "$store" "$WT_DIR" "the spawn did not pre-register the worktree in agy's trust store" + assert_agy_trusted "$store" "/home/someone/elsewhere" "the spawn dropped an existing trustedWorkspaces entry" + [ "$(agy_store_value "$store" model)" = '"Gemini 3.8 Flash (High)"' ] \ + || fail "the spawn did not preserve an unrelated agy setting" + [ "$(cat "$CASE_DIR/agy.state")" = busy ] \ + || fail "the spawn reported success before the pane reached a busy turn (state: $(cat "$CASE_DIR/agy.state"))" + enters=$(count_enter_sends "$CASE_DIR/tmux-calls.log") + [ "$enters" -eq 1 ] \ + || fail "a pre-trusted worktree must receive only the launch Enter, got $enters Enter sends" + assert_not_contains "$(cat "$CASE_DIR/tmux-calls.log")" "kill-window" \ + "a successful agy spawn must never tear down the endpoint it just launched" + pass "fm-spawn: agy pre-registers the worktree and launches straight into a busy turn" +} + +test_agy_dialog_despite_registration_is_answered_once() { + local id rec out rc enters + id="agy-vendor-z10-$$" + rec=$(make_agy_spawn_case vendor-dialog "$id") + read_agy_spawn_record "$rec" + out=$(FM_FAKE_AGY_IGNORE_TRUST=1 run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" \ + "$FAKEBIN_DIR" "$id" --model gemini-3.8-flash-low) + rc=$? + expect_code 0 "$rc" "an agy spawn whose dialog renders despite registration should succeed" + [ "$(cat "$CASE_DIR/agy.state")" = busy ] \ + || fail "the spawn reported success before the pane reached a busy turn (state: $(cat "$CASE_DIR/agy.state"))" + enters=$(count_enter_sends "$CASE_DIR/tmux-calls.log") + [ "$enters" -eq 2 ] \ + || fail "expected exactly one launch Enter plus one trust-dialog Enter, got $enters Enter sends" + pass "fm-spawn: agy answers a dialog that renders anyway exactly once, then confirms busy" +} + +test_agy_unregistered_path_ignores_busy_until_the_dialog_is_answered() { + local id rec out rc enters store before after + id="agy-race-z11-$$" + rec=$(make_agy_spawn_case race "$id") + read_agy_spawn_record "$rec" + store="$HOME_DIR/.gemini/antigravity-cli/settings.json" + printf '%s\n' '{not json' > "$store" + before=$(cat "$store") + out=$(FM_FAKE_AGY_RACE=1 run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" \ + "$FAKEBIN_DIR" "$id" --model gemini-3.8-flash-low) + rc=$? + expect_code 0 "$rc" "an agy spawn that meets the dialog after a premature busy verdict should still succeed" + assert_contains "$out" "could not pre-register agy workspace trust" \ + "a broken store did not surface the registration warning" + after=$(cat "$store") + [ "$before" = "$after" ] || fail "the spawn rewrote an unparseable agy store" + [ "$(cat "$CASE_DIR/agy.state")" = busy ] \ + || fail "the spawn reported success before the answered dialog turned busy (state: $(cat "$CASE_DIR/agy.state"))" + enters=$(count_enter_sends "$CASE_DIR/tmux-calls.log") + [ "$enters" -eq 2 ] \ + || fail "a busy verdict before the dialog must not count as ready on an unregistered path; expected the dialog Enter, got $enters Enter sends" + pass "fm-spawn: on an unregistered path a premature busy verdict waits for the dialog to be answered" +} + +test_agy_unregistered_path_without_a_dialog_fails_the_spawn() { + local id rec out rc store + id="agy-nodialog-z12-$$" + rec=$(make_agy_spawn_case nodialog "$id") + read_agy_spawn_record "$rec" + store="$HOME_DIR/.gemini/antigravity-cli/settings.json" + printf '%s\n' '{not json' > "$store" + rc=0 + out=$(FM_FAKE_AGY_ASSUME_TRUSTED=1 run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" \ + "$FAKEBIN_DIR" "$id" --model gemini-3.8-flash-low) || rc=$? + [ "$rc" -ne 0 ] || fail "a busy verdict on an unregistered path with no dialog must not pass the gate" + assert_contains "$out" "never showed its folder-trust dialog on an unregistered worktree" \ + "the failure did not name the unconfirmed workspace" + assert_not_contains "$out" "spawned $id" "an unconfirmed workspace still reported a successful spawn" + [ "$(count_enter_sends "$CASE_DIR/tmux-calls.log")" -eq 1 ] \ + || fail "the gate must not send Enter into a pane that shows no dialog" + assert_contains "$(cat "$CASE_DIR/tmux-calls.log")" "kill-window" \ + "a failed agy readiness gate left its launched endpoint running" + assert_grep 'failed: agy never showed its folder-trust dialog' "$HOME_DIR/state/$id.status" \ + "a failed agy readiness gate did not record the failure in the task status" + pass "fm-spawn: a busy verdict on an unregistered path without a dialog fails and closes the endpoint" +} + +test_agy_pre_trusted_path_that_never_turns_busy_fails_the_spawn() { + local id rec out rc + id="agy-idle-z13-$$" + rec=$(make_agy_spawn_case idle "$id") + read_agy_spawn_record "$rec" + rc=0 + out=$(FM_FAKE_AGY_IGNORE_TRUST=1 FM_FAKE_AGY_ANSWER=stuck run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" \ + "$FAKEBIN_DIR" "$id" --model gemini-3.8-flash-low) || rc=$? + [ "$rc" -ne 0 ] || fail "a dialog that never turns into a busy turn must fail the spawn" + assert_contains "$out" "did not start processing its brief after the folder-trust dialog was answered" \ + "a stuck trust dialog failed without its concrete reason" + [ "$(count_enter_sends "$CASE_DIR/tmux-calls.log")" -eq 2 ] \ + || fail "the gate must answer the dialog exactly once and never hammer Enter" + assert_contains "$(cat "$CASE_DIR/tmux-calls.log")" "kill-window" \ + "a failed agy readiness gate left its launched endpoint running" + pass "fm-spawn: an agy dialog that never turns busy fails the spawn and closes the endpoint" +} + +test_agy_missing_binary_refuses_before_pane_creation() { + local id rec out rc + id="agy-missing-z5-$$" + rec=$(make_agy_spawn_case missing "$id") + read_agy_spawn_record "$rec" + rm "$FAKEBIN_DIR/agy" + rc=0 + out=$(run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$id") || rc=$? + [ "$rc" -ne 0 ] || fail "a missing agy executable should refuse the spawn" + assert_contains "$out" "agy executable not found on PATH" "missing agy diagnostic lacked its concrete reason" + [ -s "$CASE_DIR/launch.log" ] && fail "a missing agy executable created a launch command" || true + pass "fm-spawn: a missing agy executable refuses before pane creation" +} + +test_agy_secondmate_is_refused() { + local id rec out rc + id="agy-secondmate-z6-$$" + rec=$(make_agy_spawn_case secondmate-refuse "$id") + read_agy_spawn_record "$rec" + rc=0 + out=$(HOME="$HOME_DIR" FM_ROOT_OVERRIDE='' FM_HOME="$HOME_DIR" \ + FM_STATE_OVERRIDE="$HOME_DIR/state" FM_DATA_OVERRIDE="$HOME_DIR/data" \ + FM_PROJECTS_OVERRIDE="$HOME_DIR/projects" FM_CONFIG_OVERRIDE="$HOME_DIR/config" \ + FM_SPAWN_NO_GUARD=1 PATH="$FAKEBIN_DIR:$BASE_PATH" \ + "$SPAWN" "$id" --secondmate agy 2>&1) || rc=$? + [ "$rc" -ne 0 ] || fail "an agy secondmate spawn should be refused" + assert_contains "$out" "agy is a verified crewmate/scout adapter only" \ + "agy secondmate refusal lacked its concrete reason" + pass "fm-spawn: agy cannot be launched as a secondmate" +} + +test_agy_spawn_arms_no_busy_wiring() { + local id rec out rc statedir + id="agy-nowiring-z7-$$" + rec=$(make_agy_spawn_case nowiring "$id") + read_agy_spawn_record "$rec" + out=$(run_agy_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$id" \ + --model gemini-3.8-flash-low) + rc=$? + expect_code 0 "$rc" "agy spawn should succeed" + statedir="$HOME_DIR/state" + [ -e "$statedir/$id.busy-gen" ] && fail "agy spawn armed a busy generation nothing could clear" || true + for sidecar in "$statedir/$id.agy-"*; do + [ -e "$sidecar" ] || continue + fail "agy spawn left an adapter sidecar behind: $sidecar" + done + pass "fm-spawn: agy arms no busy wiring and writes no sidecar" +} + +test_agy_ancestry_detects_the_native_command_name +test_agy_ancestry_rejects_unrelated_mentions +test_agy_claims_no_inherited_launcher_marker +test_agy_control_mechanics_are_the_verified_ones +test_agy_busy_tail_needs_the_pinned_status_row +test_agy_busy_signatures_are_harness_scoped +test_agy_classify_reports_unknown_when_the_marker_scrolls_out +test_agy_tmux_names_the_native_binary_an_agent +test_herdr_done_with_live_registry_stays_live +test_herdr_registered_status_over_a_shell_only_pane_is_stale_not_live +test_herdr_shell_first_with_live_registry_stays_live +test_herdr_lone_unregistered_pane_is_agent_free +test_herdr_malformed_and_failed_reads_stay_unknown +test_agy_launch_carries_the_brief_with_model_effort_and_autonomy +test_agy_effort_xhigh_is_recorded_but_omitted +test_agy_unlisted_model_refuses_before_pane_creation +test_agy_unreachable_listing_launches_unvalidated +test_agy_hung_listing_is_cut_off_and_launches +test_agy_zero_model_timeout_is_clamped_to_the_default_bound +test_agy_trust_registers_the_logical_and_resolved_worktree_paths +test_agy_trust_creates_a_missing_store +test_agy_trust_refuses_out_of_scope_paths +test_agy_fresh_worktree_is_pre_trusted_and_launches_without_a_dialog +test_agy_dialog_despite_registration_is_answered_once +test_agy_unregistered_path_ignores_busy_until_the_dialog_is_answered +test_agy_unregistered_path_without_a_dialog_fails_the_spawn +test_agy_pre_trusted_path_that_never_turns_busy_fails_the_spawn +test_agy_missing_binary_refuses_before_pane_creation +test_agy_secondmate_is_refused +test_agy_spawn_arms_no_busy_wiring diff --git a/tests/fm-agy-signals-live-e2e.test.sh b/tests/fm-agy-signals-live-e2e.test.sh new file mode 100755 index 00000000000..59957fe793d --- /dev/null +++ b/tests/fm-agy-signals-live-e2e.test.sh @@ -0,0 +1,193 @@ +#!/usr/bin/env bash +# Live drift guard for the Antigravity CLI adapter's vendor-controlled surface: +# process name, trust dialog, rendered busy/interrupt/exit behavior. +# Opt-in because it submits real prompts (no echo provider exists for agy). +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +AGY_BIN=$(command -v agy 2>/dev/null || true) +REAL_TMUX=$(command -v tmux 2>/dev/null || true) +LAB= +SOCKET="fm-agy-signals-$$" +SESSION=agy-signals +TARGET="$SESSION:agy" + +cleanup() { + [ -n "$REAL_TMUX" ] && "$REAL_TMUX" -L "$SOCKET" kill-server >/dev/null 2>&1 || true + [ -z "$LAB" ] || rm -rf -- "$LAB" +} + +fail() { + printf 'not ok - %s\n' "$1" >&2 + cleanup + exit 1 +} + +pass() { + printf 'ok - %s\n' "$1" +} + +fm_live_gate opt-in FM_AGY_SIGNALS_LIVE agy tmux +[ -n "$AGY_BIN" ] || fail "agy is not installed" + +LAB=$(mktemp -d "${TMPDIR:-/tmp}/fm-agy-signals.XXXXXX") || fail "could not create the isolated agy lab" +trap cleanup EXIT +mkdir -p "$LAB/workspace" +git -C "$LAB/workspace" init -q || fail "could not initialize the isolated agy workspace" +git -C "$LAB/workspace" config user.email "guard@local" || fail "could not configure the isolated agy workspace" +git -C "$LAB/workspace" config user.name "guard" || fail "could not configure the isolated agy workspace" +git -C "$LAB/workspace" commit -q --allow-empty -m init || fail "could not seed the isolated agy workspace" +WORKSPACE=$(cd "$LAB/workspace" && pwd -P) || fail "could not resolve the isolated agy workspace" + +# The worker runs under a throwaway HOME holding a copy of ~/.gemini (the +# method recorded in docs/verification/agy.md), so its trust answer and every +# other agy write land in the lab store, never the operator's real one. +AGY_HOME="$LAB/home" +mkdir -p "$AGY_HOME" || fail "could not create the throwaway agy HOME" +[ -d "$HOME/.gemini" ] || fail "no ~/.gemini to stage for the throwaway agy HOME" +cp -R "$HOME/.gemini" "$AGY_HOME/.gemini" || fail "could not stage the throwaway agy credential copy" + +# shellcheck source=/dev/null +. "$ROOT/bin/fm-busy-lib.sh" +# shellcheck source=/dev/null +. "$ROOT/bin/fm-composer-lib.sh" + +"$REAL_TMUX" -L "$SOCKET" new-session -d -s "$SESSION" -n control -c "$WORKSPACE" \ + || fail "could not start the isolated tmux server" +"$REAL_TMUX" -L "$SOCKET" new-window -d -t "$SESSION:" -n agy -c "$WORKSPACE" \ + || fail "could not open the isolated agy window" + +capture() { + "$REAL_TMUX" -L "$SOCKET" capture-pane -p -t "$TARGET" -S -100 2>/dev/null || true +} + +# The launch prompt asks for a computed answer (12345+67890=80235) so the +# awaited token never appears in the echoed launch line itself, where a plain +# reply token would false-positive on the shell echo (including across tmux +# wrapped rows). +"$REAL_TMUX" -L "$SOCKET" send-keys -t "$TARGET" -l \ + "HOME=\"$AGY_HOME\" $AGY_BIN --prompt-interactive \"Add 12345 and 67890. Reply with exactly the sum and nothing else\" --model gemini-3.8-flash-low --effort low --dangerously-skip-permissions" \ + || fail "could not type the agy launch line" +"$REAL_TMUX" -L "$SOCKET" send-keys -t "$TARGET" Enter \ + || fail "could not submit the agy launch line" + +# A fresh workspace stops on the folder-trust dialog. Answer the preselected +# safe choice once it renders. The answer appends the workspace to +# trustedWorkspaces in the throwaway HOME's copy of the agy settings store. +screen= +for _ in $(seq 1 150); do + screen=$(capture) + case "$screen" in + *"Do you trust the contents of this project?"*|*80235*|*80,235*) break ;; + esac + sleep 0.5 +done +case "$screen" in + *"Do you trust the contents of this project?"*) + "$REAL_TMUX" -L "$SOCKET" send-keys -t "$TARGET" Enter \ + || fail "could not answer the agy trust dialog" + ;; +esac + +# The initial turn executes and its reply lands; the busy footer must render +# while it is in flight so the portable matcher has live text to prove. +# Trivial turns were observed taking one to two minutes (cold start plus model +# latency), so these windows are generous; the guard is opt-in. +busy_live= +for _ in $(seq 1 240); do + screen=$(capture) + if printf '%s' "$screen" | fm_busy_agy_tail_busy; then busy_live=1; break; fi + case "$screen" in *80235*|*80,235*) break ;; esac + sleep 1 +done +[ -n "$busy_live" ] || fail "fm_busy_agy_tail_busy never matched the real agy turn in flight" +pass "the real agy busy footer matches fm_busy_agy_tail_busy in flight" + +for _ in $(seq 1 480); do + screen=$(capture) + case "$screen" in *80235*|*80,235*) break ;; esac + sleep 0.5 +done +reply=$(capture) +case "$reply" in + *80235*|*80,235*) pass "the real agy worker processed its launch prompt" ;; + *) fail "the real agy worker never answered its launch prompt" ;; +esac +# The reply can render while the turn is still finishing: the busy footer stays +# pinned until the idle composer replaces it, so wait for the settled idle row +# before asserting what the settled pane must not match. The wait itself +# refreshes $screen: the reply-wait loop above can legitimately break on a +# frame that still carries the pinned busy footer, and asserting on that stale +# frame would fail every run whose reply lands mid-turn. +idle_settled= +for _ in $(seq 1 120); do + screen=$(capture) + case "$screen" in *"? for shortcuts"*) idle_settled=1; break ;; esac + sleep 0.5 +done +[ -n "$idle_settled" ] || fail "the agy composer never settled to its idle footer after the reply" +# Scope to the visible tail the same way the owners do: mid-turn busy rows stay +# in scrollback after the turn settles and must not count as still busy. +printf '%s' "$screen" | grep -v '^[[:space:]]*$' | tail -12 | fm_busy_lines_match agy \ + && fail "harness=agy matched its own idle footer as busy" || true +printf '%s' "$screen" | fm_busy_agy_tail_busy \ + && fail "the settled agy footer still matches the busy signature" || true + +# The dialog can outlive the turn it gated, so a still-rendered dialog must be +# dismissed before steering anything: typed text would land in it instead of +# the composer. +if case "$(capture)" in *"Do you trust the contents of this project?"*) true ;; *) false ;; esac; then + "$REAL_TMUX" -L "$SOCKET" send-keys -t "$TARGET" Enter \ + || fail "could not dismiss the residual agy trust dialog" + idle= + for _ in $(seq 1 120); do + case "$(capture)" in *"? for shortcuts"*) idle=1; break ;; esac + sleep 0.5 + done + [ -n "$idle" ] || fail "the agy composer never went idle after the trust answer" +fi + +# Interrupt a genuinely long turn: poll until busy is observed, then send +# exactly one Escape and wait only for the Interrupted row it prints; a busy +# footer that merely disappears is not cancellation and no further Escape is +# sent, so a turn that survives one Escape fails this guard. +"$REAL_TMUX" -L "$SOCKET" send-keys -t "$TARGET" -l \ + "Write a 1500-word essay on the history of glass" \ + || fail "could not type the long agy prompt" +"$REAL_TMUX" -L "$SOCKET" send-keys -t "$TARGET" Enter \ + || fail "could not submit the long agy prompt" +for _ in $(seq 1 100); do + screen=$(capture) + printf '%s' "$screen" | fm_busy_agy_tail_busy && break + sleep 0.5 +done +printf '%s' "$screen" | fm_busy_agy_tail_busy \ + || fail "the long agy turn never showed its busy footer" +"$REAL_TMUX" -L "$SOCKET" send-keys -t "$TARGET" Escape \ + || fail "could not send Escape to the real agy turn" +cancelled= +for _ in $(seq 1 120); do + screen=$(capture) + case "$screen" in *Interrupted*) cancelled=1; break ;; esac + sleep 0.5 +done +[ -n "$cancelled" ] || fail "a single Escape never cancelled the real agy turn" +pass "a single Escape cancels the real agy turn" + +"$REAL_TMUX" -L "$SOCKET" send-keys -t "$TARGET" -l "/quit" \ + || fail "could not type the agy exit command" +"$REAL_TMUX" -L "$SOCKET" send-keys -t "$TARGET" Enter \ + || fail "could not submit the agy exit command" +gone= +for _ in $(seq 1 60); do + current=$("$REAL_TMUX" -L "$SOCKET" display-message -p -t "$TARGET" '#{pane_current_command}' 2>/dev/null || true) + case "$current" in *agy*) sleep 0.5 ;; *) gone=1; break ;; esac +done +[ -n "$gone" ] || fail "/quit never stopped the real agy process" +pass "/quit stops the real agy process" + +cleanup +trap - EXIT diff --git a/tests/fm-bootstrap.test.sh b/tests/fm-bootstrap.test.sh index 23eac5b67f5..afbb0db67e1 100755 --- a/tests/fm-bootstrap.test.sh +++ b/tests/fm-bootstrap.test.sh @@ -1135,6 +1135,10 @@ pi max effort is accepted^{"rules":[{"when":"deep coding","use":{"harness":"pi", pi-signed max effort is accepted^{"rules":[{"when":"signed coding","use":{"harness":"pi-signed","model":"openai-codex/gpt-5.6-sol","effort":"max"}}]}^empty^ muse shared efforts are accepted^{"rules":[{"when":"muse low","use":{"harness":"muse","effort":"low"}},{"when":"muse medium","use":{"harness":"muse","effort":"medium"}},{"when":"muse high","use":{"harness":"muse","effort":"high"}},{"when":"muse xhigh","use":{"harness":"muse","effort":"xhigh"}},{"when":"muse max","use":{"harness":"muse","effort":"max"}}]}^empty^ unsupported muse ultra effort is flagged^{"rules":[{"when":"muse ultra","use":{"harness":"muse","effort":"ultra"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: muse:ultra +agy model profile is accepted^{"rules":[{"when":"agy work","use":{"harness":"agy","model":"gemini-3.8-flash-high"}}]}^empty^ +agy low medium high efforts are accepted^{"rules":[{"when":"agy low","use":{"harness":"agy","effort":"low"}},{"when":"agy medium","use":{"harness":"agy","effort":"medium"}},{"when":"agy high","use":{"harness":"agy","effort":"high"}}]}^empty^ +unsupported agy xhigh effort is flagged^{"rules":[{"when":"agy xhigh","use":{"harness":"agy","effort":"xhigh"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: agy:xhigh +unsupported agy max effort is flagged^{"rules":[{"when":"agy max","use":{"harness":"agy","effort":"max"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: agy:max unsupported opencode effort is flagged^{"rules":[{"when":"opencode work","use":{"harness":"opencode","model":"anthropic/claude-sonnet-4-5","effort":"high"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: opencode:high kimi model profile is accepted^{"rules":[{"when":"kimi work","use":{"harness":"kimi","model":"kimi-code/k3"}}]}^empty^ unsupported kimi effort is flagged^{"rules":[{"when":"kimi work","use":{"harness":"kimi","model":"kimi-code/k3","effort":"high"}}]}^exact^CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: kimi:high diff --git a/tests/fm-send-agy-confirm.test.sh b/tests/fm-send-agy-confirm.test.sh new file mode 100755 index 00000000000..1a5a5240509 --- /dev/null +++ b/tests/fm-send-agy-confirm.test.sh @@ -0,0 +1,165 @@ +#!/usr/bin/env bash +# fm-send typed-plane submit-confirm budget for agy targets. +# +# A typed send to an explicit tmux agy endpoint is acknowledged only by the +# submit core's idle-to-busy transition poll: agy's bare `>` composer verdict +# is `unknown` (dead-shell rule), so the poll watching the pane's verified +# `esc to cancel` busy footer is the only proof a landed Enter can get. agy +# renders that footer ~1.5s after Enter for a short steer and ~4s for a +# multi-line brief (live-measured on agy 1.2.1), while the shared default +# confirm budget is 3 retries x 0.4s - so fm-send used to exit 1 "not +# submitted" for a message that landed and ran, inviting a duplicate resend. +# fm-send now gives agy typed targets a longer default budget (20 retries, +# ~8s at the default cadence); an explicit FM_SEND_RETRIES still wins and every +# other harness keeps the +# shared 3-retry default. These tests pin that behavior hermetically (stubbed +# tmux + sleep, no real agent): the fake tmux renders the busy footer only +# from the BUSY_AT-th plain pane capture, so the number of logged 0.4s waits +# stands in for wall-clock latency and each case is deterministic: +# 1. agy target, busy footer at the 5th poll (short-steer latency): the send +# exits 0 (confirmed idle-to-busy), and the sleep log shows the poll +# reaching that read. +# 2. agy target, busy footer only at the 15th poll (the live-measured long +# brief latency): the default budget still reaches it and exits 0. +# 3. agy target with an explicit FM_SEND_RETRIES=3: the operator knob wins +# and the send keeps the loud exit-1 verdict=unknown refusal. +# 4. agy target whose busy footer never renders: no confirmation is +# fabricated - exit 1 verdict=unknown. +# 5. claude target, same late-busy pane: the shared 3-retry default is +# untouched, so the send still exits 1 verdict=unknown. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +SEND="$ROOT/bin/fm-send.sh" + +TMP_ROOT=$(fm_test_tmproot fm-send-agy-confirm) + +# A fake tmux that models agy's late busy render, plus a fake sleep that +# records every requested duration (one per line) into FM_SLEEP_LOG instead of +# sleeping. The styled capture (-e) always shows agy's idle bare-`>` composer +# (verdict `unknown`); the plain capture - the one fm_pane_busy_state polls - +# shows the idle screen until its BUSY_AT-th call and the verified `esc to +# cancel` busy row from then on. The BUSY_AT threshold is read from the +# per-case dir so cases are independent. +make_stubs() { # <dir> <busy-at> -> echoes fakebin dir + local dir=$1 busy_at=$2 fb="$1/fakebin" + mkdir -p "$fb" + cat > "$fb/tmux" <<SH +#!/usr/bin/env bash +set -u +cnt_file="$dir/plain.count" +case "\${1:-}" in + send-keys) exit 0 ;; + display-message) + for a in "\$@"; do + case "\$a" in + *cursor_y*) printf '0\n'; exit 0 ;; + *pane_tty*) printf '\n'; exit 0 ;; + esac + done + printf 'fakepane\n'; exit 0 ;; + capture-pane) + styled=0 + for a in "\$@"; do [ "\$a" = -e ] && styled=1; done + if [ "\$styled" = 1 ]; then + printf '> \n? for shortcuts\n' + exit 0 + fi + n=\$(( \$(cat "\$cnt_file" 2>/dev/null || echo 0) + 1 )) + printf '%s' "\$n" > "\$cnt_file" + if [ "\$n" -ge $busy_at ]; then + printf '> \n? for shortcuts\n ⏺ 5s · esc to cancel · gemini-3.8-flash-low\n' + else + printf '> \n? for shortcuts\n' + fi + exit 0 ;; + list-windows) printf 'win\n'; exit 0 ;; +esac +exit 0 +SH + chmod +x "$fb/tmux" + cat > "$fb/sleep" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "${1:-}" >> "$FM_SLEEP_LOG" +exit 0 +SH + chmod +x "$fb/sleep" + printf '%s\n' "$fb" +} + +# run_send <harness> <busy-at> <extra-env-assignments...>: build a fresh home +# whose recorded task targets sess:win on the tmux backend with <harness> meta, +# then run the real fm-send typed plane against the stubs. FM_SEND_SETTLE=0 +# strips the post-submit pause so the sleep log holds only the popup settle +# plus the 0.4 submit waits, keeping the poll arithmetic visible. FM_ROOT_OVERRIDE +# points at the case dir so fm-guard's tangle check stays silent. Emits +# "rc <exit>" and leaves the send's stderr in $dir/err and the sleep log in +# $dir/sleep.log for the caller to assert on. +run_send() { # <harness> <busy-at> [env=val ...] + local harness=$1 busy_at=$2 dir fb log + shift 2 + dir="$TMP_ROOT/case-$RANDOM-$RANDOM"; mkdir -p "$dir/state" + fb=$(make_stubs "$dir" "$busy_at") + log="$dir/sleep.log"; : > "$log" + fm_write_meta "$dir/state/agyw.meta" "window=sess:win" "harness=$harness" + ( + export FM_GATE_REFUSE_BYPASS=1 FM_SEND_SETTLE=0 + export PATH="$fb:$PATH" FM_ROOT_OVERRIDE="$dir" FM_HOME="$dir" FM_SLEEP_LOG="$log" + for a in "$@"; do eval "export $a"; done + "$SEND" sess:win 'Append steer1 line to notes.md' 2>"$dir/err" + printf 'rc %s\n' "$?" + ) +} + +# agy, default budget, busy footer renders at the 5th poll (the 6th plain +# capture): the raised default (20 retries) must reach that read and exit 0. +# Under the old shared default (3 retries) this exact shape exited 1 +# "verdict=unknown" - the regression this suite pins. +out=$(run_send agy 6) +expect_code 0 "$(printf '%s' "$out" | sed -n 's/^rc //p')" \ + "agy typed send with late busy footer confirms idle-to-busy and exits 0" +grep -q 'not submitted' "$TMP_ROOT"/*/err 2>/dev/null && \ + fail "agy typed send: refusal text present despite confirmed submit" +pass "agy typed send: no not-submitted refusal on confirmed idle-to-busy" +case_dir=$(printf '%s\n' "$TMP_ROOT"/case-* | head -1) +settles=$(grep -cv '^0\.4$' "$case_dir/sleep.log" || true) +waits=$(grep -c '^0\.4$' "$case_dir/sleep.log" || true) +[ "$settles" = 1 ] || fail "agy typed send: expected exactly 1 non-wait sleep (popup settle), got $settles" +[ "$waits" = 5 ] || fail "agy typed send: expected the poll to reach the 5th busy read (5 x 0.4s: Enter wait + 4 poll waits), got $waits" +pass "agy typed send: sleep log shows the confirm poll running to the late busy render" + +# agy with an explicit FM_SEND_RETRIES=3: the operator knob wins over the agy +# default, the budget expires before the late footer, and the loud refusal +# boundary is preserved. +out=$(run_send agy 6 'FM_SEND_RETRIES=3') +expect_code 1 "$(printf '%s' "$out" | sed -n 's/^rc //p')" \ + "agy typed send honors an explicit FM_SEND_RETRIES=3" +grep -q 'verdict=unknown' "$TMP_ROOT"/*/err || fail "agy typed send FM_SEND_RETRIES=3: expected verdict=unknown refusal" +pass "agy typed send: explicit FM_SEND_RETRIES=3 keeps the exit-1 verdict=unknown refusal" + +# agy whose busy footer never renders: the raised budget must time out into +# the same loud refusal, never fabricate a confirmation. +out=$(run_send agy 999) +expect_code 1 "$(printf '%s' "$out" | sed -n 's/^rc //p')" \ + "agy typed send with no busy footer refuses exit 1" +grep -q 'verdict=unknown' "$TMP_ROOT"/*/err || fail "agy typed send never-busy: expected verdict=unknown refusal" +pass "agy typed send: never-rendering busy footer still refuses with verdict=unknown" + +# agy, busy footer renders only at the 15th poll: the live long-brief case. +# The default budget must still reach that read and exit 0; under the shared +# 3-retry default this shape refused for a message that landed. +out=$(run_send agy 16) +expect_code 0 "$(printf '%s' "$out" | sed -n 's/^rc //p')" \ + "agy typed send with long-brief late busy footer confirms and exits 0" +pass "agy typed send: long-brief render (15th poll) still confirms idle-to-busy" + +# claude on the identical late-busy pane: the shared 3-retry default is +# untouched, so the same latency still refuses - the raised budget is +# agy-scoped, not a global slowdown. +out=$(run_send claude 6) +expect_code 1 "$(printf '%s' "$out" | sed -n 's/^rc //p')" \ + "claude typed send keeps the shared 3-retry default" +grep -q 'verdict=unknown' "$TMP_ROOT"/*/err || fail "claude typed send: expected verdict=unknown refusal" +pass "claude typed send: late busy footer still refuses (agy budget is agy-scoped)" From b518a256e89c4f3b4072670af5ac8ee5188672da Mon Sep 17 00:00:00 2001 From: NewAiCoder-bot <iamacodernow-bot@theinbtw.com> Date: Sat, 12 Sep 2026 22:28:03 -0400 Subject: [PATCH 24/31] feat(afk): add quiet supervision mode for a present captain (#4337) * feat(afk): add quiet supervision mode for a present captain Adds a first-class quiet supervision mode alongside /afk for kunchenguid/firstmate#2356: the same away-mode daemon, injection, busy/composer guards, classification policy, and reliability properties, but the captain staying present and chatting no longer exits it - only an explicit /quiet off does. state/.afk's first line now declares its mode (away, the default, or quiet); fm_afk_mode() in bin/fm-wake-lib.sh is the single reader, falling back to away for missing/empty/unreadable/unrecognized content (including the legacy bare-epoch-timestamp format written before mode existed) so nothing regresses. fm_afk_flag_write() preserves the on-disk mode on a bare refresh (no explicit mode given) rather than defaulting to away, which is what keeps the daemon's own redundant terminal-side re-write from silently resetting a captain's quiet mode back to away underneath them. New .agents/skills/quiet/SKILL.md is a thin wrapper cross-referencing /afk for every shared mechanism, per the one-owner rule. AGENTS.md gains the state/.afk table entry and section 8's exit-trigger line. bin/fm-supervision-instructions.sh, bin/fm-session-start.sh, and bin/fm-guard.sh's stale-watcher banner all become mode-aware so a quiet-mode captain is never misdirected to /afk in captain-facing text. Closes #2356 * no-mistakes(review): Fix AFK epoch parsing and quiet-mode digest wording for two-line flag * no-mistakes(document): Fix turnend-guard.md daemon-ownership contract for quiet mode --------- Co-authored-by: NewAiCoder <claude@theinbtw.com> Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com> --- .agents/skills/afk/SKILL.md | 1 + .../firstmate-coding-guidelines/SKILL.md | 2 +- .agents/skills/quiet/SKILL.md | 78 +++++++++++++ AGENTS.md | 17 +-- README.md | 1 + bin/fm-afk-launch.sh | 9 +- bin/fm-afk-return.sh | 3 + bin/fm-afk-start.sh | 22 +++- bin/fm-guard.sh | 1 + bin/fm-session-start.sh | 22 +++- bin/fm-supervision-instructions.sh | 27 ++++- bin/fm-wake-lib.sh | 22 ++++ docs/documentation-audiences.json | 4 + docs/turnend-guard.md | 12 +- tests/fm-afk-launch.test.sh | 110 ++++++++++++++++++ tests/fm-afk-return.test.sh | 18 +++ tests/fm-guard-stale-banner.test.sh | 17 +++ tests/fm-session-start.test.sh | 44 +++++++ tests/fm-supervision-instructions.test.sh | 26 +++++ 19 files changed, 409 insertions(+), 27 deletions(-) create mode 100644 .agents/skills/quiet/SKILL.md diff --git a/.agents/skills/afk/SKILL.md b/.agents/skills/afk/SKILL.md index 1de10bfddec..0046d63e020 100644 --- a/.agents/skills/afk/SKILL.md +++ b/.agents/skills/afk/SKILL.md @@ -77,6 +77,7 @@ No `/back` is needed. The first genuine message is the return signal: - Re-invoking `/afk` while already away -> stay away (refresh); this does **not** trigger an exit. Bias ambiguous cases toward exit: a present captain beats token savings, and a false exit is self-correcting (the captain re-runs `/afk`). +When the captain wants this same token-saving supervision while staying present and chatting - ordinary messages should NOT exit it - that is `/quiet` (kunchenguid/firstmate#2356), not `/afk`. ## Orthogonal to approval authority diff --git a/.agents/skills/firstmate-coding-guidelines/SKILL.md b/.agents/skills/firstmate-coding-guidelines/SKILL.md index 6a50e6f944f..0ed6d4b8f52 100644 --- a/.agents/skills/firstmate-coding-guidelines/SKILL.md +++ b/.agents/skills/firstmate-coding-guidelines/SKILL.md @@ -53,7 +53,7 @@ That is the trigger condition for loading the skill, plus any safety-critical fa Everything else - the procedure, the mechanism, the surrounding detail - moves out completely. Do not leave a partial restatement behind "just in case". A partial copy is exactly the duplication the one-owner rule forbids. -The model to copy is `AGENTS.md` section 8's "Away-mode stub": it keeps only the marker format, the ownership-transfer rule, and the exit condition inline, and points everything else at the `/afk` skill. +The model to copy is `AGENTS.md` section 8's "Away-mode and quiet-mode stub": it keeps only the marker format, the ownership-transfer rule, and the exit condition inline, and points everything else at the `/afk` and `/quiet` skills. ## Size discipline diff --git a/.agents/skills/quiet/SKILL.md b/.agents/skills/quiet/SKILL.md new file mode 100644 index 00000000000..1c57b6700fc --- /dev/null +++ b/.agents/skills/quiet/SKILL.md @@ -0,0 +1,78 @@ +--- +name: quiet +description: >- + Enter quiet supervision mode when the captain invokes /quiet or asks for quiet mode, quiet-while-present, or fewer routine wake turns while they stay in the session. + It sets the same durable away/quiet-mode flag as /afk, in `quiet` mode, so the sub-supervisor daemon self-handles routine wakes and escalates captain-relevant events exactly as away mode does, but ordinary captain chat does NOT exit it - only an explicit `/quiet off` does. +user-invocable: true +metadata: + internal: true +--- + +# quiet + +Quiet supervision mode (kunchenguid/firstmate#2356): the same token-saving +daemon tradeoff as `/afk`, made explicit for a captain who is staying, +watching the session, and does not want to exit the mode just by chatting. + +This skill is a thin wrapper. +Every mechanism below - the daemon, its injection, its busy/composer guards, +its classification policy, its reliability properties - is owned once by the +`afk` skill and is IDENTICAL in quiet mode; nothing here restates it. +The only things quiet mode changes are which mode the flag declares and what +exits it. + +## What it does + +1. **Enter the lifecycle through `bin/fm-afk-launch.sh`, exactly as `/afk` + does, with `FM_AFK_MODE=quiet` set first.** + Follow the `afk` skill's "What it does" steps 1-3 verbatim (terminal- + backed vs harness-native entry, daemon-already-running refresh, never + arming a separate `fm-watch.sh`) with one addition: export + `FM_AFK_MODE=quiet` in the shell that invokes `bin/fm-afk-launch.sh start` + (or `start-native`), so `state/.afk`'s first line reads `quiet` instead of + `away`. + Leaving `FM_AFK_MODE` unset on a bare refresh of an already-running quiet + daemon is also correct and does nothing wrong: `fm_afk_flag_write` + preserves the on-disk mode when no explicit mode is given, so a plain + `/afk`-shaped refresh call never resets quiet back to away underneath the + captain. + +2. **Acknowledge** in `AGENTS.md` section 9 language: "Captain, quiet mode is + active; I will batch routine updates and surface only decisions, failures, + credentials, or review-ready work - ordinary chat will not exit this, say + `/quiet off` when you want normal per-wake responses back." + +## How to exit quiet mode + +Unlike `/afk`, ordinary chat is never the exit signal - that is the entire +point of this mode (AGENTS.md section 8's away-mode stub, quiet branch). + +- Only an explicit `/quiet off` (or the captain plainly asking to leave quiet + mode / resume normal supervision) exits it: run `bin/fm-afk-return.sh` + unchanged, exactly the procedure `/afk`'s "How to exit afk" section + documents for its own return path (correct-ordered daemon shutdown, + durable wake presentation and acknowledgement, escalation/wedge evidence, + and the return-catch-up gate). + That script does not read or care about the flag's mode, so it needs no + quiet-specific variant. +- A marked daemon escalation, or a message beginning `/quiet` while already + in quiet mode (refresh, not exit) -> stay in quiet mode and process it, the + same two carve-outs `/afk` documents for away mode. +- Every other message while in quiet mode is simply answered as ordinary + work; the flag and daemon are left untouched. + +## Orthogonal to approval authority + +Identical to `/afk`: quiet mode changes how aggressively firstmate surfaces +things, never who approves what. +A PR ready for merge keeps the merge authority from `AGENTS.md` section 7, and +a needs-decision finding keeps the `ask-user-authority` policy. + +## Must not hide a decision or a failure + +Per the issue's own author triage: quiet mode is presentation only. +Progress, retries, and internal mechanics stay below deck exactly as in away +mode, but review-ready work, findings, decisions, failures, and credentials +escalate every time, through the same classification policy `/afk` owns. +Quiet mode is opt-in and never the unconsented default; only an explicit +`/quiet` invocation enters it. diff --git a/AGENTS.md b/AGENTS.md index 7d58297bb07..f4940d9915e 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -142,7 +142,7 @@ state/ runtime records and signals; gitignored .status-presentation-cursor .status-presentation-lock fleet-wide per-task status identity plus independent annotation and outcome-backstop byte offsets, with a serialization lock preventing already-presented lines from replaying while preserving delayed signal annotations; owned by fm-classify-lib.sh, with each task's row retired by teardown .afk-contract the away-posture record: the captain's verbatim away words, expected return, reach profile, spend cap, and structured mandate clauses; written only by bin/fm-afk-contract.sh after the captain confirms the read-back, archived under afk-contracts/ at return; its presence IS the away posture in every harness; its sibling .afk-contract.lock serializes actions authorized by the live record (contract: bin/fm-afk-contract.sh) afk-contracts/ archived away-posture records: one final record per away window keyed by entry time, plus any superseded mandates from that window - .afk durable away-mode daemon flag on the harnesses that still launch the daemon (never on Pi); present = sub-supervisor may inject escalations (set by the daemon entry, cleared on user return) + .afk durable away/quiet-mode daemon flag on the harnesses that still launch the daemon (never on Pi); present = sub-supervisor may inject escalations, first line `away` (default, set by /afk, cleared on user return) or `quiet` (set by /quiet, cleared only on explicit /quiet off) per the single owner fm_afk_mode() in bin/fm-wake-lib.sh .watch.lock .wake-queue.lock watcher singleton and queue serialization locks .claude-autoarm.lock .claude-autoarm-epoch .claude-autoarm-failure-notified .claude-autoarm-failure-alarmed .turnend-claude-blocks .turnend-claude-blocks.lock Claude Stop auto-arm single-flight, epoch, failure-episode, attended-alarm, guard-budget, and budget-lock records; never touch .cursor-park-owner .cursor-park-owner.lock .turnend-cursor-blocks Cursor stop-hook owner record, publication and commit lock, and bounded repair-nag budget; never touch @@ -243,7 +243,7 @@ For an ordinary direct report whose endpoint is dead or metadata has no window, For a dead secondmate direct report, load `secondmate-provisioning` and reconcile only that secondmate, never its whole child tree from the main home. Each secondmate reconciles work already in its own home and then idles; recovery never authorizes it to invent work. -If away mode is present, load `/afk`; where its daemon runs, let the daemon own supervision rather than arming another cycle, and on Pi keep the ordinary supervision session, which runs in both postures. +If `state/.afk` is present, load `/afk` in away mode or `/quiet` in quiet mode (`bin/fm-wake-lib.sh`'s `fm_afk_mode`); where its daemon runs, let the daemon own supervision rather than arming another cycle, and on Pi keep the ordinary supervision session, which runs in both postures. Surface only captain-relevant decisions, review-ready PRs, failures, and credential needs; otherwise resume the emitted supervision protocol silently. A restart must be a non-event because durable state and live backend inventory, not conversation memory, are authoritative. @@ -446,19 +446,20 @@ Queued wakes must be presented before other action and acknowledged only after h The spawn assertion and generated ship brief must both enforce that project work starts in an isolated disposable worktree, never the primary checkout. Harness-aware turn-end guards are structural backstops, not permission to omit the live cycle. -### Away-mode stub +### Away-mode and quiet-mode stub Invoke the `/afk` skill when the captain says `/afk`, says they are going afk, `state/.afk-contract` or `state/.afk` exists, an incoming message starts with `FM_INJECT_MARK`, or any `state/.subsuper-*` marker is involved. -The skill owns the daemon procedure; these safety facts remain inline: +Invoke the `/quiet` skill instead when the captain says `/quiet` or asks for quiet mode, or `state/.afk` already exists in quiet mode (`fm_afk_mode` in `bin/fm-wake-lib.sh`). +Each skill owns its own daemon procedure, which is otherwise identical; these safety facts remain inline for both: - Every current daemon injection uses the `away-supervisor` kind from `bin/fm-operational-input.sh` after `FM_OPERATIONAL_PREFIX` (U+2063 INVISIBLE SEPARATOR followed by `FIRSTMATE_OP: `), while the `/afk` skill owns legacy bare-marker compatibility. - `state/.afk-contract` is the away posture, written only after the captain confirms the read-back of their away words; entry announces hold-for-return only, and the record's clauses are recorded, not executed, in this release. - While `state/.afk` exists, the daemon owns supervision; do not arm a separate watcher. The daemon is never launched on Pi, where the ordinary supervision session continues under the record. -- A marked message while away mode is active is internal escalation and does not exit away mode. -- A message beginning `/afk` refreshes away mode. -- Any other unmarked message means the captain returned; load `/afk`, run the return owner, and do not process that message as ordinary work until its durable catch-up gate clears. -- Away mode never expands approval authority for merges, ask-user findings, destructive actions, irreversible actions, or security-sensitive choices. +- A marked message while away or quiet mode is active is internal escalation and does not exit that mode. +- A message beginning `/afk` refreshes away mode; a message beginning `/quiet` refreshes quiet mode. +- Any other unmarked message means the captain returned in away mode (load `/afk`, run the return owner, and do not process that message as ordinary work until its durable catch-up gate clears), or, in quiet mode, is simply answered as ordinary work with the flag and daemon left untouched until an explicit `/quiet off`. +- Away and quiet mode never expand approval authority for merges, ask-user findings, destructive actions, irreversible actions, or security-sensitive choices. - Bias ambiguous input toward exit because a present captain takes precedence. ### Stuck-worker trigger diff --git a/README.md b/README.md index 40441e999bc..ec6c92e4e36 100644 --- a/README.md +++ b/README.md @@ -184,6 +184,7 @@ Claude and grok use the slash form shown here; codex uses the same names with `$ | Skill | What it does | | ------------------ | -------------------------------------------------------------------------------------------------------------------------------------------- | | `/afk` | Enter away-mode supervision: the sub-supervisor self-handles routine notifications in bash, escalates captain-relevant events and bounded declared-external-wait rechecks as batched digests, and actively alerts if delivery gets stuck while you step away | +| `/quiet` | Enter quiet supervision mode: the same token-saving sub-supervisor tradeoff as `/afk`, for a captain who is staying and chatting - ordinary messages do not exit it, only an explicit `/quiet off` does | | `/ahoy` | Recap visible session events since the prior real captain message plus visibly unanswered captain decisions, then guide the captain through any open decisions one at a time in agent-judged impact order; fall back to Bearings when invoked as the session's first real captain message | | `/bearings` | Generate a concise four-section chat digest from bounded fleet state, including registered remote-home ledgers; use `/bearings file` to also replace today's dated report in `data/`, and add `include PRs` for live GitHub enrichment | | `/updatefirstmate` | Fast-forward the running firstmate and its secondmates, then persist and restart every live mate successfully left on the target commit - including already-current homes - with an honest re-read nudge only when restart cannot be proven | diff --git a/bin/fm-afk-launch.sh b/bin/fm-afk-launch.sh index 2466356f5bd..a85e37a8a1e 100755 --- a/bin/fm-afk-launch.sh +++ b/bin/fm-afk-launch.sh @@ -73,6 +73,9 @@ # terminal (default bin/fm-afk-start.sh), so a topology test can run a harmless # placeholder instead of a real daemon. FM_SUPERVISOR_TARGET/FM_SUPERVISOR_BACKEND # override the captured captain pane/backend (an isolated lab pane in tests). +# FM_AFK_MODE (away|quiet, default away) declares which mode a `start` entry +# requests; leave it unset for a plain refresh of an already-running daemon +# so its current mode is preserved (bin/fm-afk-start.sh fm_afk_flag_write). set -u FM_AFK_LAUNCH_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -247,7 +250,11 @@ fm_afk_launch_record_write() { # <backend> <target> <extra> } fm_afk_launch_flag_write() { - fm_afk_flag_write "$FM_AFK_LAUNCH_STATE" + # FM_AFK_MODE is the ONE place a caller declares which mode this entry + # requests (away, the unset default, or quiet - kunchenguid/firstmate#2356); + # fm_afk_flag_write itself preserves the on-disk mode when it is unset, so + # a plain /afk refresh of an already-quiet daemon never resets it. + fm_afk_flag_write "$FM_AFK_LAUNCH_STATE" "${FM_AFK_MODE:-}" } # Read the recorded terminal into FM_AFK_REC_BACKEND/FM_AFK_REC_TARGET. The third diff --git a/bin/fm-afk-return.sh b/bin/fm-afk-return.sh index 7773c7d325c..51f6d7ff8ff 100755 --- a/bin/fm-afk-return.sh +++ b/bin/fm-afk-return.sh @@ -140,6 +140,9 @@ window_start_epoch() { fi if [ -z "$epoch" ] && [ -f "$STATE/.afk" ]; then flag=$(head -1 "$STATE/.afk" 2>/dev/null || true) + case "$flag" in + ''|*[!0-9]*) flag=$(sed -n '2p' "$STATE/.afk" 2>/dev/null || true) ;; + esac case "$flag" in ''|*[!0-9]*) ;; *) epoch=$flag ;; esac fi case "$epoch" in ''|*[!0-9]*) printf '' ;; *) printf '%s' "$epoch" ;; esac diff --git a/bin/fm-afk-start.sh b/bin/fm-afk-start.sh index e86c54f170a..e268d2d61e0 100755 --- a/bin/fm-afk-start.sh +++ b/bin/fm-afk-start.sh @@ -3,8 +3,8 @@ # foreground process when one is not already alive. # # Usage: fm-afk-start.sh -# Sets state/.afk unless FM_AFK_STATE_PREPARED=1, checks -# state/.supervise-daemon.lock, and: +# Sets state/.afk (mode preserved on refresh, see fm_afk_flag_write) unless +# FM_AFK_STATE_PREPARED=1, checks state/.supervise-daemon.lock, and: # - prints "afk: daemon already running pid=<pid>" then exits 0 when that # lock is held by a live daemon (a REFRESH: no stale-artifact clear); # - otherwise clears any prior away session's stale escalation artifacts @@ -110,12 +110,24 @@ daemon_lock_held_by_live_daemon() { daemon_pid_matches "$pid" "$owner" } -fm_afk_flag_write() { # <state-dir> - local state=$1 lock="$1/.cursor-park-owner.lock" pending attempt=0 status=1 +fm_afk_flag_write() { # <state-dir> [mode] + local state=$1 requested_mode=${2:-} lock="$1/.cursor-park-owner.lock" \ + pending attempt=0 status=1 mode mkdir -p "$state" || return 1 [ ! -d "$state/.afk" ] || return 1 + # An explicit mode is a caller's deliberate request (a fresh /afk or /quiet + # entry). Omitted means "just refresh" (an already-running daemon, or + # recovery re-entering generically) and PRESERVES whatever mode is already + # on disk via fm_afk_mode - which itself falls back to "away" when nothing + # is on disk yet, so a genuinely fresh unspecified entry still defaults + # away. This is what keeps a refresh from silently flipping a captain's + # quiet mode back to away underneath them (kunchenguid/firstmate#2356). + case "$requested_mode" in + away|quiet) mode=$requested_mode ;; + *) mode=$(fm_afk_mode "$state") ;; + esac pending=$(mktemp "$state/.afk.pending.XXXXXX") || return 1 - date '+%s' > "$pending" || { rm -f "$pending"; return 1; } + { printf '%s\n' "$mode"; date '+%s'; } > "$pending" || { rm -f "$pending"; return 1; } while [ "$attempt" -lt 50 ]; do attempt=$((attempt + 1)) if fm_lock_try_acquire "$lock"; then diff --git a/bin/fm-guard.sh b/bin/fm-guard.sh index a1f2f1cd488..ba9ee330465 100755 --- a/bin/fm-guard.sh +++ b/bin/fm-guard.sh @@ -221,6 +221,7 @@ if [ "$watcher_healthy" = false ]; then fix=$("$SCRIPT_DIR/fm-supervision-instructions.sh" \ --read-only "$READ_ONLY" \ --afk "$afk" \ + --afk-mode "$(fm_afk_mode "$STATE")" \ --x-mode "$x_mode" \ --queue-pending "$queue_arg" \ --repair-line 2>/dev/null || printf '%s\n' 'Repair missing watcher supervision according to the session-start operating block.') diff --git a/bin/fm-session-start.sh b/bin/fm-session-start.sh index 59b9044ac58..3a994b86030 100755 --- a/bin/fm-session-start.sh +++ b/bin/fm-session-start.sh @@ -749,6 +749,7 @@ fi stage supervision-instructions AFK_PRESENT=0 [ -e "$STATE/.afk" ] && AFK_PRESENT=1 +AFK_MODE=$(fm_afk_mode "$STATE") X_MODE_PRESENT=0 [ -f "$CONFIG/x-mode.env" ] && X_MODE_PRESENT=1 @@ -789,6 +790,7 @@ fi --harness "$PRIMARY_HARNESS" \ --read-only "$READ_ONLY" \ --afk "$AFK_PRESENT" \ + --afk-mode "$AFK_MODE" \ --x-mode "$X_MODE_PRESENT" # --- 5. read-once contract ------------------------------------------------- @@ -879,12 +881,20 @@ if [ -f "$STATE/.afk-contract" ]; then printf 'present - away posture recorded at %s (hold-for-return only; bin/fm-afk-contract.sh readback for the mandate)' \ "$("$SCRIPT_DIR/fm-afk-contract.sh" field entered 2>/dev/null || printf unknown)" if [ -e "$STATE/.afk" ]; then - printf '; the away daemon owns the watcher.\n' + if [ "$AFK_MODE" = quiet ]; then + printf '; the quiet daemon owns the watcher.\n' + else + printf '; the away daemon owns the watcher.\n' + fi else printf '; no daemon runs, the ordinary supervision session continues.\n' fi elif [ -e "$STATE/.afk" ]; then - printf 'present - away-mode supervision is active; the daemon owns the watcher (legacy flag with no posture record).\n' + if [ "$AFK_MODE" = quiet ]; then + printf 'present - quiet-mode supervision is active; the daemon owns the watcher, only an explicit /quiet off exits it (legacy flag with no posture record).\n' + else + printf 'present - away-mode supervision is active; the daemon owns the watcher (legacy flag with no posture record).\n' + fi else printf 'absent\n' fi @@ -948,6 +958,14 @@ This session did not acquire the fleet lock. Stay read-only: do not arm, drain, spawn, steer, merge, or repair fleet state from here. Only a session with verified fleet-lock ownership may perform mutable follow-up. +EOF +elif [ "$AFK_PRESENT" -eq 1 ] && [ "$AFK_MODE" = quiet ]; then + cat <<'EOF' +Quiet mode is active. Follow the supervision operating instructions block +above: load /quiet and ensure the daemon is running, because the daemon owns +watcher supervision. Ordinary captain chat does not exit it; only an +explicit /quiet off does. + EOF elif [ "$AFK_PRESENT" -eq 1 ]; then cat <<'EOF' diff --git a/bin/fm-supervision-instructions.sh b/bin/fm-supervision-instructions.sh index 94316cfb03c..d5de85133a7 100755 --- a/bin/fm-supervision-instructions.sh +++ b/bin/fm-supervision-instructions.sh @@ -13,16 +13,19 @@ DOC_DIR="$REPO_ROOT/docs/supervision-protocols" HARNESS= READ_ONLY=0 AFK=0 +AFK_MODE=away X_MODE=0 REPAIR_LINE=0 QUEUE_PENDING=0 usage() { cat <<'EOF' -Usage: fm-supervision-instructions.sh [--harness <name>] [--read-only 0|1] [--afk 0|1] [--x-mode 0|1] [--repair-line] [--queue-pending 0|1] +Usage: fm-supervision-instructions.sh [--harness <name>] [--read-only 0|1] [--afk 0|1] [--afk-mode away|quiet] [--x-mode 0|1] [--repair-line] [--queue-pending 0|1] Print the current primary harness's supervision operating instructions. With --repair-line, print one concise repair instruction for guard and hook messages. +--afk-mode only matters when --afk 1 (present); it selects the away-mode vs +quiet-mode (kunchenguid/firstmate#2356) wording, and defaults to away. EOF } @@ -50,6 +53,14 @@ while [ "$#" -gt 0 ]; do AFK=$(bool_value "$2") shift 2 ;; + --afk-mode) + [ "$#" -gt 1 ] || { echo "error: --afk-mode requires away or quiet" >&2; exit 2; } + case "$2" in + away|quiet) AFK_MODE=$2 ;; + *) AFK_MODE=away ;; + esac + shift 2 + ;; --x-mode) [ "$#" -gt 1 ] || { echo "error: --x-mode requires 0 or 1" >&2; exit 2; } X_MODE=$(bool_value "$2") @@ -125,7 +136,11 @@ repair_line() { return 0 fi if [ "$AFK" -eq 1 ]; then - printf '%s\n' 'Away mode owns watcher supervision; load /afk and ensure the daemon is running instead of starting normal supervision directly.' + if [ "$AFK_MODE" = quiet ]; then + printf '%s\n' 'Quiet mode owns watcher supervision; load /quiet and ensure the daemon is running instead of starting normal supervision directly.' + else + printf '%s\n' 'Away mode owns watcher supervision; load /afk and ensure the daemon is running instead of starting normal supervision directly.' + fi return 0 fi @@ -210,9 +225,13 @@ else printf '%s\n' '- Lock: held by this session; this session owns normal supervision unless away mode says otherwise.' fi if [ "$AFK" -eq 1 ]; then - printf '%s\n' '- Away mode: active; load /afk and keep normal harness supervision paused while the daemon owns the watcher.' + if [ "$AFK_MODE" = quiet ]; then + printf '%s\n' '- Quiet mode: active; load /quiet and keep normal harness supervision paused while the daemon owns the watcher. Ordinary captain chat does NOT exit it - only an explicit /quiet off does.' + else + printf '%s\n' '- Away mode: active; load /afk and keep normal harness supervision paused while the daemon owns the watcher.' + fi else - printf '%s\n' '- Away mode: inactive.' + printf '%s\n' '- Away/quiet mode: inactive.' fi if [ "$X_MODE" -eq 1 ]; then printf '%s%s%s\n' '- X mode: active; source ' "$x_mode_env" ' before launching any watcher process so the 30s cadence is inherited.' diff --git a/bin/fm-wake-lib.sh b/bin/fm-wake-lib.sh index 54d3770ace3..e3582cbc4da 100755 --- a/bin/fm-wake-lib.sh +++ b/bin/fm-wake-lib.sh @@ -327,6 +327,28 @@ fm_afk_daemon_owns_supervision() { [ "$current" = "$recorded" ] } +# fm_afk_mode <state> +# The single owner of reading state/.afk's declared mode. Always prints +# exactly one of "away" or "quiet" and always succeeds - every caller gets a +# definitive answer, never an error to handle. Presence/liveness stays owned +# by fm_afk_daemon_owns_supervision and the raw `-e "$state/.afk"` checks +# throughout the tree; this is the mode of an ALREADY-present flag. +# "away" (today's return-on-any-unmarked-message behavior) is the safe +# default: missing, empty, unreadable, or unrecognized content, and the +# legacy bare-epoch-timestamp content written before mode existed, all read +# as "away". Only an exact first-line "quiet" ever reads as "quiet" - +# kunchenguid/firstmate#2356's standing captain-present quiet mode, entered +# only through /quiet and exited only through an explicit /quiet off +# (AGENTS.md section 8's away-mode stub). +fm_afk_mode() { + local state=$1 mode + mode=$(head -n 1 "$state/.afk" 2>/dev/null) || { printf '%s\n' away; return 0; } + case "$mode" in + quiet) printf '%s\n' quiet ;; + *) printf '%s\n' away ;; + esac +} + # fm_watcher_supervision_verdict <state> <watch-path> [grace] [home] [root] # Model-aware "is supervision healthy right now" verdict for the pull warning # guard (bin/fm-guard.sh), NOT the arm layer or the turn-end guard. Sets: diff --git a/docs/documentation-audiences.json b/docs/documentation-audiences.json index 1e66c6dbc9c..618caa941ec 100644 --- a/docs/documentation-audiences.json +++ b/docs/documentation-audiences.json @@ -232,6 +232,10 @@ "path": ".agents/skills/project-management/SKILL.md", "audience": "agent-runtime" }, + { + "path": ".agents/skills/quiet/SKILL.md", + "audience": "agent-runtime" + }, { "path": ".agents/skills/quota-array-dispatch/SKILL.md", "audience": "agent-runtime" diff --git a/docs/turnend-guard.md b/docs/turnend-guard.md index 579225d2970..1f0788c8ea8 100644 --- a/docs/turnend-guard.md +++ b/docs/turnend-guard.md @@ -15,7 +15,7 @@ Do not infer this guard's scope, loop safety, or compatibility tradeoffs for tho The turn-end guard closes the remaining gap at the primary's own turn boundary. When work, a process-event source, a registered custom check, or Relay polling needs supervision at that boundary and no identity-matched watcher has a fresh beacon, the harness integration must either block the turn end or force one bounded follow-up that uses the recovery instruction from the emitted session-start protocol. The mid-turn pull warning uses the model-aware supervision verdict described below, while the turn-end guard keeps the PID-strict watcher predicate. -Away mode is the one place the turn-end guard accepts a different supervisor: while `state/.afk` exists the away-mode daemon owns supervision, so a live identity-matched daemon with a fresh beacon satisfies that boundary in place of a watcher process holding the lock. +Away and quiet mode are the one place the turn-end guard accepts a different supervisor: while `state/.afk` exists, in either mode (`bin/fm-wake-lib.sh`'s `fm_afk_mode`), the daemon owns supervision, so a live identity-matched daemon with a fresh beacon satisfies that boundary in place of a watcher process holding the lock. The guard remains a backstop; [`watcher-continuity.md`](watcher-continuity.md) owns normal continuity. ## Guard predicates @@ -47,13 +47,13 @@ Without that proof an unheld lock alarms exactly as it did before, so an unloade Under every persistent-watcher harness a live identity-matched watcher with a fresh beacon is still required, so the pull guard keeps the same strict semantics there. Its banner names the true failing condition, either a missing live watcher process or a genuinely stale beacon with its real age, and keys the once-per-episode dedup on that condition rather than the beacon mtime. -While `state/.afk` exists the away-mode daemon (`bin/fm-supervise-daemon.sh`) owns supervision and runs the watcher one-shot: the watcher exits on every wake and the daemon starts its replacement, so a turn boundary regularly lands in a hand-off where no watcher process holds the lock and nothing is wrong. -The turn-end guard therefore accepts `fm_afk_daemon_owns_supervision` from `bin/fm-wake-lib.sh` as proof of supervision on that path: away mode must be active, and this home's `state/.supervise-daemon.lock` must name a live pid whose current process identity still matches the identity the daemon recorded for itself. +While `state/.afk` exists the daemon (`bin/fm-supervise-daemon.sh`) owns supervision and runs the watcher one-shot, in either away or quiet mode: the watcher exits on every wake and the daemon starts its replacement, so a turn boundary regularly lands in a hand-off where no watcher process holds the lock and nothing is wrong. +The turn-end guard therefore accepts `fm_afk_daemon_owns_supervision` from `bin/fm-wake-lib.sh` as proof of supervision on that path: `state/.afk` must exist (the predicate does not distinguish away from quiet mode), and this home's `state/.supervise-daemon.lock` must name a live pid whose current process identity still matches the identity the daemon recorded for itself. That is the same identity discipline the watcher lock uses, so a recycled pid, a lock left behind by a killed daemon, and a daemon that never recorded its identity all fail it. -A daemon that cannot record its own identity at startup logs a warning and keeps running, because a supervisor must not refuse to run over an unreadable `ps`; that warning is what names the cause when the guard then keeps blocking away-mode turn boundaries for the rest of that daemon's life. +A daemon that cannot record its own identity at startup logs a warning and keeps running, because a supervisor must not refuse to run over an unreadable `ps`; that warning is what names the cause when the guard then keeps blocking away/quiet-mode turn boundaries for the rest of that daemon's life. The proof covers ownership only, never freshness: the guard still requires a fresh beacon, so a daemon that stops restarting its watcher still blocks once the beacon passes grace, and a home with no daemon and no watcher blocks exactly as it did before. That beacon check uses the poll-derived grace described below rather than the flat `FM_GUARD_GRACE` default, because the daemon starts a fresh one-shot watcher only after it finishes handling the previous wake, and that handling can legitimately outrun a fixed 300-second window under load (a slow registered check, a busy supervisor pane) with the daemon perfectly healthy throughout. -With away mode off the daemon lock proves nothing and the strict watcher predicate is unchanged. +With `state/.afk` absent the daemon lock proves nothing and the strict watcher predicate is unchanged. `FM_STATE_OVERRIDE` wins over `FM_HOME/state`, and `FM_HOME` wins over repository-root `state/`. `FM_GUARD_GRACE` controls beacon freshness and defaults to 300 seconds. @@ -66,7 +66,7 @@ A fixed 300-second grace default stops correctly bounding staleness once a home' That hook and `bin/fm-watch.sh`'s own pre-acquisition staleness check (the "lock held by live pid but heartbeat is stale" refusal) both derive their default grace from the configured poll instead of a bare constant: `max(300, FM_POLL + 60)`, so the default never drops below the historical 300-second floor for the common short-poll case but grows with the poll cadence once that cadence would otherwise outrun it. `fm_poll_derived_grace` in `bin/fm-wake-lib.sh` is the single owner of that formula. The auto-arm hook additionally exports its resolved `FM_GUARD_GRACE` when it forks `bin/fm-watch-arm.sh`, so the arm wrapper and the watcher it may start judge staleness with the exact same value the hook just judged it with, whether that value came from an operator override or the poll-derived default. -`bin/fm-turnend-guard.sh`'s away-mode branch (`fm_afk_daemon_owns_supervision`, above) also derives its beacon grace from `fm_poll_derived_grace` rather than falling back to the bare 300-second default, for the same reason: the daemon's watcher-restart cadence there is not a fixed poll loop, so a flat grace misreads a daemon that is genuinely still cycling as down. +`bin/fm-turnend-guard.sh`'s daemon-ownership branch (`fm_afk_daemon_owns_supervision`, above, covering both away and quiet mode) also derives its beacon grace from `fm_poll_derived_grace` rather than falling back to the bare 300-second default, for the same reason: the daemon's watcher-restart cadence there is not a fixed poll loop, so a flat grace misreads a daemon that is genuinely still cycling as down. Every other direct `FM_GUARD_GRACE` reader (`bin/fm-guard.sh`, the strict-watcher checks in `bin/fm-turnend-guard.sh` and its harness-specific wrappers, `bin/fm-wake-lib.sh`) still falls back to the bare 300-second default unless `FM_GUARD_GRACE` is set explicitly in the environment. ## Harness integrations diff --git a/tests/fm-afk-launch.test.sh b/tests/fm-afk-launch.test.sh index a660e227d57..a8ca8e71033 100755 --- a/tests/fm-afk-launch.test.sh +++ b/tests/fm-afk-launch.test.sh @@ -280,6 +280,112 @@ unit_fresh_vs_refresh() { rm -rf "$st" } +# --------------------------------------------------------------------------- +# UNIT 2a: away/quiet mode plumbing (kunchenguid/firstmate#2356). fm_afk_mode +# is the single owner of reading the mode; these pin its write side +# (fm_afk_launch_flag_write / fm_afk_flag_write) against the exact double- +# write risk a live entry hits - the launcher writes the flag, then the +# terminal-side fm-afk-start.sh entry re-writes it a second time on every +# real (non-native) entry, per UNIT 2 above. +# --------------------------------------------------------------------------- +read_mode() { # <state-dir> + bash -c '. "$1"; fm_afk_mode "$2"' _ "$ROOT/bin/fm-wake-lib.sh" "$1" +} + +unit_mode_explicit_write() { + local st out + st=$(mktemp -d "${TMPDIR:-/tmp}/fm-afk-mode-explicit.XXXXXX") + mkdir -p "$st/state" + FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" FM_AFK_MODE=quiet \ + bash -c '. "$1"; fm_afk_launch_flag_write' _ "$LAUNCH" + out=$(read_mode "$st/state") + if [ "$out" = quiet ]; then + pass "mode: a fresh entry with FM_AFK_MODE=quiet writes quiet" + else + fail "mode: explicit FM_AFK_MODE=quiet fresh entry wrote '$out' instead of quiet" + fi + rm -rf "$st" +} + +unit_mode_fresh_defaults_away() { + local st out + st=$(mktemp -d "${TMPDIR:-/tmp}/fm-afk-mode-default.XXXXXX") + mkdir -p "$st/state" + FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" \ + bash -c '. "$1"; fm_afk_launch_flag_write' _ "$LAUNCH" + out=$(read_mode "$st/state") + if [ "$out" = away ]; then + pass "mode: a fresh entry with FM_AFK_MODE unset defaults to away" + else + fail "mode: fresh unset-mode entry wrote '$out' instead of away" + fi + rm -rf "$st" +} + +unit_mode_refresh_preserves_quiet() { + local st sleep_pid lock out + st=$(mktemp -d "${TMPDIR:-/tmp}/fm-afk-mode-preserve.XXXXXX") + mkdir -p "$st/state" + printf 'quiet\n%s\n' "$(date '+%s')" > "$st/state/.afk" + sleep 600 & + sleep_pid=$! + lock="$st/state/.supervise-daemon.lock" + mkdir -p "$lock" + printf '%s' "$sleep_pid" > "$lock/pid" + ( . "$ROOT/bin/fm-wake-lib.sh"; fm_pid_identity "$sleep_pid" > "$lock/pid-identity" 2>/dev/null ) || true + # The exact real-entry shape: a bare direct re-write with no explicit mode, + # simulating the terminal-side fm-afk-start.sh redundant write that would + # silently clobber quiet back to away if it were not preserve-on-refresh. + FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" "$START" >/dev/null 2>&1 + out=$(read_mode "$st/state") + if [ "$out" = quiet ]; then + pass "mode: a bare refresh (FM_AFK_MODE unset) of an already-running quiet daemon preserves quiet, never resets to away" + else + fail "mode: refresh incorrectly changed quiet mode to '$out'" + fi + kill "$sleep_pid" 2>/dev/null || true + wait "$sleep_pid" 2>/dev/null || true + rm -rf "$st" +} + +unit_mode_garbage_and_legacy_content_reads_away() { + local st out + st=$(mktemp -d "${TMPDIR:-/tmp}/fm-afk-mode-garbage.XXXXXX") + mkdir -p "$st/state" + + : > "$st/state/.afk" + out=$(read_mode "$st/state") + if [ "$out" = away ]; then + pass "mode: an empty (legacy pre-mode) flag reads as away" + else + fail "mode: empty flag read as '$out' instead of away" + fi + + date '+%s' > "$st/state/.afk" + out=$(read_mode "$st/state") + if [ "$out" = away ]; then + pass "mode: a bare-epoch-timestamp (legacy pre-mode) flag reads as away" + else + fail "mode: legacy timestamp flag read as '$out' instead of away" + fi + + printf 'nonsense-mode\n' > "$st/state/.afk" + out=$(read_mode "$st/state") + if [ "$out" = away ]; then + pass "mode: unrecognized content falls back to away" + else + fail "mode: unrecognized content read as '$out' instead of away" + fi + + out=$(read_mode "$st/state/missing") + if [ "$out" = away ]; then + pass "mode: a missing flag reads as away" + else + fail "mode: missing flag read as '$out' instead of away" + fi + rm -rf "$st" +} + # --------------------------------------------------------------------------- # UNIT 3: exit ordering - fm_afk_launch_stop SIGTERMs the daemon WHILE .afk is # still present (so its flush is not a no-op), and clears .afk last. @@ -1087,6 +1193,10 @@ unit_failed_daemon_launch_preserves_confirmed_record unit_stop_archives_the_record_last unit_relative_paths_are_absolute_before_daemon_launch unit_fresh_vs_refresh +unit_mode_explicit_write +unit_mode_fresh_defaults_away +unit_mode_refresh_preserves_quiet +unit_mode_garbage_and_legacy_content_reads_away unit_stop_ordering unit_stop_rejects_reused_pid unit_failed_start_rolls_back_state diff --git a/tests/fm-afk-return.test.sh b/tests/fm-afk-return.test.sh index 4b76e7f9947..d655a49946d 100755 --- a/tests/fm-afk-return.test.sh +++ b/tests/fm-afk-return.test.sh @@ -306,6 +306,23 @@ test_away_reentry_refuses_pending_return_gate() { pass "away-mode re-entry fails closed while the prior return catch-up is pending" } +test_return_is_mode_agnostic_for_quiet_mode() { + # kunchenguid/firstmate#2356's /quiet off calls this exact script, unchanged + # - it must behave identically whether state/.afk declares "away" or + # "quiet", since return_guard/return_reconcile only ever test presence. + local dir out + dir="$TMP_ROOT/quiet-mode-return" + install_runner "$dir" + printf 'quiet\n%s\n' "$(date +%s)" > "$dir/home/state/.afk" + : > "$dir/home/state/.fake-drain" + + out=$(run_return "$dir" begin) || fail "return did not succeed cleanly against a quiet-mode flag: $out" + assert_contains "$out" 'catch-up clear' "quiet-mode return did not announce ordinary work may proceed" + [ ! -e "$dir/home/state/.afk" ] || fail "quiet-mode return left the mode flag behind" + [ "$(wc -l < "$dir/home/stop.log" | tr -d ' ')" -eq 1 ] || fail "quiet-mode return did not stop the daemon exactly once" + pass "/quiet off's return path behaves identically for a quiet-content flag as for a legacy away-content one" +} + test_check_retries_recorded_terminal_teardown() { local dir gate out rc dir="$TMP_ROOT/terminal-teardown" @@ -754,6 +771,7 @@ test_explicit_reclassification_requires_durable_reason test_captain_decision_does_not_masquerade_as_firstmate_blocker test_evidence_publication_failure_preserves_wake_for_redrain test_away_reentry_refuses_pending_return_gate +test_return_is_mode_agnostic_for_quiet_mode test_check_retries_recorded_terminal_teardown test_unreadable_superseded_archive_keeps_return_gated test_missing_final_archive_keeps_retained_contract_gated diff --git a/tests/fm-guard-stale-banner.test.sh b/tests/fm-guard-stale-banner.test.sh index 0f74ea130b0..9115acd3f7d 100755 --- a/tests/fm-guard-stale-banner.test.sh +++ b/tests/fm-guard-stale-banner.test.sh @@ -169,6 +169,22 @@ test_first_stale_call_prints_full_banner() { pass "fm-guard stale banner: first stale call prints the full actionable banner" } +test_full_banner_names_quiet_mode_when_active() { + # kunchenguid/firstmate#2356: the banner's repair line must not misdirect a + # captain in quiet mode to /afk - fm-guard.sh threads the flag's declared + # mode through to fm-supervision-instructions.sh's --afk-mode. + local dir home out + dir=$(make_guard_case quiet-mode-banner) + home=$(case_home "$dir") + printf 'quiet\n%s\n' "$(date '+%s')" > "$home/state/.afk" + out=$(run_guard_case "$dir") + assert_contains "$out" "Quiet mode owns watcher supervision; load /quiet" \ + "full banner did not name /quiet for an active quiet-mode flag" + assert_not_contains "$out" "Away mode owns watcher supervision" \ + "full banner misdirected a quiet-mode captain to /afk" + pass "fm-guard stale banner: repair line is quiet-mode-aware, not hardcoded to away mode" +} + test_repeated_same_episode_prints_reminder_only() { local dir out1 out2 marker lines dir=$(make_guard_case repeated-stale) @@ -870,6 +886,7 @@ test_pi_harness_routes_itself_to_the_extension_model() { } test_first_stale_call_prints_full_banner +test_full_banner_names_quiet_mode_when_active test_repeated_same_episode_prints_reminder_only test_pi_harness_routes_itself_to_the_extension_model test_extension_handoff_with_live_session_is_healthy diff --git a/tests/fm-session-start.test.sh b/tests/fm-session-start.test.sh index cee8c0e6091..260d44d553a 100755 --- a/tests/fm-session-start.test.sh +++ b/tests/fm-session-start.test.sh @@ -2424,6 +2424,48 @@ EOF pass "next step delegates watcher ownership to the AFK daemon" } +test_next_step_quiet_mode_delegates_to_daemon() { + local rec root home fakebin out + rec=$(new_world next-step-quiet) + IFS='|' read -r root home fakebin <<EOF +$rec +EOF + make_fake_toolchain "$fakebin" + make_fake_ps_claude "$fakebin" + printf 'quiet\n%s\n' "$(date '+%s')" > "$home/state/.afk" + + out=$(run_session_start "$home" "$root" "$fakebin:$BASE_PATH") + + assert_contains "$out" "quiet-mode supervision is active" "AFK digest did not report quiet mode for a quiet-content flag" + assert_contains "$out" "only an explicit /quiet off exits it" "AFK digest lost the explicit-only exit rule" + assert_contains "$out" "Quiet mode is active" "next step did not switch to quiet-mode guidance" + assert_contains "$out" "load /quiet" "next step did not name the /quiet skill" + assert_contains "$out" "- Quiet mode: active" "supervision block did not include active quiet state" + assert_not_contains "$out" "Away mode is active" "quiet-mode flag was misreported as away mode" + assert_not_contains "$out" " bin/fm-watch-arm.sh" "quiet next step still told the agent to arm the watcher directly" + + pass "next step delegates watcher ownership to the daemon in quiet mode, distinctly from away mode" +} + +test_next_step_afk_legacy_empty_flag_defaults_away() { + local rec root home fakebin out + rec=$(new_world next-step-afk-legacy) + IFS='|' read -r root home fakebin <<EOF +$rec +EOF + make_fake_toolchain "$fakebin" + make_fake_ps_claude "$fakebin" + : > "$home/state/.afk" + + out=$(run_session_start "$home" "$root" "$fakebin:$BASE_PATH") + + assert_contains "$out" "away-mode supervision is active" "a legacy empty .afk flag was not read as away mode" + assert_contains "$out" "Away mode is active" "a legacy empty .afk flag did not drive away-mode next-step guidance" + assert_not_contains "$out" "Quiet mode" "a legacy empty .afk flag leaked quiet-mode text" + + pass "a legacy empty .afk flag (written before mode existed) still reads as away mode" +} + test_supervision_block_exactly_one_and_pi_diagnostic() { local rec root home fakebin out block_count wake_line sup_line context_line rec=$(new_world pi-supervision-block) @@ -2657,6 +2699,8 @@ test_backlog_compact_tasks_axi_unavailable_uses_manual_fallback test_fleet_digest_empty_fleet test_next_step_sources_x_mode_cadence test_next_step_afk_delegates_to_daemon +test_next_step_quiet_mode_delegates_to_daemon +test_next_step_afk_legacy_empty_flag_defaults_away test_supervision_block_exactly_one_and_pi_diagnostic test_pi_signed_primary_uses_pi_extensions_without_identity_normalization test_pi_diagnostic_rejects_stale_loaded_marker diff --git a/tests/fm-supervision-instructions.test.sh b/tests/fm-supervision-instructions.test.sh index 2d1ddb61e59..6d6a974aaa4 100755 --- a/tests/fm-supervision-instructions.test.sh +++ b/tests/fm-supervision-instructions.test.sh @@ -42,6 +42,31 @@ test_conditional_stanzas() { pass "renderer includes read-only, afk, and effective x-mode current-state stanzas" } +test_quiet_mode_stanzas() { + local home config out + home="$TMP_ROOT/quiet-home" + config="$TMP_ROOT/quiet-config" + mkdir -p "$home/state" "$config" + out=$(FM_HOME="$home" FM_CONFIG_OVERRIDE="$config" "$RENDER" --harness codex --afk 1 --afk-mode quiet) + assert_contains "$out" "- Quiet mode: active" "quiet stanza missing" + assert_contains "$out" "load /quiet" "quiet stanza did not name the /quiet skill" + assert_contains "$out" "Ordinary captain chat does NOT exit it" "quiet stanza lost the explicit-only exit rule" + assert_not_contains "$out" "- Away mode: active" "quiet mode incorrectly rendered as away mode" + out=$(FM_HOME="$home" "$RENDER" --harness codex --afk 1 --afk-mode quiet --repair-line) + assert_contains "$out" "Quiet mode owns watcher supervision; load /quiet" "quiet repair line did not name /quiet" + + out=$(FM_HOME="$home" "$RENDER" --harness codex --afk 1) + assert_contains "$out" "- Away mode: active" "omitting --afk-mode did not default to away (regression)" + assert_not_contains "$out" "Quiet mode" "omitting --afk-mode leaked quiet-mode text" + + out=$(FM_HOME="$home" "$RENDER" --harness codex --afk 1 --afk-mode not-a-real-mode) + assert_contains "$out" "- Away mode: active" "unrecognized --afk-mode value did not fall back to away" + + out=$(FM_HOME="$home" "$RENDER" --harness codex --afk 0) + assert_contains "$out" "- Away/quiet mode: inactive" "inactive stanza missing" + pass "renderer's away/quiet stanzas are mode-aware, default to away, and fall back safely on garbage input" +} + test_repair_lines() { local home out home="$TMP_ROOT/repair-home" @@ -196,6 +221,7 @@ test_pi_snippet_uses_effective_extension_path() { test_selected_harness_block_only test_unknown_fallback test_conditional_stanzas +test_quiet_mode_stanzas test_repair_lines test_cross_harness_ordinary_continuation_and_repair_matrix test_pi_signed_preserves_identity_with_pi_supervision_protocol From 0962d4a02375986894b45c366b36b452e1abc948 Mon Sep 17 00:00:00 2001 From: NewAiCoder <iamacodernow@theinbtw.com> Date: Sun, 13 Sep 2026 02:26:09 -0400 Subject: [PATCH 25/31] fix(bin): let verified harness ancestry outrank retained markers (#3) (#3578) * fix(bin): let verified harness ancestry outrank retained markers (#3) * fix(bin): let a structural harness ancestor outrank a retained marker bin/fm-harness.sh treated a verified environment marker as unconditionally authoritative, so a Codex session started from an environment that had retained CLAUDECODE=1 detected as claude. Session start then emitted Claude's Stop-owned supervision protocol to a Codex primary, and every turn end was blocked for missing Claude recovery. The defect is the precedence boundary, not any one harness. codex, opencode, kimi, and muse publish no identity marker at all, so with markers winning outright any retained CLAUDECODE renamed them; the Cursor-before-Claude ordering was a point patch on the same class of problem, and the launch-time marker clearing only ever covered sessions fm-spawn started. Markers and ancestry are now separate evidence layers that detect_own arbitrates: - no ancestry match, or no marker: the single available layer answers, unchanged; - same harness family: the marker's finer verdict stands, so a launch-selected pi-signed is not flattened to pi by an ancestry walk that can only see the shared launcher name; - different harness with a structural (command-name) ancestor: ancestry wins, because only ancestry proves who owns the process tree; - different harness with only a bare-interpreter script-path match: the marker wins, since a harness-shaped path in some node process's arguments is weaker evidence than a harness publishing its own identity. The correction is symmetric: a retained CURSOR_AGENT no longer renames a claude worker nested under cursor either. Adds fm-harness.sh ancestry [<pid>], ancestry evidence with no marker layer, so a real harness process can be asked what the walk makes of it. tests/fm-harness-precedence.test.sh is the portable regression, built from real renamed processes with no harness installed. Every case drives the two layers apart and asserts each alone as well as the combination, so no case can pass vacuously; it also pins Codex's real two-process install topology, since the fix depends on the native binary being what a tool subprocess meets first. The opt-in drift guard gains the matching live half: each installed harness's real running process must still be identified by the ancestry walk, and it fails naming the harness and version when a release changes that name. Documentation follows the corrected contract in the script header, the harness-adapters detection section, the codex, opencode, kimi, and cursor references, and a dated verification record. * fix(tests): drop the unused argument pass-through in the shim-topology helper bin/fm-lint.sh refused the branch: run_shim declared a `[ancestry]` argument and forwarded "$@", but every call site that varies the environment or passes the ancestry subcommand invokes the shim entry point directly, so the helper is only ever called with no arguments (ShellCheck SC2120/SC2119). Behavior is unchanged: with no arguments "$@" expanded to nothing. * fix(bin): examine the top of the process chain instead of assuming init harness_ancestry stopped as soon as the next pid was 1, on the assumption that pid 1 is always init and can never be a harness. Inside a PID namespace that assumption inverts: the harness itself is pid 1, so the walk never examined the one process that proves who owns the tree, reported no ancestry at all, and handed the verdict straight back to a retained marker. A real Codex session under `codex sandbox`, holding CLAUDECODE=1 and CLAUDE_CODE_ENTRYPOINT=cli, is exactly that shape: it resolved claude and rendered Claude's Stop-owned supervision protocol even with the marker-vs-ancestry precedence boundary in place. The same probe now resolves codex and renders the Codex foreground checkpoint. A host's real pid 1 (init, systemd, launchd) matches no harness name, so examining it costs one ps call and can introduce no false positive; the walk still stops once that top process has been read, and a non-numeric or zero ppid still ends it. tests/fm-harness-precedence.test.sh pins the namespace shape with a fake ps that reports every process as bash with ppid 1 and pid 1 as the harness. The case asserts the marker still answers alone when pid 1 is host-shaped, so it cannot pass vacuously, and it fails against the previous stop condition. * docs(verification): record the real-Codex retained-marker evidence The existing record proved the precedence boundary with the portable regression and recorded each installed harness's process name behind the ancestry walk, but it had no evidence from a real Codex process actually holding a retained Claude marker, which is the failure the boundary exists for. Adds the dated before/after result from codex-cli 0.152.0 under `codex sandbox`, with the exact command and the decisive verdict and rendered protocol on each side, and records the second boundary that shape exposed: the walk must examine the top of the process chain, because inside a PID namespace the harness is pid 1. Refreshes the portable regression's observed output for the case it gained. * no-mistakes(review): blind ancestry in marker-pinned harness tests * no-mistakes(review): blind ancestry in the Pi guard-routing test * no-mistakes(review): classify precedence suite, dedupe ps stub, soften claims * no-mistakes(review): model the spawn-and-wait Codex shim topology * no-mistakes(document): correct stale muse marker-clearing detection claims * no-mistakes: apply CI fixes * fix(bin): examine the top of the chain in the lock and nudge walks too The pid-1 defect corrected in bin/fm-harness.sh survived unchanged in the two other harness-ancestry walks, on the exact topology the branch verified against a real Codex process. bin/fm-session-lock-lib.sh's fm_harness_ancestry_pids stopped as soon as the next pid was 1, so a firstmate whose harness is pid 1 of its own PID namespace could not find that harness at all and did not recognize its own session lock. bin/fm-sessionstart-nudge.sh carried the same stop plus a blanket rejection of a lock pid of 1, so the same session was told to run session start again on every turn. Both walks now compare the top process before stopping, matching the shape used in bin/fm-harness.sh. For the lock walk this is safe because fm_harness_process_matches rejects a host's real pid 1. For the nudge, `kill -0` still gates the lock pid, and on a host an unprivileged `kill -0 1` fails, so a lock file that wrongly names pid 1 leaves the hook silent rather than acting on init. Each walk gains one regression case. The lock case drives a deterministic process table whose pid 1 is the harness and asserts a host-shaped pid 1 still finds nothing, so it cannot pass vacuously. The nudge case needs a real PID namespace, because the builtin `kill -0` gate cannot be reached through a fake ps, and it first proves the same fixture nudges with no lock present; it skips explicitly where unprivileged namespaces are unavailable. * no-mistakes(review): assert comm-strength detection from subprocess vantage in drift guard * fix(bin): verify the live harness guard at the strength the guarantee needs The marker-versus-ancestry boundary this branch ships is a strength claim: detect_own hands an args-strength verdict straight back to a retained foreign marker, so a harness is only protected where the ancestry walk reaches it at comm strength. The installed-harness drift guard probed the pane process alone. Under an interpreter shim the pane process IS the shim, whose own script path is args strength, while the native binary that carries comm strength is its child. The guard therefore observed args for Codex, passed, and would have kept passing if a release stopped spawning that native child at all, while real sessions silently regressed to the original bug. fm-harness.sh gains `ancestry-subtree`, which asks the walk from the pane process and every descendant of it, the vantage a tool subprocess actually occupies. The guard now requires comm strength somewhere in that set and requires every vantage to name the same harness. This supersedes the preceding commit's in-guard leaf walk, which reached the same vantage but left the logic inside the test file, where CI could not pin it and nothing else could reuse it. A harness-dependent check needs both halves: `tests/fm-harness-precedence.test.sh` now carries a portable case proving the subtree probe reaches a strength the top-of-session probe cannot, mutation checked twice, once against the pre-change script and once by disabling descendant enumeration. The subtree walk also avoids depending on tty and process-group semantics that differ between Linux and macOS. Verified live: codex-cli 0.152.0 reports [args codex;comm codex] and Claude Code 2.1.257 reports [comm claude]. * no-mistakes(review): narrow drift guard to the upward vantage path * no-mistakes(review): judge only comm-strength vantages in drift guard * no-mistakes(document): drop duplicated rationale in detection precedence evidence * no-mistakes(review): fix pid-1 nudge case vacuity and descent no-arg expansion * no-mistakes(document): drop branch-relative phrasing in detection precedence evidence * no-mistakes(review): guard remaining empty positional expansions in fm-harness * no-mistakes(document): scope cursor marker-ordering claim to the marker layer * no-mistakes(review): Prefer comm-strength leaves in equal-depth descent ties * no-mistakes(document): Document comm-strength descent tie-break --------- * no-mistakes(review): Blind ancestry in stale gemini/rovo marker-precedence tests * no-mistakes(document): Add missing equal-depth-tie test line to precedence evidence transcript * no-mistakes(review): Fix stale/vacuous agy precedence test, add agy to precedence suite and docs * no-mistakes(document): Fix stale kimi.md marker doc missed by ancestry-precedence fix --------- Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com> --- .agents/skills/harness-adapters/SKILL.md | 3 +- .../references/harness/agy.md | 2 +- .../references/harness/codex.md | 1 + .../references/harness/cursor.md | 1 + .../references/harness/kimi.md | 2 +- .../references/harness/muse.md | 2 +- .../references/harness/opencode.md | 1 + bin/fm-harness.sh | 434 +++++++--- bin/fm-session-lock-lib.sh | 7 +- bin/fm-sessionstart-nudge.sh | 13 +- bin/fm-spawn.sh | 7 +- bin/fm-test-run.sh | 1 + docs/verification/muse.md | 3 +- docs/verification/runtime-backends.md | 106 ++- tests/fm-agy-harness.test.sh | 29 +- tests/fm-cursor-harness.test.sh | 44 +- tests/fm-gemini-harness.test.sh | 13 +- tests/fm-guard-stale-banner.test.sh | 11 +- ...fm-harness-liveness-drift-live-e2e.test.sh | 100 ++- tests/fm-harness-precedence.test.sh | 761 ++++++++++++++++++ tests/fm-kimi-harness.test.sh | 16 +- tests/fm-rovo-harness.test.sh | 13 +- tests/fm-secondmate-harness.test.sh | 41 +- tests/fm-session-lock-ancestry.test.sh | 47 ++ tests/fm-session-start.test.sh | 9 +- tests/fm-sessionstart-nudge.test.sh | 35 + tests/fm-turnend-guard.test.sh | 14 +- tests/fm-watcher-lock.test.sh | 14 +- tests/fm-x-mode.test.sh | 10 +- tests/lib.sh | 28 + 30 files changed, 1578 insertions(+), 190 deletions(-) create mode 100755 tests/fm-harness-precedence.test.sh diff --git a/.agents/skills/harness-adapters/SKILL.md b/.agents/skills/harness-adapters/SKILL.md index 0c3a0353411..ca4f1233245 100644 --- a/.agents/skills/harness-adapters/SKILL.md +++ b/.agents/skills/harness-adapters/SKILL.md @@ -39,7 +39,8 @@ Muse, Gemini, and AGY are verified only for crewmate and scout work, never a sec ## Detection -`../../../bin/fm-harness.sh` prints firstmate's own harness from verified environment markers, then process ancestry. +`../../../bin/fm-harness.sh` prints firstmate's own harness from verified environment markers and process ancestry, and owns how they combine. +A marker names its harness, but a structural ancestor of a different harness outranks it, because a marker is ordinary environment state a child or a multiplexer can retain while ancestry is what proves who owns the process tree. Only `FM_PI_HARNESS=pi-signed` at the launch boundary together with `PI_CODING_AGENT=true` selects Pi-signed; shared unmarked launcher ancestry remains Pi. omp publishes no marker of its own; `FM_OMP_HARNESS=omp` is Firstmate's launch marker and the anchored process name `omp` is its ancestry evidence, as `references/harness/omp.md` records. `../../../bin/fm-spawn.sh` owns worker marker establishment, while the README launch command owns the signed-primary boundary. diff --git a/.agents/skills/harness-adapters/references/harness/agy.md b/.agents/skills/harness-adapters/references/harness/agy.md index 0f38ca16ba6..406dcb1b070 100644 --- a/.agents/skills/harness-adapters/references/harness/agy.md +++ b/.agents/skills/harness-adapters/references/harness/agy.md @@ -39,7 +39,7 @@ The unauthenticated failure mode was not observed, so treat any auth prompt or r ## Detection Detected by ancestry alone: `../../../../../bin/fm-harness.sh` matches the anchored process name `agy`, never `*agy*`. -No environment marker is promoted: `AGENT=1` observed on a live TUI is an inherited launcher value, not an agy identity, and agy does not clear an inherited `CLAUDECODE`, so the spawn clears foreign markers at the launch boundary and the ancestry arm decides. +No environment marker is promoted: `AGENT=1` observed on a live TUI is an inherited launcher value, not an agy identity, and agy does not clear an inherited `CLAUDECODE` - but a structural agy ancestor now outranks that retained marker, which `../../../../../bin/fm-harness.sh` decides without depending on the spawn's own launch-boundary marker clearing. agy is deliberately absent from the session-lock name vocabulary in `../../../../../bin/fm-session-lock-lib.sh`, where muse, gemini, and rovo are also absent: a crewmate-only adapter must never own a home session lock. ## Worker busy state and turn end diff --git a/.agents/skills/harness-adapters/references/harness/codex.md b/.agents/skills/harness-adapters/references/harness/codex.md index 5fb95b8e494..368afadddf9 100644 --- a/.agents/skills/harness-adapters/references/harness/codex.md +++ b/.agents/skills/harness-adapters/references/harness/codex.md @@ -14,6 +14,7 @@ Verified on 2026-06-11 with codex-cli 0.139.0 unless a fact gives a newer versio | Model flag | `--model <model>`. | | Effort flag | `-c 'model_reasoning_effort="<low\|medium\|high\|xhigh>"'`, verified on codex-cli 0.142.1 whose installed schema contains `model_reasoning_effort`, active config uses it, and bundled catalog advertises only these four values while omitting `max`. | | Model discovery | Open the current interactive session's `/model` picker. | +| Marker | None; identity comes from ancestry, and `../../../bin/fm-harness.sh` is what keeps a retained foreign `CLAUDECODE` from renaming it. Verified on 2026-09-01 with codex-cli 0.152.0: the pane process is the `node` npm shim and the native `codex` binary runs as its foreground child, so a tool subprocess reaches the native name directly while the shim itself is identified from its script path. | A directory trust dialog appears on the first run for a repository root: "Do you trust the contents of this directory?" Accept it with Enter and verify the instructions begin processing. diff --git a/.agents/skills/harness-adapters/references/harness/cursor.md b/.agents/skills/harness-adapters/references/harness/cursor.md index 3048a0a8347..0bdede0f20a 100644 --- a/.agents/skills/harness-adapters/references/harness/cursor.md +++ b/.agents/skills/harness-adapters/references/harness/cursor.md @@ -28,6 +28,7 @@ The slash popup consumes the first Enter; that Enter closes it and a genuine sec Cursor does not clear inherited `CLAUDECODE`, so a Cursor worker under Claude carries both markers. `../../../bin/fm-harness.sh` tests Cursor first, and launch also clears foreign markers. Both remain necessary: sanitization covers Firstmate launches, ordering covers hand-started sessions. +That ordering settles the marker layer only, and a nearer Claude ancestor still outranks a retained Cursor marker. Cursor is a bundled Node script, so tmux can report bare `node` while `ps -o comm=` carries its install path. Bare `node` matches nothing; `../../../bin/fm-cursor-lib.sh` proves identity from Cursor's name or install tree in path or argv zero. diff --git a/.agents/skills/harness-adapters/references/harness/kimi.md b/.agents/skills/harness-adapters/references/harness/kimi.md index 8b61d813e48..8799f8bcfcd 100644 --- a/.agents/skills/harness-adapters/references/harness/kimi.md +++ b/.agents/skills/harness-adapters/references/harness/kimi.md @@ -16,7 +16,7 @@ Verified on 2026-07-25 with Kimi Code CLI 0.29.1. | Autonomy | `--auto`; `-y` and `--yolo` are weaker and are not used. | | Trust dialog | None observed on a clean first launch in a fresh pooled worktree. | | Slash submission | One Enter submits, with no popup swallow or settle hazard. | -| Environment marker | None; detection uses process ancestry command name `kimi`. | +| Environment marker | None; identity comes from process ancestry command name `kimi`, which `../../../bin/fm-harness.sh` keeps a retained foreign marker from overriding. | | Composer | Bordered box with a bare `>` prompt glyph and no observed ghost or placeholder text. | | Effort | No verified reasoning-effort flag; `references/common/model-and-effort.md` owns unsupported-value handling. | diff --git a/.agents/skills/harness-adapters/references/harness/muse.md b/.agents/skills/harness-adapters/references/harness/muse.md index a0a9df3e004..d00258e1f45 100644 --- a/.agents/skills/harness-adapters/references/harness/muse.md +++ b/.agents/skills/harness-adapters/references/harness/muse.md @@ -17,7 +17,7 @@ The router owns Muse's task-kind boundary. | Resume | `muse resume --last` or `muse resume <session-uuid>`; bare `muse resume` opens a picker. | | Autonomy | `--yolo` disables approval and sandbox and trusts the workspace. | | Trust | Dialog `Do you trust this workspace?`, choice `1 Trust and continue` preselected for Enter; `--yolo` suppresses it, which fresh task paths require. | -| Marker | None; detect anchored `muse-bin-*` ancestry after clearing foreign primary markers, while `MUSE_CURRENT_SESSION_LOG` is a path rather than identity and its export to tools is unverified. | +| Marker | None; identity comes from anchored `muse-bin-*` ancestry, which `../../../bin/fm-harness.sh` keeps a retained foreign marker from overriding, while `MUSE_CURRENT_SESSION_LOG` is a path rather than identity and its export to tools is unverified. | | Composer | Bordered `⟩`, truecolor `38;2;90;160;255`, luminance about 149.9 and narrowly above ghost threshold 128; typed text is `38;2;204;211;219`, about 209.8, with no observed placeholder or ghost. | | Effort | `--reasoning-effort`, default `high`, accepts `none\|minimal\|low\|medium\|high\|xhigh\|ultra`; shared values expose low through xhigh, explicit captain `max` maps to `ultra`, and `none` or `minimal` remain unreachable. | diff --git a/.agents/skills/harness-adapters/references/harness/opencode.md b/.agents/skills/harness-adapters/references/harness/opencode.md index 8b37a8d35ad..ca9ff18b3f5 100644 --- a/.agents/skills/harness-adapters/references/harness/opencode.md +++ b/.agents/skills/harness-adapters/references/harness/opencode.md @@ -15,6 +15,7 @@ Verified on 2026-06-11 across versions 1.15.7 through 1.17.6, with busy-queue be | Effort flag | None for Firstmate's interactive `opencode --prompt` launch verified on 1.17.6; `opencode run` has `--variant`, but that is not this path. | | Model discovery | Run `opencode models [provider]` to list available provider/model identifiers. | | Trust dialog | None. | +| Marker | None; OpenCode publishes no identity marker, so `../../../bin/fm-harness.sh` identifies it from process ancestry. | OpenCode can auto-upgrade in the background, and the running TUI can exit mid-task. That behavior was observed live during an upgrade from 1.15.7 to 1.17.3. diff --git a/bin/fm-harness.sh b/bin/fm-harness.sh index ecc19935197..7989643f1b6 100755 --- a/bin/fm-harness.sh +++ b/bin/fm-harness.sh @@ -19,13 +19,42 @@ # codex-native/<id>. Other efforts retain # their adapter's existing policy. Native # Codex validates model support at startup. +# fm-harness.sh ancestry [<pid>] print "<strength> <harness>" for the nearest +# harness process at or above <pid> (default this +# process), or nothing when the walk finds none. +# Ancestry evidence only, with no marker layer, so +# a real harness process can be asked what the walk +# makes of it (tests/fm-harness-liveness-drift-live-e2e.test.sh). +# fm-harness.sh ancestry-descent [<pid>] [<leaf-pid>...] +# print each DISTINCT "<strength> <harness>" the walk +# reaches from the vantages on the UPWARD path +# between the deepest descendant of <pid> and <pid> +# itself, deepest first. Same evidence-only purpose +# as `ancestry`, asked from the vantage point a tool +# subprocess actually occupies rather than from the +# top of the session, which is the only place a +# harness behind an interpreter shim can be seen at +# comm strength. Optional <leaf-pid> values restrict +# which descendants may be chosen as the deepest one, +# so a caller that knows the terminal's foreground +# process group can keep a backgrounded process out +# of the selection. # config/secondmate-harness format: a single line "<harness> [<model>] [<effort>]", # whitespace-separated. A bare "<harness>" (today's format) behaves exactly as before: # harness only, no model/effort. Only the first non-empty, non-comment line is parsed. # Model/effort come ONLY from this file - config/crew-harness stays a bare adapter # name and is never parsed for a model. -# Detection layers: verified environment markers first, then process ancestry. -# Record each newly verified env marker here. +# Detection evidence and precedence: +# Markers - verified environment variables a harness publishes about itself. +# Cheap and unambiguous about WHICH harness set them, but they are +# ordinary environment state: a child inherits them, and a terminal +# multiplexer can replay a stale one into an unrelated session. +# Ancestry - the nearest harness process in this process's parent chain. This +# is the structural fact about who actually owns the process tree, +# so it is what settles a disagreement. +# detect_own is the single owner of how the two combine; harness_marker and +# harness_ancestry only report evidence. Record each newly verified env marker +# in harness_marker, and each newly verified command name in harness_ancestry. set -u SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -38,24 +67,18 @@ CONFIG="${FM_CONFIG_OVERRIDE:-$FM_HOME/config}" # shellcheck source=bin/fm-gemini-lib.sh . "$SCRIPT_DIR/fm-gemini-lib.sh" -detect_own() { - # Layer 1: environment markers for verified harnesses. - # Keep marker detection before ancestry detection as an explicit precedence rule. - # Claude, Pi, Grok, and Cursor set verified markers of their own; codex, - # opencode, Kimi, and Muse are markerless, so a foreign marker retained in a terminal - # multiplexer's stored environment can silently misidentify one of them before - # ancestry is consulted. This is a precedence hazard, not evidence that - # CLAUDECODE inheritance into a kimi child was observed; it was not observed. - # Cursor is checked BEFORE claude, deliberately. cursor-agent does NOT clear - # an inherited CLAUDECODE, so a cursor worker launched from a claude primary - # carries BOTH markers and whichever is tested first wins. Cursor's own - # markers are unambiguous when present, so ordering them first is what makes - # the verdict correct; bin/fm-spawn.sh additionally clears the foreign markers - # at the launch boundary. Both are kept: the launch sanitization only covers - # sessions fm-spawn started, while this ordering also covers a cursor session - # a human started by hand. Verified live on cursor-agent 2026.08.11-e8db854: - # CURSOR_INVOKED_AS=cursor-agent is set on the agent process itself, and - # CURSOR_AGENT=1 is set for the child/tool processes this script runs as. +# Print the harness named by a verified environment marker, or nothing when no +# marker is present. Markers only report what the environment CLAIMS; detect_own +# decides whether that claim survives contradicting ancestry. +harness_marker() { + # Cursor is tested BEFORE claude, deliberately. cursor-agent does NOT clear an + # inherited CLAUDECODE, so a cursor session started by hand from a claude + # primary carries BOTH markers and whichever is tested first wins. This + # ordering only settles the case where ancestry finds nothing to arbitrate + # with; a nearer claude ancestor still outranks both in detect_own. + # Verified live on cursor-agent 2026.08.11-e8db854: CURSOR_INVOKED_AS=cursor-agent + # is set on the agent process itself, and CURSOR_AGENT=1 is set for the + # child/tool processes this script runs as. [ "${CURSOR_AGENT:-}" = "1" ] && { echo cursor; return; } [ "${CURSOR_INVOKED_AS:-}" = "cursor-agent" ] && { echo cursor; return; } # Gemini is checked BEFORE claude for exactly cursor's reason above: the @@ -110,98 +133,21 @@ detect_own() { # identified, and any rule that must be RELIABLE under grok has to test the hook # markers too (see .claude/settings.json Stop entries, docs/turnend-guard.md). [ "${GROK_AGENT:-}" = "1" ] && { echo grok; return; } - # muse (Muse Code) publishes no harness-identity marker of its own. The only - # MUSE_* variable it is documented to hand a child is MUSE_CURRENT_SESSION_LOG, - # a per-session log PATH rather than an identity, and its export to tool - # subprocesses is unverified (verified: muse 0.1.0-R708.1), so muse is detected - # by ancestry alone below. Do NOT promote MUSE_CURRENT_SESSION_LOG to a marker - # without verifying it reaches children AND that it cannot survive in a - # multiplexer's stored environment, which is the precedence hazard above. - # Layer 2: walk the parent chain and match the command name. - local pid=$$ comm args argv0 - for _ in 1 2 3 4 5 6 7 8; do - comm=$(ps -o comm= -p "$pid" 2>/dev/null) || break - argv0=$(fm_cursor_argv0_for_pid "$pid" "$comm" 2>/dev/null || true) - if fm_cursor_process_matches "$comm" '' "$argv0"; then - echo cursor - return - fi - if fm_gemini_path_is_gemini "$comm"; then - echo gemini - return - fi - case "$(basename -- "$comm")" in - # gemini precedes claude here for the same precedence reason as the - # marker layer above, so a gemini worker under a claude primary is never - # read as claude. This arm covers a natively-named gemini binary only. - # It does NOT reach the currently installed CLI, which is a node bundle - # (~/.local/bin/gemini -> @google/gemini-cli/bundle/gemini.js): modern - # Node on Linux reports `comm` as MainThread rather than node (measured - # on Node v24.20.0), so neither this arm nor the node interpreter arm - # below matches a live gemini process. GEMINI_CLI above is therefore - # load-bearing for gemini rather than a fast path, which is why gemini - # is not offered as a primary or secondmate harness. Do NOT add - # MainThread to the interpreter arm to close this: that would make the - # args of EVERY node process searchable and let an unrelated node - # command carrying a harness name in its arguments claim an identity. - *claude*) echo claude; return ;; - *codex*) echo codex; return ;; - *opencode*) echo opencode; return ;; - *grok*) echo grok; return ;; - kimi) echo kimi; return ;; - rovo) echo rovo; return ;; - # muse's installed launcher ~/.local/bin/muse execs ~/.local/bin/muse-bin-<version> - # (verified in the published launcher, muse 0.1.0-R708.1), so the live process - # name carries the version and CHANGES on every auto-update. Match the stable - # prefix rather than any exact name. Deliberately anchored, never *muse*, so - # unrelated commands (musescore, amuse) cannot be misread as this harness. - muse|muse-bin-*) echo muse; return ;; - pi-signed) echo pi; return ;; - pi) echo pi; return ;; - # omp is a Bun-compiled single binary whose process name is exactly `omp` - # (verified, omp 18.1.11: `ps -o comm=` reports omp from both its `!` - # bash path and the model's bash tool). Anchored, never *omp*, so ompd, - # comp, and similar unrelated commands are not misread as this harness. - # It sits above the node*|python* interpreter fallback deliberately: the - # optional claude-bridge extension runs a nested executable literally - # named `claude` with its own node child, and that fallback's *claude* - # args glob would otherwise claim it if that subtree were ever walked. - omp) echo omp; return ;; - # agy (Antigravity CLI) is a Go-compiled single binary whose process name - # is exactly `agy` (verified, agy 1.2.0: `ps -o comm=` reports agy and - # Herdr's process-info reports name agy with argv[0] agy). Anchored, never - # *agy*, so unrelated commands cannot be misread as this harness. agy - # publishes no harness-identity marker of its own (a live 1.2.0 TUI - # carries no AGY_* or ANTIGRAVITY_* variable; AGENT=1 seen there is an - # inherited launcher value, not an agy identity), so like muse it is - # detected by ancestry alone. - agy) echo agy; return ;; - node*|python*) - # Bare interpreter: match the harness name in its script path. - args=$(ps -o args= -p "$pid" 2>/dev/null) - if fm_gemini_args_are_gemini "$args"; then - echo gemini - return - fi - case "$args" in - *claude*) echo claude; return ;; - *codex*) echo codex; return ;; - *opencode*) echo opencode; return ;; - *grok*) echo grok; return ;; - *" pi "*|*/pi) echo pi; return ;; - esac ;; - esac - pid=$(ps -o ppid= -p "$pid" 2>/dev/null | tr -d ' ') - if [ -z "$pid" ] || [ "$pid" -le 1 ]; then - break - fi - done - echo unknown + # codex, opencode, kimi, muse, and agy publish no harness-identity marker at all, so + # they are never named here and are identified by ancestry alone. That is the + # whole reason a foreign marker must not outrank ancestry: with markers winning + # unconditionally, any retained CLAUDECODE would silently rename one of them. + # muse's only documented child variable is MUSE_CURRENT_SESSION_LOG, a + # per-session log PATH rather than an identity, and its export to tool + # subprocesses is unverified (verified: muse 0.1.0-R708.1). Do NOT promote it + # to a marker without verifying it reaches children AND that it cannot survive + # in a multiplexer's stored environment. + return 0 } # True when an exact `omp` process sits within eight parents of this one. The -# same anchored match as the ancestry walk in detect_own, kept separate so the -# marker precedence above can demand real process evidence. +# same anchored match as the ancestry walk below, kept separate so the marker +# precedence above can demand real process evidence before trusting FM_OMP_HARNESS. ancestry_names_omp() { local pid=$$ comm for _ in 1 2 3 4 5 6 7 8; do @@ -213,6 +159,259 @@ ancestry_names_omp() { return 1 } +# Print "<strength> <harness>" when one process identifies a harness, or nothing. +# Strength records how the match was made: +# comm - the ancestor's own executable name identifies the harness. This is a +# structural fact about the running program, so it outranks a marker. +# args - a bare interpreter matched only because a harness name appears in the +# script path it was handed. This is the weakest inference in this file +# (any node process holding a harness-shaped path matches it), so it is +# used only when no marker is present. +harness_process_verdict() { # <pid> + local pid=$1 comm args argv0 + comm=$(ps -o comm= -p "$pid" 2>/dev/null) || return 0 + argv0=$(fm_cursor_argv0_for_pid "$pid" "$comm" 2>/dev/null || true) + if fm_cursor_process_matches "$comm" '' "$argv0"; then + echo "comm cursor" + return + fi + if fm_gemini_path_is_gemini "$comm"; then + echo "comm gemini" + return + fi + case "$(basename -- "$comm")" in + # gemini precedes claude here for the same precedence reason as the + # marker layer above, so a gemini worker under a claude primary is never + # read as claude. This arm covers a natively-named gemini binary only. + # It does NOT reach the currently installed CLI, which is a node bundle + # (~/.local/bin/gemini -> @google/gemini-cli/bundle/gemini.js): modern + # Node on Linux reports `comm` as MainThread rather than node (measured + # on Node v24.20.0), so neither this arm nor the node interpreter arm + # below matches a live gemini process. GEMINI_CLI above is therefore + # load-bearing for gemini rather than a fast path, which is why gemini + # is not offered as a primary or secondmate harness. Do NOT add + # MainThread to the interpreter arm to close this: that would make the + # args of EVERY node process searchable and let an unrelated node + # command carrying a harness name in its arguments claim an identity. + *claude*) echo "comm claude"; return ;; + *codex*) echo "comm codex"; return ;; + *opencode*) echo "comm opencode"; return ;; + *grok*) echo "comm grok"; return ;; + kimi) echo "comm kimi"; return ;; + rovo) echo "comm rovo"; return ;; + # muse's installed launcher ~/.local/bin/muse execs ~/.local/bin/muse-bin-<version> + # (verified in the published launcher, muse 0.1.0-R708.1), so the live process + # name carries the version and CHANGES on every auto-update. Match the stable + # prefix rather than any exact name. Deliberately anchored, never *muse*, so + # unrelated commands (musescore, amuse) cannot be misread as this harness. + muse|muse-bin-*) echo "comm muse"; return ;; + # Both Pi identities share this launcher name. Ancestry can only prove the + # FAMILY; only the launch-boundary marker selects the signed identity, which + # is why detect_own keeps a marker that agrees on the family. + pi-signed) echo "comm pi"; return ;; + pi) echo "comm pi"; return ;; + # omp is a Bun-compiled single binary whose process name is exactly `omp` + # (verified, omp 18.1.11: `ps -o comm=` reports omp from both its `!` + # bash path and the model's bash tool). Anchored, never *omp*, so ompd, + # comp, and similar unrelated commands are not misread as this harness. + # It sits above the node*|python* interpreter fallback deliberately: the + # optional claude-bridge extension runs a nested executable literally + # named `claude` with its own node child, and that fallback's *claude* + # args glob would otherwise claim it if that subtree were ever walked. + omp) echo "comm omp"; return ;; + # agy (Antigravity CLI) is a Go-compiled single binary whose process name + # is exactly `agy` (verified, agy 1.2.0: `ps -o comm=` reports agy and + # Herdr's process-info reports name agy with argv[0] agy). Anchored, never + # *agy*, so unrelated commands cannot be misread as this harness. agy + # publishes no harness-identity marker of its own (a live 1.2.0 TUI + # carries no AGY_* or ANTIGRAVITY_* variable; AGENT=1 seen there is an + # inherited launcher value, not an agy identity), so like muse it is + # detected by ancestry alone. + agy) echo "comm agy"; return ;; + node*|python*) + # Bare interpreter: match the harness name in its script path. + args=$(ps -o args= -p "$pid" 2>/dev/null) + if fm_gemini_args_are_gemini "$args"; then + echo "args gemini" + return + fi + case "$args" in + *claude*) echo "args claude"; return ;; + *codex*) echo "args codex"; return ;; + *opencode*) echo "args opencode"; return ;; + *grok*) echo "args grok"; return ;; + *" pi "*|*/pi) echo "args pi"; return ;; + esac ;; + esac +} + +# Print the verdict for the NEAREST harness process in the parent chain, or +# nothing when the walk finds none. The nearest match wins, so a worker nested +# inside another harness resolves to its own harness. +harness_ancestry() { # [<pid>] + local pid=${1:-$$} verdict + for _ in 1 2 3 4 5 6 7 8; do + verdict=$(harness_process_verdict "$pid") + [ -z "$verdict" ] || { echo "$verdict"; return; } + pid=$(ps -o ppid= -p "$pid" 2>/dev/null | tr -d ' ') + # Stop only once the walk has EXAMINED the top of the chain. Inside a PID + # namespace the harness itself is pid 1 - a container, or the `codex sandbox` + # this boundary was proven in - so breaking as soon as the next pid is 1 + # skips the one process that identifies the session and hands the verdict + # straight back to a retained marker. A host's real pid 1 (init, systemd, + # launchd) matches no harness name above, so examining it costs one ps call + # and can introduce no false positive. + case "$pid" in '' | *[!0-9]*) break ;; esac + [ "$pid" -ge 1 ] || break + done + return 0 +} + +# Print the pids on the UPWARD path between the deepest descendant of <root> and +# <root> itself, deepest first. Optional <eligible-leaf-pid> values restrict which +# descendants may be chosen as that deepest one; with none given every descendant +# is eligible. Bounded to the same eight levels harness_ancestry climbs, so a deep +# or pathological tree cannot make this walk unbounded. +process_descent_path() { # <root> [<eligible-leaf-pid>...] + local root=${1:-$$} eligible any hit pairs frontier next pid child parent verdict + local parents='' depth=0 best best_depth=0 best_strength='' hops=0 + case "$root" in '' | *[!0-9]*) return 0 ;; esac + shift 2>/dev/null || true + eligible=" ${*+$*} " + any=0 + [ "$#" -eq 0 ] && any=1 + pairs=$(ps -eo pid=,ppid= 2>/dev/null) || { printf '%s\n' "$root"; return 0; } + best=$root + frontier=$root + while [ -n "$frontier" ] && [ "$depth" -lt 8 ]; do + next= + for pid in $frontier; do + while read -r child parent; do + [ "$parent" = "$pid" ] || continue + [ "$child" != "$pid" ] || continue + parents="$parents $child:$pid" + next="$next $child" + if [ "$any" = 1 ]; then + hit=1 + else + case "$eligible" in + *" $child "*) hit=1 ;; + *) hit=0 ;; + esac + fi + if [ "$hit" = 1 ]; then + verdict=$(harness_process_verdict "$child") + if [ $((depth + 1)) -gt "$best_depth" ]; then + best=$child + best_depth=$((depth + 1)) + best_strength=${verdict%% *} + # At equal depth, prefer the leaf whose own executable reaches comm + # strength. Otherwise an earlier MCP interpreter carrying a foreign + # harness path can hide a native harness sibling purely through ps + # ordering. This repairs the chosen path's comm-strength guarantee; + # args-strength foreign verdicts remain excluded from cross-checking. + elif [ $((depth + 1)) -eq "$best_depth" ] \ + && [ "$best_strength" != comm ] && [ "${verdict%% *}" = comm ]; then + best=$child + best_strength='comm' + fi + fi + done <<EOF +$pairs +EOF + done + frontier=$next + depth=$((depth + 1)) + done + + pid=$best + while [ -n "$pid" ] && [ "$hops" -le 8 ]; do + printf '%s\n' "$pid" + [ "$pid" != "$root" ] || break + parent= + case "$parents" in + *" $pid:"*) + parent=${parents##*" $pid:"} + parent=${parent%% *} ;; + esac + pid=$parent + hops=$((hops + 1)) + done +} + +# Print each DISTINCT "<strength> <harness>" verdict harness_ancestry reaches from +# the vantages on the upward path between the deepest descendant of <root> and +# <root>, one per line, deepest first. +# +# Why a descent path and not <root> alone: detect_own always runs from a TOOL +# SUBPROCESS inside a session, never from the process at the top of it, and that +# difference decides whether a retained foreign marker can rename the session. A +# harness that ships as an interpreter shim spawning its native binary as a CHILD +# is only args strength when asked from the shim, and detect_own hands an +# args-strength verdict straight back to the marker; the native child is comm +# strength and outranks it. Asking from below is what puts the question at the +# vantage point a real session uses, so a guard built on this can assert the +# strength the shipped guarantee actually depends on +# (tests/fm-harness-liveness-drift-live-e2e.test.sh). +# +# Why the upward path and not the whole subtree: harness_ancestry only ever climbs, +# so a SIBLING branch is a vantage firstmate's own detection can never occupy. A +# harness-spawned MCP server running as `node <home>/.claude/mcp/<server>.js` matches +# *claude* on its script path in the bare-interpreter branch above and would report a +# foreign harness from a process no real tool subprocess ever asks from. +harness_ancestry_descent() { # <root> [<eligible-leaf-pid>...] + local pid verdict seen= + for pid in $(process_descent_path "$@"); do + verdict=$(harness_ancestry "$pid") + [ -n "$verdict" ] || continue + case "$seen" in *"|$verdict|"*) continue ;; esac + seen="$seen|$verdict|" + printf '%s\n' "$verdict" + done +} + +# Collapse a verdict to the harness FAMILY its evidence can actually prove, so a +# marker's more specific verdict and ancestry's coarser one are not read as a +# disagreement. Only Pi has two identities behind one launcher name. +harness_family() { + case "$1" in + pi-signed) printf 'pi\n' ;; + *) printf '%s\n' "$1" ;; + esac +} + +# Combine the two evidence layers. The precedence boundary, in one rule: a +# marker names its harness, but only ancestry proves which harness owns this +# process tree, so a structural (comm) ancestor of a DIFFERENT harness wins. +# - No ancestry match: the marker is the only evidence there is. +# - No marker: ancestry is the only evidence there is. +# - Same family: keep the marker's verdict, which is the more specific one +# (pi-signed, which ancestry can only see as pi). +# - Different harness, structural ancestor: ancestry wins. This is what stops +# an inherited or multiplexer-retained CLAUDECODE from renaming a markerless +# codex, opencode, kimi, or muse session, and symmetrically stops a retained +# CURSOR_AGENT from renaming a claude worker nested under cursor. +# - Different harness, interpreter-args ancestor only: the marker wins, because +# a harness-shaped path in some node process's arguments is weaker evidence +# than a harness publishing its own identity. +detect_own() { + local marker ancestry strength harness + marker=$(harness_marker) + ancestry=$(harness_ancestry) + if [ -z "$ancestry" ]; then + if [ -n "$marker" ]; then echo "$marker"; else echo unknown; fi + return + fi + strength=${ancestry%% *} + harness=${ancestry#* } + [ -n "$marker" ] || { echo "$harness"; return; } + if [ "$(harness_family "$marker")" = "$(harness_family "$harness")" ]; then + echo "$marker" + return + fi + if [ "$strength" = comm ]; then echo "$harness"; else echo "$marker"; fi +} + # Resolve the effective crewmate harness: config/crew-harness (a bare adapter # name) wins; absent or "default" mirrors firstmate's own harness. resolve_crew() { @@ -299,6 +498,23 @@ validate_native_effort() { case "${1:-}" in validate-native-effort) shift; validate_native_effort "$@" ;; + ancestry) + case "${2:-}" in + ''|*[!0-9]*) [ -z "${2:-}" ] || { echo "error: ancestry takes a numeric pid" >&2; exit 2; } ;; + esac + harness_ancestry "${2:-$$}" + ;; + ancestry-descent) + shift + for arg in ${1+"$@"}; do + case "$arg" in + ''|*[!0-9]*) echo "error: ancestry-descent takes numeric pids" >&2; exit 2 ;; + esac + done + descent_pid="${1:-$$}" + [ "$#" -eq 0 ] || shift + harness_ancestry_descent "$descent_pid" ${1+"$@"} + ;; crew) resolve_crew ;; secondmate) resolve_secondmate ;; secondmate-model) resolve_secondmate_model ;; diff --git a/bin/fm-session-lock-lib.sh b/bin/fm-session-lock-lib.sh index c2a117b0b84..91c901f820b 100644 --- a/bin/fm-session-lock-lib.sh +++ b/bin/fm-session-lock-lib.sh @@ -122,7 +122,12 @@ fm_harness_ancestry_pids() { break fi pid=$(ps -o ppid= -p "$pid" 2>/dev/null | tr -d ' ') - [ -n "$pid" ] && [ "$pid" -gt 1 ] || break + # Examine the top of the chain before stopping. Inside a PID namespace the + # harness itself is pid 1, so stopping as soon as the next pid is 1 hides the + # very process this walk exists to find. A host's real pid 1 (init, systemd, + # launchd) is not harness-shaped, so fm_harness_process_matches rejects it. + case "$pid" in '' | *[!0-9]*) break ;; esac + [ "$pid" -ge 1 ] || break done [ "$printed" -eq 1 ] } diff --git a/bin/fm-sessionstart-nudge.sh b/bin/fm-sessionstart-nudge.sh index fccf775dd95..a12aa3e4628 100755 --- a/bin/fm-sessionstart-nudge.sh +++ b/bin/fm-sessionstart-nudge.sh @@ -25,13 +25,22 @@ lock_is_in_ancestry() { [ -f "$STATE/.lock" ] || return 1 IFS= read -r lock_pid < "$STATE/.lock" 2>/dev/null || return 1 case "$lock_pid" in - ''|*[!0-9]*|1) return 1 ;; + # A lock pid of 1 is legitimate inside a PID namespace, where the harness + # holding the home lock IS pid 1, so it is no longer rejected outright; the + # liveness check below still gates it. On a host, a lock file that wrongly + # names pid 1 can now make this hook conclude the lock is already held and + # stay silent, which is the safe direction for a SessionStart hook whose only + # outputs are one nudge line or nothing. + ''|*[!0-9]*) return 1 ;; esac kill -0 "$lock_pid" 2>/dev/null || return 1 for _ in 1 2 3 4 5 6 7 8; do [ "$pid" = "$lock_pid" ] && return 0 pid=$(ps -o ppid= -p "$pid" 2>/dev/null | tr -d ' ') - [ -n "$pid" ] && [ "$pid" -gt 1 ] || return 1 + # Stop only after the top of the chain has been compared, for the same + # namespace reason as bin/fm-session-lock-lib.sh's walk. + case "$pid" in '' | *[!0-9]*) return 1 ;; esac + [ "$pid" -ge 1 ] || return 1 done return 1 } diff --git a/bin/fm-spawn.sh b/bin/fm-spawn.sh index 2188298b37e..366f0c1be00 100755 --- a/bin/fm-spawn.sh +++ b/bin/fm-spawn.sh @@ -1687,7 +1687,12 @@ launch_template() { # plugin engine is off in the default build, so firstmate folds muse's own # session event log instead (bin/fm-busy-lib.sh), bound by the sidecar # written below. Nothing to place in the template for it. - # codex, opencode, and kimi are also markerless and share this inherited-marker hazard; changing their verified launch boundaries belongs in follow-up work. + # codex, opencode, and kimi are markerless too and inherit foreign markers the + # same way, but detection no longer depends on this launch-side clearing: + # bin/fm-harness.sh lets a markerless harness's structural ancestor outrank an + # inherited marker. The clearing stays on the cursor and muse templates as the + # verified launch behavior their evidence records, not as the only thing + # standing between a retained marker and a misidentified worker. muse) printf '%s' 'env -u CLAUDECODE -u PI_CODING_AGENT -u GROK_AGENT -u FM_PI_HARNESS XDG_CONFIG_HOME=__MUSECONFIG__ XDG_DATA_HOME=__MUSEDATA__ MUSE_EXPERIMENTAL_FOREIGN_PERSONAL_CONTEXT_KILL=on __MUSEBIN__ --yolo __MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' ;; # rovo (Atlassian Rovo CLI): a positional brief is dead-on-arrival - rovo # loads, never enters a working state, and drops back to an idle shell within diff --git a/bin/fm-test-run.sh b/bin/fm-test-run.sh index e08ea75b444..e69855f3875 100755 --- a/bin/fm-test-run.sh +++ b/bin/fm-test-run.sh @@ -278,6 +278,7 @@ family_for_basename() { fm-composer-ghost.test.sh|fm-composer-lib.test.sh|\ fm-crew-state.test.sh|fm-captain-hold-lifecycle.test.sh|\ fm-documentation-audiences.test.sh|fm-ensure-agents-md.test.sh|fm-grok-harness.test.sh|\ + fm-harness-precedence.test.sh|\ fm-kimi-harness.test.sh|fm-muse-harness.test.sh|fm-rovo-harness.test.sh|fm-agy-harness.test.sh|fm-omp-harness.test.sh|fm-herdr-lab.test.sh|fm-lint.test.sh|\ fm-lint-workflows.test.sh|\ fm-operational-input.test.sh|fm-pi-primary-types.test.sh|\ diff --git a/docs/verification/muse.md b/docs/verification/muse.md index 38d12653a6a..ba0d52c234b 100644 --- a/docs/verification/muse.md +++ b/docs/verification/muse.md @@ -48,7 +48,8 @@ $ grep -nE 'muse-bin|exec ' launcher.sh `ps -o comm= -p <pid>` returns the full executable path, whose basename is `muse-bin-<version>`. That is why both `bin/fm-harness.sh` and `bin/backends/tmux.sh` match the anchored prefix `muse-bin-*` rather than an exact name, and why neither can rely on an install-path component: `~/.local/bin/muse-bin-<version>` contains no `muse` path component. -The Muse launch clears `CLAUDECODE`, `PI_CODING_AGENT`, `GROK_AGENT`, `FM_PI_HARNESS`, `CURSOR_AGENT`, and `CURSOR_INVOKED_AS` before the worker starts so foreign primary markers cannot override the versioned ancestry. +The Muse launch clears `CLAUDECODE`, `PI_CODING_AGENT`, `GROK_AGENT`, `FM_PI_HARNESS`, `CURSOR_AGENT`, and `CURSOR_INVOKED_AS` before the worker starts, which is the verified launch behavior rather than what detection depends on. +[Harness detection precedence](runtime-backends.md#harness-detection-precedence) owns why a retained foreign marker cannot override the versioned ancestry. [`runtime-backends.md`](runtime-backends.md#agent-liveness-name-sources) owns the resulting tmux liveness verdict and its relationship to the portable decoy regression. diff --git a/docs/verification/runtime-backends.md b/docs/verification/runtime-backends.md index 80e032d706a..078eb069549 100644 --- a/docs/verification/runtime-backends.md +++ b/docs/verification/runtime-backends.md @@ -6,6 +6,107 @@ This record contains reusable version-scoped evidence for active runtime guarant The backend guides own current setup, safety boundaries, and limitations. Exact task chronology, branch names, temporary homes, local paths, process ids, thread ids, and delivery transcripts remain in private reports or PR evidence. +## Harness detection precedence + +Firstmate's own harness comes from two kinds of evidence, and `bin/fm-harness.sh` owns how they combine: an environment marker names its harness, and the nearest harness process in the parent chain proves who owns the process tree. +A marker alone is not proof of ownership, because it is ordinary environment state that a child inherits and a terminal multiplexer can replay into an unrelated session. +Verified on 2026-09-02 on Linux 7.1.12 with the portable regression, which builds every case from real renamed processes and no installed harness: + +```sh +bin/fm-test-run.sh tests/fm-harness-precedence.test.sh +``` + +Observed output: + +```text +ok - a markerless harness keeps its identity under an inherited foreign marker +ok - a harness that publishes a marker inside its own process tree is unchanged +ok - with ancestry silent, the marker layer and its cursor-first ordering still decide +ok - a retained cursor marker does not rename a nested claude worker +ok - an agreeing marker keeps Pi's finer identity that ancestry cannot prove +ok - an interpreter script-path match answers alone but never outranks a marker +ok - a native harness binary under an interpreter shim decides at comm strength +ok - a harness that is pid 1 of its own namespace is examined, not skipped +ok - the descent probe reaches comm strength where the top-of-session probe sees only args +ok - the descent probe reports no verdict from a sibling branch detection cannot reach +ok - a foreign args-only verdict at the deepest vantage leaves the comm-strength identity intact +ok - equal-depth descent ties prefer the comm-strength leaf regardless of spawn order +ok - session start renders the Codex protocol for a Codex primary holding a retained CLAUDECODE +FM_TEST_SUMMARY total=1 failed=0 skipped_gate=0 duration_ms=3666 +``` + +Before that boundary existed, a Codex session started from an environment that had retained `CLAUDECODE=1` reported `claude`, and session start emitted Claude's Stop-owned supervision protocol to a Codex primary. +The same live shape, reproduced with a real process named `codex` and no installed harness, now reports `codex` with the marker present and `claude` with the marker present and ancestry blinded, which is what proves the case is not vacuous. + +### A real Codex session holding a retained Claude marker + +The portable regression builds its process tree from renamed executables, so the same guarantee is proven again against the real installed Codex. +`codex sandbox` runs a command under the installed native binary with no model turn, inside a PID namespace where that binary is pid 1 and the command is pid 2. +Verified on 2026-09-01 with codex-cli 0.152.0 on Linux 7.1.10, with both Claude markers retained in the launching environment: + +```sh +CLAUDECODE=1 CLAUDE_CODE_ENTRYPOINT=cli codex sandbox bash -c \ + 'cd <checkout> && bin/fm-harness.sh; bin/fm-harness.sh ancestry; bin/fm-supervision-instructions.sh' +``` + +Against the parent commit, with `CLAUDECODE=1` and `CLAUDE_CODE_ENTRYPOINT=cli` confirmed present in the probe's own environment and the chain reading pid 2 `bash` to pid 1 `codex`: + +```text +verdict=claude +SUPERVISION OPERATING INSTRUCTIONS - primary harness: claude +Mode: Claude Stop-hook-owned supervision. +``` + +With the current boundaries in place, from the same command and the same process chain: + +```text +verdict=codex +ancestry=comm codex +SUPERVISION OPERATING INSTRUCTIONS - primary harness: codex +Mode: Codex foreground checkpoint. +``` + +Two boundaries are load-bearing here, and the marker-versus-ancestry precedence above is only the first. +The walk also used to stop as soon as the next pid was 1, on the assumption that pid 1 is always init. +That assumption inverts inside a PID namespace, where the harness is pid 1: the walk returned no ancestry at all, so the retained marker won by default even with precedence corrected. +The walk now examines that top process before stopping, which costs one `ps` call and can introduce no false positive, because a host's real pid 1 (init, systemd, launchd) matches no harness name. +The portable regression asserts both directions of that case: a host-shaped pid 1 still leaves the marker to answer, and a harness at pid 1 outranks it. + +Run on the host under Claude Code 2.1.252 with the same two markers set, the same probe reports `claude`, `comm claude`, and Claude's Stop-owned protocol, so the correction does not trade one misidentification for its inverse. + +### Real harness process names behind the walk + +The detection half of the opt-in drift guard asks the ancestry walk what it makes of each INSTALLED harness's real running process: + +```sh +FM_HARNESS_LIVENESS_DRIFT=1 bin/fm-test-run.sh tests/fm-harness-liveness-drift-live-e2e.test.sh +``` + +The guard probes the upward path between the deepest foreground descendant of the pane process and the pane process itself, and reports each distinct verdict that vantage set produces. +Observed on 2026-09-02 for the harnesses installed on that machine: + +```text +# claude 2.1.258 (Claude Code): title='claude' foreground=[claude ] +# claude 2.1.258 (Claude Code): ancestry verdicts=[comm claude] +# codex codex-cli 0.152.0: title='node' foreground=[node codex ] +# codex codex-cli 0.152.0: ancestry verdicts=[comm codex;args codex] +``` + +The verdicts are reported deepest first, so Codex's native child answers before the shim above it. +When eligible foreground descendants tie at the greatest depth, the probe prefers a leaf whose own verdict reaches comm strength; if none does, it keeps the first leaf, so process-table ordering cannot hide an equally deep native harness binary behind an args-strength interpreter. + +Codex ships as a `node` npm shim that spawns its native `codex` binary as a foreground child, which is why its two verdicts differ: the pane process is identified only from the shim's script path, and the native child is what carries the process name. +That difference is the reason the guard cannot probe the pane process alone. +The guarantee this guard holds is a strength claim, not only an identity one, because `detect_own` hands an args-strength verdict straight back to a retained foreign marker. +A pane-only probe would have observed `args codex`, passed, and gone on passing if a later release stopped spawning the native child, while real sessions silently regressed to the original bug. +Probing from below asks the question from the vantage a tool subprocess actually occupies, so the guard can require comm strength somewhere in the session and require every comm-strength vantage to name the same harness. +The vantage set stops at the upward path rather than the whole subtree, because `harness_ancestry` only ever climbs and a sibling branch is therefore a vantage firstmate's own detection can never occupy. +The reject-other-harness cross-check judges comm-strength vantages only, because an args-strength verdict is path-ambiguous by construction: a harness-spawned MCP server running as `node <home>/.claude/mcp/<server>.js` answers `args claude` purely from the `.claude` path component, and such a server is normally a child of the agent binary, so it can be the deepest descendant and sit on this path. +That narrowing changes only which vantages the cross-check judges; the comm-strength requirement itself is unchanged. +A single-process harness has no descendant that adds a distinct verdict, which is why `claude` reports one. +The portable regression pins every half without any harness installed: `tests/fm-harness-precedence.test.sh` asserts that this two-process topology decides at comm strength, that the descent probe reaches a strength the top-of-session probe cannot, that a sibling branch answering a foreign harness contributes no verdict, that a foreign args-only verdict at the deepest vantage leaves the comm-strength identity intact, and that equal-depth ties choose the comm-strength leaf regardless of process ordering. +The run did not reach `opencode`, `pi`, `pi-signed`, `grok`, `kimi`, or `muse`, which were not installed, and stopped at the same pre-existing liveness failure for `cursor` 3.18.9, whose resolved binary on that machine is the editor rather than `cursor-agent`; those adapters are unverified by this run. + ## tmux Foreground-process behavior was verified on 2026-07-07 with tmux 3.6a on macOS. @@ -1441,8 +1542,9 @@ Read from the live agent process and from a tool subprocess it spawned: | `CURSOR_CONVERSATION_ID=<uuid>` | child/tool processes | | `AGENT_TRANSCRIPTS=<projects-root>/<slug>/agent-transcripts` | child/tool processes | -Cursor does not clear an inherited `CLAUDECODE`, so ordering decides the verdict. -With both markers set, `bin/fm-harness.sh` reports `cursor`; with `CLAUDECODE` alone it still reports `claude`. +Cursor does not clear an inherited `CLAUDECODE`, so ordering decides the verdict within the marker layer. +With both markers set and ancestry silent, `bin/fm-harness.sh` reports `cursor`; with `CLAUDECODE` alone it still reports `claude`. +[Harness detection precedence](#harness-detection-precedence) owns what happens when a structural ancestor of another harness is present, which outranks either marker. ### Composer diff --git a/tests/fm-agy-harness.test.sh b/tests/fm-agy-harness.test.sh index 6acaf588d79..d13f23c625d 100755 --- a/tests/fm-agy-harness.test.sh +++ b/tests/fm-agy-harness.test.sh @@ -7,8 +7,9 @@ # carries no AGY_* variable; AGENT=1 there is inherited launcher state), # so detection is ancestry alone on the anchored process name `agy`. # 2. The anchored match must never claim unrelated commands containing the -# fragment, and an inherited CLAUDECODE still outranks ancestry until the -# spawn clears it - the clearing is load-bearing, not cosmetic. +# fragment, and a structural agy ancestor now outranks a retained or +# inherited CLAUDECODE - tests/fm-harness-precedence.test.sh owns the +# general boundary. # 3. The launch carries the brief via --prompt-interactive with --model, # --effort, and --dangerously-skip-permissions; a requested model a # reachable `agy models` omits refuses loudly instead of wedging a pane, @@ -118,19 +119,29 @@ SH } test_agy_claims_no_inherited_launcher_marker() { - local out + local fakebin out # AGENT=1 was observed on a live agy TUI as inherited launcher state, so it # must never promote to an agy identity the way GEMINI_CLI does for gemini. out=$(AGENT=1 "$HARNESS") [ "$out" != agy ] \ || fail "an inherited AGENT=1 must never claim the agy identity, got '$out'" # Drive the hazard the other way: agy does not clear an inherited CLAUDECODE, - # so the marker still wins over a real agy ancestor until the spawn clears it - # at the launch boundary. Pin both halves so neither can rot silently. - out=$(CLAUDECODE=1 FAKE_PS_COMM=agy FAKE_PS_ARGS='agy --prompt-interactive hi' \ - PATH="$(fm_fakebin "$TMP_ROOT/anc-claude"):$PATH" "$HARNESS") - [ "$out" = claude ] \ - || fail "an inherited CLAUDECODE must still outrank agy ancestry, got '$out'" + # so a structural agy ancestor must still outrank the retained marker rather + # than being renamed away from it. Pin both halves so neither can rot + # silently. + fakebin=$(fm_fakebin "$TMP_ROOT/anc-claude") + cat > "$fakebin/ps" <<'SH' +#!/usr/bin/env bash +case "$*" in + *"comm="*) printf '%s\n' agy; exit 0 ;; + *"args="*) printf '%s\n' 'agy --prompt-interactive hi'; exit 0 ;; +esac +exit 1 +SH + chmod +x "$fakebin/ps" + out=$(CLAUDECODE=1 PATH="$fakebin:$PATH" "$HARNESS") + [ "$out" = agy ] \ + || fail "a structural agy ancestor must outrank an inherited CLAUDECODE, got '$out'" pass "fm-harness.sh: no inherited launcher marker claims the agy identity" } diff --git a/tests/fm-cursor-harness.test.sh b/tests/fm-cursor-harness.test.sh index cbdb3047a91..4c59d1fdfbe 100755 --- a/tests/fm-cursor-harness.test.sh +++ b/tests/fm-cursor-harness.test.sh @@ -16,7 +16,9 @@ # 2. An unrelated `node`/`agent` pane classifies `other`, which the liveness # callers fold into `ambiguous` - NEVER `dead`. # 3. Cursor's env marker outranks an inherited CLAUDECODE, because cursor does -# not clear it and whichever marker is tested first wins. +# not clear it and whichever marker is tested first wins. That ordering +# settles the marker layer only: a nearer claude ancestor still outranks +# both (tests/fm-harness-precedence.test.sh owns that boundary). # 4. The transcript fold brackets a turn: a trailing turn_ended is idle, a # later role:user is busy, and an unresolvable binding is unknown. # 5. Cursor is a crewmate/scout adapter only and refuses a secondmate launch. @@ -175,25 +177,49 @@ test_tmux_classifies_cursor_pane_without_inferring_dead() { # --- 3. Detection ordering --------------------------------------------------- +# The marker ordering decides only when ancestry has nothing to say, so this +# case runs against a fake ps that reports a bash chain terminating at pid 1. +# Without it the suite would assert against whatever harness actually launched +# it, and the verdicts below would be about the runner rather than the ordering. test_cursor_marker_outranks_inherited_claudecode() { - local out + local out fakebin base_path + base_path=${FM_TEST_BASE_PATH:-/usr/bin:/bin:/usr/sbin:/sbin} + fakebin=$(fm_fakebin "$TMP_ROOT/marker-ordering") + fm_fake_blind_ancestry "$fakebin" # This is the exact hazard: cursor does NOT clear an inherited CLAUDECODE, so - # a cursor worker under a claude primary carries both markers. - out=$(CLAUDECODE=1 CURSOR_AGENT=1 "$HARNESS") + # a cursor session started by hand under a claude primary carries both markers. + out=$(PATH="$fakebin:$base_path" CLAUDECODE=1 CURSOR_AGENT=1 "$HARNESS") [ "$out" = cursor ] || fail "CLAUDECODE + CURSOR_AGENT must detect cursor, got '$out'" - out=$(CLAUDECODE=1 CURSOR_INVOKED_AS=cursor-agent "$HARNESS") + out=$(PATH="$fakebin:$base_path" CLAUDECODE=1 CURSOR_INVOKED_AS=cursor-agent "$HARNESS") [ "$out" = cursor ] || fail "CLAUDECODE + CURSOR_INVOKED_AS must detect cursor, got '$out'" # Both cursor markers stand alone, and neither steals a plain claude session. - out=$(env -u CLAUDECODE CURSOR_AGENT=1 "$HARNESS") + out=$(env -u CLAUDECODE PATH="$fakebin:$base_path" CURSOR_AGENT=1 "$HARNESS") [ "$out" = cursor ] || fail "CURSOR_AGENT alone must detect cursor, got '$out'" - out=$(env -u CURSOR_AGENT -u CURSOR_INVOKED_AS CLAUDECODE=1 "$HARNESS") + out=$(env -u CURSOR_AGENT -u CURSOR_INVOKED_AS PATH="$fakebin:$base_path" \ + CLAUDECODE=1 "$HARNESS") [ "$out" = claude ] || fail "CLAUDECODE alone must still detect claude, got '$out'" # A CURSOR_* variable that is not the invocation identity proves nothing. - out=$(env -u CURSOR_AGENT CLAUDECODE=1 CURSOR_API_ENDPOINT=https://example \ + out=$(env -u CURSOR_AGENT PATH="$fakebin:$base_path" CLAUDECODE=1 \ + CURSOR_API_ENDPOINT=https://example \ CURSOR_INVOKED_AS=something-else "$HARNESS") [ "$out" = claude ] \ || fail "an unrelated CURSOR_* setting must not claim the cursor identity, got '$out'" - pass "fm-harness.sh: cursor's marker outranks an inherited CLAUDECODE" + # The ordering is a marker-layer tiebreak, not a licence to overrule the + # process tree: with a real cursor-agent ancestor the two agree, and with a + # real claude ancestor the retained cursor marker loses. + local tree_dir + tree_dir="$TMP_ROOT/marker-ordering-trees" + mkdir -p "$tree_dir" + cp "$(command -v bash)" "$tree_dir/cursor-agent" + cp "$(command -v bash)" "$tree_dir/claude" + out=$(env -u CLAUDECODE "$tree_dir/cursor-agent" -c \ + "r=\$(CURSOR_AGENT=1 \"$HARNESS\"); printf '%s' \"\$r\"") + [ "$out" = cursor ] || fail "a real cursor-agent ancestor must detect cursor, got '$out'" + out=$("$tree_dir/claude" -c \ + "r=\$(CLAUDECODE=1 CURSOR_AGENT=1 \"$HARNESS\"); printf '%s' \"\$r\"") + [ "$out" = claude ] \ + || fail "a retained CURSOR_AGENT must not rename a real claude ancestor, got '$out'" + pass "fm-harness.sh: cursor's marker outranks an inherited CLAUDECODE when ancestry is silent" } test_harness_ancestry_rejects_cursor_named_node_script() { diff --git a/tests/fm-gemini-harness.test.sh b/tests/fm-gemini-harness.test.sh index bca10321f62..c570413aee2 100644 --- a/tests/fm-gemini-harness.test.sh +++ b/tests/fm-gemini-harness.test.sh @@ -31,19 +31,22 @@ HARNESS="$ROOT/bin/fm-harness.sh" TMP_ROOT=$(fm_test_tmproot fm-gemini-harness) test_gemini_marker_outranks_inherited_claudecode() { - local out + local out fakebin base_path + base_path=${FM_TEST_BASE_PATH:-/usr/bin:/bin:/usr/sbin:/sbin} + fakebin=$(fm_fakebin "$TMP_ROOT/marker-ordering") + fm_fake_blind_ancestry "$fakebin" # This is the exact hazard: gemini does not clear an inherited CLAUDECODE, so # a gemini worker under a claude primary carries both markers at once. - out=$(CLAUDECODE=1 GEMINI_CLI=1 "$HARNESS") + out=$(PATH="$fakebin:$base_path" CLAUDECODE=1 GEMINI_CLI=1 "$HARNESS") [ "$out" = gemini ] || fail "CLAUDECODE + GEMINI_CLI must detect gemini, got '$out'" # Drive the two signals apart so the case above cannot go quietly vacuous: # each marker alone must still produce its own verdict. - out=$(env -u CLAUDECODE GEMINI_CLI=1 "$HARNESS") + out=$(env -u CLAUDECODE PATH="$fakebin:$base_path" GEMINI_CLI=1 "$HARNESS") [ "$out" = gemini ] || fail "GEMINI_CLI alone must detect gemini, got '$out'" - out=$(env -u GEMINI_CLI CLAUDECODE=1 "$HARNESS") + out=$(env -u GEMINI_CLI PATH="$fakebin:$base_path" CLAUDECODE=1 "$HARNESS") [ "$out" = claude ] || fail "CLAUDECODE alone must still detect claude, got '$out'" # Cursor's marker still outranks gemini's, preserving the documented order. - out=$(CURSOR_AGENT=1 GEMINI_CLI=1 "$HARNESS") + out=$(PATH="$fakebin:$base_path" CURSOR_AGENT=1 GEMINI_CLI=1 "$HARNESS") [ "$out" = cursor ] || fail "CURSOR_AGENT must still outrank GEMINI_CLI, got '$out'" pass "fm-harness.sh: gemini's marker outranks an inherited CLAUDECODE" } diff --git a/tests/fm-guard-stale-banner.test.sh b/tests/fm-guard-stale-banner.test.sh index 9115acd3f7d..5f06c15f739 100755 --- a/tests/fm-guard-stale-banner.test.sh +++ b/tests/fm-guard-stale-banner.test.sh @@ -857,11 +857,15 @@ test_extension_live_watcher_is_healthy_without_ownership_evidence() { # The cases above pin the model. This one takes the end-user path instead: no # FM_SUPERVISION_MODEL at all, so bin/fm-harness.sh must route a Pi primary to the # extension model on its own. Without that routing the tolerance would never reach -# a real Pi home. The foreign markers are cleared because fm-harness.sh tests them -# ahead of Pi, and the host running this suite may carry one. +# a real Pi home. Pinning Pi takes both halves of the evidence: the foreign markers +# are cleared because the host running this suite may carry one, and the ancestry +# walk is blinded because a structural ancestor of a different harness outranks the +# Pi marker, so the harness this suite was launched from would otherwise answer. test_pi_harness_routes_itself_to_the_extension_model() { - local dir home out pid harness + local dir home out pid harness blind local -a pi_env + blind=$(fm_fakebin "$TMP_ROOT/pi-routing-blind") + fm_fake_blind_ancestry "$blind" for harness in pi pi-signed; do pi_env=(PI_CODING_AGENT=true) [ "$harness" = pi ] || pi_env+=(FM_PI_HARNESS=pi-signed) @@ -873,6 +877,7 @@ test_pi_harness_routes_itself_to_the_extension_model() { touch "$home/state/.last-watcher-beat" out=$(env -u CLAUDECODE -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u GROK_AGENT -u FM_SUPERVISION_MODEL \ "${pi_env[@]}" \ + PATH="$blind:$PATH" \ FM_ROOT_OVERRIDE="$(case_root "$dir")" \ FM_HOME="$home" \ FM_GUARD_GRACE=999 \ diff --git a/tests/fm-harness-liveness-drift-live-e2e.test.sh b/tests/fm-harness-liveness-drift-live-e2e.test.sh index a885d7f975f..ce93335274d 100755 --- a/tests/fm-harness-liveness-drift-live-e2e.test.sh +++ b/tests/fm-harness-liveness-drift-live-e2e.test.sh @@ -1,15 +1,23 @@ #!/usr/bin/env bash # tests/fm-harness-liveness-drift-live-e2e.test.sh - default-on drift guard proving # every INSTALLED harness is still classified `alive` by the tmux liveness -# probe (bin/backends/tmux.sh). +# probe (bin/backends/tmux.sh) AND still identified by the harness-detection +# ancestry walk (bin/fm-harness.sh). # -# Why this file exists: liveness classification depends on how a harness names -# its own process, which is a surface the harness vendor controls and changes -# without notice. Claude Code began reporting its version string as its process -# name and became unattributable, which silently degraded supervision. A -# regression that only a real harness release can cause needs a check that runs -# real harnesses; a stubbed agent cannot see it, and neither can a table of -# names transcribed from a previous release. +# Why this file exists: both verdicts depend on how a harness names its own +# process, which is a surface the harness vendor controls and changes without +# notice. Claude Code began reporting its version string as its process name and +# became unattributable, which silently degraded supervision. A regression that +# only a real harness release can cause needs a check that runs real harnesses; +# a stubbed agent cannot see it, and neither can a table of names transcribed +# from a previous release. +# +# Detection carries the same exposure for a second reason: a structural ancestor +# now outranks an environment marker (bin/fm-harness.sh owns that boundary), so +# a harness whose process name stops matching no longer merely loses a fast +# path - the walk keeps climbing and can reach a DIFFERENT harness that really +# is further up the tree. This guard is what catches that at the release that +# causes it. # # Each harness is launched bare, with no prompt, so this consumes no model # tokens. The launch uses whatever credentials the harness already has; an @@ -139,6 +147,82 @@ for harness in claude codex opencode pi pi-signed grok kimi cursor muse; do note "$harness $version: title='$title' foreground=[$comms]" pass "harness liveness: $harness $version classifies alive" + + # Detection: ask the ancestry walk what it makes of this real harness process. + # Both Pi identities share one launcher name, so ancestry can only ever prove + # the family; only the launch-boundary marker selects the signed identity. + expect_harness=$harness + [ "$harness" = pi-signed ] && expect_harness=pi + pane_pid=$("$REAL_TMUX" -L "$SOCKET" display-message -p -t "$target" '#{pane_pid}' 2>/dev/null | tr -d ' ') + [ -n "$pane_pid" ] || fail "$harness ($version): could not read the pane pid for the detection probe" + # Probe from BELOW the pane process, not the pane process alone. The shipped + # guarantee is a strength claim: detect_own hands an args-strength verdict back + # to a retained foreign marker, so a harness is only protected where the walk + # reaches it at comm strength. A harness that ships as a thin interpreter shim + # spawning its native binary as a CHILD is args strength from the pane process + # and comm strength from below that child - which is where firstmate's own + # detection actually runs, as a tool subprocess. Probing only the pane would + # therefore pass on evidence the guarantee does not rest on, and would keep + # passing if a release stopped spawning the native child at all. + # + # The vantage set is the UPWARD path from the deepest foreground descendant, not + # every descendant in the subtree, because harness_ancestry only ever climbs: a + # sibling branch is a vantage firstmate's own detection can never occupy. + # Restricting the deepest descendant to the pane tty's foreground process group + # keeps a process left running in the background out of the selection as well. + # + # The reject-other-harness cross-check below judges COMM-strength vantages only. + # An args-strength verdict is path-ambiguous by construction: harness_ancestry's + # bare-interpreter branch matches a harness name anywhere in the script path, so a + # harness-spawned MCP server running as `node <home>/.claude/mcp/<server>.js` + # answers `args claude` purely from the .claude path component, and such a server + # is normally a child of the agent binary rather than a sibling of it, so it can + # be the deepest descendant and sit ON this path. That ambiguity is the sole source + # of the false failure; a comm-strength verdict carries the real process name and + # cannot be produced that way. The comm-strength REQUIREMENT is unchanged - some + # vantage on the path must still name the expected harness at comm strength, + # because detect_own hands an args-strength verdict straight back to a retained + # foreign marker. + # The native binary can take a moment to appear, so poll for it. + pane_tty=$("$REAL_TMUX" -L "$SOCKET" display-message -p -t "$target" '#{pane_tty}' 2>/dev/null | tr -d ' ') + verdicts= + for _ in $(seq 1 150); do + fg_pids= + if [ -n "$pane_tty" ]; then + fg_pids=$(LC_ALL=C ps -t "${pane_tty#/dev/}" -o pid=,pgid=,tpgid= 2>/dev/null \ + | while read -r fg_pid fg_pgid fg_tpgid; do + [ -n "$fg_pid" ] || continue + [ "$fg_pgid" = "$fg_tpgid" ] || continue + printf '%s ' "$fg_pid" + done) + fi + # shellcheck disable=SC2086 # deliberate: the foreground pids are separate arguments + verdicts=$("$ROOT/bin/fm-harness.sh" ancestry-descent "$pane_pid" $fg_pids 2>/dev/null || true) + case "$verdicts" in *"comm $expect_harness"*) break ;; esac + sleep 0.2 + done + + drift_context="Observed process title '$title'; observed foreground process names [$comms]; observed ancestry verdicts [$(printf '%s' "$verdicts" | tr '\n' ';')]." + + [ -n "$verdicts" ] || fail \ + "DETECTION DRIFT: $harness $version is running but the ancestry walk reports nothing from the pane process or any vantage below it, so firstmate cannot identify this session at all. $drift_context Teach bin/fm-harness.sh's harness_ancestry the name this release actually reports." + + SAW_COMM=0 + while read -r strength named; do + [ -n "$strength" ] || continue + [ "$strength" = comm ] || continue + [ "$named" = "$expect_harness" ] || fail \ + "DETECTION DRIFT: $harness $version is running but a comm-strength vantage point on the upward path through its own session resolves to '$named', not '$expect_harness'. bin/fm-harness.sh lets a structural ancestor outrank an environment marker, so an unmatched process name can resolve to a DIFFERENT harness further up the tree instead of merely losing a fast path. $drift_context Teach bin/fm-harness.sh's harness_ancestry the name this release actually reports." + SAW_COMM=1 + done <<EOF +$verdicts +EOF + + [ "$SAW_COMM" = 1 ] || fail \ + "DETECTION DRIFT: $harness $version is identified only at interpreter-args strength, from no vantage point on the upward path through its session at comm strength. detect_own hands an args-strength verdict back to a retained foreign marker, so a stale CLAUDECODE would silently rename this session even though this guard sees the right identity. $drift_context Restore a process name bin/fm-harness.sh's harness_ancestry can match structurally, or teach it the name this release reports." + + note "$harness $version: ancestry verdicts=[$(printf '%s' "$verdicts" | tr '\n' ';')]" + pass "harness detection: $harness $version is identified by the ancestry walk at comm strength" CHECKED=$((CHECKED + 1)) done diff --git a/tests/fm-harness-precedence.test.sh b/tests/fm-harness-precedence.test.sh new file mode 100755 index 00000000000..0d4999984a3 --- /dev/null +++ b/tests/fm-harness-precedence.test.sh @@ -0,0 +1,761 @@ +#!/usr/bin/env bash +# Behavior tests for bin/fm-harness.sh's marker-vs-ancestry precedence boundary, +# and for the supervision protocol session start selects from it. +# +# The bug this pins: a Codex session started from an environment that had +# retained CLAUDECODE=1 detected as claude, because a verified marker outranked +# ancestry unconditionally. Session start then emitted Claude's Stop-owned +# supervision protocol to a Codex primary, which blocked every turn end. +# +# Every case drives the two evidence layers APART deliberately and asserts each +# one alone as well as the combination, so no case can pass vacuously if a layer +# silently stops working: +# marker alone - ancestry blinded by a fake ps, proving the marker is live +# and is what the old precedence would have returned. +# ancestry alone - marker cleared, proving the ancestry signal is live. +# both together - the precedence verdict this file exists to pin. +# The fake ps blinds only the ancestry walk, and the suite proves that rather +# than assuming it. fm-harness.sh reads process ancestry through ps alone; the +# one source it does not read through ps is the Cursor argv[0] probe, which on +# Linux reads /proc directly and on macOS falls back to ps. Either way it +# resolves the harmless real path of a bash-named process, which +# fm_cursor_process_matches rejects. The no-marker case below asserts `unknown` +# under the fake ps on whichever platform the run happens on, and it is exactly +# that case that fails if the blinding ever leaks a real ancestor through. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +# This suite states the markers it means to test in every case. Drop the ambient +# ones so a verdict never depends on which harness launched the suite. +unset CLAUDECODE PI_CODING_AGENT FM_PI_HARNESS GROK_AGENT CURSOR_AGENT CURSOR_INVOKED_AS + +HARNESS="$ROOT/bin/fm-harness.sh" +RENDER="$ROOT/bin/fm-supervision-instructions.sh" +TMP_ROOT=$(fm_test_tmproot fm-harness-precedence) +BASE_PATH=${FM_TEST_BASE_PATH:-/usr/bin:/bin:/usr/sbin:/sbin} + +# A real process named after a harness, asked for its verdict from a child. +# The command substitution around the probe is load-bearing: a bare `-c <cmd>` +# lets the shell exec the probe in place, which REPLACES the harness-named +# process the walk is supposed to find. +under_process() { # <named-executable> [VAR=VAL ...] + local bin=$1 + shift + env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS "$@" \ + "$bin" -c "r=\$(\"$HARNESS\"); printf '%s' \"\$r\"" +} + +# A fake ps that reports a bash ancestor terminating at pid 1, so the ancestry +# layer proves nothing and only the marker layer can answer. +blind_ancestry_bin() { # <dir> + local fakebin + fakebin=$(fm_fakebin "$1") + cat > "$fakebin/ps" <<'SH' +#!/usr/bin/env bash +case "$*" in + *'ppid='*) printf '%s\n' 1 ;; + *) printf '%s\n' bash ;; +esac +SH + chmod +x "$fakebin/ps" + printf '%s\n' "$fakebin" +} + +# A fake ps that models a PID NAMESPACE: every process reports bash with ppid 1, +# and pid 1 reports whatever FM_TEST_PID1_COMM names. This is what a harness +# looks like from inside a container or `codex sandbox`, where the harness is +# pid 1 of its own namespace rather than a child of a shell. +namespace_ancestry_bin() { # <dir> + local fakebin + fakebin=$(fm_fakebin "$1") + cat > "$fakebin/ps" <<'SH' +#!/usr/bin/env bash +pid= +prev= +for a in "$@"; do + [ "$prev" = -p ] && pid=$a + prev=$a +done +if [ "$pid" = 1 ]; then + comm=${FM_TEST_PID1_COMM:-init} + ppid=0 +else + comm=bash + ppid=1 +fi +case "$*" in + *'ppid='*) printf '%s\n' "$ppid" ;; + *) printf '%s\n' "$comm" ;; +esac +SH + chmod +x "$fakebin/ps" + printf '%s\n' "$fakebin" +} + +# Run the harness script under a fake ps, with the ambient markers dropped so +# each case states its own. +under_fake_ps() { # <fakebin> <VAR=VAL ...> -- [harness args] + local fakebin=$1 + shift + local -a assignments=() + while [ "$#" -gt 0 ] && [ "$1" != -- ]; do + assignments+=("$1") + shift + done + [ "${1:-}" = -- ] && shift + env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS "${assignments[@]}" \ + PATH="$fakebin:$BASE_PATH" "$HARNESS" "$@" +} + +with_blind_ancestry() { # <fakebin> [VAR=VAL ...] + local fakebin=$1 + shift + env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS "$@" \ + PATH="$fakebin:$BASE_PATH" "$HARNESS" +} + +named_bin() { # <dir> <name> + mkdir -p "$1" + cp "$(command -v bash)" "$1/$2" + printf '%s\n' "$1/$2" +} + +# --- 1. A foreign marker never renames a markerless harness ----------------- + +# codex, opencode, kimi, muse, and agy publish no identity marker, so before +# this boundary existed ANY retained marker renamed them outright. This is the +# reported live failure, generalized to every markerless adapter and to both +# foreign markers that can be retained. +test_markerless_ancestry_outranks_foreign_marker() { + local dir fakebin bin got name + dir="$TMP_ROOT/markerless" + fakebin=$(blind_ancestry_bin "$dir/blind") + for name in codex opencode kimi muse-bin-0.1.0 agy; do + bin=$(named_bin "$dir/$name-tree" "$name") + local expect=$name + case "$name" in muse-bin-*) expect=muse ;; esac + + got=$(under_process "$bin") + [ "$got" = "$expect" ] \ + || fail "$name ancestry alone resolved '$got', expected $expect (the ancestry signal is not live)" + + got=$(with_blind_ancestry "$fakebin" CLAUDECODE=1) + [ "$got" = claude ] \ + || fail "an inherited CLAUDECODE alone resolved '$got', expected claude (the marker signal is not live)" + + got=$(under_process "$bin" CLAUDECODE=1) + [ "$got" = "$expect" ] \ + || fail "$name ancestry with an inherited CLAUDECODE resolved '$got', expected $expect" + + got=$(under_process "$bin" CURSOR_AGENT=1) + [ "$got" = "$expect" ] \ + || fail "$name ancestry with an inherited CURSOR_AGENT resolved '$got', expected $expect" + done + pass "a markerless harness keeps its identity under an inherited foreign marker" +} + +# --- 2. A genuine harness in its own process tree still wins ---------------- + +test_genuine_marker_and_ancestry_agree() { + local dir bin got + dir="$TMP_ROOT/genuine" + + bin=$(named_bin "$dir/claude-tree" claude) + got=$(under_process "$bin" CLAUDECODE=1) + [ "$got" = claude ] || fail "a genuine claude session resolved '$got', expected claude" + + bin=$(named_bin "$dir/cursor-tree" cursor-agent) + got=$(under_process "$bin" CURSOR_AGENT=1) + [ "$got" = cursor ] || fail "a genuine cursor session resolved '$got', expected cursor" + got=$(under_process "$bin" CURSOR_INVOKED_AS=cursor-agent) + [ "$got" = cursor ] || fail "a genuine cursor session (launcher marker) resolved '$got', expected cursor" + + bin=$(named_bin "$dir/grok-tree" grok) + got=$(under_process "$bin" GROK_AGENT=1) + [ "$got" = grok ] || fail "a genuine grok session resolved '$got', expected grok" + # grok 1.0.0 hook processes carry no GROK_AGENT at all, so ancestry alone must + # still answer for them. + got=$(under_process "$bin") + [ "$got" = grok ] || fail "an unmarked grok hook process resolved '$got', expected grok" + + pass "a harness that publishes a marker inside its own process tree is unchanged" +} + +# Cursor is the case that motivated the pre-existing marker ordering: a cursor +# session started by hand under a claude primary carries BOTH markers. Ancestry +# is silent about which owns the tree there, so the ordering still decides. +test_cursor_ordering_still_decides_when_ancestry_is_silent() { + local fakebin got + fakebin=$(blind_ancestry_bin "$TMP_ROOT/cursor-ordering") + got=$(with_blind_ancestry "$fakebin" CLAUDECODE=1 CURSOR_AGENT=1) + [ "$got" = cursor ] || fail "both markers with no ancestry resolved '$got', expected cursor" + got=$(with_blind_ancestry "$fakebin" CLAUDECODE=1 CURSOR_INVOKED_AS=cursor-agent) + [ "$got" = cursor ] || fail "both markers (launcher form) with no ancestry resolved '$got', expected cursor" + got=$(with_blind_ancestry "$fakebin" CLAUDECODE=1) + [ "$got" = claude ] || fail "CLAUDECODE with no ancestry resolved '$got', expected claude" + got=$(with_blind_ancestry "$fakebin") + [ "$got" = unknown ] \ + || fail "no marker and no ancestry resolved '$got', expected unknown" + pass "with ancestry silent, the marker layer and its cursor-first ordering still decide" +} + +# The symmetric half of the same bug: cursor-agent's marker reaches a nested +# claude worker's environment, and the nearer claude ancestor must win. +test_retained_cursor_marker_does_not_rename_a_nested_claude() { + local bin got + bin=$(named_bin "$TMP_ROOT/nested-claude" claude) + got=$(under_process "$bin" CURSOR_AGENT=1 CLAUDECODE=1) + [ "$got" = claude ] \ + || fail "a claude tree carrying a retained CURSOR_AGENT resolved '$got', expected claude" + got=$(under_process "$bin" CURSOR_INVOKED_AS=cursor-agent CLAUDECODE=1) + [ "$got" = claude ] \ + || fail "a claude tree carrying a retained cursor launcher marker resolved '$got', expected claude" + pass "a retained cursor marker does not rename a nested claude worker" +} + +# --- 3. Pi keeps the marker's more specific identity ------------------------ + +# Both Pi identities share the launcher name, so ancestry can only prove the +# family. A marker that agrees on the family must keep its finer verdict rather +# than being flattened to pi by the ancestry walk. +test_pi_signed_survives_agreeing_ancestry() { + local bin got + bin=$(named_bin "$TMP_ROOT/pi-tree" pi) + got=$(under_process "$bin" PI_CODING_AGENT=true FM_PI_HARNESS=pi-signed) + [ "$got" = pi-signed ] || fail "signed Pi over pi ancestry resolved '$got', expected pi-signed" + got=$(under_process "$bin" PI_CODING_AGENT=true) + [ "$got" = pi ] || fail "plain Pi over pi ancestry resolved '$got', expected pi" + got=$(under_process "$bin") + [ "$got" = pi ] || fail "unmarked pi ancestry resolved '$got', expected pi" + + bin=$(named_bin "$TMP_ROOT/pi-signed-tree" pi-signed) + got=$(under_process "$bin" PI_CODING_AGENT=true FM_PI_HARNESS=pi-signed) + [ "$got" = pi-signed ] \ + || fail "signed Pi over shared signed-wrapper ancestry resolved '$got', expected pi-signed" + got=$(under_process "$bin") + [ "$got" = pi ] \ + || fail "unmarked signed-wrapper ancestry resolved '$got', expected pi" + pass "an agreeing marker keeps Pi's finer identity that ancestry cannot prove" +} + +# --- 4. The weakest ancestry signal does not outrank a marker --------------- + +# A bare interpreter matched only by a harness name inside the script path it +# was handed is the weakest inference in fm-harness.sh: any node process holding +# a harness-shaped path matches it. It answers when nothing else does, but it +# must not overturn a harness publishing its own identity. +test_interpreter_args_match_does_not_outrank_a_marker() { + local dir node script got + dir="$TMP_ROOT/weak-args" + node=$(named_bin "$dir" node) + script="$dir/codex-tool.sh" + cat > "$script" <<SH +r=\$("$HARNESS"); printf '%s' "\$r" +SH + + got=$(env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS "$node" "$script") + [ "$got" = codex ] \ + || fail "an unmarked interpreter holding a codex-shaped script path resolved '$got', expected codex" + + got=$(env -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS CLAUDECODE=1 "$node" "$script") + [ "$got" = claude ] \ + || fail "a published CLAUDECODE lost to a codex-shaped script path, resolving '$got'" + pass "an interpreter script-path match answers alone but never outranks a marker" +} + +# The real Codex install topology, modelled because the fix depends on it. codex +# ships as a `node` npm shim that spawns its native `codex` binary as a child and +# waits, so BOTH are in a tool subprocess's parent chain and the native name is +# the nearer one. Verified live on 2026-09-01 with codex-cli 0.152.0, whose pane +# foreground process names were [node codex]. What this case pins is that rule +# and nothing wider: a native harness binary nearer than an interpreter decides +# at comm strength, so the shim's own script path never gets to hand the verdict +# back to a retained marker. The strength assertion below is what keeps that +# non-vacuous - reaching the node shim instead would answer 'args codex'. +# A fixture cannot notice a vendor topology change; the opt-in live drift guard +# (tests/fm-harness-liveness-drift-live-e2e.test.sh) owns that. +test_native_child_of_an_interpreter_shim_decides_at_comm_strength() { + local dir node native probe entry got + dir="$TMP_ROOT/shim-topology" + node=$(named_bin "$dir" node) + native=$(named_bin "$dir/vendor" codex) + + # The probe forks so the command substitution's child is what asks, exactly as + # a tool subprocess of a real harness would. + probe="$dir/probe.sh" + cat > "$probe" <<'SH' +r=$("$FM_TEST_HARNESS" "$@") +printf '%s' "$r" +SH + # The shim SPAWNS its native binary and waits, so the node process stays alive + # above it and the walk meets the native binary first. + entry="$dir/codex-cli-entry.sh" + cat > "$entry" <<'SH' +"$FM_TEST_NATIVE" "$FM_TEST_PROBE" "$@" & +wait "$!" +SH + + # No arguments: the two cases that vary the environment or the subcommand call + # the shim entry point directly below, so this helper stays the plain no-marker + # launch. + run_shim() { + env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ + FM_TEST_HARNESS="$HARNESS" FM_TEST_NATIVE="$native" FM_TEST_PROBE="$probe" \ + "$node" "$entry" + } + + got=$(run_shim) + [ "$got" = codex ] \ + || fail "the shim topology without a marker resolved '$got', expected codex" + + got=$(env CLAUDECODE=1 FM_TEST_HARNESS="$HARNESS" FM_TEST_NATIVE="$native" \ + FM_TEST_PROBE="$probe" "$node" "$entry") + [ "$got" = codex ] \ + || fail "the real Codex shim topology with a retained CLAUDECODE resolved '$got', expected codex" + + got=$(env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ + FM_TEST_HARNESS="$HARNESS" FM_TEST_NATIVE="$native" FM_TEST_PROBE="$probe" \ + "$node" "$entry" ancestry) + [ "$got" = "comm codex" ] \ + || fail "the native child must decide at comm strength, got '$got'" + pass "a native harness binary under an interpreter shim decides at comm strength" +} + +# --- 5. A harness that is pid 1 of its own namespace ------------------------ + +# The walk used to stop as soon as the NEXT pid was 1, on the assumption that +# pid 1 is always init. Inside a PID namespace that assumption inverts: the +# harness itself is pid 1, so the one process that proves who owns the tree was +# never examined and a retained marker won by default. Verified against the real +# installed Codex, which runs as pid 1 under `codex sandbox`. +test_harness_at_namespace_pid1_is_examined() { + local fakebin got + fakebin=$(namespace_ancestry_bin "$TMP_ROOT/namespace-pid1") + + # Non-vacuity, both directions: with a host-shaped pid 1 the marker is the + # only evidence and must still answer, so the case below cannot pass by the + # ancestry layer simply matching everything. + got=$(under_fake_ps "$fakebin" FM_TEST_PID1_COMM=init CLAUDECODE=1 --) + [ "$got" = claude ] \ + || fail "a host-shaped pid 1 resolved '$got', expected claude (the marker layer is not live)" + + got=$(under_fake_ps "$fakebin" FM_TEST_PID1_COMM=codex --) + [ "$got" = codex ] \ + || fail "a Codex session at namespace pid 1 resolved '$got' with no marker, expected codex" + + got=$(under_fake_ps "$fakebin" FM_TEST_PID1_COMM=codex CLAUDECODE=1 --) + [ "$got" = codex ] \ + || fail "a Codex session at namespace pid 1 holding a retained CLAUDECODE resolved '$got', expected codex" + + got=$(under_fake_ps "$fakebin" FM_TEST_PID1_COMM=codex CLAUDECODE=1 -- ancestry) + [ "$got" = "comm codex" ] \ + || fail "the namespace pid 1 harness must decide at comm strength, got '$got'" + + pass "a harness that is pid 1 of its own namespace is examined, not skipped" +} + +# --- 6. The vantage point a probe asks from decides what strength it can see -- + +# The shipped guarantee is a strength claim, not just an identity one: detect_own +# hands an args-strength verdict back to a retained marker, so a harness is only +# protected where the walk reaches it at comm strength. Which strength is even +# REACHABLE depends on where the question is asked from. Under an interpreter +# shim the top of the session is the shim, whose own script path is args +# strength, while the native binary that carries comm strength is its CHILD. +# firstmate's own detect_own always runs from a tool subprocess below that child, +# so it sees comm; a guard that probed only the top of a real session would +# observe args, pass, and never notice a vendor release that stopped spawning the +# native child at all. `ancestry-descent` is what lets a probe ask from the same +# vantage a real session occupies, and this case pins that it reaches strictly +# further than the top-of-session probe does. +test_descent_probe_reaches_a_strength_the_top_of_session_cannot() { + local dir node native hold entry ready shim_pid got waited + dir="$TMP_ROOT/descent-vantage" + node=$(named_bin "$dir" node) + native=$(named_bin "$dir/vendor" codex) + ready="$dir/ready" + + # The native binary parks until the test releases it, so the whole topology is + # still standing while the probes run. + hold="$dir/hold.sh" + cat > "$hold" <<'SH' +touch "$FM_TEST_READY" +while [ -e "$FM_TEST_READY" ]; do sleep 0.05; done +SH + # The shim spawns its native binary and waits, so the node process stays alive + # ABOVE it exactly as the real Codex npm shim does. + entry="$dir/codex-cli-entry.sh" + cat > "$entry" <<'SH' +"$FM_TEST_NATIVE" "$FM_TEST_HOLD" & +wait "$!" +SH + + env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ + FM_TEST_NATIVE="$native" FM_TEST_HOLD="$hold" FM_TEST_READY="$ready" \ + "$node" "$entry" & + shim_pid=$! + + waited=0 + while [ ! -e "$ready" ] && [ "$waited" -lt 200 ]; do + sleep 0.05 + waited=$((waited + 1)) + done + [ -e "$ready" ] || { rm -f "$ready"; kill "$shim_pid" 2>/dev/null; fail "the shim fixture never reached its native child"; } + + # Non-vacuity: the top-of-session vantage really is limited to args strength + # here, which is the whole reason the descent probe has something to add. + got=$("$HARNESS" ancestry "$shim_pid") + [ "$got" = "args codex" ] \ + || fail "the shim's own vantage should see only 'args codex', got '$got'; the descent case proves nothing if the top of the session already reaches comm strength" + + got=$("$HARNESS" ancestry-descent "$shim_pid") + case "$got" in + *"comm codex"*) ;; + *) fail "the descent probe did not reach the native child at comm strength, got '$got'" ;; + esac + + # No vantage point inside the session may name a DIFFERENT harness, or a guard + # built on this probe would accept a tree it should have rejected. + while read -r strength named; do + [ -n "$strength" ] || continue + [ "$named" = codex ] \ + || fail "a vantage point inside the codex fixture reported '$strength $named'" + done <<EOF +$got +EOF + + rm -f "$ready" + wait "$shim_pid" 2>/dev/null || true + pass "the descent probe reaches comm strength where the top-of-session probe sees only args" +} + +# The other half of the vantage question: which vantages a probe must NOT ask +# from. harness_ancestry only ever climbs, so firstmate's own detection can never +# occupy a SIBLING branch of the process that runs it. A harness routinely spawns +# such branches - an MCP server started as `node <home>/.claude/mcp/<server>.js` +# matches *claude* on its script path in the bare-interpreter branch of the walk - +# and a probe that reported every descendant would answer a foreign harness from a +# process no real tool subprocess can ask from. The descent probe asks only the +# vantages on the upward path from the deepest descendant, which is exactly the set +# detection itself can reach. +test_descent_probe_ignores_a_sibling_branch_the_walk_cannot_reach() { + local dir node native worker mcp_script block hold entry ready fifo + local shim_pid mcp_pid got waited + dir="$TMP_ROOT/descent-sibling" + node=$(named_bin "$dir" node) + native=$(named_bin "$dir/vendor" codex) + worker=$(named_bin "$dir/vendor" worker) + ready="$dir/ready" + fifo="$dir/fifo" + mkdir -p "$dir/.claude/mcp" + mkfifo "$fifo" + + # Both leaves park on a fifo nothing ever writes, so they hold their position in + # the tree without spawning children of their own and the depths stay fixed. + block="$dir/block.sh" + cat > "$block" <<'SH' +read -r _ < "$FM_TEST_FIFO" +SH + # The MCP server is the sibling branch: a bare interpreter whose script path + # carries a harness name it does not belong to. + mcp_script="$dir/.claude/mcp/foo.js" + cp "$block" "$mcp_script" + + # The native binary keeps a child of its own, so the deepest descendant is + # unambiguously on the codex branch rather than tied with the sibling. + hold="$dir/hold.sh" + cat > "$hold" <<'SH' +"$FM_TEST_WORKER" "$FM_TEST_BLOCK" & +printf '%s\n' "$!" > "$FM_TEST_DIR/worker.pid" +touch "$FM_TEST_READY" +wait +SH + entry="$dir/codex-cli-entry.sh" + cat > "$entry" <<'SH' +"$FM_TEST_NATIVE" "$FM_TEST_HOLD" & +printf '%s\n' "$!" > "$FM_TEST_DIR/native.pid" +"$FM_TEST_NODE" "$FM_TEST_MCP" & +printf '%s\n' "$!" > "$FM_TEST_DIR/mcp.pid" +wait +SH + + env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ + FM_TEST_DIR="$dir" FM_TEST_NODE="$node" FM_TEST_NATIVE="$native" \ + FM_TEST_WORKER="$worker" FM_TEST_HOLD="$hold" FM_TEST_BLOCK="$block" \ + FM_TEST_MCP="$mcp_script" FM_TEST_READY="$ready" FM_TEST_FIFO="$fifo" \ + "$node" "$entry" & + shim_pid=$! + + waited=0 + while { [ ! -e "$ready" ] || [ ! -s "$dir/mcp.pid" ] || [ ! -s "$dir/worker.pid" ]; } \ + && [ "$waited" -lt 200 ]; do + sleep 0.05 + waited=$((waited + 1)) + done + release_sibling_fixture() { + kill "$(cat "$dir/worker.pid" 2>/dev/null)" "$(cat "$dir/mcp.pid" 2>/dev/null)" \ + "$(cat "$dir/native.pid" 2>/dev/null)" "$shim_pid" 2>/dev/null || true + wait "$shim_pid" 2>/dev/null || true + } + { [ -e "$ready" ] && [ -s "$dir/mcp.pid" ] && [ -s "$dir/worker.pid" ]; } \ + || { release_sibling_fixture; fail "the sibling fixture never reached both of its leaves"; } + mcp_pid=$(cat "$dir/mcp.pid") + + # Non-vacuity: the sibling really does answer a foreign harness when asked, so a + # probe that reported every descendant would have reported claude here. + got=$("$HARNESS" ancestry "$mcp_pid") + [ "$got" = "args claude" ] \ + || { release_sibling_fixture; fail "the sibling MCP process reported '$got', expected 'args claude'; this case proves nothing unless that branch really names a foreign harness"; } + + got=$("$HARNESS" ancestry-descent "$shim_pid") + case "$got" in + *"comm codex"*) ;; + *) release_sibling_fixture; fail "the descent probe did not reach the native child at comm strength, got '$got'" ;; + esac + while read -r strength named; do + [ -n "$strength" ] || continue + [ "$named" = codex ] \ + || { release_sibling_fixture; fail "the descent probe reported '$strength $named' from a sibling branch the ancestry walk can never climb through"; } + done <<EOF +$got +EOF + + release_sibling_fixture + pass "the descent probe reports no verdict from a sibling branch detection cannot reach" +} + +# The deeper shape the case above cannot reach, and the reason the live guard's +# reject-other-harness cross-check judges COMM-strength vantages only. A harness +# spawns its MCP servers from the AGENT BINARY, not from the npm shim, so the real +# Codex topology is shim -> native codex -> mcp server: the server inherits its +# parent's process group, passes the foreground filter, and is the deepest eligible +# descendant, which puts its own `args claude` vantage ON the descent path rather +# than off it. An args-strength verdict is path-ambiguous by construction - the +# bare-interpreter branch of the walk matches a harness name anywhere in the script +# path - so it is the comm-strength verdicts that carry a real process name and are +# the ones worth cross-checking. This case pins that the path still reaches +# `comm codex`, that every comm-strength vantage on it names codex, and that an +# `args claude` vantage really is present, which is what a cross-check applied to +# args strength would have rejected. +test_descent_probe_tolerates_an_args_only_foreign_verdict_at_the_deepest_vantage() { + local dir node native mcp_script hold entry ready fifo + local shim_pid mcp_pid got waited saw_comm + dir="$TMP_ROOT/descent-deep-mcp" + node=$(named_bin "$dir" node) + native=$(named_bin "$dir/vendor" codex) + ready="$dir/ready" + fifo="$dir/fifo" + mkdir -p "$dir/.claude/mcp" + mkfifo "$fifo" + + mcp_script="$dir/.claude/mcp/foo.js" + cat > "$mcp_script" <<'SH' +read -r _ < "$FM_TEST_FIFO" +SH + + # The native binary is what starts the MCP server, so the server sits BELOW it and + # is the deepest descendant of the whole tree. + hold="$dir/hold.sh" + cat > "$hold" <<'SH' +"$FM_TEST_NODE" "$FM_TEST_MCP" & +printf '%s\n' "$!" > "$FM_TEST_DIR/mcp.pid" +touch "$FM_TEST_READY" +wait +SH + entry="$dir/codex-cli-entry.sh" + cat > "$entry" <<'SH' +"$FM_TEST_NATIVE" "$FM_TEST_HOLD" & +printf '%s\n' "$!" > "$FM_TEST_DIR/native.pid" +wait +SH + + env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ + FM_TEST_DIR="$dir" FM_TEST_NODE="$node" FM_TEST_NATIVE="$native" \ + FM_TEST_HOLD="$hold" FM_TEST_MCP="$mcp_script" FM_TEST_READY="$ready" \ + FM_TEST_FIFO="$fifo" \ + "$node" "$entry" & + shim_pid=$! + + waited=0 + while { [ ! -e "$ready" ] || [ ! -s "$dir/mcp.pid" ]; } && [ "$waited" -lt 200 ]; do + sleep 0.05 + waited=$((waited + 1)) + done + release_deep_mcp_fixture() { + kill "$(cat "$dir/mcp.pid" 2>/dev/null)" "$(cat "$dir/native.pid" 2>/dev/null)" \ + "$shim_pid" 2>/dev/null || true + wait "$shim_pid" 2>/dev/null || true + } + { [ -e "$ready" ] && [ -s "$dir/mcp.pid" ]; } \ + || { release_deep_mcp_fixture; fail "the deep MCP fixture never reached its server process"; } + mcp_pid=$(cat "$dir/mcp.pid") + + got=$("$HARNESS" ancestry "$mcp_pid") + [ "$got" = "args claude" ] \ + || { release_deep_mcp_fixture; fail "the MCP server reported '$got', expected 'args claude'; this case proves nothing unless the deepest vantage really answers a foreign harness"; } + + got=$("$HARNESS" ancestry-descent "$shim_pid") + case "$got" in + *"args claude"*) ;; + *) release_deep_mcp_fixture; fail "the descent path did not include the MCP server's foreign args verdict, got '$got'; a cross-check restricted to comm strength is untested unless that vantage is on the path" ;; + esac + case "$got" in + *"comm codex"*) ;; + *) release_deep_mcp_fixture; fail "the descent probe did not reach the native binary at comm strength, got '$got'" ;; + esac + + saw_comm=0 + while read -r strength named; do + [ -n "$strength" ] || continue + [ "$strength" = comm ] || continue + [ "$named" = codex ] \ + || { release_deep_mcp_fixture; fail "a comm-strength vantage on the descent path reported '$named', expected codex"; } + saw_comm=1 + done <<EOF +$got +EOF + [ "$saw_comm" = 1 ] \ + || { release_deep_mcp_fixture; fail "no comm-strength vantage on the descent path, so the guard's strength requirement would reject this tree"; } + + release_deep_mcp_fixture + pass "a foreign args-only verdict at the deepest vantage leaves the comm-strength identity intact" +} + +# Two equally deep foreground leaves must not let ps ordering decide whether the +# chosen path reaches comm strength. The foreign MCP interpreter is spawned first +# in one pass and the native codex binary first in the other; both must resolve to +# the native leaf while the single-path shape remains intact. +test_descent_probe_prefers_comm_strength_when_deepest_leaves_tie() { + local order dir node native mcp_script block entry ready fifo + local shim_pid mcp_pid native_pid got waited + for order in mcp-first native-first; do + dir="$TMP_ROOT/descent-equal-$order" + node=$(named_bin "$dir" node) + native=$(named_bin "$dir/vendor" codex) + ready="$dir/ready" + fifo="$dir/fifo" + mkdir -p "$dir/.claude/mcp" + mkfifo "$fifo" + + block="$dir/block.sh" + cat > "$block" <<'SH' +read -r _ < "$FM_TEST_FIFO" +SH + mcp_script="$dir/.claude/mcp/foo.js" + cp "$block" "$mcp_script" + entry="$dir/codex-cli-entry.sh" + cat > "$entry" <<'SH' +if [ "$FM_TEST_ORDER" = mcp-first ]; then + "$FM_TEST_NODE" "$FM_TEST_MCP" & + printf '%s\n' "$!" > "$FM_TEST_DIR/mcp.pid" + "$FM_TEST_NATIVE" "$FM_TEST_BLOCK" & + printf '%s\n' "$!" > "$FM_TEST_DIR/native.pid" +else + "$FM_TEST_NATIVE" "$FM_TEST_BLOCK" & + printf '%s\n' "$!" > "$FM_TEST_DIR/native.pid" + "$FM_TEST_NODE" "$FM_TEST_MCP" & + printf '%s\n' "$!" > "$FM_TEST_DIR/mcp.pid" +fi +touch "$FM_TEST_READY" +wait +SH + + env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ + FM_TEST_ORDER="$order" FM_TEST_DIR="$dir" FM_TEST_NODE="$node" \ + FM_TEST_NATIVE="$native" FM_TEST_BLOCK="$block" FM_TEST_MCP="$mcp_script" \ + FM_TEST_READY="$ready" FM_TEST_FIFO="$fifo" "$node" "$entry" & + shim_pid=$! + + waited=0 + while { [ ! -e "$ready" ] || [ ! -s "$dir/mcp.pid" ] || [ ! -s "$dir/native.pid" ]; } \ + && [ "$waited" -lt 200 ]; do + sleep 0.05 + waited=$((waited + 1)) + done + mcp_pid=$(cat "$dir/mcp.pid" 2>/dev/null || true) + native_pid=$(cat "$dir/native.pid" 2>/dev/null || true) + release_equal_depth_fixture() { + kill "$mcp_pid" "$native_pid" "$shim_pid" 2>/dev/null || true + wait "$shim_pid" 2>/dev/null || true + } + { [ -e "$ready" ] && [ -n "$mcp_pid" ] && [ -n "$native_pid" ]; } \ + || { release_equal_depth_fixture; fail "the $order equal-depth fixture never reached both leaves"; } + + got=$("$HARNESS" ancestry "$mcp_pid") + [ "$got" = "args claude" ] \ + || { release_equal_depth_fixture; fail "the $order MCP leaf reported '$got', expected 'args claude'"; } + got=$("$HARNESS" ancestry "$native_pid") + [ "$got" = "comm codex" ] \ + || { release_equal_depth_fixture; fail "the $order native leaf reported '$got', expected 'comm codex'"; } + + got=$("$HARNESS" ancestry-descent "$shim_pid" "$mcp_pid" "$native_pid") + case "$got" in + "comm codex"*) ;; + *) release_equal_depth_fixture; fail "the $order equal-depth tie did not choose the comm-strength native leaf, got '$got'" ;; + esac + case "$got" in + *"args claude"*) release_equal_depth_fixture; fail "the $order equal-depth tie chose the foreign args-strength leaf" ;; + esac + + release_equal_depth_fixture + done + pass "equal-depth descent ties prefer the comm-strength leaf regardless of spawn order" +} + +# --- 7. Session start's supervision protocol follows the corrected verdict --- + +# The consequence the captain actually hit: the wrong verdict emitted Claude's +# Stop-owned protocol to a Codex primary, so every turn end was blocked for +# missing Claude recovery. +test_supervision_protocol_follows_corrected_verdict() { + local dir home fakebin bin got + dir="$TMP_ROOT/supervision" + home="$dir/home" + mkdir -p "$home/state" "$home/config" + bin=$(named_bin "$dir/codex-tree" codex) + fakebin=$(blind_ancestry_bin "$dir/blind") + + got=$(env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS CLAUDECODE=1 FM_HOME="$home" \ + PATH="$fakebin:$BASE_PATH" "$RENDER") + assert_contains "$got" "primary harness: claude" \ + "with ancestry blinded, the retained marker must still render claude (the case is otherwise vacuous)" + + got=$(env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS CLAUDECODE=1 FM_HOME="$home" \ + "$bin" -c "r=\$(\"$RENDER\"); printf '%s' \"\$r\"") + assert_contains "$got" "primary harness: codex" \ + "a Codex primary carrying a retained CLAUDECODE did not render the Codex protocol" + assert_contains "$got" "Mode: Codex foreground checkpoint." \ + "the rendered block is not Codex's foreground-checkpoint protocol" + assert_not_contains "$got" "Mode: Claude Stop-hook-owned supervision." \ + "the rendered block still carries Claude's Stop-owned protocol" + pass "session start renders the Codex protocol for a Codex primary holding a retained CLAUDECODE" +} + +test_markerless_ancestry_outranks_foreign_marker +test_genuine_marker_and_ancestry_agree +test_cursor_ordering_still_decides_when_ancestry_is_silent +test_retained_cursor_marker_does_not_rename_a_nested_claude +test_pi_signed_survives_agreeing_ancestry +test_interpreter_args_match_does_not_outrank_a_marker +test_native_child_of_an_interpreter_shim_decides_at_comm_strength +test_harness_at_namespace_pid1_is_examined +test_descent_probe_reaches_a_strength_the_top_of_session_cannot +test_descent_probe_ignores_a_sibling_branch_the_walk_cannot_reach +test_descent_probe_tolerates_an_args_only_foreign_verdict_at_the_deepest_vantage +test_descent_probe_prefers_comm_strength_when_deepest_leaves_tie +test_supervision_protocol_follows_corrected_verdict diff --git a/tests/fm-kimi-harness.test.sh b/tests/fm-kimi-harness.test.sh index 5235cb0d15a..518a614b7eb 100755 --- a/tests/fm-kimi-harness.test.sh +++ b/tests/fm-kimi-harness.test.sh @@ -5,10 +5,11 @@ set -u # shellcheck source=tests/lib.sh . "$(dirname "${BASH_SOURCE[0]}")/lib.sh" -# bin/fm-harness.sh checks verified ENV markers before ancestry. A suite run -# from inside Cursor, Claude, Pi, or Grok inherits those markers, which outrank -# the fake ancestry the detection cases set up. Drop the ambient markers so the -# asserted verdict does not depend on which harness launched the suite. +# bin/fm-harness.sh answers from environment markers and process ancestry. A +# suite run from inside Cursor, Claude, Pi, or Grok inherits those markers and +# its own real ancestry, either of which can decide a case the detection cases +# meant to control. Drop the ambient markers so the asserted verdict does not +# depend on which harness launched the suite. unset CLAUDECODE PI_CODING_AGENT FM_PI_HARNESS GROK_AGENT CURSOR_AGENT CURSOR_INVOKED_AS SPAWN="$ROOT/bin/fm-spawn.sh" @@ -552,10 +553,13 @@ SH -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u GEMINI_CLI \ PATH="$fakebin:$BASE_PATH" FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh") [ "$out" = kimi ] || fail "kimi ancestry detection returned '$out'" + # Kimi publishes no identity marker, so an inherited CLAUDECODE used to rename + # it outright. A structural kimi ancestor now outranks that marker; + # tests/fm-harness-precedence.test.sh owns the general boundary. out=$(env -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u GEMINI_CLI \ CLAUDECODE=1 PATH="$fakebin:$BASE_PATH" FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh") - [ "$out" = claude ] || fail "verified env-marker precedence changed, got '$out'" - pass "fm-harness: markerless kimi is detected by ancestry after env-marker precedence" + [ "$out" = kimi ] || fail "an inherited CLAUDECODE renamed markerless kimi, got '$out'" + pass "fm-harness: markerless kimi keeps its ancestry identity under an inherited marker" } test_kimi_session_lock_identity() { diff --git a/tests/fm-rovo-harness.test.sh b/tests/fm-rovo-harness.test.sh index bdfd17d2f26..021d6ff0c59 100644 --- a/tests/fm-rovo-harness.test.sh +++ b/tests/fm-rovo-harness.test.sh @@ -5,7 +5,9 @@ set -u # shellcheck source=tests/lib.sh . "$(dirname "${BASH_SOURCE[0]}")/lib.sh" -# bin/fm-harness.sh checks verified ENV markers before ancestry. A suite run +# bin/fm-harness.sh checks verified ENV markers before ancestry, but that +# ordering settles the marker layer only: a structural (comm-strength) +# ancestor of a different harness still outranks either marker. A suite run # from inside Cursor, Claude, Pi, or Grok inherits those markers, which outrank # the fake ancestry the detection cases set up. Drop the ambient markers so the # asserted verdict does not depend on which harness launched the suite. @@ -386,8 +388,15 @@ SH PATH="$fakebin:$BASE_PATH" FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh") [ "$out" = rovo ] || fail "rovo's ROVODEV_CLI marker did not outrank an inherited CLAUDECODE, got '$out'" + # CLAUDECODE alone, with no rovo marker, is a marker-layer question, not an + # ancestry one: blind the walk so the rovo-resolving fake ps above (needed + # for the markerless-ancestry and marker+ancestry cases) cannot also decide + # this assertion, matching the sibling-file pattern. + local blind_fakebin + blind_fakebin=$(fm_fakebin "$dir/blind-ancestry") + fm_fake_blind_ancestry "$blind_fakebin" out=$(env -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ - CLAUDECODE=1 PATH="$fakebin:$BASE_PATH" FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh") + CLAUDECODE=1 PATH="$blind_fakebin:$BASE_PATH" FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh") [ "$out" = claude ] || fail "verified env-marker precedence changed, got '$out'" pass "fm-harness: rovo's markers outrank an inherited CLAUDECODE, and markerless ancestry still resolves rovo" } diff --git a/tests/fm-secondmate-harness.test.sh b/tests/fm-secondmate-harness.test.sh index 9a6ebe2df1f..f4d9546e0b4 100755 --- a/tests/fm-secondmate-harness.test.sh +++ b/tests/fm-secondmate-harness.test.sh @@ -51,12 +51,11 @@ set -u . "$ROOT/bin/fm-config-inherit-lib.sh" # The harness-detection cases below fake `ps` so process ancestry is fully -# controlled, but bin/fm-harness.sh checks verified ENV markers before ancestry. -# A suite run from inside one of those harnesses inherits its marker, and the -# highest-precedence one wins over everything these cases set up: with an -# ambient CLAUDECODE=1, the pi-signed ancestry case resolves "claude". Drop the -# ambient markers so what this suite asserts does not depend on which harness it -# was launched from; every case states the marker it means to test. +# controlled, but bin/fm-harness.sh also reads verified ENV markers. A suite run +# from inside one of those harnesses inherits its marker, and it wins over +# everything these cases set up wherever ancestry is silent. Drop the ambient +# markers so what this suite asserts does not depend on which harness it was +# launched from; every case states the marker it means to test. unset CLAUDECODE PI_CODING_AGENT FM_PI_HARNESS GROK_AGENT CURSOR_AGENT CURSOR_INVOKED_AS BASE_PATH=${FM_TEST_BASE_PATH:-/usr/bin:/bin:/usr/sbin:/sbin} @@ -69,12 +68,20 @@ export FM_BACKEND=tmux # Several cases here resolve claude, so every spawn below pins a throwaway HOME # with an empty CLAUDE_CONFIG_DIR and puts node on the spawn's PATH; without the # first, this suite would write the developer's real ~/.claude.json. +# Dropping the ambient markers is only half the isolation: a structural ancestor +# outranks a marker, so a case that PINS detect_own with CLAUDECODE=1 also has to +# blind the ancestry walk, or the harness this suite was launched from answers +# instead of the pin. BLIND_BIN goes AFTER a case's own fakebin in PATH, so a +# fixture that deliberately supplies its own ps or a harness-named ancestor keeps +# it (tests/fm-harness-precedence.test.sh owns the precedence boundary itself). +BLIND_BIN=$(fm_fakebin "$TMP_ROOT/blind-ancestry") +fm_fake_blind_ancestry "$BLIND_BIN" # =========================================================================== # A) fm-harness.sh secondmate resolution + fallback (deterministic detect_own) # =========================================================================== -# detect_own is pinned to claude via CLAUDECODE=1 so the "fall through to own" -# cases are reproducible. Each row sets crew-harness / secondmate-harness in a +# detect_own is pinned to claude via CLAUDECODE=1 over a blinded ancestry walk so +# the "fall through to own" cases are reproducible on any host harness. Each row sets crew-harness / secondmate-harness in a # fresh config dir (a literal '-' means leave the file absent) and asserts BOTH # the secondmate resolution AND that crew resolution is unchanged (backward-compat). # <label>^<crew-harness>^<secondmate-harness>^<expect-secondmate>^<expect-crew> @@ -89,8 +96,8 @@ test_harness_resolution() { mkdir -p "$cfg" [ "$crew" = "-" ] || printf '%s\n' "$crew" > "$cfg/crew-harness" [ "$sm" = "-" ] || printf '%s\n' "$sm" > "$cfg/secondmate-harness" - got_sm=$(CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" secondmate) - got_crew=$(CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" crew) + got_sm=$(PATH="$BLIND_BIN:$BASE_PATH" CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" secondmate) + got_crew=$(PATH="$BLIND_BIN:$BASE_PATH" CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" crew) [ "$got_sm" = "$exp_sm" ] || fail "$label: secondmate resolved '$got_sm', expected '$exp_sm'" [ "$got_crew" = "$exp_crew" ] || fail "$label: crew resolved '$got_crew', expected '$exp_crew'" done <<'ROWS' @@ -146,9 +153,9 @@ test_secondmate_model_effort_tokens() { cfg="$case_dir/config" mkdir -p "$cfg" [ "$line" = ABSENT ] || printf '%b\n' "$line" > "$cfg/secondmate-harness" - got_h=$(CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" secondmate) - got_m=$(CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" secondmate-model) - got_e=$(CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" secondmate-effort) + got_h=$(PATH="$BLIND_BIN:$BASE_PATH" CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" secondmate) + got_m=$(PATH="$BLIND_BIN:$BASE_PATH" CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" secondmate-model) + got_e=$(PATH="$BLIND_BIN:$BASE_PATH" CLAUDECODE=1 FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh" secondmate-effort) [ "$got_h" = "$exp_harness" ] || fail "$label: harness resolved '$got_h', expected '$exp_harness'" [ "$got_m" = "$exp_model" ] || fail "$label: model resolved '$got_m', expected '$exp_model'" [ "$got_e" = "$exp_effort" ] || fail "$label: effort resolved '$got_e', expected '$exp_effort'" @@ -452,8 +459,8 @@ make_seeded_home() { # spawn_secondmate <world> <id> <home> [explicit-harness] # Runs fm-spawn.sh in secondmate mode. FM_ROOT is the real repo (so fm-harness.sh -# resolves), the primary config dir is <world>/home/config, and CLAUDECODE pins -# detect_own. stderr is discarded (the local-HEAD ff sync harmlessly skips a +# resolves), the primary config dir is <world>/home/config, and CLAUDECODE over a +# blinded ancestry walk pins detect_own. stderr is discarded (the local-HEAD ff sync harmlessly skips a # non-worktree home). Inspect <world>/home/state/<id>.meta and <home>/config after. spawn_secondmate() { local world=$1 id=$2 home=$3 harness=${4:-} fakebin @@ -464,7 +471,7 @@ spawn_secondmate() { local spawn_args=("$id" "$home") [ -n "$harness" ] && spawn_args+=("$harness") spawn_args+=(--secondmate) - PATH="$fakebin:$BASE_PATH" TMUX='' CLAUDECODE=1 \ + PATH="$fakebin:$BLIND_BIN:$BASE_PATH" TMUX='' CLAUDECODE=1 \ FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$world/home" HOME="$world/home/user-home" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$world/home/state" FM_DATA_OVERRIDE="$world/home/data" \ FM_PROJECTS_OVERRIDE="$world/home/projects" FM_CONFIG_OVERRIDE="$world/home/config" \ @@ -687,7 +694,7 @@ spawn_secondmate_capture() { mkdir -p "$world/home/state" "$world/home/data" fakebin=$(make_launch_capturing_tmux "$world/tmux-$id") : > "$launchlog" - PATH="$fakebin:$BASE_PATH" TMUX='' CLAUDECODE=1 \ + PATH="$fakebin:$BLIND_BIN:$BASE_PATH" TMUX='' CLAUDECODE=1 \ FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$world/home" HOME="$world/home/user-home" CLAUDE_CONFIG_DIR='' \ FM_STATE_OVERRIDE="$world/home/state" FM_DATA_OVERRIDE="$world/home/data" \ FM_PROJECTS_OVERRIDE="$world/home/projects" FM_CONFIG_OVERRIDE="$world/home/config" \ diff --git a/tests/fm-session-lock-ancestry.test.sh b/tests/fm-session-lock-ancestry.test.sh index 26086948526..dbf1e683f77 100755 --- a/tests/fm-session-lock-ancestry.test.sh +++ b/tests/fm-session-lock-ancestry.test.sh @@ -87,6 +87,52 @@ SH pass "session-lock: a version-named Claude Code session is identified from its install path and argv[0]" } +# A harness that is pid 1 of its own PID namespace - a container, or the +# `codex sandbox` this shape was verified in - used to be invisible: the walk +# stopped as soon as the NEXT pid was 1, so the one process that identifies the +# session was never examined and the session could not recognize its own lock. +test_harness_at_namespace_pid1_is_examined() { + local dir fakebin got + dir="$TMP_ROOT/namespace-pid1" + fakebin=$(fm_fakebin "$dir") + mkdir -p "$dir/state" + cat > "$fakebin/ps" <<'SH' +#!/usr/bin/env bash +set -u +field= pid= +while [ "$#" -gt 0 ]; do + case "$1" in + -o) field=$2; shift 2 ;; + -p) pid=$2; shift 2 ;; + *) shift ;; + esac +done +case "$pid:$field" in + 1:comm=) printf '%s\n' "${FM_TEST_PID1_COMM:-claude}" ;; + 1:args=) printf '%s\n' "${FM_TEST_PID1_COMM:-claude}" ;; + 1:ppid=) printf '%s\n' 0 ;; + *:comm=) printf '%s\n' bash ;; + *:args=) printf '%s\n' 'bash /repo/bin/fm-watch.sh' ;; + *:ppid=) printf '%s\n' 1 ;; +esac +SH + chmod +x "$fakebin/ps" + printf '1\n' > "$dir/state/.lock" + + # Non-vacuity: with a host-shaped pid 1 the same table must find nothing, so + # this case cannot pass by the walk matching everything it reaches. + if FM_TEST_PID1_COMM=systemd lib_eval "$fakebin" 'fm_harness_ancestry_pid' >/dev/null 2>&1; then + fail "a host-shaped pid 1 was read as a harness process" + fi + + got=$(lib_eval "$fakebin" 'fm_harness_ancestry_pid') \ + || fail "the harness at namespace pid 1 was not found in the ancestry at all" + [ "$got" = 1 ] || fail "ancestry resolved '$got', expected the namespace harness pid 1" + lib_eval "$fakebin" "fm_session_lock_owned_by_self '$dir/state'" \ + || fail "the session holding the lock at namespace pid 1 did not recognize itself as the owner" + pass "session-lock: a harness that is pid 1 of its own namespace is examined, not skipped" +} + test_ordinary_paths_are_never_harness_processes() { local dir fakebin shape dir="$TMP_ROOT/ordinary-paths" @@ -359,6 +405,7 @@ test_e2e_daemon_parented_version_named_session_keeps_its_lock() { } test_version_named_session_is_identified_on_both_platforms +test_harness_at_namespace_pid1_is_examined test_ordinary_paths_are_never_harness_processes test_harness_beyond_a_gap_never_owns_the_lock test_competing_version_named_session_is_seen_as_live diff --git a/tests/fm-session-start.test.sh b/tests/fm-session-start.test.sh index 260d44d553a..b3d6aba1458 100755 --- a/tests/fm-session-start.test.sh +++ b/tests/fm-session-start.test.sh @@ -215,10 +215,15 @@ make_fake_ps_claude() { make_fake_ps_harness() { local fakebin=$1 harness=$2 - cat > "$fakebin/ps" <<'SH' + cat > "$fakebin/ps" <<SH #!/usr/bin/env bash set -u -harness=${FM_FAKE_HARNESS:-claude} +# The ancestry this stub reports defaults to the harness the fixture was built +# for, so a case that builds a pi (or codex) fixture gets pi (or codex) ancestry +# without having to repeat it per run; FM_FAKE_HARNESS still overrides it. +harness=\${FM_FAKE_HARNESS:-$harness} +SH + cat >> "$fakebin/ps" <<'SH' pid= previous= for argument in "$@"; do diff --git a/tests/fm-sessionstart-nudge.test.sh b/tests/fm-sessionstart-nudge.test.sh index 5b6cf779bdc..03bedf85502 100755 --- a/tests/fm-sessionstart-nudge.test.sh +++ b/tests/fm-sessionstart-nudge.test.sh @@ -131,6 +131,40 @@ test_owned_lock_is_silent() { pass "fm-sessionstart-nudge: a lock holder in process ancestry is already run" } +# A firstmate running inside a PID namespace - a container, or `codex sandbox` - +# holds its home lock from a harness that IS pid 1, so this hook must recognize +# that owner. The old walk rejected a lock pid of 1 outright and stopped before +# comparing pid 1, so the hook nudged a session that had already run. +# A fake ps cannot reach this path: `kill -0` is a shell builtin gating the lock +# pid, and on a host `kill -0 1` fails for an unprivileged user, so the case +# needs a real namespace where pid 1 is this user's own process. +test_namespace_pid1_lock_holder_is_silent() { + local root="$TMP_ROOT/namespace-pid1" out status=0 + if ! command -v unshare >/dev/null 2>&1 \ + || ! unshare -rpf --mount-proc true >/dev/null 2>&1; then + printf '# skip: unprivileged PID namespaces are unavailable here, so the namespace pid 1 lock owner is unverified\n' + return 0 + fi + make_primary "$root" + + # Non-vacuity: inside the same namespace, with no lock at all, the hook must + # still produce its nudge, so silence below means the owner was recognized. + out=$(unshare -rpf --mount-proc bash -c \ + "FM_GATE_REFUSE_BYPASS=0 FM_ROOT_OVERRIDE='$root' FM_HOME='$root' '$NUDGE'; exit \$?") || status=$? + expect_code 0 "$status" "namespace nudge without a lock" + [ "$out" = "$NUDGE_LINE" ] \ + || fail "the namespace fixture did not nudge without a lock, so its silence proves nothing: $out" + + printf '1\n' > "$root/state/.lock" + status=0 + out=$(unshare -rpf --mount-proc bash -c \ + "FM_GATE_REFUSE_BYPASS=0 FM_ROOT_OVERRIDE='$root' FM_HOME='$root' '$NUDGE'; exit \$?") || status=$? + expect_code 0 "$status" "namespace pid 1 lock nudge" + [ -z "$out" ] \ + || fail "a lock held by the harness at namespace pid 1 was not recognized, got: $out" + pass "fm-sessionstart-nudge: a lock holder that is pid 1 of its own namespace is already run" +} + test_opencode_plugin_delivers_exact_nudge_once() { local root="$TMP_ROOT/opencode-primary" out status=0 make_primary "$root" @@ -1021,6 +1055,7 @@ test_unmarked_linked_worktree_is_silent test_linked_secondmate_primary_nudges test_missing_state_is_silent test_owned_lock_is_silent +test_namespace_pid1_lock_holder_is_silent test_opencode_plugin_delivers_exact_nudge_once test_run_startup_runs_the_full_digest test_run_clear_and_compact_reemit diff --git a/tests/fm-turnend-guard.test.sh b/tests/fm-turnend-guard.test.sh index 1fd43c8d6c3..f0245f6a827 100755 --- a/tests/fm-turnend-guard.test.sh +++ b/tests/fm-turnend-guard.test.sh @@ -22,6 +22,14 @@ fm_git_identity fmtest fmtest@example.invalid REQUIRED_REASON='watcher supervision needs Stop-owned automatic recovery; inspect the hook registration and startup status before ending the turn' AWAY_REQUIRED_REASON='Away mode owns watcher supervision' +# REQUIRED_REASON is the CLAUDE repair line, so the cases asserting it need +# detect_own to answer claude. CLAUDECODE=1 alone no longer pins that - a +# structural ancestor of a different harness outranks a marker - so those +# invocations also blind the ancestry walk. Only per-pid comm/args/ppid queries +# are answered here; watcher liveness still reaches the real ps. +BLIND_BIN=$(fm_fakebin "$TMP_ROOT/blind-ancestry") +fm_fake_blind_ancestry "$BLIND_BIN" + # --- PREDICATE: bin/fm-supervision-lib.sh ----------------------------------- test_predicate_healthy_no_inflight() { @@ -260,7 +268,7 @@ make_secondmate_linked_home_dir() { run_hook() { local dir=$1 stop_active=$2 home home=$(cd "$dir" && pwd) - printf '{"stop_hook_active":%s}' "$stop_active" | CLAUDECODE=1 FM_HOME="$home" bash "$dir/bin/fm-turnend-guard.sh" 2>&1 + printf '{"stop_hook_active":%s}' "$stop_active" | PATH="$BLIND_BIN:$PATH" CLAUDECODE=1 FM_HOME="$home" bash "$dir/bin/fm-turnend-guard.sh" 2>&1 } nonexistent_pid() { @@ -438,7 +446,7 @@ test_hook_blocks_from_fm_home_state() { home="$TMP_ROOT/hook-fm-home-op" mkdir -p "$home/state" : > "$home/state/task1.meta" - out=$(printf '{"stop_hook_active":false}' | CLAUDECODE=1 FM_HOME="$home" bash "$dir/bin/fm-turnend-guard.sh" 2>&1); status=$? + out=$(printf '{"stop_hook_active":false}' | PATH="$BLIND_BIN:$PATH" CLAUDECODE=1 FM_HOME="$home" bash "$dir/bin/fm-turnend-guard.sh" 2>&1); status=$? expect_code 2 "$status" "hook must inspect the active FM_HOME state dir" assert_contains "$out" "$REQUIRED_REASON" "block reason must contain the exact required instruction" pass "fm-turnend-guard: blocks from active FM_HOME state, not only repo-root state" @@ -497,7 +505,7 @@ test_hook_uses_state_override() { state="$TMP_ROOT/hook-state-override-active" mkdir -p "$home/state" "$state" : > "$state/task1.meta" - out=$(printf '{"stop_hook_active":false}' | CLAUDECODE=1 FM_HOME="$home" FM_STATE_OVERRIDE="$state" bash "$dir/bin/fm-turnend-guard.sh" 2>&1); status=$? + out=$(printf '{"stop_hook_active":false}' | PATH="$BLIND_BIN:$PATH" CLAUDECODE=1 FM_HOME="$home" FM_STATE_OVERRIDE="$state" bash "$dir/bin/fm-turnend-guard.sh" 2>&1); status=$? expect_code 2 "$status" "hook must let FM_STATE_OVERRIDE win over FM_HOME/state" assert_contains "$out" "$REQUIRED_REASON" "block reason must contain the exact required instruction" pass "fm-turnend-guard: uses FM_STATE_OVERRIDE ahead of FM_HOME/state" diff --git a/tests/fm-watcher-lock.test.sh b/tests/fm-watcher-lock.test.sh index 77fd4fbcca3..c0d37f0cf61 100755 --- a/tests/fm-watcher-lock.test.sh +++ b/tests/fm-watcher-lock.test.sh @@ -124,19 +124,25 @@ test_guard_warnings() { # warning follows it, and the guidance is repair-after-drain (never the # old conflicting "restart NOW first"). # (2) a fresh watcher and an empty queue: total silence. - local dir state err first banner_line queue_line pid identity + local dir state err first banner_line queue_line pid identity blind dir=$(make_case guard) + # The repair line the cases below assert is the CLAUDE one, so detect_own has to + # answer claude. A marker alone no longer pins that: a structural ancestor of a + # different harness outranks it, so the harness this suite was launched from + # would otherwise choose the wording. Blind the ancestry walk as well; every + # other ps query (watcher liveness below) still reaches the real ps. + blind=$(fm_fakebin "$dir/blind") + fm_fake_blind_ancestry "$blind" state="$dir/state" err="$dir/guard.err" # (1) watcher down (no beacon) + two in-flight tasks + a queued wake. # FM_ROOT_OVERRIDE points the worktree-tangle check at a non-git dir so it stays # inert here; this case is about the watcher-down banner, not the tangle guard. - # Pin Claude so the host test runner's harness ancestry cannot change this fixture. printf 'project=x\n' > "$state/task.meta" printf 'project=y\n' > "$state/task2.meta" append_wake "$state" heartbeat heartbeat heartbeat || fail "guard heartbeat append failed" - CLAUDECODE=1 PI_CODING_AGENT='' GROK_AGENT='' FM_ROOT_OVERRIDE="$dir" FM_STATE_OVERRIDE="$state" FM_GUARD_GRACE=1 "$ROOT/bin/fm-guard.sh" 2> "$err" >/dev/null || fail "guard failed" + PATH="$blind:$PATH" CLAUDECODE=1 PI_CODING_AGENT='' GROK_AGENT='' FM_ROOT_OVERRIDE="$dir" FM_STATE_OVERRIDE="$state" FM_GUARD_GRACE=1 "$ROOT/bin/fm-guard.sh" 2> "$err" >/dev/null || fail "guard failed" first=$(grep -v '^[[:space:]]*$' "$err" | head -1) case "$first" in '●'*) ;; @@ -162,7 +168,7 @@ test_guard_warnings() { mkdir -p "$dir/config" printf 'project=x\n' > "$state/task.meta" : > "$dir/config/x-mode.env" - CLAUDECODE=1 PI_CODING_AGENT='' GROK_AGENT='' FM_ROOT_OVERRIDE="$dir" FM_STATE_OVERRIDE="$state" FM_GUARD_GRACE=1 "$ROOT/bin/fm-guard.sh" 2> "$err" >/dev/null || fail "guard failed" + PATH="$blind:$PATH" CLAUDECODE=1 PI_CODING_AGENT='' GROK_AGENT='' FM_ROOT_OVERRIDE="$dir" FM_STATE_OVERRIDE="$state" FM_GUARD_GRACE=1 "$ROOT/bin/fm-guard.sh" 2> "$err" >/dev/null || fail "guard failed" grep -F "source '$dir/config/x-mode.env' first" "$err" >/dev/null || fail "guard repair line did not source the X-mode cadence config" # (2) live watcher plus fresh beacon, empty queue -> silence. diff --git a/tests/fm-x-mode.test.sh b/tests/fm-x-mode.test.sh index ff15f6a95b9..9f45252bc70 100755 --- a/tests/fm-x-mode.test.sh +++ b/tests/fm-x-mode.test.sh @@ -948,8 +948,14 @@ test_reply_text_file_and_stdin() { } test_bootstrap_opt_out_cleanup() { - local home out + local home out blind home="$TMP_ROOT/boot-optout"; mkdir -p "$home" + # The remediation wording asserted below is the CLAUDE one, so detect_own has to + # answer claude. A marker alone no longer pins that - a structural ancestor of a + # different harness outranks it - so blind the ancestry walk too, or the harness + # this suite was launched from picks the wording. + blind=$(fm_fakebin "$TMP_ROOT/boot-optout-blind") + fm_fake_blind_ancestry "$blind" # Opt in, artifacts appear. printf 'FMX_PAIRING_TOKEN=tok-out\n' > "$home/.env" FM_HOME="$home" "$ROOT/bin/fm-bootstrap.sh" >/dev/null 2>&1 @@ -957,7 +963,7 @@ test_bootstrap_opt_out_cleanup() { assert_present "$home/config/x-mode.env" "opt-in must create the cadence config" # Opt out: empty the token, re-run bootstrap -> artifacts removed + one off line. printf 'FMX_PAIRING_TOKEN=\n' > "$home/.env" - out=$(CLAUDECODE=1 FM_HOME="$home" "$ROOT/bin/fm-bootstrap.sh" 2>/dev/null) + out=$(PATH="$blind:$PATH" CLAUDECODE=1 FM_HOME="$home" "$ROOT/bin/fm-bootstrap.sh" 2>/dev/null) assert_contains "$out" "FMX: X mode off" "opt-out must announce X mode off when it removed artifacts" assert_contains "$out" "watcher supervision needs Stop-owned automatic recovery" "opt-out remediation must use neutral automatic-recovery guidance" assert_not_contains "$out" "is broken" "opt-out remediation claimed an unverified mechanism failure" diff --git a/tests/lib.sh b/tests/lib.sh index 65edb6defd9..4429362ae55 100644 --- a/tests/lib.sh +++ b/tests/lib.sh @@ -396,6 +396,34 @@ SH chmod +x "$fakebin/fm-crash-inject" } +# fm_fake_blind_ancestry <fakebin> +# Blind the parent-chain walks: a query of the FIELD-FIRST per-pid form those walks +# use - `ps -o comm=|args=|ppid= -p <pid>`, the shape in bin/fm-harness.sh, +# bin/fm-session-lock-lib.sh, bin/fm-sessionstart-nudge.sh and bin/fm-backend.sh's +# cmux ancestor detection - reports a bash ancestor terminating at pid 1, so ancestry +# proves nothing and the marker a case sets is the only evidence left. A case that pins +# its harness with a marker (CLAUDECODE=1 and friends) needs this, because a structural +# ancestor of a DIFFERENT harness outranks a marker - without it, the harness the SUITE +# was launched from decides the verdict. +# Every other ps query reaches the real ps untouched, and the pid-first form is +# deliberately among them: bin/fm-tmux-lib.sh and bin/backends/tmux.sh read pane and +# cursor identity with `ps -p <pid> -o args=`, so intercepting that shape too would make +# a pane assertion under a PATH-wide blind read `bash` and reject every cursor pane. +fm_fake_blind_ancestry() { + local fakebin=$1 real_ps + real_ps=$(command -v ps) || return 1 + cat > "$fakebin/ps" <<SH +#!/usr/bin/env bash +case "\$*" in + '-o comm= -p '*) printf '%s\n' bash ;; + '-o args= -p '*) printf '%s\n' bash ;; + '-o ppid= -p '*) printf '%s\n' 1 ;; + *) exec "$real_ps" "\$@" ;; +esac +SH + chmod +x "$fakebin/ps" +} + # fm_fake_version_tool <fakebin> <tool> <override-env-var> <default-version> # The stub answers `--version` with <override-env-var> when that variable is set # and non-empty, and with <default-version> otherwise; every other invocation From 305dfffa9a53ad9dc4d1364574c2fe0ccc9049fd Mon Sep 17 00:00:00 2001 From: NewAiCoder-bot <iamacodernow-bot@theinbtw.com> Date: Sun, 13 Sep 2026 06:20:20 -0400 Subject: [PATCH 26/31] fix(bin): pre-approve external CLAUDE.md import dialog for spawned workers (#3944) Claude Code's external-imports check (hasClaudeMdExternalIncludesApproved) reads only the canonical git-root project entry in ~/.claude.json, which its own worktree-to-primary-checkout canonicalization means is never the task worktree fm-claude-trust.sh registered. The trust dialog kept working previously only because its check has an ancestor-walk fallback that happens to reach the worktree entry; the external-imports check has no such fallback. Verified by disassembling the installed claude binary and reproducing in an isolated three-way tmux launch: identical flags registered only at the worktree key still showed the external-imports dialog, and registering them at the primary checkout key suppressed both dialogs. fm-claude-trust.sh now registers all three flags on both the worktree entry and the primary-checkout entry in one atomic write, and refuses when the <project> argument is not itself a primary checkout (its own write target would then be wrong). Extends the harness-adapters Claude reference and the trust test suite. Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com> --- .../references/harness/claude.md | 20 +- CONTRIBUTING.md | 2 +- bin/fm-claude-trust.sh | 215 ++++++++++++++++-- tests/fm-claude-trust.test.sh | 152 +++++++++++++ 4 files changed, 361 insertions(+), 28 deletions(-) diff --git a/.agents/skills/harness-adapters/references/harness/claude.md b/.agents/skills/harness-adapters/references/harness/claude.md index 437eb77c072..f29720ca411 100644 --- a/.agents/skills/harness-adapters/references/harness/claude.md +++ b/.agents/skills/harness-adapters/references/harness/claude.md @@ -16,16 +16,24 @@ Busy hooks verified 2026-07-28 on Claude Code 2.1.220. ## Workspace trust -Claude gates a folder it has never seen behind an interactive workspace-trust dialog, so every fresh task worktree would hit it, and so would every secondmate home no operator has opened by hand. +Claude gates a folder it has never seen behind an interactive workspace-trust dialog (titled "Quick safety check: Is this a project you created or one you trust?"), so every fresh task worktree would hit it, and so would every secondmate home no operator has opened by hand. `--dangerously-skip-permissions` does not cover that gate: `claude --help` records that the dialog is skipped only in non-interactive mode, through `-p` or a non-TTY stdout, and a spawned pane is interactive. Every claude spawn therefore pre-registers the directory its pane starts in before launch, and the dialog does not appear: the task worktree for a ship or scout, and the home itself for a `--secondmate` spawn, in either seeded shape (a leased worktree or a standalone clone). -`../../../bin/fm-claude-trust.sh` records `hasTrustDialogAccepted` for that path in `${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json` and owns the structural scope test each shape must pass, and `../../../bin/fm-spawn.sh` refuses the spawn when the registration fails rather than launching an agent that would wedge. -Never try to answer the trust dialog with a key. -Firstmate's key plane carries only Enter, Escape, and C-c with no arrow navigation, so it cannot move a dialog's selection at all, and the observed rendering starts on `No, exit`, which means a sent Enter ends the session instead of accepting. -A visible trust dialog means pre-registration did not take effect, so inspect the store and the spawn's error output rather than sending keys. +A second, separate dialog - "Allow external CLAUDE.md file imports?" - renders whenever a loaded CLAUDE.md chain reaches outside the project tree, which every crewmate's does through the captain's own `~/.claude/CLAUDE.md` importing `~/.claude/RTK.md`. +`--setting-sources project,local` (the minimal worker tool surface) does not suppress it either, and it gates the pane exactly like the trust dialog: cursor on "No, disable external imports", no way to move the selection from firstmate's steering plane. -The once-per-machine bypass-permissions confirmation is a separate dialog, scoped to the machine rather than the path, and pre-registration does not address it. +`../../../bin/fm-claude-trust.sh` records `hasTrustDialogAccepted` for both the worktree and its primary checkout in `${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json` for a ship or scout spawn; a secondmate spawn registers only its own home entry, since a secondmate home has no separate primary-checkout entry to carry import consent forward from. +For a ship or scout spawn, the external-imports flags (`hasClaudeMdExternalIncludesApproved`, `hasClaudeMdExternalIncludesWarningShown`) are carried forward alongside the trust flag only when the primary checkout's project entry already carries an explicit `hasClaudeMdExternalIncludesApproved===true` from a prior interactive session - the common first-spawn case is a project claude has never been asked about, so those two flags are left unwritten and the import dialog still renders, even though trust registers normally. +When the project entry instead already carries an explicit decline (`===false`), the whole registration refuses - including the trust flag - rather than manufacture consent the human never gave, so that spawn wedges on the trust dialog before it would even reach the import one. +The why-two-entries mechanism and the consent-gating logic live in the script's own header comment, which is the one owner for that contract; the fact worth repeating here is that `../../../bin/fm-spawn.sh` refuses the spawn when the trust flag fails to land, rather than launching a worker that would wedge on that dialog. + +Never try to answer either dialog with a key. +Firstmate's key plane carries only Enter, Escape, and C-c with no arrow navigation, so it cannot move a dialog's selection at all, and both dialogs render with the cursor on their declining option, which means a sent Enter ends the session instead of accepting. +A visible trust dialog means pre-registration did not take effect (or the project entry already carries an explicit decline) - inspect the store and the spawn's error output rather than sending keys. +A visible external-imports dialog is expected, not a failure signal, whenever the project entry has no prior explicit approval on record - the common first-spawn case; `fm-control.sh <id> interrupt` delivers Escape, which dismisses whichever of the two is on screen without answering it, and is the safe way to clear a wedged pane for inspection. + +The once-per-machine bypass-permissions confirmation is a third, separate dialog, scoped to the machine rather than the path, and pre-registration does not address it. Never send Enter to that one either: it was observed rendering in the same shape as the trust dialog, with the selection on `No, exit` and the footer `Enter to confirm . Esc to cancel`, so Enter ends the session rather than accepting. Firstmate cannot move a selection with Enter, Escape, and C-c alone, so it cannot accept this dialog at all, and an operator accepts it once per machine instead. Inspect the pane to identify which dialog is on screen, and report it rather than answering it. diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index de4cc759a35..9567425893b 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -52,7 +52,7 @@ See the [no-mistakes quick start](https://kunchenguid.github.io/no-mistakes/star It pins one exact shellcheck version and one exact actionlint version and refuses to run under any other. Print the shellcheck pin with `bin/fm-lint.sh --required-version` and the actionlint pin with `bin/fm-lint-workflows.sh --required-version`. Use `bin/fm-install-shellcheck.sh` and `bin/fm-install-actionlint.sh` to install those exact builds locally; each installer's header owns its destination usage and supported platforms. -- Harness-adapter ownership spans detection in `bin/fm-harness.sh`, launch and hook mechanics in `bin/fm-spawn.sh`, spawn-time workspace-trust pre-registration in `bin/fm-claude-trust.sh` and `bin/fm-agy-trust.sh`, semantic busy sources and trust gates in `bin/fm-busy-lib.sh`, delivery-only rendered guards in `bin/fm-composer-lib.sh`, cleanup in `bin/fm-teardown.sh`, and facts in the skill tree rooted at `.agents/skills/harness-adapters/SKILL.md`; the `firstmate-coding-guidelines` skill owns the validation policy for checks that depend on those harnesses. +- Harness-adapter ownership spans detection in `bin/fm-harness.sh`, launch and hook mechanics in `bin/fm-spawn.sh`, spawn-time Claude workspace-trust and external-CLAUDE.md-import pre-approval in `bin/fm-claude-trust.sh`, semantic busy sources and trust gates in `bin/fm-busy-lib.sh`, delivery-only rendered guards in `bin/fm-composer-lib.sh`, cleanup in `bin/fm-teardown.sh`, and facts in the skill tree rooted at `.agents/skills/harness-adapters/SKILL.md`; the `firstmate-coding-guidelines` skill owns the validation policy for checks that depend on those harnesses. - Changes to runtime session backends (`bin/fm-backend.sh`, `bin/backends/`, and the scripts that dispatch through them) keep current setup and limits in the relevant backend guide and active empirical evidence in [`docs/verification/runtime-backends.md`](docs/verification/runtime-backends.md). - [`docs/documentation-audiences.md`](docs/documentation-audiences.md) and its machine-consumed inventory own prose classification; run `bin/fm-doc-audience-check.sh` after documentation changes. - In Markdown, put each full sentence on its own line. diff --git a/bin/fm-claude-trust.sh b/bin/fm-claude-trust.sh index 770dd79cfb5..07762cf9a0a 100755 --- a/bin/fm-claude-trust.sh +++ b/bin/fm-claude-trust.sh @@ -2,7 +2,11 @@ # Pre-register Claude Code's workspace trust for the directory a claude spawn is # about to launch into - the isolated task worktree of a ship or scout crewmate, # or the seeded home of a secondmate - so the agent reaches its brief or charter -# instead of wedging on the trust dialog. +# instead of wedging on the trust dialog. In worktree mode it also carries +# forward the external-CLAUDE.md-import approval, but only when the primary +# checkout already holds standing consent for it - see the consent-gating +# block below for why that dialog is otherwise left for the worker to wedge +# on rather than answered on the human's behalf. # # Usage: fm-claude-trust.sh <worktree> <project> # fm-claude-trust.sh --secondmate-home <home> <id> @@ -22,7 +26,52 @@ # Enter, Escape and C-c with no arrow navigation, so firstmate cannot answer it # and must not try - pressing Enter would select exit. The agent wedges before # it ever reads the brief. Registering the trust before launch is the only -# control that reaches an interactive pane. +# control that reaches an interactive pane. The same reasoning covers Claude +# Code's separate "Allow external CLAUDE.md file imports?" dialog, which +# `--setting-sources project,local` (firstmate PR 10's minimal worker tool +# surface) stopped suppressing: it renders whenever a loaded CLAUDE.md chain +# reaches outside the project tree - which every crewmate's does, through the +# captain's own `~/.claude/CLAUDE.md` importing `~/.claude/RTK.md` - and it is +# gated the same fail-closed way as trust: cursor on "No, disable", no arrow +# navigation from firstmate's steering plane. Only worktree mode reaches this +# second dialog's flags: a secondmate home has no separate "project" entry to +# carry consent forward from, so its registration stays trust-only. +# +# TWO PROJECT-CONFIG ENTRIES IN WORKTREE MODE, NOT ONE. Registering both flags +# on the worktree entry alone (the original trust-only design) leaves the +# external-imports dialog showing. Verified 2026-09-06 by disassembling the +# installed `claude` binary and reproducing in an isolated three-way tmux +# launch: Claude Code's own trust check (`Rde`) reads the canonical +# project-root entry first and, failing that, falls back to an ancestor walk +# from the worktree upward that DOES reach the worktree's own entry - which is +# why the trust dialog kept working after PR 10. The external-imports check +# (`es`/`F1e`) has no such fallback: it reads ONLY the canonical project-root +# entry, and that root is never the worktree - Claude Code's own git-root +# canonicalization (`Fr`/`Se`) walks a linked worktree's `.git` file through +# its `commondir` pointer back to the PRIMARY CHECKOUT, exactly the <project> +# argument this script already receives for the worktree-mode scope test +# below. So the trust flag is registered on BOTH the worktree entry (for +# trust's ancestor-walk fallback and defense in depth) and the project entry +# (the trust check's first, canonical-shaped, look); the two external-imports +# flags land on those same two entries only when the project entry already +# carries standing consent (see the consent-gating block below) - the project +# entry is the only place the external-imports check ever looks. Registering +# the project entry is a write to the launching user's OWN Claude config +# store, keyed by a project PATH the scope test below has already verified is +# real - not a write to the project's tracked content, so hard rule 1 does not +# apply, same as the existing worktree-entry write. +# +# THAT SAME PROJECT ENTRY IS ALSO THE LAUNCHING HUMAN'S OWN INTERACTIVE +# CONFIG, though, so this registration must never overwrite a decision the +# human already made there. If the project entry already carries +# hasClaudeMdExternalIncludesApproved===false - Claude Code only ever writes +# that on an explicit "No, disable" answer - the whole registration refuses +# rather than flipping it, because doing so would grant every future +# interactive session in that checkout silent external-file inclusion the +# human declined, permanently and without being asked. The worktree entry is +# left unwritten too: the spawn wedges on the dialog, which is the honest +# outcome given a standing decline, not registered trust with a stripped +# consent record. # # THE SCOPE TEST IS THE SAFETY PROPERTY, and it is STRUCTURAL rather than a # path policy. Each mode has its own, because the two directories have entirely @@ -34,7 +83,18 @@ # own word: a primary checkout (git dir == common dir), a worktree of an # unrelated repo, a subdirectory of a worktree, a plain directory, and a home # directory are each refused. Refusal is a non-zero exit, never a warning and -# never a silent skip. +# never a silent skip. When <project> is itself a linked worktree (a +# secondmate home spawned from, rather than as, the primary checkout), +# refusing outright would wedge a relaunch that is otherwise perfectly valid: +# its own common dir already IS the primary checkout's own git dir (git's +# git-common-dir answer never changes by which worktree asks), so the +# checkout is derived structurally from it - its parent directory in the +# standard non-bare, non-GIT_DIR-overridden layout this script already +# requires elsewhere - and verified, never assumed: the candidate's own +# resolved git dir must equal that common dir, the same primary-checkout +# definition used throughout, or this refuses rather than guess. The +# consent-gated external-imports flags land on that resolved canonical +# checkout, never on the linked-worktree argument itself. # # The test is deliberately NOT a treehouse or orca path prefix. Treehouse's # root is configurable (--root, TREEHOUSE_ROOT, config, and a relative @@ -75,15 +135,21 @@ # # Home-level trust is broader than worktree trust, since the pane starts in the # home and the secondmate works across it, so it is granted on that seed -# evidence alone and never on a caller's word about what a path is. +# evidence alone and never on a caller's word about what a path is. It is +# trust-only: a secondmate home has no separate primary-checkout "project" +# argument to gate external-imports consent against, so the two import flags +# are never written there. # -# Only the launching user's own store is written: the projects entry for the -# registered path in ${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json, which must be a -# regular file this uid owns. Every unrelated key and project entry is -# preserved, and the replacement is atomic. fm-spawn.sh forwards CLAUDE_CONFIG_DIR -# onto the claude launch verbatim rather than resolving it, and the pane starts -# in the registered directory, so only an absolute value names the same store on -# both sides; a relative one is refused below rather than guessed at. +# Only the launching user's own store is written. In worktree mode: the +# projects entries for the worktree path and the resolved canonical project +# path in ${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json, which must be a regular +# file this uid owns; every unrelated key and project entry is preserved, and +# both entries land in one atomic replacement. In secondmate-home mode: the +# single projects entry for the registered home path, same store, same atomic +# replacement. fm-spawn.sh forwards CLAUDE_CONFIG_DIR onto the claude launch +# verbatim rather than resolving it, and the pane starts in the registered +# directory, so only an absolute value names the same store on both sides; a +# relative one is refused below rather than guessed at. set -u # Path resolution here must answer from the filesystem, never from the caller's # environment, because the refusals below are the safety property. CDPATH would @@ -204,6 +270,35 @@ if [ "$MODE" = worktree ]; then PROJ_COMMON=$(common_dir_of "$PROJ_REAL") || true [ -n "$PROJ_COMMON" ] || refuse "project '$PROJ_REAL' is not inside a git repository" [ "$WT_COMMON" = "$PROJ_COMMON" ] || refuse "'$TARGET_REAL' is not a worktree of project '$PROJ_REAL'" + + # The external-imports flags must land on the primary checkout - its own git + # dir equals the common dir - because that is exactly the path Claude Code's + # own git-root canonicalization collapses every linked worktree to. When + # <project> is itself a linked worktree (a secondmate home spawned from, + # rather than as, the primary checkout), refusing outright would wedge a + # relaunch that is otherwise perfectly valid: PROJ_COMMON already IS that + # primary checkout's own git dir (git's git-common-dir answer never changes + # by which worktree asks), so the checkout is derived structurally from it - + # its parent directory in the standard non-bare, non-GIT_DIR-overridden + # layout this script already requires elsewhere - and verified, never + # assumed: the candidate's own resolved git dir must equal PROJ_COMMON, the + # same primary-checkout definition used above, or this refuses rather than + # guess. + PROJ_GIT_DIR=$(git -C "$PROJ_REAL" rev-parse --absolute-git-dir 2>/dev/null) || true + [ -n "$PROJ_GIT_DIR" ] || refuse "project '$PROJ_REAL' has no resolvable git directory" + PROJ_GIT_DIR=$(real_dir "$PROJ_GIT_DIR") || true + [ -n "$PROJ_GIT_DIR" ] || refuse "project '$PROJ_REAL' has an unresolvable git directory" + if [ "$PROJ_GIT_DIR" = "$PROJ_COMMON" ]; then + PROJ_CANON=$PROJ_REAL + else + PROJ_CANON=$(real_dir "$(dirname -- "$PROJ_COMMON")") || true + [ -n "$PROJ_CANON" ] \ + || refuse "project '$PROJ_REAL' is a linked worktree whose primary checkout could not be resolved" + CANON_GIT_DIR=$(git -C "$PROJ_CANON" rev-parse --absolute-git-dir 2>/dev/null) || true + CANON_GIT_DIR=$(real_dir "${CANON_GIT_DIR:-}") || true + [ -n "$CANON_GIT_DIR" ] && [ "$CANON_GIT_DIR" = "$PROJ_COMMON" ] \ + || refuse "project '$PROJ_REAL' is a linked worktree whose primary checkout could not be resolved" + fi else # The seed evidence, in the order that names the most useful reason first: the # marker decides whether this is a secondmate home at all, the id decides @@ -287,11 +382,44 @@ fi # attempts, and it must fail loudly rather than report a trust it did not leave. # ponytail: fingerprint-and-refuse, not a lock; flock is absent on macOS and # cannot stop a vendor session's own rewrite anyway. -if ! node - "$STORE" "$TARGET_REAL" <<'NODE' +# +# In worktree mode every flag lands on both the worktree entry and the project +# entry in the same read-modify-write attempt, so a single rename either +# records all of it or none of it - there is no state where the worktree entry +# is fresh and the project entry stale, or the other way round. In +# secondmate-home mode only the single home entry is written. +# +# The two external-imports flags (worktree mode only) are gated separately +# from the trust flag, because they are a CONSENT grant, not a pre-approval +# this script is allowed to manufacture. Claude Code only ever writes +# hasClaudeMdExternalIncludesApproved itself, on an explicit interactive +# answer; this script's own job is to keep a worker from wedging on a dialog, +# never to answer that dialog on the human's behalf. So the import flags land +# on the project entry - the only place the imports check ever reads (see the +# disassembly note above) - only when that entry ALREADY carries +# hasClaudeMdExternalIncludesApproved===true, i.e. the human already said yes +# at some point and this write is a same-value refresh, not new consent from +# an absent flag. When it is not already true (including plain absent, the +# common case for a project claude has never asked about), the import flags +# are left untouched on both entries: writing them to the worktree entry alone +# would be a pure no-op (the imports check never reads it) that only obscures +# the real state, so trust still registers normally but the import dialog is +# left exactly as undecided as it already was - the worker wedges on it, the +# same honest outcome as an explicit decline, rather than a spawn spending +# consent the human was never asked for. +TRUST_FLAG='hasTrustDialogAccepted' +IMPORT_FLAGS='["hasClaudeMdExternalIncludesApproved","hasClaudeMdExternalIncludesWarningShown"]' +if [ "$MODE" = worktree ]; then + WRITE_ARGS=("$STORE" "$MODE" "$TARGET_REAL" "$PROJ_CANON" "$TRUST_FLAG" "$IMPORT_FLAGS") +else + WRITE_ARGS=("$STORE" "$MODE" "$TARGET_REAL" "" "$TRUST_FLAG" "$IMPORT_FLAGS") +fi +if ! node - "${WRITE_ARGS[@]}" <<'NODE' const fs = require("node:fs"); const path = require("node:path"); const crypto = require("node:crypto"); -const [store, target] = process.argv.slice(2); +const [store, mode, target, project, trustFlag, importFlagsJson] = process.argv.slice(2); +const importFlags = JSON.parse(importFlagsJson); const readStore = () => { try { return fs.readFileSync(store); @@ -302,6 +430,32 @@ const readStore = () => { }; const fingerprint = (buf) => buf === null ? "absent" : crypto.createHash("sha256").update(buf).digest("hex"); +const setFlags = (projects, key, flags) => { + let entry = projects[key]; + if (entry === undefined || entry === null || typeof entry !== "object" || Array.isArray(entry)) { + entry = {}; + } + for (const flag of flags) entry[flag] = true; + projects[key] = entry; +}; +const flagsLanded = (projects, key, flags) => + flags.every((flag) => projects?.[key]?.[flag] === true); +// The project entry is the launching user's OWN interactive config, not a +// throwaway worktree, so a spawn must never silently reverse a decision the +// human already recorded there. hasClaudeMdExternalIncludesApproved===false +// is exactly that decision (Claude Code only ever writes it on an explicit +// "No, disable" answer); flipping it to true would grant every future +// interactive session in that checkout silent external-file inclusion the +// human declined. Refuse the whole registration instead of overriding it - +// the worktree entry is not written either, so the spawn wedges on the +// dialog rather than the human's consent being spent without being asked. +const declinedExternalImports = (projects, key) => + projects?.[key]?.hasClaudeMdExternalIncludesApproved === false; +// True only on an explicit prior "Yes, allow" answer - the sole state this +// script may treat as standing consent to refresh. Absent, or any other +// value, is NOT consent (see the block comment above this script's node call). +const approvedExternalImports = (projects, key) => + projects?.[key]?.hasClaudeMdExternalIncludesApproved === true; const attempt = () => { const original = readStore(); const before = fingerprint(original); @@ -320,12 +474,23 @@ const attempt = () => { if (projects === null || typeof projects !== "object" || Array.isArray(projects)) { throw new Error(`${store} has a non-object "projects" value`); } - let entry = projects[target]; - if (entry === undefined || entry === null || typeof entry !== "object" || Array.isArray(entry)) { - entry = {}; + let keys; + if (mode === "worktree") { + if (declinedExternalImports(projects, project)) { + throw new Error( + `project entry for ${project} in ${store} already declined external CLAUDE.md imports; refusing to override that consent`, + ); + } + const carryImportConsent = approvedExternalImports(projects, project); + const targetFlags = carryImportConsent ? [trustFlag, ...importFlags] : [trustFlag]; + const projectFlags = carryImportConsent ? [trustFlag, ...importFlags] : [trustFlag]; + setFlags(projects, target, targetFlags); + setFlags(projects, project, projectFlags); + keys = [[target, targetFlags], [project, projectFlags]]; + } else { + setFlags(projects, target, [trustFlag]); + keys = [[target, [trustFlag]]]; } - entry.hasTrustDialogAccepted = true; - projects[target] = entry; // Unpredictable name plus an exclusive create: the config directory may be // writable by another local account, and a predictable path could be // pre-created there as a symlink that a plain write would follow into some @@ -347,7 +512,8 @@ const attempt = () => { if (!renamed) fs.rmSync(tmp, { force: true }); } const back = JSON.parse(fs.readFileSync(store, "utf8")); - return back.projects?.[target]?.hasTrustDialogAccepted === true ? "recorded" : "dropped"; + const landed = keys.every(([key, flags]) => flagsLanded(back.projects, key, flags)); + return landed ? "recorded" : "dropped"; }; try { for (let i = 0; i < 3; i += 1) { @@ -362,11 +528,18 @@ try { console.error(`error: ${err.message}`); process.exit(1); } -console.error(`error: ${store} did not retain trust for ${target} after 3 attempts`); +console.error(`error: ${store} did not retain trust for ${target}${project ? ` and ${project}` : ""} after 3 attempts`); process.exit(1); NODE then - refuse "could not record trust for '$TARGET_REAL' in '$STORE'" + if [ "$MODE" = worktree ]; then + refuse "could not record trust for '$TARGET_REAL' and project '$PROJ_CANON' in '$STORE'" + else + refuse "could not record trust for '$TARGET_REAL' in '$STORE'" + fi fi echo "trusted: $TARGET_REAL" +if [ "$MODE" = worktree ]; then + echo "trusted (project root): $PROJ_CANON" +fi diff --git a/tests/fm-claude-trust.test.sh b/tests/fm-claude-trust.test.sh index 893840c1901..c040682516d 100755 --- a/tests/fm-claude-trust.test.sh +++ b/tests/fm-claude-trust.test.sh @@ -68,6 +68,39 @@ assert_store_value() { # <store> <expected-json> <msg> <key...> [ "$actual" = "$expected" ] || fail "$msg (expected $expected, got $actual)" } +# assert_all_flags <store> <path> <msg>: all three registered flags - trust, +# external-includes approved, external-includes warning-shown - are true on +# the project entry at <path>. The external-imports flags are the ones the +# running app reads only from the PROJECT-root entry, never the worktree +# entry, so this is what actually proves the dialog is suppressed. +assert_all_flags() { + local store=$1 key=$2 msg=$3 + node -e ' + const j=JSON.parse(require("node:fs").readFileSync(process.argv[1],"utf8")); + const e=(j.projects||{})[process.argv[2]]||{}; + const flags=["hasTrustDialogAccepted","hasClaudeMdExternalIncludesApproved","hasClaudeMdExternalIncludesWarningShown"]; + process.exit(flags.every((f)=>e[f]===true)?0:1); + ' "$store" "$key" || fail "$msg" +} + +# assert_trust_only_no_import_consent <store> <path> <msg>: the entry at +# <path> carries hasTrustDialogAccepted===true but NEITHER external-imports +# flag is true - the shape a registration must leave behind when the project +# entry had no prior explicit "Yes, allow" for external CLAUDE.md imports, so +# a spawn never manufactures that consent from an absent flag. +assert_trust_only_no_import_consent() { + local store=$1 key=$2 msg=$3 + node -e ' + const j=JSON.parse(require("node:fs").readFileSync(process.argv[1],"utf8")); + const e=(j.projects||{})[process.argv[2]]||{}; + const trustOk = e.hasTrustDialogAccepted === true; + const noImportConsent = + e.hasClaudeMdExternalIncludesApproved !== true && + e.hasClaudeMdExternalIncludesWarningShown !== true; + process.exit(trustOk && noImportConsent ? 0 : 1); + ' "$store" "$key" || fail "$msg" +} + # A PATH carrying the tools the scope test needs but no node, so the # missing-interpreter path is exercised without disturbing the real PATH. node_free_path() { # <case-dir> -> a bin dir holding the script's own tools but no node @@ -140,6 +173,97 @@ test_fresh_worktree_is_trusted() { pass "fm-claude-trust.sh: a fresh task worktree is trusted" } +# The trust dialog is read only from the PROJECT-root entry, never the +# worktree entry (Claude Code's own git-root canonicalization collapses every +# linked worktree to its primary checkout for that check, with no +# ancestor-walk fallback the way the trust check has), so this proves both +# entries carry the trust flag after one registration. External-imports +# consent is a SEPARATE grant this script never manufactures: on a genuinely +# fresh project (no prior interactive answer at all) neither entry may carry +# hasClaudeMdExternalIncludesApproved or hasClaudeMdExternalIncludesWarningShown +# - see test_registration_carries_forward_existing_import_consent below for +# the case where the project already said yes. +test_fresh_worktree_also_trusts_the_project_root_without_import_consent() { + local rec out + rec=$(make_case fresh-project) + read_case "$rec" + out=$(run_trust "$CONFIG" "$WT" "$PROJ") + expect_code 0 $? "a fresh linked worktree must be trusted: $out" + assert_contains "$out" "$PROJ" "registration did not report the project root it also trusted" + assert_trust_only_no_import_consent "$CONFIG/.claude.json" "$WT" \ + "the worktree entry either lost trust or gained unearned import consent" + assert_trust_only_no_import_consent "$CONFIG/.claude.json" "$PROJ" \ + "the project-root entry either lost trust or gained unearned import consent" + pass "fm-claude-trust.sh: a fresh registration trusts the project root without manufacturing import consent" +} + +# The Greptile-flagged regression this pins: a project entry that already +# carries an explicit "Yes, allow" (hasClaudeMdExternalIncludesApproved===true) +# is exactly the standing consent this script may refresh - and refreshing it +# is what actually suppresses the external-imports dialog for the worker, +# since that check reads only the project entry (see the disassembly note at +# the top of fm-claude-trust.sh), never the worktree one. +test_registration_carries_forward_existing_import_consent() { + local rec store + rec=$(make_case import-consent-carried) + read_case "$rec" + store="$CONFIG/.claude.json" + cat > "$store" <<JSON +{"hasCompletedOnboarding":true,"projects":{"$PROJ":{"hasTrustDialogAccepted":true,"hasClaudeMdExternalIncludesApproved":true,"hasClaudeMdExternalIncludesWarningShown":true}}} +JSON + run_trust "$CONFIG" "$WT" "$PROJ" >/dev/null || fail "registration failed against a project that already approved external imports" + assert_all_flags "$store" "$WT" \ + "the worktree entry did not carry the refreshed import consent" + assert_all_flags "$store" "$PROJ" \ + "the project-root entry lost its own already-granted import consent" + pass "fm-claude-trust.sh: carries forward a project's already-granted import consent to the worktree entry" +} + +# The project-root entry is the same store the launching user's interactive +# claude sessions read and write (it is usually already present, carrying +# unrelated keys such as allowedTools or MCP config), so preservation must +# hold there exactly as it holds for the worktree entry. +test_project_root_entry_preserves_other_keys() { + local rec store + rec=$(make_case project-preserve) + read_case "$rec" + store="$CONFIG/.claude.json" + cat > "$store" <<JSON +{"hasCompletedOnboarding":true,"projects":{"$PROJ":{"hasTrustDialogAccepted":false,"allowedTools":["Read"]}}} +JSON + run_trust "$CONFIG" "$WT" "$PROJ" >/dev/null || fail "registration failed against an existing project entry" + assert_trust_only_no_import_consent "$store" "$PROJ" \ + "the project-root entry did not gain trust, or gained unearned import consent it had never been asked for" + assert_store_value "$store" '["Read"]' "the project entry's unrelated settings were lost" projects "$PROJ" allowedTools + pass "fm-claude-trust.sh: preserves unrelated keys on the project-root entry" +} + +# hasClaudeMdExternalIncludesApproved===false on the project-root entry is a +# human's explicit "No, disable" answer, recorded in the SAME store their own +# interactive sessions read. A spawn must never flip that to true on their +# behalf: doing so would grant every later interactive session in that +# checkout silent external-file inclusion the human declined. The whole +# registration refuses instead, and the store - including the worktree entry, +# which is never reached - must come back byte-for-byte unchanged. +test_project_root_entry_declined_external_imports_is_not_overridden() { + local rec store out before after + rec=$(make_case project-decline) + read_case "$rec" + store="$CONFIG/.claude.json" + cat > "$store" <<JSON +{"hasCompletedOnboarding":true,"projects":{"$PROJ":{"hasTrustDialogAccepted":true,"hasClaudeMdExternalIncludesApproved":false,"hasClaudeMdExternalIncludesWarningShown":true,"allowedTools":["Read"]}}} +JSON + before=$(cat "$store") + out=$(run_trust "$CONFIG" "$WT" "$PROJ") + expect_code 1 $? "a project that already declined external imports must be refused: $out" + assert_contains "$out" "declined external CLAUDE.md imports" \ + "the refusal did not name the declined-consent reason" + after=$(cat "$store") + [ "$before" = "$after" ] || fail "the store was modified despite the refusal" + assert_not_trusted "$store" "$WT" "the worktree entry was registered despite the refusal" + pass "fm-claude-trust.sh: refuses to override a project's declined external-imports consent" +} + test_registration_is_idempotent() { local rec out count rec=$(make_case idempotent) @@ -312,6 +436,29 @@ test_worktree_subdirectory_is_refused() { pass "fm-claude-trust.sh: refuses a subdirectory of the worktree" } +# The write target the external-imports flags depend on is only correct when +# it names the primary checkout. When <project> is itself a linked worktree +# (a secondmate home spawned from, rather than as, the primary checkout), +# writing the flags at that worktree's own path would land them at a key +# Claude Code's git-root canonicalization never reads, silently reproducing +# the bug this script exists to close - so this resolves the argument +# structurally to its primary checkout instead of refusing it. +test_project_argument_that_is_itself_a_worktree_resolves_to_the_primary_checkout() { + local rec out proj_wt + rec=$(make_case nested-project) + read_case "$rec" + proj_wt="$CASE_DIR/proj-wt" + git -C "$PROJ" worktree add --quiet -b wt-proj-wt "$proj_wt" + out=$(run_trust "$CONFIG" "$WT" "$proj_wt") + expect_code 0 $? "a project argument that is itself a linked worktree must resolve to its primary checkout: $out" + assert_contains "$out" "$PROJ" "the outcome did not name the resolved primary checkout" + assert_trust_only_no_import_consent "$CONFIG/.claude.json" "$PROJ" \ + "the resolved primary checkout either lost trust or gained unearned import consent" + assert_not_trusted "$CONFIG/.claude.json" "$proj_wt" \ + "the linked worktree argument itself was recorded as the project root" + pass "fm-claude-trust.sh: a project argument that is itself a linked worktree resolves to the primary checkout" +} + test_unrelated_store_content_is_preserved() { local rec store rec=$(make_case preserve) @@ -645,6 +792,10 @@ test_secondmate_spawn_fails_closed_when_home_trust_cannot_be_recorded() { } test_fresh_worktree_is_trusted +test_fresh_worktree_also_trusts_the_project_root_without_import_consent +test_registration_carries_forward_existing_import_consent +test_project_root_entry_preserves_other_keys +test_project_root_entry_declined_external_imports_is_not_overridden test_registration_is_idempotent test_primary_checkout_is_refused test_cdpath_cannot_defeat_the_primary_checkout_refusal @@ -656,6 +807,7 @@ test_non_git_directory_is_refused test_missing_directory_is_refused test_foreign_project_worktree_is_refused test_worktree_subdirectory_is_refused +test_project_argument_that_is_itself_a_worktree_resolves_to_the_primary_checkout test_unrelated_store_content_is_preserved test_symlinked_store_to_a_foreign_owned_target_is_refused test_symlinked_store_to_an_owned_target_is_accepted From ecfe071981b1460c83b4dc82897e3d22e98f8503 Mon Sep 17 00:00:00 2001 From: NewAiCoder-bot <iamacodernow-bot@theinbtw.com> Date: Sun, 13 Sep 2026 06:37:26 -0400 Subject: [PATCH 27/31] fix(afk-return): treat an acked watcher-down marker as no gap (#4355) The marker lifecycle (fm-wake-lib.sh _fm_recovery_marker_ack) leaves state/.watcher-down behind in an acked:* state after a downtime episode is handled. health_snapshot's presence check reported that as an open gap on every later return, so a handled episode kept surfacing as a false GAP forever. --- bin/fm-afk-return.sh | 13 ++++++++++++- tests/fm-afk-return.test.sh | 19 +++++++++++++++++++ 2 files changed, 31 insertions(+), 1 deletion(-) diff --git a/bin/fm-afk-return.sh b/bin/fm-afk-return.sh index 51f6d7ff8ff..0587fa6d347 100755 --- a/bin/fm-afk-return.sh +++ b/bin/fm-afk-return.sh @@ -315,7 +315,18 @@ health_snapshot() { # <evidence-file> local evidence=$1 beat_age lines="" beat_age=$(fm_path_age "$STATE/.last-watcher-beat") if [ -e "$STATE/.watcher-down" ]; then - lines="GAP: watcher downtime was detected during the away window (recovery marker present)" + # The marker survives past its episode in an acked:* state + # (fm-wake-lib.sh _fm_recovery_marker_ack); only pending:* and + # announced:* mean the downtime is still open. A marker this read + # cannot parse is treated the same as an open gap, conservatively. + if fm_recovery_marker_snapshot "$STATE/.watcher-down"; then + case "$FM_RECOVERY_MARKER_TOKEN" in + acked:*) : ;; + *) lines="GAP: watcher downtime was detected during the away window (recovery marker present)" ;; + esac + else + lines="GAP: watcher downtime was detected during the away window (recovery marker present)" + fi fi if [ -e "$STATE/.afk" ] && ! fm_afk_daemon_owns_supervision "$STATE"; then lines="$lines diff --git a/tests/fm-afk-return.test.sh b/tests/fm-afk-return.test.sh index d655a49946d..2322687d68a 100755 --- a/tests/fm-afk-return.test.sh +++ b/tests/fm-afk-return.test.sh @@ -677,6 +677,24 @@ test_return_brief_health_leads_with_a_gap() { pass "the return brief leads with supervisor health and names every detected gap" } +test_return_brief_does_not_report_an_acked_watcher_down_marker_as_a_gap() { + local dir out + dir="$TMP_ROOT/brief-acked-marker" + install_runner "$dir" + contract_in "$dir" propose >/dev/null 2>&1 || fail "could not propose the away-posture record" + contract_in "$dir" confirm >/dev/null 2>&1 || fail "could not write the away-posture record" + # An episode that was detected and fully handled during the away window + # leaves the marker behind in an acked state (fm-wake-lib.sh + # _fm_recovery_marker_ack); that is not an open gap. + printf 'acked:downtime:fixture-generation\n' > "$dir/home/state/.watcher-down" + touch "$dir/home/state/.last-watcher-beat" + : > "$dir/home/state/.fake-drain" + out=$(run_return "$dir" begin) || fail "a clean fleet with only a handled marker should clear the gate: $out" + assert_not_contains "$out" 'GAP: watcher downtime was detected' "an acked recovery marker was reported as an open gap" + assert_contains "$out" 'no detected gap' "a fully acked window was not reported as clean" + pass "the return brief does not report an already-acked watcher-down marker as an open gap" +} + test_return_brief_without_a_record_reports_the_legacy_flag() { local dir out dir="$TMP_ROOT/brief-legacy" @@ -784,4 +802,5 @@ test_failed_held_listing_keeps_catchup_gated test_unreadable_status_file_keeps_catchup_gated test_return_guard_refuses_while_the_record_exists test_return_brief_health_leads_with_a_gap +test_return_brief_does_not_report_an_acked_watcher_down_marker_as_a_gap test_return_brief_without_a_record_reports_the_legacy_flag From b182d0f908b78d08c7ccb8dce3775bdca8c5d657 Mon Sep 17 00:00:00 2001 From: NewAiCoder-bot <iamacodernow-bot@theinbtw.com> Date: Sun, 13 Sep 2026 10:26:45 -0400 Subject: [PATCH 28/31] fix(bin): rebind fm-procevent-when trust bindings after a self-update (#4361) * fix(update): rebind fm-procevent-when watches after a self-update A self-update fast-forwards bin/ in place, changing an armed watch's action executable bytes with no tampering involved. The watch's trust binding was hashed at arm time, so the very next fire was refused as not matching the registered binding and the watch died silently. Add fm-procevent-when.sh rebind-all: it re-hashes and republishes the trust binding for every watch whose action executable lives under FM_ROOT, using the same spec/trust validation as an ordinary fire, and leaves any watch whose action lives outside FM_ROOT untouched. Wire it into fm-update.sh right after a successful fast-forward, for both the primary home and any local secondmate home that advances. * no-mistakes(review): Canonicalize FM_ROOT for rebind-all's containment check * no-mistakes(document): Document fm-update.sh's automatic watch rebind and its verification evidence * no-mistakes(lint): fix(tests): double-quote printf scripts to satisfy shellcheck SC2016 * no-mistakes(review): Reload trust binding from disk before firing to reach live pollers * no-mistakes(review): Lock the fire-time trust reload against rebind_one's publish race * no-mistakes(document): Document rebind-all's self-update guarantee and its two review-round test rows --------- Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com> --- bin/fm-procevent-when.sh | 140 ++++++++++- bin/fm-update.sh | 32 ++- docs/configuration.md | 3 +- docs/verification/process-event-sources.md | 5 + tests/fm-procevent-when.test.sh | 256 +++++++++++++++++++++ tests/fm-update.test.sh | 37 +++ 6 files changed, 467 insertions(+), 6 deletions(-) diff --git a/bin/fm-procevent-when.sh b/bin/fm-procevent-when.sh index c67539f27c9..76f11f7df68 100755 --- a/bin/fm-procevent-when.sh +++ b/bin/fm-procevent-when.sh @@ -10,6 +10,7 @@ # fm-procevent-when.sh terminal <result-file> # fm-procevent-when.sh source-id <name> # fm-procevent-when.sh retire <name> +# fm-procevent-when.sh rebind-all # fm-procevent-when.sh run <source-id> # # arm Bind a (condition, action) pair as process-event source @@ -49,6 +50,19 @@ # record, and fired marker. Idempotent. Captured results and their # handled acknowledgements are never touched. Warns when the action # had already fired without a captured outcome. +# rebind-all Refresh the trust binding of every registered watch whose action +# executable lives under this repo (FM_ROOT), re-hashing it against +# its CURRENT on-disk bytes. A self-update fast-forwards bin/ in +# place, which changes those bytes with no tampering involved; left +# alone, the next fire is refused as not matching the registered +# trust binding, and the watch dies silently. rebind-all is meant to +# run right after such an update. It still validates each watch's +# existing spec and trust chain exactly as an ordinary fire would +# (a watch already broken for some other reason is reported, not +# silently patched over), and it never touches an action executable +# outside FM_ROOT: rebinding follows this repo's own tracked +# update, never an arbitrary swapped action. Idempotent: a watch +# whose action bytes already match its binding is left alone. # run The blocking child the generic runner executes; never run it in a # conversational turn. It polls the condition on the registered # cadence, requires the stable count of consecutive trues, claims a @@ -78,6 +92,7 @@ SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" +FM_ROOT_REAL=$(cd "$FM_ROOT" 2>/dev/null && pwd -P) || FM_ROOT_REAL=$FM_ROOT # shellcheck source=bin/fm-pr-lib.sh . "$SCRIPT_DIR/fm-pr-lib.sh" @@ -367,7 +382,7 @@ cmd_run() { emit_doc "$sid" rejected "cannot stage command output; nothing was executed" 0 '' '' exit 0 fi - trap 'rm -f -- "$out"' EXIT + trap 'fm_procevent_source_lock_release "$sid"; rm -f -- "$out"' EXIT while :; do now=$(date +%s) @@ -415,15 +430,29 @@ cmd_run() { exit 0 fi - # Revalidate the registered action bytes immediately before claiming the - # fire. A changed or unavailable executable must never be run. + # Reload the trust binding from disk immediately before claiming the fire, + # rather than trusting the value cached at spec_load time when this poll + # loop started: a rebind-all can run (e.g. after a self-update) while this + # process is still polling, and only a fresh read sees its rebound hash. + # rebind_one publishes the spec and trust files as two separate renames, so + # the lock brackets this reload exactly as it brackets that publish, + # keeping the reader from observing a torn intermediate state. local current_action_hash + if ! fm_procevent_source_lock_acquire "$sid"; then + emit_doc "$sid" rejected "refused without executing anything: cannot lock the watch source" "$polls" '' '' + exit 0 + fi + if ! spec_load "$sid"; then + emit_doc "$sid" rejected "refused without executing anything: $SPEC_ERROR" "$polls" '' '' + exit 0 + fi current_action_hash=$(fm_pr_sha256 "${ACT_ARGV[0]}") || current_action_hash= if [ "$current_action_hash" != "$SPEC_ACTION_SHA256" ]; then emit_doc "$sid" rejected \ "refused without executing the action: its bytes do not match the registered trust binding" "$polls" '' '' exit 0 fi + fm_procevent_source_lock_release "$sid" # Claim the fire durably and exclusively BEFORE the action, so no restart or # concurrent runner can ever run the action a second time. @@ -473,6 +502,110 @@ cmd_terminal() { [ "$(cmd_classify "$file")" != unknown ] } +# --- rebind-all --------------------------------------------------------------- + +# publish_spec <sid> <device> <action_hash>: write and hash-bind a spec from +# the SPEC_* scalars and COND_ARGV/ACT_ARGV a prior spec_load already +# populated, using the given action hash. Mirrors cmd_arm's write block; the +# only caller today is rebind_one, refreshing action_sha256 alone. +publish_spec() { + local sid=$1 device=$2 action_hash=$3 tmp trust_tmp hash + tmp=$(umask 077; mktemp "$WHEN_DIR/.spec.XXXXXX") || return 1 + { + printf 'fm-when-spec-v1\n' + printf 'armed=%s\n' "$SPEC_ARMED" + printf 'interval=%s\n' "$SPEC_INTERVAL" + printf 'stable=%s\n' "$SPEC_STABLE" + printf 'deadline=%s\n' "$SPEC_DEADLINE" + printf 'condition_timeout=%s\n' "$SPEC_CONDITION_TIMEOUT" + printf 'action_timeout=%s\n' "$SPEC_ACTION_TIMEOUT" + printf 'error_budget=%s\n' "$SPEC_ERROR_BUDGET" + printf 'action_sha256=%s\n' "$action_hash" + printf 'condition_argc=%s\n' "${#COND_ARGV[@]}" + printf 'action_argc=%s\n' "${#ACT_ARGV[@]}" + printf 'argv:\n' + printf '%s\n' "${COND_ARGV[@]}" + printf '%s\n' "${ACT_ARGV[@]}" + } > "$tmp" || { rm -f -- "$tmp"; return 1; } + chmod 0600 "$tmp" || { rm -f -- "$tmp"; return 1; } + hash=$(fm_pr_sha256 "$tmp") || { rm -f -- "$tmp"; return 1; } + trust_tmp=$(umask 077; mktemp "$WHEN_DIR/.trust.XXXXXX") || { rm -f -- "$tmp"; return 1; } + printf 'fm-when-trust-v1\n%s\n' "$hash" > "$trust_tmp" || { rm -f -- "$tmp" "$trust_tmp"; return 1; } + chmod 0600 "$trust_tmp" || { rm -f -- "$tmp" "$trust_tmp"; return 1; } + mv -f -- "$tmp" "$(spec_file "$sid")" || { rm -f -- "$tmp" "$trust_tmp"; return 1; } + mv -f -- "$trust_tmp" "$(trust_file "$sid")" || { rm -f -- "$(spec_file "$sid")" "$trust_tmp"; return 1; } + if ! fm_pr_private_file_valid "$(spec_file "$sid")" 600 "$device" \ + || ! fm_pr_private_file_valid "$(trust_file "$sid")" 600 "$device"; then + rm -f -- "$(spec_file "$sid")" "$(trust_file "$sid")" + return 1 + fi +} + +# rebind_one <source-id>: 0 = rebound, 1 = failed (reported to stderr), 2 = +# unchanged or the action lives outside FM_ROOT (skipped, not an error). +rebind_one() { + local sid=$1 action_path action_hash device + if ! fm_procevent_source_lock_acquire "$sid"; then + printf 'skip: %s (cannot lock)\n' "$sid" >&2 + return 1 + fi + if ! spec_load "$sid"; then + printf 'skip: %s (%s)\n' "$sid" "$SPEC_ERROR" >&2 + fm_procevent_source_lock_release "$sid" + return 1 + fi + if ! action_path=$(action_executable "${ACT_ARGV[0]}"); then + printf 'skip: %s (action executable is unavailable: %s)\n' "$sid" "${ACT_ARGV[0]}" >&2 + fm_procevent_source_lock_release "$sid" + return 1 + fi + case "$action_path" in + "$FM_ROOT_REAL"/*) ;; + *) fm_procevent_source_lock_release "$sid"; return 2 ;; + esac + if ! action_hash=$(fm_pr_sha256 "$action_path"); then + printf 'skip: %s (cannot hash the action executable)\n' "$sid" >&2 + fm_procevent_source_lock_release "$sid" + return 1 + fi + if [ "$action_hash" = "$SPEC_ACTION_SHA256" ]; then + fm_procevent_source_lock_release "$sid" + return 2 + fi + if ! device=$(fm_pr_file_device "$WHEN_DIR"); then + printf 'skip: %s (cannot inspect the watch directory)\n' "$sid" >&2 + fm_procevent_source_lock_release "$sid" + return 1 + fi + if ! publish_spec "$sid" "$device" "$action_hash"; then + printf 'skip: %s (could not publish the refreshed trust binding)\n' "$sid" >&2 + fm_procevent_source_lock_release "$sid" + return 1 + fi + fm_procevent_source_lock_release "$sid" + printf 'rebound: %s\n' "$sid" + return 0 +} + +cmd_rebind_all() { + [ "$#" -eq 0 ] || usage + local spec sid rebound=0 skipped=0 failed=0 rc + [ -d "$WHEN_DIR" ] || { printf 'no watches registered\n'; return 0; } + for spec in "$WHEN_DIR"/when-*.spec; do + [ -e "$spec" ] || continue + sid=$(basename "$spec" .spec) + rebind_one "$sid" + rc=$? + case "$rc" in + 0) rebound=$((rebound + 1)) ;; + 2) skipped=$((skipped + 1)) ;; + *) failed=$((failed + 1)) ;; + esac + done + printf 'rebind-all: %s rebound, %s unchanged or out of scope, %s failed\n' "$rebound" "$skipped" "$failed" + [ "$failed" -eq 0 ] +} + # --- retire ------------------------------------------------------------------ cmd_retire() { @@ -499,6 +632,7 @@ case "${1-}" in terminal) shift; cmd_terminal "$@" ;; source-id) shift; cmd_source_id "$@" ;; retire) shift; cmd_retire "$@" ;; + rebind-all) shift; cmd_rebind_all "$@" ;; ''|-h|--help|help) usage ;; *) die "unknown command: $1" ;; esac diff --git a/bin/fm-update.sh b/bin/fm-update.sh index 621f82f7022..ce8aa279874 100755 --- a/bin/fm-update.sh +++ b/bin/fm-update.sh @@ -51,6 +51,14 @@ # A positively dead or missing endpoint has no agent to replace and is left to # the ordinary startup recovery. # +# A fast-forward that lands changes bytes under bin/ in place, which desyncs +# the trust binding of any locally armed fm-procevent-when watch whose action +# executable lives in the updated repo; left alone, the watch's next fire +# would be wrongly refused. After each home's own update (primary and every +# local secondmate), this script best-effort runs that home's own +# fm-procevent-when.sh rebind-all to republish those bindings against the new +# bytes; a failure there is swallowed rather than failing the update. +# # Usage: fm-update.sh [--help] set -eu @@ -78,8 +86,19 @@ fi reread_firstmate="no" ff_target "$FM_ROOT" "firstmate" origin no no -if [ "$FF_STATUS" = "updated" ] && [ -n "$FF_INSTR" ]; then - reread_firstmate="yes" +if [ "$FF_STATUS" = "updated" ]; then + if [ -n "$FF_INSTR" ]; then + reread_firstmate="yes" + fi + # A fast-forward changes bin/'s bytes out from under any locally armed + # fm-procevent-when watch's trust binding, with no tampering involved; left + # alone, the very next fire is refused and the watch dies silently. Refresh + # every such watch now, right after the update that broke it. FM_ROOT_OVERRIDE + # is passed explicitly rather than relying on the script's own location: this + # process's own FM_ROOT is the repo that was just updated, which is not + # always where this very script file happens to live (FM_ROOT_OVERRIDE, as + # this test suite uses to point fm-update.sh at a fixture checkout). + FM_HOME="$FM_HOME" FM_ROOT_OVERRIDE="$FM_ROOT" "$SCRIPT_DIR/fm-procevent-when.sh" rebind-all || true fi # --- secondmates ----------------------------------------------------------- @@ -137,6 +156,15 @@ claim_settled_secondmate() { # <id> # bin/fm-ff-lib.sh calls this for each local home it left AT the base with a live # endpoint - status "updated" or "current" alike. A skipped home never gets here. fm_ff_after_secondmate_settled() { # <id> <home> <window> <status> <instr> + # Same bin/-changed-out-from-under-a-watch problem as the primary home + # above, for a local secondmate's own worktree; "current" means bin/ did + # not move there this pass, so there is nothing to rebind. Run the + # secondmate's OWN copy of the script, explicitly overriding FM_ROOT to its + # own worktree rather than letting an outer FM_ROOT_OVERRIDE (this process's + # own, if the caller set one) leak into the child and misscope it. + if [ "${4:-}" = "updated" ] && [ -x "$2/bin/fm-procevent-when.sh" ]; then + FM_HOME="$2" FM_ROOT_OVERRIDE="$2" "$2/bin/fm-procevent-when.sh" rebind-all || true + fi claim_settled_secondmate "$1" } diff --git a/docs/configuration.md b/docs/configuration.md index d10624d2a77..e1797073646 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -793,7 +793,8 @@ Real feedback, ended and missing sessions, any other `SERVER_ERROR`, and that sa An already-armed Lavish source keeps its registered listener command until it is retired and armed again, so re-arm a live board once to adopt this retry policy. The `when` adapter (`bin/fm-procevent-when.sh`) turns this channel into a condition->action primitive: it registers a deterministic condition and a deterministic action once, its blocking child polls the condition without waking firstmate, and a stable true fires the action at most once before one terminal outcome is durably captured and published as a wake that remains eligible for re-announcement until handled. -The (condition, action) spec is stored privately under `state/when/` and hash-bound by a trust record the same way `bin/fm-check-register.sh` binds a custom check, while the spec separately binds the resolved action executable's bytes; a mutated or unregistered spec or a changed action executable is refused before the action runs. +The (condition, action) spec is stored privately under `state/when/` and hash-bound by a trust record the same way `bin/fm-check-register.sh` binds a custom check, while the spec separately binds the resolved action executable's bytes; a mutated or unregistered spec or a changed action executable is refused before the action runs, and that binding is reloaded from disk immediately before each fire rather than trusted from when polling started. +A repo update that fast-forwards an in-repo action's bytes in place would otherwise desync every already-armed watch's trust binding with no tampering involved; `bin/fm-procevent-when.sh rebind-all` re-hashes and republishes the binding for every registered watch whose action lives under `FM_ROOT`, including one already polling, so it keeps firing across such an update instead of being refused on its next fire. Every failure path - a mutated spec or action executable, a condition error past its budget, an expired deadline, a failed action, or an earlier fire whose outcome was never captured - produces a terminal captured outcome that wakes firstmate rather than a silent retry, and a durable single-fire marker claimed before the action makes restarts and re-polls unable to fire it twice. The adapter automates only the exact deterministic subset: anything needing judgment, and anything destructive, irreversible, or security-sensitive, keeps the ordinary check-fires-then-firstmate-decides flow, and the adapter's header and `--help` own its commands, flags, and outcome document. diff --git a/docs/verification/process-event-sources.md b/docs/verification/process-event-sources.md index 7e7c4312a28..52c0829aa19 100644 --- a/docs/verification/process-event-sources.md +++ b/docs/verification/process-event-sources.md @@ -141,6 +141,11 @@ Exercised by `tests/fm-procevent.test.sh` against a fake blocking source whose c | condition->action process bounds | the same suite proves action timeout terminates descendants and command-output staging remains within `FM_WHEN_OUTPUT_TAIL_BYTES` while the command runs | | silent failure handling | a nonzero exit with no output publishes nothing and leaves the source registered for retry | | inertness | a home with no registered source generates no state, starts no process, and does not need supervision | +| rebind-all refreshes an in-repo trust binding, leaves an out-of-repo one alone | after a simulated self-update rewrites an armed watch's in-repo action executable's bytes, `rebind-all` republishes exactly that watch's trust binding against the new bytes and leaves a watch whose action lives outside `FM_ROOT` untouched byte-for-byte; the rebound watch then fires cleanly against the new bytes instead of being refused, and a second `rebind-all` with nothing changed rebinds nothing (`tests/fm-procevent-when.test.sh`) | +| rebind-all matches FM_ROOT reached through a symlink | `FM_ROOT` and the resolved action executable are each canonicalized before the containment comparison, so a watch whose action is reached through a symlinked checkout path is still recognized as in-repo and rebound rather than silently skipped as out of scope (`tests/fm-procevent-when.test.sh`) | +| self-update rebinds a locally armed watch | `fm-update.sh` runs its own home's `rebind-all` best-effort immediately after a successful fast-forward of the primary repo or a local secondmate, so a watch armed against an in-repo action keeps firing across the update with no separate operator step (`tests/fm-update.test.sh`) | +| rebind-all reaches a watch already polling when the update lands | `run`'s poll loop calls `spec_load` once before entering its loop and would otherwise compare fire-time bytes against that stale in-memory hash forever; the fire-time check instead reloads the trust binding from disk immediately before firing, so a watch armed before a self-update still fires against the rebound bytes instead of being rejected as stale (`tests/fm-procevent-when.test.sh`) | +| fire-time reload serializes against rebind_one's publish | `publish_spec` renames the new spec into place and the new trust into place as two separate renames, never one atomic swap; the fire-time reload takes the same per-sid source lock `rebind_one` holds across that publish, so it can never observe the torn combination of rebound spec bytes next to a still-old trust record and instead waits for the publish to finish (`tests/fm-procevent-when.test.sh`) | | absent extension registry parity | `tests/fm-extension-binding.test.sh` drives `list` and `verify` in a fresh home while the current directory contains project files and Pi packages and an environment variable names fake package data; both commands report no bindings, create no home path, and discover nothing outside `config/extensions.d` | | complete package and binding identity | the same suite drives the public bind and verify commands through manifest duplicate/unknown/version failures, project and task-copy confinement, canonical path and symlink rejection, hard-link rejection, owner/mode checks, a non-executable entrypoint, binding mode drift, complete-tree mutation, exact executable mutation, and a missing executable; the foreign-owner fixture executes when the platform permits constructing another uid and otherwise reports that privilege limitation, while ordinary non-privileged CI does not exercise it or claim it ran | | external evidence write confinement | the same suite substitutes `state/procevent/` and `state/procevent-inbox/` with post-registration symlinks and proves an external start fails before bytes reach either outside target; it proves public lifecycle entry, environment, paths, and descriptors cannot forge capture authority; it proves live-generation claim release removes pending or consumed capture reservations only from the recorded revalidated state root, while a generation independently proved gone may leave an unreachable token-keyed reservation rather than wedging ownership; and it proves the absent-registry built-in capture path retains its legacy state-path behavior | diff --git a/tests/fm-procevent-when.test.sh b/tests/fm-procevent-when.test.sh index 396c48df8d9..36ddc955f78 100755 --- a/tests/fm-procevent-when.test.sh +++ b/tests/fm-procevent-when.test.sh @@ -17,6 +17,8 @@ set -u ROOT=$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd) TMP_ROOT=$(fm_test_tmproot fm-procevent-when-tests) export FM_PROCEVENT_CLAIM_ROOT="$TMP_ROOT/claims" +# shellcheck source=bin/fm-pr-lib.sh +. "$ROOT/bin/fm-pr-lib.sh" pe() { FM_HOME="$1" "$ROOT/bin/fm-procevent.sh" "${@:2}"; } when() { FM_HOME="$1" "$ROOT/bin/fm-procevent-when.sh" "${@:2}"; } @@ -390,4 +392,258 @@ assert_absent "$ACTION_TAMPER_LOG" "the mutated action was not executed" assert_absent "$H/state/when/when-action-tamper.fired" "no fire was claimed for mutated action bytes" pass "mutated action bytes are refused before claiming the fire" +# --- rebind-all refreshes a watch's action hash after a self-update ---------- +# A self-update fast-forwards bin/ in place, changing an in-repo action +# executable's bytes with no tampering involved. Without rebind-all the next +# fire is refused as "does not match the registered trust binding" (see the +# mutated-action-bytes case above); rebind-all exists to follow that update +# and republish a trust binding that matches the new bytes, but only for an +# action living under the simulated repo root, never for one outside it. +H="$TMP_ROOT/h-rebind"; new_home "$H" +REPO_ROOT="$TMP_ROOT/rebind-repo" +mkdir -p "$REPO_ROOT/bin" +IN_REPO_ACT="$REPO_ROOT/bin/act.sh" +cat > "$IN_REPO_ACT" <<'SH' +#!/usr/bin/env bash +log=$1 +echo v1 >> "$log" +SH +chmod +x "$IN_REPO_ACT" +OUT_OF_REPO_ACT="$TMP_ROOT/rebind-outside-act.sh" +cat > "$OUT_OF_REPO_ACT" <<'SH' +#!/usr/bin/env bash +log=$1 +echo v1 >> "$log" +SH +chmod +x "$OUT_OF_REPO_ACT" +when_ro() { FM_HOME="$1" FM_ROOT_OVERRIDE="$REPO_ROOT" "$ROOT/bin/fm-procevent-when.sh" "${@:2}"; } + +when_ro "$H" arm rebind-in-repo --interval 0.1 --stable 1 \ + --condition true --action "$IN_REPO_ACT" "$TMP_ROOT/rebind-in-repo.log" >/dev/null +when_ro "$H" arm rebind-out-of-repo --interval 0.1 --stable 1 \ + --condition true --action "$OUT_OF_REPO_ACT" "$TMP_ROOT/rebind-out-of-repo.log" >/dev/null + +SPEC_IN="$H/state/when/when-rebind-in-repo.spec" +TRUST_IN="$H/state/when/when-rebind-in-repo.trust" +TRUST_OUT="$H/state/when/when-rebind-out-of-repo.trust" + +OUT=$(when_ro "$H" rebind-all) || fail "rebind-all failed with nothing to rebind: $OUT" +assert_contains "$OUT" "0 rebound, 2 unchanged or out of scope, 0 failed" \ + "rebind-all should be a no-op before any action bytes change" + +# Simulate the self-update: rewrite both action scripts' bytes in place. +OLD_IN_REPO_SHA=$(fm_pr_sha256 "$IN_REPO_ACT") +cat > "$IN_REPO_ACT" <<'SH' +#!/usr/bin/env bash +log=$1 +echo v2 >> "$log" +echo "action ran v2 against $log" +SH +chmod +x "$IN_REPO_ACT" +NEW_HASH=$(fm_pr_sha256 "$IN_REPO_ACT") +[ "$OLD_IN_REPO_SHA" != "$NEW_HASH" ] || fail "test fixture error: mutation did not change the in-repo action's hash" +printf "#!/usr/bin/env bash\necho v2 >> \"\$1\"\n" > "$OUT_OF_REPO_ACT" +chmod +x "$OUT_OF_REPO_ACT" + +OLD_TRUST_OUT=$(cat "$TRUST_OUT") +OUT=$(when_ro "$H" rebind-all) || fail "rebind-all reported a failure: $OUT" +assert_contains "$OUT" "rebound: when-rebind-in-repo" "the in-repo watch was rebound" +assert_contains "$OUT" "1 rebound, 1 unchanged or out of scope, 0 failed" \ + "exactly the in-repo watch should rebind; the out-of-repo one stays out of scope" +[ "$(cat "$TRUST_OUT")" = "$OLD_TRUST_OUT" ] \ + || fail "rebind-all must never touch a watch whose action lives outside FM_ROOT" + +grep -qx "action_sha256=$NEW_HASH" "$SPEC_IN" \ + || fail "rebind-all did not record the action's current bytes in the spec" +SPEC_HASH=$(fm_pr_sha256 "$SPEC_IN") +TRUST_WANT=$(sed -n '2p' "$TRUST_IN") +[ "$SPEC_HASH" = "$TRUST_WANT" ] \ + || fail "the republished spec must still match its own trust binding" + +# The watch actually works again: a fresh run fires cleanly against the new +# bytes instead of being rejected. +pe "$H" reconcile >/dev/null +wait_for_result "$H" "when-rebind-in-repo" || fail "the rebound watch captured no outcome" +RESULT=$(first_result "$H" "when-rebind-in-repo") +assert_grep 'status: fired' "$RESULT" "the rebound watch fires instead of being rejected" +assert_grep 'action ran v2 against' "$RESULT" "the fired action ran the new bytes, not a stale copy" +pass "rebind-all refreshes an in-repo watch's trust binding after a self-update and leaves an out-of-repo one alone" + +# --- rebind-all matches an action reached through a symlinked FM_ROOT ------- +H="$TMP_ROOT/h-rebind-symlink"; new_home "$H" +REPO_REAL="$TMP_ROOT/rebind-symlink-real" +mkdir -p "$REPO_REAL/bin" +REPO_LINK="$TMP_ROOT/rebind-symlink-link" +ln -s "$REPO_REAL" "$REPO_LINK" +SYMLINK_ACT="$REPO_LINK/bin/act.sh" +cat > "$REPO_REAL/bin/act.sh" <<'SH' +#!/usr/bin/env bash +log=$1 +echo v1 >> "$log" +SH +chmod +x "$REPO_REAL/bin/act.sh" +when_symlink_ro() { FM_HOME="$1" FM_ROOT_OVERRIDE="$REPO_LINK" "$ROOT/bin/fm-procevent-when.sh" "${@:2}"; } + +when_symlink_ro "$H" arm rebind-symlink --interval 0.1 --stable 1 \ + --condition true --action "$SYMLINK_ACT" "$TMP_ROOT/rebind-symlink.log" >/dev/null + +cat > "$REPO_REAL/bin/act.sh" <<'SH' +#!/usr/bin/env bash +log=$1 +echo v2 >> "$log" +SH +chmod +x "$REPO_REAL/bin/act.sh" + +OUT=$(when_symlink_ro "$H" rebind-all) || fail "rebind-all reported a failure through a symlinked FM_ROOT: $OUT" +assert_contains "$OUT" "rebound: when-rebind-symlink" \ + "rebind-all must rebind an action reached through a symlinked FM_ROOT, not report it out of scope" +pass "rebind-all matches FM_ROOT through a symlinked checkout path" + +# --- rebind-all reaches a watch whose poller is already running ------------- +# The self-update race the fire-time revalidation targets: `run` calls +# spec_load once before entering its poll loop and caches the action hash in +# memory for the rest of its life. If the self-update (and its rebind-all) +# land while that poll loop is still running, only rewriting the on-disk spec +# and trust is not enough - the fire-time check must re-read the binding from +# disk, or the still-running poller compares against its stale in-memory hash +# and rejects a perfectly legitimate post-update fire. +H="$TMP_ROOT/h-live-rebind"; new_home "$H" +REPO_ROOT="$TMP_ROOT/live-rebind-repo" +mkdir -p "$REPO_ROOT/bin" +LIVE_ACT="$REPO_ROOT/bin/act.sh" +cat > "$LIVE_ACT" <<'SH' +#!/usr/bin/env bash +echo v1 >> "$1" +echo "v1 ran against $1" +SH +chmod +x "$LIVE_ACT" +LIVE_TRIGGER="$TMP_ROOT/live-rebind-trigger" +LIVE_COUNTER="$TMP_ROOT/live-rebind-count" +LIVE_LOG="$TMP_ROOT/live-rebind.log" +when_live_ro() { FM_HOME="$1" FM_ROOT_OVERRIDE="$REPO_ROOT" "$ROOT/bin/fm-procevent-when.sh" "${@:2}"; } + +when_live_ro "$H" arm live-rebind --interval 0.1 --stable 1 \ + --condition "$COND" "$LIVE_TRIGGER" "$LIVE_COUNTER" \ + --action "$LIVE_ACT" "$LIVE_LOG" >/dev/null + +# Start the poller now, before the simulated self-update, so its one-time +# spec_load caches the pre-update (v1) action hash in memory. +pe "$H" reconcile >/dev/null +wait_for_file "$LIVE_COUNTER" || fail "the live-rebind poller never evaluated its condition" + +# Simulate the self-update while that poller is still running: rewrite the +# action's bytes in place, then rebind-all republishes the on-disk trust +# binding to match. The already-running poller's in-memory hash is untouched. +cat > "$LIVE_ACT" <<'SH' +#!/usr/bin/env bash +echo v2 >> "$1" +echo "v2 ran against $1" +SH +chmod +x "$LIVE_ACT" +OUT=$(when_live_ro "$H" rebind-all) || fail "rebind-all reported a failure during a live poll: $OUT" +assert_contains "$OUT" "rebound: when-live-rebind" "the live watch's trust binding was rebound on disk" + +# Let the condition go true; the still-running poller must pick up the fresh +# binding at fire time instead of comparing against its stale cached hash. +: > "$LIVE_TRIGGER" +wait_for_result "$H" when-live-rebind || fail "the live poller never captured an outcome after rebind-all" +RESULT=$(first_result "$H" when-live-rebind) +assert_grep 'status: fired' "$RESULT" \ + "a watch whose poller was already running when rebind-all ran must still fire, not be rejected as stale" +assert_grep 'v2 ran against' "$RESULT" "the fired action ran the post-update bytes, not the ones cached at poll start" +pass "rebind-all reaches a watch whose run process was already polling when the self-update landed" + +# --- the fire-time reload never observes rebind_one's publish mid-rename ---- +# publish_spec is not an atomic swap: it renames the new spec into place, then +# separately renames the new trust into place. A `run` process reloading the +# binding at fire time must serialize against that window instead of reading +# a spec already rebound to v2 next to a trust record still bound to v1 - the +# exact torn combination that would otherwise report the rebind itself as a +# trust violation. This test builds that torn state under a held per-sid lock +# (the same lock rebind_one takes) so the reload's timing is deterministic, +# not a race that only sometimes reproduces. +H="$TMP_ROOT/h-torn-race"; new_home "$H" +TORN_ACT="$TMP_ROOT/torn-act.sh" +cat > "$TORN_ACT" <<'SH' +#!/usr/bin/env bash +echo v1 >> "$1" +echo "v1 ran against $1" +SH +chmod +x "$TORN_ACT" +TORN_TRIGGER="$TMP_ROOT/torn-race-trigger" +TORN_COUNTER="$TMP_ROOT/torn-race-count" +TORN_LOG="$TMP_ROOT/torn-race.log" +when "$H" arm torn-race --interval 0.05 --stable 1 \ + --condition "$COND" "$TORN_TRIGGER" "$TORN_COUNTER" \ + --action "$TORN_ACT" "$TORN_LOG" >/dev/null +SID=$(when "$H" source-id torn-race) +SPEC_TORN="$H/state/when/$SID.spec" +TRUST_TORN="$H/state/when/$SID.trust" + +# Start the poller now, with the condition still false, so reconcile's own +# brief use of this same per-sid lock (to claim and launch the source) is +# already done and released well before the holder below ever takes it. +pe "$H" reconcile >/dev/null +wait_for_file "$TORN_COUNTER" || fail "the torn-race poller never evaluated its condition" + +# Simulate the self-update, then build the rebound (v2) spec+trust pair ahead +# of time exactly as publish_spec would (same fields, only action_sha256 +# differs), so the background holder below only performs the two renames. +cat > "$TORN_ACT" <<'SH' +#!/usr/bin/env bash +echo v2 >> "$1" +echo "v2 ran against $1" +SH +chmod +x "$TORN_ACT" +NEW_HASH=$(fm_pr_sha256 "$TORN_ACT") +NEW_SPEC="$TMP_ROOT/torn-race-new.spec" +sed "s/^action_sha256=.*/action_sha256=$NEW_HASH/" "$SPEC_TORN" > "$NEW_SPEC" +NEW_SPEC_HASH=$(fm_pr_sha256 "$NEW_SPEC") +NEW_TRUST="$TMP_ROOT/torn-race-new.trust" +printf 'fm-when-trust-v1\n%s\n' "$NEW_SPEC_HASH" > "$NEW_TRUST" +chmod 0600 "$NEW_SPEC" "$NEW_TRUST" + +TORN_READY="$TMP_ROOT/torn-ready" +TORN_RELEASE="$TMP_ROOT/torn-release" +rm -f "$TORN_READY" "$TORN_RELEASE" +parent=$$ +FM_HOME="$TMP_ROOT/torn-race-lock-helper-home" bash -c ' + . "$1/bin/fm-pr-lib.sh" + . "$1/bin/fm-wake-lib.sh" + . "$1/bin/fm-procevent-lib.sh" + fm_procevent_source_lock_acquire "$2" || exit 1 + trap "fm_procevent_source_lock_release \"$2\"" EXIT + mv -f -- "$3" "$5" + printf "ready\n" > "$6" + while [ ! -e "$7" ]; do + kill -0 "$8" 2>/dev/null || exit 0 + sleep 0.02 + done + mv -f -- "$4" "$9" +' _ "$ROOT" "$SID" "$NEW_SPEC" "$NEW_TRUST" "$SPEC_TORN" "$TORN_READY" "$TORN_RELEASE" "$parent" "$TRUST_TORN" & +HOLDER_PID=$! + +wait_for_file "$TORN_READY" || fail "the torn-write holder never installed the rebound spec" +grep -qx "action_sha256=$NEW_HASH" "$SPEC_TORN" \ + || fail "test fixture error: the torn window did not actually install the rebound spec" +[ "$(sed -n '2p' "$TRUST_TORN")" != "$NEW_SPEC_HASH" ] \ + || fail "test fixture error: the trust file was rebound before the torn window began" + +# The still-running poller now sees its condition go true and reaches the +# fire-time reload while the torn state above is live and the lock is held. +: > "$TORN_TRIGGER" +sleep 0.3 +if first_result "$H" "$SID" >/dev/null 2>&1; then + fail "the reload must block on the source lock instead of reading the torn spec/trust pair" +fi + +: > "$TORN_RELEASE" +wait "$HOLDER_PID" 2>/dev/null || true +wait_for_result "$H" "$SID" || fail "the watch never captured an outcome after the torn window closed" +RESULT=$(first_result "$H" "$SID") +assert_grep 'status: fired' "$RESULT" \ + "the reload must wait past the torn spec/trust window, not reject a legitimate rebind mid-publish" +assert_grep 'v2 ran against' "$RESULT" "the fired action ran the rebound (v2) bytes, not a rejection from a torn read" +pass "the fire-time reload never observes rebind_one's spec/trust publish mid-rename" + printf 'all fm-procevent-when tests passed\n' diff --git a/tests/fm-update.test.sh b/tests/fm-update.test.sh index 39bce4dafba..bfc3ea143b7 100755 --- a/tests/fm-update.test.sh +++ b/tests/fm-update.test.sh @@ -471,6 +471,42 @@ test_unsafe_secondmate_home_skipped_before_git_update() { pass "T11 unsafe secondmate home is not fast-forwarded" } +# --- T12: a self-update rebinds a locally armed watch on the primary -------- +# A self-update fast-forwards bin/ in place, changing bytes an armed +# fm-procevent-when watch's trust binding was hashed against with no +# tampering involved; without a rebind the very next fire would be refused. +test_primary_update_rebinds_local_watch() { + local w before_hash after_hash out spec + w=$(new_world t12) + mkdir -p "$w/seed/bin" + printf "#!/usr/bin/env bash\necho v1 >> \"\$1\"\n" > "$w/seed/bin/watched-action.sh" + chmod +x "$w/seed/bin/watched-action.sh" + git -C "$w/seed" add -A + git -C "$w/seed" commit -qm add-watched-action + git -C "$w/seed" push -q origin main + git -C "$w/main" pull -q origin main + + FM_ROOT_OVERRIDE="$w/main" FM_HOME="$w/home" "$ROOT/bin/fm-procevent-when.sh" \ + arm rebind-primary --interval 60 --stable 1 \ + --condition true --action "$w/main/bin/watched-action.sh" "$w/rebind.log" >/dev/null + spec="$w/home/state/when/when-rebind-primary.spec" + before_hash=$(grep '^action_sha256=' "$spec") + + printf "#!/usr/bin/env bash\necho v2 >> \"\$1\"\n" > "$w/seed/bin/watched-action.sh" + git -C "$w/seed" add -A + git -C "$w/seed" commit -qm bump-watched-action + git -C "$w/seed" push -q origin main + + out=$(run_update "$w") + + assert_contains "$out" "firstmate: updated " "the primary still advanced" + assert_contains "$out" "rebound: when-rebind-primary" "the primary self-update rebound its own locally armed watch" + after_hash=$(grep '^action_sha256=' "$spec") + [ "$before_hash" != "$after_hash" ] \ + || fail "the watch's trust binding was not refreshed to match the updated action bytes" + pass "T12 a self-update rebinds a locally armed watch on the primary" +} + test_updates_main_and_secondmate test_reread_gate_is_instruction_only test_bin_only_advance_restarts @@ -485,5 +521,6 @@ test_registry_backstop_dedup_and_self_exclusion test_firstmate_wrong_branch_skipped test_firstmate_detached_head_skipped test_unsafe_secondmate_home_skipped_before_git_update +test_primary_update_rebinds_local_watch echo "# all fm-update tests passed" From 4107a1839ee51774a410da0e2429a84909a41c36 Mon Sep 17 00:00:00 2001 From: Sungin Kim <sunginapp@gmail.com> Date: Mon, 14 Sep 2026 17:17:38 +0000 Subject: [PATCH 29/31] fix(sync): bind expected policy exception to reviewed merge ancestry Apply the explicit sync landing decision relayed by Firstmate. Preserve all other check and authority gates, and pin scout golden capability input for CI. --- .agents/skills/sync-upstream/SKILL.md | 8 +- bin/fm-pr-merge.sh | 89 +++++++++++++++++++++- docs/fork-divergence.md | 2 + tests/fm-brief.test.sh | 4 + tests/fm-pr-merge.test.sh | 105 ++++++++++++++++++++++++++ 5 files changed, 202 insertions(+), 6 deletions(-) diff --git a/.agents/skills/sync-upstream/SKILL.md b/.agents/skills/sync-upstream/SKILL.md index 27c1365e586..6ded28772ba 100644 --- a/.agents/skills/sync-upstream/SKILL.md +++ b/.agents/skills/sync-upstream/SKILL.md @@ -69,8 +69,12 @@ A sync PR is expected to carry the red `PR must be raised via no-mistakes` check The round deliberately uses `direct-PR` because no-mistakes rebases onto `origin/main`, which would replay and linearize a merge-only branch. Every other required check must pass, the PR must be mergeable, and the body must carry the applicability table and fork-survival evidence before it is reported ready. -Register the ready PR through the normal task lifecycle and stop for the configured merge authority. -When that authority later approves landing, the normal merge handler must use: +Firstmate has standing authority to land upstream-sync rounds after its review of tests, lint, fork-preservation evidence and passing substantive CI. +The worker still reports the ready PR through the normal task lifecycle and stops without merging. +After that review, Firstmate records the exact round review declaration in the task metadata as specified by `github_verified_upstream_sync` in [`bin/fm-pr-merge.sh`](../../../bin/fm-pr-merge.sh). +That guard owns the narrow expected-policy-check exception and its identity and ancestry proofs; any new head needs a fresh review declaration. +Every other merge gate, including the existing away-authority gate, remains applicable. +The landing handler must use: ```sh bin/fm-pr-merge.sh <task-id> <full-PR-url> -- --merge diff --git a/bin/fm-pr-merge.sh b/bin/fm-pr-merge.sh index 8c6f9c8479e..13fee57563d 100755 --- a/bin/fm-pr-merge.sh +++ b/bin/fm-pr-merge.sh @@ -18,7 +18,8 @@ # --match-head-commit, so a push that lands between that read and the merge # fails the merge instead of landing commits nothing verified. Reading that # state needs gh and jq, and either one absent stops the merge before any -# state is recorded. An attended --allow-red <check-name> may be passed once, +# state is recorded. github_verified_upstream_sync below owns the narrow +# reviewed upstream-sync policy exception. An attended --allow-red <check-name> may be passed once, # with the name as a separate argument; it waives only checks with that exact # name, still requires every other check green, and still binds the head. It is # refused while the away-posture record exists, and it never @@ -801,14 +802,87 @@ github_checks_not_green() { ' 2>/dev/null || return 1 } +# A review declaration is task metadata, not a general red-check waiver. +# Firstmate records exactly one upstream_sync_review=<base>:<target>:<head> +# after reviewing this round's tests, lint, fork survival and substantive CI. +# These full commit IDs pin the fork base, upstream endpoint and reviewed head. +# Only a direct-PR fm-upstream-sync-* task in HelloWorldSungin/firstmate can +# consume it, and only with --merge. The task worktree must prove the live +# branch/head, fork push target, upstream ancestry and exactly one two-parent +# merge from that base to that endpoint, followed only by linear fix commits. +# The base must still equal the live default tip. No fetch or metadata write +# happens here. Missing, duplicate, stale or unreadable proof leaves checks red. +# This exempts only completed FAILURE CheckRuns named exactly +# "PR must be raised via no-mistakes"; pending/cancelled checks and status +# contexts never qualify. All ordinary authority, hold and merge gates remain. +github_upstream_sync_task() { + [ "$PR_HOST/$PR_PATH" = github.com/HelloWorldSungin/firstmate ] || return 1 + case "$ID" in fm-upstream-sync-*) ;; *) return 1 ;; esac + [ "$(sed -n 's/^mode=//p' "$META")" = direct-PR ] +} + +github_verified_upstream_sync() { + local json=$1 live_head=$2 review wt base target reviewed rest + local origin upstream branch merges merge parents upstream_line + github_upstream_sync_task || return 1 + [ "$FM_PR_GITHUB_CALLER_METHOD" = merge ] || return 1 + [ "${#ALLOW_RED[@]}" -eq 0 ] || return 1 + review=$(sed -n 's/^upstream_sync_review=//p' "$META") + wt=$(sed -n 's/^worktree=//p' "$META") + [ -d "$wt" ] || return 1 + base=${review%%:*}; rest=${review#*:} + target=${rest%%:*}; reviewed=${rest#*:} + fm_pr_head_valid "$base" && fm_pr_head_valid "$target" \ + && fm_pr_head_valid "$reviewed" || return 1 + [ "$review" = "$base:$target:$reviewed" ] || return 1 + [ "$reviewed" = "$live_head" ] && [ "$base" = "$github_judged_default_tip" ] || return 1 + branch=$(git -C "$wt" symbolic-ref --quiet --short HEAD) || return 1 + [ "$branch" = "fm/$ID" ] || return 1 + [ "$(git -C "$wt" rev-parse --verify HEAD)" = "$reviewed" ] || return 1 + origin=$(git -C "$wt" remote get-url --push --all origin) || return 1 + case "$origin" in + https://github.com/HelloWorldSungin/firstmate.git|git@github.com:HelloWorldSungin/firstmate.git) ;; + *) return 1 ;; + esac + upstream=$(git -C "$wt" remote get-url upstream) || return 1 + case "$upstream" in + https://github.com/kunchenguid/firstmate.git|git@github.com:kunchenguid/firstmate.git) ;; + *) return 1 ;; + esac + upstream_line=$(git -C "$wt" rev-list --first-parent refs/remotes/upstream/main) || return 1 + printf '%s\n' "$upstream_line" | grep -qxF "$target" || return 1 + git -C "$wt" merge-base --is-ancestor "$base" "$reviewed" || return 1 + if git -C "$wt" merge-base --is-ancestor "$target" "$base"; then return 1; fi + # First-parent traversal excludes upstream's own merge commits. + merges=$(git -C "$wt" rev-list --first-parent --min-parents=2 "$base..$reviewed") || return 1 + [ -n "$merges" ] || return 1 + merge=$merges + parents=$(git -C "$wt" show -s --format=%P "$merge" 2>/dev/null) || return 1 + [ "$parents" = "$base $target" ] || return 1 + printf '%s' "$json" | jq -e --arg branch "$branch" ' + .headRefName == $branch + and .headRepository.nameWithOwner == "HelloWorldSungin/firstmate" + and ([.statusCheckRollup[] | select(.name == "PR must be raised via no-mistakes")] | length > 0) + and all(.statusCheckRollup[]; + if (.name == "PR must be raised via no-mistakes" or .context == "PR must be raised via no-mistakes") + then .__typename == "CheckRun" and .status == "COMPLETED" + and (.conclusion == "FAILURE" or .conclusion == "SUCCESS" or .conclusion == "NEUTRAL" or .conclusion == "SKIPPED") + else true end) + ' >/dev/null 2>&1 +} + # Pre-merge conditions for a GitHub pull request, read from one live view. # Sets FM_PR_MERGE_HEAD to the verified head on success. github_verify_mergeable() { - local json fields line red name covered + local json fields line red name covered sync_policy=false local total=0 named=0 refusals='' local state='' draft='' mergeable='' merge_state='' live_head='' base='' - if ! json=$(gh pr view "$URL" --json state,isDraft,mergeable,mergeStateStatus,headRefOid,baseRefName,statusCheckRollup 2>/dev/null) \ + if github_upstream_sync_task && [ "$FM_PR_GITHUB_CALLER_METHOD" != merge ]; then + echo 'error: upstream-sync tasks require explicit --merge to preserve upstream parentage' >&2 + return 1 + fi + if ! json=$(gh pr view "$URL" --json state,isDraft,mergeable,mergeStateStatus,headRefOid,baseRefName,statusCheckRollup,headRefName,headRepository 2>/dev/null) \ || [ -z "$json" ]; then echo "error: could not read the GitHub pull request state before merging" >&2 return 1 @@ -873,10 +947,17 @@ FIELDS || refusals="$refusals - mergeStateStatus is DIRTY (conflicts) " + if github_verified_upstream_sync "$json" "$live_head"; then + sync_policy=true + printf 'verified: reviewed upstream-sync graph at %s permits only the completed no-mistakes policy failure\n' "$live_head" >&2 + fi uncovered='' while IFS= read -r name; do [ -n "$name" ] || continue covered=0 + if [ "$sync_policy" = true ] && [ "$name" = 'PR must be raised via no-mistakes' ]; then + covered=1 + fi if [ "${#ALLOW_RED[@]}" -gt 0 ]; then for check in "${ALLOW_RED[@]}"; do [ "$check" = "$name" ] && covered=1 @@ -897,7 +978,7 @@ EOF [ -z "$uncovered" ] || printf 'error: these checks are not green: %s\n' "$uncovered" >&2 return 1 fi - printf 'verified: %s is open and mergeable, with every required check green at head %s\n' \ + printf 'verified: %s is open and mergeable, with every unwaived check green at head %s\n' \ "$URL" "$live_head" >&2 FM_PR_MERGE_HEAD=$live_head FM_PR_GITHUB_BASE=$base diff --git a/docs/fork-divergence.md b/docs/fork-divergence.md index ee6cc88caf2..6894d2bac58 100644 --- a/docs/fork-divergence.md +++ b/docs/fork-divergence.md @@ -403,6 +403,8 @@ This fork carries the optional drift detector, bootstrap diagnostic, sync-round The feature is inert when no `upstream` git remote exists so upstream users do not acquire fork behavior merely by taking another change. The detector fetches only into a disposable repository and never changes the source repository's objects, refs, index, branch, or worktree. A sync request dispatches an isolated merge task and PR rather than merging in the primary copy or extending `/updatefirstmate` with merge behavior. +The standing Firstmate review and landing authority lives in [`sync-upstream`](../.agents/skills/sync-upstream/SKILL.md); the exact-head expected-policy exception is owned by `github_verified_upstream_sync` in [`bin/fm-pr-merge.sh`](../bin/fm-pr-merge.sh) and exercised by [`tests/fm-pr-merge.test.sh`](../tests/fm-pr-merge.test.sh). +This deliberately preserves direct-PR upstream parentage while retaining every substantive-check, publishing, task-hold, default-tip and away-authority gate. ### Repository-local validation evidence diff --git a/tests/fm-brief.test.sh b/tests/fm-brief.test.sh index 6e3c5ce75b2..bee49113fae 100755 --- a/tests/fm-brief.test.sh +++ b/tests/fm-brief.test.sh @@ -294,6 +294,10 @@ brief_fingerprint() { test_no_issue_briefs_match_exact_goldens() { local home actual id + local PATH="$TMP_ROOT/golden-bin:$PATH" + mkdir -p "$TMP_ROOT/golden-bin" + printf '%s\n' '#!/usr/bin/env bash' 'echo "lavish-axi 0.1.62"' > "$TMP_ROOT/golden-bin/lavish-axi" + chmod +x "$TMP_ROOT/golden-bin/lavish-axi" home="$TMP_ROOT/no-issue-golden-home" actual="$TMP_ROOT/no-issue-golden.actual" write_registry "$home" diff --git a/tests/fm-pr-merge.test.sh b/tests/fm-pr-merge.test.sh index 97fd46ea0c8..d80acf57427 100755 --- a/tests/fm-pr-merge.test.sh +++ b/tests/fm-pr-merge.test.sh @@ -4821,6 +4821,111 @@ SH test_degraded_view_cannot_answer_the_post_merge_question +# Real Git histories prove the exception's parentage; only forge transport is +# mocked. Every refusal must leave the forge merge command uncalled. +test_reviewed_upstream_sync_policy_exception() { + local case_dir scenario wt base target head id meta rc saved_tip url method filter + saved_tip=$DEFAULT_TIP + for scenario in attended away fix-forward absent-review stale-review wrong-target duplicate-review \ + wrong-mode wrong-task wrong-repo wrong-push wrong-upstream side-endpoint wrong-branch wrong-live-branch \ + wrong-live-repo wrong-local-head moved-base squash green-squash pending cancelled status-context lint-red lint-pending \ + extra-merge no-grant; do + case_dir=$(make_case "sync-policy-$scenario") + wt="$case_dir/wt" + id=fm-upstream-sync-2026-09-14-tail + url=https://github.com/HelloWorldSungin/firstmate/pull/96 + git init -q "$wt" + git -C "$wt" commit -qm root --allow-empty + git -C "$wt" checkout -qb upstream + git -C "$wt" commit -qm upstream --allow-empty + target=$(git -C "$wt" rev-parse HEAD) + git -C "$wt" checkout -qb "fm/$id" HEAD~1 + git -C "$wt" commit -qm fork --allow-empty + base=$(git -C "$wt" rev-parse HEAD) + git -C "$wt" merge -q --no-ff -m sync "$target" + if [ "$scenario" = extra-merge ]; then + git -C "$wt" checkout -qb extra "$base" + git -C "$wt" commit -qm extra --allow-empty + git -C "$wt" checkout -q "fm/$id" + git -C "$wt" merge -q --no-ff -m extra extra + fi + [ "$scenario" != fix-forward ] || git -C "$wt" commit -qm fix --allow-empty + head=$(git -C "$wt" rev-parse HEAD) + git -C "$wt" remote add origin https://github.com/HelloWorldSungin/firstmate.git + git -C "$wt" remote add upstream https://github.com/kunchenguid/firstmate.git + git -C "$wt" update-ref refs/remotes/upstream/main "$target" + if [ "$scenario" = side-endpoint ]; then + git -C "$wt" checkout -qb upstream-main "$target~1" + git -C "$wt" commit -qm mainline --allow-empty + git -C "$wt" merge -q --no-ff -m side-target "$target" + git -C "$wt" update-ref refs/remotes/upstream/main HEAD + git -C "$wt" checkout -q "fm/$id" + fi + DEFAULT_TIP=$base + add_gh_mocks "$case_dir" "$head" + write_github_red_json "$case_dir" "$head" 'PR must be raised via no-mistakes' + jq --arg branch "fm/$id" '. + {headRefName:$branch,headRepository:{nameWithOwner:"HelloWorldSungin/firstmate"}}' \ + "$case_dir/github-view.json" > "$case_dir/view.tmp" + mv "$case_dir/view.tmp" "$case_dir/github-view.json" + [ "$scenario" != wrong-task ] || id=ordinary-task + meta="$case_dir/state/$id.meta" + mv "$case_dir/state/task-x1.meta" "$meta" + sed 's/^mode=.*/mode=direct-PR/' "$meta" > "$case_dir/meta.tmp" + mv "$case_dir/meta.tmp" "$meta" + printf 'upstream_sync_review=%s:%s:%s\n' "$base" "$target" "$head" >> "$meta" + method=--merge + case "$scenario" in + away|no-grant) write_away_record "$case_dir" --grant "$id" + [ "$scenario" != no-grant ] || write_away_record "$case_dir" ;; + absent-review) sed '/^upstream_sync_review=/d' "$meta" > "$case_dir/meta.tmp"; mv "$case_dir/meta.tmp" "$meta" ;; + stale-review|wrong-target) + sed '/^upstream_sync_review=/d' "$meta" > "$case_dir/meta.tmp"; mv "$case_dir/meta.tmp" "$meta" + if [ "$scenario" = stale-review ]; then + printf 'upstream_sync_review=%s:%s:%s\n' "$base" "$target" "$base" >> "$meta" + else + printf 'upstream_sync_review=%s:%s:%s\n' "$base" "$base" "$head" >> "$meta" + fi ;; + duplicate-review) printf 'upstream_sync_review=%s:%s:%s\n' "$base" "$target" "$head" >> "$meta" ;; + wrong-mode) printf 'mode=no-mistakes\n' >> "$meta" ;; + wrong-repo) url=https://github.com/example/repo/pull/96 ;; + wrong-push) git -C "$wt" remote set-url --push origin https://github.com/kunchenguid/firstmate.git ;; + wrong-upstream) git -C "$wt" remote set-url upstream https://github.com/example/repo.git ;; + wrong-branch) git -C "$wt" checkout -qb unrelated ;; + wrong-local-head) git -C "$wt" commit -qm unreviewed --allow-empty ;; + moved-base) DEFAULT_TIP=$target ;; + squash|green-squash) method=--squash ;; + esac + case "$scenario" in + green-squash) filter='.statusCheckRollup[0].conclusion="SUCCESS"' ;; + pending) filter='.statusCheckRollup[0].status="IN_PROGRESS"' ;; + cancelled) filter='.statusCheckRollup[0].conclusion="CANCELLED"' ;; + status-context) filter='.statusCheckRollup += [{__typename:"StatusContext",context:"PR must be raised via no-mistakes",state:"PENDING"}]' ;; + lint-red) filter='.statusCheckRollup += [{__typename:"CheckRun",name:"lint",status:"COMPLETED",conclusion:"FAILURE"}]' ;; + lint-pending) filter='.statusCheckRollup += [{__typename:"CheckRun",name:"lint",status:"QUEUED",conclusion:null}]' ;; + wrong-live-branch) filter='.headRefName="unrelated"' ;; + wrong-live-repo) filter='.headRepository.nameWithOwner="example/repo"' ;; + *) filter='.' ;; + esac + jq "$filter" "$case_dir/github-view.json" > "$case_dir/view.tmp" + mv "$case_dir/view.tmp" "$case_dir/github-view.json" + rc=0 + run_pr_merge "$case_dir" "$id" "$url" -- "$method" \ + > "$case_dir/stdout" 2> "$case_dir/stderr" || rc=$? + case "$scenario" in + attended|away|fix-forward) + expect_code 0 "$rc" "sync-policy-$scenario: verified round should merge: $(cat "$case_dir/stderr")" + assert_logged_gh_merge "$case_dir" 96 HelloWorldSungin/firstmate --merge ;; + *) + [ "$rc" -ne 0 ] || fail "sync-policy-$scenario: invalid proof or check was accepted" + assert_no_grep 'pr merge' "$case_dir/gh.log" "sync-policy-$scenario: forge merge ran" ;; + esac + done + DEFAULT_TIP=$saved_tip + pass "fm-pr-merge limits the sync exception to reviewed graphs and completed policy failures" +} + +test_reviewed_upstream_sync_policy_exception + test_github_red_checks_refuse_and_allow_red_waives_named test_superseded_failed_check_run_no_longer_refuses test_check_runs_never_supersede_status_contexts From f515d1e725988fa0c9aa23957fc0966368af5210 Mon Sep 17 00:00:00 2001 From: Sungin Kim <sunginapp@gmail.com> Date: Mon, 14 Sep 2026 18:04:28 +0000 Subject: [PATCH 30/31] fix(ci): overlap isolated tests and settle fixture maintenance --- .github/workflows/ci.yml | 6 ++++-- bin/fm-test-run.sh | 9 +++++---- docs/fm-test-isolation-proof.md | 15 +++++++++++++++ docs/fm-test-portable-shards.md | 3 ++- docs/fork-divergence.md | 1 + tests/fm-remote-secondmate-trace-context.test.sh | 8 ++++++-- 6 files changed, 33 insertions(+), 9 deletions(-) diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 1de4c2e7e6c..99605b31a80 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -62,6 +62,8 @@ jobs: # Two duration-balanced portable parallel shards of the Phase 2 proven-isolated # set only. Composition owner: bin/fm-test-run.sh (docs/fm-test-portable-shards.md). + # Two admitted workers overlap this isolated work inside each existing job. + # docs/fm-test-isolation-proof.md owns the concurrency evidence. tests-portable-parallel-1: name: Behavior portable parallel 1 runs-on: ubuntu-latest @@ -101,7 +103,7 @@ jobs: run: | set -eu mkdir -p "$RUNNER_TEMP/fm-test" - bin/fm-test-run.sh --lane portable-parallel-1 \ + bin/fm-test-run.sh --lane portable-parallel-1 --jobs 2 \ --fail-on-gate-skip 'Pi extension typecheck prerequisite not found' \ --json "$RUNNER_TEMP/fm-test/fm-test-timing-portable-parallel-1.json" - name: Upload shard 1 timing artifact @@ -140,7 +142,7 @@ jobs: run: | set -eu mkdir -p "$RUNNER_TEMP/fm-test" - bin/fm-test-run.sh --lane portable-parallel-2 \ + bin/fm-test-run.sh --lane portable-parallel-2 --jobs 2 \ --json "$RUNNER_TEMP/fm-test/fm-test-timing-portable-parallel-2.json" - name: Upload shard 2 timing artifact if: always() diff --git a/bin/fm-test-run.sh b/bin/fm-test-run.sh index 92fa72f432c..660de20e9c2 100755 --- a/bin/fm-test-run.sh +++ b/bin/fm-test-run.sh @@ -205,10 +205,11 @@ PER_SCRIPT_TIMEOUT_SET= # with about nine minutes left for setup and runner variability. The script # bound stays unchanged; the serial job adopts upstream's 30-minute cap. # real-Herdr is 420s less its ~35s mean slot plus 480s, inside its 1200s step cap. -# portable-parallel is tighter: CI runs that lane serially, so ~545s less its -# ~45s mean slot plus 480s is ~980s, already past its 600s job cap before setup. -# That lane can therefore lose per-script attribution to a job cancellation. -# Raising the script bound would not remedy that enclosing-job limit. +# Portable-parallel CI overlaps the admitted isolated scripts, so its serial +# hint sum is no longer a job wall-time estimate. A late-starting hung script +# can still exhaust the enclosing job cap before its own bound, losing the +# artifact to job cancellation. The workflow owns worker count and job caps; +# raising the script bound would not remedy that enclosing-job limit. # # It is a guard, not a speed control: a HUNG script becomes a bounded failure # instead of an unbounded suite, which is the shape that silently outruns a diff --git a/docs/fm-test-isolation-proof.md b/docs/fm-test-isolation-proof.md index 6f0766a16bc..4afb7507660 100644 --- a/docs/fm-test-isolation-proof.md +++ b/docs/fm-test-isolation-proof.md @@ -20,6 +20,21 @@ This record owns concurrent isolation evidence for the portable parallel candida | failed | 0 | | wall duration | 113278 ms | +## Portable pool with two workers + +Verified on 2026-09-14 on Linux x86_64 with Git 2.54.0, Pi 0.85.1, TypeScript 7.0.2 and Ruby 3.4.9 available on PATH. +The command was `taskset -c 0,1 bin/fm-test-isolation-proof.sh --jobs 2`, with `--json` directed to a private evidence file. +The exact completion output was: + +```text +FM_ISOLATION_SUMMARY total=24 failed=0 concurrency=2 duration_ms=434026 +``` + +Every candidate ran without a gate skip, and the proof's Git-configuration and temporary-root isolation checks passed. +After completion, the proof root was absent and no live process retained a `TMPDIR` beneath it. +The two-CPU affinity bounds this local concurrency observation; it is not a measurement of hosted-runner speed or a guarantee of CI wall-time headroom. +[The CI workflow](../.github/workflows/ci.yml) owns its worker count and deadlines, while [portable shards](fm-test-portable-shards.md) explains how to interpret lane measurements. + ## Candidate set - `tests/fm-arm-pretool-check.test.sh` diff --git a/docs/fm-test-portable-shards.md b/docs/fm-test-portable-shards.md index fcc2768402e..99824d0e551 100644 --- a/docs/fm-test-portable-shards.md +++ b/docs/fm-test-portable-shards.md @@ -5,7 +5,8 @@ ## Verification inputs -Balance hints come from serial runs of the real lanes on `ubuntu-latest`. +Balance hints come from serial measurements of the real lanes on `ubuntu-latest`. +The current workflow overlaps the admitted isolated scripts within each parallel job; [`.github/workflows/ci.yml`](../.github/workflows/ci.yml) owns its worker count and unchanged job caps. The concurrent isolation proof in [fm-test-isolation-proof.md](fm-test-isolation-proof.md) establishes concurrency safety, not serial CI duration. Local timings are not interchangeable with CI timings: platform and machine load can affect each script differently and change their relative weights. diff --git a/docs/fork-divergence.md b/docs/fork-divergence.md index 6894d2bac58..18821a49156 100644 --- a/docs/fork-divergence.md +++ b/docs/fork-divergence.md @@ -204,6 +204,7 @@ The rationale beside the constant owns why 480s and what the bound costs each CI Upstream `kunchenguid/firstmate#4151` refreshes shared duration hints and raises the portable serial job cap to 30 minutes. The fork retains eight shards and takes per-script maxima across both parents, preserving its fork-only timing hints. The runner still owns the 480-second per-script bound and the updated margin arithmetic. +The fork overlaps already-admitted isolated scripts inside its existing portable parallel jobs; [`.github/workflows/ci.yml`](../.github/workflows/ci.yml) owns worker count and caps, with current concurrency evidence in [`docs/fm-test-isolation-proof.md`](../docs/fm-test-isolation-proof.md). Watcher triage cases are partitioned between `tests/fm-watch-triage.test.sh` and `tests/fm-watch-triage-waits.test.sh`, with shared case definitions in `tests/watch-triage-helpers.sh`, so suite growth does not weaken the per-script bound. diff --git a/tests/fm-remote-secondmate-trace-context.test.sh b/tests/fm-remote-secondmate-trace-context.test.sh index 8efa24291cd..2be7e5c7116 100755 --- a/tests/fm-remote-secondmate-trace-context.test.sh +++ b/tests/fm-remote-secondmate-trace-context.test.sh @@ -96,7 +96,11 @@ git -C "$REMOTE_ROOT" init -q -b main git -C "$REMOTE_ROOT" config user.email test@example.com git -C "$REMOTE_ROOT" config user.name Test git -C "$REMOTE_ROOT" add . -git -C "$REMOTE_ROOT" commit -qm 'remote fixture root' +# Complete fixture maintenance before the remote route clones this repository. +# A detached repack can remove loose objects while a local clone copies them. +# The gc setting also covers Git versions predating maintenance.autoDetach. +git -C "$REMOTE_ROOT" -c maintenance.autoDetach=false -c gc.autoDetach=false \ + commit -qm 'remote fixture root' cat > "$FAKEBIN/fake-ssh" <<'SH' #!/usr/bin/env bash @@ -307,4 +311,4 @@ try_flag 'requires a non-empty value' \ --secondmate --traceparent= pass "delivery: a parent-supplied carrier is accepted only for a secondmate launch and only as a strict W3C value" -echo "ALL TESTS PASSED" +printf '\nall fm-remote-secondmate-trace-context tests passed\n' From 66da6fb840582c7a2c256d6341dba8768938ed87 Mon Sep 17 00:00:00 2001 From: Sungin Kim <sunginapp@gmail.com> Date: Mon, 14 Sep 2026 18:27:53 +0000 Subject: [PATCH 31/31] test(procevent): settle completed poll before retirement assertion --- tests/fm-procevent.test.sh | 9 +++++++++ 1 file changed, 9 insertions(+) diff --git a/tests/fm-procevent.test.sh b/tests/fm-procevent.test.sh index 60ab8dc3125..c75f4c54ac5 100755 --- a/tests/fm-procevent.test.sh +++ b/tests/fm-procevent.test.sh @@ -2264,6 +2264,15 @@ assert_present "$HMISS/state/procevent/$missing_id.source" \ MISSING_CAPTURED=$(first_result "$HMISS" "$missing_id" || true) assert_contains "$("$ROOT/bin/fm-procevent-lavish.sh" classify "$MISSING_CAPTURED")" artifact-missing \ "the recurring result from a deleted artifact classifies as artifact-missing" +# Capture precedes runner exit and claim release. This case tests retiring the +# registered id after completed polls, not racing the runner's identity proof +# against its exit; live-runner retirement has separate blocking-source cases. +for _ in $(seq 1 100); do + [ ! -e "$FM_PROCEVENT_CLAIM_ROOT/$missing_id.claim" ] && break + sleep 0.1 +done +assert_absent "$FM_PROCEVENT_CLAIM_ROOT/$missing_id.claim" \ + "the completed missing-artifact runner releases its claim before retirement" # Retire by source id stops the recurring wakes (the wake text carries the id). # Retire returns only after the runner it owns is stopped and its claim # released, so the count taken after it is a settled baseline.