Repository navigation
fix(bin): retire stalled owned watcher child with bounded TERM/KILL - #2320
mayankasthana wants to merge 18 commits into
Conversation
…h handling beacon
…platform divergence
SIGKILL is never held pending for a stopped process on Linux: the bounded retirement's KILL kills the stopped watcher immediately, the arm's wait reaps it before the expected-pid hardening runs, and the retirement takes the released-lock shape (stale-beacon-retired), not the release-failed shape. The old Linux case block asserted the opposite and failed deterministically on the ubuntu runner; the platform-gated block had never run during macOS local validation.
|
Speaking as Kun's firstmate: Scheduled 3:10pm PT 8/23 pass. VISION.md read in full from current main VISION (inspected the stalled-owned-child retirement in Class: corrective. Security: none. No workflow-file / secret / injection risk. TERM/KILL is scoped to the arm-owned watcher process group after a stale-beacon predicate; fail-loud, not silent. Overlap / HOLD: CI / NM: HEAD Workflows: already approved historically (CI completed SUCCESS on 2026-08-17). Run IDs: 32033358684 (CI), 32033358622 (Require no-mistakes). No pending first-time-fork approval. What would help this PR land: rebase onto current Land-eligible rec: NO (merge conflicts; 63 behind; NM attestation mismatch; stale CI vs current main). Captain-flag NOW: no. Waiting on the author to rebase off current main, clear the conflicts, and re-stamp no-mistakes for the new HEAD. Not a captain-decision hold. |
…talled-child-retiremen-5d # Conflicts: # bin/fm-watch.sh # docs/watcher-continuity.md
…er-continuity docs
|
|
Speaking as Kun's firstmate: Scheduled 11:10am PT 8/24 pass. VISION.md read in full from current main VISION (re-inspected stalled-owned-child retirement in Class: corrective. Security: none. No workflow-file / secret / injection. TERM/KILL is scoped to the arm-owned watcher process group after a stale-beacon predicate; fail-loud. Greptile still flags a successor-lock race in Overlap: THIS HEAD vs last pass: conflicts cleared. GitHub mergeable=MERGEABLE, mergeStateStatus=UNSTABLE. ahead 14 / behind 0 (was ahead 10 / behind 63). I will not resolve conflicts — they are already gone — and I will not conflict-fix via cloud agent because the PR is still not otherwise auto-merge-ready. CI / NM: HEAD Workflows approved this pass: CI 32757334876, Require no-mistakes 32757334813. Not green at comment time. Land-eligible rec: NO (NM attestation mismatch vs THIS HEAD; CI not yet green). Captain-flag NOW: no. Waiting-on-author to regenerate |
|
Speaking as Kun's firstmate: recirc. Conflicts look cleared vs last pass (now MERGEABLE). Still not auto-eligible. class=corrective. Bounded TERM/KILL of an arm-owned stalled watcher child after a stale-beacon predicate. Restores intended liveness. VISION.md: restart-as-non-event aligns. Honest interface aligns. Scripts align. Authority n/a. Spine aligns. Vendor aligns. Scope aligns. This HEAD: Overlap: Waiting on author for a HEAD-matching attestation and green CI on this SHA. Not a captain-decision hold. |
Resolve the stalled-owned-child retirement against the landed stall-bound eviction (kunchenguid#5594): both knobs coexist (FM_WATCH_STALL_RETIRE_TIMEOUT for the owned-child retirement, fm_watcher_stall_bound for the attach-follow bound), the interrupt path keeps main's cleanup-ready guard before the bounded retirement, and the lost-race regression now arms through a copied bin dir with a never-exiting watcher child since main no longer runs the pre-lock check migration. clear_stale_recorded_watcher_lock now holds the steal mutex across a final ownership recheck and the removal, so a concurrent successor steal can no longer lose its freshly acquired lock in the verify window.
| # legitimately reclaimed as a dead-pid lock - a second winner in the | ||
| # ledger. Staying alive until all losers have finished makes the | ||
| # single-winner invariant hold deterministically. | ||
| while [ "$(wc -l < "$4" 2>/dev/null || true)" -lt "$5" ]; do |
There was a problem hiding this comment.
Multiple winners hang the test
If two contenders acquire the lock, both winners wait for 39 losing contributions, but only 38 contenders can contribute one. Neither winner exits, so the test hangs until an external timeout instead of reporting the lock failure. The stale-lock concurrency test has the same wait condition.
Intent
Land PR #2320 for issue #2251: retire a stalled owned watcher child with a bounded TERM/KILL sequence so a wedged watcher no longer requires a session restart. This update re-applies the fix onto current main (resolving the overlap with the landed stall-bound eviction #5594), closes the successor-lock race Greptile flagged by holding the steal mutex across the stale-lock removal, reworks the lost-race regression for the current startup sequence, and regenerates the pipeline attestation for the new head.
What Changed
bin/fm-watch-arm.shno longer treats a live PID as permanent health for a watcher it forked:wait_owned_childre-applies the identity-bound beacon predicate after initial readiness and, once the shared stale-beacon grace is reached, retires the child through the boundedretire_watch_childcontract (TERM, then KILL to the child's isolated process group afterFM_WATCH_STALL_RETIRE_TIMEOUT, default 2s, invalid or zero values falling back to 2). The arm then publishes the existing watcher-down recovery episode and exits with a typedFAILEDline so a persistent adapter can run its existing bounded retry instead of a primary session restart. The lost-race stand-down path and the confirmation-timeout path now run the same bounded retirement instead of an unboundedwait.clear_stale_recorded_watcher_locktakes the watcher's steal mutex and rechecks the recorded PID and identity under it before removing the stale lock, so a concurrent arm that steals the same stale lock can no longer race the removal; the release is refused (and recorded in the cycle-exit ledger) when the child survived the bound or no longer matches the expected owner.bin/fm-watch.shrefreshes the handling-successor beacon while it waits out a pending downtime marker. Docs gain theFM_WATCH_STALL_RETIRE_TIMEOUTknob indocs/configuration.mdand the retirement/release contract indocs/watcher-continuity.md; tests add a real-process SIGSTOP retirement counterfactual on both platforms and a bounded lost-race stand-down regression, with.claude/hooks/added to.gitignore.Risk Assessment
Testing
Verification targeted the change's runtime surface: the watcher-arm retirement contract in
bin/fm-watch-arm.sh(plus the smallbin/fm-watch.shhandling-successor pre-loop block the branch carries). Three focused suites that drive real processes ran green (42 + 31 + 2 cases, no failures): fm-watcher-lock, fm-watch-arm, and fm-watch-recovery-loop. On top of those I drove the product by hand twice in disposable homes — first against the real arm and the real watcher it forks, where a SIGSTOPped, stale-beacon owned watcher was retired in 3s (SIGKILL, exit 137 in the ledger), the stale singleton lock was released,state/.watcher-downbecamepending:downtime:*, and the arm exited 1 with its typed line; the same home's next arm then surfacedcheck: rearm-resurfaceand exited 0, i.e. no session restart; and a healthy owned watcher was left untouched for 14s (grace 5s, retirement bound 3s) with the arm alive and the beacon advancing. Second, adversarial work on the successor-lock race: a concurrent successor that steals the lock during the arm's retirement window kept its lock untouched, and a live foreign holder of the steal mutex made the arm refuse the removal by name rather than delete a lock it could not verify. Reviewer-visible evidence is the two CLI transcripts plus the three suite logs in the evidence directory; there is no UI surface in this change, so no screenshot/video artifact applies. No product defect surfaced; the one surface I could not exercise is the harness-owned auto-arm path (a live Pi / OpenCode / Claude Stop-hook session), which this host cannot stand up — its entry pointbin/fm-claude-stop-autoarm.shis unchanged by this branch, and the arm contract it depends on is what the transcripts above exercised.bash tests/fm-watcher-lock.test.sh/test_stopped_watcher_is_retired_and_rearms_without_session_restartcheck: rearm-resurfaceand exited 0; also the recovery half of `test_stopped_watcher_is_retired_and_rearms_without_s…bash tests/fm-watch-arm.test.sh/test_attached_arm_follows_a_slow_live_holderandtest_attached_arm_hands_a_stalled_holder_to_its_replacement(arm-suite.log); the watcher-side eviction is also…bash tests/fm-watch-arm.test.sh/test_lost_race_child_stand_down_is_bounded— driven against a copied bin dir whose fm-watch.sh never exits, asserting the arm exits on its own with `stalled befor…watcher: FAILED - watcher pid=11671 ... recovery state could not release stale ownership, ledger reason=stale-beacon-release-failed, lock still…bash tests/fm-watcher-lock.test.shcasesarm cleans child watcher and temp output on HUP,arm defers TERM until startup watcher can run its lock cleanup, `arm TERM bounds wait for stalled startu…bash tests/fm-watch-recovery-loop.test.sh— 2/2 ok: a resurfacing handling successor stays alive and supervises instead of going blind (recovery-loop-suite.log)Evidence: Live drive: wedged owned watcher retired, same-session recovery, healthy child not retired
Source: Live drive: wedged owned watcher retired, same-session recovery, healthy child not retired
Evidence: Live adversarial drive of the successor-lock race (both halves)
Source: Live adversarial drive of the successor-lock race (both halves)
Evidence: fm-watcher-lock suite (42 cases, includes the stalled-watcher retirement regression)
Source: fm-watcher-lock suite (42 cases, includes the stalled-watcher retirement regression)
Evidence: fm-watch-arm suite (31 cases, includes the lost-race and attached-holder bounds)
Source: fm-watch-arm suite (31 cases, includes the lost-race and attached-holder bounds)
Evidence: fm-watch-recovery-loop suite (handling successor still supervises)
Source: fm-watch-recovery-loop suite (handling successor still supervises)
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
bin/fm-watch-arm.sh:50- The merged header now contradicts itself about how long a started (owned) arm tolerates a slow child. Lines 41-47 (added by this change) say a live identity-matched child whose beacon reaches the shared stale-beacon grace is retired with the bounded TERM/KILL sequence, and docs/watcher-continuity.md:382 documents the same trigger ("reaches the shared stale-beacon grace"). But line 50 still carries upstream's clause "as a started arm waits out a slow child, until the lock changes or the beacon reaches fm_watcher_stall_bound" — i.e. 3x grace (fm_watcher_stall_bound, bin/fm-wake-lib.sh:150-154) — for the started path. The implementation uses GRACE (1x), so the started arm now retires at ~300s what the attached path still follows to ~900s. Concretely: a watcher whose main-loop check runs 300-900s without touching state/.last-watcher-beat (FM_WATCHER_STALE_GRACE/WATCHER_STALL_BOUND tolerate exactly that) is killed by this arm although the same evidence on an attached holder is followed. Same-class stale text: line 41 "On started it waits the child and propagates the wake reason", now superseded by lines 42-47. Remedy is a comment correction (align line 50/41 with the grace-based owned-child retirement); if the 3x tolerance was actually intended for owned children, the code needs the change instead - that choice is the author's.bin/fm-watch-arm.sh:797- start_owned_watchdog (line 763) introduces a second state-dir artifact, "$child_out.liveness", but only the two branches inside wait_owned_child remove it (lines 962 and 970). cleanup_child (line 793) removes just $child_out, so a HUP/TERM/INT handled while the watchdog has already written the stale beacon age leaves state/.watch-arm-output.XXXXXX.liveness behind. That window is reachable: the watchdog writes the file immediately, then sleeps STALL_RETIRE_TIMEOUT, and for a TERM-resistant child the arm keeps the file for up to STALL_RETIRE_TIMEOUT+3s before its own stale branch tears it down, so a signal landing in that window takes handle_arm_signal -> cleanup_child -> exit and the file survives. The repo already asserts on exactly this name pattern (tests/fm-watcher-lock.test.sh:1163: ! ls "$state"/.watch-arm-output.*), so the leak is an intermittent failure of that assertion as well as state-dir litter. Fix: remove the liveness sibling from cleanup_child too (guarded, since watchdog_status is empty on the confirm-timeout path).bin/fm-watch-arm.sh:762- Component: the arm-owned liveness watchdog subshell (start_owned_watchdog/stop_owned_watchdog, lines 762-791) plus its "$child_out.liveness" status file and its TERM/KILL trap. No intent requirement needs a second process: it exists only to evaluate owned_child_has_stale_beacon and then TERM/KILL the isolated group, which is a parallel copy of a rule wait_owned_child already owns. Its original justification ("Keep the arm itself in a raw wait so an actionable child close propagates immediately") no longer holds - commit 72c5d2a replaced that raw wait with the ATTACH_POLL poll loop the arm now runs (lines 916-924), so the arm is already free to evaluate that same predicate inline and call the existing retire_watch_child, which already implements the bounded TERM/KILL/group-sweep/reap shape. Net removal: ~30 lines, one background process, one state file, and the trap/sleep/wait choreography, with no behavior the intent requires (actionable closes are already detected only to ATTACH_POLL granularity, and the status file is the sole thing arming the retire deadline). This is an implementation choice, not a defect, so it needs the author's call.🔧 Fix applied.
7 issues (4 warnings, 3 infos) still open:
bin/fm-watch-arm.sh:50- The merged header now contradicts itself about how long a started (owned) arm tolerates a slow child. Lines 41-47 (added by this change) say a live identity-matched child whose beacon reaches the shared stale-beacon grace is retired with the bounded TERM/KILL sequence, and docs/watcher-continuity.md:382 documents the same trigger ("reaches the shared stale-beacon grace"). But line 50 still carries upstream's clause "as a started arm waits out a slow child, until the lock changes or the beacon reaches fm_watcher_stall_bound" — i.e. 3x grace (fm_watcher_stall_bound, bin/fm-wake-lib.sh:150-154) — for the started path. The implementation uses GRACE (1x), so the started arm now retires at ~300s what the attached path still follows to ~900s. Concretely: a watcher whose main-loop check runs 300-900s without touching state/.last-watcher-beat (FM_WATCHER_STALE_GRACE/WATCHER_STALL_BOUND tolerate exactly that) is killed by this arm although the same evidence on an attached holder is followed. Same-class stale text: line 41 "On started it waits the child and propagates the wake reason", now superseded by lines 42-47. Remedy is a comment correction (align line 50/41 with the grace-based owned-child retirement); if the 3x tolerance was actually intended for owned children, the code needs the change instead - that choice is the author's.bin/fm-watch-arm.sh:797- start_owned_watchdog (line 763) introduces a second state-dir artifact, "$child_out.liveness", but only the two branches inside wait_owned_child remove it (lines 962 and 970). cleanup_child (line 793) removes just $child_out, so a HUP/TERM/INT handled while the watchdog has already written the stale beacon age leaves state/.watch-arm-output.XXXXXX.liveness behind. That window is reachable: the watchdog writes the file immediately, then sleeps STALL_RETIRE_TIMEOUT, and for a TERM-resistant child the arm keeps the file for up to STALL_RETIRE_TIMEOUT+3s before its own stale branch tears it down, so a signal landing in that window takes handle_arm_signal -> cleanup_child -> exit and the file survives. The repo already asserts on exactly this name pattern (tests/fm-watcher-lock.test.sh:1163: ! ls "$state"/.watch-arm-output.*), so the leak is an intermittent failure of that assertion as well as state-dir litter. Fix: remove the liveness sibling from cleanup_child too (guarded, since watchdog_status is empty on the confirm-timeout path).bin/fm-watch-arm.sh:762- Component: the arm-owned liveness watchdog subshell (start_owned_watchdog/stop_owned_watchdog, lines 762-791) plus its "$child_out.liveness" status file and its TERM/KILL trap. No intent requirement needs a second process: it exists only to evaluate owned_child_has_stale_beacon and then TERM/KILL the isolated group, which is a parallel copy of a rule wait_owned_child already owns. Its original justification ("Keep the arm itself in a raw wait so an actionable child close propagates immediately") no longer holds - commit 72c5d2a replaced that raw wait with the ATTACH_POLL poll loop the arm now runs (lines 916-924), so the arm is already free to evaluate that same predicate inline and call the existing retire_watch_child, which already implements the bounded TERM/KILL/group-sweep/reap shape. Net removal: ~30 lines, one background process, one state file, and the trap/sleep/wait choreography, with no behavior the intent requires (actionable closes are already detected only to ATTACH_POLL granularity, and the status file is the sole thing arming the retire deadline). This is an implementation choice, not a defect, so it needs the author's call.bin/fm-watch.sh:2630- Component introduced by this change that no intent requirement needs: the FM_WATCH_HANDLING_SUCCESSOR pre-loop wait (if [ "${FM_WATCH_HANDLING_SUCCESSOR:-0}" = 1 ]; then touch beat; while handling_wait -lt 600; do touch beat; fm_recovery_marker_snapshot; case pending:downtime:*;; *) break;; esac; sleep 0.05; done; [ handling_wait -lt 600 ] || WATCHER_RECOVERY_PENDING=1; fi). It is not part of the arm-retirement fix; it arrived through this branch's merges of current main (base 96876db had it at bin/fm-watch.sh:816, it was never authored by a branch commit, and upstream deleted it as a bug in 3f03533 / PR fix: bound recovery announcements and preserve supervision #2733 precisely because it could keep a handling successor out of its poll loop). The block's original rationale (keep the beacon fresh so the arm's liveness watchdog did not retire a waiting handling successor) died in this run's fix round, which removed that watchdog entirely (commit ea786da). The tree's own contract now contradicts the block's existence: docs/watcher-continuity.md:195 states "A handling successor does not re-announce. It enters its poll loop immediately and keeps scanning signals, stale panes, and checks." Effect today is mostly latent, because bin/fm-watch.sh:2462 (main's arm_check) rewrites apending:downtime:episode toannounced:downtime:before this block runs, so thecaseusually breaks on the first pass; the live residue is the narrow window between that arm_check and this block, where a republication resolves the episode topending:downtime:again and the successor then holds the marker lock each iteration (bin/fm-wake-lib.sh:846-853) for up to 600 iterations - measured in ~30-55s - before it starts supervising. Remedy (the remedy, not the defect, is what needs authorisation: it deletes merge-preserved behaviour rather than adding to it): delete lines 2630-2644 and accept upstream's version of the hunk, so the successor enters the poll loop immediately as the in-tree doc and fix: bound recovery announcements and preserve supervision #2733 require. No intent requirement in the supplied goal ("re-applies the fix onto current main") needs this block, and its beacon-refresh purpose is already provided by the poll loop's own touch at bin/fm-watch.sh:2682.bin/fm-watch-arm.sh:737- The new retirement predicateowned_child_has_stale_beacondecides a child is stalled with[ "$(fm_path_age "$BEAT")" -ge "$GRACE" ], where GRACE is the arm's fixed defaultGRACE=${FM_GUARD_GRACE:-300}(line 134). The watcher layer derives its own staleness from the poll cadence:WATCHER_STALE_GRACE=${FM_WATCHER_STALE_GRACE:-${FM_GUARD_GRACE:-$(fm_poll_derived_grace "$POLL")}}(bin/fm-watch.sh:277), and bin/fm-wake-lib.sh:125-140 documents that a fixed 300 is known-wrong once the cadence reaches it. The arm already derives the other bound that way (STALL_BOUND=$(fm_watcher_stall_bound), line 154, = 3x the derived grace), so the two thresholds can diverge by 3x. Concrete sequence: a home configured FM_POLL=600 (a quiet/long-poll home; fm_poll_derived_grace exists for exactly this), armed by Pi/OpenCode - neither .pi/extensions/fm-primary-pi-watch.ts nor .opencode/plugins/fm-primary-watch-arm.js sets FM_GUARD_GRACE, so the arm keeps GRACE=300 while the watcher's stale grace is 660 and its stall bound is 1980. The arm forks the watcher, confirms it healthy (beacon fresh, age<300), prints "watcher: started", then polls every ATTACH_POLL; the watcher touches the beacon once per main-loop iteration (bin/fm-watch.sh:2682), so at t=300s the arm sees age>=300 while the watcher is still mid-cycle and healthy by its own standard (and would only be evicted at 1980). The arm now TERMs/KILLs it, publishes downtime and exits 1 - every cycle, forever - which is supervision churn of the same class this change exists to remove. Before this change the same GRACE fed only the readiness gate, where the same divergence was benign (the confirmation loop samples right after each touch). Sibling sites of the same invariant: the readiness gatefm_watcher_healthy "$STATE" "$WATCH" "$GRACE"(bin/fm-watch-arm.sh:343) and the attach-hold bound derived at line 154 (attach_and_wait, line 444) already disagree with line 737 for any FM_POLL>240 home. Earliest shared boundary: derive the grace once at line 134 the same way fm_watcher_stall_bound does (FM_WATCHER_STALE_GRACE -> FM_GUARD_GRACE -> fm_poll_derived_grace "${FM_POLL:-15}") so readiness, the owned-child retirement and the attach bound all read one value; the alternative narrower fix is to have every arm spawner pass FM_GUARD_GRACE, which the Pi/OpenCode paths do not do today. This is a correctness fix to the new predicate's input, not new state or a new subsystem.bin/fm-watch-arm.sh:424- Sibling sites of the contract text the previous fix round (ea786da) corrected in the header but did not finish, all of which still assert that a started arm waits out a slow child: (1) bin/fm-watch-arm.sh:424-428, the attach_and_wait comment - "while the holder is alive and the lock still names it under the same identity, it is a slow cycle, which a started arm tolerates by waiting on its child, so this arm keeps following it" - is now false: the started path retires its child once the beacon reaches GRACE (wait_owned_child, line 870), well before STALL_BOUND, so the parallel justification for the attached path no longer exists (the header's replacement rationale is that an attached arm owns no child to wait on, lines 48-51). (2) The status-line enumeration at bin/fm-watch-arm.sh:28-37, which the same round rewrote adjacent to, still omits the retirement lines this change introduces: "watcher: FAILED - watcher pid=<N> stopped advancing its beacon ..." (lines 884 and 890) and "watcher: FAILED - our child pid=<N> stalled before standing down ..." (line 945); these are the change's primary new outputs and callers grep them (bin/fm-claude-stop-autoarm.sh:386-390). (3) tests/fm-watcher-lock.test.sh:1454 and 1476 still name "the arm's own stale-beacon watchdog" (a comment used to justify FM_GUARD_GRACE=30, and a fail message), a process the same round deleted, which makes a future failure message point at a component that does not exist. Comment/contract text only; no behaviour.tests/fm-watch-arm.test.sh:1- The merge commit d0fc0a8 silently flipped this test's mode (git diff --summary 1f3e769..HEAD => "mode change 100755 => 100644 tests/fm-watch-arm.test.sh"); it is the only mode change in the change set and upstream main keeps 100755. No branch commit authored the flip, so it is a merge artifact rather than intent. Impact is low because bin/fm-test-run.sh runs scripts asbash "$script"(line 2474) and no script executes this file directly, but it diverges from every sibling test and breaks any direct./tests/fm-watch-arm.test.shinvocation. Fix: restore mode 100755.✅ **Test** - passed
✅ No issues found.
bash tests/fm-watcher-lock.test.sh/test_stopped_watcher_is_retired_and_rearms_without_session_restartcheck: rearm-resurfaceand exited 0; also the recovery half of `test_stopped_watcher_is_retired_and_rearms_without_s…bash tests/fm-watch-arm.test.sh/test_attached_arm_follows_a_slow_live_holderandtest_attached_arm_hands_a_stalled_holder_to_its_replacement(arm-suite.log); the watcher-side eviction is also…bash tests/fm-watch-arm.test.sh/test_lost_race_child_stand_down_is_bounded— driven against a copied bin dir whose fm-watch.sh never exits, asserting the arm exits on its own with `stalled befor…watcher: FAILED - watcher pid=11671 ... recovery state could not release stale ownership, ledger reason=stale-beacon-release-failed, lock still…bash tests/fm-watcher-lock.test.shcasesarm cleans child watcher and temp output on HUP,arm defers TERM until startup watcher can run its lock cleanup, `arm TERM bounds wait for stalled startu…bash tests/fm-watch-recovery-loop.test.sh— 2/2 ok: a resurfacing handling successor stays alive and supervises instead of going blind (recovery-loop-suite.log)bash tests/fm-watcher-lock.test.sh— 42 cases, all ok, includingtest_stopped_watcher_is_retired_and_rearms_without_session_restart,test_arm_term_bounds_wait_for_stalled_startup,test_arm_hup_cleans_child_and_temp_output,test_live_stalled_watch_lock_is_replaced_past_hard_boundbash tests/fm-watch-arm.test.sh— 31 cases, all ok, includingtest_lost_race_child_stand_down_is_bounded,test_attached_arm_follows_a_slow_live_holder,test_attached_arm_hands_a_stalled_holder_to_its_replacementbash tests/fm-watch-recovery-loop.test.sh— 2 cases, all ok (handling successor enters its poll loop and supervises)bash live-retirement-drive.sh— manual live drive of the real arm + real watcher: S1 bounded retirement, S2 same-session recovery, S3 no over-retirement of a healthy childbash live-successor-lock-race.sh— manual live adversarial drive: A concurrent successor's lock survives the stale-lock release; B a live steal-mutex holder forces the typedrecovery state could not release stale ownershiprefusalgit status --porcelainafter all runs — clean; no stray fm-watch/fm-watch-arm processes left behind✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.