From 10b93b2cc6f4241e87fccaee2e357c33a7347a53 Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Wed, 26 Aug 2026 23:55:39 -0700 Subject: [PATCH 01/25] fix(bin): recover Claude auto-arm from hung claims (#3156) * fix(bin): make Claude auto-arm continuity self-heal past a hung claim On a Claude primary, a Stop-hook auto-arm process that hung mid-arm held the single-flight owner lock with its epoch ledger frozen at outcome=arming, and the abandonment proof read any live lock holder in arming as legitimately deciding forever. Every later Stop firing exited 0 at the lock, the turn-end guard kept deferring to the hung owner as recovery under way, and the watcher was never auto-re-armed again for the rest of the session - supervision survived only on manual arms and lapsed between them (the 2026-08-26 watcher flap). Corrections layered onto the lock-held-across-arm shape each reopened the same concurrency class one level down, so this replaces the claim machinery wholesale with a generation-based optimistic design: - The epoch ledger's monotonic sequence IS the claim generation; the two-line entry (classic epoch record plus the claimant's MANDATORY pid-identity) is the claim. Every firing defers to a live OPEN claim: outcome arming, owner alive, identity recomputes and matches, and not stuck (entry and watcher beacon both older than the guard grace). - A finished, dead, identity-mismatched, identityless, or stuck claim is superseded by simply taking the next generation - no signalling or revocation of a steady-state predecessor. - No mutex is held across arming or output; the owner lock survives only as a micro-mutex around individual ledger writes. A superseded owner goes completely silent: ownership is re-verified before every arm invocation, episode-state mutation, ledger write, and continuation. - The irrevocable commit point of a translation is the exit status (the harness delivers the collected stderr only on exit 2), so the owned terminal ledger write is the atomic commit: the winning generation exits 2 unconditionally after it, a refused one exits 0 silently even after printing, and the once-per-episode failure notice commits in the same owned critical section as the winning failed write. Two bounded residuals are documented accepted intent: an owner dying between its owned write and its own exit, and a hung old-build owner resuming during the one legacy upgrade window. - The pre-generation lock-holding claim shape keeps defer-or-reclaim behavior through a legacy shim: a live identity-verified stuck owner is retired via TERM (with a queued TERM sufficient when the owner is stopped) before its lock is removed, an unverified or identityless pid is never signalled but never blocks a proven-abandoned reclaim, and the lock's identity evidence is grafted into the ledger (mtime-preserving) so pid-reuse protection survives the lock. - The guard reads the same predicates for recovery ownership and its terminal fail-open (which re-checks for a live open claim under the held locks before committing the attended alarm), with ledger reads anchored to line 1 so the identity line can never confuse them. Behavioral regression coverage exercises all three edge classes through the real hook and guard - a live open claim defers with no lock held, a stuck claim is superseded and the home re-arms, and an end-to-end run with a genuinely hung owner shows a concurrent firing deferring promptly mid-arm, a later firing superseding the stuck owner, and the superseded owner exiting silently without a second translation - plus the identityless/reused-pid loopholes, the superseded-owner arm boundary, and the legacy TERM, SIGSTOP, and signal-free reclaim paths. * no-mistakes(review): Refuse auto-arm commits when notice marker creation fails * no-mistakes(review): Make episode reset atomic with generation ownership * no-mistakes(document): Update auto-arm generation and commit documentation --- bin/fm-claude-stop-autoarm.sh | 182 +++++++----- bin/fm-turnend-guard.sh | 67 +++-- bin/fm-wake-lib.sh | 375 ++++++++++++++++++++----- docs/configuration.md | 2 +- docs/supervision-protocols/claude.md | 2 +- docs/turnend-guard.md | 20 +- docs/watcher-continuity.md | 10 +- tests/fm-claude-stop-autoarm.test.sh | 397 ++++++++++++++++++++++++++- tests/fm-turnend-guard.test.sh | 75 +++++ 9 files changed, 942 insertions(+), 188 deletions(-) diff --git a/bin/fm-claude-stop-autoarm.sh b/bin/fm-claude-stop-autoarm.sh index 89ce011f6bb..762b1a3dcf4 100755 --- a/bin/fm-claude-stop-autoarm.sh +++ b/bin/fm-claude-stop-autoarm.sh @@ -20,21 +20,31 @@ # translation time so a mid-cycle AFK transition is honored). # - Need: arms only while work is in flight (state/*.meta) or X mode has a # relay poll to run (state/x-watch.check.sh); an idle home exits 0. -# - Single-flight: Claude does not dedupe async hooks, so a home-scoped owner -# lock (state/.claude-autoarm.lock) admits exactly one owner; every other -# concurrent firing exits 0 without translating, which keeps one event -# epoch on exactly one recovery turn. A lock left behind by a claim whose -# ledger outcome is already terminal, or whose recorded pid-identity no -# longer matches its live pid, is reclaimed once rather than deferred to -# forever (fm_autoarm_claim_abandoned in bin/fm-wake-lib.sh). +# - Single-flight: Claude does not dedupe async hooks, so exactly one +# GENERATION owner arms per event epoch: the epoch ledger's monotonic +# sequence is the claim generation, every firing defers (exit 0) to a live +# open claim, and a stuck, dead, identity-mismatched, or finished claim is +# superseded by taking the next generation instead of being unlocked or +# revoked. No mutex is ever held across arming or output - the owner lock +# survives only as the micro-mutex serializing individual ledger writes - +# and a superseded owner goes completely silent: ownership is re-verified +# before every arm invocation, episode-state mutation, ledger write, and +# continuation (fm_autoarm_claim_open/fm_autoarm_claim_next in +# bin/fm-wake-lib.sh own the contract, including the legacy shim for a +# pre-generation lock). # - Foreground arm: the owner runs bin/fm-watch-arm.sh in the FOREGROUND of # this hook-owned process tree (never shell &); Claude owns the process # group, so its timeout/session teardown kills arm and watcher together. # - Translation: while supervision is still needed and AFK remains inactive, # an actionable arm close (signal:/stale:/check:/heartbeat) prints one # rewake banner to stderr and exits 2, which wakes Claude even while idle -# ("Stop hook feedback"). A close that reports no actionable reason is -# benign when a live identity-matched watcher still has a fresh beacon. +# ("Stop hook feedback"). The irrevocable commit point is the EXIT STATUS: +# the harness delivers the collected stderr only on exit 2, so an owned +# terminal commit decides the exit. Markerless outcomes commit with the +# ledger write; the failure notice additionally requires its marker write. +# A refused generation exits 0 silently even after printing. A close that +# reports no actionable reason is benign when a live identity-matched +# watcher still has a fresh beacon. # - Failure handling: a typed failure is rechecked against the same live, # fresh watcher predicate and retried a bounded number of times in this # hook. Only an exhausted failure with no verified watcher emits one @@ -42,10 +52,11 @@ # exit 2 to guarantee the next Stop-owned retry without repeating notice, # until the synchronous guard has consumed its attended fail-open. # -# The epoch ledger state/.claude-autoarm-epoch records the latest claim and -# outcome so the synchronous Stop guard (bin/fm-turnend-guard.sh --claude) can -# allow a stop whose recovery this hook already owns, instead of forcing a -# duplicate continuation for the same event epoch. The failure marker +# The epoch ledger state/.claude-autoarm-epoch records the latest claim +# generation and outcome so the synchronous Stop guard +# (bin/fm-turnend-guard.sh --claude) can allow a stop whose recovery this hook +# already owns, instead of forcing a duplicate continuation for the same event +# epoch. The failure marker # state/.claude-autoarm-failure-notified deduplicates the last-resort notice, # and state/.claude-autoarm-failure-alarmed bounds the attended fail-open and # suppresses any later automatic continuation in that unresolved episode. @@ -64,7 +75,6 @@ STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" CONFIG="${FM_CONFIG_OVERRIDE:-$FM_HOME/config}" GRACE=${FM_GUARD_GRACE:-300} OWNER_LOCK="$STATE/.claude-autoarm.lock" -EPOCH="$STATE/.claude-autoarm-epoch" FAILURE_NOTICE="$STATE/.claude-autoarm-failure-notified" FAILURE_ALARM="$STATE/.claude-autoarm-failure-alarmed" AUTOARM_ATTEMPTS=${FM_CLAUDE_AUTOARM_ATTEMPTS:-2} @@ -134,49 +144,51 @@ if [ "$RECOVER_SESSION_LOCK" -eq 1 ]; then fm_session_lock_owned_by_self "$STATE" || exit 0 fi -# --- single-flight owner claim ------------------------------------------------ +# --- single-flight generation claim -------------------------------------------- # Claude runs one background process per firing with no dedupe. Exactly one -# owner foregrounds the arm and translates its close; every other firing exits -# 0 so one watcher cycle maps to at most one exit-2 rewake. -# -# A claim whose own ledger entry or recorded pid-identity proves its supervision -# decision already finished is abandoned, not in flight: deferring to it forever -# is what leaves a home unsupervised with no watcher and no lock -# (fm_autoarm_claim_abandoned in bin/fm-wake-lib.sh owns that proof and its -# race-free reclaim). Reclaim it once and retry; anything still genuinely -# deciding keeps the lock and this firing stays inert. -if ! fm_lock_try_acquire "$OWNER_LOCK"; then - fm_autoarm_release_abandoned "$STATE" || exit 0 - fm_lock_try_acquire "$OWNER_LOCK" || exit 0 -fi -# Record WHO this claim is before publishing the role both Stop participants read -# as ownership. A bare pid the operating system later hands to an unrelated live -# process is exactly what makes a killed claim look in flight forever, in the two -# shapes the ledger cannot settle: an entry still reading arming, and no entry at -# all. Best effort; a home whose identity cannot be recorded keeps the ledger-only -# boundary rather than losing its claim. -fm_autoarm_claim_record_identity "$STATE" || true -if ! fm_lock_set_role "$OWNER_LOCK" autoarm; then - fm_lock_release "$OWNER_LOCK" - exit 0 +# generation owner arms and translates per event epoch: every firing defers to +# a live open claim, and a stuck, dead, identity-mismatched, or finished claim +# is superseded by taking the next generation (fm_autoarm_claim_open and +# fm_autoarm_claim_next in bin/fm-wake-lib.sh own the contract). No mutex is +# held past this point. A micro-mutex contention with a bare hold is another +# participant's short ledger section and the next Stop firing simply retries, +# while a role-carrying hold is a legacy lock-holding claim from a +# pre-generation build (or the guard's own terminal-check), which the legacy +# shim defers to while genuinely deciding and reclaims once when proven +# abandoned. +fm_autoarm_claim_open "$STATE" "$GRACE" && exit 0 +fm_autoarm_claim_next "$STATE" "$GRACE" +CLAIM_RC=$? +if [ "$CLAIM_RC" -ne 0 ]; then + [ "$CLAIM_RC" -eq 2 ] && exit 0 + ROLE=$(fm_lock_role "$OWNER_LOCK" 2>/dev/null || true) + [ -n "$ROLE" ] || exit 0 + fm_autoarm_release_abandoned "$STATE" "$GRACE" || exit 0 + fm_autoarm_claim_next "$STATE" "$GRACE" || exit 0 fi -trap 'fm_lock_release "$OWNER_LOCK"' EXIT +MY_GEN=$FM_AUTOARM_MY_GEN +[ -n "$MY_GEN" ] || exit 0 -write_epoch() { # - local outcome=$1 seq tmp - seq=$(sed -n 's/^epoch=\([0-9][0-9]*\) .*/\1/p' "$EPOCH" 2>/dev/null || true) - case "$seq" in - ''|*[!0-9]*) seq=0 ;; - esac - seq=$((seq + 1)) - tmp="$EPOCH.tmp.$$" - printf 'epoch=%s owner_pid=%s outcome=%s updated_at=%s\n' \ - "$seq" "${BASHPID:-$$}" "$outcome" "$(date +%s)" > "$tmp" 2>/dev/null \ - && mv -f "$tmp" "$EPOCH" 2>/dev/null - rm -f "$tmp" 2>/dev/null || true +# Commit (optionally with the once-per-episode notice marker) for +# this generation. Success means this generation's translation WINS and the +# caller exits 2 unconditionally. Markerless outcomes commit with the owned +# ledger write; a notice wins only when its following marker write succeeds in +# the same hold. Failure means refused or unverifiable: the caller goes silent +# (cleanup, exit 0) - the harness discards the collected stderr on exit 0, so +# even an already-printed banner is never delivered by a losing generation. +autoarm_commit() { # [marker-file] + if [ -n "${2:-}" ]; then + fm_autoarm_write_owned "$STATE" "$MY_GEN" "$1" "$2" + else + fm_autoarm_write_owned "$STATE" "$MY_GEN" "$1" + fi } -write_epoch arming +# Best-effort ownership-checked record for exit-0 paths, where supersession +# changes nothing about the action taken. +autoarm_record() { # + fm_autoarm_write_owned "$STATE" "$MY_GEN" "$1" >/dev/null 2>&1 || true +} # X mode cadence: source the generated config so an X instance polls at its # 30s cadence (fm-bootstrap.sh x_mode_setup contract). @@ -195,6 +207,13 @@ ACTIONABLE=0 HEALTHY=0 attempt=0 while [ "$attempt" -lt "$AUTOARM_ATTEMPTS" ]; do + # A superseded owner must not start or attach another watcher or mutate any + # watcher/wake state: re-verify generation ownership before every arm + # invocation, first attempt and retries alike. + if ! fm_autoarm_still_owner "$STATE" "$MY_GEN"; then + [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true + exit 0 + fi attempt=$((attempt + 1)) OUT=$(mktemp "$STATE/.claude-autoarm-output.XXXXXX") || OUT= if [ -n "$OUT" ]; then @@ -206,7 +225,7 @@ while [ "$attempt" -lt "$AUTOARM_ATTEMPTS" ]; do # AFK may have appeared mid-cycle: the daemon owns triage now, so suppress # every subsequent classification and handoff. if [ -e "$STATE/.afk" ]; then - write_epoch afk + autoarm_record afk [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true exit 0 fi @@ -231,56 +250,85 @@ done # The need may have vanished mid-cycle (fleet torn down, X opted out): nothing # left to supervise, so close quietly instead of waking the model. if ! need_supervision; then - write_epoch clean + autoarm_record clean [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true exit 0 fi if [ "$HEALTHY" -eq 1 ]; then - if fm_failure_episode_reset "$STATE"; then - write_epoch clean + fm_autoarm_reset_owned "$STATE" "$MY_GEN" + RESET_RC=$? + if [ "$RESET_RC" -eq 0 ]; then + autoarm_record clean + [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true + exit 0 + fi + if [ "$RESET_RC" -eq 2 ]; then [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true exit 0 fi - write_epoch failed-suppressed + if autoarm_commit failed-suppressed; then + [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true + [ -e "$FAILURE_ALARM" ] && exit 0 + exit 2 + fi [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true - [ -e "$FAILURE_ALARM" ] && exit 0 - exit 2 + exit 0 fi # After the synchronous guard has consumed the episode's attended fail-open, # do not create another exit-2 continuation that could defeat it. if [ -e "$FAILURE_ALARM" ]; then - write_epoch failed-suppressed + autoarm_record failed-suppressed [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true exit 0 fi if [ "$ACTIONABLE" -eq 1 ]; then - write_epoch rewake + # Cheap early-out before composing the banner; the real commit decision is + # the owned terminal write below. + if ! fm_autoarm_still_owner "$STATE" "$MY_GEN"; then + [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true + exit 0 + fi { printf 'firstmate watcher wake - one supervision event needs a handling turn now.\n' [ -n "$OUT" ] && grep -E '^(signal:|stale:|check:|heartbeat)' "$OUT" 2>/dev/null | head -8 printf 'Run bin/fm-wake-drain.sh first, handle the wake, then run its exact WAKE_ACK_REQUIRED --ack-through command. Until that post-handling acknowledgement, interruption leaves the wake durable for idempotent re-handling. This Stop hook owns watcher continuity: when the handling turn ends, the next needed cycle arms automatically - do NOT run bin/fm-watch-arm.sh after an ordinary wake.\n' } >&2 + if autoarm_commit rewake; then + [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true + exit 2 + fi [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true - exit 2 + exit 0 fi # Notify only once for this continuous failure episode; every later invocation # still exits 2 so Claude must continue into another Stop-owned retry without -# creating a repeated operator notice or manual-arm loop. +# creating a repeated operator notice or manual-arm loop. The notice marker +# commits in the same owned critical section as the winning failed write, so a +# losing generation can neither consume nor deliver it. if [ ! -e "$FAILURE_NOTICE" ]; then - write_epoch failed + if ! fm_autoarm_still_owner "$STATE" "$MY_GEN"; then + [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true + exit 0 + fi { printf 'firstmate watcher auto-arm FAILED - the Stop-owned automatic supervision mechanism is broken after %s bounded attempts, and no live watcher with a fresh beacon was verified.\n' "$attempt" [ -n "$OUT" ] && grep -E '^(watcher:|signal:|stale:|check:|heartbeat)' "$OUT" 2>/dev/null | head -8 printf 'Do not launch a manual background arm from this notice; investigate the automatic Stop hook and watcher startup before ending blind.\n' } >&2 - : > "$FAILURE_NOTICE" 2>/dev/null || true + if autoarm_commit failed "$FAILURE_NOTICE"; then + [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true + exit 2 + fi + [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true + exit 0 +fi +if autoarm_commit failed-suppressed; then [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true exit 2 fi -write_epoch failed-suppressed [ -z "$OUT" ] || rm -f "$OUT" 2>/dev/null || true -exit 2 +exit 0 diff --git a/bin/fm-turnend-guard.sh b/bin/fm-turnend-guard.sh index 43e70457060..7d9601308af 100755 --- a/bin/fm-turnend-guard.sh +++ b/bin/fm-turnend-guard.sh @@ -51,9 +51,9 @@ # auto-arm (bin/fm-claude-stop-autoarm.sh), which fires on the same Stop event: # 1. a live identity-matched watcher with a fresh beacon allows immediately; # 2. otherwise wait briefly (FM_CLAUDE_AUTOARM_SYNC_WAIT_MS, default 800ms) -# for the auto-arm to claim this home (state/.claude-autoarm.lock owner -# alive, with a supervision decision still open rather than a claim its own -# ledger entry or recorded pid-identity already settles as finished) or to +# for the auto-arm to claim this home (a live OPEN generation claim in the +# state/.claude-autoarm-epoch ledger - fm_autoarm_claim_open - or a legacy +# build's lock-holding claim under the legacy abandonment proof) or to # record a fresh actionable exit-2 outcome # (state/.claude-autoarm-epoch) for this event epoch - either proof allows # without consuming a continuation, so one event epoch yields exactly one recovery turn; @@ -210,8 +210,8 @@ fi budget_account_current_epoch() { local current_epoch outcome old_session old_count old_epoch tmp initialized fm_lock_try_acquire "$BUDGET_LOCK" || return 1 - current_epoch=$(sed -n 's/^epoch=\([0-9][0-9]*\) .*/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) - outcome=$(sed -n 's/^.*outcome=\([a-z][a-z-]*\) .*$/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) + current_epoch=$(sed -n '1s/^epoch=\([0-9][0-9]*\) .*/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) + outcome=$(sed -n '1s/^.*outcome=\([a-z][a-z-]*\) .*$/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) initialized=0 COUNT=0 if [ -f "$BUDGET_FILE" ]; then @@ -259,21 +259,29 @@ budget_account_current_epoch() { autoarm_owns_recovery() { local pid role outcome age fm_watcher_healthy "$STATE" "$WATCH" "$GRACE" "$FM_HOME" && return 0 + # A live OPEN generation claim owns recovery: the ledger names a live, + # identity-matched owner still arming that is not stuck (fm_autoarm_claim_open + # in bin/fm-wake-lib.sh owns that predicate). A finished, dead, + # identity-mismatched, or stuck claim deliberately fails it and falls + # through, because treating such a claim as ownership is what let a dead + # watcher go unnoticed for turn after turn; the outcome cases below still + # cover a claim that finished moments ago, so a genuine handoff is not + # duplicated, while a stale one now reaches the block. + if fm_autoarm_claim_open "$STATE" "$GRACE"; then + [ ! -e "$FAILURE_NOTICE" ] || budget_account_current_epoch || true + return 0 + fi + # Legacy shim: a pre-generation build's claim holds the owner lock with the + # autoarm role for its whole cycle; defer to it under the legacy abandonment + # proof so an upgrade mid-session cannot double-arm. pid=$(cat "$OWNER_LOCK/pid" 2>/dev/null || true) role=$(fm_lock_role "$OWNER_LOCK" 2>/dev/null || true) - # A live auto-arm owner is only evidence of ownership while its supervision - # decision is still open. Once its own ledger entry records a terminal outcome, - # or its recorded pid-identity stops matching the pid holding the lock, the lock - # is abandoned, and treating it as ownership is what let a dead watcher go - # unnoticed for turn after turn. Fall through instead: the outcome cases below - # still cover a claim that finished moments ago, so a genuine handoff is not - # duplicated, while a stale one now reaches the block. if fm_pid_alive "$pid" && [ "$role" = autoarm ] \ - && ! fm_autoarm_claim_abandoned "$STATE"; then + && ! fm_autoarm_claim_abandoned "$STATE" "$GRACE"; then [ ! -e "$FAILURE_NOTICE" ] || budget_account_current_epoch || true return 0 fi - outcome=$(sed -n 's/^.*outcome=\([a-z][a-z-]*\) .*$/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) + outcome=$(sed -n '1s/^.*outcome=\([a-z][a-z-]*\) .*$/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) case "$outcome" in rewake) age=$(fm_path_age "$STATE/.claude-autoarm-epoch") @@ -305,20 +313,24 @@ terminal_fail_open() { [ "$COUNT" -gt "$BLOCK_BUDGET" ] || return 1 failure_episode_verified || return 1 [ ! -e "$FAILURE_ALARM" ] || return 1 + # A live open generation claim is a concurrent recovery decision to step + # aside for, exactly like the legacy live-owner case below. + fm_autoarm_claim_open "$STATE" "$GRACE" && return 2 if ! fm_lock_try_acquire "$OWNER_LOCK"; then pid=$(cat "$OWNER_LOCK/pid" 2>/dev/null || true) role=$(fm_lock_role "$OWNER_LOCK" 2>/dev/null || true) - # Same abandonment test as autoarm_owns_recovery: a claim whose ledger entry - # is already terminal, or whose recorded pid-identity no longer matches the - # live pid, is not a concurrent owner to step aside for. Stepping aside for one - # here allows the stop silently, and the episode's one attended alarm would - # never fire, so clear the abandoned claim and let this decision finish - # instead. Failing to clear it re-blocks rather than allowing. + # Same legacy abandonment test as autoarm_owns_recovery: a claim whose + # ledger entry is already terminal, or whose recorded pid-identity no + # longer matches the live pid, is not a concurrent owner to step aside + # for. Stepping aside for one here allows the stop silently, and the + # episode's one attended alarm would never fire, so clear the abandoned + # claim and let this decision finish instead. Failing to clear it + # re-blocks rather than allowing. if fm_pid_alive "$pid" && [ "$role" = autoarm ] \ - && ! fm_autoarm_claim_abandoned "$STATE"; then + && ! fm_autoarm_claim_abandoned "$STATE" "$GRACE"; then return 2 fi - fm_autoarm_release_abandoned "$STATE" || return 1 + fm_autoarm_release_abandoned "$STATE" "$GRACE" || return 1 fm_lock_try_acquire "$OWNER_LOCK" || return 1 fi if ! fm_lock_set_role "$OWNER_LOCK" terminal-check; then @@ -352,6 +364,15 @@ terminal_fail_open() { fm_lock_release "$OWNER_LOCK" return 2 fi + # Re-check for a live open generation claim now that both locks are held: a + # claimant that published "arming" between the pre-check above and the lock + # acquisition is active recovery, and alarming over it would fire the + # episode's one attended fail-open while a continuation is under way. + if fm_autoarm_claim_open "$STATE" "$GRACE"; then + fm_lock_release "$BUDGET_LOCK" + fm_lock_release "$OWNER_LOCK" + return 2 + fi if ! (set -C; : > "$FAILURE_ALARM") 2>/dev/null; then fm_lock_release "$BUDGET_LOCK" fm_lock_release "$OWNER_LOCK" @@ -366,7 +387,7 @@ failure_episode_verified() { local outcome [ ! -e "$STATE/.afk" ] || return 1 [ -e "$FAILURE_NOTICE" ] || return 1 - outcome=$(sed -n 's/^.*outcome=\([a-z][a-z-]*\) .*$/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) + outcome=$(sed -n '1s/^.*outcome=\([a-z][a-z-]*\) .*$/\1/p' "$STATE/.claude-autoarm-epoch" 2>/dev/null || true) case "$outcome" in failed|failed-suppressed) return 0 ;; *) return 1 ;; diff --git a/bin/fm-wake-lib.sh b/bin/fm-wake-lib.sh index 7588f52cbe9..686fdd5e5e9 100755 --- a/bin/fm-wake-lib.sh +++ b/bin/fm-wake-lib.sh @@ -982,48 +982,79 @@ fm_failure_episode_reset() { return 0 } -# --- Claude Stop auto-arm claim abandonment ---------------------------------- +# --- Claude Stop auto-arm generation claims ----------------------------------- # Both Stop-event participants (bin/fm-claude-stop-autoarm.sh and -# bin/fm-turnend-guard.sh --claude) stand down for whoever holds the auto-arm's -# single-flight owner lock, on the premise that a live holder is still deciding -# supervision. A holder that has already FINISHED that decision but never -# released the lock turns the courtesy into indefinite silence: every later -# async firing exits at the lock, the epoch ledger freezes at its last outcome, -# and each following turn end allows a blind stop while nothing re-arms the -# watcher. Observed 2026-08-14: one delivered rewake, then a beacon that went -# 40 minutes without a beat, no watcher lock at all, two workers in flight, and -# both of their reports unread until an operator drained the queue by hand. +# bin/fm-turnend-guard.sh --claude) coordinate through the epoch ledger +# state/.claude-autoarm-epoch, whose monotonic epoch sequence IS the claim +# generation. This is an optimistic, generation-based single-flight design: # -# One abandonment proof is the ledger, not pid liveness, because both ways a -# finished claim keeps a live pid - reuse of the recorded pid, and a hook still -# blocked writing its rewake banner - look alive: +# - The CURRENT claim is the ledger's latest entry: line 1 is the classic +# "epoch=N owner_pid=P outcome=O updated_at=T" record, and line 2 is the +# claiming process's pid-identity, the same identity every other +# supervision lock in this repo records (fm_pid_identity above). The +# identity is MANDATORY: a claimant that cannot record it does not claim +# (continuity falls to the synchronous guard), and the identity is read +# from the ledger entry alone - never substituted from any lock - so a +# reused pid can never authenticate someone else's stale entry. +# - A claim is OPEN (fm_autoarm_claim_open) while its outcome is "arming", +# its owner pid is alive, its recorded identity successfully recomputes +# and matches that pid, and it is not STUCK - stuck meaning both the +# ledger entry and the watcher beacon (state/.last-watcher-beat) are older +# than the guard grace, which proves the owner hung mid-arm with nothing +# supervising (every legitimate arming phase with no watcher is bounded in +# seconds, while a healthy hours-long cycle keeps the beacon beating). +# - Every firing DEFERS (exits 0) to an open claim; anything else - a +# terminal outcome, a dead or identity-mismatched owner, a stuck owner, an +# identityless entry, or no claim at all - lets the next firing take +# generation N+1 (fm_autoarm_claim_next). Taking a newer generation IS the +# reclaim: a steady-state predecessor is never signalled or revoked. +# - NO mutex is ever held across a blocking step. The owner lock +# state/.claude-autoarm.lock survives only as a micro-mutex serializing +# individual ledger reads-then-writes (a few non-blocking file +# operations); a holder that dies inside the hold is reclaimed by +# fm_lock_try_acquire's ordinary dead-owner steal. +# - A superseded owner goes COMPLETELY silent - cleanup only. Ownership is +# re-verified before every side effect: each arm invocation, each +# episode-state mutation, each ledger write, and each continuation. +# - The irrevocable commit point of a translation is the EXIT STATUS: the +# harness delivers the collected stderr banner only on exit 2 and discards +# it on exit 0. Markerless outcomes commit with the owned terminal ledger +# write. The once-per-episode failure notice commits only when its marker is +# created after the winning "failed" write in the same owned critical +# section. A superseded generation or failed required-marker creation is +# refused and exits 0 silently even after printing; a later generation +# supersedes the terminal entry and retries the notice. # -# 1. the owner lock exists and carries the auto-arm role, -# 2. its recorded pid is numeric, -# 3. the ledger's owner_pid is exactly that pid, and -# 4. the ledger's outcome is present and is not "arming". +# This structurally removes the failure classes the lock-held-across-arm +# design produced: a hung owner deferring every later firing forever (observed +# 2026-08-26: a hook hung mid-arm with its ledger frozen at "arming" kept the +# watcher from ever being auto-re-armed again; and 2026-08-14: a finished +# claim whose leftover lock silenced both participants for 40 beacon-less +# minutes), a reclaim mutex held across a blocking banner write recreating the +# same unreclaimable-live-owner shape, and a reclaimed-but-alive owner racing +# its replacement to translate one close twice. # -# Condition 3 is what makes reclaiming race-free. A fresh claimant creates the -# lock BEFORE it writes "arming", so until it does the ledger still names the -# PREVIOUS owner and the two pids cannot match; a just-started claim is never -# mistaken for an abandoned one. Condition 4 treats "arming" as in progress no -# matter how old, because the owner foregrounds fm-watch-arm.sh for the whole -# watcher cycle, which legitimately runs for hours. +# Two bounded residuals are ACCEPTED INTENT, because closing them absolutely +# would require a mutex held across output or steady-state revocation, both +# deliberately rejected: (1) an owner that dies between its owned terminal +# write and its own process exit leaves a committed outcome whose banner was +# never delivered (process-death territory; the durable wake queue retains the +# underlying event), and (2) a hung old-build owner that resumes during the +# one legacy upgrade window may add one duplicate continuation. Each residual +# costs at most one extra exit-2 continuation turn absorbed by the durable +# idempotent wake queue. A claim misread as stuck in a pathological race +# (e.g. a beacon read right at system wake) likewise yields at most one extra +# arm that the watcher singleton dedupes, while the superseded owner still +# goes silent. # -# The ledger alone cannot prove every abandonment, though: an entry still reading -# "arming", or no entry at all, says nothing about a recorded pid the operating -# system has since handed to an unrelated live process - the same lapse, reached -# when a session teardown kills a claim's whole process group before it can record -# any outcome or run its release trap. So the claim also records the pid-identity -# every other supervision lock in this repo records (fm_pid_identity above, used by -# state/.watch.lock, the supervise-daemon lock, and the AFK launch lock), and a -# recorded identity that no longer matches the live pid is abandonment on its own, -# whatever the ledger says. That identity is written BEFORE the auto-arm role is -# published, and every participant requires that role first, so a claim that is -# genuinely mid-flight is never read as identity-less. A claim carrying no recorded -# identity at all (an older build, a hand-edited lock) keeps exactly the -# ledger-only reasoning above, and an identity that cannot be recomputed for the -# live pid proves nothing either way, so it falls through to the ledger too. +# fm_autoarm_claim_abandoned / fm_autoarm_release_abandoned below survive as +# the LEGACY shim for a lock-holding claim from a pre-generation build (the +# lock carries a role file only in that legacy shape, and in the guard's own +# short terminal-check hold): a live legacy owner still defers per the legacy +# proof, and a proven-abandoned one is reclaimed once through the steal mutex +# - with an identity-verified live owner retired via TERM first, because old +# code cannot re-check generations - so an upgrade mid-session can neither +# double-arm nor deadlock behind a hung legacy hook. _fm_autoarm_epoch_field() { # local file=$1 field=$2 tok local -a toks=() @@ -1039,42 +1070,188 @@ _fm_autoarm_epoch_field() { # return 1 } -# Record the claiming process's pid-identity inside the auto-arm owner lock, the -# way every other supervision lock in this repo records it. Best effort by design: -# a platform where fm_pid_identity cannot answer keeps the ledger-only reasoning -# rather than losing the claim, and a record that cannot be completed leaves NO -# identity file behind, so a partial write can never read as a mismatch against -# its own live owner. Call it before publishing the auto-arm role. -fm_autoarm_claim_record_identity() { # - local state=$1 lock pid held identity back +# Parse the current ledger claim. Sets FM_AUTOARM_GEN, FM_AUTOARM_OWNER, +# FM_AUTOARM_OUTCOME, and FM_AUTOARM_IDENTITY (line 2 of the entry, and ONLY +# line 2 - identity is never substituted from a lock, so a transient +# micro-mutex hold or a reused pid can never authenticate a stale entry). +fm_autoarm_ledger_read() { # + local state=$1 epoch + epoch="$state/.claude-autoarm-epoch" + FM_AUTOARM_GEN= + FM_AUTOARM_OWNER= + FM_AUTOARM_OUTCOME= + FM_AUTOARM_IDENTITY= + FM_AUTOARM_GEN=$(_fm_autoarm_epoch_field "$epoch" epoch) || return 1 + FM_AUTOARM_OWNER=$(_fm_autoarm_epoch_field "$epoch" owner_pid) || return 1 + FM_AUTOARM_OUTCOME=$(_fm_autoarm_epoch_field "$epoch" outcome) || return 1 + case "$FM_AUTOARM_GEN" in + ''|*[!0-9]*) return 1 ;; + esac + FM_AUTOARM_IDENTITY=$(sed -n '2p' "$epoch" 2>/dev/null || true) + return 0 +} + +# True while the CURRENT ledger claim is open and healthy - the defer predicate +# both Stop participants use. Open means: outcome "arming", a live owner whose +# mandatory recorded identity recomputes and matches its pid, and not stuck +# (the contract comment above owns the stuck proof). fm_path_age reports an +# absent beacon as ancient, which is exactly right: arming for a full grace +# window without producing a first beat is the same hang. An identityless +# entry is never open: real generation claims always record identity, a legacy +# build's entry gets its deference from its held role-carrying lock through +# the legacy shim, and anything else must not defer. +fm_autoarm_claim_open() { # [grace] + local state=$1 grace=${2:-${FM_GUARD_GRACE:-300}} epoch current + epoch="$state/.claude-autoarm-epoch" + case "$grace" in + ''|*[!0-9]*|0) grace=300 ;; + esac + fm_autoarm_ledger_read "$state" || return 1 + [ "$FM_AUTOARM_OUTCOME" = arming ] || return 1 + fm_pid_alive "$FM_AUTOARM_OWNER" || return 1 + [ -n "$FM_AUTOARM_IDENTITY" ] || return 1 + current=$(fm_pid_identity "$FM_AUTOARM_OWNER" 2>/dev/null) || return 1 + [ -n "$current" ] || return 1 + [ "$current" = "$FM_AUTOARM_IDENTITY" ] || return 1 + if [ "$(fm_path_age "$epoch")" -ge "$grace" ] \ + && [ "$(fm_path_age "$state/.last-watcher-beat")" -ge "$grace" ]; then + return 1 + fi + return 0 +} + +# Atomically publish this process as the owner of generation N+1, under one +# short micro-mutex hold. Returns 0 with FM_AUTOARM_MY_GEN set on success, 2 +# when a competing claimant won the race (the ledger holds an open claim), and +# 1 when the micro-mutex is contended, the mandatory identity cannot be +# computed, or the write failed. +fm_autoarm_claim_next() { # [grace] + local state=$1 grace=${2:-${FM_GUARD_GRACE:-300}} lock epoch pid gen identity tmp lock="$state/.claude-autoarm.lock" - # Resolve the pid into a variable FIRST: expanding ${BASHPID:-$$} inside the - # command substitution below would resolve it in that subshell, recording the - # identity of a process that exits immediately and leaving every later reader - # with a permanent mismatch against the real owner. + epoch="$state/.claude-autoarm-epoch" + FM_AUTOARM_MY_GEN= + # Resolve the pid into a variable FIRST: expanding ${BASHPID:-$$} inside a + # command substitution would resolve it in that subshell, recording the + # identity of a process that exits immediately. pid=${BASHPID:-$$} - # The identity must describe the pid the lock publishes, so record it only for a - # lock this process actually holds (the same ownership test as fm_lock_set_role). - held=$(cat "$lock/pid" 2>/dev/null || true) - [ "$held" = "$pid" ] || return 1 identity=$(fm_pid_identity "$pid" 2>/dev/null) || return 1 [ -n "$identity" ] || return 1 - if ! printf '%s\n' "$identity" > "$lock/pid-identity" 2>/dev/null; then - rm -f "$lock/pid-identity" 2>/dev/null || true + fm_lock_try_acquire "$lock" || return 1 + if fm_autoarm_claim_open "$state" "$grace"; then + fm_lock_release "$lock" + return 2 + fi + gen=$(_fm_autoarm_epoch_field "$epoch" epoch 2>/dev/null || true) + case "$gen" in + ''|*[!0-9]*) gen=0 ;; + esac + gen=$((gen + 1)) + tmp="$epoch.tmp.$pid" + if ! printf 'epoch=%s owner_pid=%s outcome=arming updated_at=%s\n%s\n' \ + "$gen" "$pid" "$(date +%s)" "$identity" > "$tmp" 2>/dev/null \ + || ! mv -f "$tmp" "$epoch" 2>/dev/null; then + rm -f "$tmp" 2>/dev/null || true + fm_lock_release "$lock" + return 1 + fi + fm_lock_release "$lock" + # shellcheck disable=SC2034 # Read by callers after the claim succeeds. + FM_AUTOARM_MY_GEN=$gen + return 0 +} + +# Write a new outcome for a generation this process still owns, re-verified +# under the micro-mutex so a superseded owner can never clobber a newer claim. +# With a fourth argument, create that marker after the ledger rename in the same +# owned critical section (the once-per-episode failure notice). A marker failure +# refuses the commit even though its terminal ledger entry remains; marker-first +# ordering could permanently suppress a notice whose ledger write never won. +# Returns 0 committed, 2 refused (superseded or required-marker failure), and 1 +# unable (bounded contention or ledger-write failure). +fm_autoarm_write_owned() { # [marker-file] + local state=$1 gen=$2 outcome=$3 marker=${4:-} lock epoch pid identity tmp i + lock="$state/.claude-autoarm.lock" + epoch="$state/.claude-autoarm-epoch" + pid=${BASHPID:-$$} + i=0 + while ! fm_lock_try_acquire "$lock"; do + [ "$i" -lt 20 ] || return 1 + sleep 0.02 + i=$((i + 1)) + done + if ! fm_autoarm_ledger_read "$state" \ + || [ "$FM_AUTOARM_GEN" != "$gen" ] || [ "$FM_AUTOARM_OWNER" != "$pid" ]; then + fm_lock_release "$lock" + return 2 + fi + identity=$FM_AUTOARM_IDENTITY + tmp="$epoch.tmp.$pid" + if ! { + printf 'epoch=%s owner_pid=%s outcome=%s updated_at=%s\n' \ + "$gen" "$pid" "$outcome" "$(date +%s)" + [ -z "$identity" ] || printf '%s\n' "$identity" + } > "$tmp" 2>/dev/null || ! mv -f "$tmp" "$epoch" 2>/dev/null; then + rm -f "$tmp" 2>/dev/null || true + fm_lock_release "$lock" return 1 fi - back=$(cat "$lock/pid-identity" 2>/dev/null || true) - if [ "$back" != "$identity" ]; then - rm -f "$lock/pid-identity" 2>/dev/null || true + if [ -n "$marker" ] && ! : > "$marker" 2>/dev/null; then + fm_lock_release "$lock" + return 2 + fi + fm_lock_release "$lock" + return 0 +} + +# Lockless pre-side-effect ownership check: true while the ledger still names +# owned by this process. A superseded owner must go silent instead of +# arming, mutating shared state, or emitting. +fm_autoarm_still_owner() { # + local state=$1 gen=$2 pid + pid=${BASHPID:-$$} + fm_autoarm_ledger_read "$state" || return 1 + [ "$FM_AUTOARM_GEN" = "$gen" ] && [ "$FM_AUTOARM_OWNER" = "$pid" ] +} + +fm_autoarm_reset_owned() { # + local state=$1 gen=$2 lock pid + lock="$state/.claude-autoarm.lock" + pid=${BASHPID:-$$} + fm_lock_try_acquire "$lock" || return 2 + if ! fm_autoarm_ledger_read "$state" \ + || [ "$FM_AUTOARM_GEN" != "$gen" ] || [ "$FM_AUTOARM_OWNER" != "$pid" ]; then + fm_lock_release "$lock" + return 2 + fi + if ! fm_failure_episode_reset "$state"; then + fm_lock_release "$lock" return 1 fi + fm_lock_release "$lock" return 0 } -fm_autoarm_claim_abandoned() { # - local state=$1 epoch lock role pid owner outcome recorded current +# LEGACY shim (see the contract comment above): the abandonment proof for a +# lock-holding claim from a pre-generation build, recognizable by the role +# file only such claims and the guard's short terminal-check hold publish. +# A live legacy owner defers per this proof; a finished, identity-mismatched, +# or stuck one is abandoned: +# +# 1. the owner lock exists and carries the auto-arm role, +# 2. its recorded pid is numeric, +# 3. a recorded pid-identity that no longer matches the live pid is +# abandonment on its own (pid reuse after a group kill), and otherwise +# 4. the ledger's owner_pid is exactly that pid and its outcome is present +# and either is not "arming", or is "arming" while both the ledger entry +# and the watcher beacon are older than the guard grace (the same stuck +# proof as fm_autoarm_claim_open). +fm_autoarm_claim_abandoned() { # [grace] + local state=$1 grace=${2:-${FM_GUARD_GRACE:-300}} epoch lock role pid owner outcome recorded current lock="$state/.claude-autoarm.lock" epoch="$state/.claude-autoarm-epoch" + case "$grace" in + ''|*[!0-9]*|0) grace=300 ;; + esac [ -e "$lock" ] || [ -L "$lock" ] || return 1 role=$(fm_lock_role "$lock") [ "$role" = autoarm ] || return 1 @@ -1091,26 +1268,84 @@ fm_autoarm_claim_abandoned() { # [ "$owner" = "$pid" ] || return 1 outcome=$(_fm_autoarm_epoch_field "$epoch" outcome) || return 1 case "$outcome" in - ''|arming) return 1 ;; + '') return 1 ;; + arming) + [ "$(fm_path_age "$epoch")" -ge "$grace" ] || return 1 + [ "$(fm_path_age "$state/.last-watcher-beat")" -ge "$grace" ] || return 1 + return 0 + ;; esac return 0 } -# Remove a proven-abandoned auto-arm claim so the next claimant can arm. -# The proof is re-verified while holding the lock's steal mutex, which is the -# same serialization fm_lock_try_acquire uses for stale-owner reclaim: while it -# is held no other process can publish the primary lock, so the window between -# proving abandonment and removing the lock cannot swallow a genuine new claim. -fm_autoarm_release_abandoned() { # - local state=$1 lock steal +# Remove a proven-abandoned legacy claim so the next claimant can arm. The +# proof is re-verified while holding the lock's steal mutex, the same +# serialization fm_lock_try_acquire uses for stale-owner reclaim: while it is +# held no other process can publish the primary lock, so the window between +# proving abandonment and removing the lock cannot swallow a genuine new +# claim. +# +# Old-build code cannot re-check generations, so a LIVE proven-abandoned +# legacy owner whose recorded identity is verified to match its pid is retired +# with TERM before the lock is removed: once the TERM is successfully queued +# the process can never resume normal execution (delivery precedes any further +# user code when it continues), so a short bounded wait for observed exit is a +# courtesy, not a requirement. A pid is never signalled without a verified +# matching identity; when the kill itself fails or the identity stops matching +# mid-procedure (pid reuse), the reclaim refuses. Missing identity evidence +# never blocks the reclaim of a proven-abandoned claim - it only disables the +# TERM and the ledger graft below, keeping the documented bounded +# upgrade-window residual instead of the deadlock. +fm_autoarm_release_abandoned() { # [grace] + local state=$1 grace=${2:-${FM_GUARD_GRACE:-300}} lock steal epoch lock_pid recorded current owner line1 tmp i lock="$state/.claude-autoarm.lock" steal="$lock.steal" - fm_autoarm_claim_abandoned "$state" || return 1 + epoch="$state/.claude-autoarm-epoch" + fm_autoarm_claim_abandoned "$state" "$grace" || return 1 fm_lock_try_acquire "$steal" || return 1 - if ! fm_autoarm_claim_abandoned "$state"; then + if ! fm_autoarm_claim_abandoned "$state" "$grace"; then fm_lock_release "$steal" return 1 fi + lock_pid=$(cat "$lock/pid" 2>/dev/null || true) + recorded=$(cat "$lock/pid-identity" 2>/dev/null || true) + if [ -n "$recorded" ] && fm_pid_alive "$lock_pid" \ + && current=$(fm_pid_identity "$lock_pid" 2>/dev/null) \ + && [ -n "$current" ] && [ "$current" = "$recorded" ]; then + # A live pid still answering to the recorded identity IS the genuine + # legacy owner (proven stuck or blocked after a terminal write): retire it + # before removing its lock, because old-build code cannot re-check + # generations. A pid the recorded identity does NOT verify - reused, + # unverifiable, or never recorded - is NEVER signalled; those shapes are + # reclaimed as-is, which is safe exactly because the recorded owner is + # gone or was never provably this process. + if ! kill -TERM "$lock_pid" 2>/dev/null; then + fm_lock_release "$steal" + return 1 + fi + i=0 + while [ "$i" -lt 20 ] && fm_pid_alive "$lock_pid"; do + sleep 0.05 + i=$((i + 1)) + done + fi + # Preserve the legacy lock's identity evidence in the ledger before the lock + # disappears, keeping the ledger's original mtime so the stuck proof's age + # window is not silently reopened. Best effort. + if [ -n "$recorded" ] && [ -n "$lock_pid" ] \ + && owner=$(_fm_autoarm_epoch_field "$epoch" owner_pid 2>/dev/null) \ + && [ "$owner" = "$lock_pid" ] \ + && [ -z "$(sed -n '2p' "$epoch" 2>/dev/null)" ]; then + line1=$(sed -n '1p' "$epoch" 2>/dev/null || true) + tmp="$epoch.tmp.${BASHPID:-$$}" + if [ -n "$line1" ] \ + && printf '%s\n%s\n' "$line1" "$recorded" > "$tmp" 2>/dev/null \ + && touch -r "$epoch" "$tmp" 2>/dev/null \ + && mv -f "$tmp" "$epoch" 2>/dev/null; then + : + fi + rm -f "$tmp" 2>/dev/null || true + fi fm_lock_remove_path "$lock" || true fm_lock_release "$steal" [ -e "$lock" ] || [ -L "$lock" ] || return 0 diff --git a/docs/configuration.md b/docs/configuration.md index fd66239b436..9df7bb77373 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -734,7 +734,7 @@ FM_PF_RETRY_BACKOFF_SECS=900 # seconds before the next attempt after a retryab FM_LOCK_STALE_AFTER=2 # seconds before dead-pid lock records can be reclaimed; mid-acquire locks keep at least 2s grace FM_GUARD_GRACE=300 # seconds before guard warnings, arm health checks, and the primary turn-end guard treat a watcher beacon as stale FM_CLAUDE_AUTOARM_ATTEMPTS=2 # bounded Stop-owned arm attempts per Claude auto-arm cycle; accepted values are 1, 2, or 3 -FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=800 # milliseconds the --claude turn-end guard waits for watcher health, a role-verified Stop auto-arm claim, or a fresh epoch before deciding recovery ownership or failure progression +FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=800 # milliseconds the --claude turn-end guard waits for watcher health, an open Stop auto-arm generation claim, or a fresh epoch before deciding recovery ownership or failure progression FM_CLAUDE_AUTOARM_EPOCH_FRESH=15 # seconds a recorded auto-arm outcome remains eligible for the current event epoch's recovery or failure decision FM_CLAUDE_TURNEND_BLOCK_BUDGET=3 # consecutive --claude guard re-blocks before the verified one-time attended fail-open; safely below Claude Code's 8-block override FM_ARM_CONFIRM_TIMEOUT=10 # seconds fm-watch-arm waits to confirm a fresh watcher before reporting FAILED; default 30 on Git Bash/MSYS diff --git a/docs/supervision-protocols/claude.md b/docs/supervision-protocols/claude.md index 1e5033a55ed..f0d631f6493 100644 --- a/docs/supervision-protocols/claude.md +++ b/docs/supervision-protocols/claude.md @@ -19,7 +19,7 @@ When this session owns supervision and away mode is not active: [`watcher-continuity.md`](../watcher-continuity.md) owns the exact session-lock recovery boundary. 8. The turn-end guard (`bin/fm-turnend-guard.sh --claude`) remains the final backstop. It requires the PID-strict live-watcher and fresh-beacon predicate at the Stop boundary, while the mid-turn pull guard accepts a fresh beacon without a live process under Claude's between-turns auto-arm model. - It allows the stop when a watcher is healthy or the role-verified auto-arm owns recovery, while fresh failure epochs advance the bounded one-time attended fail-open progression described in [`turnend-guard.md`](../turnend-guard.md). + It allows the stop when a watcher is healthy or an open auto-arm generation claim owns recovery, while fresh failure epochs advance the bounded one-time attended fail-open progression described in [`turnend-guard.md`](../turnend-guard.md). 9. Waiting on the hook-owned cycle is silent: do not send idle progress while the watcher is parked. The watcher itself remains `bin/fm-watch.sh`, and `bin/fm-watch-arm.sh` remains the verified arm wrapper that the Stop hook foregrounds. diff --git a/docs/turnend-guard.md b/docs/turnend-guard.md index c9852326f7a..134c2f5dc41 100644 --- a/docs/turnend-guard.md +++ b/docs/turnend-guard.md @@ -72,14 +72,16 @@ In the default Codex mode, a true value lets the second stop finish after one fo Claude runs the guard with `--claude`, which ignores `stop_hook_active` and cooperates with the Stop-owned auto-arm. Claude Code sets `stop_hook_active=true` on every stop after any stop-hook continuation, including `asyncRewake` rewakes, which re-opened the 2026-07-21 blind window under the default one-shot behavior. -The Claude mode waits up to `FM_CLAUDE_AUTOARM_SYNC_WAIT_MS` (default 800 milliseconds) and allows the stop when the watcher is healthy, `state/.claude-autoarm.lock` has a live `autoarm` role owner whose supervision decision is still open and whose eventual failure must exit 2, or `state/.claude-autoarm-epoch` contains a fresh actionable rewake owned by this event epoch. -A live owner counts as that proof only while its decision is open, which the ledger settles: an entry naming that owner's own pid with any outcome other than `arming` means the claim already finished, so the lock is abandoned rather than in flight. -The guard then stops reading it as recovery under way, the terminal check clears it instead of stepping aside for it, and the next Stop-owned firing reclaims it and arms rather than deferring. -Without that boundary a cycle that armed, delivered one rewake, and exited left both Stop participants deferring to its leftover lock indefinitely, so on 2026-08-14 a home with two tasks in flight and a beacon 40 minutes cold ended every turn blind until an operator intervened. -An `arming` entry stays in flight however old it is, because the owner foregrounds the arm for the whole watcher cycle. -The shapes the ledger cannot settle are settled by identity instead: the claim records the same `pid-identity` file every other supervision lock records, before it publishes its `autoarm` role, so a recorded identity that no longer matches the pid holding the lock proves abandonment on its own even while the entry still reads `arming` or no ledger entry exists at all. -That covers a claim whose process group was killed before it could record any outcome and whose pid the operating system later handed to an unrelated live process. -A claim carrying no recorded identity keeps the ledger-only boundary, and a failed reclaim re-blocks rather than allowing a blind stop. +The Claude mode waits up to `FM_CLAUDE_AUTOARM_SYNC_WAIT_MS` (default 800 milliseconds) and allows the stop when the watcher is healthy, the auto-arm's generation claim is open, or `state/.claude-autoarm-epoch` contains a fresh actionable rewake owned by this event epoch. +The claim is the ledger entry itself: the epoch sequence in `state/.claude-autoarm-epoch` is a monotonic claim generation, line 1 is the classic epoch record, and line 2 records the claiming process's mandatory pid-identity (`fm_autoarm_claim_open` and `fm_autoarm_claim_next` in `bin/fm-wake-lib.sh` own the contract). +A claim is open while its outcome is `arming`, its owner pid is alive, its recorded identity successfully recomputes and matches that pid, and it is not stuck - stuck meaning the entry and the watcher beacon are both older than the guard grace, which proves the owner hung mid-arm (a healthy hours-long foregrounded cycle keeps the beacon beating, and every arming phase with no watcher is bounded in seconds). +Anything else - a finished outcome, a dead or identity-mismatched owner, a stuck owner, an identityless entry, or no entry - lets the next Stop-owned firing take the next generation and arm; taking a newer generation is the reclaim, and a steady-state predecessor is never signalled or revoked. +No mutex is held across arming or output: `state/.claude-autoarm.lock` survives only as a micro-mutex serializing individual ledger writes, and a superseded owner goes completely silent - ownership is re-verified before every arm invocation, episode-state mutation, ledger write, and continuation. +The irrevocable commit point of a translation is the exit status, because the harness delivers the collected stderr banner only on exit 2, so an owned terminal commit decides the exit: markerless outcomes commit with the ledger write, while the once-per-episode failure notice commits only when its marker is created after the winning failed write in the same critical section. +A generation whose required marker cannot be created is refused and exits 0 silently even after printing; its terminal ledger entry is superseded by a later firing, which retries the notice. +Without those boundaries a cycle that armed, delivered one rewake, and exited left both Stop participants deferring to its leftover lock indefinitely (2026-08-14: two tasks in flight, a beacon 40 minutes cold, every turn blind until an operator intervened), and a hook that hung mid-arm kept a live pid on the lock so the watcher was never auto-re-armed again (2026-08-26). +Two bounded residuals are accepted intent, each costing at most one extra continuation turn absorbed by the durable idempotent wake queue: an owner that dies between its owned terminal write and its own process exit, and a hung old-build owner that resumes during the one legacy upgrade window. +A legacy build's lock-holding claim (recognizable by its `autoarm` role file) still defers or reclaims under the legacy abandonment proof, with a live identity-verified stuck owner retired via TERM before its lock is removed and an unverified pid never signalled, so an upgrade mid-session can neither double-arm nor deadlock, and a failed reclaim re-blocks rather than allowing a blind stop. Fresh `failed` and `failed-suppressed` outcomes enter or advance the failure progression instead of acting as unconditional recovery proof. The auto-arm itself rechecks the healthy watcher predicate and retries a bounded number of times before reporting a genuine failure. The first fresh exhausted-failure epoch preserves its handoff without consuming a blocked-stop count, while later fresh failed epochs advance the same monotonic progression instead of resetting it. @@ -157,7 +159,7 @@ That warning uses `bin/fm-supervision-instructions.sh --repair-line`, so it alwa ## Regression coverage -`tests/fm-turnend-guard.test.sh` covers the predicate, main and secondmate primary scope, child-worktree exclusion, `FM_HOME` and `FM_STATE_OVERRIDE` precedence, the live-lock and fresh-beacon guard predicate, the cooperative `--claude` claim wait, monotonic failed-epoch progression, bounded attended fail-open, post-alarm continuation suppression, positive recovery reset, the abandoned auto-arm claim cases that must block or clear instead of allowing a blind stop, Pi logical-run latching, missing-`jq` behavior, all five primary registrations, Grok native and legacy selection, typed field precedence, malformed input, and exactly-one-path safety. +`tests/fm-turnend-guard.test.sh` covers the predicate, main and secondmate primary scope, child-worktree exclusion, `FM_HOME` and `FM_STATE_OVERRIDE` precedence, the live-lock and fresh-beacon guard predicate, the cooperative `--claude` open-generation claim wait, monotonic failed-epoch progression, bounded attended fail-open, post-alarm continuation suppression, positive recovery reset, generation and legacy claim cases that must block or clear instead of allowing a blind stop, Pi logical-run latching, missing-`jq` behavior, all five primary registrations, Grok native and legacy selection, typed field precedence, malformed input, and exactly-one-path safety. `tests/fm-guard-stale-banner.test.sh` covers the pull-guard predicate, including the persistent-model fresh-leftover-beacon negative control, the auto-arm model's healthy fresh-beacon-without-a-watcher case and stale-beacon alarm, and the extension model's live-watcher path, ownership-qualified fresh hand-off, held-lock failures, independently broken ownership signals, stale-beacon alarm, queued-wake warning, and Pi and pi-signed harness routing. It also covers true-reason banner wording and reason-keyed episode dedup surviving a beacon mtime change. `tests/fm-cursor-primary.test.sh` covers the Cursor park end to end over real processes with no harness installed: each tracked Claude-shaped entrypoint standing down on a Cursor payload, both follow-up sources, the bounded repair nag and its reset, the nested loop bounds, supersession, away-mode and lock-ownership inertness, Pi-host stand-down without Cursor identity and continued parking when `PI_CODING_AGENT` leaks alongside `CURSOR_AGENT` or `CURSOR_INVOKED_AS`, child-worktree exclusion, and that the adapter never exits 2. diff --git a/docs/watcher-continuity.md b/docs/watcher-continuity.md index 0655085ce28..be43542f2ab 100644 --- a/docs/watcher-continuity.md +++ b/docs/watcher-continuity.md @@ -16,8 +16,8 @@ A numeric session-lock owner that fails the shared `fm_harness_pid_alive` predic The stale-owner claim occurs only after the existing AFK and supervision-need gates pass. After each non-actionable arm close, the hook rechecks the identity-matched watcher lock and fresh beacon before retrying a bounded number of times. A cycle-end failure is benign when that live-watcher predicate is true, and the hook suppresses the arm output and continues silently. -Only an exhausted failure with no verified watcher emits one last-resort notice for the continuous failure episode; later consecutive Stop cycles exit 2 to guarantee another Stop-owned retry without repeating the notice until the turn-end guard consumes the attended fail-open. -The Claude turn-end guard owns the monotonic failure progression, one-time attended fail-open, post-alarm continuation suppression, and positive recovery reset described in [`turnend-guard.md`](turnend-guard.md#harness-integrations). +Only an exhausted failure with no verified watcher commits one last-resort notice for the continuous failure episode; a refused notice commit stays silent for a later retry, and after a successful notice later Stop cycles exit 2 without repeating it until the turn-end guard consumes the attended fail-open. +The Claude turn-end guard owns that notice commit contract, the monotonic failure progression, one-time attended fail-open, post-alarm continuation suppression, and positive recovery reset described in [`turnend-guard.md`](turnend-guard.md#harness-integrations). While supervision is still needed and away mode remains inactive, an actionable close wakes the idle session through exit 2. ## Actionable wake ordering @@ -32,7 +32,7 @@ After the configured retry bound is exhausted, it delivers the original wake wit This is deliberate Option B ordering: the fleet is protected before the model handles the wake whenever restoration succeeds, but the model is never left blind when it does not. Claude's Stop hook starts the successor arm at the next Stop after the handling turn, rather than before notification as Pi and OpenCode do. -The durable wake queue preserves actionable events during the residual active-turn window, and the bounded turn-end guard enforces recovery at Stop when no watcher is live and no auto-arm claim is still deciding, so a leftover claim whose own decision already finished cannot suppress it ([`turnend-guard.md`](turnend-guard.md#harness-integrations) owns that boundary). +The durable wake queue preserves actionable events during the residual active-turn window, and the bounded turn-end guard enforces recovery at Stop when no watcher is live and no open generation claim is still deciding, so a finished, hung, or identity-mismatched claim cannot suppress it ([`turnend-guard.md`](turnend-guard.md#harness-integrations) owns that boundary). The recovery-episode contract below owns once-per-generation announcement. A handling successor does not re-announce; it enters its poll loop immediately and keeps scanning signals, stale panes, and checks. The model no longer re-arms after ordinary wakes. @@ -105,9 +105,9 @@ The same suite covers ordinary same-process session replacement for `/new`, `/re `tests/fm-watcher-lock.test.sh` covers verified-successor attach, recovery publication before stale-lock removal, the typed self-eviction failure, bounded and successor-linked lifecycle rows, and a SIGSTOP counterfactual that distinguishes a live PID from a stale beacon before classifying termination. `tests/fm-subagent-pretool-check.test.sh` proves Claude retains only the non-status Bash seatbelts. `tests/fm-claude-stop-autoarm.test.sh` covers the auto-arm's scope, stale and live session owners, unchanged AFK and need boundaries, single-flight, bounded failure retries, benign live-watcher cycle ends, one-notice failure episodes, and exit-2 translation. -It also covers abandoned single-flight claims: a claim the ledger shows already finished, and one whose recorded pid-identity no longer matches its live pid while the ledger still reads arming or is absent entirely, are both reclaimed so a lapsed home re-arms, while an identity-matched claim still arming, one the ledger does not name, and the guard's own terminal check keep the gate closed ([`turnend-guard.md`](turnend-guard.md) owns that boundary). +It also covers generation-claim single-flight, stuck-claim supersession, superseded-owner silence, notice-marker refusal and retry, ownership-atomic episode reset, and the legacy upgrade shim; [`turnend-guard.md`](turnend-guard.md) owns those behavior contracts. `FM_CLAUDE_LIVE_E2E=1 tests/fm-claude-stop-autoarm-live-e2e.test.sh` starts with the reproduced stale-lock state, runs session start first, completes two tokenless cycles, and checks the competing-live-owner negative control. -`tests/fm-turnend-guard.test.sh` covers the cooperative `--claude` guard, including monotonic failed-epoch progression, the integrated bounded fail-open, post-alarm continuation suppression, and positive recovery reset; [`turnend-guard.md`](turnend-guard.md#regression-coverage) lists that suite's full coverage, including the abandoned-claim cases. +`tests/fm-turnend-guard.test.sh` covers the cooperative `--claude` guard, including monotonic failed-epoch progression, the integrated bounded fail-open, post-alarm continuation suppression, and positive recovery reset; [`turnend-guard.md`](turnend-guard.md#regression-coverage) lists that suite's full generation and legacy claim coverage. ## Active limits and verification diff --git a/tests/fm-claude-stop-autoarm.test.sh b/tests/fm-claude-stop-autoarm.test.sh index ff095c32912..042d04ba947 100755 --- a/tests/fm-claude-stop-autoarm.test.sh +++ b/tests/fm-claude-stop-autoarm.test.sh @@ -114,6 +114,16 @@ SH echo "$$" >> "$FM_HOME/state/arm-ran" printf 'watcher: FAILED - cycle ended without an actionable reason\n' exit 1 +SH + ;; + reset-boundary) + cat > "$dir/bin/fm-watch-arm.sh" <<'SH' +#!/usr/bin/env bash +echo "$$" >> "$FM_HOME/state/arm-ran" +: > "$FM_HOME/state/arm-waiting" +while [ ! -e "$FM_HOME/state/arm-release" ]; do sleep 0.02; done +printf 'watcher: FAILED - cycle ended without an actionable reason\n' +exit 1 SH ;; slow-actionable) @@ -124,6 +134,26 @@ sleep 2 printf 'watcher: started pid=%s (beacon fresh)\n' "$$" printf 'signal: task.status done: slow fixture\n' exit 0 +SH + ;; + blocking-actionable) + cat > "$dir/bin/fm-watch-arm.sh" <<'SH' +#!/usr/bin/env bash +echo "$$" >> "$FM_HOME/state/arm-ran" +sleep 6 +printf 'watcher: started pid=%s (beacon fresh)\n' "$$" +printf 'stale: fixture-win actionable\n' +exit 0 +SH + ;; + supersede-then-fail) + cat > "$dir/bin/fm-watch-arm.sh" <<'SH' +#!/usr/bin/env bash +echo "$$" >> "$FM_HOME/state/arm-ran" +printf 'epoch=999 owner_pid=1 outcome=arming updated_at=%s\nfixture-superseder-identity\n' "$(date +%s)" \ + > "$FM_HOME/state/.claude-autoarm-epoch" +printf 'watcher: FAILED - no live watcher with a fresh beacon\n' +exit 1 SH ;; meta-vanishes) @@ -155,7 +185,21 @@ SH } epoch_outcome() { - sed -n 's/^.*outcome=\([a-z][a-z-]*\) .*$/\1/p' "$1/state/.claude-autoarm-epoch" 2>/dev/null || true + sed -n '1s/^.*outcome=\([a-z][a-z-]*\) .*$/\1/p' "$1/state/.claude-autoarm-epoch" 2>/dev/null || true +} + +# Run the hook in the background under the fake harness, output captured to a +# file. Sets RUN_AUTOARM_BG_PID (a direct child of the calling shell, so the +# caller can `wait` on it for the hook's exit status). +RUN_AUTOARM_BG_PID= +run_autoarm_bg() { + local dir=$1 out=$2 + printf '%s\n' '{"session_id":"sess-autoarm","stop_hook_active":false}' \ + | FM_HOME="$dir" "$FAKE_CLAUDE" -c ' + printf "%s\n" "$$" > "$FM_HOME/state/.lock" + "$FM_HOME/bin/fm-claude-stop-autoarm.sh" + ' > "$out" 2>&1 & + RUN_AUTOARM_BG_PID=$! } watcher_identity() { @@ -409,6 +453,34 @@ test_failed_cycles_notify_once_and_keep_retrying() { pass "auto-arm: consecutive failures keep Stop-owned retry without repeating notice" } +test_failure_notice_marker_write_refuses_delivery_and_retries() { + local dir marker out1 out2 out3 status1 status2 status3 gen1 delivered + dir=$(make_primary_dir "$TMP_ROOT/failed-marker-refusal") + : > "$dir/state/task.meta" + write_arm_fixture "$dir" failed + marker="$dir/state/.claude-autoarm-failure-notified" + ln -s "$dir/state/missing/notice" "$marker" + + out1=$(run_autoarm "$dir" 2>/dev/null); status1=$? + expect_code 0 "$status1" "an unrecordable failure notice must refuse delivery" + [ -L "$marker" ] || fail "the failed marker write unexpectedly replaced its dangling symlink" + [ "$(epoch_outcome "$dir")" = failed ] || fail "the refused generation must leave its terminal ledger outcome" + gen1=$(epoch_field "$dir" epoch) + + rm -f "$marker" + out2=$(run_autoarm "$dir" 2>/dev/null); status2=$? + out3=$(run_autoarm "$dir" 2>/dev/null); status3=$? + expect_code 2 "$status2" "a successor must retry and deliver after the marker path is restored" + expect_code 2 "$status3" "a later failure must retain the Stop-owned retry" + [ "$(epoch_field "$dir" epoch)" -gt "$gen1" ] || fail "the successor did not supersede the refused terminal entry" + assert_present "$marker" "the successful successor did not record the failure notice" + assert_contains "$out2" "automatic supervision mechanism is broken" "the successful successor did not deliver the failure notice" + [ -z "$out3" ] || fail "the firing after the successful marker commit repeated the notice: $out3" + delivered=$(printf '%s\n%s\n' "$out2" "$out3" | grep -c 'automatic supervision mechanism is broken' || true) + [ "$delivered" -eq 1 ] || fail "the restored episode delivered $delivered failure notices instead of one" + pass "auto-arm: marker-write refusal defers delivery until one successor commits the notice" +} + test_unverified_clean_close_exhausts_retries() { local dir out status dir=$(make_primary_dir "$TMP_ROOT/clean") @@ -500,6 +572,46 @@ test_positive_recovery_budget_contention_preserves_episode() { pass "auto-arm: budget contention preserves the episode and forces a reset retry" } +test_owner_mutex_contention_preserves_failure_episode_reset() { + local dir out hook_pid status watcher watcher_id holder i + dir=$(make_primary_dir "$TMP_ROOT/reset-owner-contention") + : > "$dir/state/task.meta" + : > "$dir/state/.turnend-claude-blocks" + : > "$dir/state/.claude-autoarm-failure-notified" + : > "$dir/state/.claude-autoarm-failure-alarmed" + write_arm_fixture "$dir" reset-boundary + sleep 60 & + watcher=$! + watcher_id=$(watcher_identity "$dir" "$watcher") || fail "could not identify reset-contention watcher" + record_watcher_lock "$dir" "$watcher" "$watcher_id" + touch "$dir/state/.last-watcher-beat" + out="$dir/state/hook.out" + run_autoarm_bg "$dir" "$out" + hook_pid=$RUN_AUTOARM_BG_PID + i=0 + while [ ! -e "$dir/state/arm-waiting" ]; do + [ "$i" -lt 50 ] || fail "healthy owner never reached the reset boundary" + sleep 0.05 + i=$((i + 1)) + done + sleep 60 & + holder=$! + mkdir -p "$dir/state/.claude-autoarm.lock" + printf '%s\n' "$holder" > "$dir/state/.claude-autoarm.lock/pid" + : > "$dir/state/arm-release" + wait "$hook_pid"; status=$? + expect_code 0 "$status" "owner-mutex contention at reset must close quietly" + [ ! -s "$out" ] || fail "owner-mutex contention at reset produced output: $(cat "$out")" + assert_present "$dir/state/.turnend-claude-blocks" "contended reset deleted the block budget" + assert_present "$dir/state/.claude-autoarm-failure-notified" "contended reset deleted the failure notice" + assert_present "$dir/state/.claude-autoarm-failure-alarmed" "contended reset deleted the attended alarm" + kill "$holder" "$watcher" 2>/dev/null || true + wait "$holder" 2>/dev/null || true + wait "$watcher" 2>/dev/null || true + rm -rf "$dir/state/.claude-autoarm.lock" + pass "auto-arm: owner-mutex contention preserves successor episode state" +} + test_arms_for_x_mode_poll_need_without_inflight() { local dir out status dir=$(make_primary_dir "$TMP_ROOT/x-need") @@ -534,7 +646,7 @@ test_single_flight_admits_exactly_one_owner() { pass "auto-arm: concurrent firings admit one owner and one rewake translation" } -# --- abandoned single-flight claim recovery ----------------------------------- +# --- abandoned single-flight claim recovery (legacy shim) ---------------------- # The 2026-08-14 lapse: one cycle armed, beat its beacon, delivered a single # rewake, and exited, leaving its owner lock behind with a live pid. The single # flight gate then turned every later firing into exit 0, so with two tasks in @@ -543,6 +655,13 @@ test_single_flight_admits_exactly_one_owner() { # enough to prove that: the ledger naming that same pid with a finished outcome, # or a recorded pid-identity the live pid no longer matches, is what distinguishes # an abandoned claim from one still deciding. +# +# These fixtures fabricate the LOCK-HOLDING claim shape a pre-generation build +# leaves behind, so this section pins the legacy shim: a live legacy owner +# still defers the gate, and an abandoned one is reclaimed once so the home +# re-arms - with an identity-verified live owner retired via TERM first, and +# an identityless one reclaimed without any signalling. The generation-claim +# section below pins the current contract. # Fabricate a held owner lock: . Plain-dir shape on purpose - # the hook must reclaim whatever a crashed or blocked owner left behind. @@ -589,6 +708,7 @@ test_abandoned_owner_claim_is_reclaimed_and_rearms() { record_autoarm_owner "$dir" "$pid" record_autoarm_epoch "$dir" 464 "$pid" rewake out=$(run_autoarm "$dir" 2>/dev/null); status=$? + kill -0 "$pid" 2>/dev/null || fail "an identityless abandoned owner must be reclaimed without being signalled" kill "$pid" 2>/dev/null || true wait "$pid" 2>/dev/null || true expect_code 2 "$status" "a claim whose ledger outcome is already terminal must be reclaimed, not deferred to forever" @@ -602,7 +722,7 @@ test_abandoned_owner_claim_is_reclaimed_and_rearms() { pass "auto-arm: an abandoned owner claim is reclaimed so a lapsed cycle re-arms" } -test_arming_claim_is_never_reclaimed() { +test_arming_claim_with_fresh_beacon_is_never_reclaimed() { local dir out status pid dir=$(make_primary_dir "$TMP_ROOT/arming-claim") : > "$dir/state/task1.meta" @@ -610,18 +730,44 @@ test_arming_claim_is_never_reclaimed() { sleep 60 & pid=$! record_autoarm_owner "$dir" "$pid" - # An owner foregrounds the arm for the whole watcher cycle, so "arming" is in - # progress no matter how old its ledger entry is. + # An owner foregrounds the arm for the whole watcher cycle, so an old "arming" + # entry is still in progress while its watcher keeps beating the beacon. record_autoarm_epoch "$dir" 464 "$pid" arming + : > "$dir/state/.last-watcher-beat" out=$(run_autoarm "$dir" 2>/dev/null); status=$? kill "$pid" 2>/dev/null || true wait "$pid" 2>/dev/null || true - expect_code 0 "$status" "a claim still arming must keep the single-flight gate closed" + expect_code 0 "$status" "a legacy claim still arming under a fresh beacon must keep the single-flight gate closed" [ -z "$out" ] || fail "deferring to an arming claim produced output: $out" assert_absent "$dir/state/arm-ran" "an arming claim was stolen and double-armed" [ "$(epoch_field "$dir" epoch)" = 464 ] || fail "deferred firing rewrote the arming ledger entry" assert_present "$dir/state/.claude-autoarm.lock" "an arming claim lost its owner lock" - pass "auto-arm: an owner still arming is never reclaimed, however long the cycle runs" + pass "auto-arm: a legacy owner still arming is never reclaimed while its watcher keeps beating" +} + +# The other legitimate legacy arming shape: a claim that JUST started arming +# after a real lapse, so the beacon is long stale but the entry is fresh. The +# arm's bounded startup window must never be stolen out from under it. +test_fresh_arming_claim_with_stale_beacon_is_never_reclaimed() { + local dir out status pid + dir=$(make_primary_dir "$TMP_ROOT/fresh-arming-claim") + : > "$dir/state/task1.meta" + write_arm_fixture "$dir" actionable + sleep 60 & + pid=$! + record_autoarm_owner "$dir" "$pid" + record_autoarm_owner_identity "$dir" "$pid" || fail "could not record a claim pid-identity" + printf 'epoch=464 owner_pid=%s outcome=arming updated_at=%s\n' "$pid" "$(date +%s)" \ + > "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.last-watcher-beat" + out=$(run_autoarm "$dir" 2>/dev/null); status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 0 "$status" "a freshly arming legacy claim must keep the single-flight gate closed even after a long lapse" + [ -z "$out" ] || fail "deferring to a fresh arming claim produced output: $out" + assert_absent "$dir/state/arm-ran" "a fresh arming claim was stolen and double-armed" + assert_present "$dir/state/.claude-autoarm.lock" "a fresh arming claim lost its owner lock" + pass "auto-arm: a fresh legacy arming claim is never reclaimed while its startup window is still open" } test_claim_not_named_by_the_ledger_is_never_reclaimed() { @@ -648,9 +794,11 @@ test_claim_not_named_by_the_ledger_is_never_reclaimed() { # The same unrecoverable lapse, reached where the ledger cannot prove it: a session # teardown kills the claim's whole process group before it records any outcome, so -# the entry still reads "arming" (in flight however old, by contract) while the -# recorded pid is later handed to an unrelated live process. Only the identity the -# claim recorded inside its own lock separates that from a real arm in progress. +# the entry still reads "arming" while the recorded pid is later handed to an +# unrelated live process. Only the identity the claim recorded inside its own lock +# separates that from a real arm in progress, so keep the beacon fresh here: this +# case must reclaim on the identity leg alone, not the stuck-arming leg. The +# reclaim must not signal the unrelated live process that inherited the number. test_pid_reused_arming_claim_is_reclaimed_and_rearms() { local dir out status pid dir=$(make_primary_dir "$TMP_ROOT/reused-pid-arming") @@ -662,7 +810,9 @@ test_pid_reused_arming_claim_is_reclaimed_and_rearms() { record_autoarm_owner "$dir" "$pid" record_autoarm_owner_identity "$dir" "$$" || fail "could not record a claim pid-identity" record_autoarm_epoch "$dir" 464 "$pid" arming + : > "$dir/state/.last-watcher-beat" out=$(run_autoarm "$dir" 2>/dev/null); status=$? + kill -0 "$pid" 2>/dev/null || fail "the unrelated live process inheriting the number must never be signalled" kill "$pid" 2>/dev/null || true wait "$pid" 2>/dev/null || true expect_code 2 "$status" "a claim whose recorded identity no longer matches its live pid must be reclaimed, arming entry or not" @@ -700,7 +850,8 @@ test_pid_reused_claim_with_no_ledger_is_reclaimed_and_rearms() { # The negative control for the identity leg: a claim whose recorded identity still # matches the process holding the lock is genuinely in flight, so an arm that has -# legitimately been running for hours must keep the single-flight gate closed. +# legitimately been running for hours - its watcher beating the whole time - must +# keep the single-flight gate closed. test_identity_matched_arming_claim_is_never_reclaimed() { local dir out status pid dir=$(make_primary_dir "$TMP_ROOT/identity-matched-arming") @@ -711,6 +862,7 @@ test_identity_matched_arming_claim_is_never_reclaimed() { record_autoarm_owner "$dir" "$pid" record_autoarm_owner_identity "$dir" "$pid" || fail "could not record a claim pid-identity" record_autoarm_epoch "$dir" 464 "$pid" arming + : > "$dir/state/.last-watcher-beat" out=$(run_autoarm "$dir" 2>/dev/null); status=$? kill "$pid" 2>/dev/null || true wait "$pid" 2>/dev/null || true @@ -743,6 +895,217 @@ test_terminal_check_claim_is_never_reclaimed() { pass "auto-arm: the guard's terminal-check claim is never reclaimed" } +# A proven-stuck legacy owner that is still ALIVE and identity-verified is +# retired with TERM before its lock is removed, because old-build code cannot +# re-check generations and would otherwise resume and act after supersession. +test_stuck_live_legacy_owner_is_retired_and_reclaimed() { + local dir out status pid + dir=$(make_primary_dir "$TMP_ROOT/legacy-term") + : > "$dir/state/task1.meta" + write_arm_fixture "$dir" actionable + sleep 60 & + pid=$! + record_autoarm_owner "$dir" "$pid" + record_autoarm_owner_identity "$dir" "$pid" || fail "could not record a claim pid-identity" + record_autoarm_epoch "$dir" 464 "$pid" arming + touch -t 202001010000 "$dir/state/.last-watcher-beat" + out=$(run_autoarm "$dir" 2>/dev/null); status=$? + expect_code 2 "$status" "a proven-stuck identity-verified live legacy owner must be retired and reclaimed" + kill -0 "$pid" 2>/dev/null && fail "the stuck legacy owner was reclaimed without being retired" + wait "$pid" 2>/dev/null || true + [ -e "$dir/state/arm-ran" ] || fail "the reclaimed home did not re-arm" + assert_contains "$out" "firstmate watcher wake" "the reclaimed cycle must still translate its wake" + assert_absent "$dir/state/.claude-autoarm.lock" "reclaim left the legacy owner lock behind" + pass "auto-arm: a stuck live legacy owner is retired via TERM and its lock reclaimed" +} + +# The SIGSTOP counterfactual: a stopped legacy owner survives the bounded +# retirement wait with TERM queued, and the reclaim must proceed anyway - a +# pending TERM on the verified owner is retirement-safe because delivery +# precedes any further user code when the process continues. +test_stopped_legacy_owner_is_reclaimed_with_term_pending() { + local dir out status pid i + dir=$(make_primary_dir "$TMP_ROOT/legacy-term-stopped") + : > "$dir/state/task1.meta" + write_arm_fixture "$dir" actionable + sleep 60 & + pid=$! + record_autoarm_owner "$dir" "$pid" + record_autoarm_owner_identity "$dir" "$pid" || fail "could not record a claim pid-identity" + record_autoarm_epoch "$dir" 464 "$pid" arming + touch -t 202001010000 "$dir/state/.last-watcher-beat" + kill -STOP "$pid" 2>/dev/null || fail "could not stop the legacy owner fixture" + out=$(run_autoarm "$dir" 2>/dev/null); status=$? + expect_code 2 "$status" "a stopped legacy owner with TERM queued must not block the reclaim forever" + [ -e "$dir/state/arm-ran" ] || fail "the reclaimed home did not re-arm past the stopped owner" + assert_absent "$dir/state/.claude-autoarm.lock" "reclaim left the stopped owner's lock behind" + kill -CONT "$pid" 2>/dev/null || true + i=0 + while [ "$i" -lt 40 ] && kill -0 "$pid" 2>/dev/null; do + sleep 0.05 + i=$((i + 1)) + done + kill -0 "$pid" 2>/dev/null && fail "the queued TERM did not retire the owner on continue" + wait "$pid" 2>/dev/null || true + pass "auto-arm: a SIGSTOPped legacy owner is reclaimed with TERM pending and dies on continue" +} + +# --- generation claims: optimistic single-flight and supersession -------------- +# The current claim is the two-line ledger entry itself (line 1 the classic +# epoch record, line 2 the owner's MANDATORY pid-identity); no lock is held +# across arming or output. A live open claim defers every firing; a stuck, +# dead, identity-mismatched, identityless, or finished claim is superseded by +# taking the next generation; a superseded owner goes completely silent. + +# Fabricate a v2 generation claim: +# . The identity of is recorded as line 2 (the +# claim's own pid for a matched claim, another pid to reproduce pid reuse). +record_autoarm_v2_claim() { + local dir=$1 gen=$2 owner=$3 outcome=$4 identity_pid=$5 identity + identity=$(fm_test_pid_identity "$identity_pid") || return 1 + [ -n "$identity" ] || return 1 + printf 'epoch=%s owner_pid=%s outcome=%s updated_at=1\n%s\n' \ + "$gen" "$owner" "$outcome" "$identity" > "$dir/state/.claude-autoarm-epoch" +} + +# A live open generation claim needs no lock to keep the gate closed: the +# ledger alone defers a concurrent firing, however old the entry, while the +# watcher keeps beating the beacon. +test_open_generation_claim_defers_without_any_lock() { + local dir out status pid + dir=$(make_primary_dir "$TMP_ROOT/v2-open-claim") + : > "$dir/state/task1.meta" + write_arm_fixture "$dir" actionable + sleep 60 & + pid=$! + record_autoarm_v2_claim "$dir" 464 "$pid" arming "$pid" || fail "could not record a v2 claim" + touch -t 202001010000 "$dir/state/.claude-autoarm-epoch" + : > "$dir/state/.last-watcher-beat" + assert_absent "$dir/state/.claude-autoarm.lock" "this case must start with no owner lock at all" + out=$(run_autoarm "$dir" 2>/dev/null); status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 0 "$status" "a live open generation claim must keep the single-flight gate closed with no lock held" + [ -z "$out" ] || fail "deferring to an open generation claim produced output: $out" + assert_absent "$dir/state/arm-ran" "an open generation claim was superseded and double-armed" + [ "$(epoch_field "$dir" epoch)" = 464 ] || fail "deferred firing rewrote the open claim's ledger entry" + pass "auto-arm: a live open generation claim defers concurrent firings with no lock held" +} + +# The 2026-08-26 watcher flap in the generation model: a live, identity-matched +# owner whose ledger entry and watcher beacon are both older than grace is +# stuck, and the next firing supersedes it by taking the next generation. +test_stuck_generation_claim_is_superseded_and_rearms() { + local dir out status pid + dir=$(make_primary_dir "$TMP_ROOT/v2-stuck-claim") + : > "$dir/state/task1.meta" + : > "$dir/state/task2.meta" + write_arm_fixture "$dir" actionable + sleep 60 & + pid=$! + record_autoarm_v2_claim "$dir" 464 "$pid" arming "$pid" || fail "could not record a v2 claim" + touch -t 202001010000 "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.last-watcher-beat" + out=$(run_autoarm "$dir" 2>/dev/null); status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 2 "$status" "a live owner stuck arming past grace with a beacon just as stale must be superseded, not deferred to forever" + [ -e "$dir/state/arm-ran" ] || fail "a stuck generation claim left the home unarmed with work in flight" + assert_contains "$out" "firstmate watcher wake" "the superseding generation must still translate its wake" + [ "$(epoch_field "$dir" epoch)" -gt 464 ] || fail "superseding claim did not advance the frozen ledger: $(epoch_field "$dir" epoch)" + [ "$(epoch_field "$dir" owner_pid)" != "$pid" ] || fail "superseding claim left the stuck owner on the ledger" + assert_absent "$dir/state/.claude-autoarm.lock" "the generation claim left a lock held after finishing" + pass "auto-arm: a hung generation owner with no watcher beat is superseded so re-arming self-heals" +} + +# Identity is mandatory at read time: a bare identityless one-line arming +# ledger naming an unrelated live pid is NOT an open claim - it must neither +# defer the hook nor survive as the current entry, whatever the beacon says. +test_identityless_ledger_never_defers() { + local dir out status pid + dir=$(make_primary_dir "$TMP_ROOT/v2-identityless-ledger") + : > "$dir/state/task1.meta" + write_arm_fixture "$dir" actionable + sleep 60 & + pid=$! + printf 'epoch=464 owner_pid=%s outcome=arming updated_at=1\n' "$pid" \ + > "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.claude-autoarm-epoch" + : > "$dir/state/.last-watcher-beat" + out=$(run_autoarm "$dir" 2>/dev/null); status=$? + kill -0 "$pid" 2>/dev/null || fail "the unrelated live pid on an identityless ledger must never be signalled" + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 2 "$status" "an identityless arming ledger must be superseded, never deferred to" + [ -e "$dir/state/arm-ran" ] || fail "an identityless ledger left the home unarmed" + [ "$(epoch_field "$dir" epoch)" -gt 464 ] || fail "the identityless entry was not superseded: $(epoch_field "$dir" epoch)" + pass "auto-arm: an identityless arming ledger never defers the gate (reused-pid loophole closed)" +} + +# A superseded owner must not start or attach another watcher: when its claim +# is superseded between arm attempts, the retry boundary goes silent instead +# of invoking the arm again. +test_superseded_owner_never_reinvokes_the_arm() { + local dir out status count + dir=$(make_primary_dir "$TMP_ROOT/v2-superseded-arm-boundary") + : > "$dir/state/task1.meta" + write_arm_fixture "$dir" supersede-then-fail + out=$(run_autoarm "$dir" 2>/dev/null); status=$? + expect_code 0 "$status" "an owner superseded between arm attempts must exit 0 silently" + [ -z "$out" ] || fail "a superseded owner produced output at the arm boundary: $out" + count=$(wc -l < "$dir/state/arm-ran" | tr -d ' ') + [ "$count" -eq 1 ] || fail "a superseded owner re-invoked the arm, saw $count arms" + [ "$(epoch_field "$dir" epoch)" = 999 ] || fail "a superseded owner rewrote its successor's ledger entry: $(epoch_field "$dir" epoch)" + pass "auto-arm: a superseded owner never re-invokes the arm and leaves its successor's claim untouched" +} + +# End-to-end regression for all three concurrency edge classes at once, with a +# REAL hook process hung mid-arm: +# 1. no mutex across blocking steps - while owner A is mid-arm, a concurrent +# firing B defers promptly instead of queueing on any lock; +# 2. stuck-owner supersession - once A's claim and the beacon age past grace +# while A is still alive arming, firing C takes the next generation and +# translates its own close (exit 2); +# 3. no double-translation - when A's arm finally returns, A finds itself +# superseded and goes completely silent (exit 0, no banner, no ledger +# write), so one supersession episode produces exactly one translation. +test_superseded_owner_goes_silent_and_never_double_translates() { + local dir a_out a_pid b_out b_status c_out c_status a_status i count + dir=$(make_primary_dir "$TMP_ROOT/v2-superseded-silence") + : > "$dir/state/task1.meta" + write_arm_fixture "$dir" blocking-actionable + a_out="$dir/state/a.out" + run_autoarm_bg "$dir" "$a_out" + a_pid=$RUN_AUTOARM_BG_PID + i=0 + while [ "$(epoch_outcome "$dir")" != arming ] || [ ! -e "$dir/state/arm-ran" ]; do + [ "$i" -lt 50 ] || fail "owner A never published its arming claim" + sleep 0.1 + i=$((i + 1)) + done + b_out=$(run_autoarm "$dir" 2>/dev/null); b_status=$? + expect_code 0 "$b_status" "a firing during a live open claim must defer promptly (no mutex is held across arming)" + [ -z "$b_out" ] || fail "deferring firing produced output: $b_out" + count=$(wc -l < "$dir/state/arm-ran" | tr -d ' ') + [ "$count" -eq 1 ] || fail "deferring firing must not arm, saw $count arms" + # A is still alive mid-arm; make its claim stuck-shaped. + kill -0 "$a_pid" 2>/dev/null || fail "owner A finished before the supersession could be exercised" + touch -t 202001010000 "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.last-watcher-beat" + c_out=$(run_autoarm "$dir" 2>/dev/null); c_status=$? + expect_code 2 "$c_status" "the superseding generation must translate its own close" + assert_contains "$c_out" "firstmate watcher wake" "the superseding generation must carry the rewake banner" + wait "$a_pid" + a_status=$? + expect_code 0 "$a_status" "the superseded owner must exit 0 instead of double-translating" + [ ! -s "$a_out" ] || fail "the superseded owner emitted output after losing its generation: $(cat "$a_out")" + [ "$(epoch_field "$dir" epoch)" = 2 ] || fail "the superseded owner advanced the ledger past its successor: $(epoch_field "$dir" epoch)" + [ "$(epoch_outcome "$dir")" = rewake ] || fail "the superseding generation's outcome was overwritten: $(epoch_outcome "$dir")" + count=$(wc -l < "$dir/state/arm-ran" | tr -d ' ') + [ "$count" -eq 2 ] || fail "expected exactly the owner and superseder arms, saw $count" + pass "auto-arm: a superseded owner goes silent - one supersession episode, one translation, no held mutex" +} + test_need_vanished_mid_cycle_closes_quietly() { local dir out status dir=$(make_primary_dir "$TMP_ROOT/vanished") @@ -798,19 +1161,29 @@ test_actionable_close_rewakes_with_reason test_actionable_close_with_live_successor_rewakes_once test_failed_close_rewakes_with_failure_banner test_failed_cycles_notify_once_and_keep_retrying +test_failure_notice_marker_write_refuses_delivery_and_retries test_unverified_clean_close_exhausts_retries test_post_alarm_actionable_close_is_suppressed test_benign_cycle_end_with_live_watcher_is_silent test_positive_recovery_budget_contention_preserves_episode +test_owner_mutex_contention_preserves_failure_episode_reset test_arms_for_x_mode_poll_need_without_inflight test_single_flight_admits_exactly_one_owner test_abandoned_owner_claim_is_reclaimed_and_rearms -test_arming_claim_is_never_reclaimed +test_arming_claim_with_fresh_beacon_is_never_reclaimed +test_fresh_arming_claim_with_stale_beacon_is_never_reclaimed test_claim_not_named_by_the_ledger_is_never_reclaimed test_pid_reused_arming_claim_is_reclaimed_and_rearms test_pid_reused_claim_with_no_ledger_is_reclaimed_and_rearms test_identity_matched_arming_claim_is_never_reclaimed test_terminal_check_claim_is_never_reclaimed +test_stuck_live_legacy_owner_is_retired_and_reclaimed +test_stopped_legacy_owner_is_reclaimed_with_term_pending +test_open_generation_claim_defers_without_any_lock +test_stuck_generation_claim_is_superseded_and_rearms +test_identityless_ledger_never_defers +test_superseded_owner_never_reinvokes_the_arm +test_superseded_owner_goes_silent_and_never_double_translates test_need_vanished_mid_cycle_closes_quietly test_afk_mid_cycle_suppresses_rewake test_active_in_marked_secondmate_home diff --git a/tests/fm-turnend-guard.test.sh b/tests/fm-turnend-guard.test.sh index 5c21f4d4306..54cfcdae861 100755 --- a/tests/fm-turnend-guard.test.sh +++ b/tests/fm-turnend-guard.test.sh @@ -1342,6 +1342,7 @@ test_hook_claude_mode_blocks_on_pid_reused_arming_claim() { printf '%s\n' "$identity" > "$dir/state/.claude-autoarm.lock/pid-identity" printf 'epoch=464 owner_pid=%s outcome=arming updated_at=1\n' "$pid" > "$dir/state/.claude-autoarm-epoch" touch -t 202001010000 "$dir/state/.claude-autoarm-epoch" + : > "$dir/state/.last-watcher-beat" out=$(FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=200 run_hook_claude "$dir" true); status=$? kill "$pid" 2>/dev/null || true wait "$pid" 2>/dev/null || true @@ -1351,6 +1352,77 @@ test_hook_claude_mode_blocks_on_pid_reused_arming_claim() { pass "fm-turnend-guard --claude: a claim whose pid was reused stops counting as recovery even while its entry reads arming" } +# The legacy stuck-arming shape (the 2026-08-26 flap): a live identity-matched +# lock-holding owner frozen at arming past grace with a beacon just as stale +# must not count as recovery under way. +test_hook_claude_mode_blocks_on_stuck_arming_claim() { + local dir out status pid identity + dir=$(make_primary_dir "$TMP_ROOT/hook-claude-stuck-arming-claim") + : > "$dir/state/task1.meta" + : > "$dir/state/task2.meta" + sleep 60 & + pid=$! + record_autoarm_owner "$dir" "$pid" + identity=$(fm_test_pid_identity "$pid") || fail "could not compute a claim pid-identity" + printf '%s\n' "$identity" > "$dir/state/.claude-autoarm.lock/pid-identity" + printf 'epoch=464 owner_pid=%s outcome=arming updated_at=1\n' "$pid" > "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.last-watcher-beat" + out=$(FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=200 run_hook_claude "$dir" true); status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 2 "$status" "a live owner stuck arming past grace with a stale beacon must not pass for recovery under way" + assert_contains "$out" "TURN WOULD END BLIND" "stuck-arming claim block must carry the blind-turn banner" + assert_contains "$out" "2 task(s) in flight" "stuck-arming claim block must name the unsupervised work" + pass "fm-turnend-guard --claude: a hung owner frozen at arming with no watcher beat no longer allows a blind stop" +} + +# The generation model's ownership proof: a live open ledger claim (two-line +# entry, identity-matched owner, watcher still beating) owns recovery with no +# lock held at all. +test_hook_claude_mode_allows_on_open_generation_claim() { + local dir out status pid identity + dir=$(make_primary_dir "$TMP_ROOT/hook-claude-open-generation") + : > "$dir/state/task1.meta" + sleep 60 & + pid=$! + identity=$(fm_test_pid_identity "$pid") || fail "could not compute a claim pid-identity" + printf 'epoch=464 owner_pid=%s outcome=arming updated_at=1\n%s\n' "$pid" "$identity" \ + > "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.claude-autoarm-epoch" + : > "$dir/state/.last-watcher-beat" + [ ! -e "$dir/state/.claude-autoarm.lock" ] || fail "this case must start with no owner lock at all" + out=$(run_hook_claude "$dir" false); status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 0 "$status" "--claude mode must allow when a live open generation claim owns recovery" + [ -z "$out" ] || fail "open-generation-claim allow produced output: $out" + pass "fm-turnend-guard --claude: a live open generation claim owns recovery with no lock held" +} + +# The same claim gone stuck (entry and beacon both past grace) stops counting +# as recovery even though its owner is alive and identity-matched. +test_hook_claude_mode_blocks_on_stuck_generation_claim() { + local dir out status pid identity + dir=$(make_primary_dir "$TMP_ROOT/hook-claude-stuck-generation") + : > "$dir/state/task1.meta" + : > "$dir/state/task2.meta" + sleep 60 & + pid=$! + identity=$(fm_test_pid_identity "$pid") || fail "could not compute a claim pid-identity" + printf 'epoch=464 owner_pid=%s outcome=arming updated_at=1\n%s\n' "$pid" "$identity" \ + > "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.claude-autoarm-epoch" + touch -t 202001010000 "$dir/state/.last-watcher-beat" + out=$(FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=200 run_hook_claude "$dir" true); status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 2 "$status" "a stuck generation claim must not pass for recovery under way" + assert_contains "$out" "TURN WOULD END BLIND" "stuck-generation-claim block must carry the blind-turn banner" + assert_contains "$out" "2 task(s) in flight" "stuck-generation-claim block must name the unsupervised work" + pass "fm-turnend-guard --claude: a stuck generation claim no longer allows a blind stop" +} + # The same abandoned claim on the terminal path: stepping aside for it allowed the # stop silently AND spent no attended alarm, so a genuinely broken automatic # mechanism stayed invisible. The guard must clear the claim and finish instead. @@ -1732,6 +1804,9 @@ test_hook_claude_mode_terminal_boundary_excludes_starting_owner test_hook_claude_mode_allows_on_fresh_rewake_epoch test_hook_claude_mode_blocks_on_abandoned_autoarm_claim test_hook_claude_mode_blocks_on_pid_reused_arming_claim +test_hook_claude_mode_blocks_on_stuck_arming_claim +test_hook_claude_mode_allows_on_open_generation_claim +test_hook_claude_mode_blocks_on_stuck_generation_claim test_hook_claude_mode_terminal_fail_open_clears_abandoned_claim test_hook_claude_mode_preserves_fresh_failed_progression test_hook_claude_mode_integrated_monotonic_fail_open From 7ee0c192e9d664b022361bd4609303bebfc7de14 Mon Sep 17 00:00:00 2001 From: Wojciech Kawecki Date: Thu, 27 Aug 2026 16:49:55 +0200 Subject: [PATCH 02/25] fix(bin): verify the real GitHub merge outcome instead of reporting an unproved merge (#3064) * fix(pr): verify GitHub merge outcome * no-mistakes(review): Captain, fixed forge-only merge verification, queue guidance, metadata propagation * no-mistakes(document): Correct forge-specific merge documentation * no-mistakes(review): Captain: forge-only queue fix, focused tests pass * no-mistakes(review): Captain: suppress closed-state guidance and prove parent regression * no-mistakes(review): Captain: remove history proof; retain executable regressions * no-mistakes(document): Clarify GitHub recording timing in architecture docs * no-mistakes(document): Clarify outcome-aware PR merge recording documentation * no-mistakes: apply CI fixes * Revert "no-mistakes: apply CI fixes" This reverts commit c326cfa9430c6173eedc8ff7f27d19d0552daf01. The automatic CI repair round removed the up-front `gh` prerequisite check while keeping the `gh` dependency: `bin/fm-pr-merge.sh` still calls `gh api graphql` for the outcome read and `gh api` for the branch-rules read. That left the same hard requirement without the clear named error, and review immediately raised a new finding for exactly the failure the check prevents - `gh-axi pr merge` landing the merge while the follow-up read fails, so the PR metadata is never recorded. The check is also symmetric with the GitLab arm directly above it, which already refuses up front when `glab` or `jq` is missing, on the stated principle that a missing tool should be a named prerequisite rather than a merge that is armed and then refused for an unexplained reason. The workflows this round was chasing sit at `action_required` because this is a fork pull request; no code change can turn them green. * fix(pr): keep PR bookkeeping when a merge outcome read fails On the GitHub path a merge call that returned success was followed by `github_read_outcome || exit 1`, so a transient API failure, rate limit, or network blip during the read dropped out of the script before `record_pr_metadata` ever ran. The merge could have landed while `pr=` went unrecorded and the merge poll was never armed - bookkeeping lost on a real merge. The failure path just above already recorded metadata before exiting, so the error path was more careful than the success one. Record the PR before that refusal. Recording arms the later merge poll and is not a success claim, which is the same reasoning that keeps `record_pr_metadata` on the gh-axi failure path. The refusal itself is unchanged: exit stays non-zero and the message still names the concrete observed state. Metadata is withheld only when the read succeeds and proves the pull request neither merged nor queued. Pin it with a case that stubs `gh api graphql` into failure after a successful `gh-axi pr merge`, asserting both the non-zero exit and the recorded metadata. * no-mistakes(review): Aggregate queue rules and report conflicts explicitly * fix(pr): keep the merge abstraction reachable and its bookkeeping intact Two holes remained in the outcome-verified GitHub merge path, both on installations where gh-axi is present but gh is not. The verification preflight refused before bin/fm-pr-merge.sh ever reached the configured gh-axi merge abstraction, so an installation without gh could no longer merge at all. gh-axi now performs the merge unconditionally and the queue-aware gh read became an optional enrichment: with gh on PATH its GraphQL view still separates merged from queued, and without gh the gh-axi view still proves a landed merge while every outcome it cannot prove refuses. The PR metadata recording sat behind the outcome read, so a merge that landed before that read failed lost pr= and its merge poll. Recording now happens once, before either forge call, which arms the poll without claiming a landed outcome and leaves teardown a PR identity to verify against no matter how the read ends. Rebasing onto main also restored the durable merge-outcome reporting and the GitLab landed-state confirmation that the conflict resolution dropped. Tests pin each fix through the executable interface: the merge abstraction is reached and verified with gh absent, a failed fallback read keeps its bookkeeping, and a mock that snapshots the task meta during the forge call proves pr= is recorded before the merge can land. * no-mistakes(review): fix(pr): de-dup queue methods, fall back on failed gh read, refresh contracts * no-mistakes(review): fix(pr): quote forge output and explain armed auto-merge on refusal * no-mistakes(review): fix(pr): claim auto-merge armed only when the forge accepted it * no-mistakes(review): fix(pr): tell the operator what each GitHub refusal could not observe * no-mistakes(review): fix(pr): gate every forge-acceptance claim on a successful merge * no-mistakes(document): align merge docs with verified GitHub outcome contract --- AGENTS.md | 2 +- bin/fm-pr-merge.sh | 413 +++++++++++- docs/architecture.md | 7 +- docs/gitlab-merge-watch.md | 2 +- docs/scripts.md | 2 +- tests/fm-pr-check-security.test.sh | 10 + tests/fm-pr-merge.test.sh | 1013 +++++++++++++++++++++++++++- 7 files changed, 1405 insertions(+), 44 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 673258c2bcb..89b40f466c8 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -336,7 +336,7 @@ Delivery mode and `yolo` are orthogonal. Never merge a red PR under either setting; destructive, irreversible, and security-sensitive merges still escalate. Without a current explicit captain instruction that states the concrete merge, that default stands, and standing `yolo` cannot authorize a red merge; section 1 owns when such an instruction overrides a Firstmate-written standing rule within its exact scope. Load `ask-user-authority` before deciding any ask-user finding; the implementation worker never answers its own finding. -Use `bin/fm-pr-merge.sh` for every task PR merge so merge metadata is recorded, and use `bin/fm-merge-local.sh` for approved local-only landing; never call a lower-level merge command around their guards. +Use `bin/fm-pr-merge.sh` for every task PR merge so merge metadata is recorded and an unproved merge is refused instead of reported as landed, and use `bin/fm-merge-local.sh` for approved local-only landing; never call a lower-level merge command around their guards. After an autonomous merge, give the captain a one-line full-URL or local-main outcome. ### Validate diff --git a/bin/fm-pr-merge.sh b/bin/fm-pr-merge.sh index d35bc9f30fb..3e61b33f7bc 100755 --- a/bin/fm-pr-merge.sh +++ b/bin/fm-pr-merge.sh @@ -8,6 +8,35 @@ # # Merge method on GitHub defaults to --squash when the caller passes none of # --squash, --merge, --rebase, or --method after the optional -- separator. +# The gh-axi merge abstraction always performs the merge; the outcome read that +# follows it never becomes a prerequisite for reaching that abstraction. After +# gh-axi returns success, GitHub's live state is read back and accepted only +# when the pull request is merged or in the merge queue. gh's GraphQL API +# supplies that queue-aware read when gh is on PATH; when gh is absent or its +# read fails, gh-axi's own view still proves a landed merge, and every outcome +# it cannot prove refuses, reporting the single failed read when gh is absent +# and naming both failed reads when gh is present and its own read failed. +# If the pull request remains open and the base branch has an effective +# merge_queue rule, the refusal names the queue's configured merge method and +# the exact -- --auto -- retry flags, unless the caller already passed +# that method with --auto to a merge command that returned success, in which +# case it reports instead that the accepted request has not entered the queue +# and the queue state has to be re-checked. +# No method is selected for the caller in any case. A rules response that names +# no queue rule, one that could not be read, rules that disagree, and a method +# this script does not recognise are four distinct outcomes and are reported +# apart, because each one leaves the operator somewhere different. +# A caller-requested --auto that leaves the pull request neither merged nor +# queued is refused the same way and says auto-merge was armed with nothing +# landed or queued yet, or, when the merge command itself failed, that auto-merge +# was only requested; both are read from the caller's own arguments rather than +# from the forge's prose. The observed state is judged the same way whichever +# read produced it, and a refusal built on the gh-axi view says the merge queue +# could not be observed at all rather than implying an unqueued pull request. +# Every refusal that follows a merge command which returned success quotes that +# command's own output, marked as the forge's text and kept apart from this +# script's verdict, including the refusal for an outcome that cannot be read; +# a merge command that failed keeps its original error surfaced raw and first. # GitLab adds no method flag at all: its merge method is the project's own # setting, which the merge API applies, and imposing squash there would override # that convention rather than mirror the GitHub default. @@ -28,10 +57,10 @@ # short-option cluster such as -yR, because the repository comes only from the # URL, nor --sha on GitLab because the head comes only from the live read. # -# After the forge command, this script confirms the PR is actually merged before -# reporting it; an auto-merge-queued or unconfirmed request leaves the poll armed -# and records no landed outcome. bin/fm-merge-outcome-lib.sh owns a confirmed -# merge's destination, normal-case deduplication, and at-least-once recovery. +# On GitLab, this script confirms the MR is actually merged before reporting it; +# an auto-merge-queued or unconfirmed request leaves the poll armed and records +# no landed outcome. bin/fm-merge-outcome-lib.sh owns a confirmed merge's +# destination, normal-case deduplication, and at-least-once recovery. # A landed merge whose outcome cannot be written is reported loudly rather than # misreported as a failed merge. # Usage: fm-pr-merge.sh [-- ] @@ -84,6 +113,47 @@ caller_has_merge_method() { return 1 } +# The merge method the caller's own extra arguments named, in the --flag, +# --method and --method= forms caller_has_merge_method accepts. +caller_merge_method() { + local arg method='' pending=false + for arg in "$@"; do + if [ "$pending" = true ]; then + method=$arg + pending=false + continue + fi + case "$arg" in + --squash) method=squash ;; + --merge) method=merge ;; + --rebase) method=rebase ;; + --method) pending=true ;; + --method=*) method=${arg#--method=} ;; + esac + done + printf '%s' "$method" +} + +# Whether the caller's own extra arguments asked for auto-merge, including the +# --flag=value spelling the forge's flag parser accepts. --disable-auto cancels +# the request, and gh exposes no short option that could bundle either flag. +caller_requested_auto_merge() { + local arg requested=1 + for arg in "$@"; do + case "$arg" in + --auto) requested=0 ;; + --auto=*) + case "${arg#--auto=}" in + [tT]|[tT][rR][uU][eE]|1) requested=0 ;; + *) requested=1 ;; + esac + ;; + --disable-auto) requested=1 ;; + esac + done + return "$requested" +} + reject_repo_overrides() { local arg for arg in "$@"; do @@ -147,12 +217,6 @@ if [ "$PROVIDER" = gitlab ]; then RECORDED_HEAD=$(grep '^pr_head=' "$META" | tail -1 | cut -d= -f2- || true) fi -"$SCRIPT_DIR/fm-pr-check.sh" "$ID" "$URL" -grep -qxF "pr=$URL" "$META" || { - echo "error: PR metadata recording failed" >&2 - exit 1 -} - # Pre-merge conditions for a GitLab merge request, read from one live view of # the merge request. Sets FM_PR_MERGE_HEAD to the verified head on success and # returns non-zero after reporting every condition that failed. @@ -254,22 +318,291 @@ FIELDS FM_PR_MERGE_HEAD=$live_head } -github_confirm_merged() { +# Read one live GitHub pull request view after gh-axi returns. The selected +# fields distinguish a landed pull request from a merge-queue entry and retain +# the concrete state needed for a refusal. gh supplies the complete queue-aware +# view when available; gh-axi remains the degradation path that can prove a +# landed merge without making gh a prerequisite for the merge abstraction. +FM_PR_GITHUB_STATE= +FM_PR_GITHUB_MERGED= +FM_PR_GITHUB_QUEUED= +FM_PR_GITHUB_BASE= +FM_PR_GITHUB_QUEUE_OBSERVED=false +github_read_outcome_with_gh() { + local fields line + local total=0 named=0 + local state='' merged='' queued='' base='' + + # shellcheck disable=SC2016 # GraphQL variables are literal query syntax. + if ! fields=$(gh api graphql \ + -f query='query($owner:String!,$repo:String!,$number:Int!){repository(owner:$owner,name:$repo){pullRequest(number:$number){state merged isInMergeQueue baseRefName}}}' \ + -F "owner=$PR_OWNER" -F "repo=$PR_REPO" -F "number=$PR_NUMBER" \ + --jq '.data.repository.pullRequest | "state=" + (.state // ""), "merged=" + (.merged | tostring), "queued=" + (.isInMergeQueue | tostring), "base=" + (.baseRefName // "")' \ + 2>/dev/null) || [ -z "$fields" ]; then + return 1 + fi + while IFS= read -r line; do + total=$((total + 1)) + case "$line" in + state=*) state=${line#state=} ;; + merged=*) merged=${line#merged=} ;; + queued=*) queued=${line#queued=} ;; + base=*) base=${line#base=} ;; + *) continue ;; + esac + named=$((named + 1)) + done </dev/null); then - printf 'actionable: GitHub accepted the merge request for %s but its landed state could not be confirmed; the merge poll remains armed\n' \ - "$URL" >&2 - return 2 + return 1 fi if ! state=$(printf '%s\n' "$output" | awk ' $1 == "state:" { count++; value=$2 } END { if (count == 1 && value != "") print value; else exit 1 } '); then - printf 'actionable: GitHub accepted the merge request for %s but its landed state could not be confirmed; the merge poll remains armed\n' \ + return 1 + fi + case "$state" in + merged) + FM_PR_GITHUB_STATE=MERGED + FM_PR_GITHUB_MERGED=true + FM_PR_GITHUB_QUEUED=false + ;; + *) + FM_PR_GITHUB_STATE=$state + FM_PR_GITHUB_MERGED=false + FM_PR_GITHUB_QUEUED=unknown + ;; + esac + FM_PR_GITHUB_BASE= + FM_PR_GITHUB_QUEUE_OBSERVED=false +} + +github_read_outcome() { + if ! command -v gh >/dev/null 2>&1; then + github_read_outcome_with_gh_axi && return 0 + echo "error: could not read the GitHub pull request outcome after the merge attempt; PR metadata and merge poll remain recorded" >&2 + return 1 + fi + # Only a failed gh read falls back. A gh read that completes and reports the + # pull request as neither merged nor queued is a concrete outcome, not a + # missing one, so it keeps its own refusal. The gh-axi view cannot observe the + # merge queue, so it can only turn this into a proved merge or into a refusal. + github_read_outcome_with_gh && return 0 + if github_read_outcome_with_gh_axi && [ "$FM_PR_GITHUB_MERGED" = true ]; then + return 0 + fi + echo "error: could not read the GitHub pull request outcome after the merge attempt: the gh read failed and the gh-axi view could not prove the outcome either; PR metadata and merge poll remain recorded" >&2 + return 1 +} + +github_urlencode_path_segment() { + local LC_ALL=C input=$1 encoded='' char octet hex + while [ -n "$input" ]; do + char=${input%"${input#?}"} + input=${input#?} + case "$char" in + [-._~a-zA-Z0-9]) encoded=$encoded$char ;; + *) + printf -v octet '%d' "'$char" + [ "$octet" -ge 0 ] || octet=$((octet + 256)) + printf -v hex '%02X' "$octet" + encoded=$encoded%$hex + ;; + esac + done + printf '%s' "$encoded" +} + +# Read the effective merge-queue method for the observed base branch. The four +# situations the refusal has to keep apart - no queue rule, a rules response +# that could not be read, several rules that disagree, and a rule whose method +# this script does not recognise - are reported as a status rather than folded +# into one failure, because each one means something different to the operator. +FM_PR_GITHUB_QUEUE_METHOD= +FM_PR_GITHUB_QUEUE_METHODS= +FM_PR_GITHUB_QUEUE_STATUS=unreadable +github_read_queue_method() { + local methods line candidate method='' count=0 branch_path + local unrecognised=false conflicting=false + FM_PR_GITHUB_QUEUE_METHOD= + FM_PR_GITHUB_QUEUE_METHODS= + FM_PR_GITHUB_QUEUE_STATUS=unreadable + command -v gh >/dev/null 2>&1 || return 0 + [ -n "$FM_PR_GITHUB_BASE" ] || return 0 + branch_path=$(github_urlencode_path_segment "$FM_PR_GITHUB_BASE") + if ! methods=$(gh api \ + --paginate "repos/$PR_OWNER/$PR_REPO/rules/branches/$branch_path" \ + --jq '.[] | select(.type == "merge_queue") | "merge_method=" + (.parameters.merge_method // "")' \ + 2>/dev/null); then + return 0 + fi + while IFS= read -r line; do + [ -n "$line" ] || continue + case "$line" in + merge_method=*) candidate=${line#merge_method=} ;; + *) return 0 ;; + esac + count=$((count + 1)) + case "$candidate" in + MERGE|SQUASH|REBASE) ;; + *) unrecognised=true ;; + esac + if [ -z "$FM_PR_GITHUB_QUEUE_METHODS" ] && [ "$count" -eq 1 ]; then + FM_PR_GITHUB_QUEUE_METHODS=$candidate + else + case ",$FM_PR_GITHUB_QUEUE_METHODS," in + *",$candidate,"*) ;; + *) + FM_PR_GITHUB_QUEUE_METHODS="$FM_PR_GITHUB_QUEUE_METHODS,$candidate" + conflicting=true + ;; + esac + fi + method=$candidate + done <&2 + return 1 + } +} + +FM_PR_GITHUB_AUTO_REQUESTED=false +FM_PR_GITHUB_MERGE_ACCEPTED=false +FM_PR_GITHUB_CALLER_METHOD= + +# The single gate every statement about what the forge accepted, armed, or +# reported has to pass. A merge command that failed accepted nothing, so no +# such statement may be made on its path, and routing them all through one +# predicate keeps a later one from being written without the gate. +github_merge_command_succeeded() { + [ "$FM_PR_GITHUB_MERGE_ACCEPTED" = true ] +} + +github_report_forge_output() { + local output=$1 line + github_merge_command_succeeded || return 0 + [ -n "$output" ] || return 0 + echo "error: the merge command's own output follows, quoted; it is the forge CLI's report, not this script's verdict:" >&2 + while IFS= read -r line; do + printf 'error: > %s\n' "$line" >&2 + done <&2 + else + printf 'error: base branch %s requires the merge queue; retry with: %s %s %s -- --auto --%s\n' \ + "$FM_PR_GITHUB_BASE" "$0" "$ID" "$URL" "$queue_method" >&2 + fi + ;; + conflicting) + printf 'error: base branch %s has conflicting merge queue methods (%s); exact retry flags are ambiguous\n' \ + "$FM_PR_GITHUB_BASE" "${FM_PR_GITHUB_QUEUE_METHODS//,/, }" >&2 + ;; + unrecognised) + methods_display=${FM_PR_GITHUB_QUEUE_METHODS//,/, } + [ -n "$methods_display" ] || methods_display='' + printf 'error: base branch %s requires the merge queue, but its configured merge method (%s) is not one this script recognises, so exact retry flags cannot be named\n' \ + "$FM_PR_GITHUB_BASE" "$methods_display" >&2 + ;; + unreadable) + printf 'error: the branch rules for base branch %s could not be read, so a merge queue requirement can be neither confirmed nor ruled out here\n' \ + "${FM_PR_GITHUB_BASE:-}" >&2 + ;; + esac +} + +github_report_unmerged_outcome() { + printf 'error: GitHub merge outcome was not successful: state=%s, merged=%s, isInMergeQueue=%s\n' \ + "$FM_PR_GITHUB_STATE" "$FM_PR_GITHUB_MERGED" "$FM_PR_GITHUB_QUEUED" >&2 + if ! github_state_is_open || [ "$FM_PR_GITHUB_MERGED" != false ] \ + || [ "$FM_PR_GITHUB_QUEUED" = true ]; then + return 0 + fi + if [ "$FM_PR_GITHUB_AUTO_REQUESTED" = true ]; then + if github_merge_command_succeeded; then + printf 'error: auto-merge was requested and armed for %s, but nothing is merged or in the merge queue yet, so this run refuses instead of reporting an unproved merge\n' \ + "$URL" >&2 + else + printf 'error: auto-merge was requested for %s, but the merge command itself failed, so nothing was enabled, merged or queued\n' \ + "$URL" >&2 + fi + fi + if [ "$FM_PR_GITHUB_QUEUE_OBSERVED" != true ]; then + printf 'error: the merge queue could not be observed for %s because the queue-aware read was unavailable, so a pull request already in the merge queue cannot be told apart from one that never entered it; re-check the pull request'"'"'s merge queue state before retrying\n' \ "$URL" >&2 - return 2 + return 0 fi - [ "$state" = merged ] + github_report_queue_rules } gitlab_confirm_merged() { @@ -290,16 +623,54 @@ gitlab_confirm_merged() { [ "$state" = merged ] } +# Record before either forge call. This arms the merge poll without claiming a +# landed outcome, so even a provider read failure after a real merge cannot +# leave teardown without the PR identity it needs to verify the result. +record_pr_metadata || exit 1 + case "$PROVIDER" in github) + merge_output= merge_args=() if ! caller_has_merge_method "$@"; then merge_args=(--squash) fi - gh-axi pr merge "$PR_NUMBER" --repo "$PR_OWNER/$PR_REPO" "${merge_args[@]+"${merge_args[@]}"}" "$@" - github_confirm_rc=0 - github_confirm_merged || github_confirm_rc=$? - [ "$github_confirm_rc" -eq 0 ] || exit 0 + if caller_requested_auto_merge "$@"; then + FM_PR_GITHUB_AUTO_REQUESTED=true + fi + FM_PR_GITHUB_CALLER_METHOD=$(caller_merge_method "$@") + if merge_output=$(gh-axi pr merge "$PR_NUMBER" --repo "$PR_OWNER/$PR_REPO" \ + "${merge_args[@]+"${merge_args[@]}"}" "$@" 2>&1); then + FM_PR_GITHUB_MERGE_ACCEPTED=true + else + merge_status=$? + [ -z "$merge_output" ] || printf '%s\n' "$merge_output" >&2 + if github_read_outcome; then + if [ "$FM_PR_GITHUB_MERGED" != true ] && [ "$FM_PR_GITHUB_QUEUED" != true ]; then + github_report_unmerged_outcome + else + printf 'actionable: the merge command for %s failed, but the pull request reads back as state=%s, merged=%s, isInMergeQueue=%s\n' \ + "$URL" "$FM_PR_GITHUB_STATE" "$FM_PR_GITHUB_MERGED" "$FM_PR_GITHUB_QUEUED" >&2 + fi + fi + exit "$merge_status" + fi + if ! github_read_outcome; then + github_report_forge_output "$merge_output" + exit 1 + fi + if [ "$FM_PR_GITHUB_MERGED" = true ]; then + printf 'verified: %s is merged (state=%s, merged=%s, isInMergeQueue=%s)\n' \ + "$URL" "$FM_PR_GITHUB_STATE" "$FM_PR_GITHUB_MERGED" "$FM_PR_GITHUB_QUEUED" + elif [ "$FM_PR_GITHUB_QUEUED" = true ]; then + printf 'verified: %s is queued (state=%s, merged=%s, isInMergeQueue=%s)\n' \ + "$URL" "$FM_PR_GITHUB_STATE" "$FM_PR_GITHUB_MERGED" "$FM_PR_GITHUB_QUEUED" + exit 0 + else + github_report_forge_output "$merge_output" + github_report_unmerged_outcome + exit 1 + fi ;; gitlab) gitlab_verify_mergeable || exit 1 diff --git a/docs/architecture.md b/docs/architecture.md index 66261873270..0b18c16f0dc 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -275,7 +275,12 @@ The helper requires a full canonical URL and rejects malformed URLs or repo over A `https://github.com///pull/` URL invokes `gh-axi pr merge --repo /`, defaults to `--squash`, and preserves explicit merge-method flags. A `https:////-/merge_requests/` URL (see [docs/gitlab-merge-watch.md](gitlab-merge-watch.md)) invokes `glab mr merge -R https:///`, so the instance comes from the URL, and adds no merge-method flag because the project's own merge method applies. That path merges only after one live read of the merge request confirms it is open, mergeable, conflict-free, with blocking discussions resolved and a successful pipeline at the current head, and it binds the merge to that verified head; recorded metadata is never the authority for those conditions because a rebase leaves it stale. -After either forge command returns, the script confirms the PR or MR is actually merged; an auto-merge-queued or unconfirmed request records no landed outcome and leaves its poll armed. +After either forge command returns, the script confirms the PR or MR actually landed, and only a confirmed landing records a landed outcome; a queued or unconfirmed request records none and leaves its poll armed. +On GitLab an auto-merge-queued or unconfirmed request is reported without failing the run. +On GitHub an outcome that is neither merged nor queued is refused loudly and non-zero, naming the observed state, and a base branch that requires the merge queue is refused with the concrete retry flags its configured method requires rather than having a merge method chosen on the caller's behalf. +When the forge already accepted exactly those flags and the pull request still has not entered the queue, that refusal points at the queue state to re-check instead of echoing back the flags the caller just ran. +An auto-merge request is held to the same standard: `--auto` that leaves the pull request neither merged nor queued is refused rather than reported as success. +Every GitHub refusal states what it could not observe as plainly as what it did, so an unreadable branch-rule response, an unrecognised queue method, and a merge queue no available read can see are each named rather than left to look like a base branch with no queue at all. A confirmed merge leaves a durable role-routed outcome instead of living only in the merging agent's memory, and [`bin/fm-merge-outcome-lib.sh`](../bin/fm-merge-outcome-lib.sh)'s header owns its destination, shape, identity, normal-case deduplication, and at-least-once recovery. The same emitter handles a merge firstmate performed and one its poll detected, while the watcher immediately delivers the emitter's local actionable poll row. Teardown is fail-closed for ship worktrees: dirty worktrees refuse, and committed work must be landed before the worktree is returned. diff --git a/docs/gitlab-merge-watch.md b/docs/gitlab-merge-watch.md index 79dc138e1f6..3a66aacaf35 100644 --- a/docs/gitlab-merge-watch.md +++ b/docs/gitlab-merge-watch.md @@ -208,7 +208,7 @@ No armed watch is lost by upgrading. ## Merging a merge request -`bin/fm-pr-merge.sh` now merges a GitLab merge request through the same recording and the same guards a GitHub pull request gets. +`bin/fm-pr-merge.sh` now merges a GitLab merge request through the shared recording helper and GitLab's own live pre-merge guards. Every run below used a throwaway `FM_HOME`, so no live task record was touched, and a `glab` wrapper that refused any `merge` subcommand outright, so no merge could reach the forge even if a check were wrong. That wrapper is why the open fixture merge request could be used as evidence at all: it is `mergeable` with discussions resolved, so the pipeline conditions are the only thing between it and a real merge. diff --git a/docs/scripts.md b/docs/scripts.md index 758b9c1d614..e45d23c86d2 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -115,7 +115,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-pr-poll.sh` | Provide the byte-static watcher program for validated PR/MR-poll sidecars | | `fm-pr-check-migrate.sh` | Quarantine older task polls without execution and rebuild only canonical polls | | `fm-pr-check.sh` | Record validated `pr=` and `pr_head=` values, then atomically arm a static merge poll | -| `fm-pr-merge.sh` | Record PR metadata, then merge a task's canonical full GitHub or GitLab URL | +| `fm-pr-merge.sh` | Record PR metadata, merge a task's canonical full GitHub or GitLab URL, then refuse an outcome it cannot prove landed or queued | | `fm-merge-outcome-lib.sh` | Publish a confirmed merge's durable, role-routed supervision outcome | | `fm-promote.sh` | Promote a scout task in place to a protected ship task with an explicit delivery mode | | `fm-teardown.sh` | Fail-closed teardown: return landed ship worktrees, require completed scout deliverables, retire secondmate homes | diff --git a/tests/fm-pr-check-security.test.sh b/tests/fm-pr-check-security.test.sh index 700dd74bf0d..f39e941d411 100755 --- a/tests/fm-pr-check-security.test.sh +++ b/tests/fm-pr-check-security.test.sh @@ -80,6 +80,16 @@ SH cat > "$fakebin/gh" <<'SH' #!/usr/bin/env bash printf '%s\n' "$*" >> "$FM_TEST_GH_LOG" +case "${1:-} ${2:-}" in + "api graphql") + printf '%s\n' \ + 'state=MERGED' \ + 'merged=true' \ + 'queued=false' \ + 'base=main' + exit 0 + ;; +esac case " $* " in *" headRefOid "*) printf '%s\n' "${FM_TEST_GH_HEAD:-0123456789abcdef0123456789abcdef01234567}" ;; *" state "*) diff --git a/tests/fm-pr-merge.test.sh b/tests/fm-pr-merge.test.sh index d3842939ce4..f75bc58c07b 100755 --- a/tests/fm-pr-merge.test.sh +++ b/tests/fm-pr-merge.test.sh @@ -1,12 +1,12 @@ #!/usr/bin/env bash # Tests for bin/fm-pr-merge.sh: the one path firstmate uses to merge a task's -# PR, which must always record pr= and any available pr_head= into the task's -# meta before merging so fm-teardown.sh's landed-check has a PR reference to -# verify against, even on repos with no PR CI where the usual "checks green" -# fm-pr-check.sh trigger never fires. +# PR, which must record pr= and any available pr_head= into the task's meta so +# fm-teardown.sh's landed-check has a PR reference to verify against, even on +# repos with no PR CI where the usual "checks green" fm-pr-check.sh trigger +# never fires. # # Matrix: -# (a) merge records pr= and pr_head= before merging, and merges +# (a) a verified merge records pr= and pr_head= # (b) merge is refused when gh-axi pr merge itself fails (no silent success) # (c) extra gh-axi pr merge args are forwarded after number and --repo # (d) merge is refused before gh-axi when task meta is missing @@ -24,17 +24,50 @@ # (o) glab or jq absent refuses before any state is recorded # (p) --sha in extra GitLab args fails fast, and still forwards on GitHub # (q) a GitLab refusal still leaves pr= recorded and the merge poll armed -# (r) a successful merge in a secondmate home reports the landed PR upward +# (r) GitHub success is accepted only after the PR is read back as merged +# (s) an open GitHub PR that is neither merged nor queued fails verification +# (t) a GitHub PR in the merge queue is reported as queued, not merged +# (u) a queue-required refusal names the exact compatible retry flags +# (v) a failed poll setup cannot be reported as a verified GitHub merge +# (w) a zero-exit queue-required refusal keeps merge semantics unchanged +# (x) an unreadable outcome after a successful merge call keeps the PR +# recorded and the merge poll armed +# (y) agreeing queue rules still produce exact retry flags +# (z) conflicting queue rules report ambiguous retry guidance +# (aa) gh-axi remains usable when gh is absent +# (ab) a landed merge whose fallback outcome read fails keeps its poll armed +# (ac) a successful merge in a secondmate home reports the landed PR upward # once, on the route its parent binding names, and a repeat merge of the # same PR does not duplicate that line -# (s) a refused or failed merge reports nothing -# (t) a successful merge in a main home leaves a durable wake naming the PR -# (u) a secondmate home with no usable parent binding says so loudly instead +# (ad) a refused or failed merge reports nothing +# (ae) a successful merge in a main home leaves a durable wake naming the PR +# (af) a secondmate home with no usable parent binding says so loudly instead # of merging in silence -# (v) an accepted queued GitHub merge emits nothing and leaves its poll armed -# (w) an accepted queued GitLab merge emits nothing and leaves its poll armed -# (x) an uncommitted marker retry never loses the durable outcome -# (y) distinct merged PRs for a reused task each survive queue deduplication +# (ag) an accepted queued GitHub merge emits nothing and leaves its poll armed +# (ah) an accepted queued GitLab merge emits nothing and leaves its poll armed +# (ai) an uncommitted marker retry never loses the durable outcome +# (aj) distinct merged PRs for a reused task each survive queue deduplication +# (ak) pr= is already recorded when the forge call that can land the merge runs +# (al) a failed gh read falls back to the gh-axi view, which can prove a merge +# (am) a failed merge command still names an outcome read that proves a landed +# or queued pull request, without masking the forge failure +# (an) a refusal after a zero-exit merge quotes the forge's own output, marked +# apart from the wrapper's verdict and never leaked to stdout +# (ao) a caller-requested auto-merge on a queue-less base refuses and says +# auto-merge is armed with nothing merged or queued yet +# (ap) a caller-requested auto-merge whose merge command failed refuses +# without ever claiming auto-merge was armed +# (aq) an outcome read that fails after a zero-exit merge still quotes the +# forge's own output, the only evidence left +# (ar) auto-merge with the queue's own method that is still unqueued refuses +# without echoing back the flags just used, and names the next step +# (as) a caller method the queue does not use still gets exact retry flags +# (at) an unrecognised queue method still names the queue requirement and +# guesses no method +# (au) unreadable branch rules are reported apart from a queue-less base +# (av) a base branch with no queue rule says nothing about a merge queue +# (aw) a refusal built on the gh-axi view says the merge queue could not be +# observed, and judges that view's state like the queue-aware one set -u # shellcheck source=tests/lib.sh @@ -55,6 +88,7 @@ MR_HEAD=aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa MR_STALE_HEAD=bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb JQ_BIN=$(command -v jq) || fail "these tests read glab's JSON with the real jq, which was not found" +REAL_MV=$(command -v mv) || fail "these tests need mv to simulate a failed poll publish" # Build a fresh sandbox for one test case: a state dir with a task meta and a # fakebin with a gh-axi mock that records how it was invoked. Echoes the case dir. @@ -69,6 +103,13 @@ make_case() { "project=$case_dir/project" \ "kind=ship" \ "mode=no-mistakes" + printf '%s\n' \ + 'state=MERGED' \ + 'merged=true' \ + 'queued=false' \ + 'base=main' > "$case_dir/github-outcome" + : > "$case_dir/github-rules" + : > "$case_dir/gh.log" # No worktree/project on disk; fm-pr-check.sh tolerates a worktree it cannot # stat and simply skips the pr_head lookup via `gh` in that case, so give it # one that resolves for cases that want pr_head recorded. @@ -83,6 +124,7 @@ add_gh_mocks() { #!/usr/bin/env bash printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" case "${1:-} ${2:-}" in + "pr merge") printf 'merged:\n number: %s\n status: ok\n' "${3:-}" ;; "pr view") [ "$#" -eq 5 ] && [ "${4:-}" = --repo ] || exit 2 printf 'pull_request:\n number: %s\n state: %s\n' "$3" "${FM_TEST_GH_MERGE_STATE:-merged}" @@ -92,12 +134,21 @@ exit 0 SH cat > "$case_dir/fakebin/gh" <> "\$FM_TEST_GH_LOG" case "\${1:-} \${2:-}" in "pr view") case " \$* " in *headRefOid*) printf '%s\n' '$head' ; exit 0 ;; esac ;; + "api graphql") + cat "\$FM_TEST_GH_OUTCOME" + exit 0 + ;; + api\ *) + cat "\$FM_TEST_GH_RULES" + exit 0 + ;; esac exit 0 SH @@ -113,16 +164,81 @@ add_gh_mocks_merge_fails() { printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" case "${1:-} ${2:-}" in "pr merge") echo "error: pr merge failed" >&2 ; exit 1 ;; -esac -exit 0 + esac + exit 0 SH cat > "$case_dir/fakebin/gh" <<'SH' #!/usr/bin/env bash +printf '%s\n' "$*" >> "$FM_TEST_GH_LOG" +case "${1:-} ${2:-}" in + "api graphql") + cat "$FM_TEST_GH_OUTCOME" + exit 0 + ;; + api\ *) + cat "$FM_TEST_GH_RULES" + exit 0 + ;; +esac exit 0 SH chmod +x "$case_dir/fakebin/gh-axi" "$case_dir/fakebin/gh" } +# gh mock that still answers fm-pr-check.sh's head lookup but cannot answer the +# outcome read, so a merge call that returned success is followed by a live +# state nothing can prove. Args: case_dir head_sha +add_gh_mock_outcome_read_fails() { + local case_dir=$1 head=$2 + cat > "$case_dir/fakebin/gh" <> "\$FM_TEST_GH_LOG" +case "\${1:-} \${2:-}" in + "pr view") + case " \$* " in + *headRefOid*) printf '%s\n' '$head' ; exit 0 ;; + esac + ;; + "api graphql") + echo 'error: could not reach the GitHub API' >&2 + exit 1 + ;; +esac +exit 0 +SH + chmod +x "$case_dir/fakebin/gh" +} + +# gh-axi mock that merges but cannot answer its own view, so a case can prove +# what happens when neither reader can establish the outcome. Args: case_dir +add_gh_axi_mock_view_fails() { + local case_dir=$1 + cat > "$case_dir/fakebin/gh-axi" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" +case "${1:-} ${2:-}" in + "pr merge") printf 'merged:\n number: %s\n status: ok\n' "${3:-}" ;; + "pr view") exit 1 ;; +esac +exit 0 +SH + chmod +x "$case_dir/fakebin/gh-axi" +} + +add_failing_poll_publish_mv() { + local case_dir=$1 + cat > "$case_dir/fakebin/mv" <<'SH' +#!/usr/bin/env bash +for arg in "$@"; do + case "$arg" in + */.fm-pr-poll-data.*) exit 1 ;; + esac +done +exec "$FM_TEST_REAL_MV" "$@" +SH + chmod +x "$case_dir/fakebin/mv" +} + # glab mock recording every invocation together with the GITLAB_HOST it was # given, so a test can prove the instance came from the URL. `mr view` answers # from the case's JSON payload; marker files in the case dir drive the failure @@ -241,6 +357,11 @@ run_pr_merge() { FM_HOME="${FM_TEST_HOME:-$ROOT}" \ FM_STATE_OVERRIDE="$case_dir/state" \ FM_TEST_GH_AXI_LOG="$case_dir/gh-axi.log" \ + FM_TEST_GH_LOG="$case_dir/gh.log" \ + FM_TEST_GH_OUTCOME="$case_dir/github-outcome" \ + FM_TEST_GH_RULES="$case_dir/github-rules" \ + FM_TEST_META_AT_MERGE="$case_dir/meta-at-merge" \ + FM_TEST_REAL_MV="$REAL_MV" \ FM_TEST_GLAB_LOG="$case_dir/glab.log" \ FM_TEST_GLAB_JSON="$case_dir/mr.json" \ PATH="$case_dir/fakebin:$PATH" \ @@ -253,7 +374,16 @@ run_pr_merge() { return "$rc" } -test_records_pr_and_head_before_merging() { +write_github_outcome() { + local case_dir=$1 state=$2 merged=$3 queued=$4 base=$5 + printf '%s\n' \ + "state=$state" \ + "merged=$merged" \ + "queued=$queued" \ + "base=$base" > "$case_dir/github-outcome" +} + +test_verified_merge_records_pr_and_head() { local case_dir rc case_dir=$(make_case records-before-merge) mkdir -p "$case_dir/wt" @@ -273,7 +403,47 @@ test_records_pr_and_head_before_merging() { "records-before-merge: pr_head= was not recorded" grep -qxF 'pr merge 9 --repo example/repo --squash' "$case_dir/gh-axi.log" \ || fail "records-before-merge: gh-axi pr merge was not invoked with number, --repo, and default --squash" - pass "fm-pr-merge records pr= and pr_head= before invoking gh-axi pr merge" + pass "fm-pr-merge records pr= and pr_head= for a verified GitHub merge" +} + +# The forge call is the point of no return: once gh-axi has merged, nothing this +# script does afterwards can un-merge it. Proving pr= is already in the task's +# meta at that moment is what makes a later failure unable to lose the merge. +test_pr_metadata_is_recorded_before_the_forge_call() { + local case_dir rc + case_dir=$(make_case records-ahead-of-forge-call) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 5151515151515151515151515151515151515151 + cat > "$case_dir/fakebin/gh-axi" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" +case "${1:-} ${2:-}" in + "pr merge") + cat "$FM_STATE_OVERRIDE/task-x1.meta" > "$FM_TEST_META_AT_MERGE" + printf 'merged:\n number: %s\n status: ok\n' "${3:-}" + ;; + "pr view") + printf 'pull_request:\n number: %s\n state: merged\n' "$3" + ;; +esac +exit 0 +SH + chmod +x "$case_dir/fakebin/gh-axi" + : > "$case_dir/gh-axi.log" + : > "$case_dir/meta-at-merge" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/62 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 0 "$rc" "records-ahead-of-forge-call: fm-pr-merge should succeed" + assert_grep 'pr merge 62 --repo example/repo --squash' "$case_dir/gh-axi.log" \ + "records-ahead-of-forge-call: the merge abstraction was never invoked" + assert_grep 'pr=https://github.com/example/repo/pull/62' "$case_dir/meta-at-merge" \ + "records-ahead-of-forge-call: the merge ran before pr= was recorded" + pass "fm-pr-merge records pr= before the forge call can land the merge" } test_merge_failure_propagates_after_recording() { @@ -295,6 +465,785 @@ test_merge_failure_propagates_after_recording() { pass "fm-pr-merge propagates a real merge failure without silently succeeding" } +test_github_merged_outcome_is_verified() { + local case_dir rc + case_dir=$(make_case github-verified-merged) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 1010101010101010101010101010101010101010 + : > "$case_dir/gh-axi.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/51 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 0 "$rc" "github-verified-merged: a merged PR should succeed" + assert_grep 'verified: https://github.com/example/repo/pull/51 is merged' \ + "$case_dir/stdout" "github-verified-merged: success was not reported as verified" + assert_grep 'api graphql' "$case_dir/gh.log" \ + "github-verified-merged: the PR outcome was not read back after merging" + pass "fm-pr-merge verifies a genuinely merged GitHub pull request" +} + +test_github_verified_merge_requires_poll_recording() { + local case_dir rc + case_dir=$(make_case github-poll-recording-fails) + add_gh_mocks "$case_dir" 1111111111111111111111111111111111111111 + add_failing_poll_publish_mv "$case_dir" + : > "$case_dir/gh-axi.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/55 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-poll-recording-fails: poll setup failure should fail the merge wrapper" + assert_grep 'error: could not publish PR poll' "$case_dir/stderr" \ + "github-poll-recording-fails: poll setup failure was not reported" + assert_no_grep 'verified: ' "$case_dir/stdout" \ + "github-poll-recording-fails: failed poll setup was reported as a verified merge" + assert_grep 'pr=https://github.com/example/repo/pull/55' "$case_dir/state/task-x1.meta" \ + "github-poll-recording-fails: metadata was not retained for the attempted merge" + assert_absent "$case_dir/state/task-x1.check.sh" \ + "github-poll-recording-fails: the failed poll setup left a runnable poll" + pass "fm-pr-merge refuses to claim a merge when poll recording fails" +} + +test_github_open_unqueued_outcome_refuses() { + local case_dir rc + case_dir=$(make_case github-open-unqueued) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 2020202020202020202020202020202020202020 + write_github_outcome "$case_dir" OPEN false false master + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/52 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-open-unqueued: an unproved merge must fail" + assert_grep 'state=OPEN, merged=false, isInMergeQueue=false' "$case_dir/stderr" \ + "github-open-unqueued: refusal did not name the concrete observed state" + assert_grep 'pr=https://github.com/example/repo/pull/52' "$case_dir/state/task-x1.meta" \ + "github-open-unqueued: the attempted merge lost its PR reference" + assert_present "$case_dir/state/task-x1.check.sh" \ + "github-open-unqueued: the attempted merge did not leave its poll armed" + pass "fm-pr-merge refuses a GitHub merge call that leaves the PR open and unqueued" +} + +test_github_unreadable_outcome_keeps_pr_bookkeeping() { + local case_dir rc + case_dir=$(make_case github-outcome-read-fails) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 3131313131313131313131313131313131313131 + add_gh_mock_outcome_read_fails "$case_dir" 3131313131313131313131313131313131313131 + add_gh_axi_mock_view_fails "$case_dir" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/57 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-outcome-read-fails: an unreadable outcome must fail" + assert_grep 'could not read the GitHub pull request outcome after the merge attempt' \ + "$case_dir/stderr" "github-outcome-read-fails: the unreadable outcome was not reported" + assert_grep 'the gh read failed and the gh-axi view could not prove the outcome either' \ + "$case_dir/stderr" "github-outcome-read-fails: the refusal did not name both failed reads" + assert_no_grep 'verified: ' "$case_dir/stdout" \ + "github-outcome-read-fails: an unproved merge was reported as verified" + # The merge call itself returned success, so the pull request may well have + # landed. Losing the reference here would leave teardown with nothing to + # verify against and no merge poll to catch up. + assert_grep 'pr=https://github.com/example/repo/pull/57' "$case_dir/state/task-x1.meta" \ + "github-outcome-read-fails: a successful merge call lost its PR reference" + assert_present "$case_dir/state/task-x1.check.sh" \ + "github-outcome-read-fails: no merge poll was armed for a merge that may have landed" + pass "fm-pr-merge keeps PR bookkeeping when it cannot read a successful merge call's outcome" +} + +test_github_refusal_quotes_the_forge_output() { + local case_dir rc + case_dir=$(make_case github-refusal-quotes-forge) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 6161616161616161616161616161616161616161 + cat > "$case_dir/fakebin/gh-axi" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" +case "${1:-} ${2:-}" in + "pr merge") echo "will be added to the merge queue when all requirements are met" ;; +esac +exit 0 +SH + chmod +x "$case_dir/fakebin/gh-axi" + write_github_outcome "$case_dir" OPEN false false main + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/65 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-refusal-quotes-forge: an unproved merge must fail" + assert_grep 'error: > will be added to the merge queue when all requirements are met' \ + "$case_dir/stderr" \ + "github-refusal-quotes-forge: the forge's own explanation was discarded on the refusal" + assert_grep "not this script's verdict" "$case_dir/stderr" \ + "github-refusal-quotes-forge: the forge's text was not marked as the forge's own" + assert_grep 'error: GitHub merge outcome was not successful: state=OPEN, merged=false, isInMergeQueue=false' \ + "$case_dir/stderr" "github-refusal-quotes-forge: the wrapper's own verdict was lost" + # A forge sentence about the merge queue must never stand on its own line, or + # it reads as this script's verdict rather than as quoted forge output. + ! grep -qxF 'will be added to the merge queue when all requirements are met' \ + "$case_dir/stderr" \ + || fail "github-refusal-quotes-forge: forge text was emitted as the wrapper's own line" + assert_no_grep 'will be added to the merge queue' "$case_dir/stdout" \ + "github-refusal-quotes-forge: the forge's unverified report leaked to stdout" + assert_no_grep 'verified: ' "$case_dir/stdout" \ + "github-refusal-quotes-forge: an unproved merge was reported as verified" + pass "fm-pr-merge refuses with the forge's own output quoted apart from its verdict" +} + +test_github_auto_merge_without_queue_refuses_legibly() { + local case_dir rc spelling + for spelling in --auto --auto=true; do + case_dir=$(make_case "github-auto-no-queue${spelling#--auto}") + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 7171717171717171717171717171717171717171 + write_github_outcome "$case_dir" OPEN false false main + : > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/66 \ + -- "$spelling" --merge \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-auto-no-queue: an armed but unlanded auto-merge must still fail" + assert_grep 'state=OPEN, merged=false, isInMergeQueue=false' "$case_dir/stderr" \ + "github-auto-no-queue: refusal did not name the concrete observed state" + assert_grep 'auto-merge was requested and armed for https://github.com/example/repo/pull/66' \ + "$case_dir/stderr" "github-auto-no-queue: the refusal never explained the armed auto-merge" + assert_grep 'nothing is merged or in the merge queue yet' "$case_dir/stderr" \ + "github-auto-no-queue: the refusal left the operator to infer the pending state" + grep -qxF "pr merge 66 --repo example/repo $spelling --merge" "$case_dir/gh-axi.log" \ + || fail "github-auto-no-queue: the attempted merge was changed unexpectedly" + [ "$(wc -l < "$case_dir/gh-axi.log" | tr -d '[:space:]')" = 1 ] \ + || fail "github-auto-no-queue: the wrapper attempted more than one merge" + assert_grep 'pr=https://github.com/example/repo/pull/66' "$case_dir/state/task-x1.meta" \ + "github-auto-no-queue: the attempted merge lost its PR reference" + assert_present "$case_dir/state/task-x1.check.sh" \ + "github-auto-no-queue: the attempted merge did not leave its poll armed" + done + pass "fm-pr-merge explains an armed auto-merge that landed nothing on a queue-less base" +} + +test_github_failed_merge_never_claims_armed_auto_merge() { + local case_dir rc + case_dir=$(make_case github-auto-merge-command-fails) + mkdir -p "$case_dir/wt" + add_gh_mocks_merge_fails "$case_dir" + write_github_outcome "$case_dir" OPEN false false main + : > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/67 -- --auto --merge \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-auto-merge-command-fails: the forge failure must still fail the wrapper" + assert_grep 'error: pr merge failed' "$case_dir/stderr" \ + "github-auto-merge-command-fails: the original forge error was masked" + assert_grep 'state=OPEN, merged=false, isInMergeQueue=false' "$case_dir/stderr" \ + "github-auto-merge-command-fails: refusal did not name the concrete observed state" + assert_no_grep 'armed' "$case_dir/stderr" \ + "github-auto-merge-command-fails: a failed merge command was reported as an armed auto-merge" + assert_grep 'auto-merge was requested for https://github.com/example/repo/pull/67' \ + "$case_dir/stderr" \ + "github-auto-merge-command-fails: the refusal never said auto-merge had only been requested" + assert_no_grep 'verified: ' "$case_dir/stdout" \ + "github-auto-merge-command-fails: a failed merge command was reported as verified" + pass "fm-pr-merge never reports auto-merge as armed when the merge command failed" +} + +test_github_failed_merge_with_queue_flags_never_claims_acceptance() { + local case_dir rc + case_dir=$(make_case github-failed-merge-queue-flags) + mkdir -p "$case_dir/wt" + add_gh_mocks_merge_fails "$case_dir" + write_github_outcome "$case_dir" OPEN false false main + printf 'merge_method=MERGE\n' > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/74 -- --auto --merge \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-failed-merge-queue-flags: the forge failure must still fail the wrapper" + assert_grep 'error: pr merge failed' "$case_dir/stderr" \ + "github-failed-merge-queue-flags: the original forge error was masked" + assert_grep 'state=OPEN, merged=false, isInMergeQueue=false' "$case_dir/stderr" \ + "github-failed-merge-queue-flags: refusal did not name the concrete observed state" + assert_no_grep 'was accepted with the exact flags' "$case_dir/stderr" \ + "github-failed-merge-queue-flags: a failed merge command was reported as an accepted request" + assert_no_grep 'armed' "$case_dir/stderr" \ + "github-failed-merge-queue-flags: a failed merge command was reported as an armed auto-merge" + assert_grep 'base branch main requires the merge queue; retry with:' "$case_dir/stderr" \ + "github-failed-merge-queue-flags: the failed merge command lost its concrete retry guidance" + assert_grep 'task-x1 https://github.com/example/repo/pull/74 -- --auto --merge' "$case_dir/stderr" \ + "github-failed-merge-queue-flags: the retry guidance named no queue flags" + assert_no_grep 'verified: ' "$case_dir/stdout" \ + "github-failed-merge-queue-flags: a failed merge command was reported as verified" + pass "fm-pr-merge claims no acceptance for a failed merge command carrying queue flags" +} + +test_github_accepted_queue_flags_do_not_echo_back_the_same_command() { + local case_dir rc + case_dir=$(make_case github-accepted-queue-flags) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 8181818181818181818181818181818181818181 + write_github_outcome "$case_dir" OPEN false false main + printf 'merge_method=MERGE\n' > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/68 -- --auto --merge \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-accepted-queue-flags: an unproved merge must still fail" + assert_grep 'state=OPEN, merged=false, isInMergeQueue=false' "$case_dir/stderr" \ + "github-accepted-queue-flags: refusal did not name the concrete observed state" + assert_grep 'this run refuses even though the request for https://github.com/example/repo/pull/68 was accepted with the exact flags base branch main requires (--auto --merge)' \ + "$case_dir/stderr" \ + "github-accepted-queue-flags: the refusal did not explain that the right flags were already used" + assert_grep "re-check the pull request's merge queue state" "$case_dir/stderr" \ + "github-accepted-queue-flags: the refusal named no concrete next step" + assert_no_grep 'retry with:' "$case_dir/stderr" \ + "github-accepted-queue-flags: the refusal echoed back the command that just refused" + assert_no_grep 'verified: ' "$case_dir/stdout" \ + "github-accepted-queue-flags: an unproved merge was reported as verified" + pass "fm-pr-merge does not echo back queue flags the caller already used" +} + +test_github_mismatched_queue_flags_still_name_the_retry() { + local case_dir rc + case_dir=$(make_case github-mismatched-queue-flags) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 8282828282828282828282828282828282828282 + write_github_outcome "$case_dir" OPEN false false main + printf 'merge_method=REBASE\n' > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/69 -- --auto --merge \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-mismatched-queue-flags: an unproved merge must still fail" + assert_grep 'base branch main requires the merge queue; retry with:' "$case_dir/stderr" \ + "github-mismatched-queue-flags: a caller method the queue does not use lost its retry guidance" + assert_grep '-- --auto --rebase' "$case_dir/stderr" \ + "github-mismatched-queue-flags: the exact compatible flags were not named" + pass "fm-pr-merge still names retry flags when the caller used a different method" +} + +test_github_unrecognised_queue_method_still_names_the_queue() { + local case_dir rc + case_dir=$(make_case github-unrecognised-queue-method) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 8383838383838383838383838383838383838383 + write_github_outcome "$case_dir" OPEN false false main + printf 'merge_method=FASTFORWARD\n' > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/70 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-unrecognised-queue-method: an unproved merge must fail" + assert_grep 'base branch main requires the merge queue, but its configured merge method (FASTFORWARD) is not one this script recognises' \ + "$case_dir/stderr" \ + "github-unrecognised-queue-method: a readable queue rule produced no queue mention" + assert_no_grep 'retry with:' "$case_dir/stderr" \ + "github-unrecognised-queue-method: retry flags were named for a method nothing recognises" + assert_no_grep '--auto --' "$case_dir/stderr" \ + "github-unrecognised-queue-method: a merge method was guessed for the caller" + pass "fm-pr-merge names the queue requirement even when its method is unrecognised" +} + +test_github_unreadable_queue_rules_are_not_reported_as_no_queue() { + local case_dir rc + case_dir=$(make_case github-unreadable-queue-rules) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 8484848484848484848484848484848484848484 + write_github_outcome "$case_dir" OPEN false false main + cat > "$case_dir/fakebin/gh" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$*" >> "$FM_TEST_GH_LOG" +case "${1:-} ${2:-}" in + "pr view") + case " $* " in + *headRefOid*) printf '%s\n' 8484848484848484848484848484848484848484 ; exit 0 ;; + esac + ;; + "api graphql") + cat "$FM_TEST_GH_OUTCOME" + exit 0 + ;; + api\ *) exit 1 ;; +esac +exit 0 +SH + chmod +x "$case_dir/fakebin/gh" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/71 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-unreadable-queue-rules: an unproved merge must fail" + assert_grep 'the branch rules for base branch main could not be read' "$case_dir/stderr" \ + "github-unreadable-queue-rules: an unreadable rules response read like a queue-less base" + assert_no_grep 'retry with:' "$case_dir/stderr" \ + "github-unreadable-queue-rules: retry flags were named from rules nothing could read" + pass "fm-pr-merge distinguishes unreadable branch rules from a base with no merge queue" +} + +test_github_no_queue_rule_says_nothing_about_a_queue() { + local case_dir rc + case_dir=$(make_case github-no-queue-rule) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 8585858585858585858585858585858585858585 + write_github_outcome "$case_dir" OPEN false false main + : > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/72 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-no-queue-rule: an unproved merge must fail" + assert_grep 'state=OPEN, merged=false, isInMergeQueue=false' "$case_dir/stderr" \ + "github-no-queue-rule: refusal did not name the concrete observed state" + assert_no_grep 'merge queue' "$case_dir/stderr" \ + "github-no-queue-rule: a base with no queue rule was told it requires the merge queue" + pass "fm-pr-merge says nothing about a merge queue when the base branch has no queue rule" +} + +test_github_fallback_view_refusal_says_the_queue_was_unobservable() { + local case_dir ghless_path rc + case_dir=$(make_case github-fallback-unobservable-queue) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 8686868686868686868686868686868686868686 + cat > "$case_dir/fakebin/gh-axi" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" +case "${1:-} ${2:-}" in + "pr merge") printf 'merged:\n number: %s\n status: ok\n' "${3:-}" ;; + "pr view") printf 'pull_request:\n number: %s\n state: open\n' "$3" ;; +esac +exit 0 +SH + chmod +x "$case_dir/fakebin/gh-axi" + rm "$case_dir/fakebin/gh" + ghless_path="$case_dir/path-without-gh" + mirror_path_without "$ghless_path" gh "$case_dir/fakebin" + : > "$case_dir/gh-axi.log" + + set +e + PATH="$ghless_path" run_pr_merge "$case_dir" task-x1 \ + https://github.com/example/repo/pull/73 -- --auto --merge \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-fallback-unobservable-queue: an unproved merge must fail" + assert_grep 'isInMergeQueue=unknown' "$case_dir/stderr" \ + "github-fallback-unobservable-queue: refusal did not name the concrete observed state" + assert_grep 'the merge queue could not be observed for https://github.com/example/repo/pull/73' \ + "$case_dir/stderr" \ + "github-fallback-unobservable-queue: the refusal implied an unqueued PR it could not see" + assert_grep "re-check the pull request's merge queue state" "$case_dir/stderr" \ + "github-fallback-unobservable-queue: the refusal named no concrete next step" + # The lowercase state the fallback view reports must be judged the same way + # the queue-aware read's uppercase enum is, or every explanation is skipped. + assert_grep 'auto-merge was requested and armed for https://github.com/example/repo/pull/73' \ + "$case_dir/stderr" \ + "github-fallback-unobservable-queue: the fallback view's state skipped the auto-merge explanation" + assert_no_grep 'verified: ' "$case_dir/stdout" \ + "github-fallback-unobservable-queue: an unproved merge was reported as verified" + pass "fm-pr-merge says the merge queue was unobservable when only the gh-axi view answered" +} + +test_github_unreadable_outcome_refusal_quotes_the_forge_output() { + local case_dir rc + case_dir=$(make_case github-unreadable-outcome-quotes-forge) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 8787878787878787878787878787878787878787 + cat > "$case_dir/fakebin/gh-axi" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" +case "${1:-} ${2:-}" in + "pr merge") echo "will be added to the merge queue when all requirements are met" ;; + "pr view") exit 1 ;; +esac +exit 0 +SH + chmod +x "$case_dir/fakebin/gh-axi" + add_gh_mock_outcome_read_fails "$case_dir" 8787878787878787878787878787878787878787 + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/74 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-unreadable-outcome-quotes-forge: an unreadable outcome must fail" + assert_grep 'could not read the GitHub pull request outcome after the merge attempt' \ + "$case_dir/stderr" \ + "github-unreadable-outcome-quotes-forge: the unreadable outcome was not reported" + assert_grep 'error: > will be added to the merge queue when all requirements are met' \ + "$case_dir/stderr" \ + "github-unreadable-outcome-quotes-forge: the forge's only evidence was discarded" + ! grep -qxF 'will be added to the merge queue when all requirements are met' \ + "$case_dir/stderr" \ + || fail "github-unreadable-outcome-quotes-forge: forge text was emitted as the wrapper's own line" + assert_no_grep 'verified: ' "$case_dir/stdout" \ + "github-unreadable-outcome-quotes-forge: an unproved merge was reported as verified" + assert_present "$case_dir/state/task-x1.check.sh" \ + "github-unreadable-outcome-quotes-forge: the attempted merge lost its merge poll" + pass "fm-pr-merge quotes the forge output when it cannot read the outcome either" +} + +test_github_failed_gh_read_falls_back_to_gh_axi() { + local case_dir rc + case_dir=$(make_case github-gh-read-falls-back) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 5151515151515151515151515151515151515151 + add_gh_mock_outcome_read_fails "$case_dir" 5151515151515151515151515151515151515151 + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/63 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 0 "$rc" "github-gh-read-falls-back: a merge the gh-axi view proves must succeed" + assert_grep 'pr view 63 --repo example/repo' "$case_dir/gh-axi.log" \ + "github-gh-read-falls-back: the gh-axi view was never consulted after gh's read failed" + assert_grep 'verified: https://github.com/example/repo/pull/63 is merged' \ + "$case_dir/stdout" "github-gh-read-falls-back: the proven merge was not reported" + assert_grep 'pr=https://github.com/example/repo/pull/63' "$case_dir/state/task-x1.meta" \ + "github-gh-read-falls-back: the merged PR was not recorded for teardown" + pass "fm-pr-merge falls back to the gh-axi view when gh's read fails" +} + +test_github_failed_merge_names_an_observed_landed_state() { + local case_dir rc + case_dir=$(make_case github-failed-merge-actually-landed) + mkdir -p "$case_dir/wt" + add_gh_mocks_merge_fails "$case_dir" + write_github_outcome "$case_dir" MERGED true false main + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/64 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-failed-merge-actually-landed: the forge failure must still fail the wrapper" + assert_grep 'error: pr merge failed' "$case_dir/stderr" \ + "github-failed-merge-actually-landed: the original forge error was masked" + assert_grep 'state=MERGED, merged=true, isInMergeQueue=false' "$case_dir/stderr" \ + "github-failed-merge-actually-landed: the observed landed state was never named" + assert_no_grep 'verified: ' "$case_dir/stdout" \ + "github-failed-merge-actually-landed: a failed merge command was reported as verified" + assert_grep 'pr=https://github.com/example/repo/pull/64' "$case_dir/state/task-x1.meta" \ + "github-failed-merge-actually-landed: the landed PR lost its reference" + pass "fm-pr-merge names a landed state hiding behind a failed GitHub merge command" +} + +test_github_without_gh_still_uses_gh_axi_merge() { + local case_dir ghless_path rc + case_dir=$(make_case github-without-gh) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 4141414141414141414141414141414141414141 + rm "$case_dir/fakebin/gh" + ghless_path="$case_dir/path-without-gh" + mirror_path_without "$ghless_path" gh "$case_dir/fakebin" + : > "$case_dir/gh-axi.log" + + set +e + PATH="$ghless_path" run_pr_merge "$case_dir" task-x1 \ + https://github.com/example/repo/pull/60 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 0 "$rc" "github-without-gh: gh-axi can prove a landed merge without gh" + assert_grep 'pr merge 60 --repo example/repo --squash' "$case_dir/gh-axi.log" \ + "github-without-gh: the configured merge abstraction was not invoked" + assert_grep 'pr view 60 --repo example/repo' "$case_dir/gh-axi.log" \ + "github-without-gh: the gh-axi fallback did not verify the landed state" + assert_grep 'verified: https://github.com/example/repo/pull/60 is merged' \ + "$case_dir/stdout" "github-without-gh: the fallback did not report the proven merge" + pass "fm-pr-merge reaches and verifies the gh-axi merge path without gh" +} + +test_github_without_gh_failed_read_keeps_bookkeeping() { + local case_dir ghless_path rc + case_dir=$(make_case github-without-gh-read-fails) + mkdir -p "$case_dir/wt" + cat > "$case_dir/fakebin/gh-axi" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$*" >> "$FM_TEST_GH_AXI_LOG" +case "${1:-} ${2:-}" in + "pr merge") exit 0 ;; + "pr view") exit 1 ;; +esac +exit 0 +SH + chmod +x "$case_dir/fakebin/gh-axi" + ghless_path="$case_dir/path-without-gh" + mirror_path_without "$ghless_path" gh "$case_dir/fakebin" + : > "$case_dir/gh-axi.log" + + set +e + PATH="$ghless_path" run_pr_merge "$case_dir" task-x1 \ + https://github.com/example/repo/pull/61 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-without-gh-read-fails: an unreadable outcome must fail" + assert_grep 'pr merge 61 --repo example/repo --squash' "$case_dir/gh-axi.log" \ + "github-without-gh-read-fails: the merge call did not happen before the failed read" + assert_grep 'could not read the GitHub pull request outcome after the merge attempt' \ + "$case_dir/stderr" "github-without-gh-read-fails: the failed read was not reported" + assert_grep 'pr=https://github.com/example/repo/pull/61' "$case_dir/state/task-x1.meta" \ + "github-without-gh-read-fails: a landed merge lost its PR metadata" + assert_present "$case_dir/state/task-x1.check.sh" \ + "github-without-gh-read-fails: a landed merge lost its merge poll" + pass "fm-pr-merge preserves bookkeeping when gh is absent and the fallback read fails" +} + +test_github_zero_exit_queue_required_refuses_with_exact_retry() { + local case_dir rc + case_dir=$(make_case github-zero-exit-queue-required) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 2121212121212121212121212121212121212121 + write_github_outcome "$case_dir" OPEN false false 'release/2026' + printf 'merge_method=REBASE\n' > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/56 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-zero-exit-queue-required: an unproved merge must fail" + assert_grep 'state=OPEN, merged=false, isInMergeQueue=false' "$case_dir/stderr" \ + "github-zero-exit-queue-required: refusal did not name the concrete observed state" + assert_grep 'base branch release/2026 requires the merge queue' "$case_dir/stderr" \ + "github-zero-exit-queue-required: refusal did not name the queue requirement" + assert_grep '-- --auto --rebase' "$case_dir/stderr" \ + "github-zero-exit-queue-required: refusal did not name the exact compatible flags" + assert_grep 'api --paginate repos/example/repo/rules/branches/release%2F2026' "$case_dir/gh.log" \ + "github-zero-exit-queue-required: queue rules were not read with pagination and encoded branch path" + grep -qxF 'pr merge 56 --repo example/repo --squash' "$case_dir/gh-axi.log" \ + || fail "github-zero-exit-queue-required: the attempted merge was changed unexpectedly" + [ "$(wc -l < "$case_dir/gh-axi.log" | tr -d '[:space:]')" = 1 ] \ + || fail "github-zero-exit-queue-required: the wrapper attempted more than one merge" + assert_no_grep --auto "$case_dir/gh-axi.log" \ + "github-zero-exit-queue-required: queue flags were auto-applied to the attempted merge" + assert_grep 'pr=https://github.com/example/repo/pull/56' "$case_dir/state/task-x1.meta" \ + "github-zero-exit-queue-required: the attempted merge lost its PR reference" + assert_present "$case_dir/state/task-x1.check.sh" \ + "github-zero-exit-queue-required: the attempted merge did not leave its poll armed" + pass "fm-pr-merge reports exact queue retry flags after a zero-exit false success" +} + +test_github_closed_unqueued_outcome_omits_retry_flags() { + local case_dir rc + case_dir=$(make_case github-closed-unqueued) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 2323232323232323232323232323232323232323 + write_github_outcome "$case_dir" CLOSED false false master + printf 'merge_method=MERGE\n' > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/57 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-closed-unqueued: an unproved merge must fail" + assert_grep 'state=CLOSED, merged=false, isInMergeQueue=false' "$case_dir/stderr" \ + "github-closed-unqueued: refusal did not name the concrete observed state" + assert_no_grep 'requires the merge queue' "$case_dir/stderr" \ + "github-closed-unqueued: closed PR received unusable queue guidance" + assert_no_grep '-- --auto --merge' "$case_dir/stderr" \ + "github-closed-unqueued: closed PR received retry flags" + assert_grep 'pr=https://github.com/example/repo/pull/57' "$case_dir/state/task-x1.meta" \ + "github-closed-unqueued: the attempted merge lost its PR reference" + assert_present "$case_dir/state/task-x1.check.sh" \ + "github-closed-unqueued: the attempted merge did not leave its poll armed" + pass "fm-pr-merge omits merge-queue retry guidance for a closed GitHub PR" +} + +test_github_queued_outcome_is_verified() { + local case_dir rc + case_dir=$(make_case github-verified-queued) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 3030303030303030303030303030303030303030 + write_github_outcome "$case_dir" OPEN false true master + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/53 -- --auto --merge \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 0 "$rc" "github-verified-queued: a queued PR should succeed" + assert_grep 'verified: https://github.com/example/repo/pull/53 is queued' \ + "$case_dir/stdout" "github-verified-queued: success was not reported as queued" + assert_no_grep 'merged:' "$case_dir/stdout" \ + "github-verified-queued: the forge CLI's unverified merged report leaked through" + assert_grep 'pr=https://github.com/example/repo/pull/53' "$case_dir/state/task-x1.meta" \ + "github-verified-queued: the queued PR was not recorded for teardown" + pass "fm-pr-merge accepts and accurately reports a GitHub merge-queue entry" +} + +test_github_queue_required_refusal_names_retry_flags() { + local case_dir rc + case_dir=$(make_case github-queue-required) + mkdir -p "$case_dir/wt" + add_gh_mocks_merge_fails "$case_dir" + write_github_outcome "$case_dir" OPEN false false master + printf 'merge_method=MERGE\n' > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/54 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-queue-required: an incompatible direct merge must fail" + assert_grep 'error: pr merge failed' "$case_dir/stderr" \ + "github-queue-required: the original forge failure was not preserved" + assert_grep 'base branch master requires the merge queue' "$case_dir/stderr" \ + "github-queue-required: refusal did not name the queue requirement" + grep -F -- '-- --auto --merge' "$case_dir/stderr" >/dev/null \ + || fail "github-queue-required: refusal did not name the exact compatible flags" + grep -qxF 'pr merge 54 --repo example/repo --squash' "$case_dir/gh-axi.log" \ + || fail "github-queue-required: the wrapper silently changed the attempted merge semantics" + assert_present "$case_dir/state/task-x1.check.sh" \ + "github-queue-required: the failed forge call did not leave the merge poll armed" + pass "fm-pr-merge explains how to retry with the required GitHub merge queue method" +} + +test_github_agreeing_queue_rules_keep_retry_guidance() { + local case_dir rc + case_dir=$(make_case github-agreeing-queue-rules) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 2424242424242424242424242424242424242424 + write_github_outcome "$case_dir" OPEN false false main + printf 'merge_method=REBASE\nmerge_method=REBASE\n' > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/58 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-agreeing-queue-rules: an unproved merge must fail" + assert_grep 'base branch main requires the merge queue' "$case_dir/stderr" \ + "github-agreeing-queue-rules: refusal did not name the queue requirement" + assert_grep '-- --auto --rebase' "$case_dir/stderr" \ + "github-agreeing-queue-rules: agreeing rules omitted exact retry flags" + assert_no_grep 'exact retry flags are ambiguous' "$case_dir/stderr" \ + "github-agreeing-queue-rules: agreeing rules were reported as ambiguous" + pass "fm-pr-merge aggregates agreeing merge-queue rules" +} + +test_github_conflicting_queue_rules_report_ambiguity() { + local case_dir rc + case_dir=$(make_case github-conflicting-queue-rules) + mkdir -p "$case_dir/wt" + add_gh_mocks "$case_dir" 2525252525252525252525252525252525252525 + write_github_outcome "$case_dir" OPEN false false main + printf 'merge_method=MERGE\nmerge_method=SQUASH\nmerge_method=SQUASH\n' \ + > "$case_dir/github-rules" + : > "$case_dir/gh-axi.log" + : > "$case_dir/gh.log" + + set +e + run_pr_merge "$case_dir" task-x1 https://github.com/example/repo/pull/59 \ + > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "github-conflicting-queue-rules: an unproved merge must fail" + assert_grep 'base branch main has conflicting merge queue methods (MERGE, SQUASH)' \ + "$case_dir/stderr" \ + "github-conflicting-queue-rules: conflicting methods were not named" + assert_no_grep '-- --auto --merge' "$case_dir/stderr" \ + "github-conflicting-queue-rules: an exact retry method was guessed" + assert_no_grep '-- --auto --squash' "$case_dir/stderr" \ + "github-conflicting-queue-rules: an exact retry method was guessed" + assert_no_grep 'SQUASH, SQUASH' "$case_dir/stderr" \ + "github-conflicting-queue-rules: a repeated queue method was named twice" + pass "fm-pr-merge reports ambiguity for conflicting merge-queue rules" +} + test_extra_merge_args_forwarded() { local case_dir rc case_dir=$(make_case extra-args) @@ -834,7 +1783,6 @@ test_github_still_forwards_sha_arg() { pass "fm-pr-merge leaves GitHub extra-arg handling unchanged, including --sha" } - # --- durable merge outcome --------------------------------------------------- # A merge that lands must leave a record outside the merging agent's memory. # bin/fm-merge-outcome-lib.sh owns where that record goes; these cases pin the @@ -1004,6 +1952,7 @@ test_queued_github_merge_leaves_the_poll_armed() { url=https://github.com/example/repo/pull/66 case_dir=$(make_home_case queued-github-merge) add_gh_mocks "$case_dir" 9999999999999999999999999999999999999999 + write_github_outcome "$case_dir" OPEN false true main : >"$case_dir/gh-axi.log" FM_TEST_GH_MERGE_STATE=open FM_TEST_HOME="$case_dir/home" \ @@ -1123,8 +2072,34 @@ test_secondmate_without_parent_binding_is_loud() { pass "a secondmate home that cannot report upward says so instead of merging in silence" } -test_records_pr_and_head_before_merging +test_github_zero_exit_queue_required_refuses_with_exact_retry +test_github_closed_unqueued_outcome_omits_retry_flags +test_github_agreeing_queue_rules_keep_retry_guidance +test_github_conflicting_queue_rules_report_ambiguity +test_verified_merge_records_pr_and_head +test_pr_metadata_is_recorded_before_the_forge_call test_merge_failure_propagates_after_recording +test_github_open_unqueued_outcome_refuses +test_github_unreadable_outcome_keeps_pr_bookkeeping +test_github_refusal_quotes_the_forge_output +test_github_unreadable_outcome_refusal_quotes_the_forge_output +test_github_accepted_queue_flags_do_not_echo_back_the_same_command +test_github_mismatched_queue_flags_still_name_the_retry +test_github_unrecognised_queue_method_still_names_the_queue +test_github_unreadable_queue_rules_are_not_reported_as_no_queue +test_github_no_queue_rule_says_nothing_about_a_queue +test_github_fallback_view_refusal_says_the_queue_was_unobservable +test_github_auto_merge_without_queue_refuses_legibly +test_github_failed_merge_never_claims_armed_auto_merge +test_github_failed_merge_with_queue_flags_never_claims_acceptance +test_github_failed_gh_read_falls_back_to_gh_axi +test_github_failed_merge_names_an_observed_landed_state +test_github_without_gh_still_uses_gh_axi_merge +test_github_without_gh_failed_read_keeps_bookkeeping +test_github_merged_outcome_is_verified +test_github_verified_merge_requires_poll_recording +test_github_queued_outcome_is_verified +test_github_queue_required_refusal_names_retry_flags test_extra_merge_args_forwarded test_missing_meta_refuses_before_merge test_malformed_url_refuses_before_merge From 4f89f5b5e235469d32c037b7792d6dba5bdc272d Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Thu, 27 Aug 2026 13:12:18 -0700 Subject: [PATCH 03/25] fix(pi): prevent duplicate captain outcome reports (#3184) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(pi): stop reporting one merge to the captain twice The supervision branch's captain-outcome note told main, unconditionally, that the note "is not your own earlier output" and to relay it now. When main had already reported the same event, that assertion was false and the order turned the correct response - saying nothing new - into a mechanical re-report, so the captain saw one merge reported twice in 16 seconds. Two independent changes, both needed: - The relay instruction is now conditional. It still names itself as a supervision outcome so main cannot mistake it for its own earlier answer (the silent loss that instruction exists to prevent), and it now lets main stay quiet about an outcome it has already given the captain. - The merge case is closed at its source rather than left to that judgment. One merge reaches a home on two independent paths by design - main's own permanently main-owned merge poll, and the branch's task-local status wake - and main's captain-facing text only reaches the branch's mirror at main's turn end, so the branch can escalate before it could possibly see the captain was already told. bin/fm-pr-merge-notified.sh answers that question from bin/fm-pr-lib.sh's canonical merge-notification marker, so the answer holds regardless of mirror timing. A captain outcome naming an already-published merge is delivered as the ordinary rendered note instead of opening a follow-up turn: still appended, still visible, still recorded with the verdict the branch decided, minus the wasted turn. Any error, timeout, or unreadable state relays the outcome. A duplicate announces itself; a lost outcome does not. Regression coverage drives the real delivery path in both directions: a new outcome must still reach the captain in exactly one follow-up turn even beside an unrelated published merge, and an already-published merge must open no second turn while a different PR in the same task still does. The merge path's real producer and this new consumer are exercised end to end in tests/fm-pr-merge.test.sh. Pi-only by construction: the delivery path lives in .pi/extensions, so no other harness loads it, and the new script only reads existing markers. * no-mistakes(review): Document accepted latest-marker suppression residual * no-mistakes(review): Recheck ownership before merge outcome delivery * no-mistakes(document): Document merge-outcome suppression exception * refactor(pi): drop the source-level merge suppression, keep the envelope fix The captain reviewed this branch and judged the source-level duplicate suppression overly complicated for the problem it solved, and asked for the change to be reduced to the envelope wording alone. Remove the mergeIntoMain downgrade path, bin/fm-pr-merge-notified.sh, and every test and document that existed only for it. What remains is the conditional captain-outcome instruction: main is told to stay quiet about an outcome it has already reported and to relay anything else, which covers the duplicate without a second mechanism. The silent-loss protection is untouched - the note is still typed, self-describing, and delivered as one invisible follow-up turn - and the behavioral tests still assert that, now requiring both halves of the conditional instruction. * no-mistakes(ci): Clarified in code comments and owned documentation that this is intentionally an M1-only, model-facing conditional relay fix—not source-level suppression—addressing Greptile’s mistaken scope expectation without changing runtime behavior. Net diff remains 3 files and 27 insertions. Verified with fm-pi-branch-extension tests, fm-lint, doc audience check, and git diff --check; all passed * no-mistakes(ci): Strengthened the runtime delivery test to verify the captain outcome retains its required self-description and outcome text. Verified with `bash tests/fm-pi-branch-extension.test.sh`, `bin/fm-lint.sh`, `bin/fm-doc-audience-check.sh`, and `git diff --check`; all passed. The outer pipeline can now commit and attest the new head --- .pi/extensions/fm-branch-supervision.ts | 15 +++++++++++++-- docs/pi-supervision-branch.md | 2 ++ tests/fm-pi-branch-extension.test.sh | 18 ++++++++++++++---- 3 files changed, 29 insertions(+), 6 deletions(-) diff --git a/.pi/extensions/fm-branch-supervision.ts b/.pi/extensions/fm-branch-supervision.ts index 093023fd00b..8a56bacd9b1 100644 --- a/.pi/extensions/fm-branch-supervision.ts +++ b/.pi/extensions/fm-branch-supervision.ts @@ -129,10 +129,21 @@ const MIRROR_MESSAGE_CAP = 4000; const MERGE_NOTE_BOAT = "⛵"; // Carried inside the captain note's own text because that text is the only // part of a custom message Pi gives the model (see mergeIntoMain). +// +// The relay order is CONDITIONAL on purpose. This M1 fix is deliberately a +// model-facing instruction only: mergeIntoMain must keep opening the follow-up +// turn so a genuinely new outcome cannot be suppressed before main sees it. +// The note still needs its self-description to stop main from mistaking an +// incoming outcome for its own earlier answer and silently losing the outcome. +// But an unconditional relay order is false whenever main was separately woken +// for the same event, and turns the correct response - saying nothing new - +// into a mechanical re-report. The conditional wording lets main stay quiet +// when appropriate; source-level identity suppression is outside this M1 fix. const CAPTAIN_OUTCOME_INSTRUCTION = "This is a supervision outcome delivered automatically by the supervision branch. " + - "It was not typed by the captain and it is not your own earlier output. " + - "Relay only this outcome to the captain now, in one short message, in captain outcome language. " + + "It was not typed by the captain. " + + "If you have already reported this outcome to the captain earlier in this conversation, do not report it again. " + + "Otherwise relay only this outcome to the captain now, in one short message, in captain outcome language. " + "Do not restate or repeat any earlier answer."; type MirrorItem = { tag: "captain" | "main"; text: string }; type MirrorCursor = { file: string; index: number }; diff --git a/docs/pi-supervision-branch.md b/docs/pi-supervision-branch.md index f1f04eb2122..963c7cfe2d8 100644 --- a/docs/pi-supervision-branch.md +++ b/docs/pi-supervision-branch.md @@ -55,6 +55,8 @@ Stage two is the branch's verdict on each handled event, reported through its `f The follow-up turn a `captain` verdict opens is itself the captain-visible outcome, so its merge note is delivered silently and never printed or rendered in Pi. Because Pi gives the model only a custom message's `content`, that silent note normally carries both a relay instruction and the `branch-outcome` operational kind owned by `bin/fm-operational-input.sh` inside its own text. This self-description lets main distinguish a new supervision outcome from its own earlier captain-facing answer; without it, main can mistake the outcome for that answer and re-emit the stale answer instead of relaying the outcome. +The relay instruction itself is conditional: it tells main to stay quiet about an outcome it has already given the captain in this conversation, and to relay anything else rather than restate an earlier answer. +An unconditional order made main re-report an outcome the captain already had; this M1 fix deliberately conditions main's relay without adding source-level suppression, because more than one actor can be woken for the same event. If envelope encoding fails, the note degrades to the same relay instruction as plain text rather than losing the outcome or opening another turn. A no-change heartbeat outcome explicitly reported with `task=fleet` and `silent=true` is also delivered silently with no rendered note, while every other `routine` outcome stays rendered with its sailboat prefix. The verdict criteria in the branch prompt mirror the captain-etiquette escalation list; doubt escalates. diff --git a/tests/fm-pi-branch-extension.test.sh b/tests/fm-pi-branch-extension.test.sh index bf95b4587db..d90f760c575 100644 --- a/tests/fm-pi-branch-extension.test.sh +++ b/tests/fm-pi-branch-extension.test.sh @@ -774,11 +774,20 @@ EOF body=$(./bin/fm-operational-input.sh body < "$home/state/delivered-captain-note") \ || fail "captain outcome envelope carries no readable body" case "$body" in - *"task-9: PR https://example.com/pr/9"*) ;; - *) fail "captain outcome body lost the outcome itself: $body" ;; + *"This is a supervision outcome delivered automatically by the supervision branch."*"It was not typed by the captain."*"task-9: PR https://example.com/pr/9"*) ;; + *) fail "captain outcome body lost its self-description or the outcome itself: $body" ;; esac + # Both halves of the delivered instruction matter and they pull against each + # other: main must be allowed to stay quiet about an outcome it has already + # given the captain, and must still be told to relay everything else instead + # of re-emitting its own last answer. An instruction carrying only one half + # reintroduces either the duplicate or the silent loss. case "$body" in - *"Relay only this outcome"*"Do not restate or repeat any earlier answer"*) ;; + *"already reported this outcome to the captain"*"do not report it again"*) ;; + *) fail "captain outcome body never lets main deduplicate what it already said: $body" ;; + esac + case "$body" in + *"relay only this outcome to the captain now"*"Do not restate or repeat any earlier answer"*) ;; *) fail "captain outcome body never tells main to relay it instead of repeating: $body" ;; esac # The routine note is rendered in the TUI, and its renderer reads the glyph off @@ -825,7 +834,8 @@ if (delivered.options.triggerTurn !== true || delivered.options.deliverAs !== "f if (delivered.message.content.includes("FIRSTMATE_OP:")) { throw new Error(`fallback unexpectedly carried an envelope: ${delivered.message.content}`); } -if (!delivered.message.content.includes("Relay only this outcome") || +if (!delivered.message.content.includes("relay only this outcome to the captain now") || + !delivered.message.content.includes("do not report it again") || !delivered.message.content.includes("Do not restate or repeat any earlier answer") || !delivered.message.content.includes("task-fallback: PR https://example.com/pr/fallback is ready")) { throw new Error(`fallback lost its instruction or outcome: ${delivered.message.content}`); From bca584a840011079a121fe965a222eb5a3578408 Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Thu, 27 Aug 2026 15:17:54 -0700 Subject: [PATCH 04/25] fix(bin): prioritize active pipeline-owned crew runs (#3194) * fix(bin): bind the live pipeline-owned run instead of a superseded failed row fm-crew-state.sh bound a superseded FAILED no-mistakes run to a task instead of the LIVE replacement run: the live run's pipeline-owned lane head is not a git object in the task worktree, so head-equality attribution rejected it and the coarse runs-list fallback silently continued past the RUNNING row onto an older failed row whose head equalled the stale worktree HEAD. The home summary then flipped invalid and Bearings hid the home's live work (F10). Attribution precedence now follows the daemon's own identity: - An ACTIVE run for the task's branch binds without head equality while branch_sync.state is pipeline_owned (fm_nm_run_is_pipeline_owned_active); the pipeline owning the branch is itself the attribution. - A genuinely failed run with no later run on the branch still reports failed through the unchanged head-equality path - real failures are not hidden. - In the coarse runs scan, an unresolvable head is unknown attribution and stops the scan (fm_nm_head_resolvable) instead of falling through to an older row; a resolvable-but-mismatched head keeps the historical reused-branch skip. The exemption never applies to a terminal run and requires pipeline_owned specifically, both pinned by negative-control tests. Fixture shape verified against the live incident run's real axi status output. * no-mistakes(document): Updated run-attribution documentation ownership --- AGENTS.md | 2 +- bin/fm-crew-state.sh | 34 +++++---- bin/fm-nm-run-lib.sh | 59 ++++++++++++++- bin/fm-teardown.sh | 4 +- docs/architecture.md | 6 +- docs/configuration.md | 2 +- docs/scripts.md | 2 +- tests/fm-crew-state.test.sh | 145 ++++++++++++++++++++++++++++++++++++ 8 files changed, 226 insertions(+), 28 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 89b40f466c8..125f6b1cb09 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -358,7 +358,7 @@ Send the same worker one exact decision naming the decision key, step, action, a Require the matching `resolved` event, forbid `--yes`, and require the worker to process every synchronous return until completion or a genuinely new escalation. Resume fleet supervision immediately after the decision lands. -Judge validation by the current-code-matched run step through `bin/fm-crew-state.sh`, not by shell liveness or the last status event. +Judge validation by the currently attributed run step through `bin/fm-crew-state.sh`, not by shell liveness or the last status event. Running, fixing, or CI states remain working; parked approval or fix-review states require the worker to follow the active gate help; passed or checks-passed is done; failed or cancelled is failed. A worker hand-editing, committing, aborting, or restarting during an active validation run duplicates pipeline ownership outside the supersession sequence above; steer it back to the gate response flow. The worker reports the PR when CI first becomes green rather than waiting for merge monitoring to finish. diff --git a/bin/fm-crew-state.sh b/bin/fm-crew-state.sh index 6d6a1d2906f..1687ad51d74 100755 --- a/bin/fm-crew-state.sh +++ b/bin/fm-crew-state.sh @@ -8,9 +8,8 @@ # or blocked and the crew resumes (responds to the gate, the pipeline fixes, it # re-validates), the log's last line stays stale. This helper never infers the # current state from a tail of the log: it reads the authoritative source (a -# no-mistakes run-step attributed to this crew's branch and current code -# identity, else the pane busy-signature) and reconciles the possibly-stale log -# against it. +# no-mistakes run-step attributed under bin/fm-nm-run-lib.sh's contract, else +# the pane busy-signature) and reconciles the possibly-stale log against it. # # The determinism lives entirely here - only run-step / pane / log reads plus # fixed mapping logic, no heuristics and no LLM. Output is one stable, parseable, @@ -27,14 +26,8 @@ # to the routed status log; dead/missing report the remote verdict; an # unreachable or unreadable remote reports unknown-remote, never a false # gone/dead. -# 2. Matching no-mistakes run for this crew's branch AND current code identity, -# active or terminal (from `axi status`, or the coarse `no-mistakes runs` -# fallback)? Branch name alone is not enough: a historical run on a reused -# branch whose head was rewritten or diverged must not be attributed. -# A run matches when its head equals the worktree HEAD, or the worktree HEAD -# is an ancestor of the run head (pipeline fix commits advanced the run on -# the same line of history). Local work that advanced past the run head, or -# diverged from it, invalidates attribution. +# 2. Attribute an active or terminal no-mistakes run under the branch, head, +# pipeline-custody, and newest-first rules owned by bin/fm-nm-run-lib.sh. # The run-step is AUTHORITATIVE: running/fixing -> working, ci -> working, # awaiting_approval/fix_review -> parked (with gate findings), terminal # passed/checks-passed -> done, failed/cancelled -> failed. EXCEPT: while @@ -219,7 +212,7 @@ crew_busy_verdict() { # # --- no-mistakes run lookup (authoritative when a run matches this branch) -- # trim, strip_quotes, the bounded nm_run call, nm_field's TOON parse, and the -# branch+head attribution rule below are thin wrappers over the ONE owner in +# attribution helpers below are thin wrappers over the ONE owner in # bin/fm-nm-run-lib.sh, shared with fm-teardown.sh's pre-teardown run abort. trim() { fm_nm_trim "$@"; } @@ -401,6 +394,10 @@ nm_runs_status_for_branch() { # # Same code-identity rule as axi status: skip a same-branch row whose # short-sha does not match this worktree (rewritten or advanced tip). if ! nm_coarse_head_matches_worktree "$sha"; then + # An UNRESOLVABLE head is unknown attribution, not a proven + # mismatch. Stop instead of surfacing an older, superseded row; + # the caller's pane/log fallback can answer without misattribution. + fm_nm_head_resolvable "$WT" "$sha" || return 0 continue fi printf '%s' "$st" @@ -443,12 +440,17 @@ if [ "$KIND" = ship ] && [ -n "$CREW_BRANCH" ] && command -v no-mistakes >/dev/n RUN_OUT=$(nm_run axi status) if [ -n "$RUN_OUT" ]; then run_branch=$(strip_quotes "$(nm_field branch)") - if [ -n "$run_branch" ] && [ "$run_branch" = "$CREW_BRANCH" ] && nm_run_head_matches_worktree; then + # Head equality, or the pipeline-owned-active exemption: while the + # pipeline owns this branch, the daemon's own branch attribution is + # authoritative and the lane head need not be a git object here + # (fm_nm_run_is_pipeline_owned_active in bin/fm-nm-run-lib.sh). + if [ -n "$run_branch" ] && [ "$run_branch" = "$CREW_BRANCH" ] \ + && { nm_run_head_matches_worktree || fm_nm_run_is_pipeline_owned_active "$RUN_OUT"; }; then HAVE_RUN=1 else - # The active-or-most-recent run is for another branch, or same branch with - # a rewritten/diverged head (the CLI is alive and answered; only the - # attribution missed) - try the coarse fallback. + # The active-or-most-recent run is for another branch, or its same-branch + # attribution failed (the CLI is alive and answered) - try the coarse + # fallback. # Deliberately nested inside `[ -n "$RUN_OUT" ]`: an empty/timed-out # primary call means the CLI itself did not respond, so retrying it # immediately with a second bounded call would just double the wait diff --git a/bin/fm-nm-run-lib.sh b/bin/fm-nm-run-lib.sh index 7c210c23f58..533cbeee54f 100644 --- a/bin/fm-nm-run-lib.sh +++ b/bin/fm-nm-run-lib.sh @@ -1,10 +1,11 @@ #!/usr/bin/env bash # Shared no-mistakes axi run attribution primitives. # -# ONE owner for the branch+code-identity matching rule that decides whether a -# no-mistakes run belongs to a given worktree, used by fm-crew-state.sh -# (read-only current-state reporting) and fm-teardown.sh (pre-teardown run -# abort, see its "Fix 1" header comment). Getting this wrong in either +# ONE owner for the no-mistakes run-attribution primitives used by +# fm-crew-state.sh (read-only current-state reporting) and fm-teardown.sh +# (pre-teardown run abort, see its "Fix 1" header comment). Teardown uses only +# strict branch-and-head identity; crew-state additionally permits the active +# pipeline-owned exemption defined below. Getting this wrong in either # direction is unsafe: a false negative hides a genuinely parked run, and a # false positive lets teardown act on a run it does not own. # @@ -63,6 +64,8 @@ fm_nm_field() { # # the same history advanced the run tip past local HEAD) # - run head is a strict ancestor of worktree HEAD, or diverged: no match # (local work advanced outside the run, or the branch tip was rewritten) +# fm_nm_run_is_pipeline_owned_active below carries the one exemption: a live +# run whose pipeline currently owns the branch binds without head equality. fm_nm_head_matches_worktree() { # local wt=$1 run_head=$2 local_full run_full [ -n "$run_head" ] || return 1 @@ -71,3 +74,51 @@ fm_nm_head_matches_worktree() { # [ "$run_full" = "$local_full" ] && return 0 git -C "$wt" merge-base --is-ancestor "$local_full" "$run_full" 2>/dev/null } + +# 0 if head $2 resolves to a commit object in worktree $1 at all. This +# distinguishes a PROVEN mismatch (resolvable but not current: a historical or +# diverged head fm_nm_head_matches_worktree correctly rejects) from UNKNOWN +# attribution (unresolvable: e.g. a pipeline-owned lane head that never +# reached this worktree). A caller scanning run rows newest-first must stop on +# unknown attribution rather than surface an older, superseded run. +fm_nm_head_resolvable() { # + [ -n "$2" ] || return 1 + git -C "$1" rev-parse --verify --quiet "$2^{commit}" >/dev/null 2>&1 +} + +# branch_sync.state from captured `axi status` TOON $1: the scalar directly +# under the top-level `branch_sync:` block. The first `state:` inside the +# block is the direct child (the nested local/pipeline/target/remote +# sub-blocks carry no `state:` key). Empty when the block is absent: no run +# on the current branch, another branch's run, or a CLI without branch sync. +fm_nm_branch_sync_state() { # + local s + s=$(printf '%s\n' "$1" \ + | sed -n '/^[[:space:]]*branch_sync:[[:space:]]*$/,/^[^[:space:]][^:]*:/s/^[[:space:]]\{1,\}state:[[:space:]]*\(.*\)/\1/p' \ + | head -1) + fm_nm_strip_quotes "$s" +} + +# 0 if the run in captured `axi status` TOON $1 is still in flight: no +# terminal outcome and no terminal status. +fm_nm_run_is_active() { # + local status outcome + status=$(fm_nm_strip_quotes "$(fm_nm_field "$1" status)") + outcome=$(fm_nm_strip_quotes "$(fm_nm_field "$1" outcome)") + [ -z "$outcome" ] || return 1 + case "$status" in completed|failed|cancelled) return 1 ;; esac +} + +# The one exemption to the head rule above: while the pipeline OWNS the branch +# (branch_sync.state=pipeline_owned), the daemon's own branch attribution IS +# the attribution for an ACTIVE run, and +# head equality must not be required - the pipeline's lane head is routinely +# not a git object in the task worktree (rebase and fix commits that were +# never pushed back), so the head rule rejects exactly the run that is most +# current. The exemption never applies to a terminal run: a terminal run has +# released the branch, and binding one by branch name alone is the historical +# reused-branch misattribution the head rule exists to prevent. +fm_nm_run_is_pipeline_owned_active() { # + [ "$(fm_nm_branch_sync_state "$1")" = pipeline_owned ] || return 1 + fm_nm_run_is_active "$1" +} diff --git a/bin/fm-teardown.sh b/bin/fm-teardown.sh index eaa433746c2..7673d008240 100755 --- a/bin/fm-teardown.sh +++ b/bin/fm-teardown.sh @@ -108,8 +108,8 @@ # crew's worktree, so they are not orphaned by removing the worktree. # conclude_task_no_mistakes_run attributes the active-or-most-recent run to # THIS task only when its branch AND code identity (bin/fm-nm-run-lib.sh's -# fm_nm_head_matches_worktree, the same rule bin/fm-crew-state.sh uses) both -# match this worktree, then runs `no-mistakes axi abort --run ` for +# strict fm_nm_head_matches_worktree rule) both match this worktree, then +# runs `no-mistakes axi abort --run ` for # that verified run instance. A run already terminal # (an outcome is set) or not parked at a gate is left untouched. Idempotent: # an already-aborted run reads back terminal and is skipped on retry. diff --git a/docs/architecture.md b/docs/architecture.md index 0b18c16f0dc..15b1cb10d7d 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -31,7 +31,7 @@ After successful outcome publication, the watcher immediately delivers the emitt The retirement receipt makes poll cleanup safely retryable across restarts: fixed-path recovery revalidates the same evidence, removes the runnable check first, removes its registration and data sidecars, removes the receipt last, and preserves task metadata including `pr=` and `pr_head=`. A concurrent replacement remains armed, every non-merged or invalid observation remains unchanged, and retirement never performs task or persistent-secondmate cleanup. `bin/fm-pr-lib.sh` owns the notification-marker and retirement-receipt formats plus their strict identity mechanics, [`bin/fm-merge-outcome-lib.sh`](../bin/fm-merge-outcome-lib.sh) owns role-routed publication, the local durable row, and marker ordering, and `bin/fm-watch.sh` owns immediate poll-result delivery and retirement. -No-verb wakes, such as `working:` notes and bare turn-ended signals, are benign only when `bin/fm-crew-state.sh` reports positive evidence that the crew is still working: an actively running no-mistakes step attributed to that crew's current code, or an exact busy verdict from the semantic busy-state contract. +No-verb wakes, such as `working:` notes and bare turn-ended signals, are benign only when `bin/fm-crew-state.sh` reports positive evidence that the crew is still working: a currently attributed active no-mistakes step, or an exact busy verdict from the semantic busy-state contract. A `kind=secondmate` task's status signal is the parent-directed reply stream and is never absorbed as provably working; only its bare turn-ended signal retains the ordinary absorb rule. A crew that declares `paused:` for a known external wait, or carries a verified `captain-held` transfer, is separately absorbed while idle and re-surfaced only on the longer pause cadence, rather than being treated as a possible wedge. For an ordinary crew that has stopped, the normal-mode watcher first surfaces one stale wake, then applies that same cadence to an unchanged `paused:` or durable `captain-held` endpoint only when the backend confidently reports its agent dead. @@ -56,8 +56,8 @@ The explicit resolution is written by the actor that answers, not the busy worke This home's answerer close, pending-reply escalation close, and captain-held transfer use the provenance-guarded append owned by `bin/fm-wake-lib.sh`, so they advance the watcher marker only across their own bytes when all earlier bytes were already announced; pending or interleaved foreign bytes fail toward an ordinary wake. A turn-ended-only queue row omits its historical status annotation when that status file exactly matches the same seen marker. Any direct or remaining historical annotation prints every status line unread at the presentation cursor instead of replaying only the latest line. -`bin/fm-crew-state.sh ` is the cheap current-state read for an actionable heartbeat review: it attributes a no-mistakes run, active or terminal, only when it matches the crew's branch and current code identity, then keeps that run-step authoritative even if the pane has closed. -The script header owns the exact run-head ancestry rules. +`bin/fm-crew-state.sh ` is the cheap current-state read for an actionable heartbeat review: it attributes an active or terminal no-mistakes run under the shared run-attribution contract, then keeps that run-step authoritative even if the pane has closed. +[`bin/fm-nm-run-lib.sh`](../bin/fm-nm-run-lib.sh)'s header owns the exact branch, head, pipeline-custody, and newest-first attribution rules. During no-mistakes' `ci` monitor phase, it also reads the ci step log tail because `axi status` reports both "still waiting on checks" and "checks green, waiting on merge" as `ci,running`. The most recent recognized ci log marker wins, so checks-green monitoring reports done while a later re-arm, failed-check, or issue marker returns the crew to working. Only when no matching run exists does it consult semantic busy state; exact busy reports working, exact idle permits fallback to a status-log event whose verb maps to a recognized run-state, and unknown or a dead pane stays unknown instead of trusting a stale log. diff --git a/docs/configuration.md b/docs/configuration.md index 9df7bb77373..a06dd260f7f 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -719,7 +719,7 @@ FM_WHEN_OUTPUT_TAIL_BYTES=8192 # bound on the command-output tail insid FM_CODEX_WATCH_CHECKPOINT=180 # seconds per foreground watcher checkpoint in Codex primary supervision FM_CREW_STATE_NM_TIMEOUT=10 # seconds allowed per no-mistakes query inside fm-crew-state.sh FM_TEARDOWN_NM_TIMEOUT=10 # seconds allowed per no-mistakes query or abort inside fm-teardown.sh -FM_CREW_STATE_RUNS_LIMIT=200 # recent no-mistakes run rows scanned when axi status cannot be attributed to the current code +FM_CREW_STATE_RUNS_LIMIT=200 # recent no-mistakes run rows scanned when axi status cannot be attributed directly FM_CREW_STATE_BIN=bin/fm-crew-state.sh # test override for the current-state reader used by working/paused watcher triage FMX_PAIRING_TOKEN= # Relay pairing token; .env opt-in authorizes replies and eligible lifecycle actions FMX_RELAY_URL=https://myfirstmate.io # optional Relay endpoint override, mainly for local relay development diff --git a/docs/scripts.md b/docs/scripts.md index e45d23c86d2..68945147a56 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -82,7 +82,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-supervisor-target-lib.sh` | Resolve the shared supervisor target and backend for the daemon and launcher | | `fm-supervise-daemon.sh` | Presence-gated away-mode sub-supervisor: self-handle routine wakes, guard injection by the detected primary harness, escalate batched digests, alert on failed delivery | | `fm-crew-state.sh` | Print one deterministic current-state line for a crew | -| `fm-nm-run-lib.sh` | Shared branch-and-code-identity attribution for no-mistakes runs | +| `fm-nm-run-lib.sh` | Single owner of shared no-mistakes run-attribution primitives and rules | | `fm-tangle-lib.sh` | Shared default-branch resolution and primary-checkout tangle classification | | `fm-timeout-lib.sh` | Single owner of hard-bounded command execution and its fallback watchdog | | `fm-timing-lib.sh` | Single owner of the deferred network stage's per-step elapsed-time records, inert unless a run asks for them | diff --git a/tests/fm-crew-state.test.sh b/tests/fm-crew-state.test.sh index 602b3e5cfc3..a284cbe8eb6 100755 --- a/tests/fm-crew-state.test.sh +++ b/tests/fm-crew-state.test.sh @@ -1389,6 +1389,146 @@ test_local_advanced_past_run_head_invalidates() { pass "local work advanced past run head invalidates attribution" } +# --- Run-attribution precedence for pipeline-owned lane heads ---------------- +# A live run whose pipeline OWNS the branch (branch_sync.state=pipeline_owned) +# can report a lane head that is not a git object in the task worktree. +# Every fixture head is deliberately unresolvable so only the top-level +# branch_sync exemption - never an accidental nested-field match - attributes +# the run. +run_running_pipeline_owned() { # [] + cat </dev/null + fm_write_meta "$d/state/feat-f10.meta" "window=fm:fm-feat-f10" "worktree=$d/wt" "kind=ship" + FM_FAKE_AXI_STATUS="$(run_running_pipeline_owned fm/feat-f10 f0f0f0f0)" + FM_FAKE_RUNS_LIST="$(cat < working" + assert_contains "$out" "source: run-step" "pipeline-owned live run -> run-step source" + assert_not_contains "$out" "state: failed" "superseded failed row must not surface over the live run" + pass "pipeline-owned active run binds without head equality and beats the failed row" +} + +# T1 direction 2: a genuinely-failed run with NO later run on the branch still +# surfaces as failed - hiding real failures is equally wrong. +test_failed_run_with_no_later_run_still_surfaces() { + reset_fakes + local d short; d=$(new_case f10-genuine-failure) + make_repo_on_branch "$d/wt" fm/feat-f10b + short=$(git -C "$d/wt" rev-parse --short=8 HEAD) + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-f10b.meta" "window=fm:fm-feat-f10b" "worktree=$d/wt" "kind=ship" + FM_FAKE_AXI_STATUS="$(run_failed fm/feat-f10b)" + FM_FAKE_RUNS_LIST=" failed fm/feat-f10b ${short} 2026-08-27 12:09" + local out; out=$(run_crew_state "$d" feat-f10b) + assert_contains "$out" "state: failed" "a genuinely failed run with no later run still reports failed" + assert_contains "$out" "source: run-step" "the genuine failure is run-step sourced" + pass "a genuinely failed run with no later run is not hidden" +} + +# The coarse runs-list scan: an ACTIVE row for this branch at an unresolvable +# head is unknown attribution and must STOP the scan, never fall through onto +# the older failed row (axi status answers another branch here, so attribution +# can only go through the coarse list). +test_coarse_unresolvable_active_row_never_falls_to_older_row() { + reset_fakes + local d short; d=$(new_case f10-coarse-guard) + make_repo_on_branch "$d/wt" fm/feat-f10c + short=$(git -C "$d/wt" rev-parse --short=8 HEAD) + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-f10c.meta" "window=fm:fm-feat-f10c" "worktree=$d/wt" "kind=ship" "harness=claude" + FM_FAKE_AXI_STATUS="$(run_running fm/other-crew)" + FM_FAKE_RUNS_LIST="$(cat </dev/null + fm_write_meta "$d/state/feat-f10d.meta" "window=fm:fm-feat-f10d" "worktree=$d/wt" "kind=ship" "harness=claude" + printf 'working: implementing\n' > "$d/state/feat-f10d.status" + FM_FAKE_AXI_STATUS="$(run_running_pipeline_owned fm/feat-f10d f0f0f0f0 synced)" + FM_FAKE_RUNS_LIST="" + FM_FAKE_BUSY=0 + arm_idle_record "$d/state" feat-f10d + local out; out=$(run_crew_state "$d" feat-f10d) + assert_not_contains "$out" "source: run-step" "a non-pipeline-owned unresolvable head must not bind" + assert_contains "$out" "source: status-log" "falls back to the status log without the exemption" + pass "the exemption requires branch_sync.state=pipeline_owned" +} + +# Negative control: the exemption also requires an ACTIVE run - a terminal run +# released the branch, so an inconsistent pipeline_owned label must not bind a +# terminal run by branch name alone. +test_pipeline_owned_terminal_run_not_exempt() { + reset_fakes + local d; d=$(new_case f10-terminal-not-exempt) + make_repo_on_branch "$d/wt" fm/feat-f10e + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-f10e.meta" "window=fm:fm-feat-f10e" "worktree=$d/wt" "kind=ship" "harness=claude" + printf 'working: stage 2 in progress\n' > "$d/state/feat-f10e.status" + FM_FAKE_AXI_STATUS="$(run_running_pipeline_owned fm/feat-f10e f0f0f0f0) +outcome: failed" + FM_FAKE_RUNS_LIST="" + FM_FAKE_BUSY=0 + arm_idle_record "$d/state" feat-f10e + local out; out=$(run_crew_state "$d" feat-f10e) + assert_not_contains "$out" "source: run-step" "a terminal run must not bind through the exemption" + assert_contains "$out" "source: status-log" "falls back to the status log for a terminal unresolvable head" + pass "the exemption never applies to a terminal run" +} + test_missing_run_head_falls_back_to_current_state() { reset_fakes local d out @@ -1460,6 +1600,11 @@ test_usage_error test_historical_same_branch_rewritten_head_not_current test_active_run_descendant_fix_head_remains_current test_local_advanced_past_run_head_invalidates +test_pipeline_owned_active_run_beats_superseded_failed_row +test_failed_run_with_no_later_run_still_surfaces +test_coarse_unresolvable_active_row_never_falls_to_older_row +test_non_pipeline_owned_unresolvable_head_not_attributed +test_pipeline_owned_terminal_run_not_exempt test_missing_run_head_falls_back_to_current_state echo "all fm-crew-state tests passed" From c651b590edacd261919448c032a7f5b896f1e97b Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Thu, 27 Aug 2026 20:42:57 -0700 Subject: [PATCH 05/25] fix(pi): surface requested outcomes without replaying fleet events (#3211) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(pi): surface requested supervision outcomes * no-mistakes(review): Mirror in-flight captain requests before branch dispatch * no-mistakes(review): Exercise real branch ownership and main outcome access * no-mistakes(review): Preserve request tails and align verdict guidance * no-mistakes(review): Preserve complete current captain requests * no-mistakes(review): Require visible requested outcomes and realistic classification * no-mistakes(document): Align supervision outcome documentation * no-mistakes(ci): Fixed Greptile’s runtime-ordering finding. The extension now stages Pi’s authoritative `before_agent_start` prompt before SessionManager persistence and suppresses the later duplicate entry. Updated docs and behavioral regression to reproduce real Pi ordering and verify each prompt is mirrored exactly once. Passed branch-extension tests, supervision tests, strict Pi typecheck, full lint, and diff checks * no-mistakes(review): Use canonical operational input classification * no-mistakes(review): Filter legacy operational inputs canonically * no-mistakes(document): Clarify captain request mirroring boundary * no-mistakes(ci): Fixed the CI time-boundary failure in tests/fm-public-followup.test.sh by pinning its clock, including context-registry setup. This prevents follow-up fixtures from expiring based on wall time. Verified the full regression suite passes, project-owned lint passes, and git diff checks are clean * no-mistakes(document): Clarify captain-visible supervision outcome documentation --- .pi/extensions/fm-branch-supervision.ts | 140 ++++++++---- bin/fm-branch-prompt.sh | 7 +- docs/architecture.md | 4 +- docs/configuration.md | 4 +- docs/pi-supervision-branch.md | 19 +- docs/supervision-protocols/pi.md | 6 +- tests/fm-branch-supervision.test.sh | 4 + tests/fm-pi-branch-extension.test.sh | 246 ++++++++++++++++++++-- tests/fm-public-followup.test.sh | 8 +- tests/fm-supervision-instructions.test.sh | 2 + 10 files changed, 368 insertions(+), 72 deletions(-) diff --git a/.pi/extensions/fm-branch-supervision.ts b/.pi/extensions/fm-branch-supervision.ts index 8a56bacd9b1..b4d7c1b0e9a 100644 --- a/.pi/extensions/fm-branch-supervision.ts +++ b/.pi/extensions/fm-branch-supervision.ts @@ -6,8 +6,9 @@ // real tools and reports through the fm_branch_report custom tool, which // writes the durable outcome store FIRST (bin/fm-branch-outcome.sh) and then // merges an append-only note to main's tail. Main's captain/assistant dialog -// is mirrored into the branch as read-only fm-main-mirror context at main's -// turn_end. Pi-only by construction: this file lives in .pi/extensions, so no +// is mirrored into the branch as read-only fm-main-mirror context from Pi's +// before_agent_start prompt and at main's turn_end. Pi-only by construction: this +// file lives in .pi/extensions, so no // other harness ever loads it. Supervision is default-on for every task once // this Pi session owns the fleet lock: no captain grant file is required. // Away mode (or a broken branch) keeps today's wake-to-main behavior @@ -95,7 +96,10 @@ import { FOLLOW_MAIN_VALUE, type BranchPickerItem, } from "./lib/fm-branch-model-picker.ts"; -import { encodeFirstmateOperationalInput } from "./lib/fm-operational-input.ts"; +import { + classifyFirstmateOperationalText, + encodeFirstmateOperationalInput, +} from "./lib/fm-operational-input.ts"; const extensionFile = fileURLToPath(import.meta.url); const extensionDir = dirname(extensionFile); @@ -130,21 +134,17 @@ const MERGE_NOTE_BOAT = "⛵"; // Carried inside the captain note's own text because that text is the only // part of a custom message Pi gives the model (see mergeIntoMain). // -// The relay order is CONDITIONAL on purpose. This M1 fix is deliberately a -// model-facing instruction only: mergeIntoMain must keep opening the follow-up -// turn so a genuinely new outcome cannot be suppressed before main sees it. -// The note still needs its self-description to stop main from mistaking an -// incoming outcome for its own earlier answer and silently losing the outcome. -// But an unconditional relay order is false whenever main was separately woken -// for the same event, and turns the correct response - saying nothing new - -// into a mechanical re-report. The conditional wording lets main stay quiet -// when appropriate; source-level identity suppression is outside this M1 fix. +// The note still needs to identify itself so main cannot mistake an incoming +// outcome for its own earlier answer and silently lose the outcome. Event +// ownership forbids a second fleet operation, while the captain-facing verdict +// requires a visible response and leaves its wording to main. const CAPTAIN_OUTCOME_INSTRUCTION = "This is a supervision outcome delivered automatically by the supervision branch. " + "It was not typed by the captain. " + - "If you have already reported this outcome to the captain earlier in this conversation, do not report it again. " + - "Otherwise relay only this outcome to the captain now, in one short message, in captain outcome language. " + - "Do not restate or repeat any earlier answer."; + "The fleet event is already handled: do not re-drain, re-run, or acknowledge it. " + + "This outcome is captain-facing: give the captain a visible response now. " + + "Use your judgment over the wording and how to incorporate it, not whether to surface it. " + + "An outcome that directly answers an explicit captain request is captain-facing, regardless of whether it is healthy, routine, measured, actionable, or requires a decision."; type MirrorItem = { tag: "captain" | "main"; text: string }; type MirrorCursor = { file: string; index: number }; type Verdict = "routine" | "captain"; @@ -305,15 +305,17 @@ function textOfContent(content: unknown): string { // Operational injections (watcher wakes, away-supervisor escalations, launch // briefs) are fleet machinery, not captain dialog; the report's volume // analysis counts them apart from dialog, and mirroring them would feed the -// branch its own supervision traffic back. Current injections start with the -// U+2063 operational prefix; the plain legacy form starts with FIRSTMATE. +// branch its own supervision traffic back. function isOperationalUserText(text: string): boolean { - return text.startsWith("⁣") || /^FIRSTMATE[ _]/.test(text); + return classifyFirstmateOperationalText(text) !== undefined; } function capMirrorText(text: string): string { if (text.length <= MIRROR_MESSAGE_CAP) return text; - return `${text.slice(0, MIRROR_MESSAGE_CAP)}\n[mirror truncated at ${MIRROR_MESSAGE_CAP} characters]`; + const headLength = Math.ceil(MIRROR_MESSAGE_CAP / 2); + const tailLength = MIRROR_MESSAGE_CAP - headLength; + const omitted = text.length - MIRROR_MESSAGE_CAP; + return `${text.slice(0, headLength)}\n[mirror truncated: ${omitted} characters omitted]\n${text.slice(-tailLength)}`; } function readMirrorCursor(): MirrorCursor { @@ -347,6 +349,10 @@ type ReadonlyEntries = { type MirrorCollectionState = { collectAnchor: MirrorCursor | null; pendingCursor: MirrorCursor | null; + // Pi emits before_agent_start before it appends that turn's user message to + // SessionManager. The prompt is mirrored from the event immediately, then + // this marker suppresses the same persisted entry when turn_end collects it. + stagedCaptain: { file: string; index: number; text: string } | null; }; function collectMainDialog(sessionManager: ReadonlyEntries, collection: MirrorCollectionState): MirrorItem[] { @@ -354,8 +360,20 @@ function collectMainDialog(sessionManager: ReadonlyEntries, collection: MirrorCo const entries = sessionManager.getEntries(); const anchor = collection.collectAnchor ?? readMirrorCursor(); const start = anchor.file === file ? Math.min(anchor.index, entries.length) : 0; + let currentCaptainIndex = -1; + for (let index = entries.length - 1; index >= start; index -= 1) { + const entry = entries[index]; + if (entry.type !== "message") continue; + const message = (entry as { message?: { role?: string; content?: unknown } }).message; + if (message?.role !== "user") continue; + const text = textOfContent(message.content).trim(); + if (!text || isOperationalUserText(text)) continue; + currentCaptainIndex = index; + break; + } const items: MirrorItem[] = []; - for (const entry of entries.slice(start)) { + for (let index = start; index < entries.length; index += 1) { + const entry = entries[index]; if (entry.type !== "message") continue; const message = (entry as { message?: { role?: string; content?: unknown } }).message; if (!message) continue; @@ -363,7 +381,20 @@ function collectMainDialog(sessionManager: ReadonlyEntries, collection: MirrorCo const text = textOfContent(message.content).trim(); if (!text) continue; if (message.role === "user" && isOperationalUserText(text)) continue; - items.push({ tag: message.role === "user" ? "captain" : "main", text: capMirrorText(text) }); + const staged = collection.stagedCaptain; + if ( + message.role === "user" && + staged?.file === file && + staged.index === index && + staged.text === text + ) { + collection.stagedCaptain = null; + continue; + } + items.push({ + tag: message.role === "user" ? "captain" : "main", + text: index === currentCaptainIndex ? text : capMirrorText(text), + }); } collection.collectAnchor = { file, index: entries.length }; collection.pendingCursor = collection.collectAnchor; @@ -386,7 +417,12 @@ export default function (pi: ExtensionAPI) { // serially by design). let branchChain: Promise = Promise.resolve(); const pendingMirror: MirrorItem[] = []; - const mirrorCollection: MirrorCollectionState = { collectAnchor: null, pendingCursor: null }; + const mirrorCollection: MirrorCollectionState = { + collectAnchor: null, + pendingCursor: null, + stagedCaptain: null, + }; + let currentMainSession: ReadonlyEntries | null = null; // One revision for BOTH selections: a model or effort change invalidates an // in-flight branch build exactly the same way. let branchSelectionRevision = 0; @@ -573,10 +609,11 @@ export default function (pi: ExtensionAPI) { // therefore has to carry its own identity inside `content`, or main receives // an unattributed user message written in main's own captain-facing voice // and cannot tell an incoming outcome from its own earlier answer. When that - // happens main re-emits its previous answer instead of relaying the outcome, - // and the outcome is lost. The typed operational envelope is what makes the - // note self-describing; it stays invisible to the captain because the note - // is never rendered. + // happens main can lose the outcome while deciding how to handle it. The + // typed operational envelope is what makes the note self-describing; it stays + // invisible to the captain because the note is never rendered. The + // instruction preserves the event-ownership boundary while requiring the + // captain-facing response and leaving its wording to main. // // Encoding shells out, so it can fail on a broken checkout. This file's // failure direction applies: an outcome that cannot be typed is still @@ -632,7 +669,8 @@ export default function (pi: ExtensionAPI) { parameters: Type.Object({ task: Type.String({ description: "The task id the event belongs to (or 'fleet' for fleet-wide events)" }), verdict: Type.Union([Type.Literal("routine"), Type.Literal("captain")], { - description: "captain only for what a human must see; routine otherwise", + description: + "Use captain unconditionally for an outcome that directly answers an explicit captain request, regardless of whether it is healthy, routine, measured, actionable, or requires a decision. Also use captain for work ready for review, captain-only decisions, blockers or failures after recovery is exhausted, needed credentials, and destructive, irreversible, or security-sensitive actions; use routine otherwise.", }), summary: Type.String({ description: @@ -941,6 +979,16 @@ ${context.command} }); } + function collectCurrentMainDialog(): boolean { + if (!currentMainSession) return true; + try { + pendingMirror.push(...collectMainDialog(currentMainSession, mirrorCollection)); + return true; + } catch { + return false; + } + } + function enqueueMirrorFlush(): void { if (!branch || pendingMirror.length === 0) return; const flushGeneration = generation; @@ -966,10 +1014,29 @@ ${context.command} if (!actingAsOwner()) return; // cold start pre-lock, secondary session, or shutdown if (afkActive()) return; // the away daemon owns supervision while afk if (branchBroken) return; // fail back to today's wake-to-main path + if (!collectCurrentMainDialog()) return; offer.accept(); enqueueWake(offer.message, generation); }); + pi.on?.("before_agent_start", (event, ctx) => { + rememberMainModel(ctx); + currentMainSession = ctx?.sessionManager ?? null; + if (!actingAsOwner() || !currentMainSession || !collectCurrentMainDialog()) return; + + // This event is Pi's authoritative complete current prompt. At this point + // SessionManager still contains only the preceding dialog, so relying on + // getEntries() here loses the captain request that the next wake may answer. + // Stage it verbatim and remember the future persisted index for turn_end's + // duplicate suppression. Operational extension injections are not dialog. + const prompt = event.prompt.trim(); + if (!prompt || isOperationalUserText(prompt)) return; + const file = currentMainSession.getSessionFile() ?? ""; + const index = mirrorCollection.collectAnchor?.index ?? currentMainSession.getEntries().length; + pendingMirror.push({ tag: "captain", text: prompt }); + mirrorCollection.stagedCaptain = { file, index, text: prompt }; + }); + pi.on?.("agent_start", () => { mainStreaming = true; }); @@ -980,18 +1047,16 @@ ${context.command} mainStreaming = false; }); - // Mirror at main's turn_end: collect the new captain/assistant dialog into - // the volatile queue, then deliver it through the serialized chain so it - // lands before any later wake. The durable cursor advances only in + // before_agent_start stages Pi's authoritative in-flight prompt before + // SessionManager persists it. The dispatch handler then collects any newly + // persisted dialog immediately before accepting a wake, so all context joins + // the serialized chain before that wake's branch prompt. turn_end remains + // the idle-path mirror flush. The durable cursor advances only in // flushMirror after the complete pending batch reaches the branch. pi.on?.("turn_end", (_event, ctx) => { rememberMainModel(ctx); - if (!actingAsOwner()) return; - try { - pendingMirror.push(...collectMainDialog(ctx.sessionManager, mirrorCollection)); - } catch { - return; - } + currentMainSession = ctx.sessionManager; + if (!actingAsOwner() || !collectCurrentMainDialog()) return; enqueueMirrorFlush(); }); @@ -1004,6 +1069,7 @@ ${context.command} // recorded pointer. Terminal quit simply never fires another session_start. pi.on?.("session_start", (_event, ctx) => { rememberMainModel(ctx); + currentMainSession = ctx?.sessionManager ?? null; shuttingDown = false; branchBroken = ""; generation += 1; @@ -1041,8 +1107,10 @@ ${context.command} shuttingDown = true; generation += 1; pendingMirror.length = 0; + currentMainSession = null; mirrorCollection.collectAnchor = null; mirrorCollection.pendingCursor = null; + mirrorCollection.stagedCaptain = null; if (branch) { try { branch.dispose(); diff --git a/bin/fm-branch-prompt.sh b/bin/fm-branch-prompt.sh index 6b474360d0c..71209d159e1 100755 --- a/bin/fm-branch-prompt.sh +++ b/bin/fm-branch-prompt.sh @@ -63,13 +63,16 @@ For anything it tells you to escalate, or any failure that survives the playbook # Verdict: routine or captain -Report verdict captain only for what a human must see: +Report verdict captain for any outcome that directly answers an explicit captain request. +This rule is unconditional: do not qualify it by whether the result is healthy, routine, measured, actionable, or requires a decision. +Also report verdict captain for: - work ready for review - always include the full https:// PR URL in the summary; - a decision only the captain can make, including every ask-user finding from a validation gate; - a real blocker or failure after the playbook is exhausted; - a needed credential or login; - anything destructive, irreversible, or security-sensitive. -Everything else - routine status, a successful automatic recovery, an absorbed poll, a healthy pause - is verdict routine. +Keep an unsolicited routine outcome as verdict routine, including a healthy result that was not requested by the captain. +Keep an unchanged fleet review silent as instructed above. When genuinely in doubt, choose captain: a spurious escalation costs a glance, a swallowed one costs trust. Write summaries in the captain's outcome language - the project, the fix, the PR, the worker, the blocker - never internal mechanics like wake kinds, status prefixes, worktrees, or state file names. diff --git a/docs/architecture.md b/docs/architecture.md index 15b1cb10d7d..b4ac5b43683 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -70,8 +70,8 @@ The script header owns the exact JSON schema. On a Pi primary, supervision is default-on: the watcher extension can hand eligible task-local rows from an ordinary actionable wake, plus selected fleet-wide heartbeat reviews, to a persistent in-process supervision conversation while main-only rows remain on the captain-facing path. The branch handles those rows, stores the outcome durably, and merges an append-only note back. -A captain-facing outcome instead opens exactly one follow-up turn on the captain's conversation without printing or rendering a separate note - that turn is the captain-visible result. -[docs/pi-supervision-branch.md](pi-supervision-branch.md) owns row eligibility and dispatch architecture, and every other harness keeps the wake-to-main path unchanged. +A captain-facing outcome instead opens exactly one follow-up turn on the captain's conversation without printing or rendering a separate note. +[docs/pi-supervision-branch.md](pi-supervision-branch.md) owns row eligibility and dispatch architecture, while the generated [Pi supervision protocol](supervision-protocols/pi.md) owns MAIN's captain-visible response and merged-event handling; every other harness keeps the wake-to-main path unchanged. ### Registered secondmate current state diff --git a/docs/configuration.md b/docs/configuration.md index a06dd260f7f..dc4a2e8667b 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -43,7 +43,9 @@ Away mode still declines every wake offer, and a broken branch still falls back The branch's role stays bounded exactly as the captain-approved architecture set it: it cannot merge a PR, land local work, or freshly spawn, and every existing captain gate remains unchanged. Homes on any other primary harness never load this feature and are entirely unaffected. `AGENTS.md`'s `state/` inventory routes the branch's runtime files to their format and lifecycle owners. -A captain-facing (verdict `captain`) branch outcome opens exactly one follow-up turn on main - that turn is the captain-visible result, and Pi never separately prints or renders the merge note itself. +A captain-facing (verdict `captain`) branch outcome opens exactly one follow-up turn on main, and Pi never separately prints or renders the merge note itself. +The branch prompt owns the unconditional explicit-request rule and the distinction between captain-facing, unsolicited routine, and unchanged-review outcomes. +The generated [Pi supervision protocol](supervision-protocols/pi.md) owns main's required captain-visible response, event ownership, and conversational treatment for merged outcomes. A no-change heartbeat outcome explicitly reported with `task=fleet` and `silent=true` is delivered silently with no rendered note, while every other routine outcome still appends a rendered, sailboat-prefixed note. ## Pi supervision branch model and effort (config/supervision-branch-model, config/supervision-branch-effort) diff --git a/docs/pi-supervision-branch.md b/docs/pi-supervision-branch.md index 963c7cfe2d8..80964d7ec6a 100644 --- a/docs/pi-supervision-branch.md +++ b/docs/pi-supervision-branch.md @@ -9,7 +9,7 @@ Fleet supervision on the Pi primary harness runs on a second, persistent convers Supervision is default-on: once a Pi primary session owns this home's fleet lock, the branch handles eligible task-local rows from ordinary actionable wakes plus heartbeat scans that the cheap bash-level scan flags as possibly captain-relevant, then merges each outcome back by appending a short note to the captain conversation's tail. Ordinary main-only rows remain on main even when eligible task-local rows share their queue. An unresolvable row makes the scan unsafe and returns the whole wake to main, and every watcher-failure alarm also stays on main. -Only captain-relevant branch outcomes open a turn on main - that follow-up turn is itself the captain-visible outcome, so Pi never separately prints or renders a captain-facing merge note. +Only captain-relevant branch outcomes open a turn on main; the generated [Pi supervision protocol](supervision-protocols/pi.md) requires MAIN to produce the captain-visible response in that turn, while Pi never separately prints or renders a captain-facing merge note. The design source is the captain-approved forked-supervision architecture board, a captain-private fleet record (a self-contained HTML explainer with the measured cache and judgment evidence); this document records the shape it landed as, and the delivering PR cites the board artifact itself. This feature is Pi-only by construction and changes nothing anywhere else: @@ -44,7 +44,9 @@ This feature is Pi-only by construction and changes nothing anywhere else: ## How the branch knows what the captain said -Main's captain and assistant text - never tool calls, tool results, operational injections, or the branch's own merged notes - is mirrored into the branch as read-only `fm-main-mirror` messages at main's turn end, before the next wake is handed over. +Main's captain and assistant text - never tool calls, tool results, operational injections, or the branch's own merged notes - is mirrored into the branch as read-only `fm-main-mirror` messages. +The idle path mirrors at main's turn end. +At `before_agent_start`, Pi's authoritative prompt is staged verbatim before SessionManager persists that user entry, so the complete current captain message precedes any branch wake accepted after that boundary; the later persisted copy is suppressed and older dialog entries remain bounded. The mirror cursor is durable (`state/.branch-mirror-cursor`), so a restart replays only the not-yet-mirrored dialog from main's session file, and a replacement main session re-anchors from its start. The branch prompt frames mirrored text as context for judgment, never as instructions addressed to the branch; an authorization addressed to main (for example "you may merge when green") does not relax the branch's role limits. @@ -52,14 +54,13 @@ The branch prompt frames mirrored text as context for judgment, never as instruc Stage one is unchanged: the bash watcher absorbs everything provably fine at zero token cost. Stage two is the branch's verdict on each handled event, reported through its `fm_branch_report` tool: `routine` merges without a follow-up turn, while `captain` merges with exactly one follow-up turn. -The follow-up turn a `captain` verdict opens is itself the captain-visible outcome, so its merge note is delivered silently and never printed or rendered in Pi. +The generated [Pi supervision protocol](supervision-protocols/pi.md) requires MAIN to produce the captain-visible response in the one follow-up turn a `captain` verdict opens, so its merge note is delivered silently and never printed or rendered in Pi. Because Pi gives the model only a custom message's `content`, that silent note normally carries both a relay instruction and the `branch-outcome` operational kind owned by `bin/fm-operational-input.sh` inside its own text. -This self-description lets main distinguish a new supervision outcome from its own earlier captain-facing answer; without it, main can mistake the outcome for that answer and re-emit the stale answer instead of relaying the outcome. -The relay instruction itself is conditional: it tells main to stay quiet about an outcome it has already given the captain in this conversation, and to relay anything else rather than restate an earlier answer. -An unconditional order made main re-report an outcome the captain already had; this M1 fix deliberately conditions main's relay without adding source-level suppression, because more than one actor can be woken for the same event. -If envelope encoding fails, the note degrades to the same relay instruction as plain text rather than losing the outcome or opening another turn. +This self-description lets main distinguish a new supervision outcome from its own earlier captain-facing answer; without it, main can mistake the outcome for that answer and lose the outcome while deciding how to handle it. +The generated [Pi supervision protocol](supervision-protocols/pi.md) owns main's event-ownership and conversational-treatment instructions for merged outcomes. +If envelope encoding fails, the captain-facing note degrades to the same runtime instruction as plain text rather than losing the outcome or opening another turn. A no-change heartbeat outcome explicitly reported with `task=fleet` and `silent=true` is also delivered silently with no rendered note, while every other `routine` outcome stays rendered with its sailboat prefix. -The verdict criteria in the branch prompt mirror the captain-etiquette escalation list; doubt escalates. +The branch prompt owns the verdict criteria, including its unconditional explicit-request rule; unsolicited routine outcomes remain routine sailboat notes, unchanged fleet reviews remain silent, and doubt escalates. Main can read the durable outcome store on demand through its `fm_branch_outcomes` tool. ## Heartbeat routing @@ -90,6 +91,6 @@ What is new is only the attended path: outside away mode, the branch absorbs the ## Verification -Portable regressions: `tests/fm-pi-branch-extension.test.sh` (dispatch, default-on eligibility, main-only classification, eligible-row claim lifecycle, partial pre-drain recheck, fallback, filter, mirror, model-visible captain-outcome typing and plain-instruction fallback, cache key, persistence, model pin and searchable picker, effort pin), `tests/fm-branch-supervision.test.sh` (prompt stability, store append-only, leases, guards, non-branch-home invariance), the branch-offer, heartbeat-offer, heartbeat-not-ridden-by-a-check, and main-only-check-class tests in `tests/fm-pi-watch-extension.test.sh`, the recovery test in `tests/fm-session-start.test.sh`, and the per-actor consume regression in `tests/fm-wake-queue.test.sh`. +Portable regressions: `tests/fm-pi-branch-extension.test.sh` (dispatch, default-on eligibility, main-only classification, requested-versus-unsolicited outcome delivery, pre-turn-end complete-current-request mirroring, fleet-event ownership, main outcome access, eligible-row claim lifecycle, partial pre-drain recheck, fallback, filter, model-visible captain-outcome typing and plain-instruction fallback, cache key, persistence, model pin and searchable picker, effort pin), `tests/fm-branch-supervision.test.sh` (prompt stability, store append-only, leases, guards, non-branch-home invariance), the branch-offer, heartbeat-offer, heartbeat-not-ridden-by-a-check, and main-only-check-class tests in `tests/fm-pi-watch-extension.test.sh`, the recovery test in `tests/fm-session-start.test.sh`, and the per-actor consume regression in `tests/fm-wake-queue.test.sh`. Live guard: `FM_PI_BRANCH_LIVE_E2E=1 tests/fm-pi-branch-live-e2e.test.sh` exercises the real installed Pi SDK's custom-message conversion and branch-session surfaces with no user credentials and no provider call; run it after every Pi upgrade and record the dated result in [docs/verification/runtime-backends.md](verification/runtime-backends.md). The strict typecheck in `tests/fm-pi-primary-types.test.sh` pins the extension against the installed Pi package. diff --git a/docs/supervision-protocols/pi.md b/docs/supervision-protocols/pi.md index 90bf2b29d45..2d10a05b590 100644 --- a/docs/supervision-protocols/pi.md +++ b/docs/supervision-protocols/pi.md @@ -21,10 +21,12 @@ When this session owns supervision and away mode is not active: The supervision branch is default-on (docs/pi-supervision-branch.md): whenever this session owns the fleet lock and away mode is not active, the watcher extension hands eligible task-local rows from ordinary actionable wakes, plus selected fleet-wide heartbeat reviews, to the persistent in-process supervision branch while main-only rows remain queued for this conversation. A no-change heartbeat outcome explicitly reported with `task=fleet` and `silent=true` is delivered silently with no rendered note, while every other routine outcome returns as an appended, rendered note that leads with ⛵ then the dim outcome text. -A captain-facing outcome instead opens exactly one follow-up turn on this conversation - that turn is the captain-visible result, and no separate note is printed here. +A captain-facing outcome instead opens exactly one follow-up turn on this conversation - MAIN must produce its captain-visible response in that turn, and no separate note is printed here. Before MAIN steers, controls lifecycle, or cleans up a task, claim its lease with `bin/fm-lease.sh claim ` and release it afterwards; a refused claim means the branch is acting on that task right now. This conversation still receives every other fleet-wide or unresolvable wake, the branch's wakes when it is unavailable or away mode is active, and every watcher-failure alarm regardless, so the arm and repair contract above is unchanged. -Treat a merged note or an opened captain-facing turn as already handled - do not re-drain or re-handle its event - and read the durable outcome store with the fm_branch_outcomes tool when the captain asks what happened. +Treat the merged fleet event as already handled for fleet operations: MAIN must not re-drain, re-run, or acknowledge it. +Separately, MAIN applies judgment about whether and how to surface, summarize, reference, or incorporate a merged sailboat outcome in the captain conversation; event ownership does not decide the conversational treatment. +Read the durable outcome store with the fm_branch_outcomes tool when the captain asks what happened. The turn-end guard extension lives at `__FM_PI_TURNEND_EXT__`. The watcher extension lives at `__FM_PI_EXT__`. diff --git a/tests/fm-branch-supervision.test.sh b/tests/fm-branch-supervision.test.sh index 99155a66b2a..4189254b941 100644 --- a/tests/fm-branch-supervision.test.sh +++ b/tests/fm-branch-supervision.test.sh @@ -48,6 +48,10 @@ test_branch_prompt_is_byte_stable_and_above_cache_floor() { *"stuck-crewmate-recovery"*) ;; *) fail "branch prompt lost the inlined recovery playbook" ;; esac + case "$out_a" in + *"Report verdict captain for any outcome that directly answers an explicit captain request."*"This rule is unconditional"*"Keep an unsolicited routine outcome as verdict routine"*"Keep an unchanged fleet review silent"*) ;; + *) fail "branch prompt lost the unconditional requested-outcome or routine-silence rules" ;; + esac pass "branch prompt is byte-stable across homes, cwd, timezone, and time, above the cache floor" } diff --git a/tests/fm-pi-branch-extension.test.sh b/tests/fm-pi-branch-extension.test.sh index d90f760c575..894bf82099e 100644 --- a/tests/fm-pi-branch-extension.test.sh +++ b/tests/fm-pi-branch-extension.test.sh @@ -124,7 +124,12 @@ export function createBashToolDefinition(cwd, options) { parameters: { type: "object" }, __cwd: cwd, __options: options, - execute: async () => ({ content: [], details: undefined }), + execute: async (_toolCallId, params) => { + if (!globalThis.__fmExecuteBranchBash) return { content: [], details: undefined }; + const initial = { command: String(params.command ?? ""), cwd, env: { ...process.env } }; + const context = options.spawnHook ? options.spawnHook(initial) : initial; + return globalThis.__fmExecuteBranchBash(context); + }, }; } @@ -146,6 +151,7 @@ export async function createAgentSession(options) { } session.ops.push({ kind: "prompt", text }); (globalThis.__fmPrompts ??= []).push(text); + await globalThis.__fmOnBranchPrompt?.({ session, text }); }, async sendCustomMessage(message, opts) { if (globalThis.__fmMirrorGate) { @@ -777,18 +783,20 @@ EOF *"This is a supervision outcome delivered automatically by the supervision branch."*"It was not typed by the captain."*"task-9: PR https://example.com/pr/9"*) ;; *) fail "captain outcome body lost its self-description or the outcome itself: $body" ;; esac - # Both halves of the delivered instruction matter and they pull against each - # other: main must be allowed to stay quiet about an outcome it has already - # given the captain, and must still be told to relay everything else instead - # of re-emitting its own last answer. An instruction carrying only one half - # reintroduces either the duplicate or the silent loss. + # Event ownership and conversational judgment are separate contracts. The + # delivered instruction forbids reprocessing the fleet event but leaves main + # free to decide how the outcome belongs in the captain conversation. + case "$body" in + *"The fleet event is already handled: do not re-drain, re-run, or acknowledge it."*) ;; + *) fail "captain outcome body lost the event-ownership boundary: $body" ;; + esac case "$body" in - *"already reported this outcome to the captain"*"do not report it again"*) ;; - *) fail "captain outcome body never lets main deduplicate what it already said: $body" ;; + *"This outcome is captain-facing: give the captain a visible response now."*"Use your judgment over the wording and how to incorporate it, not whether to surface it."*) ;; + *) fail "captain outcome body made visibility optional or removed wording judgment: $body" ;; esac case "$body" in - *"relay only this outcome to the captain now"*"Do not restate or repeat any earlier answer"*) ;; - *) fail "captain outcome body never tells main to relay it instead of repeating: $body" ;; + *"An outcome that directly answers an explicit captain request is captain-facing"*"regardless of whether it is healthy, routine, measured, actionable, or requires a decision."*) ;; + *) fail "captain outcome body lost the unconditional explicit-request rule: $body" ;; esac # The routine note is rendered in the TUI, and its renderer reads the glyph off # the front of this same string, so it must stay plain text. @@ -798,6 +806,203 @@ EOF pass "a captain outcome reaches main's model as typed, self-describing input while routine notes stay plain" } +test_requested_healthy_outcome_and_unsolicited_routine_outcome_delivery() { + local repo home out status + repo="$TMP_ROOT/requested-outcome-root" + home="$TMP_ROOT/requested-outcome-home" + mkdir -p "$home/state" "$home/config" + install_pi_branch_extension_fixture "$repo" + PLUGIN="$repo/.pi/extensions/fm-branch-supervision.ts" FM_HOME="$home" FM_ROOT_OVERRIDE="$ROOT" \ + DRIVER_PRELUDE="$DRIVER_PRELUDE" node --input-type=module > "$TMP_ROOT/node-output" 2>&1 <<'EOF' +const prelude = process.env.DRIVER_PRELUDE; +await eval(`(async () => { ${prelude}; globalThis.__t = { fire, dispatch, settle, sentToMain, outcomeScript, mainTools, home, realRoot }; })()`); +const { fire, dispatch, settle, sentToMain, outcomeScript, mainTools, home, realRoot } = globalThis.__t; +import { existsSync, readFileSync } from "node:fs"; +import { spawnSync } from "node:child_process"; + +const fleetOperations = []; +globalThis.__fmExecuteBranchBash = async (context) => { + const actor = spawnSync( + "bash", + ["-c", '. "$1"; fm_lease_actor', "_", `${realRoot}/bin/fm-lease-lib.sh`], + { encoding: "utf8", cwd: context.cwd, env: context.env }, + ); + if (actor.status !== 0) throw new Error(`branch bash actor resolution failed: ${actor.stderr}`); + const result = spawnSync("bash", ["-c", context.command], { + encoding: "utf8", + cwd: context.cwd, + env: context.env, + }); + fleetOperations.push({ command: context.command, actor: actor.stdout.trim(), status: result.status }); + return { + content: [{ type: "text", text: `${result.stdout}${result.stderr}` }], + details: { stdout: result.stdout, stderr: result.stderr, exitCode: result.status, actor: actor.stdout.trim() }, + isError: result.status !== 0, + }; +}; + +async function runFleetCommand(session, args) { + const bash = session.options.customTools.find((tool) => tool.name === "bash"); + const command = ["bin/fm-wake-drain.sh", ...args].join(" "); + const result = await bash.execute(`fleet-${fleetOperations.length}`, { command }, undefined, undefined, {}); + if (result.isError) throw new Error(`fleet command failed: ${JSON.stringify(result)}`); + return result.details; +} + +function directlyRequestsResourceReport(mirror) { + const latestCaptain = mirror.at(-1) ?? ""; + const words = new Set(latestCaptain.toLowerCase().split(/[^a-z0-9]+/).filter(Boolean)); + const requestsDelivery = ["give", "provide", "send", "show"].some((word) => words.has(word)); + const namesReport = ["report", "status", "measurement"].some((word) => words.has(word)); + const namesResources = ["resource", "resources", "cpu", "memory"].some((word) => words.has(word)); + return requestsDelivery && namesReport && namesResources; +} + +globalThis.__fmOnBranchPrompt = async ({ session }) => { + const mirror = session.ops + .filter((op) => op.kind === "custom" && op.message.customType === "fm-main-mirror") + .map((op) => op.message.content); + const directlyRequested = directlyRequestsResourceReport(mirror); + const drained = await runFleetCommand(session, []); + const ack = drained.stderr.match(/--ack-through ([0-9]+) --recovery-generation ([A-Za-z0-9._-]+)/); + if (!ack) throw new Error(`drain did not return its acknowledgement command: ${drained.stderr}`); + const report = session.options.customTools.find((tool) => tool.name === "fm_branch_report"); + const verdictDescription = report.parameters.properties.verdict.description; + if (!verdictDescription.includes("unconditionally") || + !verdictDescription.includes("directly answers an explicit captain request") || + !verdictDescription.includes("regardless of whether it is healthy, routine, measured, actionable, or requires a decision")) { + throw new Error(`branch provider received conflicting verdict semantics: ${verdictDescription}`); + } + const result = await report.execute( + `resource-result-${fleetOperations.length}`, + { + task: "task-resource", + verdict: directlyRequested ? "captain" : "routine", + summary: "healthy resource report: CPU 12%, memory 41%", + wake: "signal: healthy resource result", + }, + undefined, + undefined, + {}, + ); + if (result.isError) throw new Error(`branch report failed: ${JSON.stringify(result)}`); + await runFleetCommand(session, ["--ack-through", ack[1], "--recovery-generation", ack[2]]); +}; + +const explicitRequest = "Please give me a fresh mini system-resource report."; +const longRequests = [ + `${explicitRequest}${" head context".repeat(500)}`, + `${"middle context ".repeat(250)}${explicitRequest}${" middle context".repeat(250)}`, + `${"tail context ".repeat(500)}${explicitRequest}`, +]; +const requestedPrompts = [...longRequests, "FIRSTMATE give me a fresh system-resource report."]; +// Match Pi's real AgentSession.prompt ordering: before_agent_start receives +// the expanded prompt before _runAgentPrompt appends its user message to the +// SessionManager. Keeping entries stale at the hook boundary is the regression. +const entries = []; +const mainCtx = { + model: { provider: "anthropic", id: "main-model" }, + sessionManager: { + getSessionFile: () => `${home}/main.jsonl`, + getEntries: () => entries, + }, +}; +const operational = spawnSync( + "bash", + [`${realRoot}/bin/fm-operational-input.sh`, "encode", "watcher"], + { encoding: "utf8", input: "operational watcher injection" }, +); +if (operational.status !== 0) throw new Error(`could not create operational input: ${operational.stderr}`); +fire("before_agent_start", { prompt: operational.stdout }, mainCtx); +entries.push({ type: "message", message: { role: "user", content: operational.stdout } }); +const unsolicitedPrompt = "Please keep responses concise while monitoring the fleet."; +fire("before_agent_start", { prompt: unsolicitedPrompt }, mainCtx); +entries.push({ type: "message", message: { role: "user", content: unsolicitedPrompt } }); +fire("agent_start", {}, mainCtx); +fire("agent_end", {}, mainCtx); +const legacyOperational = "⁣FIRSTMATE_OP: give me a fresh system-resource report."; +fire("before_agent_start", { prompt: legacyOperational }, mainCtx); +entries.push({ type: "message", message: { role: "user", content: legacyOperational } }); +fire("agent_start", {}, mainCtx); +const unsolicited = dispatch("signal: healthy resource result"); +if (!unsolicited.accepted) throw new Error("branch did not accept the unsolicited result"); +await settle(() => fleetOperations.length === 2, "unsolicited result acknowledgement"); +if (sentToMain.length !== 1 || sentToMain[0].options.triggerTurn) { + throw new Error(`unsolicited healthy result opened a main turn: ${JSON.stringify(sentToMain)}`); +} +const sailboat = sentToMain[0]; +if (sailboat.message.display !== true || !sailboat.message.content.startsWith("⛵ task-resource:")) { + throw new Error(`unsolicited healthy result was not a rendered sailboat note: ${JSON.stringify(sailboat)}`); +} + +const outcomes = mainTools.find((tool) => tool.name === "fm_branch_outcomes"); +if (!outcomes) throw new Error("main did not receive its outcome-reading permission surface"); +const visibleToMain = await outcomes.execute("main-reads-sailboat", { recent: 1 }, undefined, undefined, {}); +const mainOutcomeText = visibleToMain.content.map((item) => item.text ?? "").join("\n"); +if (visibleToMain.isError || !mainOutcomeText.includes("healthy resource report: CPU 12%, memory 41%")) { + throw new Error(`main could not use the sailboat content through its existing permission path: ${JSON.stringify(visibleToMain)}`); +} +if (fleetOperations.length !== 2) throw new Error("main's outcome read reprocessed the fleet event"); + +for (let index = 0; index < requestedPrompts.length; index += 1) { + const content = requestedPrompts[index]; + if (index < longRequests.length && content.length <= 4000) { + throw new Error(`request fixture ${index} did not exceed the mirror bound`); + } + fire("before_agent_start", { prompt: content }, mainCtx); + // Pi persists this only after every before_agent_start handler has returned. + entries.push({ type: "message", message: { role: "user", content } }); + fire("agent_start", {}, mainCtx); + const requested = dispatch("signal: healthy resource result"); + if (!requested.accepted) throw new Error(`branch did not accept requested result ${index}`); + await settle(() => fleetOperations.length === 4 + (index * 2), `requested result ${index} acknowledgement`); + const deliveredRequestMirror = globalThis.__fmSessions[0].ops + .filter((op) => op.kind === "custom" && op.message.customType === "fm-main-mirror") + .at(-1)?.message.content; + if (deliveredRequestMirror !== `[captain] ${content}`) { + throw new Error(`pre-turn-end mirror changed long captain request ${index}`); + } + const turns = sentToMain.filter((sent) => sent.options.triggerTurn === true); + if (turns.length !== index + 1 || turns.at(-1).options.deliverAs !== "followUp") { + throw new Error(`requested result ${index} did not open exactly one main turn: ${JSON.stringify(sentToMain)}`); + } +} +const mirroredCaptainText = globalThis.__fmSessions[0].ops + .filter((op) => op.kind === "custom" && op.message.customType === "fm-main-mirror") + .map((op) => op.message.content); +for (const content of [unsolicitedPrompt, ...requestedPrompts]) { + const copies = mirroredCaptainText.filter((text) => text === `[captain] ${content}`).length; + if (copies !== 1) throw new Error(`current captain prompt was mirrored ${copies} times instead of once`); +} +if (mirroredCaptainText.some((text) => + text.includes("operational watcher injection") || text.includes("FIRSTMATE_OP: give me a fresh system-resource report") +)) { + throw new Error("canonical current or legacy operational input entered captain mirror context"); +} +if ((globalThis.__fmPrompts ?? []).length !== 5) throw new Error("a handled fleet wake was rerun"); +if (sentToMain.length !== 5) throw new Error(`one result was reprocessed into ${sentToMain.length} main messages`); +if (fleetOperations.length !== 10 || fleetOperations.some((operation) => operation.status !== 0)) { + throw new Error(`fleet event ownership repeated or failed work: ${JSON.stringify(fleetOperations)}`); +} +if (fleetOperations.some((operation) => operation.actor !== "branch")) { + throw new Error(`main took fleet-event ownership: ${JSON.stringify(fleetOperations)}`); +} +if (existsSync(`${home}/state/.wake-queue`) && readFileSync(`${home}/state/.wake-queue`, "utf8") !== "") { + throw new Error("acknowledged fleet wake remained queued for another owner"); +} +const rows = readFileSync(`${home}/state/branch-outcomes.jsonl`, "utf8").trim().split("\n").map((line) => JSON.parse(line)); +if (rows.length !== 5 || rows[0].verdict !== "routine" || rows.slice(1).some((row) => row.verdict !== "captain")) { + throw new Error(`provider classifications were not recorded once in order: ${JSON.stringify(rows)}`); +} +if (outcomeScript(["unread"]) !== "") throw new Error("merged outcomes remained unread for redelivery"); +process.exit(0); +EOF + status=$? + out=$(cat "$TMP_ROOT/node-output") + expect_code 0 "$status" "requested and unsolicited healthy outcomes must follow their distinct public delivery paths: $out" + pass "requested and unsolicited healthy outcomes keep distinct delivery and event ownership" +} + test_captain_outcome_encoding_failure_delivers_plain_instruction() { local repo home out status repo="$TMP_ROOT/encoding-fallback-root" @@ -834,9 +1039,9 @@ if (delivered.options.triggerTurn !== true || delivered.options.deliverAs !== "f if (delivered.message.content.includes("FIRSTMATE_OP:")) { throw new Error(`fallback unexpectedly carried an envelope: ${delivered.message.content}`); } -if (!delivered.message.content.includes("relay only this outcome to the captain now") || - !delivered.message.content.includes("do not report it again") || - !delivered.message.content.includes("Do not restate or repeat any earlier answer") || +if (!delivered.message.content.includes("The fleet event is already handled: do not re-drain, re-run, or acknowledge it.") || + !delivered.message.content.includes("This outcome is captain-facing: give the captain a visible response now.") || + !delivered.message.content.includes("Use your judgment over the wording and how to incorporate it, not whether to surface it.") || !delivered.message.content.includes("task-fallback: PR https://example.com/pr/fallback is ready")) { throw new Error(`fallback lost its instruction or outcome: ${delivered.message.content}`); } @@ -1277,13 +1482,13 @@ const { fire, dispatch, settle, home } = globalThis.__t; import { existsSync, readFileSync } from "node:fs"; const entries = [ - { type: "message", message: { role: "user", content: "never merge task-7 without my word" } }, + { type: "message", message: { role: "user", content: `never merge task-7 without my word ${"h".repeat(5000)} old-history-tail` } }, { type: "message", message: { role: "assistant", content: [{ type: "text", text: "aye, holding task-7" }, { type: "toolCall", id: "t1" }] } }, { type: "message", message: { role: "user", content: "⁣FIRSTMATE_OP: v1 watcher: operational injection" } }, { type: "message", message: { role: "toolResult", content: "tool output stays in main" } }, { type: "custom", message: { role: "custom", customType: "fm-branch-merge", content: "merged note" } }, { type: "compaction", summary: "compacted" }, - { type: "message", message: { role: "user", content: `pad ${"x".repeat(5000)}` } }, + { type: "message", message: { role: "user", content: `pad ${"x".repeat(5000)}\ntail: retain this request` } }, ]; const ctx = { sessionManager: { @@ -1306,9 +1511,15 @@ if (JSON.stringify(kinds) !== JSON.stringify(["custom", "custom", "custom", "pro const mirrored = session.ops.filter((op) => op.kind === "custom").map((op) => op.message); if (mirrored.some((m) => m.customType !== "fm-main-mirror")) throw new Error("mirror used the wrong custom type"); if (mirrored.some((m) => m.display !== false)) throw new Error("mirrored context must be silent"); -if (mirrored[0].content !== "[captain] never merge task-7 without my word") throw new Error(`bad captain mirror: ${mirrored[0].content}`); +if (!mirrored[0].content.startsWith("[captain] never merge task-7 without my word") || + !mirrored[0].content.includes("[mirror truncated:") || + !mirrored[0].content.endsWith("old-history-tail")) { + throw new Error(`older captain history was not bounded: ${mirrored[0].content}`); +} if (mirrored[1].content !== "[main] aye, holding task-7") throw new Error(`bad main mirror: ${mirrored[1].content}`); -if (!mirrored[2].content.includes("[mirror truncated at 4000 characters]")) throw new Error("long dialog was not capped"); +if (mirrored[2].content !== `[captain] ${entries[6].message.content}`) { + throw new Error("current captain dialog was not preserved completely"); +} if (mirrored.some((m) => m.content.includes("operational injection") || m.content.includes("tool output") || m.content.includes("merged note"))) { throw new Error("mirror leaked operational, tool, or merge-note traffic"); } @@ -2878,6 +3089,7 @@ JS test_outcomes_tool_uses_stock_execution_and_export_consumers test_real_pi_picker_primitives_stay_bounded_and_searchable test_branch_dispatch_two_stage_filter_and_prefix_contract +test_requested_healthy_outcome_and_unsolicited_routine_outcome_delivery test_captain_outcome_encoding_failure_delivers_plain_instruction test_branch_dispatch_classifies_main_only_rows_and_writes_the_eligible_snapshot test_branch_cache_key_is_per_home_stable diff --git a/tests/fm-public-followup.test.sh b/tests/fm-public-followup.test.sh index 69a054abee8..a4f8d7bcad3 100755 --- a/tests/fm-public-followup.test.sh +++ b/tests/fm-public-followup.test.sh @@ -23,6 +23,7 @@ TEARDOWN="$ROOT/bin/fm-teardown.sh" PROMOTE="$ROOT/bin/fm-promote.sh" SESSION_START="$ROOT/bin/fm-session-start.sh" TMP_ROOT=$(fm_test_tmproot fm-public-followup) +PF_TEST_NOW=1787539200 command -v jq >/dev/null 2>&1 || { echo "skip: jq not found"; exit 0; } command -v tasks-axi >/dev/null 2>&1 || { echo "skip: tasks-axi not found"; exit 0; } @@ -97,7 +98,8 @@ run_pf() { # shift PATH="$home/fakebin:$PATH" FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$home" \ FM_STATE_OVERRIDE="$home/state" FAKE_CURL_LOG="${FAKE_CURL_LOG:-}" \ - FAKE_FOLLOWUP_CODE="${FAKE_FOLLOWUP_CODE:-200}" "$PF" "$@" + FAKE_FOLLOWUP_CODE="${FAKE_FOLLOWUP_CODE:-200}" \ + FMX_NOW_OVERRIDE="${FMX_NOW_OVERRIDE:-$PF_TEST_NOW}" "$PF" "$@" } tasks_in() { # @@ -142,7 +144,7 @@ seed_commitment() { > "$home/state/x-inbox/$request.json" chmod 700 "$home/state/x-inbox" chmod 600 "$home/state/x-inbox/$request.json" - FM_HOME="$home" bash -c \ + FM_HOME="$home" FMX_NOW_OVERRIDE="$PF_TEST_NOW" bash -c \ ". '$ROOT/bin/fm-x-lib.sh'; fmx_context_registry_set '$home/state' '$request' '$platform' 1900" \ || fail "could not retain the private request context" @@ -172,7 +174,7 @@ seed_repro_commitment() { # /dev/null || fail "add failed" tasks_in "$home" public-followup bind-work "$obligation" --relation-file "$home/relation.json" >/dev/null \ || fail "bind-work failed" - FM_HOME="$home" bash -c \ + FM_HOME="$home" FMX_NOW_OVERRIDE="$PF_TEST_NOW" bash -c \ ". '$ROOT/bin/fm-x-lib.sh'; fmx_context_registry_set '$home/state' '$request' discord 2000" \ || fail "context retain failed" run_pf "$home" register "$obligation" --relation rel-code --work-home "$work_home" \ diff --git a/tests/fm-supervision-instructions.test.sh b/tests/fm-supervision-instructions.test.sh index 377e95d152a..75ae69a6d50 100755 --- a/tests/fm-supervision-instructions.test.sh +++ b/tests/fm-supervision-instructions.test.sh @@ -170,6 +170,8 @@ test_pi_snippet_uses_effective_extension_path() { assert_contains "$out" "-e $turnend -e $watch" "pi snippet did not render both effective extension launch paths" assert_contains "$out" "The turn-end guard extension lives at \`$turnend\`" "pi snippet did not render the turn-end guard extension path" assert_contains "$out" "The watcher extension lives at \`$watch\`" "pi snippet did not render the watcher extension path" + assert_contains "$out" "MAIN must not re-drain, re-run, or acknowledge it" "pi snippet lost merged-event ownership" + assert_contains "$out" "MAIN applies judgment about whether and how to surface, summarize, reference, or incorporate a merged sailboat outcome" "pi snippet imposed a mechanical sailboat treatment" assert_not_contains "$out" "__FM_PI_EXT__" "renderer leaked the Pi extension path placeholder" assert_not_contains "$out" "__FM_PI_TURNEND_EXT__" "renderer leaked the Pi turn-end extension path placeholder" assert_not_contains "$out" "state/fm-primary-pi-watch.ts" "pi snippet kept the old generated state-relative extension path" From 1fd7ea289b7a4c23a1fd9474680ed2facd6b7dd1 Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Thu, 27 Aug 2026 21:03:46 -0700 Subject: [PATCH 06/25] feat(bin): add concurrent bounded remote transport lanes (#3210) * feat(bin): per-home remote transport lanes with cancellation, bounded send, and closed stdin All remote commands for every home on one host used to serialize through one single-job-at-a-time worker on one shared queue: a timed-out caller abandoned a staged job that kept running, retries convoyed behind it, fm-send's remote leg had no time bound, and staging captured the caller's stdin to EOF so any fm-on.sh caller with an open stdin wedged staging indefinitely. - The worker now serves one lane per staged home: same-home jobs run strictly FIFO in a new staging-sequence order while different homes run concurrently, each lane as its own top-level worker process (a backgrounded subshell does not reliably reap dead children, so a zombie group leader kept a finished command's process group signalable). Long-poll preemption is lane-scoped. - A caller that disconnects or times out cancels its job: the entrypoint marks the record on any post-staging exit and probes its parent so a dead ssh channel cancels without a signal; the worker skips cancelled queued jobs, terminates a running cancelled job's process group, and reaps the record. - fm-send's remote leg is bounded by FM_SEND_REMOTE_BUDGET (default 30s) and a bound hit exits through the existing unconfirmed-delivery contract, which stays idempotent because the remote enqueue deduplicates. - fm-on.sh defaults the remote command's stdin to /dev/null; the three payload callers pass the new --stdin flag. Abandoned .stage.* litter is age-reaped. - The job execution deadline no longer loses up to a second to clock truncation. * no-mistakes(review): Protect live stages and validate send budgets early * no-mistakes(review): Preserve sequence lock ownership during stale recovery * no-mistakes(review): Allocate job sequences at publication boundary * no-mistakes(review): Bound remote keys and extend stale lock recovery * no-mistakes(document): Document bounded remote transport behavior * no-mistakes(lint): Suppress intentional deferred-expansion lint warning * no-mistakes(ci): Fixed stale sequence-lock recovery by reconciling the counter against published job records before allocating the next sequence, preventing duplicate sequences and same-home FIFO violations. Added a behavioral regression test reproducing displacement after publication and verifying execution order. Passed fm-remote-transport-lanes.test.sh, fm-remote-job.test.sh, fm-lint.sh, and git diff --check * no-mistakes(review): Use atomic sequence claims and lossless lane keys * no-mistakes(review): Recover regressed sequence hints and rate-limit claim reaping * no-mistakes(review): Restrict worker heartbeats to serving loop * no-mistakes(review): Verify supervisor identity before lane recovery signals * no-mistakes(review): Verify tracked lane and claim owner identities * no-mistakes(document): Clarify remote lane and transport contracts * no-mistakes(ci): Fixed the CI time-boundary failure by pinning fm-public-followup tests to a deterministic clock, including context-registry setup. Verified tests/fm-public-followup.test.sh, tests/fm-remote-transport-lanes.test.sh, shellcheck, and git diff --check * no-mistakes(review): Preserve assigned lane ownership of queued jobs * no-mistakes(review): Reserve homes owned by foreign queued lanes * no-mistakes(review): Preserve completed results during crash recovery * no-mistakes(review): Harden claim cleanup, expiry, and cancellation races * no-mistakes(review): Verify process groups and reap abandoned results * no-mistakes(review): Stop leaderless groups and reap cancelled publications * no-mistakes(document): Correct remote transport lifecycle documentation * no-mistakes(lint): Quote done state comparisons for ShellCheck --- bin/fm-backlog-handoff.sh | 2 +- bin/fm-on.sh | 39 +- bin/fm-remote-entrypoint.sh | 43 ++- bin/fm-remote-home-seed.sh | 2 +- bin/fm-remote-inherit-push.sh | 2 +- bin/fm-remote-job-lib.sh | 258 +++++++++++-- bin/fm-remote-job-worker.sh | 476 ++++++++++++++++++++---- bin/fm-send.sh | 52 ++- bin/fm-test-run.sh | 1 + docs/remote-secondmates.md | 10 +- tests/fm-on.test.sh | 21 +- tests/fm-remote-job.test.sh | 15 +- tests/fm-remote-transport-lanes.test.sh | 425 +++++++++++++++++++++ tests/fm-send-remote-delivery.test.sh | 96 +++++ 14 files changed, 1304 insertions(+), 138 deletions(-) create mode 100755 tests/fm-remote-transport-lanes.test.sh diff --git a/bin/fm-backlog-handoff.sh b/bin/fm-backlog-handoff.sh index fa729c9d1b6..879de6053db 100755 --- a/bin/fm-backlog-handoff.sh +++ b/bin/fm-backlog-handoff.sh @@ -549,7 +549,7 @@ remote_deliver_outbox() { # mv -f -- "$counter_tmp" "$counter" \ || { rm -f -- "$snapshot" "$counter_tmp"; return 1; } remote_rel="state/handoff/$id.outbox.md" - if ! "$SCRIPT_DIR/fm-on.sh" "$id" fm-remote-file.sh put "$remote_rel" 1048576 \ + if ! "$SCRIPT_DIR/fm-on.sh" --stdin "$id" fm-remote-file.sh put "$remote_rel" 1048576 \ "$bytes" "$hash" "$generation" < "$snapshot"; then rm -f -- "$snapshot" echo "error: handoff transfer to $id was unavailable or completion is unknown; outbox preserved at $outbox" >&2 diff --git a/bin/fm-on.sh b/bin/fm-on.sh index 5e24f2cef1d..eff02f7c350 100755 --- a/bin/fm-on.sh +++ b/bin/fm-on.sh @@ -2,7 +2,7 @@ # Execute one tracked Firstmate command in a configured remote secondmate home. # # Usage: -# fm-on.sh [args...] +# fm-on.sh [--stdin] [args...] # # Routes come only from remote records in data/secondmates.md. A record names an # SSH config alias, remote Firstmate code root, and remote FM_HOME. A host alias @@ -11,11 +11,14 @@ # bin/fm-*.sh namespace. No per-command table exists. # # argv is encoded as one NUL-delimited stream and passed through the fixed -# fm-remote-entrypoint.sh. stdin remains the caller's stdin, stdout and stderr -# remain separate, and ssh's exit status is returned unchanged. OpenSSH never -# receives an auto-retry instruction here. Exit 255 therefore means unavailable -# transport or unknown remote completion and must be reconciled by the semantic -# caller, never blindly repeated by this layer. +# fm-remote-entrypoint.sh. The remote command's stdin is /dev/null by default, +# because remote staging captures stdin to EOF and an open caller stream would +# block staging indefinitely; a payload caller passes --stdin to forward its +# own stream as the job's bounded input. stdout and stderr remain separate, and +# ssh's exit status is returned unchanged. OpenSSH never receives an auto-retry +# instruction here. Exit 255 therefore means unavailable transport or unknown +# remote completion and must be reconciled by the semantic caller, never +# blindly repeated by this layer. # # The SSH alias keeps normal public-key and strict host-key policy in ~/.ssh. # This command explicitly disables agent forwarding, forwarding setup, and @@ -42,12 +45,17 @@ PROTOCOL=1 . "$SCRIPT_DIR/fm-secondmate-registry-lib.sh" die() { printf 'error: %s\n' "$1" >&2; exit 1; } -usage() { sed -n '2,23p' "$0" | sed 's/^# \{0,1\}//'; exit 2; } +usage() { sed -n '2,25p' "$0" | sed 's/^# \{0,1\}//'; exit 2; } encode_base64() { base64 | tr -d '\n' } +STDIN_MODE=closed +if [ "${1:-}" = --stdin ]; then + STDIN_MODE=caller + shift +fi [ "$#" -ge 2 ] || usage ROUTE=$1 COMMAND=$2 @@ -103,10 +111,15 @@ case "$ALIVE_COUNT_MAX" in ''|*[!0-9]*) die "FM_SSH_ALIVE_COUNT_MAX must be a po [ "$ALIVE_INTERVAL" -gt 0 ] || die "FM_SSH_ALIVE_INTERVAL must be a positive integer: $ALIVE_INTERVAL" [ "$ALIVE_COUNT_MAX" -gt 0 ] || die "FM_SSH_ALIVE_COUNT_MAX must be a positive integer: $ALIVE_COUNT_MAX" -"$SSH_BIN" \ - -o ForwardAgent=no \ - -o ClearAllForwardings=yes \ - -o 'SendEnv=-*' \ - -o "ServerAliveInterval=$ALIVE_INTERVAL" \ - -o "ServerAliveCountMax=$ALIVE_COUNT_MAX" \ +SSH_ARGS=( + -o ForwardAgent=no + -o ClearAllForwardings=yes + -o 'SendEnv=-*' + -o "ServerAliveInterval=$ALIVE_INTERVAL" + -o "ServerAliveCountMax=$ALIVE_COUNT_MAX" -- "$HOST" fm-remote-entrypoint.sh "$PROTOCOL" "$ROOT_B64" "$HOME_B64" "$ARGV_B64" +) +if [ "$STDIN_MODE" = caller ]; then + exec "$SSH_BIN" "${SSH_ARGS[@]}" +fi +exec "$SSH_BIN" "${SSH_ARGS[@]}" < /dev/null diff --git a/bin/fm-remote-entrypoint.sh b/bin/fm-remote-entrypoint.sh index 6763e8c955d..4549ff6e9ca 100755 --- a/bin/fm-remote-entrypoint.sh +++ b/bin/fm-remote-entrypoint.sh @@ -19,6 +19,15 @@ # disconnect remains unknown completion to fm-on.sh, which preserves OpenSSH's # exit 255 behavior. The shared library header owns job fields, bounds, PATH, # LaunchAgent contract, and worker environment. +# +# A staged job whose caller goes away is cancelled rather than abandoned: any +# exit after staging and before the published result marks the job cancelled +# (signal traps cover a delivered HUP/TERM/PIPE/INT, and the exit trap covers a +# failed bounded wait), and while waiting this process probes its parent about +# once per second, so an ssh channel that dies without delivering any signal - +# sshd exiting and reparenting this process - also cancels the job. The worker +# then skips or stops the cancelled job instead of running it to completion for +# nobody. set -eu PROTOCOL=1 @@ -74,7 +83,36 @@ sha256_file() { # [ "$#" -eq 4 ] || die "remote entrypoint expects protocol, root, home, and argv" [ "$1" = "$PROTOCOL" ] || die "incompatible remote protocol: local=$1 remote=$PROTOCOL" TMP=$(mktemp -d "${TMPDIR:-/tmp}/fm-remote-entrypoint.XXXXXX") || die "cannot create protocol staging directory" 70 -trap 'rm -rf -- "$TMP"' EXIT + +JOB_ID= +JOB_COMPLETED=0 +ACCOUNT_HOME= +ENTRYPOINT_PPID=$(ps -o ppid= -p $$ 2>/dev/null | tr -d ' ' || true) + +# The recorded parent is the ssh session process; when it disappears this +# process is reparented and the caller is provably gone. An unreadable probe +# never cancels: only an observed parent change does. +# shellcheck disable=SC2329 # Invoked by fm_remote_job_wait through FM_REMOTE_JOB_DISCONNECT_PROBE. +entrypoint_caller_connected() { + local current + case "$ENTRYPOINT_PPID" in ''|*[!0-9]*) return 0 ;; esac + current=$(ps -o ppid= -p $$ 2>/dev/null | tr -d ' ' || true) + case "$current" in ''|*[!0-9]*) return 0 ;; esac + [ "$current" = "$ENTRYPOINT_PPID" ] +} + +# shellcheck disable=SC2329 # Invoked through the EXIT trap below. +entrypoint_cleanup() { + rm -rf -- "$TMP" + if [ -n "$JOB_ID" ] && [ "$JOB_COMPLETED" -eq 0 ] && [ -n "$ACCOUNT_HOME" ]; then + fm_remote_job_cancel "$ACCOUNT_HOME" "$JOB_ID" 2>/dev/null || true + fi +} +trap entrypoint_cleanup EXIT +trap 'exit 129' HUP +trap 'exit 130' INT +trap 'exit 141' PIPE +trap 'exit 143' TERM decode_text "remote root" "$2" "$TMP/root" decode_text "remote home" "$3" "$TMP/home" @@ -138,11 +176,14 @@ if ! fm_remote_job_ensure_worker "$ROOT" "$ACCOUNT_HOME"; then die "${FM_REMOTE_JOB_ERROR:-remote job worker is unavailable; run fm-on.sh fm-remote-doctor.sh --fix}" fi if ! JOB_ID=$(fm_remote_job_stage "$ACCOUNT_HOME" "$ROOT" "$HOME_PATH" "$COMMAND" "${ARGV[@]:1}"); then + JOB_ID= die "${FM_REMOTE_JOB_ERROR:-cannot stage remote job}" 70 fi +FM_REMOTE_JOB_DISCONNECT_PROBE=entrypoint_caller_connected if ! fm_remote_job_wait "$ACCOUNT_HOME" "$JOB_ID"; then die "${FM_REMOTE_JOB_ERROR:-remote job did not complete}" 70 fi +JOB_COMPLETED=1 cat "$FM_REMOTE_JOB_STDOUT" cat "$FM_REMOTE_JOB_STDERR" >&2 RESULT=$FM_REMOTE_JOB_EXIT diff --git a/bin/fm-remote-home-seed.sh b/bin/fm-remote-home-seed.sh index a679851cbc4..7deafc40dcf 100755 --- a/bin/fm-remote-home-seed.sh +++ b/bin/fm-remote-home-seed.sh @@ -242,7 +242,7 @@ if [ "$PREFLIGHT_RC" -ne 0 ]; then fi set +e -PROVISION_OUT=$("$SCRIPT_DIR/fm-on.sh" "$ID" fm-remote-home-provision.sh < "$TMP/manifest" 2>&1) +PROVISION_OUT=$("$SCRIPT_DIR/fm-on.sh" --stdin "$ID" fm-remote-home-provision.sh < "$TMP/manifest" 2>&1) PROVISION_RC=$? set -e if [ "$PROVISION_RC" -ne 0 ]; then diff --git a/bin/fm-remote-inherit-push.sh b/bin/fm-remote-inherit-push.sh index f0d6f416d4c..ed068622986 100755 --- a/bin/fm-remote-inherit-push.sh +++ b/bin/fm-remote-inherit-push.sh @@ -80,7 +80,7 @@ while IFS= read -r rel; do [ -f "$snapshot" ] && [ ! -L "$snapshot" ] || die "inherited source snapshot is unsafe: $source" bytes=$(LC_ALL=C wc -c < "$snapshot" | tr -d ' ') hash=$(sha256_file "$snapshot") || die "cannot hash inherited source: $source" - "$SCRIPT_DIR/fm-on.sh" "$ID" fm-remote-inherit.sh put "$rel" "$bytes" "$hash" "$GENERATION" < "$snapshot" + "$SCRIPT_DIR/fm-on.sh" --stdin "$ID" fm-remote-inherit.sh put "$rel" "$bytes" "$hash" "$GENERATION" < "$snapshot" else # This loop's heredoc is its control stream, not remote command input. "$SCRIPT_DIR/fm-on.sh" "$ID" fm-remote-inherit.sh absent "$rel" 0 "$EMPTY_HASH" "$GENERATION" < /dev/null diff --git a/bin/fm-remote-job-lib.sh b/bin/fm-remote-job-lib.sh index 25d7bb73b40..f6ac2ad9b99 100755 --- a/bin/fm-remote-job-lib.sh +++ b/bin/fm-remote-job-lib.sh @@ -7,23 +7,53 @@ # isolated tests), the bounded job record, worker installation, and the remote # runtime PATH. # -# A job directory is mode 0700 and contains root, home, argv (NUL-delimited), -# stdin, stdout, stderr, queue_deadline, timeout, deadline, exit, and state. -# Stage writes state=queued last. The worker atomically claims a job with -# .claim, establishes its execution deadline, changes state to running, writes -# bounded stdout/stderr and exit, then publishes state=done last. Callers wait -# for done, relay stdout and stderr separately, then reap only their completed -# record. Input, argv, stdout, and stderr are each capped at 1048576 bytes. +# A published job directory is mode 0700 and contains root, home, argv +# (NUL-delimited), stdin, seq, stdout, stderr, queue_deadline, timeout, and +# state; deadline and exit are added as execution advances, cancel is an +# optional caller-cancellation marker, and .claim may hold owner, owner_start, +# supervisor, supervisor_start, group, group_start, and armed records while +# work executes. +# Stage writes state=queued last. seq is a queue-wide monotonic staging +# sequence reserved atomically by its persistent .seq-claims directory; the +# counter is only a forward-moving allocation hint. If the bounded hint walk +# is exhausted, allocation rescans the claims for the maximum and continues +# above it. Expired claims are reaped by an independently hourly-rate-limited +# sweep. seq is the worker's FIFO ordering key within a home, with the job id +# as the deterministic tiebreak. +# FIFO is defined over completed stagings: a stage that returns before another +# begins executes first; concurrently overlapping stagings have no relative +# ordering contract. +# The worker atomically claims a job with .claim, establishes its execution +# deadline, changes state to running, writes bounded stdout/stderr and exit, +# then publishes state=done last. Callers wait for done, relay stdout and +# stderr separately, then reap only their completed record. Input, argv, +# stdout, and stderr are each capped at 1048576 bytes. # -# The worker executes one job at a time, so a deliberately long-blocking poll -# would serialize every short interactive command behind its wait window. +# The worker serves one lane per staged home: jobs for the same home run +# strictly FIFO in seq order while lanes for different homes run concurrently, +# so one home's long job never delays another home's commands. Within a lane a +# deliberately long-blocking poll would still serialize that home's short +# interactive commands behind its wait window. # fm_remote_job_command_preemptible names the read-only long-poll class # (fm-remote-delta-read.sh, the reply-log delta read). The worker preempts a -# running preemptible job as soon as a non-preemptible job is queued and -# publishes exit 76 with emptied stdout and stderr, distinct from the poll's -# exit 75 elapsed-window-with-no-data result. The delta read is non-destructive -# and cursor-anchored, so the caller's normal re-arm re-reads the same data and -# a preempted poll loses nothing. +# running preemptible job as soon as a non-preemptible job is queued for the +# same home and publishes exit 76 with emptied stdout and stderr, distinct from +# the poll's exit 75 elapsed-window-with-no-data result. The delta read is +# non-destructive and cursor-anchored, so the caller's normal re-arm re-reads +# the same data and a preempted poll loses nothing. +# +# A caller that disconnects before its job completes cancels it instead of +# abandoning it: fm_remote_job_cancel writes a cancel marker into the record, +# the worker skips a cancelled queued job and terminates a running cancelled +# job's process group, and whichever side observes terminal publication reaps +# the finalized record because no result consumer remains. fm_remote_job_wait +# honors an optional FM_REMOTE_JOB_DISCONNECT_PROBE function name. When set, +# the probe runs about once per second; a failure cancels the job and fails +# the wait. The staging entrypoint arms it with a parent-liveness probe so an +# ssh channel +# that dies without delivering a signal still cancels the abandoned job. +# Abandoned .stage.* staging litter older than +# FM_REMOTE_JOB_STAGE_REAP_SECONDS is reaped by the worker's stale sweep. # # The worker accepts only a tracked, non-symlink executable named fm-*.sh below # its configured FM_ROOT/bin. Every child receives env -i with the composed @@ -60,12 +90,16 @@ FM_REMOTE_JOB_TIMEOUT=${FM_REMOTE_JOB_TIMEOUT:-360} FM_REMOTE_JOB_WAIT_GRACE=${FM_REMOTE_JOB_WAIT_GRACE:-30} FM_REMOTE_JOB_POLL_SECONDS=${FM_REMOTE_JOB_POLL_SECONDS:-0.05} FM_REMOTE_JOB_REAP_SECONDS=${FM_REMOTE_JOB_REAP_SECONDS:-3600} +FM_REMOTE_JOB_STAGE_REAP_SECONDS=${FM_REMOTE_JOB_STAGE_REAP_SECONDS:-600} +FM_REMOTE_JOB_SEQ_CLAIM_REAP_SECONDS=86400 +FM_REMOTE_JOB_SEQ_CLAIM_REAP_INTERVAL=3600 # shellcheck disable=SC2034 # Shared protocol constant consumed by the worker and sourcing callers. FM_REMOTE_JOB_PREEMPTED_EXIT=76 FM_REMOTE_JOB_OPERATOR_PATH= FM_REMOTE_JOB_CHILD_PATH= FM_REMOTE_JOB_STATE= FM_REMOTE_JOB_JOBS= +FM_REMOTE_JOB_SEQ_CLAIMS= FM_REMOTE_JOB_ID= FM_REMOTE_JOB_STDOUT= FM_REMOTE_JOB_STDERR= @@ -96,6 +130,7 @@ fm_remote_job_validate_settings() { case "$FM_REMOTE_JOB_WAIT_GRACE" in ''|*[!0-9]*) return 1 ;; esac [ "$FM_REMOTE_JOB_WAIT_GRACE" -le 300 ] || return 1 case "$FM_REMOTE_JOB_REAP_SECONDS" in ''|*[!0-9]*|0) return 1 ;; esac + case "$FM_REMOTE_JOB_STAGE_REAP_SECONDS" in ''|*[!0-9]*|0) return 1 ;; esac return 0 } @@ -398,6 +433,10 @@ fm_remote_job_prepare_state() { # FM_REMOTE_JOB_ERROR="remote job queue is unsafe" return 1 } + FM_REMOTE_JOB_SEQ_CLAIMS=$(fm_remote_job_safe_child_dir "$FM_REMOTE_JOB_STATE" .seq-claims) || { + FM_REMOTE_JOB_ERROR="remote job sequence claims are unsafe" + return 1 + } fm_remote_job_safe_child_dir "$FM_REMOTE_JOB_STATE" logs >/dev/null || { FM_REMOTE_JOB_ERROR="remote job log directory is unsafe" return 1 @@ -423,6 +462,20 @@ fm_remote_job_regular_bounded() { # [ "$bytes" -le "$max" ] } +fm_remote_job_remove_claim_records() { # + local claim=$1 file + [ -d "$claim" ] && [ ! -L "$claim" ] || return 1 + for file in "$claim"/owner "$claim"/owner_start "$claim"/supervisor \ + "$claim"/supervisor_start "$claim"/group "$claim"/group_start "$claim"/armed \ + "$claim"/.owner.* "$claim"/.owner_start.* "$claim"/.supervisor.* \ + "$claim"/.supervisor_start.* "$claim"/.group.* "$claim"/.group_start.* \ + "$claim"/.armed.*; do + [ -e "$file" ] || [ -L "$file" ] || continue + fm_remote_job_regular_bounded "$file" 256 || return 1 + rm -f -- "$file" || return 1 + done +} + fm_remote_job_write_state() { # queued|running|done local job=$1 value=$2 tmp case "$value" in queued|running|done) ;; *) return 1 ;; esac @@ -444,9 +497,9 @@ fm_remote_job_read_state() { # case "$value" in queued|running|'done') printf '%s\n' "$value" ;; *) return 1 ;; esac } -fm_remote_job_read_number() { # queue_deadline|timeout|deadline +fm_remote_job_read_number() { # queue_deadline|timeout|deadline|seq local job=$1 field=$2 value - case "$field" in queue_deadline|timeout|deadline) ;; *) return 1 ;; esac + case "$field" in queue_deadline|timeout|deadline|seq) ;; *) return 1 ;; esac fm_remote_job_regular_bounded "$job/$field" 32 || return 1 value=$(tr -d '\n' < "$job/$field") case "$value" in ''|*[!0-9]*) return 1 ;; esac @@ -454,9 +507,9 @@ fm_remote_job_read_number() { # queue_deadline|timeout|deadline printf '%s\n' "$value" } -fm_remote_job_write_number() { # queue_deadline|timeout|deadline +fm_remote_job_write_number() { # queue_deadline|timeout|deadline|seq local job=$1 field=$2 value=$3 tmp - case "$field" in queue_deadline|timeout|deadline) ;; *) return 1 ;; esac + case "$field" in queue_deadline|timeout|deadline|seq) ;; *) return 1 ;; esac case "$value" in ''|*[!0-9]*|0) return 1 ;; esac [ -d "$job" ] && [ ! -L "$job" ] || return 1 tmp=$(umask 077; mktemp "$job/.$field.XXXXXX") || return 1 @@ -469,8 +522,96 @@ fm_remote_job_read_deadline() { # fm_remote_job_read_number "$1" deadline } +fm_remote_job_advance_seq_hint() { # + local value=$1 counter current tmp + counter="$FM_REMOTE_JOB_STATE/seq" + current=$(cat "$counter" 2>/dev/null || true) + case "$current" in ''|*[!0-9]*) current=0 ;; esac + [ "$value" -gt "$current" ] || return 0 + tmp=$(umask 077; mktemp "$FM_REMOTE_JOB_STATE/.seqhint.XXXXXX") || return 1 + printf '%s\n' "$value" > "$tmp" || { rm -f -- "$tmp"; return 1; } + chmod 600 "$tmp" || { rm -f -- "$tmp"; return 1; } + current=$(cat "$counter" 2>/dev/null || true) + case "$current" in ''|*[!0-9]*) current=0 ;; esac + if [ "$value" -gt "$current" ]; then + mv -f -- "$tmp" "$counter" || { rm -f -- "$tmp"; return 1; } + else + rm -f -- "$tmp" + fi +} + +fm_remote_job_next_seq() { # [stage-dir destination] + local stage=${1:-} destination=${2:-} counter value claim attempt=0 recovered=0 maximum entry + [ -n "$FM_REMOTE_JOB_STATE" ] && [ -n "$FM_REMOTE_JOB_SEQ_CLAIMS" ] || return 1 + counter="$FM_REMOTE_JOB_STATE/seq" + value=$(cat "$counter" 2>/dev/null || true) + case "$value" in ''|*[!0-9]*) value=0 ;; esac + while :; do + if [ "$attempt" -ge 100000 ]; then + [ "$recovered" -eq 0 ] || return 1 + maximum=0 + for entry in "$FM_REMOTE_JOB_SEQ_CLAIMS"/*; do + [ -d "$entry" ] && [ ! -L "$entry" ] || continue + entry=${entry##*/} + case "$entry" in ''|*[!0-9]*|0) continue ;; esac + [ "$entry" -le "$maximum" ] || maximum=$entry + done + value=$maximum + attempt=0 + recovered=1 + fi + attempt=$((attempt + 1)) + value=$((value + 1)) + claim="$FM_REMOTE_JOB_SEQ_CLAIMS/$value" + if (umask 077; mkdir "$claim") 2>/dev/null; then + chmod 700 "$claim" || return 1 + fm_remote_job_advance_seq_hint "$value" || true + if [ -n "$stage" ]; then + if ! fm_remote_job_write_number "$stage" seq "$value" \ + || ! fm_remote_job_write_state "$stage" queued \ + || ! mv -- "$stage" "$destination"; then + rm -f -- "$stage/state" "$stage/seq" + return 1 + fi + rm -f -- "$destination/.owner-pid" "$destination/.owner-start" || true + fi + printf '%s\n' "$value" + return 0 + fi + [ -d "$claim" ] && [ ! -L "$claim" ] || return 1 + done +} + +fm_remote_job_cancelled() { # + [ -f "$1/cancel" ] && [ ! -L "$1/cancel" ] +} + +# Mark a job cancelled on behalf of a disconnected or abandoning caller. The +# marker never rewrites state: the worker observes it, skips a cancelled queued +# job, and stops a running cancelled job's process group. The worker reaps after +# terminal publication; if publication already won the race, this function +# reaps instead. Cancelling a job that disappeared is a harmless no-op. +fm_remote_job_cancel() { # + local account_home=$1 id=$2 job state tmp + fm_remote_job_prepare_state "$account_home" || return 1 + job=$(fm_remote_job_job_dir "$id" 2>/dev/null) || return 0 + state=$(fm_remote_job_read_state "$job" 2>/dev/null || true) + if [ "$state" = 'done' ]; then + fm_remote_job_reap "$account_home" "$id" 2>/dev/null || true + return 0 + fi + tmp=$(umask 077; mktemp "$job/.cancel.XXXXXX") || return 1 + printf 'cancelled: caller disconnected or abandoned the job\n' > "$tmp" || { rm -f -- "$tmp"; return 1; } + chmod 600 "$tmp" || { rm -f -- "$tmp"; return 1; } + mv -f -- "$tmp" "$job/cancel" || return 1 + state=$(fm_remote_job_read_state "$job" 2>/dev/null || true) + if [ "$state" = 'done' ]; then + fm_remote_job_reap "$account_home" "$id" 2>/dev/null || true + fi +} + fm_remote_job_stage() { # [args...]; stdin is captured - local account_home=$1 root=$2 home=$3 command=$4 stage id destination bytes queue_deadline + local account_home=$1 root=$2 home=$3 command=$4 stage id destination bytes queue_deadline owner_start shift 4 fm_remote_job_prepare_state "$account_home" || return 1 root=$(fm_remote_job_canonical_existing_dir "$root") || { @@ -483,13 +624,20 @@ fm_remote_job_stage() { # [args...]; stdi } case "$command" in fm-*.sh) ;; *) FM_REMOTE_JOB_ERROR="remote job command is outside the fm-*.sh namespace"; return 1 ;; esac case "$command" in */*|*..*) FM_REMOTE_JOB_ERROR="remote job command contains a path or traversal"; return 1 ;; esac + owner_start=$(fm_remote_job_process_start "$$") || { + FM_REMOTE_JOB_ERROR="cannot establish remote job staging ownership" + return 1 + } stage=$(umask 077; mktemp -d "$FM_REMOTE_JOB_JOBS/.stage.XXXXXX") || { FM_REMOTE_JOB_ERROR="cannot stage remote job" return 1 } chmod 700 "$stage" || { rm -rf -- "$stage"; return 1; } queue_deadline=$(( $(date +%s) + FM_REMOTE_JOB_QUEUE_TIMEOUT )) - if ! printf '%s\n' "$root" > "$stage/root" || + if ! printf '%s\n' "$$" > "$stage/.owner-pid" || + ! printf '%s\n' "$owner_start" > "$stage/.owner-start" || + ! chmod 600 "$stage/.owner-pid" "$stage/.owner-start" || + ! printf '%s\n' "$root" > "$stage/root" || ! printf '%s\n' "$home" > "$stage/home" || ! printf '%s\n' "$queue_deadline" > "$stage/queue_deadline" || ! printf '%s\n' "$FM_REMOTE_JOB_TIMEOUT" > "$stage/timeout" || @@ -513,19 +661,23 @@ fm_remote_job_stage() { # [args...]; stdi : > "$stage/stdout" : > "$stage/stderr" chmod 600 "$stage/stdout" "$stage/stderr" || { rm -rf -- "$stage"; return 1; } - fm_remote_job_write_state "$stage" queued || { rm -rf -- "$stage"; return 1; } id="job-${stage##*/.stage.}" fm_remote_job_safe_id "$id" || { rm -rf -- "$stage"; return 1; } destination="$FM_REMOTE_JOB_JOBS/$id" [ ! -e "$destination" ] && [ ! -L "$destination" ] || { rm -rf -- "$stage"; return 1; } - mv -- "$stage" "$destination" || { rm -rf -- "$stage"; return 1; } + if ! fm_remote_job_next_seq "$stage" "$destination" >/dev/null; then + rm -rf -- "$stage" + FM_REMOTE_JOB_ERROR="cannot allocate and publish a remote job staging sequence" + return 1 + fi # shellcheck disable=SC2034 # Sourceable API consumed by callers that do not use command substitution. FM_REMOTE_JOB_ID=$id printf '%s\n' "$id" } -fm_remote_job_wait() { # +fm_remote_job_wait() { # ; honors FM_REMOTE_JOB_DISCONNECT_PROBE local account_home=$1 id=$2 job state queue_deadline execution_timeout wait_deadline exit_value + local now next_probe=0 fm_remote_job_prepare_state "$account_home" || return 1 job=$(fm_remote_job_job_dir "$id") || { FM_REMOTE_JOB_ERROR="remote job record disappeared or became unsafe" @@ -568,10 +720,19 @@ fm_remote_job_wait() { # queued|running) ;; *) FM_REMOTE_JOB_ERROR="remote job state is invalid"; return 1 ;; esac - if [ "$(date +%s)" -ge "$wait_deadline" ]; then + now=$(date +%s) + if [ "$now" -ge "$wait_deadline" ]; then FM_REMOTE_JOB_ERROR="remote job did not complete within its bounded wait" return 1 fi + if [ -n "${FM_REMOTE_JOB_DISCONNECT_PROBE:-}" ] && [ "$now" -ge "$next_probe" ]; then + next_probe=$((now + 1)) + if ! "$FM_REMOTE_JOB_DISCONNECT_PROBE"; then + fm_remote_job_cancel "$account_home" "$id" 2>/dev/null || true + FM_REMOTE_JOB_ERROR="remote job caller disconnected; the job was cancelled" + return 1 + fi + fi sleep "$FM_REMOTE_JOB_POLL_SECONDS" done } @@ -581,14 +742,14 @@ fm_remote_job_reap() { # ; only removes an exact completed re fm_remote_job_prepare_state "$account_home" || return 1 job=$(fm_remote_job_job_dir "$id") || return 1 [ "$(fm_remote_job_read_state "$job")" = 'done' ] || return 1 - for file in root home queue_deadline timeout deadline argv stdin stdout stderr exit state; do + for file in root home queue_deadline timeout deadline seq cancel argv stdin stdout stderr exit state .owner-pid .owner-start; do [ -e "$job/$file" ] || continue [ ! -L "$job/$file" ] || return 1 rm -f -- "$job/$file" || return 1 done if [ -e "$job/.claim" ] || [ -L "$job/.claim" ]; then [ -d "$job/.claim" ] && [ ! -L "$job/.claim" ] || return 1 - rm -f -- "$job/.claim/owner" "$job/.claim/supervisor" "$job/.claim/group" "$job/.claim/armed" || return 1 + fm_remote_job_remove_claim_records "$job/.claim" || return 1 rmdir "$job/.claim" || return 1 fi rmdir "$job" @@ -600,8 +761,18 @@ fm_remote_job_path_mtime() { # if [ "$(uname -s 2>/dev/null || true)" = Darwin ]; then stat -f %m "$1" 2>/dev/null; else stat -c %Y "$1" 2>/dev/null; fi } +fm_remote_job_stage_owner_alive() { # + local stage=$1 pid recorded_start actual_start + pid=$(fm_remote_job_read_single_line "$stage/.owner-pid" 64 2>/dev/null) || return 1 + case "$pid" in ''|*[!0-9]*) return 1 ;; esac + [ "$pid" -gt 1 ] || return 1 + recorded_start=$(fm_remote_job_read_single_line "$stage/.owner-start" 256 2>/dev/null) || return 1 + actual_start=$(fm_remote_job_process_start "$pid" 2>/dev/null) || return 1 + [ "$recorded_start" = "$actual_start" ] +} + fm_remote_job_reap_stale() { # - local account_home=$1 job id state mtime now + local account_home=$1 job id state mtime now stage claim value marker tmp reap_claims=0 fm_remote_job_prepare_state "$account_home" || return 1 now=$(date +%s) for job in "$FM_REMOTE_JOB_JOBS"/job-*; do @@ -615,6 +786,39 @@ fm_remote_job_reap_stale() { # [ $((now - mtime)) -ge "$FM_REMOTE_JOB_REAP_SECONDS" ] || continue fm_remote_job_reap "$account_home" "$id" || true done + marker="$FM_REMOTE_JOB_STATE/.seq-claims-reaped" + mtime=$(fm_remote_job_path_mtime "$marker" 2>/dev/null || true) + case "$mtime" in + ''|*[!0-9]*) reap_claims=1 ;; + *) [ $((now - mtime)) -lt "$FM_REMOTE_JOB_SEQ_CLAIM_REAP_INTERVAL" ] || reap_claims=1 ;; + esac + if [ "$reap_claims" -eq 1 ]; then + tmp=$(umask 077; mktemp "$FM_REMOTE_JOB_STATE/.seqreap.XXXXXX") || tmp= + if [ -n "$tmp" ] && printf '%s\n' "$now" > "$tmp" && chmod 600 "$tmp" \ + && mv -f -- "$tmp" "$marker"; then + for claim in "$FM_REMOTE_JOB_SEQ_CLAIMS"/*; do + [ -d "$claim" ] && [ ! -L "$claim" ] || continue + value=${claim##*/} + case "$value" in ''|*[!0-9]*|0) continue ;; esac + mtime=$(fm_remote_job_path_mtime "$claim" 2>/dev/null || true) + case "$mtime" in ''|*[!0-9]*) continue ;; esac + [ $((now - mtime)) -ge "$FM_REMOTE_JOB_SEQ_CLAIM_REAP_SECONDS" ] || continue + rmdir "$claim" 2>/dev/null || true + done + else + [ -z "$tmp" ] || rm -f -- "$tmp" + fi + fi + # Staging litter a killed caller left behind is reaped after its owner is no + # longer the process that created it and the stage has exceeded the age bound. + for stage in "$FM_REMOTE_JOB_JOBS"/.stage.*; do + [ -d "$stage" ] && [ ! -L "$stage" ] || continue + fm_remote_job_stage_owner_alive "$stage" && continue + mtime=$(fm_remote_job_path_mtime "$stage" 2>/dev/null || true) + case "$mtime" in ''|*[!0-9]*) continue ;; esac + [ $((now - mtime)) -ge "$FM_REMOTE_JOB_STAGE_REAP_SECONDS" ] || continue + rm -rf -- "$stage" + done } fm_remote_job_launchagent_paths() { # diff --git a/bin/fm-remote-job-worker.sh b/bin/fm-remote-job-worker.sh index 2d7528a427f..14598eb7670 100755 --- a/bin/fm-remote-job-worker.sh +++ b/bin/fm-remote-job-worker.sh @@ -15,6 +15,13 @@ # have been committed. The library header owns the exact record fields and # lifecycle. # +# The shared library header owns lane selection, FIFO, and caller-cancellation +# contracts. This serving loop implements each active lane as a tracked, +# top-level --lane process that claims one job, records itself as the claim's +# supervisor, and runs it to publication. Shutdown stops every tracked lane and +# its recorded command group, leaving interrupted records for the replacement +# worker's orphan recovery. +# # The worker is abandoned when its configured FM_ROOT stops being a genuine # Firstmate checkout - the state a pruned no-mistakes gate worktree, a returned # pooled worktree, or a removed test fixture root leaves behind. It can never @@ -50,13 +57,17 @@ FM_ROOT=${FM_ROOT_OVERRIDE:-$(CDPATH='' cd "$SCRIPT_DIR/.." && pwd -P)} # shellcheck source=bin/fm-remote-job-lib.sh . "$SCRIPT_DIR/fm-remote-job-lib.sh" -WORKER_ACTIVE_JOB= WORKER_LOCK= WORKER_LOCK_HELD=0 WORKER_RELEASE_OWNERSHIP=1 WORKER_SUPERVISED_PID= WORKER_PREEMPTIBLE=0 WORKER_PREEMPTED=0 +WORKER_LANE_HOME= +WORKER_LANE_HOMES=() +WORKER_LANE_PIDS=() +WORKER_LANE_STARTS=() +WORKER_LANE_JOBS=() worker_error() { printf 'remote-job-worker: %s\n' "$1" >&2; } @@ -135,7 +146,7 @@ worker_quarantined_execution_stopped() { # [ ! -e "$file" ] && [ ! -L "$file" ] && continue [ ! -L "$file" ] || return 1 pid=$(worker_read_process_id "$file") || return 1 - worker_process_or_group_alive "$kind" "$pid" && return 1 + worker_recorded_execution_alive "$job" "$kind" "$pid" && return 1 done done } @@ -248,6 +259,69 @@ worker_signal_process_or_group() { # process|group esac } +worker_supervisor_identity_status() { # + local job=$1 pid=$2 recorded_start actual_start + recorded_start=$(fm_remote_job_read_single_line "$job/.claim/supervisor_start" 256 2>/dev/null) || return 2 + actual_start=$(fm_remote_job_process_start "$pid" 2>/dev/null) || { + worker_process_or_group_alive process "$pid" && return 2 + return 1 + } + [ "$recorded_start" = "$actual_start" ] && return 0 + return 1 +} + +# A leaderless live group still belongs to the recorded execution: its PGID +# cannot be reused while any old member survives, so it remains safe to signal. +# A live leader whose start identity mismatches proves PID reuse and makes the +# recorded group stale; an unreadable live leader stays indeterminate so the +# stop loop retries rather than signaling or declaring the group dead. +worker_group_identity_status() { # + local job=$1 pid=$2 recorded_start actual_start file="$1/.claim/group_start" + [ -e "$file" ] || [ -L "$file" ] || return 3 + recorded_start=$(fm_remote_job_read_single_line "$file" 256 2>/dev/null) || return 2 + actual_start=$(fm_remote_job_process_start "$pid" 2>/dev/null) || { + kill -0 "$pid" 2>/dev/null && return 2 + worker_process_or_group_alive group "$pid" && return 0 + return 1 + } + [ "$recorded_start" = "$actual_start" ] && return 0 + return 1 +} + +worker_recorded_execution_alive() { # process|group + local job=$1 kind=$2 pid=$3 identity_status + if [ "$kind" = process ]; then + worker_supervisor_identity_status "$job" "$pid" + identity_status=$? + case "$identity_status" in + 0) ;; + 1) return 1 ;; + 2) worker_process_or_group_alive process "$pid"; return ;; + esac + else + worker_group_identity_status "$job" "$pid" + identity_status=$? + case "$identity_status" in + 0|3) ;; + 1) return 1 ;; + 2) worker_process_or_group_alive group "$pid"; return ;; + esac + fi + worker_process_or_group_alive "$kind" "$pid" +} + +worker_signal_recorded_execution() { # process|group + local job=$1 kind=$2 signal=$3 pid=$4 identity_status + if [ "$kind" = process ]; then + worker_supervisor_identity_status "$job" "$pid" || return 0 + else + worker_group_identity_status "$job" "$pid" + identity_status=$? + case "$identity_status" in 0|3) ;; *) return 0 ;; esac + fi + worker_signal_process_or_group "$kind" "$signal" "$pid" +} + worker_stop_recorded_execution() { # local job=$1 kind file pid attempt still_alive for kind in process group; do @@ -255,8 +329,8 @@ worker_stop_recorded_execution() { # [ ! -e "$file" ] && [ ! -L "$file" ] && continue [ ! -L "$file" ] || return 1 pid=$(worker_read_process_id "$file") || return 1 - worker_signal_process_or_group "$kind" TERM "$pid" - worker_signal_process_or_group "$kind" KILL "$pid" + worker_signal_recorded_execution "$job" "$kind" TERM "$pid" + worker_signal_recorded_execution "$job" "$kind" KILL "$pid" wait "$pid" 2>/dev/null || true done attempt=0 @@ -267,31 +341,51 @@ worker_stop_recorded_execution() { # case "$kind" in process) file="$job/.claim/supervisor" ;; group) file="$job/.claim/group" ;; esac [ -e "$file" ] || continue pid=$(worker_read_process_id "$file") || return 1 - worker_process_or_group_alive "$kind" "$pid" && still_alive=1 + if worker_recorded_execution_alive "$job" "$kind" "$pid"; then + still_alive=1 + worker_signal_recorded_execution "$job" "$kind" TERM "$pid" + worker_signal_recorded_execution "$job" "$kind" KILL "$pid" + fi done [ "$still_alive" -eq 1 ] || break sleep 0.01 done [ "$still_alive" -eq 0 ] || return 1 - rm -f -- "$job/.claim/supervisor" "$job/.claim/group" "$job/.claim/armed" + rm -f -- "$job/.claim/supervisor" "$job/.claim/supervisor_start" \ + "$job/.claim/group" "$job/.claim/group_start" "$job/.claim/armed" +} + +# Stop every tracked lane process and its recorded command execution. The lane +# is signalled first so it cannot dispatch further work, then the job's +# recorded supervisor and group are verified stopped; a job interrupted here +# stays running-with-a-dead-owner for the replacement worker's orphan recovery, +# exactly as a crashed single-process worker's job did. +worker_lane_identity_matches() { # + local pid=$1 start=$2 actual_start + [ -n "$start" ] || return 1 + actual_start=$(fm_remote_job_process_start "$pid" 2>/dev/null) || return 1 + [ "$actual_start" = "$start" ] } worker_stop_active_execution() { - local job=${WORKER_ACTIVE_JOB:-} owner owner_pid state - if [ -n "$job" ]; then - worker_stop_recorded_execution "$job" || return 1 - else - for job in "$FM_REMOTE_JOB_JOBS"/job-*; do - [ -d "$job" ] && [ ! -L "$job" ] || continue - state=$(fm_remote_job_read_state "$job" 2>/dev/null || true) - [ "$state" = running ] || continue - owner="$job/.claim/owner" - owner_pid=$(worker_read_process_id "$owner" 2>/dev/null || true) - [ "$owner_pid" = "${BASHPID:-$$}" ] || continue - worker_stop_recorded_execution "$job" || return 1 - done - fi - WORKER_ACTIVE_JOB= + local i=0 count=${#WORKER_LANE_PIDS[@]} job pid start failed=0 + while [ "$i" -lt "$count" ]; do + pid=${WORKER_LANE_PIDS[$i]} + start=${WORKER_LANE_STARTS[$i]} + job=${WORKER_LANE_JOBS[$i]} + if worker_lane_identity_matches "$pid" "$start"; then kill -TERM "$pid" 2>/dev/null || true; fi + if worker_lane_identity_matches "$pid" "$start"; then kill -KILL "$pid" 2>/dev/null || true; fi + wait "$pid" 2>/dev/null || true + if [ -d "$job" ] && [ ! -L "$job" ]; then + worker_stop_recorded_execution "$job" || failed=1 + fi + i=$((i + 1)) + done + WORKER_LANE_HOMES=() + WORKER_LANE_PIDS=() + WORKER_LANE_STARTS=() + WORKER_LANE_JOBS=() + [ "$failed" -eq 0 ] } # Ignore, rather than restore the default disposition for, the signals this @@ -333,21 +427,40 @@ worker_exit_cleanup() { } worker_claim() { # - local job=$1 claim + local job=$1 claim pid start pid_tmp start_tmp claim="$job/.claim" [ ! -e "$claim" ] && [ ! -L "$claim" ] || return 1 (umask 077; mkdir "$claim") || return 1 - printf '%s\n' "${BASHPID:-$$}" > "$claim/owner" || { rmdir "$claim" 2>/dev/null || true; return 1; } - chmod 600 "$claim/owner" || { rm -f -- "$claim/owner"; rmdir "$claim" 2>/dev/null || true; return 1; } + pid=${BASHPID:-$$} + start=$(fm_remote_job_process_start "$pid") || { rmdir "$claim" 2>/dev/null || true; return 1; } + pid_tmp=$(umask 077; mktemp "$claim/.owner.XXXXXX") || { rmdir "$claim" 2>/dev/null || true; return 1; } + start_tmp=$(umask 077; mktemp "$claim/.owner_start.XXXXXX") || { + rm -f -- "$pid_tmp" + rmdir "$claim" 2>/dev/null || true + return 1 + } + if ! printf '%s\n' "$pid" > "$pid_tmp" || ! printf '%s\n' "$start" > "$start_tmp" \ + || ! chmod 600 "$pid_tmp" "$start_tmp" || ! mv -f -- "$start_tmp" "$claim/owner_start" \ + || ! mv -f -- "$pid_tmp" "$claim/owner"; then + rm -f -- "$pid_tmp" "$start_tmp" "$claim/owner" "$claim/owner_start" + rmdir "$claim" 2>/dev/null || true + return 1 + fi } worker_claim_owner_alive() { # - local job=$1 claim="$1/.claim" owner pid + local job=$1 claim="$1/.claim" owner pid recorded_start actual_start [ -d "$claim" ] && [ ! -L "$claim" ] || return 1 owner="$claim/owner" fm_remote_job_regular_bounded "$owner" 64 || return 1 pid=$(tr -d '\n' < "$owner") case "$pid" in ''|*[!0-9]*) return 1 ;; esac + if [ -e "$claim/owner_start" ] || [ -L "$claim/owner_start" ]; then + recorded_start=$(fm_remote_job_read_single_line "$claim/owner_start" 256 2>/dev/null) || return 1 + actual_start=$(fm_remote_job_process_start "$pid" 2>/dev/null) || return 1 + [ "$recorded_start" = "$actual_start" ] + return + fi kill -0 "$pid" 2>/dev/null } @@ -357,15 +470,23 @@ worker_clear_dead_claim() { # worker_claim_owner_alive "$job" && return 1 [ -d "$claim" ] && [ ! -L "$claim" ] || return 1 [ ! -e "$claim/owner" ] || [ ! -L "$claim/owner" ] || return 1 - rm -f -- "$claim/owner" "$claim/supervisor" "$claim/group" "$claim/armed" || return 1 + fm_remote_job_remove_claim_records "$claim" || return 1 rmdir "$claim" } -worker_recover_orphaned_job() { # - local job=$1 file - worker_claim_owner_alive "$job" && return 1 +# Reclaim a running job this serving loop does not own: a record left by a +# crashed worker, whether its lane process died with it or survived it. The +# recorded execution is stopped either way - a surviving foreign lane is not +# supervised by any owner and a second lane for its home must never start +# beside it - and the record publishes unknown completion, exactly as a +# crashed single-process worker's job always has. +worker_reclaim_running_job() { # + local job=$1 file state worker_stop_recorded_execution "$job" || return 1 + state=$(fm_remote_job_read_state "$job" 2>/dev/null) || return 1 worker_clear_dead_claim "$job" || return 1 + [ "$state" = 'done' ] && return 0 + [ "$state" = running ] || return 1 for file in .stdout.pipe .stderr.pipe; do [ ! -e "$job/$file" ] && [ ! -L "$job/$file" ] || { [ ! -L "$job/$file" ] || return 1 @@ -394,7 +515,7 @@ worker_read_text() { # } worker_publish_result() { # - local job=$1 exit_status=$2 tmp + local job=$1 exit_status=$2 tmp account_home case "$exit_status" in ''|*[!0-9]*) exit_status=125 ;; esac [ "$exit_status" -le 255 ] || exit_status=125 for tmp in stdout stderr; do @@ -404,17 +525,23 @@ worker_publish_result() { # printf '%s\n' "$exit_status" > "$tmp" || { rm -f -- "$tmp"; return 1; } chmod 600 "$tmp" || { rm -f -- "$tmp"; return 1; } mv -f -- "$tmp" "$job/exit" || { rm -f -- "$tmp"; return 1; } - fm_remote_job_write_state "$job" 'done' + fm_remote_job_write_state "$job" 'done' || return 1 + if fm_remote_job_cancelled "$job"; then + account_home=$(worker_account_home 2>/dev/null || true) + if [ -n "$account_home" ]; then + fm_remote_job_reap "$account_home" "${job##*/}" 2>/dev/null || true + fi + fi } worker_run_with_timeout() { # [args...] - local job=$1 timeout=$2 group_file armed_file group_pid rc tmp deadline next_heartbeat attempt - local timed_out=0 heartbeat_failed=0 + local job=$1 timeout=$2 group_file group_start_file armed_file group_pid group_start + local group_tmp group_start_tmp rc tmp deadline next_check attempt timed_out=0 cancelled=0 WORKER_PREEMPTED=0 shift 2 group_file="$job/.claim/group" + group_start_file="$job/.claim/group_start" armed_file="$job/.claim/armed" - WORKER_ACTIVE_JOB=$job set -m ( while [ ! -f "$armed_file" ] || [ -L "$armed_file" ]; do @@ -425,43 +552,47 @@ worker_run_with_timeout() { # [args...] ) & group_pid=$! set +m - tmp=$(umask 077; mktemp "$job/.claim/.group.XXXXXX") || { + group_start=$(fm_remote_job_process_start "$group_pid") || { worker_signal_process_or_group group KILL "$group_pid" wait "$group_pid" 2>/dev/null || true - WORKER_ACTIVE_JOB= return 125 } - printf '%s\n' "$group_pid" > "$tmp" || { - rm -f -- "$tmp" + group_tmp=$(umask 077; mktemp "$job/.claim/.group.XXXXXX") || { worker_signal_process_or_group group KILL "$group_pid" wait "$group_pid" 2>/dev/null || true - WORKER_ACTIVE_JOB= return 125 } - if ! chmod 600 "$tmp" || ! mv -f -- "$tmp" "$group_file"; then - rm -f -- "$tmp" + group_start_tmp=$(umask 077; mktemp "$job/.claim/.group_start.XXXXXX") || { + rm -f -- "$group_tmp" + worker_signal_process_or_group group KILL "$group_pid" + wait "$group_pid" 2>/dev/null || true + return 125 + } + if ! printf '%s\n' "$group_pid" > "$group_tmp" \ + || ! printf '%s\n' "$group_start" > "$group_start_tmp" \ + || ! chmod 600 "$group_tmp" "$group_start_tmp" \ + || ! mv -f -- "$group_start_tmp" "$group_start_file" \ + || ! mv -f -- "$group_tmp" "$group_file"; then + rm -f -- "$group_tmp" "$group_start_tmp" "$group_file" "$group_start_file" worker_signal_process_or_group group KILL "$group_pid" wait "$group_pid" 2>/dev/null || true - WORKER_ACTIVE_JOB= return 125 fi tmp=$(umask 077; mktemp "$job/.claim/.armed.XXXXXX") || { worker_signal_process_or_group group KILL "$group_pid" wait "$group_pid" 2>/dev/null || true - rm -f -- "$group_file" - WORKER_ACTIVE_JOB= + rm -f -- "$group_file" "$group_start_file" return 125 } if ! chmod 600 "$tmp" || ! mv -f -- "$tmp" "$armed_file"; then rm -f -- "$tmp" worker_signal_process_or_group group KILL "$group_pid" wait "$group_pid" 2>/dev/null || true - rm -f -- "$group_file" - WORKER_ACTIVE_JOB= + rm -f -- "$group_file" "$group_start_file" return 125 fi deadline=$((SECONDS + timeout)) - next_heartbeat=$((SECONDS + 1)) + next_check=$((SECONDS + 1)) while worker_process_or_group_alive group "$group_pid"; do if [ "$SECONDS" -ge "$deadline" ]; then worker_signal_process_or_group group TERM "$group_pid" @@ -469,14 +600,19 @@ worker_run_with_timeout() { # [args...] timed_out=1 break fi - if [ "$SECONDS" -ge "$next_heartbeat" ]; then - if ! worker_write_heartbeat; then + if [ "$SECONDS" -ge "$next_check" ]; then + if fm_remote_job_cancelled "$job"; then worker_signal_process_or_group group TERM "$group_pid" + attempt=0 + while worker_process_or_group_alive group "$group_pid" && [ "$attempt" -lt 20 ]; do + attempt=$((attempt + 1)) + sleep 0.05 + done worker_signal_process_or_group group KILL "$group_pid" - heartbeat_failed=1 + cancelled=1 break fi - if [ "$WORKER_PREEMPTIBLE" -eq 1 ] && worker_preempting_waiter_exists; then + if [ "$WORKER_PREEMPTIBLE" -eq 1 ] && worker_preempting_waiter_exists "$WORKER_LANE_HOME"; then worker_signal_process_or_group group TERM "$group_pid" attempt=0 while worker_process_or_group_alive group "$group_pid" && [ "$attempt" -lt 20 ]; do @@ -487,16 +623,15 @@ worker_run_with_timeout() { # [args...] WORKER_PREEMPTED=1 break fi - next_heartbeat=$((SECONDS + 1)) + next_check=$((SECONDS + 1)) fi sleep "$FM_REMOTE_JOB_POLL_SECONDS" done wait "$group_pid" 2>/dev/null rc=$? - rm -f -- "$group_file" "$armed_file" - WORKER_ACTIVE_JOB= + rm -f -- "$group_file" "$group_start_file" "$armed_file" [ "$timed_out" -eq 0 ] || return 124 - [ "$heartbeat_failed" -eq 0 ] || return 125 + [ "$cancelled" -eq 0 ] || return 130 [ "$WORKER_PREEMPTED" -eq 0 ] || return "$FM_REMOTE_JOB_PREEMPTED_EXIT" return "$rc" } @@ -508,12 +643,17 @@ worker_job_command() { # ; the first argv element of a staged record printf '%s\n' "$first" } -worker_preempting_waiter_exists() { - local job state command +worker_preempting_waiter_exists() { # + local lane_home=$1 job state command job_home for job in "$FM_REMOTE_JOB_JOBS"/job-*; do [ -d "$job" ] && [ ! -L "$job" ] || continue state=$(fm_remote_job_read_state "$job" 2>/dev/null || true) [ "$state" = queued ] || continue + fm_remote_job_cancelled "$job" && continue + # Lanes are per home, so only a waiter for this lane's own home may + # preempt; another home's queue drains through its own lane. + job_home=$(worker_read_text "$job" home 8192 2>/dev/null || true) + [ "$job_home" = "$lane_home" ] || continue command=$(worker_job_command "$job" 2>/dev/null || true) fm_remote_job_command_preemptible "$command" || return 0 done @@ -544,6 +684,9 @@ worker_run_job() { # home=$(worker_read_text "$job" home 8192) || { worker_publish_result "$job" 126; return; } root=$(fm_remote_job_canonical_existing_dir "$root") || { worker_publish_result "$job" 126; return; } home=$(fm_remote_job_canonical_home "$home") || { worker_publish_result "$job" 126; return; } + # The lane key everywhere - dispatch and the preemption scan - is the staged + # home field's exact text, so this comparison value is read the same way. + WORKER_LANE_HOME=$(worker_read_text "$job" home 8192 2>/dev/null || true) [ "$root" = "$FM_ROOT" ] || { worker_publish_result "$job" 126; return; } [ -f "$root/AGENTS.md" ] && [ ! -L "$root/AGENTS.md" ] && [ -d "$root/bin" ] && [ ! -L "$root/bin" ] || { worker_publish_result "$job" 126; return; } @@ -634,8 +777,162 @@ worker_run_job() { # worker_publish_result "$job" "$rc" || worker_error "could not publish result for ${job##*/}" } +# Finalize a cancelled record nobody waits on: publish the interrupt result so +# the record is complete, then reap it because its caller is gone. +worker_finalize_cancelled() { # + local account_home=$1 job=$2 + : > "$job/stdout" 2>/dev/null || true + printf 'remote job cancelled after its caller disconnected\n' > "$job/stderr" 2>/dev/null || true + worker_publish_result "$job" 130 || return 1 + fm_remote_job_reap "$account_home" "${job##*/}" || true +} + +worker_lane_busy() { # + local home=$1 i=0 count=${#WORKER_LANE_HOMES[@]} + while [ "$i" -lt "$count" ]; do + [ "${WORKER_LANE_HOMES[$i]}" != "$home" ] || return 0 + i=$((i + 1)) + done + return 1 +} + +worker_lane_owns_job() { # + local job=$1 i=0 count=${#WORKER_LANE_JOBS[@]} + while [ "$i" -lt "$count" ]; do + [ "${WORKER_LANE_JOBS[$i]}" != "$job" ] || return 0 + i=$((i + 1)) + done + return 1 +} + +worker_reap_finished_lanes() { + local i=0 count=${#WORKER_LANE_PIDS[@]} pid start + local live_homes=() live_pids=() live_starts=() live_jobs=() + while [ "$i" -lt "$count" ]; do + pid=${WORKER_LANE_PIDS[$i]} + start=${WORKER_LANE_STARTS[$i]} + if worker_lane_identity_matches "$pid" "$start"; then + live_homes+=("${WORKER_LANE_HOMES[$i]}") + live_pids+=("$pid") + live_starts+=("$start") + live_jobs+=("${WORKER_LANE_JOBS[$i]}") + else + wait "$pid" 2>/dev/null || true + fi + i=$((i + 1)) + done + WORKER_LANE_HOMES=() + WORKER_LANE_PIDS=() + WORKER_LANE_STARTS=() + WORKER_LANE_JOBS=() + i=0 + count=${#live_pids[@]} + while [ "$i" -lt "$count" ]; do + WORKER_LANE_HOMES+=("${live_homes[$i]}") + WORKER_LANE_PIDS+=("${live_pids[$i]}") + WORKER_LANE_STARTS+=("${live_starts[$i]}") + WORKER_LANE_JOBS+=("${live_jobs[$i]}") + i=$((i + 1)) + done +} + +# One lane's whole execution of one job, run as a background lane process: +# claim, record this process as the claim supervisor, honor a cancel that +# arrived before running, establish the deadline, run to publication, and reap +# the record when its caller cancelled and can no longer reap it. +worker_lane_execute() { # + local account_home=$1 job=$2 timeout queue_deadline deadline + local supervisor_pid supervisor_start pid_tmp start_tmp + worker_claim "$job" || return 0 + supervisor_pid=${BASHPID:-$$} + supervisor_start=$(fm_remote_job_process_start "$supervisor_pid") || { + worker_publish_result "$job" 125 || true + return 0 + } + pid_tmp=$(umask 077; mktemp "$job/.claim/.supervisor.XXXXXX") || { + worker_publish_result "$job" 125 || true + return 0 + } + start_tmp=$(umask 077; mktemp "$job/.claim/.supervisor_start.XXXXXX") || { + rm -f -- "$pid_tmp" + worker_publish_result "$job" 125 || true + return 0 + } + if ! printf '%s\n' "$supervisor_pid" > "$pid_tmp" \ + || ! printf '%s\n' "$supervisor_start" > "$start_tmp" \ + || ! chmod 600 "$pid_tmp" "$start_tmp" \ + || ! mv -f -- "$start_tmp" "$job/.claim/supervisor_start" \ + || ! mv -f -- "$pid_tmp" "$job/.claim/supervisor"; then + rm -f -- "$pid_tmp" "$start_tmp" "$job/.claim/supervisor_start" + worker_publish_result "$job" 125 || true + return 0 + fi + if fm_remote_job_cancelled "$job"; then + worker_finalize_cancelled "$account_home" "$job" || true + return 0 + fi + queue_deadline=$(fm_remote_job_read_number "$job" queue_deadline 2>/dev/null || true) + case "$queue_deadline" in ''|*[!0-9]*) worker_publish_result "$job" 126 || true; return 0 ;; esac + if [ "$(date +%s)" -ge "$queue_deadline" ]; then + worker_publish_result "$job" 124 || true + return 0 + fi + timeout=$(fm_remote_job_read_number "$job" timeout 2>/dev/null || true) + case "$timeout" in ''|*[!0-9]*) worker_publish_result "$job" 126 || true; return 0 ;; esac + if [ "$timeout" -gt 3600 ]; then + worker_publish_result "$job" 126 || true + return 0 + fi + # The deadline is measured in whole seconds from a truncated clock read, so + # the +1 keeps the granted window at least the recorded timeout instead of + # silently shaving up to a second off it. + deadline=$(( $(date +%s) + timeout + 1 )) + fm_remote_job_write_number "$job" deadline "$deadline" || { + worker_publish_result "$job" 125 || true + return 0 + } + fm_remote_job_write_state "$job" running || { + worker_publish_result "$job" 125 || true + return 0 + } + worker_run_job "$account_home" "$job" + if fm_remote_job_cancelled "$job"; then + fm_remote_job_reap "$account_home" "${job##*/}" || true + fi +} + +# Each lane runs as its own top-level worker process (--lane), not a +# backgrounded subshell: a bash subshell does not reliably reap its dead +# children, and a zombie group leader keeps its process group signalable, so a +# subshell-hosted monitor loop can believe a finished command is still running +# until the job deadline. A top-level shell is the context the monitor loop +# has always run in. +worker_start_lane() { # + local job=$1 home=$2 lane_pid lane_start + "$SCRIPT_DIR/fm-remote-job-worker.sh" --lane "${job##*/}" & + lane_pid=$! + lane_start=$(fm_remote_job_process_start "$lane_pid" 2>/dev/null || true) + WORKER_LANE_HOMES+=("$home") + WORKER_LANE_PIDS+=("$lane_pid") + WORKER_LANE_STARTS+=("$lane_start") + WORKER_LANE_JOBS+=("$job") +} + +worker_lane_main() { # + local account_home job + fm_remote_job_safe_id "$1" || { worker_error "invalid lane job id"; exit 2; } + account_home=$(worker_account_home) || { worker_error "cannot resolve account home"; exit 1; } + FM_ROOT=$(fm_remote_job_canonical_existing_dir "$FM_ROOT") || { worker_error "configured FM_ROOT is unsafe"; exit 1; } + fm_remote_job_prepare_state "$account_home" || { worker_error "$FM_REMOTE_JOB_ERROR"; exit 1; } + job=$(fm_remote_job_job_dir "$1" 2>/dev/null) || exit 0 + worker_lane_execute "$account_home" "$job" +} + worker_process_once() { # - local account_home=$1 job id state queue_deadline timeout deadline + local account_home=$1 job id state queue_deadline home seq candidates='' + local reserved_index reserved_count home_reserved + local reserved_homes=() + worker_reap_finished_lanes for job in "$FM_REMOTE_JOB_JOBS"/job-*; do [ -d "$job" ] && [ ! -L "$job" ] || continue id=${job##*/} @@ -645,38 +942,59 @@ worker_process_once() { # state=$(fm_remote_job_read_state "$job" 2>/dev/null || true) case "$state" in queued) - worker_clear_dead_claim "$job" || continue + worker_lane_owns_job "$job" && continue + if ! worker_clear_dead_claim "$job"; then + if worker_claim_owner_alive "$job"; then + home=$(worker_read_text "$job" home 8192 2>/dev/null || true) + [ -n "$home" ] && reserved_homes+=("$home") + fi + continue + fi + if fm_remote_job_cancelled "$job"; then + worker_finalize_cancelled "$account_home" "$job" || true + continue + fi queue_deadline=$(fm_remote_job_read_number "$job" queue_deadline 2>/dev/null || true) case "$queue_deadline" in ''|*[!0-9]*) worker_publish_result "$job" 126 || true; continue ;; esac if [ "$(date +%s)" -ge "$queue_deadline" ]; then worker_publish_result "$job" 124 || true continue fi + home=$(worker_read_text "$job" home 8192 2>/dev/null || true) + [ -n "$home" ] || { worker_publish_result "$job" 126 || true; continue; } + # A record staged by an older library has no seq; order it ahead of + # sequenced work as the older job it is. + seq=$(fm_remote_job_read_number "$job" seq 2>/dev/null || true) + case "$seq" in ''|*[!0-9]*) seq=0 ;; esac + candidates="$candidates$seq"$'\t'"$id"$'\t'"$home"$'\n' ;; running) - worker_recover_orphaned_job "$job" || true + worker_lane_owns_job "$job" || worker_reclaim_running_job "$job" || true continue ;; *) continue ;; esac - worker_claim "$job" || continue - timeout=$(fm_remote_job_read_number "$job" timeout 2>/dev/null || true) - case "$timeout" in ''|*[!0-9]*) worker_publish_result "$job" 126 || true; continue ;; esac - if [ "$timeout" -gt 3600 ]; then - worker_publish_result "$job" 126 || true - continue - fi - deadline=$(( $(date +%s) + timeout )) - fm_remote_job_write_number "$job" deadline "$deadline" || { - worker_publish_result "$job" 125 || true - continue - } - fm_remote_job_write_state "$job" running || { - worker_publish_result "$job" 125 || true - continue - } - worker_run_job "$account_home" "$job" done + [ -n "$candidates" ] || return 0 + while IFS=$'\t' read -r seq id home; do + [ -n "$id" ] || continue + worker_lane_busy "$home" && continue + home_reserved=0 + reserved_index=0 + reserved_count=${#reserved_homes[@]} + while [ "$reserved_index" -lt "$reserved_count" ]; do + if [ "${reserved_homes[$reserved_index]}" = "$home" ]; then + home_reserved=1 + break + fi + reserved_index=$((reserved_index + 1)) + done + [ "$home_reserved" -eq 0 ] || continue + job=$(fm_remote_job_job_dir "$id" 2>/dev/null || true) + [ -n "$job" ] || continue + [ "$(fm_remote_job_read_state "$job" 2>/dev/null || true)" = queued ] || continue + worker_start_lane "$job" "$home" + done < <(printf '%s' "$candidates" | sort -t $'\t' -k1,1n -k2,2) } main() { @@ -795,6 +1113,10 @@ case "${1:-}" in [ "$#" -eq 1 ] || { worker_error "unexpected worker arguments"; exit 2; } main ;; + --lane) + [ "$#" -eq 2 ] || { worker_error "unexpected worker arguments"; exit 2; } + worker_lane_main "$2" + ;; '') if [ "$(fm_remote_job_platform)" = linux ]; then worker_supervise_linux; else main; fi ;; diff --git a/bin/fm-send.sh b/bin/fm-send.sh index b10f381ffd6..daa638fe7a7 100755 --- a/bin/fm-send.sh +++ b/bin/fm-send.sh @@ -131,7 +131,11 @@ # FM_PENDING_REPLY_EXISTING_CORR= resend command that preserves the body # and makes a later remote enqueue deduplicate onto that same record. An # unconfirmed fire-and-forget request exits 3 and names the same delivery id to -# retry. The remote host runs no re-ring ladder of its own: a swallowed ordinary +# retry. Every remote transport attempt is bounded by FM_SEND_REMOTE_BUDGET +# seconds (default 30, and any override must be a positive integer): a bound +# hit is completion-unknown and exits through this same unconfirmed contract +# instead of waiting out a busy remote queue. +# The remote host runs no re-ring ladder of its own: a swallowed ordinary # doorbell surfaces through the parent's pending-reply recovery and escalation, # whose recovery request re-rings the remote doorbell when it is enqueued; # fire-and-forget delivery deliberately arms neither mechanism. Internal @@ -228,6 +232,8 @@ fi . "$SCRIPT_DIR/fm-wake-lib.sh" # shellcheck source=bin/fm-task-inbox-lib.sh . "$SCRIPT_DIR/fm-task-inbox-lib.sh" +# shellcheck source=bin/fm-timeout-lib.sh +. "$SCRIPT_DIR/fm-timeout-lib.sh" FM_GUARD_CONTINUE_LINE='This is a supervision warning only; the requested message WILL still be sent.' "$SCRIPT_DIR/fm-guard.sh" || true @@ -648,7 +654,15 @@ if [ "${1:-}" = "--key" ]; then key=$2 semantic_key=$(fm_send_normalize_key "$key") if [ "$TARGET_BACKEND" = remote ]; then - if ! "$SCRIPT_DIR/fm-on.sh" "$TARGET_REMOTE_ID" fm-remote-secondmate-control.sh key "$TARGET_REMOTE_ID" "$key" < /dev/null; then + FM_SEND_REMOTE_BUDGET=${FM_SEND_REMOTE_BUDGET:-30} + case "$FM_SEND_REMOTE_BUDGET" in + ''|*[!0-9]*|0) + echo "error: FM_SEND_REMOTE_BUDGET must be a positive integer: $FM_SEND_REMOTE_BUDGET" >&2 + exit 1 + ;; + esac + if ! fm_run_timed "$FM_SEND_REMOTE_BUDGET" "$SCRIPT_DIR/fm-on.sh" "$TARGET_REMOTE_ID" \ + fm-remote-secondmate-control.sh key "$TARGET_REMOTE_ID" "$key" < /dev/null; then echo "error: key '$key' not sent to remote secondmate $TARGET_REMOTE_ID; completion may be unknown" >&2 exit 1 fi @@ -660,6 +674,15 @@ if [ "${1:-}" = "--key" ]; then fm_send_record_interrupt "$semantic_key" || exit 1 else MESSAGE=$* + if [ "$TARGET_BACKEND" = remote ]; then + FM_SEND_REMOTE_BUDGET=${FM_SEND_REMOTE_BUDGET:-30} + case "$FM_SEND_REMOTE_BUDGET" in + ''|*[!0-9]*|0) + echo "error: FM_SEND_REMOTE_BUDGET must be a positive integer: $FM_SEND_REMOTE_BUDGET" >&2 + exit 1 + ;; + esac + fi # The pre-marker answer text, kept for the closing resolved note so the # durable ledger records the plain answer without marker or corr bytes. RESOLVE_ANSWER_TEXT=$MESSAGE @@ -747,6 +770,10 @@ else # 255 is safe by that idempotence; a still-lost transport preserves a # reply-bearing request's expectation, while fire-and-forget reports the # delivery id that must be reused, because the record may have landed. + # Every transport attempt is bounded by FM_SEND_REMOTE_BUDGET seconds + # (default 30, overridable) so a busy remote queue cannot hold this send + # open indefinitely; a bound hit exits through the same + # unconfirmed-delivery contract. REMOTE_META_LOCK=$(fm_meta_lock_path "$TARGET_META") || exit 1 if ! fm_task_inbox_lock_acquire "$REMOTE_META_LOCK"; then if [ "$PENDING_REPLY_CREATED" = 1 ] && [ -n "$PENDING_REPLY_CORR" ]; then @@ -781,13 +808,22 @@ else remote_completion_unknown=0 REMOTE_SEND_ARGS=("$TARGET_REMOTE_ID" "$MESSAGE") [ -z "$FIRE_AND_FORGET_ID" ] || REMOTE_SEND_ARGS+=(fire-and-forget) - "$SCRIPT_DIR/fm-on.sh" "$TARGET_REMOTE_ID" fm-remote-secondmate-control.sh send \ - "${REMOTE_SEND_ARGS[@]}" < /dev/null || remote_rc=$? - if [ "$remote_rc" -eq 255 ]; then + # Each transport attempt is bounded by FM_SEND_REMOTE_BUDGET seconds. + # fm_run_timed's 124 means the attempt was killed at the bound with remote + # completion unknown - the enqueue may have landed - so it exits through + # the same unconfirmed-delivery contract as a lost transport, without a + # retry that would only wait out the same busy remote queue again. (A + # remote job's own timeout also relays as 124; treating it as unconfirmed + # stays safe because the remote enqueue deduplicates.) + fm_run_timed "$FM_SEND_REMOTE_BUDGET" "$SCRIPT_DIR/fm-on.sh" "$TARGET_REMOTE_ID" \ + fm-remote-secondmate-control.sh send "${REMOTE_SEND_ARGS[@]}" < /dev/null || remote_rc=$? + if [ "$remote_rc" -eq 124 ]; then + remote_completion_unknown=1 + elif [ "$remote_rc" -eq 255 ]; then remote_completion_unknown=1 remote_rc=0 - "$SCRIPT_DIR/fm-on.sh" "$TARGET_REMOTE_ID" fm-remote-secondmate-control.sh send \ - "${REMOTE_SEND_ARGS[@]}" < /dev/null || remote_rc=$? + fm_run_timed "$FM_SEND_REMOTE_BUDGET" "$SCRIPT_DIR/fm-on.sh" "$TARGET_REMOTE_ID" \ + fm-remote-secondmate-control.sh send "${REMOTE_SEND_ARGS[@]}" < /dev/null || remote_rc=$? fi fm_lock_release "$REMOTE_META_LOCK" if [ "$remote_rc" -ne 0 ] && [ "$remote_completion_unknown" -eq 1 ]; then @@ -800,6 +836,8 @@ else fi if [ "$remote_rc" -eq 255 ]; then echo "error: steer to remote secondmate $TARGET_REMOTE_ID is unconfirmed (transport lost twice; remote completion unknown). Only the correlation-reusing resend below is idempotent and lands on the same remote inbox record:" >&2 + elif [ "$remote_rc" -eq 124 ]; then + echo "error: steer to remote secondmate $TARGET_REMOTE_ID is unconfirmed (the remote transport did not complete within its ${FM_SEND_REMOTE_BUDGET}s budget; remote completion unknown). Only the correlation-reusing resend below is idempotent and lands on the same remote inbox record:" >&2 else echo "error: steer to remote secondmate $TARGET_REMOTE_ID is unconfirmed (the first transport attempt had unknown completion and the retry failed). Only the correlation-reusing resend below is idempotent and lands on the same remote inbox record:" >&2 fi diff --git a/bin/fm-test-run.sh b/bin/fm-test-run.sh index 9a1d4a8be1c..c89b6e60f83 100755 --- a/bin/fm-test-run.sh +++ b/bin/fm-test-run.sh @@ -172,6 +172,7 @@ family_for_basename() { ;; fm-backlog-handoff.test.sh|fm-on.test.sh|fm-remote-backlog-handoff.test.sh|\ fm-remote-doctor.test.sh|fm-remote-job.test.sh|fm-remote-job-orphan-reap.test.sh|\ + fm-remote-transport-lanes.test.sh|\ fm-remote-reply.test.sh|fm-remote-secondmate-lifecycle-e2e.test.sh|\ fm-remote-secondmate-trace-context.test.sh|\ fm-secondmate-harness.test.sh|fm-secondmate-lifecycle-e2e.test.sh|\ diff --git a/docs/remote-secondmates.md b/docs/remote-secondmates.md index 3099854056d..5a36fef02a3 100644 --- a/docs/remote-secondmates.md +++ b/docs/remote-secondmates.md @@ -32,8 +32,10 @@ The entrypoint authorizes that bootstrap with normal git tracking when git resol After setup, every other command verifies Firstmate's account-owned remote job worker, stages the encoded argv and stdin bytes, waits for its result, and relays stdout, stderr, and the exit status separately. On macOS the worker is `dev.firstmate.remote-job`, an Aqua-scoped LaunchAgent at `~/Library/LaunchAgents/dev.firstmate.remote-job.plist` with logs under `~/Library/Logs/`. After that bootstrap every non-doctor `fm-on.sh` target runs through that worker in the remote account's GUI session, never in the SSH process or a Herdr pane. -The worker runs one staged job at a time and preempts a running reply long-poll as soon as any command other than another reply long-poll is queued, so interactive commands and startup checks are never serialized behind a poll window. +The worker serves one lane per staged home: jobs for the same home follow the staging-order contract owned by [`bin/fm-remote-job-lib.sh`](../bin/fm-remote-job-lib.sh), while different homes' lanes run concurrently so one home's long job never delays another home's commands. +Within a home's lane the worker preempts a running reply long-poll as soon as any command other than another reply long-poll is queued for that home, so interactive commands and startup checks are never serialized behind a poll window. `bin/fm-remote-job-lib.sh` owns that preemption contract and distinguishes preemption from a wait window that closes with no data, so only a genuinely quiet window proves channel freshness while either outcome can re-arm without losing data. +A caller that disconnects or whose caller-side wait expires before its job completes cancels it instead of abandoning it: cancelled queued work is skipped, cancelled running work is stopped, and the finalized record is cleaned up, so retries never convoy behind abandoned work. Linux uses the same queue and worker protocol without the Aqua-session requirement. A worker stops itself once its configured code root stops being a Firstmate checkout, so a worker started from a worktree cannot outlive that worktree, and `bin/fm-remote-job-reap-orphans.sh` clears any worker already left behind that way without ever touching one whose checkout still exists. The remote account must provide the required toolchain, the selected worker runtime, the selected session backend, and credentials that work on that host. @@ -173,8 +175,9 @@ FM_HOME= bin/fm-send.sh fm- '' The [`fm-send.sh` header](../bin/fm-send.sh) owns the exact delivery-status contract. A routed request is delivered as a durable record in the remote home's steering inbox plus a best-effort doorbell, never by typing the payload into the pane; exit 0 means the record durably exists. -An unconfirmed transport (SSH exit 255) is retried identically once and preserves this ordinary reply-bearing request's pending-reply expectation for the record that may have landed. -If it remains unconfirmed, only the exact `FM_PENDING_REPLY_EXISTING_CORR=` resend command printed by `fm-send` is safe to run later because it preserves the request body and lets the remote enqueue deduplicate onto the same record; a plain rerun mints a different correlation and is not idempotent. +Every remote transport attempt is bounded by `FM_SEND_REMOTE_BUDGET`; that header owns the setting's default and validation contract. +An unconfirmed SSH transport (exit 255) is retried identically once, while a budget expiry is not retried because completion is unknown; either outcome preserves this ordinary reply-bearing request's pending-reply expectation for the record that may have landed. +If delivery remains unconfirmed, only the exact `FM_PENDING_REPLY_EXISTING_CORR=` resend command printed by `fm-send` is safe to run later because it preserves the request body and lets the remote enqueue deduplicate onto the same record; a plain rerun mints a different correlation and is not idempotent. When deduplication finds that the worker already moved the matching record into `handled/`, the resend exits successfully without ringing the doorbell again. The remote host runs no doorbell re-ring ladder of its own; a swallowed doorbell for an ordinary reply-bearing request surfaces through the parent's pending-reply recovery and escalation, whose recovery request rings the doorbell again when it is enqueued. `fm-peek.sh` and `fm-crew-state.sh` route remote-secondmate reads to the endpoint's host instead of consulting local worktree or backend state. @@ -250,6 +253,7 @@ bin/fm-test-run.sh tests/fm-secondmate-reconcile.test.sh bin/fm-test-run.sh tests/fm-peek-remote.test.sh bin/fm-test-run.sh tests/fm-crew-state.test.sh bin/fm-test-run.sh tests/fm-remote-job.test.sh +bin/fm-test-run.sh tests/fm-remote-transport-lanes.test.sh bin/fm-test-run.sh tests/fm-remote-doctor.test.sh bin/fm-test-run.sh tests/fm-project-origin.test.sh bin/fm-test-run.sh tests/fm-remote-reply.test.sh diff --git a/tests/fm-on.test.sh b/tests/fm-on.test.sh index 790a56d5038..cde6cb3ef49 100755 --- a/tests/fm-on.test.sh +++ b/tests/fm-on.test.sh @@ -119,7 +119,8 @@ fm_on() { # The pre-feature user path had no executable transport at all. The regression # exercises the adopted public surface end to end through a deterministic SSH -# process boundary rather than checking script source. +# process boundary rather than checking script source. A payload caller passes +# --stdin explicitly; without it the remote command's stdin is /dev/null. ARGV_ACTUAL="$REMOTE_HOME/argv.bin" ARGV_EXPECTED="$TMP_ROOT/argv-expected.bin" # shellcheck disable=SC2016 # Literal shell-looking argv is the injection probe. @@ -127,7 +128,7 @@ printf '%s\0' 'plain' 'two words' '$(touch /tmp/fm-on-injected)' '' $'line one\n printf 'payload one\npayload two\n' > "$TMP_ROOT/stdin" set +e # shellcheck disable=SC2016 # Literal shell-looking argv is the injection probe. -fm_on ios fm-probe-one.sh "$ARGV_ACTUAL" 23 \ +fm_on --stdin ios fm-probe-one.sh "$ARGV_ACTUAL" 23 \ 'plain' 'two words' '$(touch /tmp/fm-on-injected)' '' $'line one\nline two' \ < "$TMP_ROOT/stdin" > "$TMP_ROOT/stdout" 2> "$TMP_ROOT/stderr" rc=$? @@ -139,7 +140,21 @@ assert_grep 'stdin: payload one' "$TMP_ROOT/stdout" "remote stdin was not preser assert_grep 'stdin: payload two' "$TMP_ROOT/stdout" "remote stdin lost its second line" assert_grep 'stderr: separate' "$TMP_ROOT/stderr" "remote stderr was not preserved separately" assert_absent /tmp/fm-on-injected "shell-looking argv was interpreted" -pass "fm-on preserves argv, stdin, stdout, stderr, and exit status without shell interpretation" +pass "fm-on --stdin preserves argv, stdin, stdout, stderr, and exit status without shell interpretation" + +# Without --stdin the remote command must see EOF even when the caller's own +# stdin holds bytes: staging captures stdin to EOF, so an open caller stream +# must never reach it by default. +set +e +fm_on ios fm-probe-one.sh "$REMOTE_HOME/argv-default.bin" 0 'default-closed' \ + < "$TMP_ROOT/stdin" > "$TMP_ROOT/stdout-default" 2> "$TMP_ROOT/stderr-default" +rc=$? +set -e +[ "$rc" -eq 0 ] || fail "the default-closed invocation did not preserve exit status (got $rc)" +if grep -q 'stdin:' "$TMP_ROOT/stdout-default"; then + fail "caller stdin crossed the transport without --stdin: $(cat "$TMP_ROOT/stdout-default")" +fi +pass "fm-on defaults the remote command's stdin to /dev/null" # A vanished remote peer must become a bounded ssh failure instead of an # indefinite hang on a half-open TCP connection, so the existing no-result -> diff --git a/tests/fm-remote-job.test.sh b/tests/fm-remote-job.test.sh index a97b2edb3e8..fb6ea8ef99b 100755 --- a/tests/fm-remote-job.test.sh +++ b/tests/fm-remote-job.test.sh @@ -627,8 +627,11 @@ RECOVERY_REFUSED_RC=$? set -e [ "$RECOVERY_REFUSED_RC" -ne 0 ] || fail "quarantine recovery ignored a recorded live process" assert_present "$RECOVERY_STATE/worker.lock/quarantine" "a live recorded process lost quarantine protection" -kill "$QUARANTINED_PROCESS_PID" 2>/dev/null || true -wait "$QUARANTINED_PROCESS_PID" 2>/dev/null || true +printf '%s\n' "$QUARANTINED_PROCESS_PID" > "$RECOVERY_JOB/.claim/owner" +printf 'stale owner identity\n' > "$RECOVERY_JOB/.claim/owner_start" +printf 'stale supervisor identity\n' > "$RECOVERY_JOB/.claim/supervisor_start" +chmod 600 "$RECOVERY_JOB/.claim/owner" "$RECOVERY_JOB/.claim/owner_start" \ + "$RECOVERY_JOB/.claim/supervisor_start" HOME="$RECOVERY_HOME" FM_ROOT_OVERRIDE="$REMOTE_ROOT" FM_REMOTE_JOB_STATE_ROOT="$RECOVERY_STATE" \ FM_REMOTE_JOB_PLATFORM_OVERRIDE=Linux "$REMOTE_ROOT/bin/fm-remote-job-worker.sh" \ > "$TMP_ROOT/recovery-worker.out" 2> "$TMP_ROOT/recovery-worker.err" & @@ -637,12 +640,16 @@ for _ in $(seq 1 300); do [ -f "$RECOVERY_STATE/worker.ready" ] && break sleep 0.05 done -assert_present "$RECOVERY_STATE/worker.ready" "a stopped quarantined execution did not permit worker recovery" +assert_present "$RECOVERY_STATE/worker.ready" "a reused supervisor pid did not permit worker recovery" assert_absent "$RECOVERY_STATE/worker.lock/quarantine" "recovered worker retained stale quarantine" +kill -0 "$QUARANTINED_PROCESS_PID" 2>/dev/null \ + || fail "worker recovery signalled a process whose supervisor identity did not match" kill -TERM "$RECOVERY_WORKER_PID" wait "$RECOVERY_WORKER_PID" 2>/dev/null || true RECOVERY_WORKER_PID= -pass "quarantine clears only after recorded execution has stopped" +kill "$QUARANTINED_PROCESS_PID" 2>/dev/null || true +wait "$QUARANTINED_PROCESS_PID" 2>/dev/null || true +pass "quarantine recovery refuses unverifiable supervisors and ignores reused pids" # A replacement stops a Linux worker by signalling its whole isolated group, and # the supervisor in that group forwards a second stop signal to the same serving diff --git a/tests/fm-remote-transport-lanes.test.sh b/tests/fm-remote-transport-lanes.test.sh new file mode 100755 index 00000000000..e4815a7f7ea --- /dev/null +++ b/tests/fm-remote-transport-lanes.test.sh @@ -0,0 +1,425 @@ +#!/usr/bin/env bash +# Behavior tests for the remote transport's per-home lanes, caller-disconnect +# cancellation, stdin default, and staging-litter reaping. +# +# Pins, against the real worker and the real fm-on -> entrypoint transport +# (through the deterministic FM_SSH_BIN seam tests/fm-on.test.sh proves +# preserves exit status): +# T9: a job for home B completes while home A runs a long job, and two +# A-jobs execute strictly in stage order even when staged rapidly. +# T3: a caller killed mid-wait cancels its job - the worker never executes a +# cancelled queued job and terminates a running cancelled job's process +# group - and a caller whose parent dies without delivering a signal +# (the dead-ssh-channel shape) cancels the same way; afterwards a burst +# of short commands completes with no convoy. +# T6: a non-payload fm-on call with an OPEN stdin pipe completes instead of +# wedging staging, and a payload caller with --stdin still delivers its +# bytes through the worker. +# Stage litter older than the reap age does not survive a worker pass while +# fresh staging does. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +ROOT=$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd -P) +# shellcheck source=bin/fm-timeout-lib.sh +. "$ROOT/bin/fm-timeout-lib.sh" + +TMP_ROOT=$(fm_test_tmproot fm-remote-transport-lanes) +mkdir -p "$TMP_ROOT" +TMP_ROOT=$(cd "$TMP_ROOT" && pwd -P) +REMOTE_ROOT="$TMP_ROOT/remote-root" +HOME_A="$TMP_ROOT/home-a" +HOME_B="$TMP_ROOT/home-b" +HOME_EDGE="$TMP_ROOT/home-a " +LOCAL_HOME="$TMP_ROOT/local-home" +ACCOUNT_HOME="$TMP_ROOT/account" +STATE_ROOT="$TMP_ROOT/remote-jobs" +FAKEBIN=$(fm_fakebin "$TMP_ROOT/fakebin") +mkdir -p "$REMOTE_ROOT/bin" "$HOME_A" "$HOME_B" "$HOME_EDGE" "$LOCAL_HOME/data" "$ACCOUNT_HOME" + +cleanup_lane_fixture() { + if [ -f "$STATE_ROOT/worker.pid" ]; then + fm_remote_job_stop_worker_tree "$(cat "$STATE_ROOT/worker.pid")" || true + fi + rm -rf -- "$TMP_ROOT" +} +trap cleanup_lane_fixture EXIT + +cp "$ROOT/bin/fm-remote-job-lib.sh" "$ROOT/bin/fm-remote-job-worker.sh" \ + "$ROOT/bin/fm-remote-entrypoint.sh" "$ROOT/bin/fm-remote-delta-read.sh" \ + "$ROOT/bin/fm-remote-secondmate-control.sh" "$ROOT/bin/fm-backend.sh" \ + "$ROOT/bin/fm-pending-reply-lib.sh" "$ROOT/bin/fm-task-inbox-lib.sh" \ + "$ROOT/bin/fm-wake-lib.sh" "$ROOT/bin/fm-marker-lib.sh" \ + "$ROOT/bin/fm-operational-input.sh" "$ROOT/bin/fm-tmux-lib.sh" \ + "$ROOT/bin/fm-composer-lib.sh" "$ROOT/bin/fm-cursor-lib.sh" \ + "$ROOT/bin/fm-classify-lib.sh" "$ROOT/bin/fm-timeout-lib.sh" \ + "$REMOTE_ROOT/bin/" +mkdir -p "$REMOTE_ROOT/bin/backends" +cp "$ROOT/bin/backends/herdr.sh" "$REMOTE_ROOT/bin/backends/herdr.sh" +printf 'fixture\n' > "$REMOTE_ROOT/AGENTS.md" +# Appends its tag to a shared log, then optionally sleeps: the log order is the +# observable execution order. +cat > "$REMOTE_ROOT/bin/fm-mark-job.sh" <<'SH' +#!/bin/bash +printf '%s\n' "$1" >> "$2" +sleep "${3:-0}" +SH +cat > "$REMOTE_ROOT/bin/fm-touch-job.sh" <<'SH' +#!/bin/bash +printf 'ran\n' > "$1" +SH +# Marks its start, sleeps, then marks completion: cancellation must leave the +# start marker without the completion marker. +cat > "$REMOTE_ROOT/bin/fm-two-phase-job.sh" <<'SH' +#!/bin/bash +printf 'started\n' > "$1" +sleep "$3" +printf 'finished\n' > "$2" +SH +cat > "$REMOTE_ROOT/bin/fm-stdin-probe.sh" <<'SH' +#!/bin/bash +while IFS= read -r line || [ -n "$line" ]; do printf 'stdin=%s\n' "$line"; done +SH +chmod +x "$REMOTE_ROOT/bin"/*.sh +git -C "$REMOTE_ROOT" init -q -b main +git -C "$REMOTE_ROOT" config user.email test@example.com +git -C "$REMOTE_ROOT" config user.name Test +git -C "$REMOTE_ROOT" add AGENTS.md bin +git -C "$REMOTE_ROOT" commit -qm 'lane transport fixture' + +# ios routes to home A, build routes to home B. +cat > "$LOCAL_HOME/data/secondmates.md" < "$FAKEBIN/fake-ssh" <<'SH' +#!/usr/bin/env bash +while [ "$#" -gt 0 ]; do + case "$1" in + -o) shift 2 ;; + --) shift; break ;; + *) exit 90 ;; + esac +done +shift 2 +exec "$FM_FAKE_REMOTE_ENTRYPOINT" "$@" +SH +chmod +x "$FAKEBIN/fake-ssh" + +export FM_REMOTE_JOB_STATE_ROOT="$STATE_ROOT" +export FM_REMOTE_JOB_PLATFORM_OVERRIDE=Linux +export FM_REMOTE_JOB_QUEUE_TIMEOUT=60 +export FM_REMOTE_JOB_TIMEOUT=30 +export FM_REMOTE_JOB_STAGE_REAP_SECONDS=1 +# shellcheck source=bin/fm-remote-job-lib.sh +. "$ROOT/bin/fm-remote-job-lib.sh" + +fm_remote_job_prepare_state "$ACCOUNT_HOME" || fail "$FM_REMOTE_JOB_ERROR" +rm -f -- "$STATE_ROOT/seq" +SEQ_PIDS=() +for i in $(seq 1 20); do + fm_remote_job_next_seq > "$TMP_ROOT/seq-$i" & + SEQ_PIDS+=("$!") +done +for pid in "${SEQ_PIDS[@]}"; do + wait "$pid" || fail "a concurrent sequence allocator failed" +done +SEQ_RESULTS=$(cat "$TMP_ROOT"/seq-* | sort -n) +SEQ_EXPECTED=$(seq 1 20) +[ "$SEQ_RESULTS" = "$SEQ_EXPECTED" ] \ + || fail "concurrent sequence claims were not unique and monotonic: $SEQ_RESULTS" +[ "$(find "$STATE_ROOT/.seq-claims" -mindepth 1 -maxdepth 1 -type d | wc -l | tr -d ' ')" = 20 ] \ + || fail "concurrent sequence allocations did not retain every durable claim" +mkdir "$STATE_ROOT/.seq-claims/999998" "$STATE_ROOT/.seq-claims/999999" +touch -t 200001010000 "$STATE_ROOT/.seq-claims/999998" +fm_remote_job_reap_stale "$ACCOUNT_HOME" || fail "sequence claim reaping failed" +assert_absent "$STATE_ROOT/.seq-claims/999998" "an expired sequence claim survived stale reaping" +assert_present "$STATE_ROOT/.seq-claims/999999" "a fresh sequence claim was reaped" +mkdir "$STATE_ROOT/.seq-claims/999997" +touch -t 200001010000 "$STATE_ROOT/.seq-claims/999997" +fm_remote_job_reap_stale "$ACCOUNT_HOME" || fail "rate-limited sequence claim reaping failed" +assert_present "$STATE_ROOT/.seq-claims/999997" "sequence claims were rescanned before the hourly interval" +touch -t 200001010000 "$STATE_ROOT/.seq-claims-reaped" +fm_remote_job_reap_stale "$ACCOUNT_HOME" || fail "expired sequence claim reaping failed" +assert_absent "$STATE_ROOT/.seq-claims/999997" "an expired sequence claim survived the next hourly scan" +rmdir "$STATE_ROOT/.seq-claims/999999" +pass "atomic sequence claims remain unique and reap only after expiry" + +fm_on() { + FM_HOME="$LOCAL_HOME" \ + FM_ROOT_OVERRIDE="$REMOTE_ROOT" \ + FM_SSH_BIN="$FAKEBIN/fake-ssh" \ + FM_FAKE_REMOTE_ENTRYPOINT="$REMOTE_ROOT/bin/fm-remote-entrypoint.sh" \ + "$ROOT/bin/fm-on.sh" "$@" +} + +job_state() { # + fm_remote_job_read_state "$STATE_ROOT/jobs/$1" 2>/dev/null || true +} + +wait_for_state() { # + local i=0 + while [ "$i" -lt 200 ]; do + [ "$(job_state "$1")" = "$2" ] && return 0 + i=$((i + 1)) + sleep 0.05 + done + return 1 +} + +HOME="$ACCOUNT_HOME" FM_ROOT_OVERRIDE="$REMOTE_ROOT" FM_REMOTE_JOB_STATE_ROOT="$STATE_ROOT" \ + FM_REMOTE_JOB_PLATFORM_OVERRIDE=Linux \ + "$REMOTE_ROOT/bin/fm-remote-job-worker.sh" > "$TMP_ROOT/worker.out" 2> "$TMP_ROOT/worker.err" & +for _ in $(seq 1 100); do + [ -f "$STATE_ROOT/worker.ready" ] && break + sleep 0.05 +done +assert_present "$STATE_ROOT/worker.ready" "the worker did not publish its readiness heartbeat" + +# T9: home B's job completes while home A runs a long job, and A's queued job +# stays strictly behind A's running job. +LOG_A="$TMP_ROOT/log-a" +LOG_B="$TMP_ROOT/log-b" +fm_remote_job_stage "$ACCOUNT_HOME" "$REMOTE_ROOT" "$HOME_A" fm-mark-job.sh a1 "$LOG_A" 4 < /dev/null > /dev/null +A1=$FM_REMOTE_JOB_ID +wait_for_state "$A1" running || fail "home A's long job did not begin running" +fm_remote_job_stage "$ACCOUNT_HOME" "$REMOTE_ROOT" "$HOME_A" fm-mark-job.sh a2 "$LOG_A" 0 < /dev/null > /dev/null +A2=$FM_REMOTE_JOB_ID +fm_remote_job_stage "$ACCOUNT_HOME" "$REMOTE_ROOT" "$HOME_EDGE" fm-mark-job.sh b1 "$LOG_B" 0 < /dev/null > /dev/null +B1=$FM_REMOTE_JOB_ID +B_BEGAN=$(date +%s) +fm_remote_job_wait "$ACCOUNT_HOME" "$B1" || fail "$FM_REMOTE_JOB_ERROR" +B_ELAPSED=$(( $(date +%s) - B_BEGAN )) +[ "$FM_REMOTE_JOB_EXIT" -eq 0 ] || fail "home B's job behind home A's long job did not complete" +[ "$B_ELAPSED" -le 3 ] || fail "home B's job waited ${B_ELAPSED}s behind home A's long job" +[ "$(job_state "$A1")" = running ] || fail "home A's long job should still be running for the FIFO assertion" +[ "$(cat "$LOG_A")" = a1 ] || fail "home A's queued job ran beside its running job: $(cat "$LOG_A")" +fm_remote_job_reap "$ACCOUNT_HOME" "$B1" || fail "home B's job could not be reaped" +fm_remote_job_wait "$ACCOUNT_HOME" "$A1" || fail "$FM_REMOTE_JOB_ERROR" +fm_remote_job_wait "$ACCOUNT_HOME" "$A2" || fail "$FM_REMOTE_JOB_ERROR" +[ "$(printf '%s' "$(cat "$LOG_A")")" = "$(printf 'a1\na2')" ] \ + || fail "home A's jobs did not execute in stage order: $(cat "$LOG_A")" +fm_remote_job_reap "$ACCOUNT_HOME" "$A1" || fail "home A's first job could not be reaped" +fm_remote_job_reap "$ACCOUNT_HOME" "$A2" || fail "home A's second job could not be reaped" +pass "lanes run homes concurrently while each home stays FIFO" + +# T9 stage order: five jobs staged in rapid succession behind a busy lane must +# execute in staging-sequence order, not the queue directory's random-id order. +: > "$LOG_A" +fm_remote_job_stage "$ACCOUNT_HOME" "$REMOTE_ROOT" "$HOME_A" fm-mark-job.sh hold "$LOG_A" 2 < /dev/null > /dev/null +HOLD=$FM_REMOTE_JOB_ID +wait_for_state "$HOLD" running || fail "the lane-holding job did not begin running" +RAPID_IDS=() +for tag in r1 r2 r3 r4 r5; do + fm_remote_job_stage "$ACCOUNT_HOME" "$REMOTE_ROOT" "$HOME_A" fm-mark-job.sh "$tag" "$LOG_A" 0 < /dev/null > /dev/null + RAPID_IDS+=("$FM_REMOTE_JOB_ID") +done +fm_remote_job_wait "$ACCOUNT_HOME" "$HOLD" || fail "$FM_REMOTE_JOB_ERROR" +fm_remote_job_reap "$ACCOUNT_HOME" "$HOLD" || true +for id in "${RAPID_IDS[@]}"; do + fm_remote_job_wait "$ACCOUNT_HOME" "$id" || fail "$FM_REMOTE_JOB_ERROR" + fm_remote_job_reap "$ACCOUNT_HOME" "$id" || true +done +[ "$(cat "$LOG_A")" = "$(printf 'hold\nr1\nr2\nr3\nr4\nr5')" ] \ + || fail "rapidly staged same-home jobs did not execute in stage order: $(tr '\n' ' ' < "$LOG_A")" +pass "same-home jobs staged in the same second execute in staging-sequence order" + +: > "$LOG_A" +fm_remote_job_stage "$ACCOUNT_HOME" "$REMOTE_ROOT" "$HOME_A" fm-mark-job.sh publish-hold "$LOG_A" 3 < /dev/null > /dev/null +PUBLISH_HOLD=$FM_REMOTE_JOB_ID +wait_for_state "$PUBLISH_HOLD" running || fail "the publication-order lane holder did not begin running" +( + { + printf 'delayed payload\n' + sleep 5 + } | fm_remote_job_stage "$ACCOUNT_HOME" "$REMOTE_ROOT" "$HOME_A" \ + fm-mark-job.sh delayed "$LOG_A" 0 +) > "$TMP_ROOT/delayed-stage-id" & +DELAYED_STAGE_PID=$! +for _ in $(seq 1 200); do + ls "$STATE_ROOT/jobs"/.stage.* >/dev/null 2>&1 && break + sleep 0.02 +done +ls "$STATE_ROOT/jobs"/.stage.* >/dev/null 2>&1 \ + || fail "the delayed stdin stage did not begin capturing" +fm_remote_job_stage "$ACCOUNT_HOME" "$REMOTE_ROOT" "$HOME_A" fm-mark-job.sh fast "$LOG_A" 0 < /dev/null > /dev/null +FAST_STAGE=$FM_REMOTE_JOB_ID +wait "$DELAYED_STAGE_PID" || fail "the delayed stdin stage failed to publish" +DELAYED_STAGE=$(cat "$TMP_ROOT/delayed-stage-id") +FAST_SEQ=$(fm_remote_job_read_number "$STATE_ROOT/jobs/$FAST_STAGE" seq) \ + || fail "the fast stage lost its sequence" +DELAYED_SEQ=$(fm_remote_job_read_number "$STATE_ROOT/jobs/$DELAYED_STAGE" seq) \ + || fail "the delayed stage lost its sequence" +[ "$FAST_SEQ" -lt "$DELAYED_SEQ" ] \ + || fail "sequence order did not follow publication order: fast=$FAST_SEQ delayed=$DELAYED_SEQ" +fm_remote_job_wait "$ACCOUNT_HOME" "$PUBLISH_HOLD" || fail "$FM_REMOTE_JOB_ERROR" +fm_remote_job_wait "$ACCOUNT_HOME" "$FAST_STAGE" || fail "$FM_REMOTE_JOB_ERROR" +fm_remote_job_wait "$ACCOUNT_HOME" "$DELAYED_STAGE" || fail "$FM_REMOTE_JOB_ERROR" +[ "$(cat "$LOG_A")" = "$(printf 'publish-hold\nfast\ndelayed')" ] \ + || fail "execution order diverged from publication sequence: $(tr '\n' ' ' < "$LOG_A")" +fm_remote_job_reap "$ACCOUNT_HOME" "$PUBLISH_HOLD" || true +fm_remote_job_reap "$ACCOUNT_HOME" "$FAST_STAGE" || true +fm_remote_job_reap "$ACCOUNT_HOME" "$DELAYED_STAGE" || true +pass "same-home sequence order follows completed staging publication" + +# T3a: a caller killed while its job is still queued cancels it; the worker +# never executes it. +fm_remote_job_stage "$ACCOUNT_HOME" "$REMOTE_ROOT" "$HOME_A" fm-mark-job.sh hold2 "$LOG_A" 4 < /dev/null > /dev/null +HOLD2=$FM_REMOTE_JOB_ID +wait_for_state "$HOLD2" running || fail "the cancellation fixture's lane holder did not begin running" +QUEUED_EFFECT="$TMP_ROOT/queued-cancel-effect" +fm_on ios fm-touch-job.sh "$QUEUED_EFFECT" > /dev/null 2>&1 & +QUEUED_CALLER=$! +QUEUED_JOB= +for _ in $(seq 1 200); do + for job in "$STATE_ROOT"/jobs/job-*; do + [ -d "$job" ] || continue + [ "${job##*/}" = "$HOLD2" ] && continue + [ "$(job_state "${job##*/}")" = queued ] && QUEUED_JOB=${job##*/} && break + done + [ -n "$QUEUED_JOB" ] && break + sleep 0.05 +done +[ -n "$QUEUED_JOB" ] || fail "the doomed caller's job never appeared in the queue" +kill -TERM "$QUEUED_CALLER" 2>/dev/null || true +wait "$QUEUED_CALLER" 2>/dev/null || true +for _ in $(seq 1 200); do + [ ! -d "$STATE_ROOT/jobs/$QUEUED_JOB" ] && break + sleep 0.05 +done +[ ! -d "$STATE_ROOT/jobs/$QUEUED_JOB" ] \ + || fail "the cancelled queued job's record survived (state: $(job_state "$QUEUED_JOB"))" +fm_remote_job_wait "$ACCOUNT_HOME" "$HOLD2" || fail "$FM_REMOTE_JOB_ERROR" +fm_remote_job_reap "$ACCOUNT_HOME" "$HOLD2" || true +sleep 1 +assert_absent "$QUEUED_EFFECT" "the worker executed a queued job whose caller was killed" +pass "a caller killed mid-wait cancels its queued job before execution" + +# T3b: a caller killed while its job is running terminates the job's process +# group instead of letting it run to completion for nobody. +RUN_START="$TMP_ROOT/running-cancel-start" +RUN_FINISH="$TMP_ROOT/running-cancel-finish" +fm_on build fm-two-phase-job.sh "$RUN_START" "$RUN_FINISH" 8 > /dev/null 2>&1 & +RUNNING_CALLER=$! +for _ in $(seq 1 200); do + [ -f "$RUN_START" ] && break + sleep 0.05 +done +assert_present "$RUN_START" "the running-cancellation fixture never started" +kill -TERM "$RUNNING_CALLER" 2>/dev/null || true +wait "$RUNNING_CALLER" 2>/dev/null || true +CANCEL_BEGAN=$(date +%s) +for _ in $(seq 1 200); do + ls "$STATE_ROOT"/jobs/job-* >/dev/null 2>&1 || break + sleep 0.05 +done +CANCEL_ELAPSED=$(( $(date +%s) - CANCEL_BEGAN )) +ls "$STATE_ROOT"/jobs/job-* >/dev/null 2>&1 \ + && fail "the cancelled running job's record survived" +[ "$CANCEL_ELAPSED" -le 6 ] || fail "running-job cancellation took ${CANCEL_ELAPSED}s" +sleep 2 +assert_absent "$RUN_FINISH" "a cancelled running job's process group ran to completion" +pass "a caller killed mid-wait stops its running job's process group" + +# T3c: a caller whose parent exits WITHOUT delivering any signal - the shape a +# dead ssh channel leaves behind - still cancels through the entrypoint's +# parent-liveness probe. +ORPHAN_START="$TMP_ROOT/orphan-cancel-start" +ORPHAN_FINISH="$TMP_ROOT/orphan-cancel-finish" +# shellcheck disable=SC2016 # Expansion is deliberately deferred to the child shell. +env FM_HOME="$LOCAL_HOME" FM_ROOT_OVERRIDE="$REMOTE_ROOT" \ + FM_SSH_BIN="$FAKEBIN/fake-ssh" \ + FM_FAKE_REMOTE_ENTRYPOINT="$REMOTE_ROOT/bin/fm-remote-entrypoint.sh" \ + FM_REMOTE_JOB_STATE_ROOT="$STATE_ROOT" FM_REMOTE_JOB_PLATFORM_OVERRIDE=Linux \ + bash -c ' + "$1/bin/fm-on.sh" build fm-two-phase-job.sh "$2" "$3" 12 >/dev/null 2>&1 & + while [ ! -f "$2" ]; do sleep 0.1; done + ' _ "$ROOT" "$ORPHAN_START" "$ORPHAN_FINISH" +assert_present "$ORPHAN_START" "the orphan-cancellation fixture never started" +ORPHAN_BEGAN=$(date +%s) +for _ in $(seq 1 300); do + ls "$STATE_ROOT"/jobs/job-* >/dev/null 2>&1 || break + sleep 0.05 +done +ORPHAN_ELAPSED=$(( $(date +%s) - ORPHAN_BEGAN )) +ls "$STATE_ROOT"/jobs/job-* >/dev/null 2>&1 \ + && fail "the orphaned caller's job record survived its disconnect" +[ "$ORPHAN_ELAPSED" -le 10 ] || fail "orphan-disconnect cancellation took ${ORPHAN_ELAPSED}s" +sleep 2 +assert_absent "$ORPHAN_FINISH" "a job abandoned by a signal-less disconnect ran to completion" +pass "a signal-less caller disconnect cancels the abandoned job through the parent probe" + +# T3: after the cancellations, a burst of short bounded commands meets its own +# budget - no convoy behind abandoned work. +BURST_BEGAN=$(date +%s) +for tag in c1 c2 c3; do + rc=0 + fm_run_timed 15 env FM_HOME="$LOCAL_HOME" FM_ROOT_OVERRIDE="$REMOTE_ROOT" \ + FM_SSH_BIN="$FAKEBIN/fake-ssh" \ + FM_FAKE_REMOTE_ENTRYPOINT="$REMOTE_ROOT/bin/fm-remote-entrypoint.sh" \ + FM_REMOTE_JOB_STATE_ROOT="$STATE_ROOT" FM_REMOTE_JOB_PLATFORM_OVERRIDE=Linux \ + "$ROOT/bin/fm-on.sh" ios fm-touch-job.sh "$TMP_ROOT/burst-$tag" >/dev/null 2>&1 || rc=$? + [ "$rc" -eq 0 ] || fail "post-cancellation burst command $tag failed with $rc" + assert_present "$TMP_ROOT/burst-$tag" "post-cancellation burst command $tag did not run" +done +BURST_ELAPSED=$(( $(date +%s) - BURST_BEGAN )) +[ "$BURST_ELAPSED" -le 12 ] || fail "the post-cancellation burst convoyed for ${BURST_ELAPSED}s" +pass "bounded reads after a cancellation meet their own budget with no convoy" + +# T6: a non-payload call with an OPEN stdin pipe completes instead of wedging +# staging on a stdin capture that never reaches EOF. +printf 'rsm\n' > "$HOME_A/.fm-secondmate-home" +printf '# fixture secondmate home\n' > "$HOME_A/AGENTS.md" +mkdir -p "$HOME_A/state" "$HOME_A/bin" +rc=0 +fm_run_timed 20 env FM_HOME="$LOCAL_HOME" FM_ROOT_OVERRIDE="$REMOTE_ROOT" \ + FM_SSH_BIN="$FAKEBIN/fake-ssh" \ + FM_FAKE_REMOTE_ENTRYPOINT="$REMOTE_ROOT/bin/fm-remote-entrypoint.sh" \ + FM_REMOTE_JOB_STATE_ROOT="$STATE_ROOT" FM_REMOTE_JOB_PLATFORM_OVERRIDE=Linux \ + "$ROOT/bin/fm-on.sh" ios fm-remote-secondmate-control.sh state rsm \ + < <(sleep 30) > "$TMP_ROOT/state-out" 2> "$TMP_ROOT/state-err" || rc=$? +[ "$rc" -ne 124 ] || fail "a control-state call with an open stdin pipe wedged staging" +assert_grep 'missing' "$TMP_ROOT/state-out" \ + "the control-state call did not complete through the worker: $(cat "$TMP_ROOT/state-err")" +pass "an open caller stdin no longer wedges a non-payload remote command" + +# A live explicit stdin stage can exceed the litter age while waiting for EOF; +# the stale sweep must retain it until its owning entrypoint publishes the job. +rc=0 +{ + printf 'slow payload one\n' + sleep 3 + printf 'slow payload two\n' +} | fm_on --stdin ios fm-stdin-probe.sh > "$TMP_ROOT/slow-payload-out" 2> "$TMP_ROOT/slow-payload-err" || rc=$? +expect_code 0 "$rc" "a live slow stdin stage must survive stale reaping: $(cat "$TMP_ROOT/slow-payload-err")" +assert_grep 'stdin=slow payload one' "$TMP_ROOT/slow-payload-out" "the slow stdin stage lost its first bytes" +assert_grep 'stdin=slow payload two' "$TMP_ROOT/slow-payload-out" "the slow stdin stage was reaped before EOF" +pass "a live explicit-stdin stage survives the staging-litter age bound" + +# T6: a payload caller with --stdin still delivers its bytes. +printf 'payload byte one\npayload byte two\n' > "$TMP_ROOT/payload" +fm_on --stdin ios fm-stdin-probe.sh < "$TMP_ROOT/payload" > "$TMP_ROOT/payload-out" 2>/dev/null \ + || fail "the --stdin payload call failed" +assert_grep 'stdin=payload byte one' "$TMP_ROOT/payload-out" "--stdin did not deliver the payload" +assert_grep 'stdin=payload byte two' "$TMP_ROOT/payload-out" "--stdin lost part of the payload" +pass "--stdin still delivers a payload caller's bytes" + +# Stage litter: an abandoned .stage.* older than the reap age does not survive +# a worker pass, while fresh staging is left alone. +OLD_STAGE="$STATE_ROOT/jobs/.stage.abandoned" +FRESH_STAGE="$STATE_ROOT/jobs/.stage.fresh" +mkdir -p "$OLD_STAGE" "$FRESH_STAGE" +touch -t 200001010000 "$OLD_STAGE" +for _ in $(seq 1 100); do + [ ! -d "$OLD_STAGE" ] && break + sleep 0.05 +done +[ ! -d "$OLD_STAGE" ] || fail "stage litter older than the reap age survived the worker pass" +assert_present "$FRESH_STAGE" "the worker reaped fresh staging that is still in use" +rmdir "$FRESH_STAGE" +pass "abandoned stage litter is reaped by age while fresh staging survives" + +echo "ALL TESTS PASSED" diff --git a/tests/fm-send-remote-delivery.test.sh b/tests/fm-send-remote-delivery.test.sh index 8a686dc9cb0..ee3736b014c 100755 --- a/tests/fm-send-remote-delivery.test.sh +++ b/tests/fm-send-remote-delivery.test.sh @@ -105,6 +105,12 @@ count=$(cat "$FM_SSH_COUNT" 2>/dev/null || echo 0) count=$((count + 1)) printf '%s\n' "$count" > "$FM_SSH_COUNT" printf '%s\n' "$*" >> "$FM_SSH_LOG" +if [ -n "${FM_FAKE_SSH_HANG:-}" ]; then + # A busy remote lane: the transport attempt never returns on its own. The + # real sleep, because the stubbed one on PATH returns immediately. + /bin/sleep "$FM_FAKE_SSH_HANG" + exit 255 +fi if [ "${FM_FAKE_SSH_AFTER_AMBIGUOUS_RC:-0}" -ne 0 ] && [ "$count" -gt 1 ]; then exit "$FM_FAKE_SSH_AFTER_AMBIGUOUS_RC" fi @@ -603,6 +609,95 @@ test_remote_transport_loss_preserves_expectation() { pass "fm-send remote: ssh 255 fails with resend-safe guidance and preserves the expectation" } +test_remote_send_budget_bounds_busy_lane() { + local dir fb ssh_log home rhome rc err began elapsed count pend delivery corr ssh_before + dir="$TMP_ROOT/remote-budget"; mkdir -p "$dir" + fb=$(make_stubs "$dir"); ssh_log="$dir/ssh.log"; : > "$ssh_log" + rhome=$(setup_remote_secondmate_home remote-budget) + home=$(setup_remote_parent_home remote-budget "$rhome") + delivery=aaaabbbbccccdddd + + rc=0 + send_env "$fb" "$home" "$ssh_log" FM_SEND_REMOTE_BUDGET=invalid \ + "$SEND" rsm --key Enter >"$dir/key-invalid.out" 2>"$dir/key-invalid.err" || rc=$? + [ "$rc" -ne 0 ] || fail "an invalid remote key budget must fail" + assert_contains "$(cat "$dir/key-invalid.err")" "must be a positive integer" \ + "an invalid remote key budget must explain its validation failure" + [ ! -f "$ssh_log.count" ] || fail "an invalid remote key budget reached the transport" + + began=$(date +%s) + rc=0 + send_env "$fb" "$home" "$ssh_log" FM_FAKE_SSH_HANG=60 FM_SEND_REMOTE_BUDGET=2 \ + "$SEND" rsm --key Enter >"$dir/key.out" 2>"$dir/key.err" || rc=$? + elapsed=$(( $(date +%s) - began )) + expect_code 1 "$rc" "a bounded remote key must preserve the existing failure contract" + [ "$elapsed" -le 15 ] || fail "the bounded remote key waited ${elapsed}s behind the busy lane" + assert_contains "$(cat "$dir/key.err")" "completion may be unknown" \ + "a bounded remote key failure must preserve its existing diagnostic" + [ "$(cat "$ssh_log.count")" = 1 ] \ + || fail "a bounded remote key must make exactly one transport attempt" + printf '0\n' > "$ssh_log.count" + + # T5: a fire-and-forget send to a mate behind a busy lane returns its + # unconfirmed result within its own budget instead of waiting the lane out. + began=$(date +%s) + rc=0 + send_env "$fb" "$home" "$ssh_log" FM_FAKE_SSH_HANG=60 FM_SEND_REMOTE_BUDGET=2 \ + "$SEND" rsm --fire-and-forget "$delivery" "reconcile your own books" \ + >"$dir/out" 2>"$dir/err" || rc=$? + elapsed=$(( $(date +%s) - began )) + err=$(cat "$dir/err") + expect_code 3 "$rc" "a budget-bounded fire-and-forget send must report unconfirmed: $err" + [ "$elapsed" -le 15 ] || fail "the bounded send waited ${elapsed}s behind the busy lane" + assert_contains "$err" "delivery-id=$delivery" \ + "the bounded unconfirmed result must name the reusable delivery id" + [ "$(cat "$ssh_log.count")" = 1 ] \ + || fail "a budget hit must not retry into the same busy lane, got $(cat "$ssh_log.count") attempts" + + # A retry with the same delivery id against the recovered lane dedups onto + # the same remote record. + send_env "$fb" "$home" "$ssh_log" \ + "$SEND" rsm --fire-and-forget "$delivery" "reconcile your own books" \ + >"$dir/retry.out" 2>"$dir/retry.err" \ + || fail "the same-delivery-id retry after the budget hit failed" + count=$(remote_inbox_records "$rhome" | grep -c . || true) + [ "$count" = 1 ] || fail "the same-delivery-id retry did not dedup onto one record, found $count" + + # A reply-bearing send names the budget and prints the correlation-reusing + # resend command, with the expectation preserved as delivery-unknown. + rc=0 + send_env "$fb" "$home" "$ssh_log" FM_FAKE_SSH_HANG=60 FM_SEND_REMOTE_BUDGET=2 \ + "$SEND" rsm "please rename the metric" >"$dir/reply.out" 2>"$dir/reply.err" || rc=$? + err=$(cat "$dir/reply.err") + [ "$rc" -ne 0 ] || fail "a budget-bounded reply-bearing send must not claim confirmed delivery" + assert_contains "$err" "within its 2s budget" \ + "the budget-bounded failure must name the budget that bounded it" + assert_contains "$err" "Only the correlation-reusing resend below is idempotent" \ + "the budget-bounded failure must print the supported safe resend boundary" + pend=$(pending_record "$home") + [ -n "$pend" ] || fail "a budget-bounded reply-bearing send must preserve its expectation" + [ "$(grep '^phase=' "$pend" | tail -1 | cut -d= -f2-)" = delivery_unknown ] \ + || fail "the preserved expectation must record unknown delivery: $(cat "$pend")" + + # Invalid transport configuration fails before a correlation-reusing resend + # mutates the preserved expectation or reaches the transport. + corr=$(fm_pending_reply_get "$pend" corr_id) + cp "$pend" "$dir/pending-before-invalid-budget" + ssh_before=$(cat "$ssh_log.count") + rc=0 + send_env "$fb" "$home" "$ssh_log" FM_SEND_REMOTE_BUDGET=invalid \ + FM_PENDING_REPLY_EXISTING_CORR="$corr" \ + "$SEND" rsm "please rename the metric" >"$dir/invalid.out" 2>"$dir/invalid.err" || rc=$? + [ "$rc" -ne 0 ] || fail "an invalid remote budget must fail the resend" + assert_contains "$(cat "$dir/invalid.err")" "must be a positive integer" \ + "an invalid remote budget must explain its validation failure" + [ "$(cat "$ssh_log.count")" = "$ssh_before" ] \ + || fail "an invalid remote budget reached the remote transport" + cmp -s "$dir/pending-before-invalid-budget" "$pend" \ + || fail "an invalid remote budget mutated the reusable pending expectation: $(cat "$pend")" + pass "fm-send remote: the remote leg is budget-bounded and stays idempotent across the bound" +} + test_local_secondmate_pending_keeps_expectation_armed() { local dir fb log home rc rec corr dir="$TMP_ROOT/local-pending-expectation"; mkdir -p "$dir" @@ -693,6 +788,7 @@ test_remote_slash_rides_inbox test_remote_real_failure_still_fails test_remote_exit3_no_longer_delivered test_remote_transport_loss_preserves_expectation +test_remote_send_budget_bounds_busy_lane test_local_pending_reports_delivered_unconfirmed test_local_pending_does_not_close_resolve_key test_local_secondmate_pending_keeps_expectation_armed From 420721401c4080d1a4f6982b0ef6769e2a749b23 Mon Sep 17 00:00:00 2001 From: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Date: Fri, 28 Aug 2026 15:56:46 -0700 Subject: [PATCH 07/25] fix(bin): accelerate and bound changed test runs (#3250) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(tests): make the changed-file map select per script and stabilize a budget flake The changed-file map's bin/ fallback resolved a direct test reference to that test's whole FAMILY. bin/fm-push-transition-lib.sh is named by exactly one real-Herdr E2E, so a one-line change to it selected all 12 real-herdr-gated scripts, including a 341s presentation E2E with no dependency on it. Resolve direct test references per script, and keep resolving consumer bin/ scripts through the curated map so recorded family-level coupling survives. Also fix a load-sensitive flake: the tool-update budget deadline is whole-second granular, so a test budget of 1 left headroom anywhere in (0, 1] seconds and the first budget check could already read as exhausted. * feat(bin): make suite wall clock a result and let a family's concurrency be proven --max-wall-ms fails a run whose wall clock exceeds the caller's budget, after reporting the per-script results. A suite that stays green while outgrowing its caller's invocation budget is the regression that got an agent killed mid-run and retried invisibly, so duration has to be a result rather than a log note. --pool on the isolation-proof harness runs the same concurrent proof over a whole family, so 'is this family safe to parallelize?' is answered by a command instead of a guess. Measured watcher-wake-lock and refused it: 3 of 18 scripts fail under concurrency on wall-clock assertions about reaching the next poll. * perf(bin): schedule the changed suite concurrently, longest first The watcher-wake-lock family is proven concurrent-safe (two clean runs, 18 candidates, 0 failures at 4 workers; docs/fm-test-isolation-proof.md), so --changed now schedules its proven-concurrent scripts with bounded parallelism and runs any unproven remainder serially afterwards, never beside them. Concurrent runs are ordered longest-hint-first. Workers are handed scripts in order, so alphabetical order started the 193s fm-watch-triage last and stranded it running alone: 395s wall against a 205s balanced four-worker sum. An explicit --jobs keeps its strict refusal, so every CI lane is unchanged. * fix(bin): bound a hung test instead of letting it hang the suite tests/fm-calm-pi-extension.test.sh was observed running 17+ minutes against a 464ms recorded hint, and the suite had no per-script bound to stop it. An unbounded suite is precisely what silently outruns a caller's invocation budget, and --max-wall-ms is evaluated after the run so it cannot end one that never finishes. --per-script-timeout-secs terminates a script that outruns it and records exit 124, so the run still completes, accounts for the script, and fails. The auto-concurrent --changed path applies 900s, far above the slowest real script (the 341s Herdr presentation E2E), so it only ever converts a hang. * no-mistakes(review): Enforce safe concurrency and descendant timeouts * no-mistakes(review): Validate empty runs and isolation proof pools * no-mistakes(review): Measure selection time in wall budget * no-mistakes(review): Reap interrupted workers and bound finalization * no-mistakes(review): Contain shutdown descendants and watchdog finalization * no-mistakes(review): Honor remaining budget and close launch races * no-mistakes(review): Restore timeout helper and simplify runner cleanup * no-mistakes(review): Record isolation pool admission metadata * no-mistakes(review): Bound Chrome reap and scope proof admission * no-mistakes(review): Align proof scheduling and preserve budget summaries * no-mistakes(review): Remove unreliable finalization watchdog * no-mistakes(review): Freeze budget duration and enforce admission caps * no-mistakes(document): Refresh test runner concurrency documentation * no-mistakes(lint): Fix ShellCheck findings in test runner scripts * no-mistakes(ci): Fixed Greptile’s concurrency-consent finding. `--changed` now remains serial by default; `--changed --jobs auto` explicitly opts into bounded concurrency and the automatic hang timeout. Updated documentation and added behavioral coverage proving serial default behavior, explicit concurrent scheduling, and refusal of `--jobs auto` outside `--changed`. Verified with `bash tests/fm-test-run.test.sh`, `bin/fm-lint.sh`, and `git diff --check` * no-mistakes(review): Restore automatic changed-suite concurrency and timeout * no-mistakes(review): Correct changed-suite contributor guidance * no-mistakes(review): Reject gate-skipped isolation proofs * no-mistakes(review): Correct automatic concurrency evidence * no-mistakes(review): Isolate nested runner process groups * no-mistakes(review): Remove unreliable signal cleanup machinery * no-mistakes(test): Narrow changed-suite selection to executable contract owners * no-mistakes(document): Document isolation proof skip and artifact semantics * no-mistakes(ci): Fixed Greptile’s concurrency-consent finding. `--changed` now remains serial by default; bounded concurrency requires explicit `--jobs auto`. Updated behavioral coverage, contributor guidance, and isolation-proof commands accordingly. Verified with `tests/fm-test-run.test.sh`, `bin/fm-doc-audience-check.sh`, `bin/fm-lint.sh`, Bash syntax checks, and `git diff --check`; all passed * no-mistakes(review): Restore plain changed-suite automatic concurrency * no-mistakes(review): Record resolved changed-suite worker count * fix(bin): keep a runner change selecting its whole curated family A pipeline fix round narrowed the curated changed-file map so bin/fm-test-run.sh and bin/fm-test-isolation-proof.sh selected only their own two contract tests, and the documentation surfaces only the audience test. That cut this branch's own changed selection from 33 scripts to 5. The runner executes every pure-contract-unit script, so its contract test passing proves its logic is right, not that the suite it drives still runs. Narrowing it also makes any wall-clock claim about the changed suite trivially true by not running the work. Only the unmapped bin/* grep fallback resolves per script; curated mappings keep their recorded family coupling. * perf(bin): admit the pure-contract-unit family to bounded concurrency A runner-file change selects pure-contract-unit, so that family decides the changed suite's wall clock. With only watcher-wake-lock admitted, 14 of its 33 selected scripts fell to the serial tail and the selection measured 327.3s against a 300s budget: the concurrent group was 19 scripts totalling 273.4s while the tail alone was 215.7s. bin/fm-test-isolation-proof.sh --pool pure-contract-unit --jobs 4 passes twice, 32 candidates, 0 failures, so the family is admitted on recorded evidence. Full 33-script plain --changed: 327.3s -> 181.8s / 178.5s / 172.7s, 0 failures, inside a 300000ms budget. Also states the per-script guard's derivation. * no-mistakes(review): Align contract-unit concurrency cap with recorded proof * no-mistakes(document): Record final changed-suite performance evidence * fix(bin): keep an empty changed selection clean on stock macOS Bash Under set -u, bash 3.2 treats "${arr[@]}" on an EMPTY array as an unbound-variable error, while bash 4.4+ makes it a harmless no-op. The concurrency work removed the early exit for an empty selection, so execution fell through to the unguarded existence loop: on stock /bin/bash 3.2.57 a contributor who changes only documentation and runs --changed got bin/fm-test-run.sh: line 1713: SCRIPTS[@]: unbound variable with exit 1 and no summary, instead of a clean total=0 pass. Restore the early exit, and guard every remaining array expansion reachable with an empty selection. The reported duration is real elapsed invocation time rather than a hardcoded zero, so a selection phase that outran --max-wall-ms still fails. Verified on this host with /bin/bash 3.2.57: exit 1 with the unbound-variable error before, exit 0 with FM_TEST_SUMMARY total=0 after. * no-mistakes(document): Document shell-bound changed-suite performance --------- Co-authored-by: Kun Chen --- CONTRIBUTING.md | 15 +- bin/fm-test-isolation-proof.sh | 174 +++++++--- bin/fm-test-run.sh | 480 +++++++++++++++++++++++--- docs/fm-test-isolation-proof.json | 1 + docs/fm-test-isolation-proof.md | 64 +++- docs/scripts.md | 4 +- tests/fm-calm-pi-extension.test.sh | 12 +- tests/fm-test-isolation-proof.test.sh | 173 +++++++++- tests/fm-test-run.test.sh | 426 ++++++++++++++++++++++- tests/fm-tool-update-check.test.sh | 10 +- 10 files changed, 1239 insertions(+), 120 deletions(-) diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index ee5824b5354..1dfbebd64dc 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -80,14 +80,17 @@ while IFS= read -r script; do /bin/bash -n "$script" || exit; done < <(bin/fm-li bin/fm-lint.sh # lint that shell surface plus GitHub workflows via pinned actionlint; the single owner CI and the no-mistakes gate both run bin/fm-test-run.sh tests/.test.sh # one script (primary local focus path, timed) bin/fm-test-run.sh --family pure-contract-unit # ordinary family-scoped local path (serial, timed) -bin/fm-test-run.sh --changed # conservative changed-file-informed set (never silent full suite) -bin/fm-test-run.sh --proven-isolated --jobs 4 # explicit local parallel of the proven set only (default is serial) +bin/fm-test-run.sh --changed # normal changed-file-informed path with automatic bounded concurrency +bin/fm-test-run.sh --changed --jobs 1 # explicit serial override +bin/fm-test-run.sh --changed --max-wall-ms 300000 # same automatic path with a post-run five-minute result check +bin/fm-test-run.sh --proven-isolated --jobs 4 # explicit local parallel of the individually proven set bin/fm-test-run.sh --lane portable-serial # portable serial remainder (watcher/AFK/tmux/stateful) bin/fm-test-run.sh --list-lanes # discover exact lane names, including the current CI serial shards bin/fm-test-run.sh --check-coverage # prove portable shards + serial + serial shards + Herdr equal the full inventory bin/fm-test-run.sh --all # deliberate complete regression (optional local full walk; not no-mistakes Test) -bin/fm-test-isolation-proof.sh --list # proven parallel candidate set (Phase 2 owner) -bin/fm-test-isolation-proof.sh --jobs 4 --json /tmp/fm-isolation-proof.json # re-run concurrent isolation proof only +bin/fm-test-isolation-proof.sh --list # proven portable parallel candidate set +bin/fm-test-isolation-proof.sh --jobs 4 --json /tmp/fm-isolation-proof.json # re-run the portable candidate proof +bin/fm-test-isolation-proof.sh --pool watcher-wake-lock --jobs 4 # re-run an admitted family proof [ ! -L CLAUDE.md ] && cmp -s CLAUDE.md - <<'EOF' @AGENTS.md @@ -96,9 +99,9 @@ EOF tmp=$(mktemp -d) && printf 'done: smoke\n' > "$tmp/smoke.status" && FM_STATE_OVERRIDE="$tmp" FM_SIGNAL_GRACE=1 FM_POLL=1 FM_HEARTBEAT=999999 bin/fm-watch-arm.sh # watcher re-arm smoke test (prints arm status, then an actionable signal) ``` -`bin/fm-test-run.sh` is the single owner of behavior-suite selection, portable CI lane composition, optional local `--jobs` for the proven-isolated set only, per-script timing markers, family totals, the coverage guard, and the optional JSON timing artifact. +`bin/fm-test-run.sh` is the single owner of behavior-suite selection, portable CI lane composition, bounded concurrency admission, per-script timing markers, family totals, the coverage guard, and the optional JSON timing artifact. Its header and `--help` own the flags, family labels, lanes, and changed-file map; this section only documents the entry points. -`bin/fm-test-isolation-proof.sh` remains the single owner of the Phase 2 concurrent isolation proof and the exact proven candidate set; see `docs/fm-test-isolation-proof.md`. +`bin/fm-test-isolation-proof.sh` remains the single owner of the portable candidate proof and reusable family proof harness; see `docs/fm-test-isolation-proof.md`. Portable shard balance evidence lives in `docs/fm-test-portable-shards.md`. Local no-mistakes Test stays intent-targeted and must not wire `commands.test` to `--all` or a `tests/*.test.sh` walk. Family selection is the ordinary local path; `--all` is deliberate full regression only. diff --git a/bin/fm-test-isolation-proof.sh b/bin/fm-test-isolation-proof.sh index 137aff8b268..e9ecd53d32d 100755 --- a/bin/fm-test-isolation-proof.sh +++ b/bin/fm-test-isolation-proof.sh @@ -1,25 +1,32 @@ #!/usr/bin/env bash -# fm-test-isolation-proof.sh - bounded concurrent isolation proof for portable -# behavior-test candidates (Phase 2 pre-shard gate). +# fm-test-isolation-proof.sh - bounded concurrent isolation proofs for portable +# behavior-test candidates and selected runner families. # -# This is the single owner of the proven parallel candidate set, the concurrent -# proof run, and the isolation checks that admitted that set. Production -# portable CI shards and bounded local fm-test-run.sh --jobs for this exact set -# are owned by bin/fm-test-run.sh (docs/fm-test-portable-shards.md). +# This is the single owner of the proven portable candidate set, the reusable +# concurrent proof run, and its isolation checks. Production portable CI shards, +# bounded local fm-test-run.sh --jobs admission, and family worker caps are owned +# by bin/fm-test-run.sh (docs/fm-test-portable-shards.md). # -# It does NOT: -# - compose production CI shard membership (fm-test-run.sh owns that partition) -# - run real Herdr, real default-server tmux, watcher lock races, AFK, live -# harnesses, or GUI backends +# It does NOT compose production CI shard membership; fm-test-run.sh owns that +# partition. The default portable pool excludes real Herdr, real default-server +# tmux, watcher lock races, AFK, live harnesses, and GUI backends. A named family +# pool instead runs that family's exact membership and inherits its prerequisites. # # Usage: -# fm-test-isolation-proof.sh [--jobs N] [--json path] [--list] +# fm-test-isolation-proof.sh [--pool ] [--jobs N] [--json path] [--list] # fm-test-isolation-proof.sh --list-exclusions # fm-test-isolation-proof.sh -h | --help # # Options: +# --pool NAME candidate pool: "portable" (default, this harness's own curated +# set) or a bin/fm-test-run.sh family name, to prove a stateful +# family that stays serial on CI but may earn bounded local +# concurrency. bin/fm-test-run.sh's list_concurrent_safe_families +# records which families passed. # --jobs N max concurrent workers (default: 4; min 1) -# --json path write a machine-readable proof artifact after the run +# --json path write a pool-scoped machine-readable proof artifact after the +# run; fm_test_run_jobs_enabled is true only for a successful +# concurrent run within that pool's recorded admission cap # --list print the proven candidate paths (one per line) and exit 0 # --list-exclusions # print basename + reason for scripts deliberately kept serial @@ -40,9 +47,11 @@ # FM_ISOLATION_SUMMARY total= failed= concurrency= duration_ms= # # Exit status is the aggregate of candidate exits: non-zero if any candidate -# fails, if isolation checks fail, or if the candidate set is empty. A script -# that fails only under concurrency must be removed from the candidate set and -# investigated; this harness never retries a failure into green. +# fails, gate-skips (first meaningful line matching ^skip:), if isolation checks +# fail, or if the candidate set is empty. A gate skip names the pool, candidate, +# and missing prerequisite and cannot admit concurrency. A script that fails +# only under concurrency must be removed from the candidate set and investigated; +# this harness never retries a failure into green. set -eu ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" @@ -52,6 +61,7 @@ JOBS=4 JSON_PATH= LIST_ONLY=0 LIST_EXCLUSIONS=0 +POOL=portable usage() { awk ' @@ -218,11 +228,20 @@ global_git_snapshot() { git config --global --list 2>/dev/null | LC_ALL=C sort || true } +detect_gate_skip() { + local file=$1 first + first=$(awk 'NF { print; exit }' "$file" 2>/dev/null || true) + case "$first" in + skip:*) printf '%s\n' "$first" ;; + *) return 1 ;; + esac +} + write_json_artifact() { - local out=$1 started=$2 finished=$3 run_id=$4 total=$5 failed=$6 concurrency=$7 duration=$8 records=$9 - python3 - "$out" "$started" "$finished" "$run_id" "$total" "$failed" "$concurrency" "$duration" "$records" <<'PY' + local out=$1 started=$2 finished=$3 run_id=$4 total=$5 failed=$6 concurrency=$7 duration=$8 records=$9 pool=${10} jobs_enabled=${11} + python3 - "$out" "$started" "$finished" "$run_id" "$total" "$failed" "$concurrency" "$duration" "$records" "$pool" "$jobs_enabled" <<'PY' import json, sys -out, started, finished, run_id, total, failed, concurrency, duration, records_path = sys.argv[1:10] +out, started, finished, run_id, total, failed, concurrency, duration, records_path, pool, jobs_enabled = sys.argv[1:12] scripts = [] with open(records_path, encoding="utf-8") as fh: for line in fh: @@ -242,6 +261,7 @@ doc = { "started_at": started, "finished_at": finished, "kind": "isolation-proof", + "pool": pool, "concurrency": int(concurrency), "summary": { "total": int(total), @@ -250,7 +270,7 @@ doc = { }, "scripts": scripts, "production_sharding_enabled": False, - "fm_test_run_jobs_enabled": False, + "fm_test_run_jobs_enabled": jobs_enabled == "1", } with open(out, "w", encoding="utf-8") as fh: json.dump(doc, fh, indent=2, sort_keys=True) @@ -278,6 +298,15 @@ while [ "$#" -gt 0 ]; do JSON_PATH=${1#--json=} shift ;; + --pool) + [ "$#" -gt 1 ] || die "--pool requires a name (portable, or a family name)" + POOL=$2 + shift 2 + ;; + --pool=*) + POOL=${1#--pool=} + shift + ;; --list) LIST_ONLY=1 shift @@ -309,11 +338,41 @@ if [ "$LIST_EXCLUSIONS" -eq 1 ]; then exit 0 fi +# The portable pool is this harness's own curated set. A family pool proves a +# stateful family that stays serial on CI but may earn bounded local +# concurrency; bin/fm-test-run.sh's list_concurrent_safe_families records which +# families passed. Membership stays empirical: a family that fails here is not +# admitted, and this harness never retries a failure into green. +pool_candidates() { + case "$POOL:$LIST_ONLY" in + portable:1) + list_parallel_candidates + ;; + portable:0) + "$ROOT/bin/fm-test-run.sh" --list-scheduled --proven-isolated + ;; + *:1) + "$ROOT/bin/fm-test-run.sh" --list --family "$POOL" \ + || die "--pool $POOL is not a known family (see bin/fm-test-run.sh --list-families)" + ;; + *) + "$ROOT/bin/fm-test-run.sh" --list-scheduled --family "$POOL" \ + || die "--pool $POOL is not a known family (see bin/fm-test-run.sh --list-families)" + ;; + esac +} + +set +e +candidate_output=$(pool_candidates) +pool_rc=$? +set -e +[ "$pool_rc" -eq 0 ] || exit "$pool_rc" + CANDIDATES=() while IFS= read -r s; do [ -n "$s" ] || continue CANDIDATES+=("$s") -done < <(list_parallel_candidates | LC_ALL=C sort -u) +done < <(printf '%s\n' "$candidate_output" | awk '!seen[$0]++') if [ "$LIST_ONLY" -eq 1 ]; then for s in "${CANDIDATES[@]+"${CANDIDATES[@]}"}"; do @@ -348,14 +407,15 @@ printf 'FM_ISOLATION_BEGIN %s concurrency=%s candidates=%s\n' \ # Worker state arrays parallel to CANDIDATES indices (1-based worker labels). declare -a WORKER_PIDS=() declare -a WORKER_IDX=() +ACTIVE_WORKERS=0 wait_one_slot() { - local pid idx work rc duration script mode - # Wait for the oldest launched worker still recorded. - pid=${WORKER_PIDS[0]} - idx=${WORKER_IDX[0]} - WORKER_PIDS=("${WORKER_PIDS[@]:1}") - WORKER_IDX=("${WORKER_IDX[@]:1}") + local slot=$1 pid idx work rc duration script mode gate_skip + pid=${WORKER_PIDS[$slot]} + idx=${WORKER_IDX[$slot]} + unset 'WORKER_PIDS[slot]' + unset 'WORKER_IDX[slot]' + ACTIVE_WORKERS=$((ACTIVE_WORKERS - 1)) set +e wait "$pid" set -e @@ -363,6 +423,10 @@ wait_one_slot() { script=${CANDIDATES[$((idx - 1))]} rc=$(cat "$work/out/exit" 2>/dev/null || echo 1) duration=$(cat "$work/out/duration_ms" 2>/dev/null || echo 0) + if [ "$rc" -eq 0 ] && gate_skip=$(detect_gate_skip "$work/out/output"); then + rc=1 + log "pool $POOL candidate gate-skipped without proving concurrency: $script: $gate_skip" + fi printf 'FM_ISOLATION_CANDIDATE_END %s %s exit=%s duration_ms=%s worker=%s\n' \ "$(now_iso)" "$script" "$rc" "$duration" "$idx" printf '%s\t%s\t%s\t%s\n' "$script" "$rc" "$duration" "$idx" >>"$RECORDS" @@ -370,13 +434,9 @@ wait_one_slot() { FAILED=$((FAILED + 1)) AGG_RC=1 log "candidate failed: $script exit=$rc" - if [ -s "$work/out/stdout" ]; then - log "--- stdout ($script) ---" - tail -n 40 "$work/out/stdout" >&2 || true - fi - if [ -s "$work/out/stderr" ]; then - log "--- stderr ($script) ---" - tail -n 40 "$work/out/stderr" >&2 || true + if [ -s "$work/out/output" ]; then + log "--- output ($script) ---" + tail -n 40 "$work/out/output" >&2 || true fi fi # Isolation: worker root must remain mode 0700 and under the proof parent. @@ -398,6 +458,29 @@ wait_one_slot() { esac } +worker_pid_is_running() { + local want=$1 running inventory="$PROOF_ROOT/running-pids" + jobs -r -p >"$inventory" + while IFS= read -r running; do + [ "$running" = "$want" ] && return 0 + done <"$inventory" + return 1 +} + +wait_one_completed_slot() { + local slot work + while :; do + for slot in "${!WORKER_PIDS[@]}"; do + work="$PROOF_ROOT/w${WORKER_IDX[$slot]}" + if [ -f "$work/out/exit" ] || ! worker_pid_is_running "${WORKER_PIDS[$slot]}"; then + wait_one_slot "$slot" + return + fi + done + sleep 0.01 + done +} + idx=0 for script in "${CANDIDATES[@]}"; do idx=$((idx + 1)) @@ -429,7 +512,7 @@ for script in "${CANDIDATES[@]}"; do FM_PROJECTS_OVERRIDE FM_CONFIG_OVERRIDE FM_BACKEND 2>/dev/null || true cd "$ROOT" || exit 1 begin_ms=$(now_ms) - bash "$script" >"$work/out/stdout" 2>"$work/out/stderr" + bash "$script" >"$work/out/output" 2>&1 rc=$? end_ms=$(now_ms) duration=$((end_ms - begin_ms)) @@ -440,17 +523,18 @@ for script in "${CANDIDATES[@]}"; do printf '%s\n' "$duration" >"$work/out/duration_ms" exit 0 ) & - WORKER_PIDS+=("$!") - WORKER_IDX+=("$idx") + WORKER_PIDS[idx]=$! + WORKER_IDX[idx]=$idx + ACTIVE_WORKERS=$((ACTIVE_WORKERS + 1)) # Bound concurrency. - while [ "${#WORKER_PIDS[@]}" -ge "$JOBS" ]; do - wait_one_slot + while [ "$ACTIVE_WORKERS" -ge "$JOBS" ]; do + wait_one_completed_slot done done -while [ "${#WORKER_PIDS[@]}" -gt 0 ]; do - wait_one_slot +while [ "$ACTIVE_WORKERS" -gt 0 ]; do + wait_one_completed_slot done GIT_AFTER=$(global_git_snapshot) @@ -487,9 +571,17 @@ if [ -n "$JSON_PATH" ]; then mkdir -p "$(dirname "$JSON_PATH")" # Stable record order for the artifact. sort -t$'\t' -k1,1 "$RECORDS" -o "$RECORDS" + jobs_enabled=0 + jobs_max=0 + if "$ROOT/bin/fm-test-run.sh" --list-concurrent-safe-families | grep -Fxq "$POOL"; then + jobs_max=$("$ROOT/bin/fm-test-run.sh" --concurrent-safe-family-jobs-max "$POOL") + fi + if [ "$AGG_RC" -eq 0 ] && [ "$JOBS" -gt 1 ] && [ "$JOBS" -le "$jobs_max" ]; then + jobs_enabled=1 + fi write_json_artifact "$JSON_PATH" \ "$RUN_STARTED_ISO" "$RUN_FINISHED_ISO" "$RUN_ID" \ - "$TOTAL" "$FAILED" "$JOBS" "$RUN_DURATION" "$RECORDS" + "$TOTAL" "$FAILED" "$JOBS" "$RUN_DURATION" "$RECORDS" "$POOL" "$jobs_enabled" log "wrote isolation proof artifact: $JSON_PATH" fi diff --git a/bin/fm-test-run.sh b/bin/fm-test-run.sh index c89b6e60f83..cd0d9e28aa8 100755 --- a/bin/fm-test-run.sh +++ b/bin/fm-test-run.sh @@ -1,6 +1,6 @@ #!/usr/bin/env bash # fm-test-run.sh - single owner of Firstmate's behavior-test runner, lane -# composition for portable CI shards, local --jobs for the proven-isolated set, +# composition for portable CI shards, local --jobs for proven-concurrent work, # timing markers, and the complete-regression coverage guard. # # Selection modes (exactly one of: --all, --family, --changed, --lane, @@ -17,7 +17,10 @@ # fm-test-run.sh --list --all # fm-test-run.sh --list --family # fm-test-run.sh --list --lane portable-parallel-1 +# fm-test-run.sh --list-scheduled --family # fm-test-run.sh --list-families +# fm-test-run.sh --list-concurrent-safe-families +# fm-test-run.sh --concurrent-safe-family-jobs-max # fm-test-run.sh --list-lanes # fm-test-run.sh --check-coverage # @@ -27,6 +30,8 @@ # Options: # --json write a deterministic timing artifact after the run # --list print selected script paths (one per line) and exit 0 +# --list-scheduled +# print selected paths longest-hint-first and exit 0 # --base with --changed, compare against this ref (default: origin/main) # --exclude-family # drop scripts whose primary family matches after selection @@ -38,10 +43,34 @@ # The required Herdr CI lane uses this so a missing pin cannot # silently pass as a gate skip. # --jobs N run the selected scripts with up to N concurrent workers. -# Default is 1 (serial). N>1 is allowed only when every -# selected script is in the proven-isolated set -# (bin/fm-test-isolation-proof.sh --list). Cap is 8. Stateful -# families never schedule under --jobs. +# Plain --changed uses min(4, cpus) workers when multiple +# selected scripts are admissible. +# N>1 is allowed only when every selected script is proven +# safe to run concurrently: individually in the proven-isolated +# set (bin/fm-test-isolation-proof.sh --list), or in a family +# carrying a recorded concurrent proof +# (list_concurrent_safe_families below). Overall cap is 8; +# family proofs may impose a lower cap. Unproven stateful +# scripts stay serial. Concurrent runs are ordered +# longest-hint-first so the slowest script is not stranded +# alone at the tail. Default is 1 (serial) except for plain +# --changed, which uses the bounded automatic scheduler. Any +# unproven remainder runs serially after that group. +# --per-script-timeout-secs N +# terminate a script that runs longer than N seconds and +# record it as exit 124 (0 disables, the default). The +# --changed applies 900s automatically: no real script +# approaches it, so it only converts a HUNG +# script into a bounded failure. --max-wall-ms is checked +# after the run and so cannot catch a hang on its own. +# External interruption cleanup is outside this runner's +# guarantee; configured per-script bounds remain authoritative. +# --max-wall-ms N fail the run when its measured invocation wall clock exceeds +# N milliseconds, including an empty selection. It is +# evaluated after selection and suite execution and cannot +# interrupt a running script; per-script hangs are +# bounded by --per-script-timeout-secs. Pathological output +# sinks that block finalization are explicitly out of scope. # -h, --help print this header # # Per-script machine-parseable markers (stdout): @@ -52,9 +81,12 @@ # FM_TEST_SUMMARY total= failed= skipped_gate= duration_ms= # FM_TEST_SUMMARY_FAMILY family= count= duration_ms= failed= # FM_TEST_SLOWEST rank= script= duration_ms= +# FM_TEST_BUDGET max_wall_ms= duration_ms= (only with --max-wall-ms) # -# Exit status is non-zero if any selected script exits non-zero or a configured -# --fail-on-gate-skip token appears. Other gate skips (first meaningful line +# Exit status is non-zero if any selected script exits non-zero, a configured +# --fail-on-gate-skip token appears, the measured duration exceeds +# --max-wall-ms, timing-artifact finalization fails, or a concurrent worker +# violates its isolation check. Other gate skips (first meaningful line # matching ^skip:) remain successful and are counted as skipped_gate. # # Family labels, the changed-file map, and production portable-shard composition @@ -67,15 +99,32 @@ # share a machine. This script owns : a lane whose disagrees with the # configured shard count is refused, so a CI matrix cannot silently drop a shard. # --changed is conservative: it over-selects related families rather than -# under-selecting, and never expands to the complete suite unless --all. +# under-selecting, and never expands to the complete suite unless --all. The one +# place it is deliberately narrow is a bin/ path with no curated family: a test +# that names it is selected as that SCRIPT, because the reference is per-script +# evidence. Consumer bin/ scripts still resolve through the curated map, so +# recorded family-level coupling still expands to the whole family. set -eu +now_ms() { + if command -v python3 >/dev/null 2>&1; then + python3 -c 'import time; print(int(time.time() * 1000))' + else + echo $(($(date +%s) * 1000)) + fi +} + +RUN_STARTED_ISO=$(date -u +%Y-%m-%dT%H:%M:%SZ) +RUN_STARTED_MS=$(now_ms) + ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" cd "$ROOT" || exit 1 MODE= LIST_ONLY=0 +LIST_SCHEDULED=0 LIST_FAMILIES=0 +LIST_CONCURRENT_SAFE_FAMILIES=0 LIST_LANES=0 CHECK_COVERAGE=0 AGGREGATE_OUT= @@ -87,7 +136,20 @@ SCRIPTS=() EXCLUDE_FAMILIES=() FAIL_ON_GATE_SKIP= JOBS=1 +JOBS_EXPLICIT=0 JOBS_MAX=8 +MAX_WALL_MS= +PER_SCRIPT_TIMEOUT_SECS=0 +# Bound applied automatically on the automatic --changed path, derived from +# measured healthy runtimes with margin rather than picked: the slowest measured +# behavior test is the 341s Herdr presentation E2E, and the slowest script in a +# runner-file changed selection is tests/fm-calm-pi-extension.test.sh at 77s +# once its Chrome reap terminates. 900s leaves roughly 2.6x headroom over the +# slowest real script, so this can only ever fire on a script that is genuinely +# stuck. It is a guard, not a speed control: a HUNG script becomes a bounded +# failure instead of an unbounded suite, which is the shape that silently +# outruns a caller's invocation budget. +CHANGED_DEFAULT_TIMEOUT_SECS=900 # How many separate-runner shards the portable serial remainder splits into. # One owner: CI lane names carry this count and are refused when they disagree. @@ -119,13 +181,14 @@ now_iso() { date -u +%Y-%m-%dT%H:%M:%SZ } -now_ms() { - if command -v python3 >/dev/null 2>&1; then - python3 -c 'import time; print(int(time.time() * 1000))' - else - # Second precision only when python3 is unavailable. - echo $(($(date +%s) * 1000)) - fi +cpu_count() { + local n + n=$(getconf _NPROCESSORS_ONLN 2>/dev/null || sysctl -n hw.ncpu 2>/dev/null || echo 1) + case "$n" in + ''|*[!0-9]*) n=1 ;; + esac + [ "$n" -ge 1 ] || n=1 + printf '%s\n' "$n" } # Primary family for one tests/*.test.sh basename. Unmapped scripts are @@ -351,6 +414,50 @@ tests/fm-composer-lib.test.sh EOF } +# Families whose scripts are proven safe to run concurrently WITH EACH OTHER +# under the bounded local scheduler. Deliberately separate from the +# proven-isolated set, which must stay exactly equal to the portable CI shard +# union (see the coverage guard); these families keep their serial CI lane and +# only gain concurrency for a local run. +# +# Membership is empirical, never assumed: +# `bin/fm-test-isolation-proof.sh --pool --jobs 4` is the owner of the +# proof, and docs/fm-test-isolation-proof.md records the dated result. +list_concurrent_safe_families() { + cat <<'EOF' +watcher-wake-lock +pure-contract-unit +EOF +} + +family_is_concurrent_safe() { + local want=$1 line + while IFS= read -r line; do + [ "$line" = "$want" ] && return 0 + done < <(list_concurrent_safe_families) + return 1 +} + +concurrent_safe_family_jobs_max() { + case "$1" in + watcher-wake-lock|pure-contract-unit) printf '4\n' ;; + *) printf '1\n' ;; + esac +} + +# A script may run under --jobs when it is individually proven isolated or is +# an exact repository member of a family carrying a recorded concurrent proof. +script_allows_concurrency() { + local s=$1 family repo_script + is_proven_isolated_script "$s" && return 0 + family=$(family_for_basename "$(basename "$s")") + family_is_concurrent_safe "$family" || return 1 + while IFS= read -r repo_script; do + [ "$repo_script" = "$s" ] && return 0 + done < <(all_repo_tests) + return 1 +} + is_proven_isolated_script() { local want=$1 line while IFS= read -r line; do @@ -679,7 +786,7 @@ run_coverage_guard() { rm -rf "$tmp" return 1 fi - printf '%s\n' "${SCRIPTS[@]}" >>"$tmp/serial_shards_raw" + printf '%s\n' "${SCRIPTS[@]+"${SCRIPTS[@]}"}" >>"$tmp/serial_shards_raw" shard=$((shard + 1)) done SCRIPTS=() @@ -894,14 +1001,66 @@ families_for_test_reference() { [ "$found" -eq 1 ] } +# Tests that name , selected as individual scripts rather than widened +# to each referencing test's whole family. A direct reference is per-script +# evidence, so it selects per script: one real-Herdr E2E sourcing a shared +# helper must not drag in every other script of that expensive family. +scripts_for_test_reference() { + local needle=$1 s + local found=0 + while IFS= read -r s; do + [ -n "$s" ] || continue + if grep -Fq "$needle" "$s"; then + printf '__script__:%s\n' "$(basename "$s")" + found=1 + fi + done < <(all_repo_tests) + [ "$found" -eq 1 ] +} + +# bin/ scripts other than itself that name . +bin_consumers_of() { + local needle=$1 b + for b in bin/*.sh bin/backends/*.sh; do + [ -f "$b" ] || continue + [ "$(basename "$b")" = "$needle" ] || ! grep -Fq "$needle" "$b" || printf '%s\n' "$b" + done +} + +# An unmapped bin/ path has no curated family of its own. Its blast radius is +# the tests that name it, plus the curated families of the bin/ scripts that +# consume it. Direct test references resolve per script (above) while consumer +# scripts resolve back through the curated map, so genuine family-level +# coupling a maintainer recorded is preserved while an incidental single-script +# reference no longer selects that script's whole family. +BIN_FALLBACK_DEPTH=0 +families_for_unmapped_bin() { + local path=$1 needle consumer out found=0 + needle=$(basename "$path") + if out=$(scripts_for_test_reference "$needle"); then + printf '%s\n' "$out" + found=1 + fi + if [ "$BIN_FALLBACK_DEPTH" -lt 2 ]; then + BIN_FALLBACK_DEPTH=$((BIN_FALLBACK_DEPTH + 1)) + while IFS= read -r consumer; do + [ -n "$consumer" ] || continue + out=$(families_for_changed_path "$consumer" | grep -v '^__unmapped__:' || true) + if [ -n "$out" ]; then + printf '%s\n' "$out" + found=1 + fi + done < <(bin_consumers_of "$needle") + BIN_FALLBACK_DEPTH=$((BIN_FALLBACK_DEPTH - 1)) + fi + [ "$found" -eq 1 ] +} + # Conservative path → family map. Over-selects rather than under-selects. # Never expands to the complete suite. families_for_changed_path() { local path=$1 fixture_ref case "$path" in - tests/fm-test-run.test.sh) - printf '%s\n' pure-contract-unit - ;; tests/fm-backend-herdr-eventwait.test.py) printf '%s\n' real-herdr-gated printf '%s\n' backend-dispatch @@ -912,6 +1071,10 @@ families_for_changed_path() { printf '%s\n' "__script__:$(basename "$path")" ;; bin/fm-test-run.sh|bin/fm-test-isolation-proof.sh) + # Deliberately the WHOLE family, not just the two contract tests. This + # runner executes every pure-contract-unit script, so a change to it is + # only proven by running them: its own contract test passing says the + # runner's logic is right, not that the suite it drives still runs. printf '%s\n' pure-contract-unit ;; bin/backends/herdr*|bin/fm-herdr-lab.sh|tests/herdr-test-safety.sh) @@ -1077,7 +1240,7 @@ families_for_changed_path() { # the fixture case above applies. Refusing on its absent mapping would # make every retirement branch unable to select its changed tests. if [ -e "$path" ]; then - families_for_test_reference "$(basename "$path")" \ + families_for_unmapped_bin "$path" \ || printf '%s\n' "__unmapped__:$path" fi ;; @@ -1178,7 +1341,7 @@ apply_exclude_families() { for s in "${SCRIPTS[@]+"${SCRIPTS[@]}"}"; do fam=$(family_for_basename "$(basename "$s")") keep=1 - for ex in "${EXCLUDE_FAMILIES[@]}"; do + for ex in "${EXCLUDE_FAMILIES[@]+"${EXCLUDE_FAMILIES[@]}"}"; do if [ "$fam" = "$ex" ]; then keep=0 break @@ -1325,20 +1488,57 @@ while [ "$#" -gt 0 ]; do --jobs) [ "$#" -gt 1 ] || die "--jobs requires a positive integer" JOBS=$2 + JOBS_EXPLICIT=1 shift 2 ;; --jobs=*) JOBS=${1#--jobs=} + JOBS_EXPLICIT=1 + shift + ;; + --max-wall-ms) + [ "$#" -gt 1 ] || die "--max-wall-ms requires a positive integer" + MAX_WALL_MS=$2 + shift 2 + ;; + --max-wall-ms=*) + MAX_WALL_MS=${1#--max-wall-ms=} + shift + ;; + --per-script-timeout-secs) + [ "$#" -gt 1 ] || die "--per-script-timeout-secs requires a whole number of seconds" + PER_SCRIPT_TIMEOUT_SECS=$2 + shift 2 + ;; + --per-script-timeout-secs=*) + PER_SCRIPT_TIMEOUT_SECS=${1#--per-script-timeout-secs=} shift ;; --list) LIST_ONLY=1 shift ;; + --list-scheduled) + LIST_SCHEDULED=1 + shift + ;; --list-families) LIST_FAMILIES=1 shift ;; + --list-concurrent-safe-families) + LIST_CONCURRENT_SAFE_FAMILIES=1 + shift + ;; + --concurrent-safe-family-jobs-max) + [ "$#" -gt 1 ] || die "--concurrent-safe-family-jobs-max requires a family name" + concurrent_safe_family_jobs_max "$2" + exit 0 + ;; + --concurrent-safe-family-jobs-max=*) + concurrent_safe_family_jobs_max "${1#--concurrent-safe-family-jobs-max=}" + exit 0 + ;; --list-lanes) LIST_LANES=1 shift @@ -1406,6 +1606,11 @@ if [ "$LIST_FAMILIES" -eq 1 ]; then exit 0 fi +if [ "$LIST_CONCURRENT_SAFE_FAMILIES" -eq 1 ]; then + list_concurrent_safe_families + exit 0 +fi + if [ "$LIST_LANES" -eq 1 ]; then list_known_lanes exit 0 @@ -1432,6 +1637,17 @@ esac [ "$JOBS" -ge 1 ] || die "--jobs must be >= 1" [ "$JOBS" -le "$JOBS_MAX" ] || die "--jobs is capped at $JOBS_MAX (got $JOBS)" +if [ -n "$MAX_WALL_MS" ]; then + case "$MAX_WALL_MS" in + ''|*[!0-9]*) die "--max-wall-ms requires a positive integer" ;; + esac + [ "$MAX_WALL_MS" -gt 0 ] || die "--max-wall-ms requires a positive integer" +fi + +case "$PER_SCRIPT_TIMEOUT_SECS" in + ''|*[!0-9]*) die "--per-script-timeout-secs requires a whole number of seconds (0 disables)" ;; +esac + case "${MODE:-}" in all) select_all @@ -1455,7 +1671,7 @@ case "${MODE:-}" in ;; scripts) # Normalize and re-add through add_script for consistent paths. - raw=("${SCRIPTS[@]}") + raw=("${SCRIPTS[@]+"${SCRIPTS[@]}"}") SCRIPTS=() for s in "${raw[@]}"; do add_script "$s" @@ -1474,31 +1690,54 @@ fi if [ -n "$FAIL_ON_GATE_SKIP" ]; then SELECTION_DESC="${SELECTION_DESC};fail-on-gate-skip=$FAIL_ON_GATE_SKIP" fi -if [ "$JOBS" -gt 1 ]; then - SELECTION_DESC="${SELECTION_DESC};jobs=$JOBS" -fi - -if [ "$LIST_ONLY" -eq 1 ]; then - for s in "${SCRIPTS[@]+"${SCRIPTS[@]}"}"; do - printf '%s\n' "$s" - done +if [ "$LIST_ONLY" -eq 1 ] || [ "$LIST_SCHEDULED" -eq 1 ]; then + if [ "$LIST_SCHEDULED" -eq 1 ]; then + for s in "${SCRIPTS[@]+"${SCRIPTS[@]}"}"; do + printf '%s\t%s\n' "$(portable_serial_weight_for "$s")" "$s" + done | LC_ALL=C sort -t"$(printf '\t')" -k1,1nr -k2,2 | cut -f2- + else + for s in "${SCRIPTS[@]+"${SCRIPTS[@]}"}"; do + printf '%s\n' "$s" + done + fi exit 0 fi +# An empty selection is a clean result, not a no-op that falls through. Exiting +# here also keeps every array expansion below off the empty-array path: under +# `set -u`, bash 3.2 (the stock macOS shell) treats "${arr[@]}" on an empty +# array as an unbound-variable error, while bash 4.4+ makes it a harmless no-op. +# A contributor on stock macOS who changes only documentation must still get +# total=0 and exit 0 rather than a crash. if [ "${#SCRIPTS[@]}" -eq 0 ]; then log "nothing to run" - printf 'FM_TEST_SUMMARY total=0 failed=0 skipped_gate=0 duration_ms=0\n' + empty_finished_ms=$(now_ms) + empty_duration=$((empty_finished_ms - RUN_STARTED_MS)) + [ "$empty_duration" -ge 0 ] || empty_duration=0 + empty_rc=0 + printf 'FM_TEST_SUMMARY total=0 failed=0 skipped_gate=0 duration_ms=%s\n' "$empty_duration" + # The budget covers the whole invocation, so a selection phase that outran it + # still fails - reporting zero work is not the same as reporting no time. + if [ -n "$MAX_WALL_MS" ]; then + printf 'FM_TEST_BUDGET max_wall_ms=%s duration_ms=%s\n' "$MAX_WALL_MS" "$empty_duration" + if [ "$empty_duration" -gt "$MAX_WALL_MS" ]; then + log "wall-clock budget exceeded: ${empty_duration}ms > ${MAX_WALL_MS}ms for $SELECTION_DESC" + empty_rc=1 + fi + fi if [ -n "$JSON_PATH" ]; then empty_rec=$(mktemp) empty_fam=$(mktemp) : >"$empty_rec" : >"$empty_fam" - started=$(now_iso) + empty_finished_iso=$(now_iso) mkdir -p "$(dirname "$JSON_PATH")" - write_json_artifact "$JSON_PATH" "$started" "$started" "empty" 0 0 0 0 "$SELECTION_DESC" "$empty_rec" "$empty_fam" + write_json_artifact "$JSON_PATH" "$RUN_STARTED_ISO" "$empty_finished_iso" \ + "fm-test-run-${RUN_STARTED_MS}-$$" 0 0 0 "$empty_duration" \ + "$SELECTION_DESC" "$empty_rec" "$empty_fam" rm -f "$empty_rec" "$empty_fam" fi - exit 0 + exit "$empty_rc" fi # Verify selected scripts exist before starting. @@ -1507,23 +1746,96 @@ for s in "${SCRIPTS[@]}"; do [ -x "$s" ] || [ -r "$s" ] || die "test script not readable: $s" done -# --jobs N>1 only for the proven-isolated set. Stateful families stay serial. -if [ "$JOBS" -gt 1 ]; then +# Plain --changed uses the bounded representative-suite scheduler; numeric +# --jobs retains the strict all-script admission rule below. +AUTO_CONCURRENCY=0 +if [ "$MODE" = changed ] && [ "$JOBS_EXPLICIT" -eq 0 ]; then + if [ "${#SCRIPTS[@]}" -gt 0 ] && [ "$PER_SCRIPT_TIMEOUT_SECS" -eq 0 ]; then + PER_SCRIPT_TIMEOUT_SECS=$CHANGED_DEFAULT_TIMEOUT_SECS + fi + auto_admissible=0 + for s in "${SCRIPTS[@]}"; do + script_allows_concurrency "$s" && auto_admissible=$((auto_admissible + 1)) + done + if [ "$auto_admissible" -gt 1 ]; then + JOBS=$(cpu_count) + [ "$JOBS" -le 4 ] || JOBS=4 + [ "$JOBS" -ge 1 ] || JOBS=1 + [ "$JOBS" -eq 1 ] || AUTO_CONCURRENCY=1 + fi +fi +if [ "$JOBS" -gt 1 ] || [ "$MODE" = changed ]; then + SELECTION_DESC="${SELECTION_DESC};jobs=$JOBS" +fi + +# An explicit --jobs names a concurrency for exactly the selection given, so an +# unproven script in it is a refusal rather than something to schedule around. +if [ "$JOBS" -gt 1 ] && [ "$AUTO_CONCURRENCY" -eq 0 ]; then for s in "${SCRIPTS[@]}"; do + if ! script_allows_concurrency "$s"; then + die "--jobs $JOBS refused: $s is not in the proven-isolated set (see bin/fm-test-isolation-proof.sh --list) and its family has no recorded concurrent proof. Unproven stateful scripts stay serial." + fi if ! is_proven_isolated_script "$s"; then - die "--jobs $JOBS refused: $s is not in the proven-isolated set (see bin/fm-test-isolation-proof.sh --list). Stateful families stay serial." + family=$(family_for_basename "$(basename "$s")") + family_jobs_max=$(concurrent_safe_family_jobs_max "$family") + [ "$JOBS" -le "$family_jobs_max" ] \ + || die "--jobs $JOBS refused: family $family is proven only up to $family_jobs_max concurrent workers" fi done fi +# Split the run into the proven-concurrent scripts and an unproven remainder. +# The remainder runs serially AFTER the concurrent group, never beside it, so an +# unproven script still never shares a machine with another test. An explicit +# --jobs refused above, so its remainder is always empty. +CONCURRENT_SCRIPTS=() +SERIAL_TAIL_SCRIPTS=() +if [ "$JOBS" -gt 1 ]; then + SCHEDULE_TMP=$(mktemp "${TMPDIR:-/tmp}/fm-test-sched.XXXXXX") + : >"$SCHEDULE_TMP" + # Two passes: the tail array must be built in this shell, so the weighted + # listing is written to a file rather than piped into sort from a loop whose + # appends would be lost in a subshell. + for s in "${SCRIPTS[@]}"; do + if script_allows_concurrency "$s"; then + # Longest first: workers are handed scripts in order, so starting the + # longest last strands it running alone at the tail. Measured over the + # watcher family, alphabetical order finished in 395s where the balanced + # four-worker sum was 205s. + printf '%s\t%s\n' "$(portable_serial_weight_for "$s")" "$s" >>"$SCHEDULE_TMP" + else + SERIAL_TAIL_SCRIPTS+=("$s") + fi + done + while IFS=$'\t' read -r _weight s; do + [ -n "$s" ] || continue + CONCURRENT_SCRIPTS+=("$s") + done < <(LC_ALL=C sort -t"$(printf '\t')" -k1,1nr -k2,2 "$SCHEDULE_TMP") + rm -f "$SCHEDULE_TMP" +fi + +if [ "$PER_SCRIPT_TIMEOUT_SECS" -gt 0 ]; then + [ -r "$ROOT/bin/fm-timeout-lib.sh" ] || die "per-script timeout helper not found: bin/fm-timeout-lib.sh" + # shellcheck source=bin/fm-timeout-lib.sh + . "$ROOT/bin/fm-timeout-lib.sh" +fi + RUN_TMP=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run.XXXXXX") RECORDS="$RUN_TMP/records.tsv" FAMILIES_TSV="$RUN_TMP/families.tsv" : >"$RECORDS" -trap 'rm -rf "$RUN_TMP"' EXIT +declare -a WORKER_PIDS=() +declare -a WORKER_IDX=() +declare -a WORKER_SCRIPTS=() + +# Invoked indirectly by the EXIT trap below. +# shellcheck disable=SC2329 +cleanup_run() { + rm -rf "$RUN_TMP" +} + +trap cleanup_run EXIT -RUN_STARTED_ISO=$(now_iso) -RUN_STARTED_MS=$(now_ms) RUN_ID="fm-test-run-${RUN_STARTED_MS}-$$" TOTAL=0 FAILED=0 @@ -1595,6 +1907,42 @@ record_script_result() { TOTAL=$((TOTAL + 1)) } +# Run