Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
27 changes: 26 additions & 1 deletion bin/fm-classify-lib.sh
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,10 @@
# to decide whether a crew that just stopped its turn or went stale is working,
# deliberately paused, or neither. Callers run it ONLY on no-verb signal handling
# and first sighting of a stale hash, never on every wake, so the per-wake triage
# stays cheap. status_open_decisions_incremental (see "incremental (cursor-backed)
# stays cheap. crew_run_step_paused makes that same bounded fm-crew-state.sh call
# a second time, but only inside the already-rare declared-pause branch its
# caller gates it behind, to tell a run-step-corroborated pause from a
# status-log-only one. status_open_decisions_incremental (see "incremental (cursor-backed)
# open-decisions fold" below) also writes: it persists a per-status-file byte
# cursor and folded open-set as a side effect, so a per-drain fleet-wide scan
# stays bounded by new appends instead of re-reading each task's whole lifetime
Expand Down Expand Up @@ -1309,6 +1312,28 @@ crew_is_paused() { # <id>
[ "$(crew_absorb_class "$1")" = paused ]
}

# 0 if crew <id>'s declared pause/captain-held wait is attributed to the
# run-step (fm-crew-state.sh's ci-monitoring-only reconciliation), as opposed
# to the no-run status-log fallback. A run-step pause has already been
# corroborated against the pipeline's own state - no gate, not fixing, no
# independently-resolved outcome - so fm-watch.sh's pause_state_class trusts
# it outright, live agent or not. A no-run fallback pause carries no such
# corroboration and instead still needs pause_state_class's agent-liveness
# recovery: a still-alive agent may be sitting at an undeclared live
# interactive gate rather than the wait it named. A second $FM_CREW_STATE_BIN
# read (crew_absorb_class already made one), but only inside the already-rare
# declared-pause branch its caller gates this behind, never per poll.
crew_run_step_paused() { # <id>
local id=$1 line state src
[ -n "$id" ] || return 1
line=$("$FM_CREW_STATE_BIN" "$id" 2>/dev/null) || return 1
case "$line" in state:*) ;; *) return 1 ;; esac
state=${line#state: }; state=${state%% *}
[ "$state" = paused ] || return 1
src=${line#*source: }; src=${src%% *}
[ "$src" = run-step ]
}

# Directories excluded from the worktree write probe below, and the depth it walks.
# The excluded set is everything a supervisor read or a package manager can write
# without the crew doing any work - .git first, so firstmate's own read-only git
Expand Down
24 changes: 23 additions & 1 deletion bin/fm-crew-state.sh
Original file line number Diff line number Diff line change
Expand Up @@ -48,7 +48,11 @@
# 3. Reconcile the status log: if its last line says needs-decision/blocked but
# the run-step shows the run moved on, the log is deterministically stale and
# is flagged superseded. A genuinely parked run plus a needs-decision log
# agree, and are reported as parked.
# agree, and are reported as parked. A last line that instead declares
# paused:/captain-held while the run's only remaining step is CI
# monitoring (never fixing, never a gate) wins over the run-step's own
# working/done reading and is reported as paused, since that is a
# declared external wait, not active work.
# 4. No run for this crew (pre-validation, or kind=scout): fall back to the
# recorded backend's pane busy state, then the status log's last line only
# when its verb maps to a recognized run-state. Decision-only events such as
Expand Down Expand Up @@ -585,6 +589,24 @@ if [ "$HAVE_RUN" = 1 ]; then
;;
esac
fi
# A crew whose only remaining pipeline activity is CI monitoring (no
# gate, not fixing, no independently-resolved outcome) is, once it has
# declared paused:/captain-held itself, intentionally waiting on an
# external dependency - upstream checks that may never post without
# maintainer approval, or the captain's own merge - not "still
# validating". Root cause of the 2026-08-27 fm-parked-run-wedge-alarm
# incident: this run-step reading otherwise stays working/done and
# outranks the declared wait every poll, so the watcher wedge-escalates
# a legitimately parked task forever. Restricting the override to the ci
# step specifically (never fixing, never a gate) keeps a genuinely
# active run's own state authoritative - a declared pause there is
# stale and must not silence it.
if [ "$CI_STEP_STATUS" = running ] \
&& [ "$has_gate" = 0 ] && [ -z "$outcome" ] \
&& status_is_paused_or_captain_held "$LOG_LINE"; then
RUN_STATE=paused
RUN_DETAIL="declared wait (ci monitoring): $RUN_DETAIL"
fi
fi
fi

Expand Down
80 changes: 60 additions & 20 deletions bin/fm-watch.sh
Original file line number Diff line number Diff line change
Expand Up @@ -27,8 +27,12 @@
# applies does the log's last line decide:
# terminal (captain-relevant) or non-terminal (no verb),
# both surfaced at once. A provably-working stale past the
# wedge threshold also surfaces, with an "escalation N"
# count in the reason; at FM_WEDGE_DEMAND_INSPECT_COUNT
# wedge threshold surfaces with an "escalation N" count in
# the reason, EXCEPT a run-step-classified working chain
# (a background no-mistakes validation, not a busy pane),
# which re-verifies the run-step at that same threshold
# and keeps absorbing for as long as it still reports
# working. At FM_WEDGE_DEMAND_INSPECT_COUNT
# consecutive escalations on the SAME pane, the reason
# also carries a "demand-deep-inspection" marker so the
# wake payload itself, not just repetition, forces a
Expand Down Expand Up @@ -538,15 +542,25 @@ clear_write_tracking() { # <window-key>
# absorbed as provably-working - repairs a missing/corrupt timer (self-heals a
# watcher restart between recording the hash and recording the timer), or
# escalates once STALE_ESCALATE_SECS have elapsed. Never re-reads the crew
# state (the costly check already ran once, at classification time). Shared by
# both places a hash can be absorbed this way: the plain non-terminal path,
# and the stale_is_terminal-overridden path (a captain-relevant status-log
# line that an active run/busy pane outranked).
# The worktree write probe runs ONLY here, inside the at-threshold branch that is
# about to escalate: at most one bounded walk per window per STALE_ESCALATE_SECS,
# never per poll.
wedge_timer_check() { # <window> <since-file> <triage-label> <escalation-count-file> <task>
local win=$1 since_file=$2 label=$3 escalation_file=$4 task=$5 since age n reason
# state by default (the costly check already ran once, at classification
# time), EXCEPT when the caller passes reverify_working=1 (the chain was
# opened by a run-step "working" classification rather than a busy pane): a
# no-mistakes background validation can legitimately sit on an idle pane far
# longer than STALE_ESCALATE_SECS, so at each threshold crossing this
# re-confirms the run is still provably working before escalating - the same
# "never escalate past positive evidence" contract crew_worktree_written_since
# already gives a busy foreground pane, applied to a background run's own
# authoritative run-step reading instead. A run-step that has moved to parked
# (awaiting_approval/fix_review) or otherwise stopped reporting working no
# longer re-verifies true, so it still escalates rather than absorbing forever.
# Shared by both places a hash can be absorbed this way: the plain
# non-terminal path, and the stale_is_terminal-overridden path (a
# captain-relevant status-log line that an active run/busy pane outranked).
# The worktree write probe and the reverify check both run ONLY here, inside
# the at-threshold branch that is about to escalate: at most one bounded check
# per window per STALE_ESCALATE_SECS, never per poll.
wedge_timer_check() { # <window> <since-file> <triage-label> <escalation-count-file> <task> [reverify_working]
local win=$1 since_file=$2 label=$3 escalation_file=$4 task=$5 reverify_working=${6:-} since age n reason
since=$(cat "$since_file" 2>/dev/null || true)
case "$since" in
''|*[!0-9]*)
Expand All @@ -559,6 +573,11 @@ wedge_timer_check() { # <window> <since-file> <triage-label> <escalation-count-
*)
age=$(( $(date +%s) - since ))
if [ "$age" -ge "$STALE_ESCALATE_SECS" ]; then
if [ -n "$reverify_working" ] && crew_is_provably_working "$task"; then
date +%s > "$since_file"
triage_log "absorbed $label timer reset (re-verified still working): $win"
return 0
fi
if crew_worktree_written_since "$task" "$STATE" "$since_file"; then
wedge_defer_writing "$win" "$since_file" "$label" "$age"
return 0
Expand Down Expand Up @@ -688,7 +707,7 @@ busy_turn_bound_check() { # <window> <task> <hash> <since-file> <escalation-fil

clear_pause_state() { # <window-key>
local key=$1
rm -f "$STATE/.paused-$key" "$STATE/.paused-rechecked-$key" "$STATE/.paused-resurfaced-$key"
rm -f "$STATE/.paused-$key" "$STATE/.paused-rechecked-$key" "$STATE/.paused-resurfaced-$key" "$STATE/.paused-run-step-$key"
}

clear_pause_tracking() { # <window-key>
Expand All @@ -699,14 +718,24 @@ clear_pause_tracking() { # <window-key>
}

# Reconcile a declared pause or captain-held status with authoritative crew state.
# After fm-crew-state has fallen back to stopped or unknown, paused classification is
# recovered only for a confidently dead ordinary crew, or for a secondmate, whose
# endpoint liveness this function deliberately never reads.
# A run-step-attributed paused verdict (fm-crew-state.sh's ci-monitoring-only
# reconciliation, which has already corroborated the wait against the
# pipeline's own state: no gate, not fixing, no independently-resolved
# outcome) is trusted immediately, live agent or not - a worker's own
# supervising loop staying alive to poll CI in the background is expected,
# not evidence the declared wait is stale. After fm-crew-state has fallen
# back to stopped or unknown, or the pause came only from the status log with
# no run to corroborate it, paused classification instead requires a
# confidently dead ordinary crew's agent (or a secondmate, whose endpoint
# liveness this function deliberately never reads): a status-log pause from a
# still-alive agent may be sitting at an undeclared live interactive gate
# rather than the wait it named, so it is surfaced instead of trusted.
pause_state_class() { # <window> <task>
local win=$1 task=$2 key last recheck_file class agent_alive kind
local win=$1 task=$2 key last recheck_file run_step_file class agent_alive kind
key=$(window_key "$win")
last=$(last_status_line "$STATE/$task.status")
recheck_file="$STATE/.paused-rechecked-$key"
run_step_file="$STATE/.paused-run-step-$key"
if ! status_is_paused_or_captain_held "$last"; then
rm -f "$recheck_file"
crew_absorb_class "$task"
Expand All @@ -717,6 +746,10 @@ pause_state_class() { # <window> <task>
# far more common no-declaration path above still costs none.
kind=$(window_kind "$win")
if [ -e "$STATE/.paused-$key" ] && [ "$(age_of "$recheck_file")" -lt "$STALE_ESCALATE_SECS" ]; then
if [ -e "$run_step_file" ]; then
printf 'paused'
return
fi
if [ "$kind" != secondmate ]; then
agent_alive=$(fm_backend_agent_alive "$(window_backend "$win")" "$win" 2>/dev/null) || agent_alive=unknown
if [ "$agent_alive" != dead ]; then
Expand All @@ -730,10 +763,17 @@ pause_state_class() { # <window> <task>
fi
class=$(crew_absorb_class "$task")
if [ "$class" = working ]; then
rm -f "$recheck_file"
rm -f "$recheck_file" "$run_step_file"
printf 'working'
return
fi
if [ "$class" = paused ] && crew_run_step_paused "$task"; then
date +%s > "$recheck_file"
: > "$run_step_file"
printf 'paused'
return
fi
rm -f "$run_step_file"
if [ "$kind" != secondmate ]; then
agent_alive=$(fm_backend_agent_alive "$(window_backend "$win")" "$win" 2>/dev/null) || agent_alive=unknown
if [ "$agent_alive" != dead ]; then
Expand Down Expand Up @@ -1457,7 +1497,7 @@ EOF
# wedge timer is running for it) - keep treating it that way
# without re-reading the crew state every poll, and without
# letting the still-captain-relevant log line re-surface it.
wedge_timer_check "$w" "$ssf" "stale (overridden terminal status)" "$ewf" "$task"
wedge_timer_check "$w" "$ssf" "stale (overridden terminal status)" "$ewf" "$task" 1
fi
# else: already surfaced as genuinely terminal on a prior poll of
# this same hash - nothing left to do (matches the original,
Expand Down Expand Up @@ -1500,12 +1540,12 @@ EOF
paused) handle_paused_stale "$w" "$task" "$h" ;;
working) clear_pause_state "$key"
printf '%s' "$h" > "$sf"
wedge_timer_check "$w" "$ssf" "non-terminal stale (provably working after a declared pause)" "$ewf" "$task"
wedge_timer_check "$w" "$ssf" "non-terminal stale (provably working after a declared pause)" "$ewf" "$task" 1
triage_log "absorbed non-terminal stale (provably working): $w" ;;
*) handle_paused_stale "$w" "$task" "$h" ;;
esac
else
wedge_timer_check "$w" "$ssf" "non-terminal stale" "$ewf" "$task"
wedge_timer_check "$w" "$ssf" "non-terminal stale" "$ewf" "$task" 1
fi
fi
fi
Expand Down
4 changes: 4 additions & 0 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,8 @@ A pane holding a file newer than the start of its own quiet window, anywhere in
That deferral re-surfaces on the same `FM_PAUSE_RESURFACE_SECS` cadence as a declared wait, with a reason naming the write evidence rather than a wedge, and it is bounded to one pruned, depth-bounded, wall-clock-bounded walk (`FM_WORKTREE_WRITE_PRUNE`, `FM_WORKTREE_WRITE_MAXDEPTH`, `FM_WORKTREE_WRITE_TIMEOUT`) taken only in the branch that was about to escalate, never on every poll.
Every absence of write evidence, including a missing worktree record, a torn-down worktree, a walk that outlives its wall-clock bound on a hung mount, and a failed walk, leaves the existing escalation schedule untouched, so a crew that writes nothing still escalates exactly as before.
A secondmate's recorded worktree is never probed for write activity, because it is a provisioned firstmate home whose own supervision keeps writing inside it whether or not the mate produces anything, so its panes keep escalating on the unchanged schedule.
A provably-working escalation whose working verdict came from a run-step reading, rather than a busy pane, is re-verified against a fresh run-step read at that same threshold before it actually escalates - the same never-escalate-past-positive-evidence contract as the worktree-write probe, applied to a background run's own authoritative reading instead of a file timestamp - so a no-mistakes validation that legitimately sits on an idle pane far longer than `FM_STALE_ESCALATE_SECS` keeps absorbing and resetting its timer for as long as the run-step still reports working; only once it stops does the escalation proceed.
That reverification is scoped to the run-step-classified wedge chain alone and leaves the separate busy-pane/no-completed-turn wedge bound (`FM_BUSY_TURN_MAX_SECS`) unchanged.
A busy pane is otherwise exempt from staleness, but only until its latest `state/<id>.turn-ended` marker reaches `FM_BUSY_TURN_MAX_SECS`, or its `state/<id>.meta` spawn record reaches that age before any turn completes; past that bound it is routed through the same wedge escalation, with the identical reason, escalation count, worktree-write deferral, and `demand-deep-inspection` marker, for inspection only - never an automatic interrupt, signal, or restart.
A crew that declared an external wait (`paused:`) or a verified captain-held transfer is the one exception to that bound: its busy verdict supplies liveness while identifying the long-running foreground call as the declared wait, so it takes the bounded `FM_PAUSE_RESURFACE_SECS` recheck instead of a wedge escalation.
Lifting the declaration restores the unchanged busy-pane wedge path, while a pane that is no longer busy returns to the existing idle declared-wait classification.
Expand All @@ -35,6 +37,8 @@ No-verb wakes, such as `working:` notes and bare turn-ended signals, are benign
A `kind=secondmate` task's status signal is the parent-directed reply stream and is never absorbed as provably working; only its bare turn-ended signal retains the ordinary absorb rule.
A crew that declares `paused:` for a known external wait, or carries a verified `captain-held` transfer, is separately absorbed while idle and re-surfaced only on the longer pause cadence, rather than being treated as a possible wedge.
For an ordinary crew that has stopped, the normal-mode watcher first surfaces one stale wake, then applies that same cadence to an unchanged `paused:` or durable `captain-held` endpoint only when the backend confidently reports its agent dead.
A pause corroborated by an authoritative run-step reading - the run's only remaining step is CI monitoring, no gate, not fixing - skips that dead-agent gate entirely and is trusted immediately, since a worker's own supervising loop staying alive to poll CI is expected, not evidence of staleness; that same run-step reconciliation reports paused instead of the run-step's own working/done reading, so a run parked purely on CI monitoring can no longer silently outrank its own declared wait and wedge-escalate forever.
A pause backed only by the status log, with no run to corroborate it, still needs the dead-agent gate, because a live-looking agent there may be sitting at an undeclared interactive gate rather than the wait it named.
Live or inconclusive liveness remains fail-open at that initial surface, and a secondmate's endpoint liveness is still never read at all; a mate is admitted to that same cadence only to serve a declared wait's bounded re-surface, so a forgotten pause or captain hold on a mate cannot rot invisibly.
Its initial normal-mode status signal still surfaces through the no-verb path, while away mode self-handles that routine signal and owns the later recheck.
Fresh stale panes use the same current-state read before trusting the status log, so an active run or a proven busy worker outranks an old captain-relevant status-log line left behind before validation.
Expand Down
Loading
Loading