Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
29 changes: 29 additions & 0 deletions bin/fm-classify-lib.sh
Original file line number Diff line number Diff line change
Expand Up @@ -1795,6 +1795,35 @@ crew_is_paused() { # <id>
[ "$(crew_absorb_class "$1")" = paused ]
}

# The note bin/fm-crew-state.sh appends to an active run-step's detail when the
# pipeline's own recency verdict says that step is still producing activity. One
# definition, written by fm-crew-state.sh and matched by the predicate below, so
# the emitted line and the classifier reading it cannot drift apart.
FM_CREW_STATE_ACTIVITY_RECENT='run activity recent'

# 0 only on POSITIVE proof that crew <id>'s OWN attributed no-mistakes run is
# still doing work: fm-crew-state.sh reports a working run-step for THIS crew and
# marks its active step's activity recent, which is the pipeline's own recency
# verdict (`axi status` prefixes last_activity with `quiet` once nothing has
# arrived), never a second threshold invented here and never the liveness of the
# shared daemon, which any other crew's run keeps up. That distinction is the
# whole point: a record left at running/fixing after a drive call was killed, or
# after the daemon exited under it, reports a working run-step while nothing
# executes it, and must NOT read as work in progress.
# The busy-pane half of crew_absorb_class's `working` is deliberately excluded: a
# caller that already holds a busy verdict cannot let that pane vouch for itself.
# Not a pure read (see crew_absorb_class), so callers run it at most once per
# STALE_ESCALATE_SECS - never per poll.
crew_nm_run_activity_is_recent() { # <id>
local id=$1 line
[ -n "$id" ] || return 1
line=$("$FM_CREW_STATE_BIN" "$id" 2>/dev/null) || true
case "$line" in
"state: working"*"source: run-step"*"$FM_CREW_STATE_ACTIVITY_RECENT"*) return 0 ;;
esac
return 1
}

# Directories excluded from the worktree write probe below, and the depth it walks.
# The excluded set is everything a supervisor read or a package manager can write
# without the crew doing any work - .git first, so firstmate's own read-only git
Expand Down
14 changes: 14 additions & 0 deletions bin/fm-crew-state.sh
Original file line number Diff line number Diff line change
Expand Up @@ -61,6 +61,12 @@
# FAILED record whose daemon an explicit probe proves down reads unknown,
# never failed: an instrument failure must not read as work failure
# (nm_daemon_probe_down).
# A working run-step also carries a positive `run activity recent` note in
# its detail while the pipeline's own recency verdict says an active step
# is still reporting (nm_run_activity_is_recent, which requires the
# captured `active_steps[]` table and never treats an absent one as
# recency). Supervisors read that note to tell an advancing run from a
# record nothing is executing.
# 3. Reconcile the status log: if its last line says needs-decision/blocked but
# the run-step shows the run moved on, the log is deterministically stale and
# is flagged superseded. A genuinely parked run plus a needs-decision log
Expand Down Expand Up @@ -743,6 +749,14 @@ if [ "$HAVE_RUN" = 1 ]; then
;;
esac

# Positive recency, for supervisors that must tell an advancing run from a
# record nothing is executing: the client's own `quiet` prefix is the verdict
# (nm_run_activity_is_recent), so the note appears only while an active step
# keeps reporting, and never for a coarse row with no steps table to read.
if [ "$RUN_STATE" = working ] && nm_run_activity_is_recent; then
RUN_DETAIL="$RUN_DETAIL${SEP}$FM_CREW_STATE_ACTIVITY_RECENT"
fi

emit "$RUN_STATE" run-step "$RUN_DETAIL"
fi

Expand Down
11 changes: 9 additions & 2 deletions bin/fm-dod-lib.sh
Original file line number Diff line number Diff line change
Expand Up @@ -198,9 +198,13 @@ fm_dod_block() { # <mode> <task-id>
# Definition of done
Delivery contract: mode=direct-PR
This task ships **direct-PR**: you raise the PR yourself, without the no-mistakes pipeline.
Do not run /no-mistakes unless firstmate explicitly instructs you to change this task's delivery path.
The task is complete only when committed on your branch.
When it is implemented and committed, push your branch and open a PR with \`gh-axi\`, then append \`done: PR {url}\` to the status file and stop.
Do NOT run /no-mistakes. The configured merge authority decides whether to merge the PR; firstmate relays the outcome.
When it is implemented and committed, push your branch and open a PR with \`gh-axi\`.
If a push, PR creation, or PR verification fails, diagnose the forge failure first, including the reported authentication, remote, branch, or API error; do not use no-mistakes as a workaround.
Before the final status, verify with \`gh-axi\` that the branch was actually pushed and that the forge reports a full \`https://...\` PR URL for that branch.
Only after those checks append \`done: PR {url}\` to the status file and stop.
The configured merge authority decides whether to merge the PR; firstmate relays the outcome.
EOF
;;
local-only)
Expand Down Expand Up @@ -237,6 +241,9 @@ So background the drive call and poll \`no-mistakes axi status\` from a separate
Where a harness's own command limit is not established, assume it bounds commands and use that same background-and-poll shape.
A killed or timed-out call is never evidence the daemon died: the daemon accepts your response immediately and runs the round in the background, so the call was only ever waiting for a read while the run kept working.
Reattach and keep going rather than reporting the pipeline blocked; rule 7 owns the checks that decide when a pipeline block is real.
After every \`no-mistakes axi respond\`, continue in the same turn with bounded calls to the structured \`no-mistakes axi status\` interface until the attributed run changes step, reaches a terminal outcome, presents a genuine ask-user decision, or rule 7's daemon checks establish a real block.
An accepted response or a status that still reports active work is not a stopping point; the same continuation rule applies after starting or reattaching to a run.
Never end your turn or promise to resume or check later while structured status shows that validation is active, unless the attributed run presents a genuine ask-user decision - escalate it and stop - or rule 7's daemon checks have established a real block.

Two firstmate-specific rules layer on top of that guidance:
- ask-user findings are never yours to answer: escalate to firstmate using rule 6's ask-user format and stop.
Expand Down
39 changes: 30 additions & 9 deletions bin/fm-watch.sh
Original file line number Diff line number Diff line change
Expand Up @@ -214,8 +214,10 @@ STALE_ESCALATE_SECS=${FM_STALE_ESCALATE_SECS:-240} # idle secs before a provabl
# non-busy stale - so it escalates via the existing stale reason, escalation
# counter, and demand-deep-inspection marker for human inspection only, never an
# automatic interrupt, signal, or restart - unless the crew declared the wait
# itself, which takes the long pause cadence instead. A completed turn touches
# turn-ended and resets the age. Set generously above any legitimate interval
# itself, which takes the long pause cadence instead, or the crew's own
# no-mistakes run reports recent activity at the escalation threshold, which
# defers that one escalation and must prove itself again for the next.
# A completed turn touches turn-ended and resets the age. Set generously above any legitimate interval
# between completed turns, including long tool calls, builds, or test runs.
BUSY_TURN_MAX_SECS=${FM_BUSY_TURN_MAX_SECS:-3600}
# A local secondmate's foreign queue is checked on every poll, but only after this
Expand Down Expand Up @@ -853,9 +855,16 @@ clear_write_tracking() { # <window-key>
# line that an active run/busy pane outranked).
# The worktree write probe runs ONLY here, inside the at-threshold branch that is
# about to escalate: at most one bounded walk per window per STALE_ESCALATE_SECS,
# never per poll.
wedge_timer_check() { # <window> <since-file> <triage-label> <escalation-count-file> <task>
local win=$1 since_file=$2 label=$3 escalation_file=$4 task=$5 since age n reason
# never per poll. `defer-live-run` adds the one other evidence a pane can offer
# there, on the same terms and for the same reason: this crew's OWN no-mistakes
# run reporting recent activity defers the escalation once, restarting the idle
# timer so the next window must prove it again. The moment that run stops
# reporting recent activity - a hung step, a record left behind after the daemon
# exited, a run that is no longer this crew's - the threshold falls straight
# through to the unchanged escalation ladder and its demand-deep-inspection
# marker, so no deferral can repeat without fresh proof of work.
wedge_timer_check() { # <window> <since-file> <triage-label> <escalation-count-file> <task> [defer-live-run]
local win=$1 since_file=$2 label=$3 escalation_file=$4 task=$5 live_run=${6-} since age n reason
since=$(cat "$since_file" 2>/dev/null || true)
case "$since" in
''|*[!0-9]*)
Expand All @@ -872,6 +881,11 @@ wedge_timer_check() { # <window> <since-file> <triage-label> <escalation-count-
wedge_defer_writing "$win" "$since_file" "$label" "$age"
return 0
fi
if [ -n "$live_run" ] && crew_nm_run_activity_is_recent "$task"; then
date +%s > "$since_file"
triage_log "absorbed $label (this crew's own no-mistakes run reports recent activity, idle ${age}s): $win"
return 0
fi
n=$(( $(cat "$escalation_file" 2>/dev/null || echo 0) + 1 ))
echo "$n" > "$escalation_file"
reason="stale: $win (idle ${age}s, possible wedge, escalation $n)"
Expand Down Expand Up @@ -947,9 +961,16 @@ handle_paused_stale() { # <window> <task> <hash>
# A busy pane past BUSY_TURN_MAX_SECS is normally a wedge suspect because a hung
# foreground call can hide behind a busy signature. A `paused:` declaration or
# verified captain-held transfer instead identifies that live foreground call as
# the expected external wait. The caller has already confirmed liveness through
# the busy verdict, so this exception does not suppress undeclared wedges or
# alter the separate non-busy classification. handle_paused_stale keeps the
# the expected external wait. An advancing no-mistakes validation is the other
# real wait - a validating worker holds ONE turn open for the whole run by
# contract, so its completed-turn age is expected to cross the bound - but that
# evidence is NOT read here: it is handed to wedge_timer_check as
# `defer-live-run`, which consults it only in the branch that is about to
# escalate. That keeps the crew-state read to at most one per window per
# STALE_ESCALATE_SECS instead of one per poll, and keeps the outcome a single
# deferral that must be re-earned rather than a standing exemption. The caller
# has already confirmed liveness through the busy verdict, so this exception does
# not suppress undeclared wedges or alter the separate non-busy classification. handle_paused_stale keeps the
# exception bounded by re-surfacing it once per PAUSE_RESURFACE_SECS. Away mode
# remains daemon-owned and receives the undecorated wake identity for its own
# classification, which is why the declaration is read before the afk branch
Expand Down Expand Up @@ -992,7 +1013,7 @@ busy_turn_bound_check() { # <window> <task> <hash> <since-file> <escalation-fil
handle_paused_stale "$win" "$task" "$h"
return 0
fi
wedge_timer_check "$win" "$since_file" "busy (no completed turn)" "$escalation_file" "$task"
wedge_timer_check "$win" "$since_file" "busy (no completed turn)" "$escalation_file" "$task" defer-live-run
return 1
}

Expand Down
7 changes: 5 additions & 2 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,9 +22,12 @@ That deferral re-surfaces on the same `FM_PAUSE_RESURFACE_SECS` cadence as a dec
Every absence of write evidence, including a missing worktree record, a torn-down worktree, a walk that outlives its wall-clock bound on a hung mount, and a failed walk, leaves the existing escalation schedule untouched, so a crew that writes nothing still escalates exactly as before.
A secondmate's recorded worktree is never probed for write activity, because it is a provisioned firstmate home whose own supervision keeps writing inside it whether or not the mate produces anything, so its panes keep escalating on the unchanged schedule.
A busy pane is otherwise exempt from staleness, but only until its latest `state/<id>.turn-ended` marker reaches `FM_BUSY_TURN_MAX_SECS`, or its `state/<id>.meta` spawn record reaches that age before any turn completes; past that bound it is routed through the same wedge escalation, with the identical reason, escalation count, worktree-write deferral, and `demand-deep-inspection` marker, for inspection only - never an automatic interrupt, signal, or restart.
A crew that declared an external wait (`paused:`) or a verified captain-held transfer is the one exception to that bound: its busy verdict supplies liveness while identifying the long-running foreground call as the declared wait, so it takes the bounded `FM_PAUSE_RESURFACE_SECS` recheck instead of a wedge escalation.
A crew that declared an external wait (`paused:`) or a verified captain-held transfer is one exception to that bound: its busy verdict supplies liveness while identifying the long-running foreground call as the declared wait, so it takes the bounded `FM_PAUSE_RESURFACE_SECS` recheck instead of a wedge escalation.
Lifting the declaration restores the unchanged busy-pane wedge path, while a pane that is no longer busy returns to the existing idle declared-wait classification.
While away mode is active, a busy pane that crosses the bound under a declared wait is handed to the daemon as the plain wake identity instead of taking that recheck in the watcher, because the daemon owns triage there and a wake already decorated as a possible wedge would override the daemon's own declared-wait verdict; an undeclared busy pane past the bound still takes the wedge escalation in away mode.
The other exception is a worker inside its own no-mistakes validation, whose contract holds one turn open for the whole run so its completed-turn age is expected to cross the bound: at each `FM_STALE_ESCALATE_SECS` threshold, a busy pane whose own attributed run reports recent activity defers that one escalation and restarts the idle timer, so the next threshold has to prove the activity again.
The proof is the pipeline's own recency verdict on the step attributed to this crew, read through `bin/fm-crew-state.sh` only in the branch that was about to escalate rather than on every poll, and never the liveness of the shared no-mistakes daemon, which any other crew's run keeps up.
A hung step, a record left behind after the daemon exited under it, and a run that is no longer this crew's all stop reporting that activity, so the threshold falls straight through to the unchanged escalation ladder and its `demand-deep-inspection` marker; this deferral exists only on the busy-pane path, and the non-busy stale classification is unchanged.
While away mode is active, a busy pane that crosses the bound under a declared wait is handed to the daemon as the plain wake identity instead of taking that recheck in the watcher, because the daemon owns triage there and a wake already decorated as a possible wedge would override the daemon's own declared-wait verdict; an undeclared busy pane past the bound still takes the wedge path in away mode, including the evidence-based deferrals above.
That handoff is keyed on the declaration itself (the status log's signature) rather than on the pane capture, so a harness footer that ticks on every poll wakes the daemon once per declaration instead of once per poll, and it clears the wedge timer, escalation count, and worktree-write deferral exactly as the normal-mode absorber does, so an undeclared busy phase's timer does not resume when the declaration lifts.
Those actionable wakes are written to a durable local queue (`state/.wake-queue`) only after generation-bound recovery evidence is published, so an interrupted watcher or handling turn can be recovered without losing the queue record.
Agent endpoint liveness and queue-consumption liveness are separate: on each poll, the primary watcher reads the oldest valid actionable row from every endpoint-recorded local secondmate home's durable wake queue without locking, consuming, or rewriting that foreign queue.
Expand Down
2 changes: 1 addition & 1 deletion docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -931,7 +931,7 @@ FM_TURNEND_CHURN_ABSORB_SECS=900 # longest one endpoint's bare turn-ends may b
FM_CAPTAIN_RE='done:|needs-decision:|blocked:|failed:|PR ready|checks green|ready in branch|merged' # captain-relevant status regex; nonterminal progress verbs remain excluded even when their prose matches
FM_CLASSIFY_PAUSED_VERB=paused # leading status verb for a declared external wait; excluded from FM_CAPTAIN_RE and distinct from blocked
FM_STALE_ESCALATE_SECS=240 # idle seconds before a provably-working stale pane escalates; stale panes whose crew is not provably working surface immediately unless admitted directly to the declared-wait cadence, while a live idle declared wait still surfaces once before that cadence bounds repeats
FM_BUSY_TURN_MAX_SECS=3600 # maximum age of a busy pane's latest state/<id>.turn-ended marker, or its state/<id>.meta spawn record before any turn completes, before the same wedge escalation used for a provably-working non-busy stale takes over; inspection-only, never an automatic interrupt or restart; a declared external wait or verified captain-held transfer takes the FM_PAUSE_RESURFACE_SECS recheck below instead
FM_BUSY_TURN_MAX_SECS=3600 # busy-turn age bound in seconds; docs/architecture.md "Event-driven supervision" owns the age sources, inspection-only escalation, and declared-wait and recent-run-activity deferrals
FM_PAUSE_RESURFACE_SECS=3600 # seconds between bounded rechecks of a declared external wait or verified captain-held transfer, and between repeated new-hash stale alarms for an ordinary crew task with an open backlog captain call; this includes a live idle pane after its first inconclusive stale wake and a live busy pane past FM_BUSY_TURN_MAX_SECS, while the away-mode daemon uses the same setting and ages its window against the crew's own latest status line rather than pane busy state
FM_SECONDMATE_WAKE_STALL_SECS=180 # minimum interval with no change of the oldest actionable foreign wake-queue row (it advances as the mate drains, and a queue reprovisioned under the same task id starts a fresh interval at whatever sequence it restarts) before an endpoint-recorded local secondmate produces one durable parent wake-loop-stall notification for that no-progress episode; a mate that is provably inside an active turn (an exact busy verdict, bounded by the same FM_BUSY_TURN_MAX_SECS above) never escalates whatever this interval says, declared external-wait pause rows are excluded, and zero or invalid values use 180
FM_WEDGE_DEMAND_INSPECT_COUNT=3 # consecutive provably-working stale escalations on the same unchanged pane before demand-deep-inspection is added
Expand Down
5 changes: 4 additions & 1 deletion docs/verification/runtime-backends.md
Original file line number Diff line number Diff line change
Expand Up @@ -818,7 +818,10 @@ The same guarded named-lab command passed on 2026-09-03 against Herdr 0.8.2 afte
It reported `steal_live=0 floor_verdict=0 default-session-tripwire=armed`, with the fleet's default session unchanged before and after.

Part C is the case the suite could not reach before: a doomed pane whose shell holds a persistent background child fails the lone-idle-shell proof on every sample, so the plan takes the plain explicit close, in the geometry where the closing workspace's right neighbour is a spacer rather than the focused anchor.
On 0.7.5 that fallback exposed a bounded four-sample wrong-focus window and restored the anchor exactly; on 0.8.0 the same fallback exposed none, which is why default-on projection is floored at 0.8.0 rather than mitigated further below it.
In the recorded runs above, the sampler observed a bounded four-sample wrong-focus window on 0.7.5 with exact anchor restoration and none on 0.8.0, supporting the 0.8.0 default-on floor.
The current Part C regression instead checks the adapter's call log for corrective `tab focus` and verifies the final exact anchor: correction must occur on a release Part A proves defective and must be absent on a focus-preserving release.
It no longer samples focus concurrently during the fallback close; the recorded sample counts are prior evidence, not output of the current guard.
Part B retains concurrent sampling of the mitigated removal path.
The suite also cross-checks its own Part A measurement against the floor classifier on whatever release it runs, so a drifted protocol-to-release mapping fails there rather than silently gating on the wrong thing.

### Presentation version floor
Expand Down
Loading
Loading